Image Composition

JP2025512670A5Pending Publication Date: 2026-03-19KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing methods for view synthesis in immersive video face challenges such as limited observation space, degradation of image quality, and errors caused by insufficient 3D video data, particularly when viewers move away from the nominal position.

Method used

The proposed solution involves a system that includes a receiver for multiple images and 3D spatial data, a view synthesis neural network, and a neural network trainer. This system generates view shift images for different view poses using the neural network, trained with images from various view poses, and produces an audio-visual data stream containing image data, scene data, and coefficient data for the neural network.

Benefits of technology

This approach improves the quality of view synthesis, enhances user experience in XR/AR/VR/MR applications, reduces data requirements, and facilitates efficient and high-quality view shifting and synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The first device comprises a first receiver 301 for receiving images of a captured scene and a second receiver 303 for receiving 3D spatial data of the scene. A view synthesis neural network 307 generates view-shifted images of the scene for different view poses from the images and spatial data. A neural network trainer 309 trains the view synthesis neural network 307 based on the images of the scene for different view poses. A generator 305 generates an audiovisual data stream including image data of the images, scene data representing the three-dimensional spatial data, and coefficient data describing the coefficients of the view synthesis neural network 307 after training. The second device receives the audiovisual data stream and configures a local neural network 403 based on the coefficient data. The local neural network 403 is then used to generate images of the scene for different view poses.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to neural network based image synthesis, and in particular, but not exclusively, to neural network based video frame synthesis. [Background technology]

[0002] 2. Description of the Related Art The variety and scope of image and video applications has increased significantly in recent years, and new services and methods for using and consuming video are continually developed and introduced.

[0003] For example, one service that is gaining popularity is the presentation of image sequences in such a way that the observer can actively interact with the system to change the parameters of the rendering.A very attractive feature in many applications is the ability to change the observer's effective viewing position and direction, e.g., to allow the observer to move and look around the displayed scene.

[0004] Such features may in particular allow a virtual reality experience to be provided to the user, whereby the user can, for example, move around in the virtual environment with (relative) freedom and dynamically change his position and where he is looking. Typically, such extended reality (XR) applications are based on a three-dimensional model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well known from gaming applications for computers and consoles, for example in the first-person shooter category. Extended reality (XR) applications include Virtual Reality (VR), Augmented Reality (AR) and Mixed Reality (MR) applications.

[0005] An example of a proposed video service or application is immersive video, where the video is played, for example, on a VR headset, to provide a three-dimensional experience. In the case of immersive video, the observer has the freedom to move around while watching the displayed scene, which can be perceived as being seen from different viewpoints. However, in many typical approaches, the amount of movement is restricted to a relatively small area around a nominal viewpoint, which typically corresponds, for example, to the viewpoint from which the video capture of the scene was performed. In such applications, three-dimensional scene information is often provided that allows high-quality viewpoint image synthesis for viewpoints relatively close to the reference viewpoint, but degrades when the viewpoint deviates too much from the reference viewpoint.

[0006] Immersive video is often referred to as six degrees of freedom (6DoF) or 3DoF+ video. MPEG Immersive Video (MIV) is a new standard where metadata to enable and standardize immersive video is used on top of existing video codecs.

[0007] A problem with immersive video is that the viewing space, the 3D space in which an observer has a sufficient quality 6DoF experience, is limited. As the observer moves outside the viewing space, degradation and errors due to the synthesis of view images become more and more noticeable, which can result in an unacceptable user experience. Errors, artifacts and inaccuracies in the generated view images can occur especially because the provided 3D video data does not provide enough information for view synthesis (e.g., deocclusion data).

[0008] For example, immersive video data may be provided in the form of multiple views, possibly accompanied by a depth data (MVD) representation of the scene. The scene may be captured by several spatially distinct cameras, and the captured images may be provided together with a depth map. However, the likelihood that such a representation does not contain sufficient image data for the de-occluded regions increases substantially as the viewpoint becomes more and more different from the reference viewpoint at which the MVD data was captured. Thus, as the observer moves away from the nominal position, image portions that should be de-occluded for the new viewpoint but are missing from the source view cannot be directly synthesized from image data describing such image portions. Also, incomplete depth maps may introduce distortions when performing view synthesis, especially as part of view warping, which is an integral part of the synthesis operation. The further the synthesized viewpoint is from the original camera viewpoint, the more severe the distortions in the synthesized view.

[0009] Most existing methods for view synthesis from multi-view images require the availability of a depth map for each source view, and synthesize a new view by combining predictions from multiple reference views, which are then performed using depth-steered rendering.

[0010] A recent alternative approach is to convert the multi-view images into a layered representation that includes transparency (e.g., Multi-Plane Image (MPI), Multi-Sphere Image (MSI), Multi-Object Surface Image (MOSI) formats, etc.) and encode / compress this layered representation. After decoding, new views can be synthesized via back-to-front layer synthesis.

[0011] Regardless of the particular 3D image format used, the process of generating view images for different view poses is a challenging process that tends to result in imperfect images. Most image synthesis algorithms tend to introduce some artifacts or errors, for example due to incomplete de-occlusion. Furthermore, most approaches tend to require complex and computationally intensive processing to generate images for different view poses with the desired quality. Also, the requirements for 3D image data that need to be generated to enable view synthesis algorithms to generate reliable, high-quality images tend to be difficult to meet, increasing data rates and making the distribution and communication of such data difficult and resource-intensive. Summary of the Invention [Problem to be solved by the invention]

[0012] Therefore, improved approaches would be advantageous, particularly approaches that allow for improved operation, greater flexibility, improved immersive user experience, reduced complexity, easier implementation, better synthetic image quality, improved rendering, greater freedom of (possibly virtual) movement for the user, improved and / or easier view synthesis for different view poses, reduced data requirements, and / or improved performance and / or operation. [Means for solving the problem]

[0013] Accordingly, the Invention seeks to preferably mitigate, reduce or eliminate one or more of the above mentioned disadvantages singly or in any combination.

[0014] According to one aspect of the present invention, there is provided an apparatus comprising: a first receiver configured to receive a plurality of images of a three-dimensional scene captured from different view poses and a second receiver configured to receive three-dimensional spatial data of the scene; a view synthesis neural network configured to generate view-shifted images of the scene for different view poses from the plurality of images and the three-dimensional spatial data; a neural network trainer configured to train the view synthesis neural network based on images of the scene for the different view poses; and a generator configured to generate an audiovisual data stream including image data for at least some of the plurality of images, scene data representative of the three-dimensional spatial data, and coefficient data describing coefficients of the view synthesis neural network after training.

[0015] The present invention allows for improved data characterizing the scene being generated, allowing for improved image synthesis. This approach allows for improved image synthesis for different view poses. In many embodiments, data can be generated that allows for efficient and high quality view shifting and synthesis based on neural networks.

[0016] The present invention can provide an improved user experience in many embodiments and scenarios: the approach enables, for example, improved XR / AR / VR / MR applications based on limited capture of a scene.

[0017] The use of neural networks for view synthesis can, in many embodiments, provide improved view synthesis and can, for example, allow for greater flexibility and reduced requirements for data representing a scene.

[0018] The multiple images may include a set of multi-view images. The images of the multiple images may be two-dimensional images. The scene may be a real-world scene and the multiple images may be captured images of the real-world scene. The images may be images captured by an image camera or a video camera. The three-dimensional spatial data may represent spatial properties of the scene and may include or consist of, for example, a 3D point cloud, a mesh of (static) background or (dynamic) foreground objects, or a depth map. The three-dimensional spatial data may be independent / unrelated to the multiple images or any view pose for the multiple images. The three-dimensional spatial data may be view pose independent. The three-dimensional spatial data may be view pose independent data that describes geometric / spatial properties of the scene. The three-dimensional spatial data may not be tied to or represented or associated with a view in the multiple images and / or view poses of the multiple images. The three-dimensional spatial data may include or consist of geometric primitives such as 3D points, planes, surfaces, lines, meshes, etc.

[0019] The server neural network may specifically be a convolutional network. The server neural network may include one or more convolutional layers. The coefficients may include coefficients of filters / kernels for one or more convolutional layers. One, more, or all of the convolutional layers of the server neural network may be hidden layers or processing layers.

[0020] In some embodiments, the server neural network may comprise one or more fully connected layers, and the coefficients may include weights for the connections of the one or more fully connected layers.

[0021] A pose can be a position, an orientation, or a position and an orientation, and can represent / indicate them.

[0022] The images of the scene at different view poses used to train the view synthesis neural network may include one or more of the multiple images received by the first receiver.

[0023] A view-shifted image may be an image that has been synthesized for a view pose that is different from the view poses of the multiple images.

[0024] According to an optional feature of the invention, a first receiver is configured to receive a set of input video sequences representing views of a three-dimensional scene from different view poses, the plurality of images being frames of the set of input video sequences, the neural network trainer is configured to dynamically train the view synthesis neural network to vary coefficients of the view synthesis neural network over time, and the generator is configured to generate an audiovisual data stream comprising frames from the set of input video sequences and time-varying coefficient data describing the time-varying coefficients of the view synthesis neural network after training.

[0025] The present invention can allow for improved data that dynamically characterizes the scene being generated, which can allow for improved video composition.

[0026] According to an optional feature of the invention, the neural network trainer is configured to train the view synthesis neural network using a training set of images including a first set of reference images for view poses not represented by at least some images of the plurality of images.

[0027] This can provide improved training and image synthesis in many scenarios, particularly as training can be enhanced to reflect characteristics over a wider area, such as across the entire desired field of view.

[0028] According to an optional feature of the invention, the first set of reference images includes at least one image selected from the group consisting of reference images generated by non-neural network view shift using a plurality of images, and reference images generated from a visual scene model of the scene.

[0029] This can provide improved training and image synthesis over larger areas, compensating / mitigating the limited capture data available. This approach can reduce capture requirements.

[0030] According to an optional feature of the invention, the training set of images includes a second set of reference images including images of the plurality of images, and the neural network trainer is configured to apply different weights to the first set of reference images and the second set of reference images.

[0031] This provides improved operation and training in many scenarios. Different weights can be implemented, for example, by applying different weights to the contributions to the cost function used to train the server neural network.

[0032] In accordance with an optional feature of the invention, the neural network trainer is configured to encode and decode at least some of the plurality of images before they are provided to the view synthesis neural network.

[0033] This can improve performance in many embodiments, for example allowing neural network view synthesis to compensate / mitigate effects and degradations resulting from the communication / distribution of audiovisual data streams.

[0034] In accordance with an optional feature of the invention, the neural network trainer is configured to encode and decode the three-dimensional spatial data before it is provided to the view synthesis neural network.

[0035] This can improve performance in many embodiments, for example allowing neural network view synthesis to compensate / mitigate artifacts and degradations resulting from the communication / distribution of audiovisual data streams.

[0036] According to an optional feature of the invention, the neural network trainer is configured to initialize the view synthesis neural network with a set of default coefficients and train the view synthesis neural network to determine modified coefficients for the view synthesis neural network, and the generator is configured to include at least some of the modified coefficients in the audiovisual data stream.

[0037] This allows for improved operation in many embodiments and scenarios, and in particular allows for reduced data rates and / or faster initialization in many embodiments.

[0038] In some embodiments, the neural network trainer can be configured to select a set of default coefficients from a plurality of sets of default coefficients in response to a plurality of images.

[0039] According to an optional feature of the invention, the generator is configured to select a subset of coefficients for transmission in dependence on a difference between the modified coefficients and the default coefficients.

[0040] This allows for improved operation in many embodiments and scenarios.

[0041] In some embodiments, the generator is configured to include image data for only a subset of the plurality of images, and to further include at least one feature map of a view synthesis neural network for images for a view pose that is different from a view pose of the subset of the plurality of images.

[0042] In some embodiments, the generator is configured to include at least one feature map of the view synthesis neural network for an image with a view pose that is different from the view poses of the plurality of images.

[0043] In some embodiments, the generator is configured to generate the audiovisual data stream such that at least two of the image data, the scene data and the coefficient data share a group of pictures (GOP) structure.

[0044] In some embodiments, the first receiver is configured to receive an input video sequence including a plurality of frames representing views of the three-dimensional scene from different view poses for each of a plurality of time points, the plurality of images including a set of frames for a time point, the neural network trainer is configured to dynamically train the view synthesis neural network to vary coefficients of the view synthesis neural network over time, and the generator is configured to generate an output video sequence including image data for at least some frames of the input video sequence and coefficient data for the different time points of the output video sequence.

[0045] According to one aspect of the present invention, there is provided an apparatus comprising: a receiver configured to receive an audiovisual data stream including image data for a plurality of images representing a three-dimensional scene captured from different view poses, three-dimensional spatial data for the scene, and coefficient data describing coefficients for a view synthesis neural network; a view synthesis neural network configured to generate view-shifted images for the scene for the different view poses from the plurality of images and the three-dimensional spatial data; and a neural network controller for setting coefficients of the neural network in response to the coefficient data.

[0046] The present invention can enable improved image synthesis. This approach allows for improved image synthesis for different view poses.

[0047] The present invention can provide an improved user experience in many embodiments and scenarios: the approach enables, for example, improved XR / AR / VR / MR applications based on limited capture of a scene.

[0048] The multiple images may include a set of multi-view images. The images of the multiple images may be two-dimensional images. The scene may be a real-world scene, and the multiple images may be captured images of the real-world scene. The images may be images captured by an image camera or a video camera. The three-dimensional spatial data may represent spatial properties of the scene, and may include or consist of, for example, a 3D point cloud, a mesh of (static) background or (dynamic) foreground objects, or a depth map.

[0049] The client neural network may specifically be a convolutional network. The client neural network may include one or more convolutional layers. The coefficients may include coefficients of filters / kernels for one or more convolutional layers. One, more, or all of the convolutional layers of the client neural network may be hidden layers or processing layers.

[0050] In some embodiments, the client neural network may comprise one or more fully connected layers, and the coefficients may include weights for the connections of the one or more fully connected layers.

[0051] A view-shifted image can be a composite image for a view pose that is different from the view poses of the multiple images.

[0052] According to an optional feature of the invention, the audiovisual data stream includes a set of video sequences including a plurality of frames representing views of a three-dimensional scene from different viewing poses, the plurality of images being frames of the plurality of frames, the coefficient data including time-varying coefficient data describing time-varying coefficients of a view synthesis neural network, and the neural network controller is configured to modify the coefficients of the view synthesis neural network in response to the time-varying coefficient data.

[0053] This allows for improved image composition in many embodiments.

[0054] According to an optional feature of the invention, the neural network controller is configured to determine interpolated coefficient values ​​for at least one time point for which no coefficient data is included in the audiovisual data stream, the interpolated coefficient values ​​being determined from coefficient values ​​of the time-varying coefficient data, and configured to set coefficients of the view synthesis neural network to the interpolated coefficient values ​​for the at least one time point.

[0055] This allows for improved image composition in many embodiments.

[0056] According to an optional feature of the invention, the audiovisual data stream includes at least one neural network feature map for an image with a view pose that is different from the view poses of the plurality of images, and the neural network controller is configured to configure the view synthesis neural network with the at least one feature map.

[0057] This allows for improved image synthesis in many embodiments, and in many scenarios it allows for reduced data rates of the audiovisual data stream.

[0058] This allows for improved image composition in many embodiments.

[0059] According to an optional feature of the invention, the neural network controller is configured to initialize the view synthesis neural network with a set of default coefficients and to overwrite the default coefficients with coefficients from the audiovisual data stream.

[0060] This allows for improved image composition in many embodiments.

[0061] In accordance with an optional feature of the invention, the neural network controller is configured to select a set of default coefficients from a plurality of sets of default coefficients in response to a plurality of images.

[0062] This allows for improved image composition in many embodiments.

[0063] In some embodiments, the neural network controller is configured to initialize the neural network with default coefficients and to override the default coefficients with coefficients determined from the coefficient data.

[0064] According to an optional feature of the invention, at least two of the image data, the scene data and the coefficient data share a Group Of Pictures (GOP) structure in the audiovisual data stream.

[0065] This allows for improved operation in many scenarios: it can facilitate the encoding and / or decoding of audiovisual data streams.

[0066] In accordance with an optional feature of the invention, the view synthesis neural network includes some tunable layers and some non-tunable layers, and the audiovisual data stream includes coefficient data only for the tunable layers.

[0067] This allows for improved image synthesis in many embodiments, and in many scenarios it allows for reduced data rates of the audiovisual data stream.

[0068] In some embodiments, the view synthesis neural network comprises a number of fixed layers with fixed coefficients and a number of variable layers with variable coefficients, the neural network trainer is configured to train only the number of adjustable layers, and the generator is configured to generate the audiovisual data stream to include coefficient data only for the variable coefficients.

[0069] According to an optional feature of the invention, the three-dimensional spatial data includes video point cloud data of the scene.

[0070] This allows for improved image composition in many embodiments.

[0071] According to one aspect of the invention, an audiovisual data stream is provided that includes image data for a plurality of images representing a three-dimensional scene captured from different view poses, three-dimensional spatial data for the scene, and coefficient data describing coefficients for a view synthesis neural network for generating view-shifted images of the scene for the different view poses from the plurality of images and the three-dimensional spatial data.

[0072] According to one aspect of the present invention, there is provided a method for generating an audiovisual data stream, the method including receiving a plurality of images of a three-dimensional scene captured from different view poses, receiving three-dimensional spatial data of the scene, a view synthesis neural network generating view-shifted images for the scene at the different view poses from the plurality of images and the three-dimensional spatial data, training the view synthesis neural network based on the images of the scene at the different view poses, and generating an audiovisual data stream including image data for at least some of the plurality of images, scene data representing the three-dimensional spatial data, and coefficient data describing coefficients of the trained view synthesis neural network.

[0073] According to one aspect of the invention, there is provided a method comprising receiving an audiovisual data stream including image data for a plurality of images representing a three-dimensional scene captured from different view poses, three-dimensional spatial data for the scene, and coefficient data describing coefficients for a view synthesis neural network, the view synthesis neural network generating view-shifted images for the scene at the different view poses from the plurality of images and the three-dimensional spatial data, and setting coefficients of the view synthesis neural network in response to the coefficient data.

[0074] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief description of the drawings]

[0075] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings in which: [Figure 1] FIG. 1 illustrates an example of elements of an image delivery system. [Diagram 2] FIG. 1 illustrates an example image capture scenario. [Diagram 3] 1 illustrates example elements of an apparatus according to some embodiments of the present invention. [Figure 4] 1 illustrates example elements of an image synthesis device according to some embodiments of the present invention. [Diagram 5] FIG. 1 illustrates an example of a view synthesis neural network. [Figure 6] FIG. 13 shows example poses for training images for training a view synthesis neural network. [Figure 7] FIG. 1 shows an example of the structure of an artificial neural network. [Figure 8] FIG. 1 shows an example of a node in an artificial neural network. [Figure 9] FIG. 2 illustrates some elements of a possible configuration of a processor for implementing elements of an apparatus according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0076] The capture, delivery and display of three-dimensional video has become increasingly popular and desirable in several applications and services. A particular approach is known as immersive video, which typically involves the provision of views of real-world scenes, and often real-time events, and allows for small observer movements, such as relatively small head movements and rotations. For example, a real-time video broadcast of a sporting event can provide a user with the impression of sitting in the stands watching the sporting event. The user can, for example, look around and have a natural experience similar to that of a spectator present at that position in the stands. In recent years, there has been an increasing popularity of display devices with position tracking and 3D interaction that support applications based on 3D capture of real-world scenes. Such display devices are well suited for immersive video applications that provide an enhanced three-dimensional user experience.

[0077] The following description focuses on immersive video applications, although it will be understood that the principles and concepts described may be used in many other applications and embodiments.

[0078] In many approaches, the immersive video can be provided locally to the viewer, for example by a stand-alone device that does not use or even have access to a remote video server. In other applications, however, the immersive application can be based on data received from a remote or central server. For example, video data can be provided to a video rendering device from a remote central server and processed locally to generate the desired immersive video experience.

[0079] 1 shows an example of an immersive video system in which a video rendering client 101 cooperates with a remote immersive video server 103 over a network 105, such as the Internet. The server 103 may be configured to support a potentially large number of client devices 101 simultaneously.

[0080] The server 103 can support immersive video experiences, for example, by transmitting three-dimensional video data describing a real-world scene. The data can specifically describe the visual features and geometric properties of the scene generated from real-time capture of the real world by a set of (possibly 3D) cameras.

[0081] To provide such services for real-world scenes, the scene is typically captured from different positions and different camera capture poses are used. As a result, the relevance and importance of multi-camera capture and, for example, 6DoF (six degrees of freedom) processing is rapidly increasing. Applications include live concerts, live sports, and telepresence. The freedom to choose one's viewpoint enriches these applications by providing a greater sense of presence than in regular video. Furthermore, one can consider immersive scenarios where observers can move through and interact with the live-captured scene. For broadcast applications, this would require real-time view synthesis in the client device. View synthesis introduces errors, and these errors vary depending on the implementation details of the algorithm.

[0082] The server 103 in Figure 1 is configured to generate an audiovisual data stream including data describing the scene, and the client 101 is configured to receive and process this audiovisual data stream to generate an output video stream that dynamically reflects changes in user pose, thereby providing an immersive video experience in which the displayed view adapts to changes in viewpoint / user pose / positioning.

[0083] In the field, the terms configuration and pose are used as general terms for position and / or orientation. For example, a combination of position and orientation of an object, a camera, a head or a view may be referred to as a pose or configuration. A configuration or pose index may thus include six values / components / degrees of freedom, each value / component typically describing an individual characteristic of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a configuration or pose may be considered or represented with fewer components, for example when one or more components are considered fixed or irrelevant (e.g., four components may provide a complete representation of the pose of an object if all objects are considered to be at the same height and have a horizontal orientation). In the following, the term pose is used to refer to a position and / or orientation that can be represented by one to six values ​​(corresponding to the maximum possible degrees of freedom). The term pose may be replaced by the term configuration. The term pose may be replaced by the term position and / or orientation. The term pose can be replaced by the terms position and orientation (if pose provides both position and orientation information), by the term position (if pose (possibly) provides position information), or by orientation (if pose (possibly) provides orientation information).

[0084] A commonly used approach to represent a scene is known as multi-view with depth (MVD) representation and capture. In such an approach, the scene is represented by multiple images with associated depth data, where the images typically represent different view poses from a limited capture area. The images can in practice be captured by using a camera rig with multiple cameras and depth sensors.

[0085] An example of such a capture system is shown in Fig. 2. The figure shows a scene to be captured, including a scene object 201 in front of a background 203. Multiple capture cameras 205 are arranged in a capture area 207. The result of the capture can be a representation of the 3D scene by multi-view images and depth representations, i.e. by images and depths provided for multiple capture poses. The multi-view images and depth representations can thus provide a representation of the 3D scene from a capture zone, where the data representing the 3D scene can provide a representation of the 3D scene from a capture zone, where the visual data provide a description of the 3D scene.

[0086] The MVD representation can be used to perform view synthesis, whereby a view image of a scene from a given view pose can be generated. The view pose may require a view shift of an image of the MVD representation to a view pose so that an image of the view of the scene from the view pose can be generated and displayed to a user. The view shift and synthesis is based on depth data, e.g., with a disparity shift between positions in the MVD image, and the view pose image depends on the depth of the corresponding object in the scene.

[0087] The quality of the generated view images depends on the image and depth information available to the view synthesis operation, and further on the amount of view shifting required.

[0088] For example, view shifting typically results in the disocclusion of parts of an image that may not be visible in the main image used for the view shift. Such holes can be filled by data from other images (if they captured the disoccluded elements), but it is also typically the case that image parts that are disoccluded for the new viewpoint are also missing from other source views. In that case, view synthesis needs to estimate the data, for example based on surrounding data. The disocclusion process is inherently prone to being a process that introduces inaccuracies, artifacts and errors. Moreover, this tends to increase with the amount of view shift, specifically, the likelihood of missing data (holes) during view synthesis increases as the image's distance from the capture pose increases.

[0089] Another possible source of distortion is imperfect depth information. Often, depth information is provided by a depth map where the depth values ​​are generated by depth estimation (e.g., disparity estimation between source images) or measurement (e.g., ranging) that is not perfect, and therefore the depth values ​​may contain errors and inaccuracies. View shift is based on depth information, and imperfect depth information introduces errors or inaccuracies into the synthesized image. The further the synthesis viewpoint is from the original camera viewpoint, the more severe the distortion of the synthetic target view image becomes.

[0090] Thus, as the view pose moves farther from the capture pose, the quality of the composite image tends to degrade: if the view pose is far enough away from the capture pose, the image quality may degrade unacceptably and a poor user experience may be experienced.

[0091] Recently, it has been proposed to perform view shifting and synthesis based on neural networks. In such applications, a neural network can be trained to synthesize view images for different view poses based on input images from a set of capture poses. The training can be based on a set of reference images provided for various view poses, and the neural network is configured to generate corresponding images for the same view poses. The difference between the reference images and the generated corresponding images is used to adapt and train the neural network. The trained neural network can then be used to dynamically generate view images for different view poses, for example, to support a user virtually moving around in a scene. In the following, an approach for representing, processing, and rendering views of a scene based on the use of neural networks is described.

[0092] 3 illustrates elements of an exemplary apparatus for generating an audiovisual data stream representing a scene. The audiovisual data stream includes data representing visual characteristics of the scene. The apparatus may specifically be a source or server that provides the audiovisual data stream to a remote device. Specifically, the apparatus may be an example of an element of server 103 of FIG. 1, and the following description focuses on describing the apparatus with reference to server 103.

[0093] Figure 4 shows elements of an exemplary apparatus for generating an image for a scene from an audiovisual data stream representing the scene, specifically from an audiovisual data stream generated by the apparatus of Figure 3. This apparatus may specifically be an end user device or a client that receives the audiovisual data stream from a remote device, such as the apparatus of Figure 3. Specifically, this apparatus may be an example of an element of client 101 of Figure 1, and the following description will focus on describing the apparatus with reference to server 103.

[0094] The server 103 has a first receiver 301 that receives multiple images of a three-dimensional scene, these images being view images of different view poses in the scene. The multiple images can be images of the scene captured from different view poses. The images can in some embodiments be images of a real scene captured by a camera at different positions and / or different orientations in the scene. In some embodiments, the images can alternatively or additionally be images of a virtual or partially virtual scene. In such cases, the images can be generated / captured, for example, by evaluating a virtual model of the scene. In the following, we focus on captured images, which are images of a real scene captured by a camera with different poses in the scene. The multiple images are called captured images or input images.

[0095] The set of input images may be provided in any suitable format, in some embodiments the images may further include depth, in particular a set of MVD images may be received.

[0096] In some embodiments, the input image may be a still image of a static scene. However, in many embodiments, the images may be frames of one or more video sequences. For example, an input video sequence for each capture position may be received. For example, a dynamic scene such as a sporting event may be captured from a set of video cameras located at a given position. Each of the cameras may provide a video sequence including frames. Thus, the first receiver 301 may receive multiple video sequences including time-sequential frames, with each video sequence frame representing a different capture pose for the scene. It will be appreciated that in some embodiments, a single video sequence may be received including frames at different capture poses at each time point.

[0097] The server 103 further comprises a second receiver 303 configured to receive three-dimensional spatial data of the scene. The spatial data may indicate spatial and geometric properties of objects in the scene. The spatial data may specifically indicate spatial properties such as the position, orientation, extent, etc. of the objects in the scene. Such objects may include background objects of the scene. For example, the spatial data may describe the spatial extent of one or more objects in the scene.

[0098] The spatial data is independent of the input images and can describe objects of the scene independent of the pose of the input images. The spatial data can be common to multiple input images and / or multiple frames of a video sequence. The spatial data can specifically be a model of a three-dimensional scene.

[0099] The spatial data may specifically be one or more of a 3D point cloud, a mesh for a (static) background or (dynamic) foreground object, or a depth map. The spatial data may, for example, correspond to a depth map of the MVD representation.

[0100] The spatial data can specifically be spatial data provided by a laser system that measures an accurate point cloud. Alternatively, the spatial data can be provided via human-guided interaction with the source view imagery, for example to interactively create a mesh using a multi-view mesh editor. Or a computer graphics mesh model (e.g., of a sports stadium) can be used, fitted (rotated and translated) to the source view data and used as a static background model. Alternatively, an object detector for a given object class (athlete) or category (floor, wall, sky) can be run, and the resulting class or category labels can be used.

[0101] The spatial data can be "artificial" for static objects, in some embodiments, using what we call a multi-view mesh editor, in which case a human injects the information (the human can be considered a sensor). Another accurate source is often a laser that generates a point cloud. Such lasers are typically much more accurate than "standard" depth sensors based on structured light or time of flight. As another example, an object detector can be used that finds and matches the geometry of a basket in a basketball scene, for example. The spatial data can be data generated by a sensor different from the sensor that generates the input image. The spatial data can typically provide accurate geometry data of the scene. The spatial data can be geometric data of the scene.

[0102] The first receiver 301 and the second receiver 303 are coupled to a data stream generator 305 configured to generate an audiovisual data stream representing the scene. The data stream generator 305 is configured to include some or possibly all of the input images in the audiovisual data stream. The data stream generator 305 can in many embodiments be configured to encode the images / frames / video according to a suitable image or video encoding algorithm.

[0103] The data stream generator 305 is further configured to generate and include scene data representative of the 3D spatial data. In many embodiments, some or all of the 3D spatial data may be included in the audiovisual data stream. In some embodiments, generating the scene data may include encoding or compressing the 3D spatial data.

[0104] One example may include an entity map containing object category labels for multi-view patches similar to those known from the MPEG immersive video standard.

[0105] The server 103 further includes a view synthesis neural network, hereafter referred to as the server neural network 307. The server neural network 307 is configured to synthesize images of different view poses based on the spatial data and some or all of the captured input images.

[0106] The server neural network 307 is configured to receive a set of images, in particular frames of a video sequence, representing a scene from different view poses, and may further receive 3D spatial data, based on which the server neural network 307 may be configured to generate output images corresponding to view images of poses different from the input view pose.

[0107] More specifically, the server neural network 307 can be a convolutional neural network.

[0108] Before being input to the neural network, the 3D spatial data and / or the multi-view data may be passed through functions / operations that depend on the source view parameters and / or the target view parameters.

[0109] An advantageous exemplary approach is described below with reference to FIG.

[0110] A depth map generator 501 that generates a depth map for a target view. 3D spatial data are generally known or specified in a common world coordinate system, rather than being tied to a given source viewpoint. They are usually very accurate. As an example, a single laser can be used that captures a point cloud. As a first step, a depth map can be generated by projecting this point cloud onto a target view. Possible holes can be filled by interpolation and / or extrapolation.

[0111] Source-view convolutional network503 The network takes a full-resolution three (color) channel source view image as input and can stack a small number of convolutional layers (e.g., two layers), each with a stride (step [pixels]) greater than one, to downscale the successive feature maps corresponding to each layer. For three layers, a stride of four would be useful. As layers are stacked, the number of feature maps (=channels) can be increased from three input layers to eight, for two layers following the input (=image) layer. It is important to realize that these layers are "live" in source view image coordinates. Note that, as is common, each (linear) convolution operation is followed by a nonlinear (so-called activation) function. We propose to use a Rectified Linear Unit (ReLU), but other approaches are of course possible.

[0112] A generator 505 that generates a flow-field from the target view to each source view. Now that a (precise) depth map for the target views has been generated, this depth map can be used to compute the so-called flow field, which maps each pixel in the source view to a corresponding pixel in each target view. If the source view parameters are precisely calibrated and the laser (or other means) provides precise 3D data, this mapping is also very precise. However, occlusions by points in the source view are not detected, and thus this needs to be resolved by the target view convolutional network (see below). The computation of the flow field and the search for source view pixels to target views can be performed at (sub-millisecond) speed on an average GPU.

[0113] A connection circuit 507 for connecting the N source view networks to a target view network 509 To allow end-to-end training, a function graph can be used where all functions are differentiable. Because the flow field is defined from the target view back to each source view, it is differentiable and can therefore be inserted as a separate layer between other layers in the neural network. Frameworks such as PyTorch allow for the definition of such a "warlayer". Specifically, for PyTorch, this layer is torch.nn.functional.grid_sample. (See: https: / / pytorch.org / docs / stable / generated / torch.nn.functional.grid_sample.html)

[0114] In this example, the output of a given source view layer, which is a 4D tensor (batch size, channels, height, width), is fetched along with the flow field to the target view where it is represented at a specified resolution. This resolution can be exactly the same as the target view resolution or a lower resolution. In many cases, the resolution can be the same as the resolution from a particular layer in the source view. In a particular example, there is both a high-resolution source view input layer (= the actual source view image) and a low-resolution output layer of the source view network (the final feature map).

[0115] Target View Convolutional Network509 The target view network receives as input the warped layers from (a subset of) the source views. This network receives as an additional input the depth map provided by the laser data. The approach can proceed in two separate branches: 1. Depth-based mixing weight computation branch: The branch with the depth map as input is followed by two convolution layers using a 3x3 kernel with an upscaling step before each convolution. The output layer is specified to output N channels at the target view resolution, where N is the number of source views used. 2. Image-based Mixture Weights Computation Branch: This branch takes the warped source view output layer (feature maps) as input, upscales it to the target view resolution, convolves it with a 3x3 kernel, and also outputs one N-channel weight tensor for each source view.

[0116] Each branch provides an independent cue (based on depth change or chrominance) on how to perform occlusion-aware blending. In this example, the simplest solution is implemented, which is to simply average these weights into a final mixing weight tensor. This tensor is normalized so that its sum over the channels equals 1, and then multiplied by the concatenated warped source view tensor. The result is summed across channels to obtain the predicted target view image. Note that this is performed for each color channel of the warped source view.

[0117] Note that for training, the source views can be used as ground truth using the known leave-one-out method.

[0118] Note that the above strategy of having a network in each source view has the advantage that the source view network can compute informative features of occlusions and texture changes after warping to the target, which can help to better suppress occluded pixels in the blending operation.

[0119] Further information on neural networks for view synthesis operations can be found, for example, in Tewari et.al., "Advances in Neural Rendering", Computing Research Repository (CoRR) volume = {abs / 2111.05849}, 2021, https: / / arxiv.org / abs / 2111.05849, arXiv, timestamp = {Tue, 16 Nov 2021 12:12:31 +0100}.

[0120] The server 103 further comprises a neural network trainer 309 configured to train the server neural network 307 .

[0121] The neural network trainer 309 may use a training set that includes many images / frames of the scene from different view poses.

[0122] The neural network trainer 309 may use any suitable training technique, including performing gradient descent based on a cost function determined from the training data / reference images.

[0123] More specifically, one of the optimizers available in frameworks such as PyTorch or TensorFlow can be used. In particular, the so-called Adam optimizer can be used, as described, for example, in https: / / pytorch.org / docs / stable / generated / torch.optim.Adam.html?highlight=adam#torch.optim.Adam or Diederik P. Kingma and Jimmy Ba Adam: "A Method for 3rd International Conference for Learning Representations, San Diego, 2015.Stochastic Optimization".

[0124] Thus, the neural network trainer 309 can train the server neural network 307 for the task of generating view images of different view poses. Specifically, based on the images / frames and 3D spatial data contained in the audiovisual data stream, the trained server neural network 307 is configured to generate view images / frames of a scene from different view poses. The connections to the nodes / artificial neurons and / or the coefficients / weights for the kernels / filters of the convolutional layers are adapted / trained to perform such view shifting with high quality.

[0125] In many embodiments (as will be described in more detail below), nodes / neurons can have / be assigned values ​​that are determined as a weighted combination of input values ​​from nodes in previous layers. The weighted combination has coefficients, which are determined by training.

[0126] In many scenarios, the node values ​​may further include bias values, which may also be determined by a training process. In some embodiments, the generated audiovisual data stream may also include data describing the bias values ​​of the nodes, in addition to the coefficient data.

[0127] In this approach, the coefficients / weights of the trained network (or at least some of these coefficients) are provided to the data stream generator 305 and included in the audiovisual data stream. Thus, the audio data stream includes coefficient data describing the coefficients of the view synthesis neural network after training.

[0128] The coefficients may include coefficients of filters / kernels for one or more convolutional layers, or weights for connections of one or more fully connected layers.

[0129] The client 101 comprises a receiver 401 which receives the audiovisual data stream generated by the server 103 .

[0130] The client 101 further comprises a view synthesis neural network, hereafter referred to as the client neural network 403. The client neural network 403 is fed with the received images / frames and the received 3D spatial data from the received audiovisual data stream, and is configured to synthesize images for different view poses based on these images and the spatial data.

[0131] The server neural network 307 can be configured to receive a set of images, which may in particular be frames of a video sequence, for different view poses, and can further receive 3D spatial data, based on which the client neural network 403 can be configured to generate output images corresponding to view images of poses different from the view poses of the received images.

[0132] In many cases, the audiovisual data stream may include a video sequence of frames representing views of a scene from different view poses. The client neural network 403 may then generate an output video sequence including frames representing views of a scene from view poses that may change dynamically. The output frames are generated by applying the structure and operations of the client neural network 403 to the received image / frame and spatial data.

[0133] The client neural network 403 and the server neural network 307 may specifically be implemented with the same structure and therefore may have the same layers, nodes / neurons, connections, etc. Comments provided with respect to the server neural network 307 also apply mutatis mutandis to the client neural network 403 and vice versa.

[0134] The client 103 further comprises a neural network controller 405 coupled to the receiver 401 and the client neural network 403. The neural network controller 405 is configured to control the setting of coefficients / weights of the client neural network 403. Specifically, the neural network controller 405 is configured to process the received coefficient data to extract neural network coefficients encoded in the audiovisual data stream. The client neural network 403 is then configured to set the coefficients of the client neural network 403 from the received coefficient data. The neural network controller 405 may set the coefficients of the client neural network 403 to values ​​included in the coefficients.

[0135] Thus, the neural network controller 405 can be configured to set up the client neural network 403 to be identical to the trained server neural network 307, such that the trained server neural network 307 is replicated in the client 103. The neural network controller 405 can then set up the client neural network 403 to function in the same way as the trained server neural network 307, thereby ensuring that the client neural network 403 can perform high quality view shifting / image / frame synthesis.

[0136] The client 103 may include a view pose generator 407 that generates a view pose for which a view image / frame is generated. For example, the view pose generator 407 may be configured to receive a view pose of an observer (particularly within a scene). This view pose may represent a position and / or orientation from which the observer views the scene, and may in particular provide a pose for which a view of the scene should be generated. Many different approaches for determining and providing a view pose are known, and it will be appreciated that any suitable approach may be used. For example, the view pose generator 407 may be configured to receive pose data from a VR headset worn by a user, from an eye tracker, etc.

[0137] The view pose generator 407 can provide a view pose to the client neural network 403, and the client neural network generates an output view image / frame for the given view pose based on the received image and 3D spatial data.

[0138] The client neural network 403 is coupled to a renderer 409 configured to render the images / frames generated by the client neural network 403. The renderer may specifically generate a display signal representing the images / frames. The display signal may be provided to a display that displays the generated view images / frames to a viewer / user.

[0139] And in this approach, the view synthesis and shifting can be based on a neural network approach, which typically results in higher quality images and / or reduced complexity. Moreover, this approach can allow the advantageous operations to be performed without requiring training of the neural network on the client side. Rather, the training can be performed on the server side. This can typically allow for improved training, since more appropriate training data is often available on the server side. Often, the server side has access to additional images for the scene that are not communicated to the client side. Furthermore, the images resulting from such operations are not needed and are often discarded. Although the inclusion of neural network and training operations to generate the view shift on the server side may increase the complexity on the server side, this is usually significantly outweighed by the operation and potentially improved quality on the client side, as well as the reduced computational requirements on the client side. This is a major advantage, since a single server may typically serve multiple, and potentially many, clients.

[0140] In this approach, a server or other source device can encode images, such as video frames, spatial data, and neural network coefficients in a bitstream. The neural network coefficients can be determined prior to encoding with the intention of maximizing view synthesis quality with available spatial data, given the limited pixel space in the video atlas. The video frames can be (a subset of) the received source views, or a transform thereof.

[0141] As a particular example, multi-view data and depth data (e.g., collected using a laser) can be input to a synthetic neural network (Server Neural Network 307). The coefficients of the neural network may be coded separately or may be coded as a video, for example, using a video atlas.

[0142] It will be understood that any suitable approach for training the server neural network 307 (and thus, implicitly, the client neural network 403) may be used depending on the particular preferences and requirements of each individual embodiment.

[0143] In many embodiments, the neural network trainer 309 is configured to train the server neural network 307 using a training set of images including images received by the first receiver 301. In particular, input images (such as frames of a video sequence) capturing a scene from different view poses may be received. Some or all of these may be provided to the data stream generator 305 and included in the audiovisual data stream. These images may be used as inputs to the neural network when performing view synthesis. The received captured images may also be used as reference images for training the server neural network 307.

[0144] The neural network trainer 309 can be specifically configured to train the server neural network 307 by minimizing a cost function that reflects the difference between synthetic images generated for different view poses and a reference image for the same view pose. In some scenarios, the synthetic view pose can correspond to a view pose where captured images are also received and these images can be used as reference images.

[0145] In many embodiments, parameter training and fitting is performed using inputs to the server neural network 307 that have been processed according to at least some of the processing performed as part of the communication to the client neural network 403 .

[0146] In particular, the neural network trainer 309 may be configured to encode and decode some or all of the input image before providing it as an input image to the client neural network 403. The encoding and decoding may coincide with the encoding performed by the data stream generator 305 and the complementary decoding performed at the client 103. Thus, the images provided to the server neural network 307 may more closely correspond to the images provided to the client neural network 403 when performing image synthesis. The encoding and decoding operations may generally include all operations performed as part of the transmission path, including, for example, pruning, packing, or pixel rate reduction.

[0147] Similarly, the neural network trainer 309 can be configured to encode and decode the three-dimensional spatial data of the scene before it is provided to the view synthesis neural network. The encoding and decoding can be the same operations that are applied to the spatial data as part of the communication to the client neural network 403. Thus, the spatial data provided to the server neural network 307 may more closely correspond to the spatial data provided to the client neural network 403.

[0148] The encoding and decoding operations can be dynamically adapted to reflect the particular encoding and decoding parameters currently being used to communicate data to the client 103. While the data provided to the client neural network 403 for synthesis may be processed to reflect effects and degradations resulting from the communication operations, the reference images are not subjected to such processing and the input images used as reference images are used unaltered.

[0149] Such an approach allows for improved training of the server neural network 307, and in particular for the client neural network 403 to be implemented based on received coefficients that are not only optimized for view shift but can also mitigate / compensate for some of the effects of communication operations. Indeed, for a given set of coding parameters, the training of the neural network uses non-ideal data that has undergone video / image compression (both image and spatial data), but still optimizes the synthesis network to predict ideal data.

[0150] Typically, the set of captured images and capture locations is relatively limited, which can affect the accuracy of training the server neural network 307. In many embodiments, the neural network trainer 309 is configured to train the server neural network 307 using a training set of images that includes reference images for view poses not represented by the input images.

[0151] For example, a set of training data can be generated to include view poses distributed over a given area, which may correspond to, for example, a target view region. For example, as shown in the top example of FIG. 6, input data can be received for five view poses. For example, a scene can be captured by five different cameras. Input images can be provided, in this case, for the five different poses, which can be provided as input images to the server neural network 307 and communicated to the client neural network 403.

[0152] However, to provide additional training data, and because this represents an increased range of view poses, reference images can be generated for other view poses than the capture pose. These additional reference images can be generated to be distributed across the expected viewing region, as shown by the example in the lower diagram of FIG. 6. Thus, not only can additional images be used for training, but they can be generated to cover a wider area, thereby providing better training of the neural network for synthesis of images across the entire expected viewing region.

[0153] The additional reference images can be generated in some embodiments by view-shifting based on the captured input images, where the view-shifting is not based on a neural network. Thus, for a given view pose in the viewing region, a view synthesis operation can be performed based on the input images to generate a view image for that view pose.

[0154] The view synthesis operation is performed using a view synthesis operation that is not based on the server neural network 307. It will be appreciated that many different view shifting and image synthesis algorithms are known to those skilled in the art and any suitable approach may be used.

[0155] In some embodiments, some or all of the additional reference images may be generated by processing a visual scene model of the scene, such as a graphics model. For example, a model may be provided for the virtual scene that can be evaluated to generate view images for different view poses. It will be appreciated that many different approaches for generating and evaluating visual scene models are known to those skilled in the art, and any suitable approach may be used.

[0156] Thus, in some embodiments, the training data may include a set of reference images corresponding to captured input images and a set of reference images for non-captured view poses. Images in the latter set of reference images are also referred to as reference images to reflect that they may typically be generated by non-neural network based view shifting or model evaluation.

[0157] In such an embodiment, the neural network trainer 309 is configured to apply different weightings to different sets of reference images. In particular, reference images corresponding to an input image can be weighted higher than simulated reference images when training and fitting the neural network. The weighting of each reference image can be done by specifically weighting the images differently when generating a cost function used for training (e.g., by a gradient descent approach). For example, an error determined for an input image may contribute more to the cost function than a corresponding error determined for a simulated image. In many embodiments, improved training can be achieved by differential weighting for different types of reference images.

[0158] The described approach can mitigate the risk that if only the input image is used as a reference image, the neural network will be trained to essentially learn to predict only the virtual view poses that are part of the capture configuration, since the source view poses are the only poses the network can be trained on. This problem can be mitigated by adding simulated reference image / example data to the training set.

[0159] This approach can specifically involve a dataset-dependent cost function with weights between zero and one that balances the observation region extrapolation quality and overfitting on the captured source view. For example, the cost function can be determined as the sum of the contributions given by for each image:

number

[0160] In some embodiments, the training of the server neural network 307 may be based on evaluating the output using a second neural network trained to evaluate the output quality of images of the client neural network 403.

[0161] In some embodiments, the server neural network 307 and / or the client neural network 403 can be initialized with a set of default coefficients. For example, a set of default coefficients that have been found to be efficient for many common scenes and scenarios can be stored. The default coefficients can be shared, for example, between the server and the client (e.g., during a previous session), and thus the same default coefficients can be loaded into the server neural network 307 and the client neural network 403.

[0162] In such a system, a neural network trainer 309 can initialize the server neural network 307 with default coefficients and can initialize training sequences based on the received video sequences / images / frames in particular. Based on the training, modified coefficients can be determined, some or all of which can be included in the audiovisual data stream and transmitted to the client.

[0163] Similarly, the neural network controller 405 can initialize the client neural network 403 with default coefficients, which can be, for example, fixed coefficients or can be, for example, coefficients stored during a previous session. Thus, when initializing a new session, the neural network controller 405 can be configured to immediately begin operation without requiring the neural network coefficients to be received from a server.

[0164] Once the modified coefficients are determined by the training operation, they can be included in the audiovisual data stream and, once received by the client, the neural network controller 405 can proceed to set the coefficients of the client neural network 403 to these modified values. Thus, the client neural network 403 can start with default coefficients, which can then be updated as modified coefficients are received. Thus, in some embodiments, the neural network controller 405 is configured to initialize the neural network with default coefficients and override the default coefficients with coefficients determined from the coefficient data.

[0165] In some embodiments, the data stream generator 305 can be configured to include all modified coefficients in the audiovisual data stream when the training operation is completed, thus including all coefficients determined for the client neural network 403. Similarly, when received at the client, the neural network controller 405 can be configured to set all coefficients of the client neural network 403 to the values ​​received from the server.

[0166] In some embodiments, only a subset of the coefficients may be communicated. Specifically, the data stream generator 305 may select a subset of coefficients for transmission depending on the difference between the modified coefficients and the default coefficients. For example, only modified coefficients that differ from the default (or previous) coefficients by more than a threshold may be transmitted to the client. Similarly, when receiving coefficient data that includes values ​​of only a subset of the coefficients, the neural network controller 405 may proceed to modify only these coefficients. Such an approach may reduce the data rate required to communicate the coefficients, may reduce the data rate of the audiovisual data stream, and / or allow for faster update rates for the neural network coefficients.

[0167] In some embodiments, the neural network trainer 309 and / or the neural network controller 405 are configured to select default coefficients from multiple sets of default coefficients. For example, several different sets of coefficients can be stored, and the neural network trainer 309 and / or the neural network controller 405 can select between them based on the received image.

[0168] For example, coefficients can be determined during a previous training sequence for capturing scenes during the day on a bright summer day, at dusk on a summer day, during a snowy winter day, during a cloudy and dry day in autumn, etc. Corresponding default coefficient sets can be stored.

[0169] When a new session is initialized, the neural network trainer 309 and / or the neural network controller 405 can analyze one or more of the initial images to determine which stored scenario is most likely to be represented. For example, the selection can be based on brightness and color parameters. If the image is very bright and contains many warm colors, a summer day coefficient can be selected, if the colors are unsaturated and the brightness is relatively low, an autumn day coefficient can be selected, and if the image is bright and the predominant color is white, a winter day coefficient can be selected.

[0170] The coefficient sets, in some embodiments, can differ for hardware types / settings (camera brands, exposure settings, etc.), etc.

[0171] Such an approach can provide improved initial performance, facilitate and improve initial training, and reduce the perceptual impact of subsequent updates to the coefficients.

[0172] In many embodiments, the described approaches can be used with video sequences, so that time-varying images can be provided. Each video sequence can be, for example, a time series of frames, and the neural network can be configured to generate frames for a given time point by applying neural network operations to input frames at that time. For example, frames closest (in time) to the time point at which the frame is to be generated can be selected in each video sequence, and neural network operations can be applied to these.

[0173] Thus, in some embodiments, the audiovisual data stream includes time-varying video data. In many embodiments, the audiovisual data stream may include time-varying coefficient data for a neural network.

[0174] For example, the neural network trainer 309 can be configured to repeatedly perform training of the server neural network 307. For example, at regular intervals, the neural network trainer 309 can control the server to perform training of the server neural network 307, resulting in new updated coefficients. These updated coefficients can then be encoded and included in the audiovisual data stream. Thus, the audiovisual data stream can include, for example, at regular intervals, coefficient data with new modified coefficients. The new modified coefficients can be further applied to the server neural network 307 and used as initial coefficients for the next training operation.

[0175] The client 103 can receive the audiovisual data stream having this time-varying coefficient data describing the time-varying coefficients of the view synthesis neural network, and can modify the neural network's coefficients in response to the time-varying coefficient data, typically by overwriting some or all of the coefficients of the client neural network 403.

[0176] In many embodiments, the neural network controller 405 can be configured to update the coefficients when new data is received. However, in some embodiments, the neural network controller 405 may be configured to update the coefficients when no coefficients are received, i.e., to apply time-varying coefficients to the client neural network 403.

[0177] In many embodiments, the neural network controller 405 can be configured to generate interpolated coefficients for time points for which no coefficient data is provided. For example, if the audiovisual data stream includes a set of new coefficients at several discrete time points, e.g., at regular time intervals, the neural network controller 405 can interpolate for times between such time points by interpolating from the received coefficients. Thus, the client neural network 403 can be updated more frequently than new coefficients are received by using the interpolated coefficients to overwrite current coefficients.

[0178] As a particular example, an audiovisual data stream may include new coefficient data at regular intervals, for example, 10 seconds. However, the neural network controller 405 may interpolate between the given coefficient values ​​to generate new modified coefficient values, for example, every second. The neural network controller 405 may then update the coefficients of the client neural network 403 every second.

[0179] Such an approach allows for significantly reduced data rates and / or reduced computational load, since the time- and computationally demanding training operations can be performed less frequently and less additional data can be generated and communicated. However, improved quality can still be achieved by providing better update rates for the client neural network 403 operations.

[0180] It will be appreciated that any suitable approach can be used to determine the interpolated coefficients. For example, the interpolation may be performed for each coefficient individually, e.g., a simple linear interpolation can be performed between the received coefficient values ​​for time points before the current time point and the received coefficient values ​​for time points after the current time point. In other embodiments, more complex interpolation operations can be used.

[0181] In some embodiments, the coefficient data may provide values ​​for all of the coefficients of the neural network, however, in some embodiments, the coefficient data may only include data for a subset of the coefficients.

[0182] In some embodiments, the neural network can be configured to have some adjustable layers and some non-adjustable layers. Non-adjustable layers can be layers whose coefficients are fixed and cannot be changed. In particular, the coefficients of non-adjustable layers are not modified or changed by training. The coefficients of non-adjustable layers can be set to predetermined default values, for example, and they are fixed and are not modified when the neural network trainer 309 performs training of the server neural network 307. For example, in the described embodiment, the source view network parameters are assumed to be constant and can be known by the client, while only the target view network parameters are transmitted.

[0183] Tunable layers have coefficients that vary, specifically, coefficients that can be modified or changed by training operations, and thus, when performing training of the server neural network 307, only the coefficients of the tunable layers are changed / optimized.

[0184] In such an embodiment, the data stream generator 305 may be configured to include coefficients only for the adjustable layers, and not for the non-adjustable layers.

[0185] The neural network controller 405 can be configured to realize the client neural network 403 at the client to correspondingly comprise one or more non-adjustable layers having fixed coefficients. Additionally, the client neural network 403 is implemented to have one or more adjustable layers. The neural network controller 405 can be configured to set the coefficients of the non-adjustable layers to default values ​​(or they can be fixedly implemented as part of the client neural network 403, for example) and set the coefficients of the adjustable layers based on coefficient data.

[0186] Such an approach can in particular be combined with a dynamic approach of time-varying coefficient data, where the time-varying coefficient data is only provided for the adjustable layers.

[0187] The use of adjustable and non-adjustable layers is highly advantageous for the described approach, since the combination of adaptive and fixed layers is particularly suitable for view synthesis based on distributed learning. In particular, partial adaptation of training allows for significantly reduced data rates in some cases, since fewer coefficients are communicated but still high quality results can be achieved.

[0188] In some embodiments, the server can be configured to transmit a feature map of the image. For example, rather than including a particular image in the audiovisual data stream, the server can proceed to extract a set of feature values ​​for a given layer of a convolutional neural network that forms the server neural network 307.

[0189] For example, in the particular example of the convolutional network mentioned above, the server can, for example, pack the low-resolution output layers of the source-view network and send it to the client using a suitable atlas approach. This has computational advantages for the client (no need to evaluate the first part of the network). On-chip HEVC decoding of video pixel data can often be faster and less power consuming than using a GPU or a dedicated neural engine on the client device.

[0190] Thus, rather than providing explicit images, feature values ​​generated by applying appropriate filters can be communicated. Feature maps can be provided specifically for poses that are not represented by any of the images / frames included in the audiovisual data stream, and thus represent the scene from additional view poses. Although the feature maps represent view images from the corresponding view poses, they typically have a much lower resolution, and therefore the data rate of the audiovisual data stream can be significantly reduced compared to transmitting a complete image.

[0191] As a specific example, an approach can be adopted that packs the low-resolution output layers of the source-view network and sends it to the client in an atlas. This can have computational advantages for the client (no need to evaluate the first part of the network). Furthermore, on-chip HEVC decoding of video pixel data can often be faster and less power consuming than using a GPU or a dedicated neural engine on the client device.

[0192] Thus, the audiovisual data stream may, in some embodiments, include at least one neural network feature map for an image with a view pose that is different from the view poses of the multiple images.

[0193] The neural network controller 405 can then be configured to extract the feature values ​​and configure the neural network with the at least one feature map. In particular, the values ​​of the nodes of the layer in which the feature map is provided can be set to the feature values ​​indicated in the feature map.

[0194] The client neural network 403 can then execute a second portion of the neural network.

[0195] Such an approach is typically advantageous over neural network based view synthesis operations.

[0196] The data stream generator 305 may be configured to generate an audiovisual data stream according to any suitable data format, structure or standard. In many embodiments, the audiovisual data stream may be generated according to a video data structure that includes video data as well as metadata or auxiliary data, which may include coefficient data and / or spatial data.

[0197] In many embodiments, the data stream generator 305 can be configured to structure the coefficient data to match the video data and structure.

[0198] Specifically, video data (image data of frames included in an audiovisual data stream) can be coded using inter-frame prediction. Video coding can employ a Group Of Pictures (GOP) structure. In such a GOP structure, frames are divided into GOPs, each GOP being a collection of frames that can be decoded without requiring information from pictures / frames belonging to another GOP. Within a given GOP, inter-frame prediction and relative coding are based only on other frames that are part of the same GOP, and not on frames of another GOP. Thus, inter-frames are coded only based on other frames of the same GOP. Each GOP typically comprises at least one intra-frame. The GOP structure may reflect the arrangement of inter-frames and intra-frames for a set of frames.

[0199] In some embodiments, the data stream generator 305 can be configured to structure at least two of the image data, scene data and coefficient data to share a GOP structure. Specifically, the data stream generator 305 can be configured to structure the scene data and / or the coefficient data to use the same GOP structure as the image / video frame.

[0200] For example, a new set of coefficients can be provided for each GOP structure, e.g., for an intra-coded frame, a set of coefficients can be provided as a set of intra-coded coefficients, and further coefficients can be provided for an inter-coded frame that are inter-coded based on the coefficients provided for other frames, specifically, coded as relative values ​​to the coefficients of the intra-coded coefficients provided for the intra-coded frame.

[0201] Such an approach allows for an efficient and practical approach to communicating coefficients.

[0202] An artificial neural network suitable for implementing a view synthesis neural network is a network of nodes organized in layers, where each node holds a node value. Figure 7 shows an example of a section of an artificial neural network.

[0203] The node value of a given node can be calculated to include contributions from some or often all nodes of the previous layer of the artificial neural network. Specifically, the node value of the node can be calculated as a weighted sum of the node values ​​of all the node outputs of the previous layer, and weights / coefficients can be determined for the node (by training). Typically, a bias may be added and the result can be subjected to an activation function. The activation function typically provides the most important role of each neuron by providing nonlinearity. Such nonlinearity and activation function provide significant influence in the learning and adaptation process of the neural network. Thus, the node value is generated as a function of the node value of the previous layer.

[0204] The artificial neural network may specifically include an input layer 701 that includes a plurality of nodes that receive input data values ​​to the artificial neural network. Thus, the node values ​​for the nodes of the input layer may typically be direct input data values ​​to the artificial neural network and therefore may not be calculated from other node values.

[0205] An artificial neural network may not include any hidden layers 703 or processing layers, or may further include one or more hidden layers 703 or processing layers. For each such layer, node values ​​are typically generated as a function of the node values ​​of the nodes in the previous layer, in particular an activation function (such as a sigmoid, ReLU, or Tanh function) may be applied, followed by a weighted combination and an added bias.

[0206] Specifically, as shown in Figure 8, each node, which may also be called a neuron, can receive input values ​​(from nodes in the previous layer) and then calculate the node value as a function of these values. Often this involves first generating a value as a linear combination of the input values, each weighted by a weight / coefficient:

number

[0207] In the described approach, coefficients / weights of linear combinations of input values ​​and / or values ​​of nodes of previous layers can be determined by training. The coefficients / weights determined by training can then be described by coefficient data included in the audiovisual data stream. In many embodiments, bias values ​​(offsets) can be determined for one or more of the nodes, and in many embodiments data describing the bias values ​​can also be included in the audiovisual data stream.

[0208] An activation function can then be applied to the resulting combination. For example, a node value l can be determined as l=f(k).

[0209] where the function can be, for example, a Rectified Linear Unit function as described in Xavier Glorot, Antoine Bordes, Yoshua Bengio Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15:315-323, 2011): f(k)=ReLU(k)=max(0,k)

[0210] Other commonly used functions include sigmoid or tanh functions. In many embodiments, node outputs or values ​​can be calculated using multiple functions. For example, both ReLU and Sigmoid functions can be combined using an activation function such that f(k)=ReLU(k)+σ(k).

[0211] Such operations can be performed by each node of the artificial neural network (typically except for the input node).

[0212] The artificial neural network further comprises an output layer 705 which provides output from the artificial neural network, i.e. the output data of the artificial neural network are the node values ​​of the output layer. As for the hidden / processing layers, the output node values ​​are generated by functions of the node values ​​of the previous layer. However, in contrast to the hidden / processing layers, where the node values ​​are typically not accessible or further used, the node values ​​of the output layer are accessible and provide the results of the operation of the artificial neural network.

[0213] Several different network structures and toolboxes for artificial neural networks have been developed, and in many embodiments, artificial neural networks can be based on the adaptation and customization of such networks. One example of a network architecture suitable for such applications is the Long short-term memory (LSTM), described in Hochreiter, Sepp, and Jurgen Schmidhuber, "Long short-term memory." Neural computation 9.8 (1997): 1735-1780.

[0214] LSTM is an architecture used for classification and regression of time-domain signals using recursive causal or bidirectional evaluation, and has been successfully applied to audio signals. For example,

number

[0215] In theory, classical (or "vanilla") artificial neural networks can track any long-term dependency in the input sequence. The problem with vanilla artificial neural networks is computational (or practical) in nature: when training vanilla artificial neural networks using backpropagation, the long-term gradients that are backpropagated can "vanish" (i.e., can become zero) or "explode" (i.e., can become infinite) due to the calculations involved in the process that use finite precision numbers. Artificial neural networks that use LSTM units partially solve the vanishing gradient problem, since the LSTM units allow the gradients to flow unchanged as well. However, LSTM networks can still suffer from the exploding gradient problem.

[0216] In some cases, the artificial neural network can be further configured to include additional contributions that allow the artificial neural network to be dynamically adapted or customized to a particular desired characteristic or feature of the output to be generated. For example, a set of values ​​can be provided to adapt the artificial neural network. These values ​​can be included by providing contributions to some nodes of the artificial neural network. These nodes can specifically be input nodes, but typically nodes of hidden layers or processing layers. Such adaptation values ​​can be weighted and summed, for example, as contributions to a weighted summation / correlation value for a given node.

[0217] The above description relates to neural network approaches that may be suitable for many embodiments and implementations. However, it will be appreciated that many other types and structures of neural networks can be used. Indeed, many different approaches for generating neural networks have been developed, including neural networks that use complex structures and processes different from those described above. This approach is not limited to any particular neural network approach, and any suitable approach can be used without detracting from the invention.

[0218] An artificial neural network is adapted for a particular purpose by a training process that is used to adapt / tune / modify the weights and other parameters (e.g., biases) of the artificial neural network. It will be appreciated that many different training process algorithms are known for training artificial neural networks. Typically, training is based on a large training set where a large number of examples of input data are provided to the network. Furthermore, the output of the artificial neural network is typically compared (directly or indirectly) to an expected or ideal outcome. A cost function can be generated to reflect the desired outcome of the training process. In a typical scenario known as supervised learning, the cost function often represents the distance between the prediction and the ground truth for a particular input data. Based on the cost function, the weights can be modified, and by repeating the process with the modified weights, the artificial neural network can be adapted to a state where the cost function is minimized.

[0219] More specifically, during the training step, a neural network may have two different information flows: from input to output (forward pass) and from output to input (backward pass). In the forward pass, data is processed by the neural network as described above, and in the backward pass, weights are updated to minimize the cost function. Typically, such backward propagation follows the gradient direction of the cost function landscape. In other words, by comparing the predicted output for a batch of data inputs with the ground truth, the direction in which the cost function is minimized can be estimated and propagated backward by updating the weights accordingly. Other approaches known for training artificial neural networks include, for example, the Levenberg-Marquardt algorithm, the conjugate gradient method, and the Newton method.

[0220] In this case, the training may in particular include a training set including a potentially large number of pairs of captured audio signals and measured respiratory waveform signals. The measured respiratory waveform signals may be generated from sensor signals of a sensor configured to measure a lung air volume dependent characteristic. Thus, training of the artificial neural network may be performed using a training set comprising linked audio / speech data / signals and sensor data / signals representing lung air volume measurements during speech.

[0221] In some embodiments, the training data is an audio signal in a time segment corresponding to the processing interval of the artificial neural network being trained, for example, the number of samples in the training audio signal may correspond to the number of samples corresponding to the input nodes of the artificial neural network being trained. Thus, each training example may correspond to one operation of the artificial neural network being trained. However, typically, to speed up the training process, a batch of training samples is considered for each step. Furthermore, many upgrades to gradient descent are also possible to speed up the convergence or to avoid local minima in the cost function landscape.

[0222] The apparatus can in particular be implemented in one or more suitably programmed processors. For example, the neural network can be implemented in one or more such suitably programmed processors. Different functional blocks, in particular the artificial neural network, can be implemented in separate processors and / or for example in the same processor. Examples of suitable processors are given below.

[0223] 9 is a block diagram illustrating an exemplary processor 900 according to an embodiment of the disclosure. The processor 900 can be used to implement one or more processors that implement the apparatus or elements thereof as described above, including, inter alia, one or more artificial neural networks. The processor 900 can be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) (wherein the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application specific integrated circuit (ASIC) (wherein the ASIC is designed to form a processor), or a combination thereof.

[0224] Processor 900 may include one or more cores 902. Core 902 may include one or more arithmetic logic units (ALUs) 904. In some embodiments, core 902 may include a floating point logic unit (FPLU) 906 and / or a digital signal processing unit (DSPU) 908 in addition to or instead of the ALU 904.

[0225] The processor 900 may include one or more registers 912 communicatively coupled to the cores 902. The registers 912 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 912 may be realized using static memory. The registers may provide data, instructions, and addresses to the cores 902.

[0226] In some embodiments, the processor 900 may include one or more levels of cache memory 910 communicatively coupled to the cores 902. The cache memory 910 may provide computer-readable instructions to the cores 902 for execution. The cache memory 910 may provide data for processing by the cores 902. In some embodiments, the computer-readable instructions may be provided to the cache memory 910 by a local memory, for example, a local memory attached to an external bus 916. The cache memory 910 may be implemented using any suitable cache memory type, such as, for example, static random access memory, dynamic random access memory, and / or any other suitable memory technology.

[0227] The processor 900 may include a controller 914 that may control inputs to the processor 900 from other processors and / or components in the system and / or outputs from the processor 900 to other processors and / or components in the system. The controller 914 may control data paths within the ALU 904, FPLU 906 and / or DSPU 908. The controller 914 may be implemented as one or more state machines, data paths and / or dedicated control logic. The gates of the controller 914 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0228] Registers 912 and cache memory 910 may communicate with controller 914 and core 902 via internal connections 920A, 920B, 920C and 920D. The internal connections may be implemented as buses, multiplexers, crossbar switches and / or any other suitable connection technology.

[0229] Input and output for the processor 900 may be provided via a bus 916, which may include one or more conductive lines. The bus 916 may be communicatively coupled to one or more components of the processor 900, such as the controller 914, the cache memory 910, and / or the registers 912. The bus 916 may be coupled to one or more components of the system.

[0230] The bus 916 can be coupled to one or more external memories. The external memory may include a read only memory (ROM) 932. The ROM 932 can be a masked ROM, an electrically programmable read only memory (EPROM), or any other suitable technology. The external memory can include a random access memory 933. The RAM 933 can be a static RAM, a battery backed up static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory can include an electrically erasable programmable read only memory (EEPROM) 935. The external memory can include a flash memory 934. The external memory can include a magnetic storage device such as a disk 936. In some embodiments, an external memory may be included in the system.

[0231] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention can optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention can be physically, functionally and logically implemented in any suitable way. Indeed functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.

[0232] In accordance with standard terminology in the field, the term pixel may be used to refer to pixel-related properties such as light intensity, depth, position, etc. of the portion / element of a scene represented by the pixel. For example, the depth of a pixel, i.e., pixel depth, may be understood to refer to the depth of the object represented by that pixel. Similarly, the brightness of a pixel, i.e., pixel luminosity, may be understood to refer to the brightness of the object represented by that pixel.

[0233] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, although certain features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0234] Furthermore, although individually recited, a plurality of means, elements, circuits or method steps may be implemented by, for example, a single circuit, unit or processor. Moreover, although individual features may be included in different claims, these may in some cases be advantageously combined, and their inclusion in different claims does not imply that the combination of features is not feasible and / or advantageous. Moreover, the inclusion of a feature in one category of claims does not imply limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Moreover, the order of features in the claims does not imply a particular order in which the features must be operated, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Moreover, a reference to the singular does not exclude a plurality. Thus, reference to "a", "an", "first", "second", etc. does not exclude a plurality. Reference signs in the claims are provided merely as a clarifying example and shall not be construed as limiting the scope of the claims in any manner.

Claims

1. A first receiver circuit configured to receive multiple images of a scene, wherein the scene is captured from different view poses, A second receiver circuit configured to receive three-dimensional spatial data of the aforementioned scene, A view synthesis neural network circuit configured to generate view shift images of the scene for different view poses from the plurality of images of the scene and the three-dimensional spatial data, A neural network trainer circuit configured to train the view synthesis neural network circuit based on the multiple images of the aforementioned scene, A generator circuit configured to generate a data stream including image data for at least one of the plurality of images of the scene, scene data representing the three-dimensional spatial data, and coefficient data describing the coefficients of the view synthesis neural network circuit after training, A device having.

2. The first receiver circuit is configured to receive a set of multiple input video sequences, the multiple input video sequences representing views of the scene from different view poses, The plurality of images in the aforementioned scene are frames of the set of input video sequences. The neural network trainer circuit is configured to dynamically train the view synthesis neural network circuit to generate coefficients for the view synthesis neural network circuit. The generator circuit is configured to generate the data stream to include a portion of the frames from the plurality of input video sequences, The generator circuit is configured to generate time-varying coefficient data, The apparatus according to claim 1, wherein the time-varying coefficient data describes the time-varying coefficients of the synthetic neural network circuit after training.

3. The neural network trainer circuit is configured to train the view synthesis neural network using multiple training images. The aforementioned plurality of training images include a first plurality of reference images, The apparatus according to claim 1, wherein the first plurality of reference images include a view pose not represented by at least one of the plurality of images of the scene.

4. The apparatus according to claim 3, wherein the first plurality of reference images include at least one image selected from a group consisting of reference images generated by non-neural network view shift using the plurality of images of the scene and reference images generated from a visual scene model of the scene.

5. The plurality of training images include a second plurality of reference images, The second set of reference images includes an image from the set of images of the scene, The apparatus according to claim 3, wherein the neural network trainer circuit is configured to apply different weights to the first plurality of reference images and the second plurality of reference images.

6. The apparatus according to claim 1, wherein the neural network trainer circuit is configured to encode and decode at least one of the plurality of images of the scene before the plurality of images of the scene are provided to the view synthesis neural network circuit.

7. The apparatus according to claim 1, wherein the neural network trainer circuit is configured to encode and decode the three-dimensional spatial data before the three-dimensional spatial data is provided to the view synthesis neural network circuit.

8. The neural network trainer circuit initializes the view synthesis neural network circuit with a plurality of default coefficients. The neural network trainer circuit is configured to train the view synthesis neural network circuit to determine the modified coefficients for the view synthesis neural network circuit. The apparatus according to claim 1, wherein the generator circuit is configured to include at least one modified coefficient in the data stream.

9. The apparatus according to claim 8, wherein the generator circuit is configured to select a subset of coefficients to transmit according to the difference between the modified coefficient and the default coefficient.

10. A receiver circuit configured to receive a data stream, the data stream comprising: image data for a plurality of images of the scene representing the scene captured from different view poses; three-dimensional spatial data of the scene; and coefficient data describing the coefficients of a view synthesis neural network. The view synthesis neural network is configured to generate view-shifted images of the scene for different view poses from the plurality of images of the scene and the three-dimensional spatial data, A neural network controller circuit configured to set the coefficients of the neural network according to the coefficient data, A device having.

11. The data stream includes multiple video sequences, The aforementioned data stream includes multiple frames, The aforementioned multiple frames represent views of the scene from the different view poses, The aforementioned multiple images of the scene are frames among the aforementioned multiple frames, The coefficient data includes coefficient data that varies over time. The aforementioned time-varying coefficient data describes the time-varying coefficients of the view synthesis neural network circuit. The apparatus according to claim 10, wherein the neural network controller circuit is configured to change the coefficients of the view synthesis neural network circuit in accordance with the time-varying coefficient data.

12. The neural network controller circuit is configured to determine interpolation coefficient values ​​for at least one time point in time when the data stream does not contain coefficient data. The interpolation coefficient value is determined from the coefficient value of the time-varying coefficient data. The apparatus according to claim 11, wherein the neural network controller circuit is configured to set the coefficients of the view synthesis neural network circuit to the interpolation coefficient values ​​for at least one time point.

13. The data stream includes at least one neural network feature map for an image with a different view pose from the view pose of the plurality of images of the scene, The apparatus according to claim 10, wherein the neural network controller circuit is configured to set up the view synthesis neural network circuit using the at least one feature map.

14. The neural network controller circuit initializes the view synthesis neural network circuit with a plurality of default coefficients. The apparatus according to claim 10, wherein the neural network controller circuit is configured to override the default coefficients with coefficients from the data stream.

15. The apparatus according to claim 14, wherein the neural network controller circuit is configured to select a plurality of default coefficients from a plurality of sets of default coefficients in accordance with the plurality of images of the scene.

16. The apparatus according to claim 10, wherein at least two of the image data, scene data, and coefficient data share a group of pictures in the data stream.

17. The aforementioned view synthesis neural network circuit has multiple adjustable layers and multiple non-adjustable layers. The apparatus according to claim 10, wherein the data stream includes coefficient data for an adjustable layer.

18. The apparatus according to claim 10, wherein the three-dimensional spatial data comprises video point cloud data for the scene.

19. A method for generating an audiovisual data stream, The steps include receiving multiple images of a scene captured from different view poses, The steps include receiving three-dimensional spatial data of the aforementioned scene, A step of generating view-shifted images of the scene for different view poses from the plurality of images of the scene and the three-dimensional spatial data, wherein the generation uses a view-compositing neural network. A step of training the view synthesis neural network based on images of the scene for different view poses, The steps include generating a data stream having image data for at least one of the plurality of images of the scene, scene data representing the three-dimensional spatial data, and coefficient data describing the coefficients of the view synthesis neural network after training, A method of having.

20. The steps include receiving a data stream containing image data for multiple images of a scene representing a scene captured from different view poses, three-dimensional spatial data of the scene, and coefficient data describing coefficients for a view synthesis neural network circuit, The steps include generating a view-shifted image of the scene for different view poses from the multiple images of the scene and the three-dimensional spatial data using the view synthesis neural network circuit, The steps include setting the coefficients of the view synthesis neural network circuit according to the coefficient data, A method of having.

21. A computer program stored in a non-temporary medium that, when executed on a processor, performs the method described in claim 19.

22. A computer program stored in a non-temporary medium that, when executed on a processor, performs the method described in claim 20.

23. A device having a processor circuit and a memory circuit, The memory circuit is configured to store instructions for the processor circuit, The processor circuit is configured to receive multiple images of a scene captured from different view poses. The processor circuit is configured to receive the three-dimensional spatial data of the scene. The processor circuit is configured to use a view synthesis neural network to generate view-shifted images of the scene for different view poses from multiple images of the scene and the three-dimensional spatial data. The processor circuit is configured to train the view synthesis neural network based on images of the scene for different view poses. The processor circuit is configured to generate a data stream, and the data stream is Image data for at least one of the plurality of images of the aforementioned scene, Scene data representing the aforementioned three-dimensional spatial data, and Coefficient data describing the coefficients of the aforementioned view synthesis neural network circuit after training, A device including a device.

24. A device having a processor circuit and a memory circuit, The memory circuit is configured to store instructions for the processor circuit, The processor circuit is configured to receive a data stream, and the data stream is Multiple images of a scene representing a scene captured from different view poses, Three-dimensional spatial data for the aforementioned scene, Coefficient data describing the coefficients of a composite neural network, Includes, The processor circuit is configured to use the view synthesis neural network to generate view-shifted images of the scene for different view poses from the multiple images of the scene and the three-dimensional spatial data. An apparatus wherein the processor circuit is configured to set the coefficients of the view synthesis neural network according to the coefficient data.