Method and apparatus for providing a rendering engine model including a neural network description embedded in a media item

By embedding the rendering engine model in the media item, the rendering problem caused by the lack of compatible software on the user device is solved, and customized processing of rendering media data is achieved without the need to install additional tools.

CN113196302BActive Publication Date: 2025-10-03QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980081877.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-17
Filing Date
2019-12-13
Publication Date
2025-10-03
Estimated Expiration
2040-04-06

AI Technical Summary

Technical Problem

When rendering different types of media files, user devices often encounter compatibility issues due to the lack of compatible media software, resulting in the inability to render media content normally, and the standardization process is cumbersome and complicated.

Method used

Embed rendering engine models, including neural network descriptions, within media items to process and render raw media data without relying on separate media rendering tools on the device.

Benefits of technology

It enables rendering of media data on different devices without the need for additional software installation, reduces standardization requirements, and supports customized processing of various media types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113196302B_ABST
    Figure CN113196302B_ABST
Patent Text Reader

Abstract

Techniques and systems are provided for providing a rendering engine model for raw media data. In some examples, the system obtains media data captured by a data capture device and embeds a rendering engine model in a media item containing the media data. The rendering engine model includes a description of a neural network configured to process the media data and generate a specific media data output, the description defining a neural network architecture of the neural network. The system then outputs the media item with the rendering engine model embedded in the media item. The rendering engine model indicates how to execute the neural network based on the description of the neural network to process the media data in the media item and generate the specific media data output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims priority to non-provisional application No. 16 / 222,885, filed on December 17, 2018, entitled “EMBEDDED RENDERING ENGINE FOR MEDIA DATA,” which is assigned to the assignee of the present application and is expressly incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to media rendering and, more particularly, to embedding a rendering engine model into a media item for rendering raw media data in the media item. Background Art

[0004] The increased versatility of media data capture products, such as digital cameras and microphones, has enabled media capture capabilities to be integrated into a variety of devices. Users can capture video, images, and / or audio from any device equipped with such media capture capabilities. Video, images, and audio can be captured for applications such as entertainment, professional use, monitoring, and automation. Media capture devices can capture raw media data, such as raw image or audio data, and generate files or streams containing the raw media data. In order to render the raw media data in a file or stream, a separate rendering engine (e.g., a decoder) suitable for the specific file or stream is necessary. For example, a video decoder is necessary for rendering and viewing content in a video file, while an image decoder is necessary for rendering and viewing content in an image file. Such a rendering engine is capable of understanding, processing, and rendering the raw media data in a file or stream.

[0005] There are a wide variety of formats and types for media data, as well as software tools for rendering different file types. Different files have different requirements, specifications, and behaviors, and are only compatible with certain media software tools. If a user wants to render a specific media file on their device, they need to install the appropriate media software for that media file. Unfortunately, in many cases, the user's device may not have the necessary media software for a specific media file. As a result, users are often unable to render media files on their devices unless they can find and install the necessary media software for each media file they want to render. This process can be frustrating, often preventing users from accessing media content, hindering their media content experience, or forcing users to search for and install the appropriate media software each time they want to render a media file that is incompatible with the software on their device. The compatibility issues caused by the different types of available media files have prompted various efforts to improve interoperability through standardization. However, standardization can be cumbersome, leading to a lengthy and complex process when adding new media functions and features. Summary of the Invention

[0006] Technology described herein can be implemented as creating media items (e.g., media files or streams), which contain original or captured media data (e.g., image data, video data, audio data, etc.) and the complete specification of a rendering engine for the original media data. A rendering engine model can be embedded in the media items together with the original media data. The rendering engine model in the media items allows a device to process and render the original media data without a separate media rendering tool (e.g., decoder, encoder, image processor, neural network processor and / or other media rendering tool). Therefore, the media items are pre-equipped with tools for rendering the original media data, regardless of the type or format of the media items, thereby avoiding the frustrating compatibility issues that typically occur when attempting to render different types of media items. Based on the rendering engine model in the media items, the user can run the rendering engine defined for the media data and render the media data from the user's device without first having to ensure that the device has a separate media rendering tool compatible with the media items to be rendered.

[0007] By including a rendering engine model or specification in a media item, any suitable processor on a device can use the raw media data and the rendering engine model in the media item to generate a rendering output. In this way, the device does not need a separate and compatible media rendering tool for a specific media item. Therefore, the media items and methods herein can eliminate or significantly reduce the demand for media item standardization, support new solutions for processing media data, and provide customized processing for different types of media data. The rendering engine model in the media items herein can be designed for the raw media data in the media item and can be customized for specific rendering intent and results.

[0008] According to at least one example, a method for creating a media item that includes media data (e.g., original or captured media data) and an embedded rendering engine model for rendering the media data is provided. The method may include obtaining media data captured by a data capture device, embedding the rendering engine model in the media item that includes the media data, and providing (e.g., sending, storing, outputting, etc.) the media item with the rendering engine model embedded in the media item to one or more devices. The one or more devices may obtain the media item and store the media item, or execute a neural network using the rendering engine model in the media item to process and render the media data in the media item. The neural network may use the media data as input to generate rendered media data output.

[0009] The media data may include image data, video data, audio data, etc. A media item may include a file, a stream, or any other type of data container or object that may contain or encapsulate media data and a rendering engine model. The rendering engine model may include a description of a neural network configured to process the media data and generate a specific media data output. The rendering engine model in the media item may indicate how to execute the neural network based on the description of the neural network to process the media data and generate a specific media data output. For example, the rendering engine model in the media item may inform one or more devices and / or any other devices that have a copy of the media item how to execute the neural network to process the media data and generate a specific media data output.

[0010] The description of a neural network may define a neural network architecture of the neural network. The neural network architecture may include, for example, a neural network structure (e.g., number of layers, number of nodes in each layer, layer interconnections, etc.), neural network filters or operations, activation functions, parameters (e.g., weights, biases, etc.), etc. The description may also define how the layers in the neural network are interconnected, how inputs to the neural network are formed, and how outputs are formed from the neural network. In addition, the description may define one or more tasks of the neural network, such as one or more custom tasks for encoding media data, decoding media data, compressing or decompressing media data, performing image processing operations (e.g., image restoration, image enhancement, demosaicing, filtering, scaling, color correction, color conversion, noise reduction, spatial filtering, image rendering, etc.), performing frame rate conversion (e.g., upconversion, downconversion), performing audio signal modification operations (e.g., generating a wideband audio signal from a narrowband audio input file), etc.

[0011] In another example, an apparatus for creating a media item comprising media data (e.g., original or captured media data) and an embedded rendering engine model for rendering the media data is provided. The example apparatus may include a memory and one or more processors configured to obtain media data captured by a data capture device, embed the rendering engine model into the media item comprising the media data, and provide (e.g., store, send, output, etc.) the media item having the rendering engine model embedded in the media item to one or more devices. The one or more devices may obtain the media item and store it, or use the rendering engine model in the media item to execute a neural network to process and render the media data in the media item. The neural network may use the media data as input to generate rendered media data output.

[0012] The media data may include image data, video data, audio data, etc. A media item may include a file, a stream, or any other type of data container or object that may contain or encapsulate media data and a rendering engine model. The rendering engine model may include a description of a neural network configured to process the media data and generate a specific media data output. The rendering engine model in the media item may indicate how to execute the neural network based on the description of the neural network to process the media data and generate a specific media data output. For example, the rendering engine model in the media item may inform one or more devices and / or any other devices that have a copy of the media item how to execute the neural network to process the media data and generate a specific media data output.

[0013] The description of a neural network may define a neural network architecture of the neural network. The neural network architecture may include, for example, a neural network structure (e.g., number of layers, number of nodes in each layer, layer interconnections, etc.), neural network filters or operations, activation functions, parameters (e.g., weights, biases, etc.), etc. The description may also define how the layers in the neural network are interconnected, how inputs to the neural network are formed, and how outputs are formed from the neural network. In addition, the description may define one or more tasks of the neural network, such as one or more custom tasks for encoding media data, decoding media data, compressing or decompressing media data, performing image processing operations (e.g., image restoration, image enhancement, demosaicing, filtering, scaling, color correction, color conversion, noise reduction, spatial filtering, image rendering, etc.), performing frame rate conversion (e.g., upconversion, downconversion), performing audio signal modification operations (e.g., generating a wideband audio signal from a narrowband audio input file), etc.

[0014] In another example, a non-transitory computer-readable medium is provided for creating a media item comprising media data (e.g., original or captured media data) and an embedded rendering engine model for rendering the media data. The non-transitory computer-readable medium may store instructions that, when executed by one or more processors, cause the one or more processors to obtain media data captured by a data capture device, embed a rendering engine model into a media item comprising the media data, and provide (e.g., store, send, output, etc.) the media item having the rendering engine model embedded in the media item to one or more devices. The one or more devices may obtain the media item and store it, or use the rendering engine model in the media item to execute a neural network to process and render the media data in the media item. The neural network may use the media data as input to generate rendered media data output.

[0015] In another example, an apparatus is provided for creating a media item that includes media data (e.g., original or captured media data) and an embedded rendering engine model for rendering the media data. The example apparatus may include components for obtaining media data captured by a data capture device, components for embedding the rendering engine model into the media item that includes the media data, and components for providing (e.g., storing, sending, outputting, etc.) the media item with the rendering engine model embedded in the media item to one or more devices. The one or more devices may obtain the media item and store it, or use the rendering engine model in the media item to execute a neural network to process and render the media data in the media item. The neural network may use the media data as input to generate rendered media data output.

[0016] The media data may include image data, video data, audio data, etc. A media item may include a file, a stream, or any other type of data container or object that may contain or encapsulate media data and a rendering engine model. The rendering engine model may include a description of a neural network configured to process the media data and generate a specific media data output. The rendering engine model in the media item may indicate how to execute the neural network based on the description of the neural network to process the media data and generate a specific media data output. For example, the rendering engine model in the media item may inform one or more devices and / or any other devices that have a copy of the media item how to execute the neural network to process the media data and generate a specific media data output.

[0017] The description of a neural network may define a neural network architecture of the neural network. The neural network architecture may include, for example, a neural network structure (e.g., number of layers, number of nodes in each layer, layer interconnections, etc.), neural network filters or operations, activation functions, parameters (e.g., weights, biases, etc.), etc. The description may also define how the layers in the neural network are interconnected, how inputs to the neural network are formed, and how outputs are formed from the neural network. In addition, the description may define one or more tasks of the neural network, such as one or more custom tasks for encoding media data, decoding media data, compressing or decompressing media data, performing image processing operations (e.g., image restoration, image enhancement, demosaicing, filtering, scaling, color correction, color conversion, noise reduction, spatial filtering, image rendering, etc.), performing frame rate conversion (e.g., upconversion, downconversion), performing audio signal modification operations (e.g., generating a wideband audio signal from a narrowband audio input file), etc.

[0018] In some aspects, the above methods, apparatus, and computer-readable media may further include embedding multiple rendering engine models in a media item. For example, the methods, apparatus, and computer-readable media may include embedding additional rendering engine models in a media item. The additional rendering engine models may include additional descriptions of additional neural networks configured to process media data and generate different media data outputs. The additional descriptions may define different neural network architectures for the additional neural networks. Based on different neural network layers, filters or operations, activation functions, parameters, etc., different neural network architectures may be customized for different operational outcomes or rendering intents.

[0019] The media items with the rendering engine model and the additional rendering engine model can be provided to one or more devices for storing the media items, processing the media items and / or rendering the media data in the media items. In some examples, the media items can be provided to one or more devices for storage. In other examples, the media items can be provided to one or more devices for processing and / or rendering the media data in the media items. For example, one or more devices can receive the media items and select one of a plurality of rendering engine models (e.g., a rendering engine model or an additional rendering engine model), and based on the selected rendering engine model, run the corresponding neural network associated with the selected rendering engine model. The one or more devices can then use the corresponding neural network to process the media data in the media items to obtain or generate media data output from the corresponding neural network.

[0020] In some aspects, the above methods, apparatus, and computer-readable media may further include generating a test neural network configured to process and render raw media data, and training the test neural network based on the media data sample. The test neural network may include a test neural network architecture, which may include a particular neural network structure (e.g., layers, nodes, interconnections, etc.), test filters or operations, test activation functions, test parameters (e.g., weights, biases, etc.), etc. Training the test neural network may include processing the media data sample using the test neural network, determining performance of the test neural network based on one or more outputs associated with the media data sample, determining one or more adjustments to the test neural network (and / or test neural network architecture) based on the performance of the test neural network, and adjusting the test neural network (e.g., test neural architecture, test parameters, test filters or operations, test activation functions, layers in the test neural network, etc.) based on the performance of the test neural network.

[0021] In some cases, determining the performance of the test neural network can include determining the accuracy of the test neural network and / or testing the loss or error in one or more outputs from the neural network. For example, determining the performance of the test neural network can include applying a loss function (e.g., a mean squared error (MSE) function) to one or more outputs to generate feedback, which can include loss or error calculations or results. The feedback can be used to identify and make adjustments to tune the test neural network.

[0022] In some cases, the training and one or more adjustments can be used to determine a neural network architecture associated with a rendering engine model in a media item. For example, a test neural network architecture and one or more adjustments to the test neural network architecture determined through training can be used to determine a specific neural network architecture and configuration that can be used as the basis for a rendering engine model embedded in a media item.

[0023] In some examples, the above methods, apparatus, and computer-readable media may include embedding an address (e.g., a uniform resource identifier (URI); a path; a network; a storage device or destination address; a link; a resource locator; etc.) in a media item to a remote rendering engine model or a remote location of the remote rendering engine model. The remote rendering engine model may include a corresponding description of a neural network configured to process media data and generate corresponding media data output. The media item with the address may be provided to one or more devices, which may use the address to retrieve the remote rendering engine model from the remote location and, based on the corresponding description in the remote rendering engine model, generate a neural network associated with the remote rendering engine model and use the neural network to process the media data in the media item to generate corresponding media data output (e.g., a rendering of the media data).

[0024] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it exhaustive or intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the entire specification, drawings, and appropriate portions of the claims of the present disclosure.

[0025] The foregoing and other features and embodiments will become more fully apparent with reference to the following description, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Illustrative embodiments of the present application are described in detail below with reference to the following drawings:

[0027] Figure 1 is a block diagram illustrating an example environment including a media processing system according to some examples;

[0028] Figure 2A and Figure 2BAn example process for generating a media item with an embedded rendering engine model according to some examples is shown;

[0029] Figure 3A and Figure 3B An example flow for processing a media item with an embedded rendering engine model and generating rendered output according to some examples is shown;

[0030] Figure 4 An example rendering engine model and an example neural network architecture defined by the rendering engine model according to some examples are shown;

[0031] Figure 5 illustrates an example use of a neural network defined by a rendering engine model for processing image data in a media item according to some examples;

[0032] Figure 6 An example implementation of a media item having an embedded address to a remote rendering engine model for media data in the media item according to some examples is shown;

[0033] Figure 7 An example process for training a neural network to identify an optimized configuration of the neural network for describing a rendering engine model of the neural network according to some examples is shown;

[0034] Figure 8 An example method for providing a rendering engine model with a media item according to some examples is shown; and

[0035] Figure 9 An example computing device architecture according to some examples is shown. DETAILED DESCRIPTION

[0036] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some can be applied in combination, which will be apparent to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that the various embodiments can be implemented without these specific details. The drawings and descriptions are not restrictive.

[0037] The following description provides only exemplary embodiments and features and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0038] Specific details are provided in the following description to provide a thorough understanding of the embodiments. However, one of ordinary skill in the art will appreciate that the embodiments may be practiced without these specific details. For example, circuits, devices, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other cases, known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0039] Furthermore, it is noted that embodiments may be described as processes depicted as flow charts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flow chart may describe operations as a sequential process, many operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. When the operations of a process are completed, the process is terminated, but there may be additional steps not included in the diagram. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0040] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying (multiple) instructions and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored, and the non-transitory medium does not include a carrier wave and / or a temporary electronic signal that is transmitted wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, a disk or tape, an optical storage medium such as a compact disc (CD) or a digital versatile disc (DVD), flash memory, a memory, or a memory device. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent any combination of a procedure, function, subroutine, program, routine, subroutine, module, software package, class, or instruction, data structure, or program statement. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be passed, forwarded, or sent via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0041] Furthermore, features and embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments that perform specific tasks (e.g., a computer program product) may be stored in a computer-readable or machine-readable medium. Processor(s) may perform specific tasks.

[0042] The disclosed technology provides a system, method, and computer-readable storage medium for generating a media item comprising raw media data and a rendering engine model, for processing the raw media data without a separate media rendering tool (e.g., a decoder, an encoder, an image processor, a neural network processor, and / or other media rendering tools). The raw media data may include media data captured by a data capture device or sensor (e.g., an image or video sensor, an audio sensor, etc.), for example, raw, ordinary, uncompressed, and / or unprocessed data (e.g., video data, image data, audio data, etc.) captured and / or output by the data capture device or sensor.

[0043] In some aspects, the raw media data may include image data from a data capture device or sensor (e.g., before or after processing by one or more components of an image signal processor). The image data may be filtered by a color filter array. In some examples, the color filter array includes a Bayer color filter array. In some aspects, the raw media data may include a patch of raw image data. A patch of raw image data may include a subset of a raw image data frame captured by the data capture device or sensor. In some aspects, the raw media data may include raw audio data. The raw audio data may be included in a raw audio file, which may contain, for example, uncompressed mono pulse codec modulated data.

[0044] In some cases, raw media data comprising raw video and / or image data may include a plurality of pixels. For example, the raw media data may include a multidimensional digital array representing raw image pixels of an image associated with the raw media data. In one example, the array may include a 128×128×11 digital array having 128 rows and 128 columns of pixel locations and 11 input values ​​per pixel location. In another illustrative example, the raw media data may include a raw image patch comprising a 128×128 array of raw image pixels.

[0045] In addition, the raw media data containing the raw image or video data may include one or more color components, or color component values ​​for each pixel. For example, in some cases, the raw media data may include a color or grayscale value for each pixel location. A color filter array may be integrated with a data capture device or sensor, or may be used in conjunction with a data capture device or sensor (e.g., placed on an associated photodiode) to convert monochrome information into color values. For example, a sensor having a color filter array (e.g., a Bayer pattern color filter array having a red, green, or blue filter at each pixel location) may be used to capture raw image data having a color at each pixel location.

[0046] In some examples, one or more devices, methods, and computer-readable media are described for generating a media item with an embedded rendering engine model. The device, method, and computer-readable medium can be implemented to create a media item, such as a media file or stream, that contains original or captured media data (e.g., image data, video data, audio data, etc.) and a complete specification of a rendering engine for the original media data. The rendering engine model can be embedded in the media item along with the original media data. Regardless of the type or format of the media item, or the requirements of the original media data in the media item, the rendering engine model can be used to process and render the original media data without the need for a separate media rendering tool.

[0047] Computing devices can use the rendering engine model in the media items to run the rendering engine defined for media data, and render the media data, without the need for a separate media rendering tool compatible with the media items to be rendered. By including a rendering engine model or specification in the media items, any suitable processor on the device can use the rendering engine model in the media items to render the original media data. Therefore, the media items and methods herein can eliminate or significantly reduce the demand for media item standardization, support new solutions to media data, and provide customized processing for different types of media data. The rendering engine model herein can be customized to process the original media data in the media items according to specific rendering intentions and achievements.

[0048] For example, a compressed audio file may include a rendering engine model that defines how to process or recover the audio stream in the audio file, while a video file may include a rendering engine model that defines how to decompress the video stream from the video file. In other examples, an image file may include a rendering engine model for rendering raw image data into an RGB (red, green, and blue) viewable image at 8x scaling, a narrowband audio file may include a rendering engine model for generating a wideband signal, and a video file may include a rendering engine model for 2x frame rate upconversion.

[0049] In some implementations, a media item can include multiple rendering engine models. Different rendering engine models can be tailored for different rendering intents and processing outcomes. For example, a media item can include raw media data, a first rendering engine model for a specific outcome (e.g., fast service with a quality tradeoff), and a second rendering engine model for a different outcome (e.g., higher quality with a speed tradeoff). Multiple rendering engine models can conveniently provide different rendering and processing capabilities and provide users with additional control over rendering and processing outcomes.

[0050] In some cases, a rendering engine model may include a description of a neural network configured to process a media item and generate a specific media data output. The neural network may function as a rendering engine for the media item. The description of the neural network may define the neural network architecture and specific tuning, as well as other parameters tailored for a specific rendering intent or desired result. For example, the description may define: the building blocks of the neural network, such as operations (e.g., 2d convolution, 1d convolution, pooling, normalization, fully connected, etc.) and activation functions (e.g., rectified linear units, exponential linear units, etc.); parameters of the neural network operations (e.g., weights or biases); how these building blocks are interconnected; how the input to the neural network is formed from media data (e.g., a 128×128 patch of pixel data from an input image); and how the output is formed from the neural network (e.g., outputting a 64x 64x 3 patch and tiling the output patch together to produce the final output).

[0051] Neural networks can be customized for various operations or rendering intents, such as data encoding and / or decoding, image / video processing, data compression and / or decompression, scaling, filtering, image restoration (e.g., noise reduction), image enhancement or modification, and so on. Neural networks can also be customized for specific rendering or processing outcomes, such as speed, quality, output accuracy, output size (e.g., image size), enhanced user experience, and so on. A neural network description can fully specify how to use the neural network to process a piece of data. For example, a raw image file can include a description of a neural network that models a camera ISP (Image Signal Processor) using the raw image data, allowing the neural network to render the result exactly as intended.

[0052] A neural network description can be implemented with a media item such as a file or stream of data. For example, a camera sensor can transmit a stream of raw data to a processor. A rendering engine model of the data can be sent by the camera sensor to the processor along with the stream. The rendering engine model can include a description of a neural network architecture with parameters for rendering a visual image from the raw sensor data. The processor can use the rendering engine model to execute the neural network and render the raw data accordingly. As another example, a compressed audio stream including the rendering engine model can be sent to the processor. The rendering engine model can be customized for decompressing audio and / or any other processing output. The rendering engine model can include a description of the neural network architecture and parameters for decompressing audio, which the processor can use to execute the neural network and decompress the audio.

[0053] The disclosed technology is described in more detail in the following disclosure. The discussion begins with a description of example systems, architectures, and methods, such as Figures 1 to 7As shown, these example systems, architectures, and methods are used to create a media item with an embedded rendering engine model, customize the rendering engine model in the media item for specific rendering intent and processing efforts, and implement a neural network using the rendering engine model. This will then be followed by a description of an example method for creating a media item with an embedded rendering engine model (e.g., Figure 8 The discussion begins with a description of an example computing device architecture (e.g. Figure 9 The present disclosure now turns to Figure 1 .

[0054] Figure 1 1 is a diagram illustrating an example computing environment 100 including a media processing system 102 and remote systems 120, 130, 140. The media processing system 102 can obtain, store, and / or generate rendering engine models and / or create media items with embedded rendering engine models as described herein. As used herein, the term "media item" can include a file, a stream or a bitstream, or any other data object or container capable of storing, containing, or encapsulating media data, a rendering engine model, and any other data.

[0055] In this illustrative example, media processing system 102 includes a computing component 104, a storage device 108, a computing engine 110, and data capture devices 112, 114. Media processing system 102 can be a computing device or a portion of multiple computing devices. In some examples, media processing system 102 can be a portion of an electronic device (device), such as a camera system (e.g., a digital camera, an IP camera, a video camera, a security camera, etc.), a phone system (e.g., a smartphone, a cellular phone, a conferencing system, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a digital media player, a game console, a streaming device, a drone, an in-vehicle computer, an Internet of Things (IoT) device, a server, a distributed system, or any other suitable electronic device(s).

[0056] In some embodiments, computing component 104, storage 108, computing engine 110, and data capture devices 112, 114 may be part of the same computing device. For example, in some cases, computing component 104, storage 108, computing engine 110, and data capture devices 112, 114 may be integrated into a smartphone, laptop computer, tablet computer, smart wearable device, gaming system, and / or any other computing device. In other embodiments, computing component 104, storage 108, computing engine 110, and data capture devices 112, 114 may be part of two or more independent computing devices. For example, computing component 104, storage 108, and computing engine 110 may be part of one computing device (e.g., a smartphone or laptop computer), and data capture devices 112, 114 may be part of or represent one or more independent computing devices (e.g., one or more independent cameras or computers).

[0057] Storage 108 may include any physical and / or logical storage device(s) for storing data. Furthermore, storage 108 may store data from any component of media processing system 102. For example, storage 108 may store data from data capture devices 112, 114 (e.g., image, video, and / or audio data), data from compute component 104 (e.g., processing parameters, calculations, processing output, data associated with compute engine 110, etc.), one or more rendering engine models, and the like. Storage 108 may also store data received by media processing system 102 from other devices, such as remote systems 120, 130, 140. For example, storage 108 may store media data captured by data capture devices 152, 154 on remote systems 120, 130, rendering engine models or parameters from remote system 140, and the like.

[0058] Data capture devices 112, 114 may include any sensor or device for capturing or recording media data (e.g., audio, image, and / or video data), such as a digital camera sensor, a video camera sensor, a smartphone camera sensor, an image / video capture device on an electronic device (e.g., a television or computer), a camera, a microphone, etc.). In some examples, data capture device 112 may be an image / video capture device (e.g., a camera, a video and / or image sensor, etc.) and data capture device 114 may be an audio capture device (e.g., a microphone).

[0059] Furthermore, the data capture devices 112, 114 can be standalone devices or part of one or more standalone computing devices, such as digital cameras, video cameras, IP cameras, smartphones, smart TVs, gaming systems, IoT devices, laptops, etc. In some examples, the data capture devices 112, 114 can include computing engines 116, 118 for processing captured data, generating rendering engine models for the captured data, creating media items containing the captured data and the rendering engine models, and / or performing any other media processing operations. In some cases, the data capture devices 112, 114 can capture media data and use the computing engines 116, 118 to locally create a media item containing the captured media data and a rendering engine model for processing and rendering the captured data as described herein. In other cases, the data capture devices 112, 114 may capture media data and send the captured media data to other devices or computing engines, such as computing component 104 and computing engine 110, for packaging the captured media data in a media item that includes at least one rendering engine model for the captured media data.

[0060] The computer component 104 may include one or more processors 106A-N (hereinafter collectively referred to as "106"), such as a central processing unit (CPU), a graphics processing unit (GPU) 114, a digital signal processor (DSP), an image signal processor (ISP), etc. The processor 106 may perform various operations, such as image enhancement, graphics rendering, augmented reality, image / video processing, sensor data processing, recognition (e.g., text recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, etc.), image stabilization, machine learning, filtering, data processing, and any of the various operations described herein. In some cases, the computing component 104 may also include other electronic circuitry or hardware, computer software, firmware, or any combination thereof to perform any of the various operations described herein.

[0061] The computing component 104 may implement a computing engine 110. The computing engine 110 may be implemented by one or more processors 106 from the computing component 104. The computing engine 110 may process media data, render media data, create a rendering engine model for processing and rendering raw or captured media data as described herein, create a media item containing raw or captured media data and a rendering engine model for the raw or captured media data, perform machine learning operations (e.g., create, configure, execute, and / or train a neural network or other machine learning system), and / or perform any other media processing and computing operations.

[0062] The computing engine 110 may include one or more media processing engines, such as a rendering engine, a front-end processing engine, an image processing engine, a digital signal processing engine, etc. The one or more media processing engines may perform various media processing operations, such as filtering, demosaicing, scaling, color correction, color conversion, noise reduction, spatial filtering, scaling, frame rate conversion, audio signal processing, noise control or removal, image enhancement, data compression and / or decompression, data encoding and / or decoding, etc. In some examples, the computing engine 110 may include multiple processing engines, which may be configured to perform the same or different computing operations.

[0063] In some cases, the computing engine 110 may receive media data (e.g., image data, video data, audio data, etc.) captured by any of the data capture devices 102, 104, 152, 154 in the computing environment 100, and receive, generate, and / or retrieve from storage one or more rendering engine models for the media data. The computing engine 110 may embed the one or more rendering engine models into a media item containing the media data, and / or embed an address (e.g., a uniform resource identifier (URI); a link; a path; a network, storage, or destination address; a resource locator; etc.) into the one or more rendering engine models in the media item containing the media data.

[0064] Remote systems 120, 130, 140 may represent client devices, such as smartphones or laptops, cloud computing environments or services, servers, IoT devices, smart devices, or any other network, device, or infrastructure. Figure 1 In the example shown, remote systems 120 and 130 represent client devices, and remote system 140 represents a server or cloud computing environment. In this example, remote systems 120 and 130 may include data capture devices 152 and 154 for capturing or recording media data such as video, images, and / or audio, and remote system 140 may include storage 156, which may serve as a repository for media data, rendering engine models, parameters, and / or other data. Storage 156 may include one or more physical and / or logical storage devices. For example, storage 156 may represent a distributed storage system.

[0065] The remote systems 120, 130, and 140 may also include a computing engine 150. The computing engine 150 may include, for example, but not limited to, an image processing engine, a digital signal processing engine, a rendering engine, a front-end processing engine, and / or any other processing or media engine. The computing engine 150 may perform various operations, such as filtering, demosaicing, scaling, color correction, color conversion, noise reduction, spatial filtering, scaling, frame rate conversion, audio signal processing, noise control or removal, image enhancement, data compression and / or decompression, data encoding and / or decoding, machine learning, and the like. In addition, the computing engine 150 may run or generate a rendering engine specified by a rendering engine model associated with a particular media item, use the rendering engine to process and render media data in the media item, generate a rendering engine model, and the like.

[0066] The media processing system 102 can communicate with remote systems 120, 130, 140 through one or more networks (e.g., a private network (e.g., a local area network, a virtual private network, a virtual private cloud, etc.), a public network (e.g., the Internet), etc.). The media processing system 102 can communicate with the remote systems 120, 130, 140 to send or receive media data (e.g., raw or captured media data, such as video, image, and / or audio data), send or receive rendering engine models, send or receive media items with embedded rendering engine models, store or retrieve rendering engine models for media items, and the like.

[0067] Although the media processing system 102, data capture devices 112, 114, and remote systems 120, 130, 140 are shown as including certain components, one of ordinary skill will understand that the media processing system 102, data capture devices 112, 114, and / or remote systems 120, 130, 140 may include more than Figure 1 For example, in some cases, the media processing system 102, data capture devices 112, 114, and remote systems 120, 130, 140 may also include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more network interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or Figure 1 Other hardware or processing devices not shown. Figure 9 Illustrative examples of computing devices and hardware components that may be implemented with media processing system 102, data capture devices 112, 114, and remote systems 120, 130, 140 are described.

[0068] Figure 2AAn example process 200 is shown for generating a media item 220 having one or more embedded rendering engine models 230, 240. In this example, a compute engine 110 on a media processing system 102 can receive media data 210 and generate a media item 220 that includes the media data 210 and one or more rendering engine models 230, 240. The media data 210 can include raw media data captured by a media capture device, such as audio, image, and / or video data captured by a data capture device 112, 114, 152, or 154.

[0069] The computing engine 110 can generate or obtain one or more rendering engine models 230, 240 configured to process and render the media data 210. Each rendering engine model 230, 240 can include a description or specification of a rendering engine configured to process the media data 210 and generate a specific media data output. The rendering engine can be customized for specific processing efforts (e.g., speed, quality, performance, etc.) and / or rendering intents (e.g., size, format, playback or rendering quality and characteristics, output configuration, etc.). The specific media data output generated by the rendering engine can therefore be based on such processing efforts and / or rendering intents.

[0070] In some examples, the rendering engine can be implemented by a neural network. Here, the rendering engine model (230 or 240) for the neural network can include a description or specification of the neural network, which can describe how to generate, configure, and execute the neural network. For example, the description or specification of the neural network can define the architecture of the neural network, such as the number of input nodes, the number and type of hidden layers, the number of nodes in each hidden layer, the number of output nodes, the filters or operations implemented by the neural network, the activation functions in the neural network, the parameters of the neural network (such as weights and biases, etc.). The description or specification of the neural network can further define how the layers in the neural network are connected to form paths of interconnected layers, how the input to the neural network is formed based on the media data 210, and how the output is formed from the neural network.

[0071] The description or specification of the neural network may also include any other information or instructions for configuring and / or executing the neural network to process the media data 210 for a particular processing outcome or rendering intent. For example, the description or specification of the neural network may define one or more customized tasks for the neural network, such as encoding the media data 210, decoding the media data 210, performing one or more compression or decompression operations on the media data 210, performing one or more image processing operations (e.g., image restoration, image enhancement, filtering, scaling, image rendering, demosaicing, color correction, resizing, etc.) on the media data 210, performing a frame rate conversion operation on the media data 210, performing an audio signal modification operation on the media data 210, etc.

[0072] After generating or obtaining the one or more rendering engine models 230, 240, the computing engine 110 can use the media data 210 and the one or more rendering engine models 230, 240 to generate a media item 220. For example, the computing engine 110 can create a file, a stream, or a data container and include or embed the media data 210 and the one or more rendering engine models 230, 240 in the file, stream, or data container. The media item 220 can represent the resulting file, stream, or data container with the media data 210 and the one or more rendering engine models 230, 240.

[0073] In some examples, media items 220 may include a single rendering engine model (e.g., 230 or 240) configured to process and render media data 210. As previously mentioned, a single rendering engine model can be customized for specific processing achievements and / or rendering intent. In other examples, media items 220 may include a plurality of rendering engine models (e.g., 230 and 240) configured to process and render media data 210 according to different processing achievements and / or rendering intent. A plurality of rendering engine models may provide extra rendering and process flexibility, control, and options. The equipment or user processing media items 220 may select a specific rendering engine model to be used for rendering the media data 210 in media items 220. Specific rendering engine model may be selected based on the processing achievements and / or rendering intent of expectation.

[0074] Once the media item 220 is generated, it is ready for processing and can be used by the computing engine 110 to render the media data 210, or can be sent to a computing device, such as a remote system (e.g., 120, 130, 140) or an internal processing device (e.g., 106A, 106B, 106N), for processing and rendering.

[0075] Figure 2B An example process 250 is shown for capturing media data 210 using a data capture device 112 and generating a media item 220 using one or more embedded rendering engine models 230, 240. In this example, the data capture device 112 can execute the process 250 and generate both the media data 210 and the media item 220 having the one or more embedded rendering engine models 230, 240. This implementation allows the media item 220 to be generated by the same device (e.g., 112) that captured the media data 210, rather than having one device capture the media data 210 and a separate or external device generate the media item 220 from the media data 210.

[0076] For example, a camera or a device equipped with a camera (e.g., a smartphone, a notebook computer, a smart TV, etc.) can capture media data 210 and generate a media item 220 comprising the media data 210 and one or more rendering engine models 230, 240. At this point, the media item 220 is ready for rendering by any processing device. Therefore, if a user wants to capture media data (e.g., images, videos, etc.) from a camera or a device equipped with a camera, and render the media data on another device, the camera or the device equipped with a camera can capture the media data 210 and prepare the media item 220 so that it is ready for rendering from any other computing device. The user will be able to render the media item 220 (when received from the camera or the device equipped with a camera) from another device without the need to install a separate decoder on the device. By this implementation, the manufacturer of the camera or the device equipped with a camera can ensure that the camera or the device equipped with a camera can capture media data and produce a final output that is ready for rendering by other devices without the need for a separate decoder tool and without (or with limited) compatibility issues.

[0077] Returning to the example process 250, the data capture device 112 can first capture media data 210. The data capture device 112 can provide the media data 210 as input to the computing engine 116, and the computing engine 116 can generate a media item 220 that includes the media data 210 and one or more rendering engine models 230, 240 for rendering the media data 210. The computing engine 116 can add or embed the media data 210 and the one or more rendering engine models 230, 240 into a file, a stream, or a data container or object to create the media item 220.

[0078] At this point, the media item 220 is ready to be rendered by the data capture device 112 or another computing device. For example, if the user wants to render the media data 210 from a separate device, the data capture device 112 can provide the media item 220 to the separate device for rendering. Figure 3A-Figure 3B To further explain, a separate device can receive media item 220 and use a rendering engine model (e.g., rendering engine model 230 or 240) in media item 220 to run a rendering engine described or modeled by the rendering engine model. The rendering engine can be configured to process and render media data 210 according to specific processing efforts and / or rendering intents. The rendering engine on the separate device can then process and render the media data 210 in media item 220 accordingly.

[0079] In example implementations where the media item 220 includes multiple rendering engine models, a user or individual device can select a particular rendering engine model based on, for example, the corresponding processing efforts and / or rendering intents associated with the various rendering engine models (e.g., 230 and 240). The individual device can render the media data 210 using the selected rendering engine model as previously described, and if different processing efforts or rendering intents are desired, a different rendering engine model can be selected and implemented to produce different rendering and / or processing efforts.

[0080] Figure 3A An example process 300 is shown for processing a media item 220 and generating a rendered output 310. In this example, the media item 220 is processed by a computing engine 110 on a media processing system 102. The media item 220 may be received by the computing engine 110 from another device or component (e.g., a data capture device 112), or generated by the computing engine 110 as previously described. Thus, in some cases, the computing engine 110 may both generate and process the media item 220.

[0081] In process 300, computing engine 110 executes a rendering engine based on a rendering engine model (e.g., 230 or 240) in media item 220. The rendering engine model may specify how to create a rendering engine. Computing engine 110 may analyze the selected rendering engine model to determine how to generate or execute the rendering engine. The rendering engine model may identify the structure, parameters, configuration, implementation information, etc. for the rendering engine, which computing engine 110 may use to generate or execute the rendering engine.

[0082] Once the compute engine 110 executes the rendering engine, it may input the media data 210 into the rendering engine, which may then process the media data 210 to produce a rendered output 310. The rendered output 310 may be a rendering of the media data 210 according to the rendering intent or configuration of the rendering engine, and / or the rendering intent or configuration reflected in the rendering engine model.

[0083] In some cases, the rendering engine can be a neural network configured to execute as a rendering engine for media data 210. Here, the rendering engine model can include a description or specification of the neural network. The description or specification of the neural network can specify the building blocks and architecture of the neural network, as well as any other information that describes how to generate or execute the neural network, such as neural network parameters (e.g., weights, biases, etc.), operations or filters in the neural network, the number of layers in the neural network and the type of layers, the number of nodes in each layer, how the layers are interconnected, how inputs are formed or processed, how outputs are formed, etc. The computing engine 110 can use the description or specification to execute the neural network defined by the description or specification. The neural network can process the media data 210 and generate a rendered output 310.

[0084] Figure 3B Another example process 350 is shown for processing a media item 220 and generating a rendered output 310. In this example, the media item 220 is processed and rendered by a remote system 130. The media processing system 102 may generate the media item 220 and send it to the remote system 130. The remote system 130 may receive the media item 220 and process it using the computing engine 150 on the remote system 130.

[0085] Computing engine 150 can use a rendering engine model (e.g., 230 or 240) in media item 220 to generate a rendering engine for media data 210. Computing engine 150 can analyze the rendering engine model to identify parameters, configuration information, instructions, etc. for the rendering engine and generate or execute the rendering engine accordingly. For example, as previously described, the rendering engine model can include a description or specification of a rendering engine (e.g., a neural network-based rendering engine), and computing engine 150 can use the description or specification to execute the rendering engine.

[0086] Once the compute engine 150 executes the rendering engine, the rendering engine can process the media data 210 in the media item 220 and generate a rendering output 310 (e.g., a rendering of the media data 210). The rendering engine can process and render the media data 210 according to the processing efforts and / or rendering intents that the rendering engine is configured to implement (e.g., as reflected in the rendering engine model). The remote system 130 can thus receive a media item 220 containing raw or captured media data (210), process the media item 220, and render the raw or captured media data in the media item 220 based on the rendering engine model (e.g., 230 or 240) in the media item 220 without using a separate decoder or media rendering software.

[0087] Figure 4An example architecture 400 of a neural network 410, defined by an example neural network description 402 in a rendering engine model 230, is shown. Neural network 410 may represent a neural network implementation of a rendering engine for rendering media data. Neural network description 402 may include a complete specification of neural network 410 (including neural network architecture 400). For example, neural network description 402 may include a description or specification of architecture 400 of neural network 410 (e.g., layers, layer interconnections, number of nodes in each layer, etc.); input and output descriptions indicating how inputs and outputs are formed or processed; indications of activation functions in the neural network, operations or filters in the neural network, etc.; neural network parameters, such as weights, biases, etc.; and the like.

[0088] Neural network 410 reflects architecture 400 defined in neural network description 402. In this example, neural network 410 includes an input layer 402 that includes input data, such as media data (e.g., 210). In one illustrative example, input layer 402 can include data representing a portion of the input media data (e.g., 210), such as a patch of data or pixels in an image corresponding to the input media data (e.g., a 128×128 data patch).

[0089] Neural network 410 includes hidden layers 404A through 404N (hereinafter collectively referred to as "404"). Hidden layers 404 may include n number of hidden layers, where n is an integer greater than or equal to 1. The number of hidden layers may include as many layers as required for the desired processing outcome and / or rendering intent. Neural network 410 also includes an output layer 406 that provides output (e.g., rendered output 310) resulting from the processing performed by hidden layers 404. In one illustrative example, output layer 406 may provide a rendering of input media data (e.g., 210). In some cases, output layer 406 may generate an output patch (e.g., a 64×64×3 patch) for each patch of input data (e.g., the 128×128 data patch in the previous example), and tile or aggregate each output patch to generate a final output that provides a rendering of the input media data.

[0090] The neural network 410 in this example is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. The information associated with the node is shared between different layers, and each layer retains the information as it is processed. In some cases, the neural network 410 may include a feedforward neural network, in which case there are no feedback connections (in which the output of the neural network is fed back to itself). In other cases, the neural network 410 may include a recursive neural network, which may have loops that allow information to be carried across nodes when reading in input.

[0091] Can exchange information between nodes by the interconnection of the node to the node between various layers.The node of input layer 402 can activate a group of nodes in the first hidden layer 404A.For example, as shown in the figure, each input node of input layer 402 is connected to each node of the first hidden layer 404A.The node of hidden layer 404A can convert the information of each input node by applying activation function to information.The information derived from conversion can then be passed to the node of next hidden layer (for example, 404B) and activate these nodes, and these nodes can perform the function that they specify themselves.Example function includes convolution, up sampling, data conversion, pooling and / or any other suitable function.Then the output of hidden layer (for example, 404B) can activate the node of next hidden layer (for example, 404N), etc.The output of last hidden layer can activate one or more nodes of output layer 406, and output is provided at this moment. In some cases, although nodes in neural network 410 (e.g., nodes 408A, 408B, 408C) are shown as having multiple output lines, a node has a single output, and all lines shown as output from a node represent the same output value.

[0092] In some cases, each node or the interconnection between nodes can have a weight, which is a set of parameters derived from training neural network 410. For example, the interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnections can have digital weights that can be tuned (e.g., based on a training data set), allowing neural network 410 to adapt to the input and learn as more data is processed.

[0093] The neural network 410 can be pre-trained to process features of the data from the input layer 402 using different hidden layers 404 to provide an output via the output layer 406. In an example where the neural network 410 is used to render an image, training data including example images can be used to train the neural network 410. For example, a training image can be input to the neural network 410, which can be processed by the neural network 410 to generate an output that can be used to tune one or more aspects of the neural network 410, such as weights, biases, etc.

[0094] In some cases, neural network 410 can use a training process called backpropagation to adjust the weights of the nodes. Backpropagation can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. For each set of training media data, this process can be repeated a certain number of times until the weights of the layer are accurately tuned.

[0095] For the example of rendering an image, a forward pass may include passing a training image through the neural network 410. Before the neural network 410 is trained, the weights may be initially randomized. The image may include, for example, an array of numbers representing pixels of the image. Each number in the array may include a value from 0 to 255 that describes the intensity of the pixel at that position in the array. In one example, the array may include a 28×28×3 array of numbers having 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, etc.).

[0096] For the first training iteration of the neural network 410, due to the random selection of weights at initialization, the output may include values ​​that are not biased towards any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each different class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). With the initial weights, the neural network 410 is unable to determine low-level features and, therefore, is unable to accurately determine what the class of the object may be. A loss function may be used to analyze the error in the output. Any suitable loss function definition may be used.

[0097] Since the actual values ​​will be different from the predicted outputs, the loss (or error) may be high for the first training data set (e.g., images). The goal of training is to minimize the amount of loss so that the predicted outputs are consistent with the target or ideal outputs. The neural network 410 can perform a backward pass by determining which inputs (weights) contribute most to the loss of the neural network 410, and can adjust the weights so that the loss is reduced and ultimately minimized.

[0098] The derivative of the loss with respect to the weights can be calculated to determine the weights that contribute most to the loss of the neural network 410. After calculating the derivative, a weight update can be performed by updating the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. The learning rate can be set to any suitable value, with higher learning rates including larger weight updates and lower values ​​indicating smaller weight updates.

[0099] Neural network 410 can include any suitable neural or deep learning network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer with multiple hidden layers between the input layer and the output layer. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. In other examples, neural network 410 can represent any other neural or deep learning network, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.

[0100] Figure 5An example use of a neural network 410 defined by a rendering engine model 230 for processing image data (eg, 210) in a media item (eg, 220) is shown.

[0101] In this example, neural network 410 includes an input layer 402, a convolutional hidden layer 404A, a pooling hidden layer 404B, a fully connected layer 404C, and an output layer 406. Neural network 410 can render input image data to generate a rendered image (e.g., rendered output 310). First, each pixel or pixel patch in the image data is treated as a neuron with learnable weights and biases. Each neuron receives some input, performs a dot product, and optionally then applies a nonlinear function. Neural network 410 can also encode certain properties into the architecture by expressing a differentiable score function from the raw image data (e.g., pixels) on one end to a class score on the other end, and process features from the target image. After rendering a portion of the image, neural network 410 can generate an average score (or z-score) for each rendered portion and average the scores within a user-defined buffer.

[0102] In some examples, input layer 404A includes raw or captured media data (e.g., 210). For example, the media data may include an array of numbers representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at that position in the array. The image may be passed through convolutional hidden layer 404A, an optional nonlinear activation layer, a pooling hidden layer 404B, and a fully connected hidden layer 406 to obtain output 310 at output layer 406. Output 310 may be a rendering of the image.

[0103] Convolutional hidden layer 404A can analyze the data from input layer 402A. Each node in convolutional hidden layer 404A can be connected to a region of nodes (e.g., pixels) in the input data (e.g., an image). Convolutional hidden layer 404A can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolution iteration of a filter is a node or neuron in convolutional hidden layer 404A. Each connection between a node and the receptive field of that node (the region of the node (e.g., pixel)) learns a weight, and in some cases, an overall bias, so that each node learns to analyze its specific local receptive field in the input image.

[0104] The convolutional nature of the convolutional hidden layer 404A is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filter of the convolutional hidden layer 404A can start from the upper left corner of the input image array and can be convolved around the input data (e.g., image). As described above, each convolution iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 404A. In each convolution iteration, the value of the filter is multiplied by the original pixel values ​​of the corresponding number of images. The products from each convolution iteration can be added together to obtain the sum for that iteration or node. Based on the receptive field of the next node in the convolutional hidden layer 404A, the process then continues at the next position in the input data (e.g., image). Processing the filter at each unique position of the input volume produces a number representing the filtering result at that position, resulting in a sum value being determined for each node of the convolutional hidden layer 404A.

[0105] The mapping from the input layer 402 to the convolutional hidden layer 404A can be called an activation map (or feature map). The activation map includes the value of each node representing the filtering result at each position of the input volume. The activation map can include an array of various summed values ​​generated from each iteration of the filter on the input volume. The convolutional hidden layer 404A can include several activation maps representing multiple feature spaces in the data (e.g., an image).

[0106] In some examples, a nonlinear hidden layer can be applied after the convolutional hidden layer 404 A. A nonlinear layer can be used to introduce nonlinearity to a system that is already computing linear operations.

[0107] The pooling hidden layer 404B can be applied after the convolutional hidden layer 404A (and, when used, after the nonlinear hidden layer). The pooling hidden layer 404B is used to simplify the information in the output from the convolutional hidden layer 404A. For example, the pooling hidden layer 404B can take each activation map output from the convolutional hidden layer 404A and generate a compressed activation map (or feature map) using a pooling function. Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 404B can use other forms of pooling functions, such as average pooling or other suitable pooling functions.

[0108] A pooling function (e.g., a max pooling filter) is applied to each activation map included in the convolutional hidden layer 404A. Figure 5 In the example shown, three pooling filters are used for the three activation maps in the convolutional hidden layer 404A. Pooling functions (e.g., max pooling) can reduce, aggregate, or concatenate feature representations in the output or input (e.g., an image). Max pooling (and other pooling methods) provides the benefit of pooling fewer features, thereby reducing the number of parameters required for subsequent layers.

[0109] The fully connected layer 404C can connect each node from the pooling hidden layer 404B to each output node in the output layer 406. The fully connected layer 404C can obtain the output of the previous pooling layer 404B (which can represent the activation map of the high-level features) and determine the features or feature representations that provide the best representation of the data. For example, the fully connected layer 404C layer can determine the high-level features that provide the best or closest representation of the data and can include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 404C and the pooling hidden layer 404B can be calculated to obtain the probabilities of different features.

[0110] Output from output layer 406 can include a rendering of the input media data (e.g., 310). In some examples, output from output layer 406 can include patches of the output, which are then tiled or combined to produce a final rendering or output (e.g., 310). Other example outputs can also be provided.

[0111] Figure 6 An example implementation of a media item 602 is shown that includes media data 210 and an embedded address 604 of a remotely stored rendering engine model 606. The rendering engine model 606 can be a specific rendering engine model capable of rendering the media data 210. In this example, the rendering engine model 606 is stored in the remote system 140. The remote system 140 can store one or more rendering engine models 606, 230, 240 for the media data 210 (and other media data) that can be accessed by the media processing system 102 (and any other device) to process and render the media data.

[0112] In some cases, the remote system 140 may store multiple rendering engine models configured for different processing efforts, rendering intents, and / or media data. For example, the remote system 140 may store multiple rendering engine models for the media data 210, each rendering engine model being customized for a different processing effort and / or rendering intent. The remote system 140 may compute and / or store many rendering engine models to provide a wide range of customization options. Furthermore, the remote system 140 may be implemented to offload the processing and resource usage for storing and / or computing rendering engine models, provide a greater number of rendering engine model options to the client, reduce the size of the media item (e.g., 602), and / or otherwise reduce the burden (processing and / or resource usage) on the client, and provide a wider range of rendering engine models with increased customization or granularity.

[0113] In some cases, the remote system 140 may have the infrastructure and capabilities, such as storage and / or computing power, to compute and / or maintain a large number of rendering engine models. For example, the remote system 140 may be a server or cloud service that contains a large repository of rendering engine models and has the computing power to generate and store rendering engine models of increasing complexity, customization, size, etc. The remote system 140 may also train rendering engine models and make the trained rendering engine models available as described herein. The remote system 140 may utilize its resources to train and tune rendering engine models and provide highly tuned rendering engine models.

[0114] exist Figure 6 In the illustrative example of , media item 602 includes media data 210 and an address 604 of a rendering engine model 606 on remote system 140. Address 604 may include, for example, a uniform resource identifier (URI), a link, a path, a resource locator, or any other address (e.g., a network, storage device, or destination address). Address 604 indicates where rendering engine model 606 is located and / or how to retrieve it. When media processing system 102 processes media item 602 (e.g., as Figure 3A ), it may use address 604 to retrieve rendering engine model 606 from remote system 140. Once media processing system 102 has retrieved rendering engine model 606 from remote system 140, media processing system 102 may use rendering engine model 606 to execute the rendering engine described by rendering engine model 606. Media processing system 102 may then use the rendering engine to process and render media data 210.

[0115] In some cases, media item 602 may include multiple addresses for multiple rendering engine models. For example, media item 602 may include the addresses of rendering engine models 606, 230, and 240 on remote system 140. This may provide a wider range of processing and rendering options for media processing system 102. Media processing system 102 (or an associated user) may select or identify a specific rendering engine model corresponding to one of the addresses to use in processing media data 210. Media processing system 102 (or an associated user) may select a specific rendering engine model based on, for example, the specific processing effort and / or rendering intent associated with the specific rendering engine model.

[0116] For illustration, media processing system 102 (or associated user) can compare the rendering engine model that is associated with the address and / or is associated with their corresponding processing achievements and / or rendering intention. Then media processing system 102 (or associated user) can select the specific rendering engine model that best matches or serves the processing achievements and / or rendering intention of expectation. Media processing system 102 can use the corresponding address in media item 602 to retrieve selected rendering engine model, and uses selected rendering engine model to execute associated rendering engine and process media data 210 as previously mentioned.

[0117] In some cases, when a media item 602 has multiple addresses and thus provides multiple options, to assist the media processing system 102 (or an associated user) in selecting a rendering engine model, each address on the media item 602 may include information about the corresponding rendering engine model associated with the address. For example, each address may include a description, a unique identifier, and / or other information about the rendering engine model to which it points. The description in the address may include, for example, but not limited to, processing efforts associated with the rendering engine model, rendering intents associated with the rendering engine model, one or more parameters of the rendering engine model, statistics associated with the rendering engine model, a summary of the rendering engine model and / or its specifications, a rating associated with the rendering engine model, suggestions or recommendations about when or how to select or implement a rendering engine model, a list of advantages and disadvantages associated with the rendering engine model, and the like.

[0118] In some cases, an address (e.g., 604) may include such information (e.g., descriptive information) even if it is the only address in media item 602. Media processing system 102 (or an associated user) may use this information to determine whether the rendering engine model associated with the address is suitable for a particular instance or meets the desired processing or rendering results. If media processing system 102 (or an associated user) determines that the rendering engine model associated with the address is inappropriate or undesirable, or if media processing system 102 (or an associated user) wants a different rendering engine model or additional options, media processing system 102 may generate a different rendering engine model as described above, or request a different rendering engine model from remote system 140. Remote system 140 may identify or calculate different rendering engine models for media processing system 102 based on the request. For example, remote system 140 may calculate different rendering engine models based on information provided in the request from media processing system 102. The information in the request may include, for example, an indication of the desired processing results and / or rendering intent, an indication of one or more desired parameters or configuration details of the rendering engine model, etc.

[0119] In some cases, media processing system 102 (or an associated user) can select a rendering engine model in different ways using or without the descriptive information in the address. For example, in order to select between rendering engine models associated with multiple addresses in media item 602, media item 602 can communicate with remote system 140 to request or retrieve information about the rendering engine model, such as a description of the corresponding processing results and / or rendering intent of the rendering engine model, a description of the corresponding specifications of the rendering engine model, statistics associated with the rendering engine model, associated rankings, etc. In another example, media processing system 102 can send a query to remote system 140 containing parameters or attributes (e.g., processing results, rendering intent, configuration parameters, etc.) describing the desired rendering engine model. Remote system 140 can receive the query and use the parameters or attributes to identify, suggest, or calculate the rendering engine model.

[0120] Figure 7 An example process 700 for training a neural network 410 to identify an optimized configuration of the neural network 410 is shown. The optimized configuration derived from the training can be used to create a rendering engine model that describes the neural network 410 having the optimized configuration. The training of the neural network 410 can be performed for various rendering engine model implementation scenarios. For example, the training of the neural network 410 can be used to train and tune the rendering engine and related rendering engine models provided by the remote system 140 to the client, such as Figure 6 As shown in .

[0121] As another example, training of neural network 410 may be used in situations where a data capture device (e.g., 112, 114, 152, 154) is adapted not only to capture media data (e.g., 210), but also to provide the media data with one or more rendering engine models for the media data, such as Figure 2B For illustration, knowing the capabilities of the data capture device 112 and the characteristics of the media data captured by the data capture device 112, the manufacturer of the data capture device 112 can design a rendering engine model suitable for rendering the media data captured by the data capture device 112, and the data capture device 112 can provide the captured media data to the rendering engine model. The rendering engine model can be pre-configured by the manufacturer and can describe a rendering engine (e.g., a neural network) that has been trained and tuned as described in process 700. The rendering engine associated with the rendering engine model can be specifically tailored to the capabilities of the data capture device 112 and / or the characteristics of the media data it generates.

[0122] Returning to process 700, media data samples 702 can be used as training inputs for neural network 410. Media data samples 702 can include n samples of image, video, and / or audio data, where n is an integer greater than or equal to 1. In some examples, media data samples 702 can include raw images or frames (e.g., images and / or video frames) captured by one or more data capture devices.

[0123] Media data samples 702 can be used to train neural network 410 to achieve a specific processing outcome and / or rendering intent. The goal can be to find the best tuning and configuration parameters (e.g., weights, biases, etc.) to achieve a specific processing outcome and / or rendering intent.

[0124] A media data sample 702 is first processed by the neural network 410 (e.g., via the input layer 402, the hidden layer 404, and the output layer 406) based on the existing weights of the nodes 408A-C or the interconnections between the nodes 408A-C in the neural network 410. The neural network 410 then outputs a rendered output 704 generated for the media data sample 702 via the output layer 406. The rendered output 704 from the neural network 410 is provided to a loss function 706, such as a mean squared error (MSE) function or any other loss function, which generates feedback 708 for the neural network 410. The feedback 708 provides a cost or error (e.g., a mean squared error) in the rendered output 704. In some cases, the cost or error is relative to a target or ideal output, such as a target rendered output for the media data sample 702.

[0125] The neural network 410 may adjust / tune the weights of the nodes 408A-C or the interconnections between the nodes 408A-C in the neural network 410 based on the feedback 708. By adjusting / tuning the weights based on the feedback 708, the neural network 410 may be able to reduce errors in rendering the output of the neural network 410 (e.g., 704) and optimize the performance and output of the neural network 410. This process may be repeated for a certain number of iterations for each set of training data (e.g., media data samples 702) until the weights in the neural network 410 are tuned to a desired level.

[0126] Example systems and concepts have been disclosed, such as Figure 8 As shown, the present disclosure now turns to an example method 800 for providing a rendering engine model with a media item. For clarity, the method 800 is referenced Figure 1 The media processing system 102, various components, and Figure 4 、 Figure 5 and Figure 78 is described with reference to a neural network 410 shown in FIG. 8 that is configured to perform the various steps of method 800. The steps outlined herein are examples and may be implemented in any combination thereof, including combinations that exclude, add, or modify certain steps.

[0127] At step 802, the media processing system 102 obtains media data (eg, 210) captured by a data capture device (eg, 112). The media data may include, for example, raw image data, raw video data, raw audio data, metadata, etc.

[0128] At step 804, the media processing system 102 embeds a rendering engine model (e.g., 230) into a media item (e.g., 220) containing media data (e.g., 210). The rendering engine model may include a description (e.g., 402) of a neural network (e.g., 410) configured to process the media data in the media item and generate a specific media data output (e.g., rendered output 310). The rendering engine model may be optimized or customized for the conditions under which the media data was captured, the characteristics of the media data, the computational complexity of decrypting the media data, providing specific processing and / or rendering options or features to the user, etc.

[0129] The description may define a neural network architecture (e.g., 400) of the neural network, such as the structure of the neural network (e.g., the number of layers, the number of nodes in each layer, the interconnection of the layers, etc.), a set of filters or operations implemented by the neural network, activation functions implemented by the neural network, parameters implemented along the paths of interconnected layers in the neural network (e.g., weights, biases, etc.), etc. In some cases, the description of the neural network may define how the layers in the neural network are interconnected, how inputs to the neural network are formed based on media data (e.g., input or data patch size and characteristics, etc.), and how outputs are formed from the neural network (e.g., output size and characteristics, etc.).

[0130] For example, the description may indicate that the input is a 128×128 data patch from media data, and the neural network outputs a 64×64×3 data patch from the media data, and the 64×64×3 output data patch is combined or tiled to obtain the final output. In addition, in some cases, the description of the neural network may define one or more tasks of the neural network, such as one or more custom tasks for encoding media data, decoding media data, compressing or decompressing media data, performing image processing operations (e.g., image restoration, image enhancement, demosaicing, filtering, scaling, color correction, color conversion, noise reduction, spatial filtering, image rendering, etc.), performing frame rate conversion (e.g., upconversion, downconversion), performing audio signal modification operations (e.g., generating a wideband audio signal from a narrowband audio input file), etc.

[0131] At step 806, media processing system 102 may provide (e.g., send, store, output, etc.) a media item (e.g., 220) comprising media data (e.g., 210) and a rendering engine model (e.g., 230) to a recipient (e.g., computing engine 110, remote system 120, remote system 130, storage device, or other recipient). In some cases, the recipient may be a device or component within media processing system 102. In other cases, the recipient may be a separate or external device, such as a server, a storage system, a client device requesting the media item, a client device identified (e.g., via an instruction or signal) by media processing system 102 as the intended recipient of the media item, etc.

[0132] The rendering engine model in the media item may include instructions for executing a neural network based on a description of the neural network to process media data (e.g., 210) in the media item (e.g., 220) and generate a specific media data output (e.g., rendered output 310). The instructions may indicate to a recipient how to execute the neural network based on the description of the neural network and how to generate a specific media output. The recipient may receive the media item and use the rendering engine model in the media item to execute the neural network to process and render the media data in the media item. The neural network may use the media data as input to generate a rendered media data output.

[0133] In some cases, the media processing system 102 may include multiple rendering engine models (e.g., 230, 240) and / or addresses (e.g., 604) in a media item. For example, the media processing system 102 may embed an additional rendering engine model (e.g., 240) in a media item (e.g., 220). The additional rendering engine model may include additional descriptions of additional neural networks configured to process media data (210) and generate different media data outputs. The additional descriptions may define different neural network architectures for the additional neural networks. Different neural network architectures may be customized for different operational outcomes based on different neural network layers, filters, activation functions, parameters, etc. The media processing system 102 may send the media item with the rendering engine model and the additional rendering engine model to a recipient for processing and rendering the media data.

[0134] A recipient may receive a media item and select one of a plurality of rendering engine models (e.g., a rendering engine model or an additional rendering engine model), and based on the selected rendering engine model, generate a corresponding neural network associated with the selected rendering engine model. The recipient may then process media data in the media item using the corresponding neural network to obtain or generate media data output from the corresponding neural network.

[0135] In some cases, method 800 may include generating a test neural network configured to process and render raw media data, and training the test neural network based on the media data sample (e.g., 702). The test neural network may include a test neural network architecture, which may include a particular neural network structure (e.g., layers, nodes, interconnections, etc.), test filters or operations, test activation functions, test parameters (e.g., weights, biases, etc.), etc. Training the test neural network may include processing the media data sample using the test neural network, determining performance of the test neural network based on one or more outputs associated with the media data sample (e.g., 704), determining one or more adjustments to the test neural network (and / or test neural network architecture) based on the performance of the test neural network (e.g., 708), and adjusting the test neural network (e.g., test neural architecture, test parameters, test filters or operations, test activation functions, layers in the test neural network, etc.) based on the performance of the test neural network.

[0136] In some cases, determining the performance of the test neural network can include determining the accuracy of the test neural network and / or the loss or error in one or more outputs from the test neural network. For example, determining the performance of the test neural network can include applying a loss function (e.g., 706) to one or more outputs to generate feedback (e.g., 708), which can include a loss or error calculation. The feedback can be used to identify and make adjustments to tune the test neural network.

[0137] In some cases, the training and one or more adjustments can be used to determine a neural network architecture associated with a rendering engine model (e.g., 230) in a media item (e.g., 220). For example, a test neural network architecture determined through training and one or more adjustments to the test neural network architecture can be used to determine a specific neural network architecture and configuration that can be used as the basis for a rendering engine model included in a media item.

[0138] In some implementations, method 800 may include embedding an address (e.g., 604) to a remote rendering engine model or a remote location of a remote rendering engine model in a media item. The remote rendering engine model may include a corresponding description of a neural network configured to process media data and generate corresponding media data output. The media item with the address may be sent to a recipient, which may use the address to retrieve the remote rendering engine model from the remote location and, based on the corresponding description in the remote rendering engine model, generate a neural network associated with the remote rendering engine model and use the neural network to process the media data in the media item to generate corresponding media data output (e.g., a rendering of the media data).

[0139] In some examples, method 800 may be performed by a computing device or apparatus, such as Figure 9 The computing device shown in or Figure 1 800 . In some cases, a computing device or apparatus may include a processor, microprocessor, microcomputer, or other component of a device configured to perform the steps of method 800 . In some examples, a computing device or apparatus may include a data capture device (e.g., 112 , 114 , 152 , 154 ) configured to capture media data (e.g., audio, image, and / or video data (e.g., video frames)). For example, a computing device may include a mobile device (e.g., a digital camera, an IP camera, a mobile phone or tablet computer including an image capture device, or other type of system with a data capture device) having a data capture device or system. In some examples, the data capture device may be separate from the computing device, in which case the computing device receives the captured media data.

[0140] In some cases, the computing device may include a display for displaying output media data (e.g., rendered images, video, and / or audio). In some cases, the computing device may include a video codec. The computing device may also include a network interface configured to transmit data, such as image, audio, and / or video data. The network interface may be configured to communicate data based on the Internet Protocol (IP) or other suitable network data.

[0141] Method 800 is illustrated as a logical flow diagram, the steps of which represent a sequence of steps or operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., which perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation or requirement, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0142] In addition, method 800 can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors through hardware or a combination thereof. As described above, the code can be stored on a computer-readable or machine-readable storage medium (e.g., in the form of a computer program including multiple instructions executable by one or more processors). The computer-readable or machine-readable storage medium can be non-transitory.

[0143] As described above, a neural network can be used to render media data. Any suitable neural network can be used to render media data. Illustrative examples of neural networks that can be used include convolutional neural networks (CNNs), autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), or any other suitable neural network.

[0144] In some examples, the decoded or rendered data can be output from the output interface to a storage device. Similarly, the decoded or rendered data can be accessed from the storage device via the input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing media data. In another example, the storage device can correspond to a file server or another intermediate storage device that can store the decoded or rendered data. The device can access the stored data from the storage device via streaming or downloading. The file server can be any type of server that can store data and send the data to the destination device. Example file servers include network servers (e.g., websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The device can access the data via any standard data connection, including an internet connection. This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.) suitable for accessing data stored on the server, or a combination of the two. Data transmission from the storage device can be streaming, downloading, or a combination thereof.

[0145] The techniques of the present disclosure can be applied to any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (e.g., Dynamic Adaptive Streaming over HTTP (DASH)), digital video on a data storage medium, decoding of digital media stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0146] In the foregoing description, various aspects of the present application have been described with reference to specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative embodiments of the present application have been described in detail herein, it should be understood that the concepts of the present invention may be implemented and used differently in other ways, and the appended claims are intended to be interpreted as including such variations. The various features and aspects of the above-mentioned subject matter may be used alone or in combination. In addition, without departing from the broader spirit and scope of this specification, the embodiments may be used in any number of environments and applications other than those described herein. Therefore, the description and drawings are to be considered illustrative, not restrictive. For illustration purposes, the methods are described in a particular order. It should be understood that, in alternative embodiments, these methods may be performed in a different order than described.

[0147] Where a component is described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0148] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≦") and greater than or equal to ("≧"), respectively, without departing from the scope of this specification.

[0149] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the features disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above based on their functions. Whether such functions are implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Technicians can implement the described functions in different ways for each specific application, but such implementation decisions should not be interpreted as causing departure from the scope of this application.

[0150] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device mobile phone, or an integrated circuit device with multiple uses (including applications in wireless communication device mobile phones and other devices). Any features described as modules or components may be implemented together in an integrated logic device, or individually as separate but interoperable logic devices. If implemented in software, these techniques may be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions that, when executed, perform one or more of the above methods. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage medium, etc. Additionally or alternatively, the techniques may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0151] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0152] Figure 9 An example computing system architecture 900 of a computing device that can implement the various techniques described herein is shown. For example, the computing system architecture 900 can be composed of Figure 1The illustrated media processing system 102 is implemented to perform the media data processing and rendering techniques described herein. The components of the computing system architecture 900 are shown in electrical communication with each other using a connection 905, such as a bus. The example computing device 900 includes a processing unit (CPU or processor) 910 and a computing device connection 905 that couples various computing device components, including computing device memory 915 (e.g., read-only memory (ROM) 920 and random access memory (RAM) 925), to the processor 910. The computing device 900 may include a cache of high-speed memory directly connected to, in close proximity to, or integrated as part of the processor 910. The computing device 900 may copy data from memory 915 and / or storage device 930 to the cache 912 for rapid access by the processor 910. In this way, the cache can provide a performance boost by avoiding delays in the processor 910 while waiting for data. These and other modules may control or be configured to control the processor 910 to perform various actions. Other computing device memories 915 may also be used. Memory 915 may include a variety of different types of memory with different performance characteristics. Processor 910 may include any general-purpose processor and hardware or software services configured to control processor 910, as well as specialized processors (e.g., service 1 932, service 2 934, and service 3 936 stored in storage device 930), where software instructions are incorporated into the processor design. Processor 910 may be a standalone system containing multiple cores or processors, a bus, a memory controller, a cache, etc. Multi-core processors may be symmetric or asymmetric.

[0153] To enable user interaction with the computing device 900, the input device 945 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice, and the like. The output device 935 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, a projector, a television, a speaker device, and the like. In some cases, a multimodal computing device can enable a user to provide multiple types of input to communicate with the computing device 900. The communication interface 940 can generally govern and manage user input and computing device output. There is no limitation on operating on any particular hardware arrangement, so the basic features herein can be easily replaced by improved hardware or firmware arrangements developed.

[0154] The storage device 930 is a non-volatile memory and may be a hard disk or other type of computer-readable medium that can store computer-accessible data, such as a magnetic tape cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a dark cartridge, a random access memory (RAM) 925, a read-only memory (ROM) 920, and a mixture thereof.

[0155] The storage device 930 may include services 932, 934, 936 for controlling the processor 910. Other hardware or software modules are contemplated. The storage device 930 may be connected to the computing device connection 905. In one aspect, a hardware module that performs a particular function may include a software component stored in a computer-readable medium that is associated with the necessary hardware components (e.g., the processor 910, the connection 905, the output device 935, etc.) to perform that function.

[0156] For clarity of explanation, in some cases, the technology may be presented as including separate functional blocks including functional blocks including devices, device components, steps or routines in methods implemented in software or a combination of hardware and software.

[0157] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0158] The methods according to the above examples can be implemented using computer-executable instructions stored in or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or group of functions. Some of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the examples include magnetic or optical disks, flash memory, USB devices equipped with non-volatile memory, network storage devices, etc.

[0159] Devices implementing the disclosed methods may include hardware, firmware, and / or software and may be implemented in any of a variety of form factors. Typical examples of these form factors include laptop computers, smartphones, small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, and the like. The functionality described herein may also be embodied in peripheral devices or plug-in cards. As another example, such functionality may be implemented on a circuit board between different chips or in different processes executed within a single device.

[0160] The instructions, the media for conveying those instructions, the computing resources for executing them, and other structures for supporting those computing resources are exemplary means for providing the functionality described in this disclosure.

[0161] Although various examples and other information are used to explain aspects within the scope of the appended claims, no limitations on the claims should be implied based on the specific features or arrangements in these examples, as one of ordinary skill in the art will be able to use these examples to derive a variety of implementations. In addition, although some subject matter may have been described in language specific to structural features and / or method steps, it should be understood that the subject matter defined in the appended claims is not necessarily limited to these described features or actions. For example, such functionality may be distributed differently or performed in components other than those identified herein. Instead, the described features and steps are disclosed as examples of components, computing devices, and methods within the scope of the appended claims.

[0162] Claim language that recites "at least one of" a group indicates that one member of the group or multiple members of the group satisfy the claim. For example, claim language that recites "at least one of A and B" refers to A, B, or A and B.

Claims

1. A method for providing a rendering engine model for media data, the method comprising: obtaining media data captured by a data capture device, the media data comprising at least one of image data, video data, and audio data; embedding the rendering engine model into a media item containing the media data, the rendering engine model comprising a description of a neural network for implementing a rendering engine, the neural network being configured to process the media data and generate a displayable specific rendering data output, the description defining a neural network architecture of the neural network; as well as The media item having the rendering engine model embedded in the media item is provided to one or more devices, the rendering engine model in the media item including instructions for executing the neural network to process the media data and generate the displayable specific rendering data output based on the description of the neural network.

2. The method of claim 1 , wherein the neural network architecture comprises a set of filters, activation functions, and parameters implemented along a path of interconnected layers in the neural network architecture, wherein the parameters comprise weights associated with one or more of the interconnected layers.

3. The method of claim 2, wherein the description of the neural network comprises: connection information defining how the interconnect layers are connected to form paths of the interconnect layers; Input information defining how to form an input to the neural network based on the media data; as well as Output information, which defines how the output is formed from the neural network.

4. The method of claim 1 , wherein the description of the neural network defines one or more customized tasks for the neural network, the one or more customized tasks comprising at least one of: encoding the media data, decoding the media data, performing one or more compression operations on the media data, performing one or more image processing operations on the media data, performing a frame rate conversion operation, and performing an audio signal modification operation. 5 . The method of claim 4 , wherein the one or more image processing operations include at least one of: an image restoration operation, an image enhancement operation, a filtering operation, a scaling operation, and an image rendering operation.

6. The method of claim 1 , wherein the media item comprises a data file or a data stream, the method further comprising: embedding an additional rendering engine model in the media item, the additional rendering engine model comprising an additional description of an additional neural network, the additional neural network configured to process the media data and generate a different displayable rendered data output, the additional description defining a different neural network architecture for the additional neural network, the different neural network architecture being tailored for a different operational outcome based on at least one of different layers, different filters, different activation functions, and different parameters defined for the different neural network architecture; as well as The media item having the rendering engine model and the additional rendering engine model embedded in the media item is provided to the one or more devices.

7. The method according to claim 6, further comprising: generating the neural network or one of the different neural networks based on the rendering engine model or one of the additional rendering engine models; as well as The media data is processed using the neural network or one of the different neural networks to generate the rendering data output or one of the different rendering data outputs.

8. The method of claim 1, wherein the media data comprises raw media data from the data capture device, wherein the data capture device comprises at least one of an image capture device, a video capture device, and an audio capture device.

9. The method according to claim 1, further comprising: embedding an additional rendering engine model in the media item, the additional rendering engine model comprising an address of a remote location of an additional description of an additional neural network configured to process the media data and generate a corresponding displayable rendered data output, the additional description defining a corresponding neural network architecture of the additional neural network, the corresponding neural network architecture being customized for a corresponding operational outcome based on at least one of a corresponding layer, a corresponding filter, a corresponding activation function, and a corresponding parameter defined for the corresponding neural network architecture; and The media item having the rendering engine model and the additional rendering engine model embedded in the media item is provided to the one or more devices.

10. The method according to claim 9, further comprising: retrieving, from the remote location, an additional description of the additional neural network based on the address; generating the additional neural network based on the additional description of the additional neural network; as well as The media data is processed using the additional neural network to generate the corresponding displayable rendering data output, the corresponding displayable rendering data output including rendered media data.

11. The method according to claim 1 , further comprising: receiving the media item having the rendering engine model embedded in the media item; generating the neural network based on the rendering engine model; as well as The media data in the media item is processed using the neural network to generate the specific rendering data output, the specific rendering data output comprising rendered media data.

12. The method of claim 1 , wherein providing the media item with the rendering engine model to one or more devices comprises at least one of storing the media item with the rendering engine model on the one or more devices, and sending the media item with the rendering engine model to the one or more devices.

13. An apparatus for providing a rendering engine model for media data, the apparatus comprising: Memory; as well as A processor configured to: obtaining media data captured by a data capture device, the media data comprising at least one of image data, video data, and audio data; inserting the rendering engine model into a media item containing the media data, the rendering engine model comprising a description of a neural network for implementing a rendering engine, the neural network being configured to process the media data and generate a specific displayable rendering data output, the description defining a neural network architecture of the neural network; as well as Outputting the media item having the rendering engine model embedded in the media item, the rendering engine model in the media item including instructions that specify how to execute the neural network to process the media data and generate the displayable specific rendering data output based on the description of the neural network.

14. The apparatus of claim 13, wherein the neural network architecture comprises a set of filters, activation functions, and parameters implemented along a path of interconnected layers in the neural network architecture, wherein the parameters comprise weights associated with one or more of the interconnected layers.

15. The apparatus of claim 14, wherein the description of the neural network comprises: connection information defining how the interconnect layers are connected to form paths of the interconnect layers; Input information defining how to form an input to the neural network based on the media data; as well as Output information, which defines how the output is formed from the neural network.

16. The apparatus of claim 13 , wherein the description of the neural network defines one or more customized tasks for the neural network, the one or more customized tasks comprising at least one of: encoding the media data, decoding the media data, performing one or more compression operations on the media data, performing one or more image processing operations on the media data, performing a frame rate conversion operation, and performing an audio signal modification operation.

17. The apparatus of claim 16, wherein the one or more image processing operations include at least one of: an image restoration operation, an image enhancement operation, a filtering operation, a scaling operation, and an image rendering operation.

18. The apparatus of claim 13, wherein the media item comprises a data file or a data stream, wherein the processor is configured to: inserting an additional rendering engine model into the media item, the additional rendering engine model comprising an additional description of an additional neural network configured to process the media data and generate a different displayable rendered data output, the additional description defining a different neural network architecture for the additional neural network, the different neural network architecture being tailored for a different operational outcome based on at least one of different layers, different filters, different activation functions, and different parameters defined for the different neural network architecture; and The media item is output with the rendering engine model and the additional rendering engine model embedded in the media item.

19. The apparatus of claim 18, wherein the processor is configured to: generating the neural network or one of the different neural networks based on the rendering engine model or one of the additional rendering engine models; and The media data is processed using the neural network or one of the different neural networks to generate the rendering data output or one of the different rendering data outputs.

20. The apparatus of claim 13, wherein the apparatus comprises at least one of: a mobile device, the data capture device, the one or more devices, and a display for displaying the specific rendered data output.

21. The apparatus of claim 13, wherein the processor is configured to: inserting an additional rendering engine model into the media item, the additional rendering engine model comprising an address of a remote location of an additional description of an additional neural network configured to process the media data and generate a corresponding displayable rendered data output, the additional description defining a corresponding neural network architecture of the additional neural network, the corresponding neural network architecture being customized for a corresponding operational outcome based on at least one of a corresponding layer, a corresponding filter, a corresponding activation function, and a corresponding parameter defined for the corresponding neural network architecture; and The media item is output with the rendering engine model and the additional rendering engine model embedded in the media item.

22. The apparatus of claim 21 , wherein the processor is configured to: retrieving, from the remote location, an additional description of the additional neural network based on the address; generating the additional neural network based on the additional description of the additional neural network; and The media data is processed using the additional neural network to generate a corresponding rendering data output, the corresponding rendering data output comprising rendered media data.

23. The apparatus of claim 13, wherein the processor is configured to: generating the neural network based on the rendering engine model in the media item; and The media data in the media item is processed using the neural network to generate the specific rendering data output, the specific rendering data output comprising rendered media data.

24. The apparatus of claim 13, wherein the apparatus comprises a mobile device.

25. The apparatus of claim 13, further comprising a data capture device for capturing the media data.

26. The apparatus of claim 13, further comprising a display for displaying one or more images.

27. A non-transitory computer-readable storage medium for providing a rendering engine model for media data, the non-transitory computer-readable storage medium comprising: Instructions stored therein, when executed by one or more processors, cause the one or more processors to: obtaining media data captured by a media data capturing device, wherein the media data includes at least one of image data, video data, and audio data; embedding the rendering engine model into a media item containing the media data, the rendering engine model comprising a description of a neural network for implementing a rendering engine, the neural network being configured to process the media data and generate a displayable specific rendering data output, the description defining a neural network architecture of the neural network; as well as Outputting the media item having the rendering engine model embedded in the media item, the rendering engine model in the media item including instructions for executing the neural network to process the media data and generate the displayable specific rendering data output based on a description of the neural network.

28. The non-transitory computer-readable storage medium of claim 27, wherein the neural network architecture comprises a set of filters, activation functions, and parameters implemented along a path of interconnected layers in the neural network architecture, wherein the parameters comprise weights associated with one or more of the interconnected layers.

29. The non-transitory computer-readable storage medium of claim 28, wherein the description of the neural network comprises: connection information defining how the interconnect layers are connected to form paths of the interconnect layers; Input information defining how to form an input to the neural network based on the media data; as well as Output information, which defines how the output is formed from the neural network.

30. The non-transitory computer-readable storage medium of claim 27, wherein the neural network is configured to perform one or more customized tasks, the one or more customized tasks comprising at least one of: encoding the media data, decoding the media data, performing one or more compression operations on the media data, performing one or more image processing operations on the media data, performing a frame rate conversion operation, and performing an audio signal modification operation.

31. The non-transitory computer-readable storage medium of claim 27, storing additional instructions that, when executed by the one or more processors, cause the one or more processors to: embedding an additional rendering engine model in the media item, the additional rendering engine model comprising an additional description of an additional neural network configured to process the media data and generate a different displayable rendered data output, the additional description defining a different neural network architecture for the additional neural network, the different neural network architecture being tailored for a different operational outcome based on at least one of different layers, different filters, different activation functions, and different parameters defined for the different neural network architecture; and The media item is output with the rendering engine model and the additional rendering engine model embedded in the media item.

Citation Information

Patent Citations

  • Editing digital images utilizing a neural network with an in-network rendering layer

    US20180253869A1

  • Consistent 3D rendering in medical imaging

    US20180260997A1

  • A method and technical equipment for video processing

    WO2018150083A1