Contextual error concealment to improve inference accuracy
Context-aware error concealment in video streaming selectively applies error masking based on data loss location and region of interest, enhancing AI inference accuracy by optimizing computational efficiency and performance.
Patent Information
- Application Number
- DE102025133015
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-22
- Filing Date
- 2025-08-19
- Publication Date
- 2026-02-26
AI Technical Summary
Video streaming is prone to data loss due to network latency and packet loss, which affects the accuracy of artificial intelligence inference processes by causing significant artifacts and latency, and conventional error masking techniques impair performance by being computationally intensive.
Implement context-aware error concealment by selectively applying error masking functions based on the location of corrupted or lost data within a video frame, particularly in regions of interest, using low-complexity methods when data loss occurs outside these regions and high-complexity methods when data loss affects critical areas.
Improves the accuracy of artificial intelligence operations by reducing computational resources and maintaining inference performance, especially in noisy networks, by dynamically selecting error masking techniques based on the frame's region of interest and data loss location.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Video streaming involves encoding video data and transmitting it over a network to a remote client device, which then decodes the video. A disadvantage of video streaming is the potential loss of video data during streaming, which can be caused by factors such as network latency and packet loss. Packet loss can cause significant artifacts or latency, negatively impacting the overall streaming experience. Furthermore, corrupted or lost data in video streams can affect the accuracy of artificial intelligence prediction and inference processes within that stream. SUMMARY
[0002] The invention is defined in the claims. For the purpose of illustrating the invention, aspects and embodiments are described herein that may or may not fall within the scope of protection of the claims.
[0003] Several examples disclose systems and methods for context-aware error concealment to improve inference accuracy. One system can identify a frame in a video stream and determine that the frame contains corrupted or lost data. The system can then generate a corrected frame by applying an error concealment function selected based on at least the location of the corrupted or lost data within the frame and a region of interest within the frame.
[0004] Embodiments of the present disclosure relate to context-aware error masking in video streams to improve the accuracy of artificial intelligence operations. The systems and methods described herein improve upon conventional error masking techniques by selectively applying error masking depending on the location of the detected damage / loss in the video data. In contrast to conventional approaches, where computationally intensive error masking functions are applied to every part of every frame, thereby impairing the performance of the artificial intelligence, the techniques described herein can be used to selectively apply error masking to portions of a frame that are most critical for the subsequent artificial intelligence operations.In some implementations or scenarios, error masking may not be applied even if corruption or data loss is detected. This significantly reduces the computational resources required to apply error masking to video stream data, improving the performance of downstream artificial intelligence operations.
[0005] At least one aspect relates to one or more processors. The one or more processors may contain one or more circuits. The one or more circuits may identify a frame of a video stream. The one or more circuits may determine that the frame contains corrupted or lost data. The one or more circuits may produce a corrected frame by applying an error masking function selected based on at least the location of the corrupted or lost data within the frame and / or a region of interest within the frame.
[0006] In some implementations, one or more circuits can select the error masking function based at least on the position of the corrupted or lost data within the region of interest in the frame. In some implementations, one or more circuits can generate the corrected frame by applying the error masking function to at least one portion of the region of interest. In some implementations, one or more circuits can apply the error masking function by providing the frame as input to a machine learning model.
[0007] In some implementations, one or more circuits can receive an encoded bitstream of the video stream. In some implementations, one or more circuits can generate the frame by decoding the encoded bitstream, with the decoding of the frame indicating the position of the corrupted or lost data. In some implementations, the error masking function is a first error masking function from a set of error masking functions. In some implementations, one or more circuits can determine that an object is detected in a predetermined number of previous frames in the video stream.
[0008] In some embodiments, the one or more circuits can select the first error-masking function from the plurality of error-masking functions, at least based on the object detected in the predetermined number of previous frames. The first error-masking function can utilize a greater amount of computing resources compared to a second error-masking function selected from the plurality of error-masking functions. In some implementations, the one or more circuits can generate the corrected frame by applying the first error-masking function to at least one portion of the frame in which the object is estimated to appear. In some implementations, the one or more circuits can determine that the corrupted or lost data is located outside the region of interest.In some implementations, one or more circuits can select the first fault cover function from the multitude of fault cover functions, at least based on the corrupted or lost data that lies outside the region of interest, with the first fault cover function using a smaller amount of computing resources compared to a second fault cover function from the multitude of fault cover functions.
[0009] At least one aspect relates to a system. The system can contain one or more processors. The system can receive a request to process a video stream comprising a plurality of frames. The system can apply an error masking function to at least one frame from the plurality of frames, the error masking function being selected based at least on the location of corrupted or lost data in the at least one frame and a region of interest in the at least one frame. The system can execute a machine learning model that uses the at least one frame as input for the request.
[0010] In some implementations, the system can determine that at least one frame contains corrupted or lost data by using a decoding process. In some implementations, the machine learning model generates a clue about an object in the at least one frame. In some implementations, the system can select a second error-masking function for at least one second frame from the multitude of frames, based at least on the clue about the object in the at least one frame.
[0011] In some implementations, the second error concealment function for the at least one second frame is selected based on an expected position of the object within that frame. In some implementations, the position of the damage or data loss is a macroblock position or a slice position within the at least one frame. In some implementations, the system can apply the error correction function by running a second machine learning model using the at least one frame as input, with the second machine learning model generating replacement information for the damaged or lost data of the at least one frame.
[0012] At least one aspect relates to a method. The method may include identifying a frame of a video stream using one or more processors. The method may include determining, using the one or more processors, that the frame contains corrupted or lost data. The method may include generating a corrected frame using the one or more processors by applying an error-masking function selected based on at least the location of the corrupted or lost data in the frame and a region of interest in the frame.
[0013] In some implementations, the procedure includes selecting the error masking function using one or more processors, at least based on the position of the corrupted or lost data located within the region of interest in the frame. In some implementations, the procedure includes generating the corrected frame using one or more processors by applying the error masking function to at least one portion of the region of interest.
[0014] The processors, systems, and / or methods described herein can be implemented by a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twin operations, a system for performing light transport simulations, a system for performing collaborative content creation for 3D assets, a system for performing deep learning operations, a system for performing generative AI operations using a large language model (LLM), a system for performing generative AI operations using a vision language model (VLM), a system implemented using an edge device, a system implemented using a robot, a system for performing conversational AI operations, a system for generating synthetic data, a system,that contains one or more virtual machines (VMs), a system that is at least partially implemented in a data center, or a system that is at least partially implemented using cloud computing resources, or that is implemented or includes such a system.
[0015] The revelation extends to all novel aspects or features described and / or illustrated herein.
[0016] Further features of the disclosure are characterized by the independent and dependent claims.
[0017] Any feature in one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0018] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0019] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.
[0020] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0021] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0022] The disclosure also provides a computer or computer system (including networked or distributed systems) with an operating system that supports a computer program for carrying out one of the methods described herein and / or for embodying one of the device or system features described herein.
[0023] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0024] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0025] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0026] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present systems and methods for context-related error concealment to improve inference accuracy are described in detail below with reference to the attached drawings, wherein: Fig. 1 is a block diagram of an exemplary system for implementing context-related error concealment to improve inference accuracy according to some embodiments of the present disclosure. Fig. 2 an exemplary data flow diagram showing how captured video data is processed by security devices in accordance with the context-related error concealment techniques described herein according to some embodiments of the present disclosure; Fig. 3 an exemplary diagram showing how damaged or lost encoded data affects video decoding according to some embodiments of the present disclosure; Fig. 4 an exemplary diagram showing how damaged or lost data may occur in a section(s) of a video image outside and inside a region of interest, according to some embodiments of the present disclosure; Fig. 5 is an exemplary diagram showing larger segments of lost or damaged data in a single image, according to some embodiments of the present disclosure; Fig. 6 is a flowchart of an exemplary procedure for implementing context-related error concealment to improve inference accuracy according to some embodiments of the present disclosure; Fig. 7 is a block diagram of an exemplary content streaming system suitable for use in the implementation of some embodiments of the present disclosure; Fig. 8 is a block diagram of an exemplary computing device suitable for use in the implementation of some embodiments of the present disclosure; and Fig. 9 is a block diagram of an exemplary data center suitable for use in the implementation of some embodiments of the present disclosure. DETAILED DESCRIPTION
[0028] This disclosure relates to systems and methods for implementing context-aware error concealment in video streams, which can be applied to improve the accuracy of machine learning operations. The error concealment techniques described herein can be performed by systems transmitting video data or data derived from video data, including security systems, artificial intelligence pipelines, or video streaming platforms in general.
[0029] These techniques are particularly useful in unreliable or noisy networks where packets can be dropped or transmitted data corrupted. Video streaming can be accomplished by transmitting packets over a streaming protocol, such as the Real-Time Streaming Protocol (RTSP). Such streaming protocols transmit video data encoded into a "bitstream" to improve throughput and accuracy. If sections of a bitstream are corrupted or lost, the corresponding frame-by-frame video information is also lost during decoding at the receiver of the video stream. For artificial intelligence systems that perform inference on decoded video data, missing or corrupted frame segments can impair detection or tracking accuracy.
[0030] Conventional error masking approaches that attempt to mitigate damage or data loss result in reduced inference performance because sophisticated and computationally intensive error masking techniques are applied to every frame in the video stream containing damaged or lost data. To address these shortcomings, the systems and methods described herein dynamically select from different error masking approaches based on the location of the lost or damaged frame data. In some implementations, error masking may not be applied even if damage or data loss is detected.
[0031] To this end, the systems and methods described herein can utilize a region of interest in the frame that corresponds to the most relevant sections of the frame for detection, inference, classification, segmentation, or other machine learning tasks. This region can be predetermined or dynamically determined and can correspond to the sections of the video stream that most likely represent objects subject to artificial intelligence operations. For example, the region of interest might be a section near the center of the video frame. To improve overall performance, the system and methods described herein can select the optimal error concealment technique depending on whether the data corruption or loss in the video frame occurs within or near one or more regions of interest in the video frame.
[0032] For example, if damage or data loss occurs outside the region(s) of interest in a video frame, low-cost error concealment techniques can be selected. In some implementations, error concealment cannot be performed if the data loss or damage occurs outside the region(s) of interest. Another example: If damage or data loss occurs in a region(s) of interest in a video frame, a high-quality, computationally intensive error concealment technique can be applied to improve the accuracy of the machine learning inference. One example of computationally intensive error concealment is the use of a generative AI model that takes a damaged frame as input and produces an uncorrupted video frame as output.
[0033] The error masking technique selected for video frames can depend on whether objects have recently been detected in one or more regions of interest within a video stream. For example, if an object in the region(s) of interest of a video stream has not been detected within a predetermined number of recent consecutive frames, low-computation error masking techniques can be selected. Conversely, if an object in one or more regions of interest of the video stream has been detected within a predetermined number of recent consecutive frames, a high-quality, computationally intensive error masking technique can be applied.
[0034] Error concealment techniques can also be selectively applied to different sections (e.g., macroblocks) of the video frame if overall data loss or corruption occurs. For example, if data packets corresponding to a slice of the video data are lost or corrupted, error concealment can be applied only to sections (e.g., macroblocks) of the video frame corresponding to one or more regions of interest, rather than to the entire video frame. Similar approaches can be used to selectively apply error concealment to sections of the region(s) of interest where an object is estimated to appear, at least based on the detection of the object in previous video frames.
[0035] With reference to Fig. 1, is Fig. 1 An exemplary computing environment comprising a system for implementing context-aware error concealment for improving inference accuracy according to some embodiments of the present disclosure. It should be noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and location.Various functions described herein, which are performed by entities, can be executed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory.
[0036] System 100 is shown to include a data processing system 102 that receives encoded video data 112, which may be an encoded bitstream of a video stream containing encoded frames 113. The data processing system 102 can implement the various techniques described herein to mask errors in the encoded video data 112 and thus improve inference accuracy. For example, the data processing system 102 can receive the encoded video data 112 from one or more computer networks. In some implementations, the data processing system 102 can access the encoded video data from a data repository or storage system. The storage system may be an external server, a distributed storage / computing environment (e.g., a cloud storage system), or some other type of storage device or system associated with the data processing system 102.In some embodiments, the storage system can be part of the data processing system 102 or be integrated into it in some other way. In such implementations, the data processing system 102 can access the encoded video data 112 from the internal working memory.
[0037] The data processing system 102 can access, retrieve, or otherwise receive the encoded video data 112 when it receives a request to process the encoded video data 112. The request can be provided by a device located outside of the data processing system 102 that communicates with it (e.g., a client device communicating over a network). In some implementations, the request can be provided in response to an input into the data processing system 102, for example, by an operator of the data processing system 102. The request can specify the encoded video data 112 to be processed or a location from which the data processing system should retrieve the encoded video data 112.
[0038] It is shown that the encoded video data 112 contains one or more encoded frames 113. The encoded video data 112 can be an encoded bitstream of video data generated by one or more video recording devices (e.g., video cameras, security cameras, built-in webcams, or smartphone cameras, etc.) or applications that generate video frames (e.g., remote gaming applications, remote desktop applications, etc.). The encoded video data 112 can be generated from any suitable source, including a video playback process, a gaming process (e.g., video output from remote-controlled video games), and other video data sources. The encoded video data 112 can contain information about the frames of the video stream, including the resolution, frame rate, or other attributes of the video stream.The encoded video data 112 can be encoded according to any suitable codec standard, including, but not limited to, codec standards such as h.264 (AVC), h.265 (HEVC), h.266 (VVC), AV1, VP8, VP9, or any other video codec that supports the segmentation of a single video frame into different geometric regions, such as slices or macroblocks. The frame data in the encoded video data 112 can be encoded as one or more encoded frames 113.
[0039] It is shown that the encoded video frames 113 are stored as part of the encoded video data 112. The decoder can consist of, or contain, software, hardware, or combinations of hardware and software. Each frame within the encoded video data 112 can be an independent unit of encoded data and can be compressed using various techniques to reduce redundancy between consecutive frames. The encoded frames 113 and / or the encoded video data 112 can contain metadata describing the structure and properties of each frame, such as frame type (I-frame, P-frame, B-frame), size, timestamp, and dependencies on other frames. In some implementations, the data for the encoded frames 113 can be stored sequentially or in a nested manner as part of the encoded video data 112.
[0040] The data processing system 102 can execute a decoder 104 for decoding the encoded video data 112. For this purpose, the decoder 104 can parse the encoded video data 112 to extract all associated video metadata, such as the frame size, frame rate, and audio sampling rate. The decoder 104 can identify the codec based on the metadata and decode the encoded video data 112 using the identified codec to produce frames of the video stream and / or audio data. In some implementations, such video metadata can be accessed separately from the encoded video data 112 (e.g., stored separately or provided separately in one or more network packets, etc.). The decoding of the encoded video data 112 may include decompression or reversing all encoding operations used to generate the encoded video data 112.
[0041] The decoded data generated from the encoded video data 112 is stored as decoded frame(s) 110. The decoded frame(s) 110 may contain raw video frame data (e.g., a collection of pixels with the resolution of the video stream). The decoded frames 110 may be stored in a frame buffer or other memory region of the data processing system 102 for processing by the machine learning model(s) 118 and / or the error masker 108. As shown in this example, the decoded frames 110 are provided to the error masker 108. In some implementations, each decoded frame 110 may be processed sequentially by the error masker 108. In some implementations, the decoded frames 110 may be processed by the error masker 108 in a batch (e.g.,as soon as a predetermined number of decoded individual images (110) have been generated).
[0042] It is shown that the decoder 104 includes an error detector 106. The error detector 106 can include software, hardware, or a combination of both, and can be used to detect errors in the encoded frames 113 while the encoded video data 112 is being decoded. The error detector 106 can detect errors caused by corrupted or lost data in the encoded video data 112. Since the encoded video data 112 may contain streamed video data from one or more video capture systems, sections of the encoded video data 112 may become corrupted or lost during transmission. In implementations where the encoded video data 112 is stored in a repository or storage system, corruption may have occurred in the repository / storage system or during transmission to the repository / storage system.
[0043] Damage to the encoded video data 112 can involve the loss of sections of one or more encoded frames 113, resulting in the inability to correctly reconstruct the information in the frames. These sections can be detected by the error detector 106 using various techniques such as parity checks, checksums, or error detection approaches that implement cyclic redundancy checks (CRC). In some implementations, the encoded video data 112 can be or contain one or more sections of an encoded bitstream, and the error detector 106 can detect instances of missing sections of encoded frames 113 by identifying missing sections of the encoded bitstream that correspond to the encoded frames 113.For example, while the decoder 104 decodes the encoded video data 112, the error detector 106 can examine the bitstream for inconsistencies or anomalies that indicate damage. This may involve accessing data corresponding to the encoded frames 113 to detect deviations from the expected patterns, including, but not limited to, discrepancies in frame size, unexpected breaks in the encoded bitstream, or incorrect checksum values.
[0044] All encoded frames 113 identified by the error detector 106 as containing missing or corrupted data can be flagged for error masking. These flags can be stored as part of the metadata in the decoded frames 110. In some implementations, decoded frames 110 flagged as containing corrupted or lost data can be stored with an indication that the corruption / data loss is present. In some implementations, the error detector 106 can determine a position (e.g., pixel coordinates, coordinates of a region, identifier of one or more macroblocks, slices, or other geometric sections, etc.) within a decoded frame 110 that has been corrupted or lost.The position(s) of a damaged or lost region(s) of a decoded single image 110 can be provided to the error concealer 108 as metadata for the decoded single image 110.
[0045] The decoder 104, when generating the decoded video data for a video frame, can provide all decoded frames 110 containing corrupted or lost data to the error concealer 108. The error concealer 108 can contain software, hardware, or a combination of hardware and software and can be executed to conceal errors detected in decoded frames 110 using a suitable error concealment function. The error concealer 108 can implement any suitable type of error concealment function to conceal regions containing corrupted or lost data in the decoded frames 108, including, but not limited to, spatial concealment (e.g., interpolation), temporal concealment (e.g.,Replacing damaged / lost pixels with corresponding pixels from previous frames), motion vector estimation-based prediction of damaged or lost data, generative machine learning techniques or combinations thereof, among others.
[0046] The error masker 108 can dynamically select one of many error masking functions for a decoded frame 110 containing corrupted or lost data, based at least on the position of the corrupted or lost data in the decoded frame 110 and a region of interest within the decoded frame 110. The region of interest can be a predetermined or dynamically identified region of pixels in the decoded frame 110. The size, position, or other attributes of the region of interest can be identified in the metadata of the encoded video data 112, extracted by the decoder 104, and provided to the error masker 108 for processing.In some implementations, the size, position, or other attributes of the decoded frame 110 can be provided in the metadata of the encoded video data 112 on a frame-by-frame basis, with each encoded frame 113 containing metadata that identifies a particular region of interest.
[0047] In some implementations, the region of interest can be determined based on at least one source of encoded video data 112. In one example, the data processing system 102 can store a data structure that maps different sources of encoded video data 112 (e.g., different recording devices, video data storage locations, client devices, etc.) to the corresponding metadata relating to regions of interest. Upon receiving or identifying encoded video data 112, the error masker 108 can access the data structure by using the source of the encoded video data 112 to identify the size, position, or attributes of the region of interest for that video source. In some implementations, a region of interest can be discontinuous, and a decoded frame 110 can contain multiple continuous regions of interest.The regions of interest can differ for different sequences of decoded frames 110. In some implementations, the size or position of a region of interest for a decoded frame 110 can be modified based on at least the output of one or more machine learning models 118, as further described herein.
[0048] The type of error-masking function used to mask lost or corrupted data in a decoded frame 110 can be selected according to conditions, including, but not limited to, the position and / or size of the lost or corrupted data in the decoded frame 110, whether the lost or corrupted data segment of the decoded frame 110 lies within (or, in some implementations, in front of) the region of interest in the decoded frame 110, and / or whether one or more objects in the region of interest were detected in one or more previous frames. Various examples of different conditions for selecting different error-masking functions are given in the context of the Fig. 3, Fig. 4 and Fig. 5 described.
[0049] With reference to Fig. 3 in connection with the components that are related to Fig. Figure 1 describes an exemplary diagram 300 illustrating how damaged or lost encoded data affects video decoding, according to some embodiments of the present disclosure. As shown, a single frame 302 contains a region 304 of interest and a region 306 with lost or damaged data. Although the region 304 of interest is provided with a horizontal line pattern, the region 304 of interest need not necessarily be visually represented in the single frame 302. Furthermore, although the region 304 of interest is shown as a rectangular shape, it should be understood that any number of regions 304 of interest can be associated with a single frame 302, and the regions 304 of interest of the single frame 302 can have any size, shape, or location.
[0050] As shown in this example, frame 302 contains a region 306 with lost or corrupted data that lies below and outside the region 304 of interest. In this example, the error concealer 108 can select an error concealment function with relatively low computational overhead, which may result in reduced quality in the region corresponding to the lost or corrupted data after the error concealment function is applied. Since, in this example, the region 306 with lost or corrupted data lies outside a region of interest, the quality of this portion of frame 302 is not necessarily relevant for downstream inference operations (e.g., object detection, segmentation, etc.).Therefore, to dynamically improve the computing power of the decoding process, an error concealment function(s) can be used which has relatively low computing requirements at the expense of output quality for the corrected region.
[0051] Non-restrictive examples of low-complexity error-covering techniques that can be selected include, but are not limited to, low-complexity interpolation techniques (e.g., spatial occlusion) or temporal occlusion techniques. Temporal occlusion techniques may involve the use of appropriate pixels from earlier decoded frames 110 stored by the error coverr 108 instead of region 306 containing lost or corrupted data. In some implementations, the error coverr 108 may not apply an error-covering function and instead allow the decoded frame 110 to be processed without modification according to the techniques described herein. Further details regarding other selection criteria for error-covering functions are given in the context of the Fig. 4 and Fig. 5 described.
[0052] With renewed reference to Fig. 1. The error masker 108 can apply the selected error masking function to the decoded frames 110 to generate a sequence of corrected frames 114. The corrected frames 114 can contain modified data from the decoded frames 110, generated by the selected error masking function. The corrected frames 114 can be generated in the same order as the decoded frames 110. All generated corrected frames 114 can be interleaved with decoded frames 110 that do not contain any corrupted or lost data, so that the sequence of frames generated by the decoder 104 is maintained in the correct temporal order of the video stream.
[0053] As described herein, the frames of the encoded video data 112 may correspond to security cameras, video recording devices, or other video creation software and represent various objects or features of interest. The regions of interest in the decoded frames 110 may correspond to the regions of the decoded frames 110 where the objects or features of interest are likely to be found. The corrected frames 114 (which may include decoded frames 110 for which error masking does not need to be applied) are provided to downstream processing tasks to detect the presence of one or more objects of interest. In this example, the corrected frames 114 are provided as input to one or more machine learning models 118.
[0054] The one or more machine learning models 118 can be or contain any type of machine learning model that has been trained / updated to process individual frames of a video stream (e.g., the decoded frames 110, the corrected frames 114, etc.). For example, the one or more machine learning models 118 can contain one or more neural networks (e.g., deep neural networks (DNNs), convolutional neural networks (DNNs), recurrent neural networks (RNNs), fully connected networks, combinations thereof, etc.) that are trained / updated for image classification, segmentation, or objection detection, among other machine learning tasks.In some implementations, the one or more machine learning models 118 may be or include other types of machine learning models, such as, among others, linear or logistic regression models, decision tree models, or support vector machine models, which process various data from the decoded frames 110 and / or corrected frames 114.
[0055] The data processing system 102 can, in some implementations, process one or more of the corrected frames 114 and / or decoded frames 110 (when no error correction is required) when they are generated to detect one or more objects or features of interest positioned within the one or more regions of interest of the respective frames. In some implementations, pixels corresponding to the regions of interest can be extracted from a frame and subsequently provided as input to the machine learning model(s) 118. In some implementations, the entirety of a frame can be provided as input to the machine learning model(s) 118.In an example where the machine learning model(s) 118 contains a neural network, the data processing system 102 can execute the machine learning model(s) 118 by providing the input data to be processed to one or more input layers or input data structures of the machine learning model(s) 118 and performing the processing operations of each layer until one or more model outputs 120 are produced.
[0056] As described herein, the machine learning models 118 may contain one or more object detection models that process at least the region(s) of interest of the corrected frames 114 and / or the decoded frames 110 (if no error correction is required). In such implementations, the model output(s) 120 may contain indications of whether one or more objects of interest were detected in the region(s) of interest in a frame. Such indications may include a label showing the presence of one or more objects / features of interest, bounding box data for the one or more objects / features of interest, and / or classifications of one or more objects / features of interest, but are not limited to those detected in a frame.The model outputs 120 can be stored in association with the single image from which they were generated. In some implementations, the model outputs 120 and / or the corresponding single images can be provided to other downstream processing systems or processes. The model outputs 120 can be provided to one or more external computing systems or stored in one or more data repositories / storage systems.
[0057] In some implementations, one or more of the model outputs 120 can be provided as input for the error masker 108. For example, model outputs 120 indicating that an object / feature of interest has been detected in a frame can be used to select error masking functions for subsequent frames in the sequence of decoded frames 110. In some implementations, the data processing system 102 can initialize or otherwise store / manage one or more counters that track the number of consecutive corrected frames 114 in which an object / feature of interest is represented within the relevant region(s) of interest.The counters can be provided to, or implemented by, the error masker 108 to dynamically select different error masking functions to improve the accuracy of the downstream inference operations performed using the machine learning models. Various examples of different approaches for selecting error masking functions for the decoded frames 110 are discussed in the context of the following. Fig. 4 and Fig. 5 described.
[0058] With reference to Fig. 4 in connection with the components that are related to Fig. Figure 1 describes an exemplary diagram 400 illustrating how corrupted or lost data may occur in a section(s) of a video image outside and inside a region of interest, according to some embodiments of the present disclosure. As shown, a single frame 402 contains a region 404 of interest and regions 406A and 406B with lost or corrupted data (sometimes referred to generally as the "region(s) 406 with lost or corrupted data"). In some implementations, one or more objects / features 408 of interest may be present in the single frame 402 (shown herein as the silhouette of a person).
[0059] In some implementations, the error concealer 108 can select an error concealment function based on whether the corrupted or lost data 406 is located before or within the region(s) 404 of interest. In such scenarios, different error concealment functions can be selected depending on whether an object 408 of interest was detected in one or more previous frames within the region of interest. For example, if corrupted or lost data 406 is located before or within the region(s) 404 of interest and an object / feature 408 of interest was not detected in one or more previous frames, a computationally weak error concealment function can be selected.Since an object 408 of interest was not detected in previous frames, it is not necessarily expected to appear in the current frame 402. Therefore, a less precise, computationally intensive error masking function can be applied to frame 402 to eliminate larger artifacts and improve the overall visualization. Using an error masking function that does not require significant processing power improves the performance of the entire computing system. Furthermore, since objects / features 408 of interest are not necessarily expected to appear in region 404 of interest (based on their absence in a predetermined number of previous consecutive frames), it is unlikely that the inference accuracy will be reduced, even if the overall quality of frame 402 is affected by the corrupted or lost data 406.
[0060] Such approaches can also be used when the corrupted or lost data 406 appears prior to the region of interest 404 (e.g., in video data generated before decoding). For example, the encoded macroblock data for frame 402 can be decoded from left to right in rows, starting with the upper left corner of frame 402 and ending with the lower right corner. In this illustrated example, the section of corrupted or lost data 406A is positioned in a row preceding the macroblocks that comprise the region of interest 404. Because the decoding process can be sequential, the information decoded from the section of corrupted or lost data 406A can affect the encoded macroblocks decoded as part of the region of interest 404 in subsequent frames.In such scenarios, error concealment in the damaged or lost data section 406A can improve the quality of subsequent frames, thereby enhancing the overall accuracy of inference across a sequence of frames. In some implementations, error concealment functions can be selected based on the proximity of the damaged or lost data section 406 to the region 404 of interest. For example, if the damaged or lost data section 406 occurs in the same row as a macroblock of a region 404 of interest, the damaged or lost data section 406 can be considered to be within the region 404 of interest for the purpose of selecting an error concealment function.
[0061] In another example, if damaged or lost data 406 is located before or within the region(s) of interest 404, and an object / feature 408 of interest has been detected in a predetermined number (e.g., four) of previous consecutive frames, a computationally intensive error concealment function can be selected to improve the quality of the frame 402. Since an object / feature 408 of interest has been detected in a predetermined number of previous frames, it is likely that the same object / feature 408 of interest will also appear in the region 404 of interest of the current frame 402.To ensure that this object / feature 408 of interest is recognized or otherwise accurately processed by the machine learning model(s) 118, a more accurate, computationally intensive error concealment function can be applied to frame 402 to eliminate larger artifacts and improve the overall visualization. Non-restrictive examples of error concealment include predicting the pixel values of missing / damaged macroblocks using pixel data and / or motion vectors from previous frames, running artificial intelligence models (e.g., machine learning models, deep learning models, generative AI models, etc.) on portions of frame 402 to replace the missing / damaged data, and / or using hybrid error concealment techniques (e.g., combinations of temporal and spatial error concealment), among others.The selective application of error concealment functions that use relatively many computing resources for certain frames, and the use of lower-accuracy, low-resource approaches in situations where the detection of objects / features of interest is unlikely, improves the system's performance when processing many consecutive frames of a video stream.
[0062] With reference to Fig. 5 in connection with the components that are related to Fig. As described in Figure 1, an exemplary diagram 500, showing larger segments with lost or damaged data in a single frame, is illustrated according to some embodiments of the present disclosure. As shown, a single frame 502 contains an upper portion of a region 504A of interest that has been lost / damaged, a lower portion of the region 504B of interest that has not been lost / damaged, and a large region 506 with lost or damaged data. The upper portion 504A and the lower portion 504B of the region of interest can be referred to together as the “region 504 of interest.” In some implementations, one or more objects / features 508 of interest may be present in the single frame 502 (shown herein as the silhouette of a person).
[0063] In some implementations, the error concealer 108 can select an error concealment function based on the amount of missing / corrupted data in the frame 502. In this example, at least one data slice represents a missing / corrupted section of the frame 502, including the upper portion of the region 504A of interest. In cases where a relatively large amount of data is missing / corrupted in a frame 502, the error concealer 108 can select an error concealment function that is applied only to certain sections (e.g., corrupted sections of the region of interest) of the frame to improve the system's computational efficiency. In one example, if an object / feature 508 of interest is missing a predetermined number (e.g.,If four errors are detected in previous consecutive frames, a high-quality, computationally intensive error concealment function is selected to improve the quality of frame 502. If error concealer 108 determines that the amount of missing / corrupted data in frame 502 exceeds a threshold, error concealer 108 can selectively apply the high-quality, computationally intensive error concealment function only to missing / corrupted macroblocks corresponding to the region of interest. In this example, the error concealment function can be applied to the upper portion of region 504A of interest.
[0064] In another example, if an object / feature 508 of interest has been detected in a predetermined number (e.g., four) of previous consecutive frames, a high-quality, computationally intensive error concealment function can be selected to improve the quality of frame 502. If the error concealer 108 determines that the amount of missing / corrupted data in frame 502 exceeds a threshold, the error concealer 108 can selectively apply the high-quality, computationally intensive error concealment function only to macroblocks predicted to contain the object 508 of interest in the current frame 502. The error concealer 108 can use the detected position of an object 508 of interest in previous frames (e.g., as specified in the model outputs 120) to estimate the position of the object 508 of interest in the current frame 502.In some implementations, the estimated position of object 508 of interest in the current frame 502 can be determined based on motion vectors generated from the preceding sequential frames. When applied to macroblocks representing object 508 of interest in previous sequential frames, these vectors can specify the estimated position of object 508 of interest in the current frame 502. In some implementations, the error-masking function can be applied only to macroblocks estimated to represent the object of interest and located within the region(s) of interest associated with frame 502. This allows for the selective application of high-quality error-masking functions to specific sections of frames (e.g., the region of interest, estimated positions of objects of interest, etc.).) and the use of lower-accuracy, low-resource approaches in other cases improves the overall performance of the system when processing many consecutive frames of a video stream.
[0065] In some implementations, the error concealer 108 can select an error concealment function that implements artificial intelligence models (e.g., machine learning models such as NVIDIA's Deep Learning Super Sampling (DLSS), generative AI models, etc.) to regenerate sections of frame 502 when an object 508 of interest is detected in a predetermined number of frames and the amount of missing / corrupted data in the frame 502 exceeds a threshold. Such a machine learning model(s) can be trained / updated to receive a frame with missing / corrupted sections as input and produce an output frame containing replacement pixels that are predicted to appear in the missing / corrupted sections.In some implementations, the machine learning model(s) can receive information from previous frames to estimate the pixel values of the missing / damaged sections in the current frame. To perform processing using previous consecutive frames, the error concealer can store / manage one or more data structures that store a sliding window of previously decoded / corrected frames, in addition to the corresponding metadata and / or model outputs associated with those frames.
[0066] With reference to Fig. Figure 2 is an exemplary data flow diagram 200 illustrating how recorded video data from security devices in an environment with one or more security cameras is processed according to some embodiments of the present disclosure. The data flow diagram 200 shows a set of streaming video sources 202, which in this example include the security cameras 203A-203N (sometimes referred to generally as "security camera(s) 203"). The streaming video sources 202 can include any type of application or device capable of generating or otherwise providing video data. The video data generated or recorded by the streaming video sources 202 can be provided to the decoder / error masking process 204, which performs any one of the functions of the decoder 104 and the error masker 108 of Fig. 1 can be implemented. In some implementations, a separate decoder / error masking process 204 can be executed for each streaming video source 202 (e.g., each security camera 203). In some implementations, the decoder / error masking process 204 can receive and process multiple video streams in parallel. The decoder / error masking process 204 can be implemented in one or more computing systems (e.g., the data processing system 102, etc.).
[0067] As described herein, the output of the decoder / error concealment process 204 can contain a sequence of corrected frames of decoded video data, which can be provided to a batch processing process 206. The batch processing process 206 can be used to aggregate decoded / corrected frame data into one or more data structures for processing by a distributed machine learning system (e.g., the data processing system 102). The aggregation of the video data can involve storing multiple frames in data structures compatible with different computing hardware, including graphics processing units (GPUs) or other distributed computing components / devices. The output of the batch processing process can be provided for processing using the machine learning operations 208.
[0068] The machine learning operations 208 can include any of the machine learning models described herein (e.g., the machine learning model(s) 118). The machine learning operations 208 can include, but are not limited to, object / feature recognition, segmentation, or classification. The machine learning operations 208 can be executed sequentially for each frame or, in some implementations, in parallel for multiple frames (e.g., as received from the output of the batch processing process 206). The machine learning operations 208 can produce outputs (e.g., the model outputs 120) that can be provided to the decoder / fault concealment process 204. For example, the decoder / fault concealment process 204 can be provided with hints (e.g., positions, etc.).) be provided to determine whether previous frames represented an object / feature of interest, so that previous frames can influence the selection of error-masking functions for subsequent frames, as described herein. The outputs of the machine learning operations 208 can be provided to one or more downstream operations 210, which may include storage in one or more data repositories / storage systems, encoding operations, video streaming operations, or other processing techniques.
[0069] Fig. Figure 6 is a flowchart illustrating a method 600 for implementing context-aware error concealment to improve inference accuracy according to some embodiments of the present disclosure. Different operations of the method 600 can be implemented by the same or different devices or entities at different times. For example, one or more first devices can implement operations relating to decoding and correcting video data, and one or more second devices can implement machine learning operations (e.g., the implementation of the machine learning models 118, etc.).
[0070] Each block of the Method 600 described herein contains a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The Method 600 can also be embodied as computer-usable instructions stored on computer storage media. The Method 600 can be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name just a few. Furthermore, the Method 600 is exemplified with respect to the systems of Fig. 1 and Fig. 2 described. However, this method 600 can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.
[0071] Procedure 600 includes in block B602 the identification of a single frame (e.g., a decoded single frame 110, etc.) of a video stream (e.g., a stream of encoded video data 112). The single frame of the video stream can be generated by decoding at least one section of an encoded bitstream. In some implementations, the single frame can be received from a decoder system (e.g., a computing system implementing decoder 104, etc.). The single frame can contain raw pixel data with a resolution and color depth specified via metadata for the video stream. The single frame can be part of a sequence of single frames of the video stream. Each single frame of the video stream can be processed according to the operations described herein. In some implementations, the single frame can be an encoded single frame (e.g., an encoded single frame 113) to be decoded according to the techniques described herein.
[0072] Procedure 600 contains in block B604 the determination of whether the single image contains damaged or lost data. This determination can be made during the decoding of the single image. To achieve this, all operations related to the decoder 104 and / or the error detector 106 can be performed. Fig. As described in section 1, various techniques such as parity checks, checksums, or advanced error detection can be used to identify whether information in the single frame is missing or corrupted. In some implementations, missing or corrupted sections of the single frame can be identified by detecting missing / corrupted sections of an encoded bitstream of the video stream. For example, the bitstream can be scanned / analyzed for inconsistencies or anomalies that indicate corrupted or missing data.
[0073] Procedure 600, in block B606, involves generating a corrected frame by applying an error masking function. This function is selected based on at least the position of the damaged or lost data within the frame and a region of interest within the frame. For example, if the lost or damaged data lies outside the region of interest of the frame, a low-quality error masking function can be chosen that does not require significant computational resources. In some implementations, no error masking function can be implemented if the lost or damaged data occurs outside the region of interest in the frame.Another example: If the lost or damaged data occurs within the region of interest (or in a region of the single image that may interfere with the decoding of the region of interest), a higher-quality and more computationally intensive error concealment function can be chosen.
[0074] In some implementations, the error masking function can be selected based on whether an object of interest is detected in a predetermined number (e.g., two, three, four, etc.) of preceding consecutive frames in the video stream. As described herein, one or more machine learning models (e.g., the machine learning models 118) can be used to detect the presence and position of objects / features of interest in a frame. In some implementations, if no object of interest is detected in a predetermined number of preceding sequential frames, and missing / corrupted data is detected within the region of interest in the frame, a low-quality error masking function that does not require significant computational resources can be selected.If an object of interest is detected in a predetermined number of preceding consecutive frames, and missing / damaged data is detected within the region of interest in the frame, a higher quality and more computationally intensive error concealment function can be selected.
[0075] The amount of missing / corrupted data in a frame can affect which error masking function is selected and / or how the error masking function is applied. For example, if an object of interest is detected in a predetermined number of previous consecutive frames and a large amount (e.g., entire slice(s), etc.) of missing / corrupted data is detected within the frames, a very high-quality and computationally more intensive error masking function can be selected. In some implementations, if an object of interest is detected in a predetermined number of previous consecutive frames and a large amount (e.g., entire slice(s), etc.)If missing / damaged data within the individual frames is detected, a high-quality error masking function is applied to macroblocks in the frame where the object / feature of interest is estimated to appear. The estimated position of the object / feature of interest can be determined based on its position in one or more previous frames.
[0076] If the amount of missing / corrupted data in the single image exceeds a threshold, some implementations can use generative machine learning models to reconstruct the missing / corrupted portions of the single image, as described herein. Once the error masking function is selected, it can be applied to the single image to generate one or more corrected single images (e.g., the corrected single images 114). The corrected single images can be processed with one or more downstream processing operations, which may include processing with a machine learning model (e.g., the machine learning models 118). EXEMPLARY CONTENT STREAMING SYSTEM
[0077] With reference to Fig. Figure 7 is an exemplary system diagram for a content streaming system 700 according to some embodiments of the present disclosure. Fig. 7 contains (a) application server 702 (which has similar components, features and / or functions to the exemplary computing device 800 of Fig. 8 may contain), Client Device(s) 704 (which may contain similar components, features and / or functions to the exemplary computing device 800 of Fig. 8 may contain), and network(s) 706 (which may be similar to the network(s) described herein). In some embodiments of the present disclosure, the system 700 can be implemented to mask errors in video streams by selectively applying error-masking functions chosen based on the positions and severity of the damaged / missing data. The application session may correspond to a game streaming application (e.g., NVIDIA GeFORCE NOW), a remote desktop application, a simulation application (e.g., autonomous or semi-autonomous vehicle simulation), computer-aided design (CAD) applications, virtual reality (VR) and / or augmented reality (AR) streaming applications, deep learning applications, and / or other application types.For example, the System 700 can be implemented to receive inputs specifying one or more features of the output to be generated using a neural network model, to provide the inputs to the model to cause the model to generate the output, and to use the output for various operations including display or simulation operations.
[0078] In the System 700, the client device(s) 704 for an application session can only receive input data in response to inputs to the input device(s) 726, transmit the input data to the application server(s) 702, receive encoded display data from the application server(s) 702, and display the display data on the display 724. This offloads the more computationally intensive calculations and processing to the application server(s) 702 (e.g., rendering—especially ray or path tracing—for the graphical output of the application session is performed by the GPU(s) of the application server(s) 702). In other words, the application session is streamed from the application server(s) 702 to the client device(s) 704, thereby reducing the graphics processing and rendering requirements of the client device(s) 704.
[0079] For example, with respect to an instantiation of an application session, a client device 704 can display a single frame of the application session on the display 724, at least based on receiving the display data from the application server(s) 702. The client device 704 can receive input at one of the input devices 726 and then generate input data. The client device 704 can transmit the input data to the application server(s) 702 via the communication interface 720 and via the network(s) 706 (e.g., the Internet), and the application server(s) 702 can receive the input data via the communication interface 718. The CPU(s) 708 can receive the input data, process the input data, and transmit data to the GPU(s) 710, which causes the GPU(s) 710 to generate a rendering of the application session.The input data can represent, for example, the movement of a user's character in a game session of a game application, firing a weapon, reloading, passing a ball, starting a vehicle, etc. The rendering component 712 can render the application session (e.g., representing the result of the input data), and the rendering capture component 714 can capture the rendering of the application session as display data (e.g., as image data capturing the rendered single frame of the application session). The rendering of the application session can include ray- or path-traced lighting and / or shadow effects, which are computed using one or more parallel processing units—such as GPUs, which further utilize one or more dedicated hardware accelerators or processing cores for performing ray- or path-tracing techniques—of the application server(s) 702.In some embodiments, one or more virtual machines (VMs) – including one or more virtual components such as vGPUs, vCPUs, etc. – can be used by the application server(s) 702 to support the application sessions. The encoder 716 can then encode the display data to produce encoded display data, and the encoded display data can be transmitted to the client device 704 via the communication interface 718 over the network(s) 706. The client device 704 can receive the encoded display data via the communication interface 720, and the decoder 722 can decode the encoded display data to produce the display data. The client device 704 can then display the display data on the display 724. EXAMPLE CALCULATION DEVICE
[0080] Fig. Figure 8 is a block diagram of an exemplary computing device(s) 800 suitable for use in implementing at least some embodiments of the present disclosure. The computing device 800 may include a connection system 802 that directly or indirectly couples the following devices: main memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply 816, one or more presentation components 818 (e.g., display(s)), and one or more logic units 820. In at least one embodiment, the computing device(s) 800 may include one or more virtual machines (VMs), and / or each of its components may include virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 808 can comprise one or more vGPUs, one or more of the CPUs 806 can comprise one or more vCPUs, and / or one or more of the logic units 820 can comprise one or more virtual logic units. Thus, a compute device 800 can contain discrete components (e.g., a complete GPU allocated to the compute device 800), virtual components (e.g., a portion of a GPU allocated to the compute device 800), or a combination thereof.
[0081] Although the various blocks of Fig. Where components 8 are shown connected via the connection system 802, this is not intended as a limitation and is for clarity only. In some embodiments, for example, a presentation component 818, such as a display device, may be considered an I / O component 814 (e.g., if the display is a touchscreen). As another example, the CPUs 806 and / or GPUs 808 may contain memory (e.g., the memory 804 may constitute a memory device in addition to the memory of the GPUs 808, the CPUs 806, and / or other components). In other words, the computing device of Fig. Section 8 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are covered by the scope of protection of the computing device of Fig. 8 are being considered.
[0082] The 802 interconnect system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 802 interconnect system can be arranged in various topologies, including, but not limited to, bus, star, ring, mesh, tree, or hybrid topologies. The 802 interconnect system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the 806 CPU can be directly connected to the 804 memory.Furthermore, the CPU 806 can be directly connected to the GPU 808. In the case of a direct or point-to-point connection between components, the connection system 802 can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the computing device 800.
[0083] The 804 main memory can contain any of a variety of computer-readable media. Computer-readable media can be any available media that the 800 computing device can access. Computer-readable media can include both volatile and non-volatile media, and removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.
[0084] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other types of data. For example, main memory can store 804 computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may, but are not limited to, include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that the Computing Device 800 can access. As used herein, computer storage media do not per se include signals.
[0085] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other types of data in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media may include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the foregoing should also be included in the scope of protection of the computer-readable media.
[0086] The CPU(s) 806 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 800 to perform one or more of the procedures and / or processes described herein. The CPU(s) 806 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a plurality of software threads simultaneously. The CPU(s) 806 can contain any type of processor and may contain different types of processors depending on the type of Computing Device 800 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of Computing Device 800, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The Computing Device 800 can contain one or more CPUs 806, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0087] In addition to or as an alternative to the CPU(s) 806, the CPU(s) 808 may be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 800 to perform one or more of the procedures and / or processes described herein. One or more of the GPU(s) 808 may be an integrated GPU (e.g., with one or more of the CPU(s) 806) and / or one or more of the GPU(s) 808 may be a discrete GPU. In embodiments, one or more of the GPU(s) 808 may be a coprocessor of one or more of the CPU(s) 806. The GPU(s) 808 may be used by the Computing Device 800 to render graphics (e.g., 3D graphics) or to perform general-purpose calculations. The GPU(s) 808 can be used, for example, for general-purpose computing on GPUs (GPGPU).The GPU(s) 808 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 808 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 806 received via a host interface). The GPU(s) 808 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the 804 main memory. The GPU(s) 808 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU 808 can generate pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU 808 can have its own dedicated memory or share memory with other GPUs.
[0088] In addition to or as an alternative to the CPU(s) 806 and / or the GPU(s) 808, the logic unit(s) 820 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 800 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 806, the GPU(s) 808, and / or the logic unit(s) 820 may discretely or jointly perform any combination of the methods, processes, and / or sections thereof. One or more of the logic units 820 may be part of and / or integrated within one or more of the CPU(s) 806 and / or the GPU(s) 808, and / or one or more of the logic units 820 may be discrete components or otherwise separate from the CPU(s) 806 and / or the GPU(s) 808.In embodiments, one or more of the logic units 820 can be a co-processor of one or more of the CPUs 806 and / or one or more of the GPUs 808.
[0089] Examples of Logic Unit(s) 820 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Image Processing Units (IPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), and Arithmetic Logic Units. ALUs), application-specific integrated circuits (ASICs),Floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or PCI Express (PCIe) elements, and / or the like.
[0090] The communication interface 810 can include one or more receivers, transmitters, and / or transmit-receivers that enable the computing device 800 to communicate with other computers over an electronic network, including wired and / or wireless communication. The communication interface 810 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit(s) 820 and / or the communication interface 810 can include one or more data processing units (DPUs) to directly transmit data received over a network and / or the interconnection system 802 to one or more GPUs 808 (e.g.,to transfer a working memory thereof. In some embodiments, a plurality of computing devices 800 or components thereof, which may be similar or different in various respects, can be communicatively coupled to transmit and receive data for carrying out various operations described herein, such as to facilitate a reduction in latency.
[0091] The I / O ports 812 enable the Computing Device 800 to be logically coupled with other devices, including the I / O Components 814, the Presentation Component(s) 818, and / or other components, some of which may be built into (e.g., integrated with) the Computing Device 800. Illustrative I / O Components 814 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O Components 814 can provide a Natural User Interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be routed to a suitable network element for further processing, such as modifying and registering images.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as described in more detail below) associated with a display of the Computing Device 800. The Computing Device 800 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 800 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 800 to render immersive augmented reality or virtual reality.
[0092] The 816 power supply can include a hardwired power supply, a battery power supply, or a combination of both. The 816 power supply can power the 800 computing device to enable the operation of the 800 computing device's components.
[0093] The 818 presentation component(s) can include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The 818 presentation component(s) can receive data from other components (e.g., the 808 GPU(s), the 806 CPU(s), DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0094] Fig. Figure 9 illustrates an exemplary data center 900 that can be used in at least one embodiment of the present disclosure, such as for implementing the system 100, the operations associated with Fig. 2 described, or in one or more examples of the data center 900. The data center 900 can contain an infrastructure layer 910 of the data center, a framework layer 920, a software layer 930 and / or an application layer 940.
[0095] As in Fig. As shown in Figure 9, the infrastructure layer 910 of the data center can contain a resource orchestrator 912, clustered computer resources 914 and node computer resources (“node CRs”) 916(1)-916(N), where “N” is any positive integer. In at least one embodiment, the Node CRs 916(1)-916(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid-state memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules, and / or cooling modules, etc. In some embodiments, one or more Node CRs mayThe node CRs 916(1)-916(N) correspond to a server that has one or more of the aforementioned computing resources. Furthermore, in some embodiments, the node CRs 916(1)-916(N) may contain one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the node CRs 916(1)-916(N) may correspond to a virtual machine (VM).
[0096] In at least one embodiment, the grouped compute resources 914 can contain separate groupings of node CRs 916, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node CRs 916 within grouped compute resources 914 can contain grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 916, including the CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources for supporting one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.
[0097] The resource orchestrator 912 can configure or otherwise control one or more node CRs 916(1)-916(N) and / or grouped compute resources 914. In at least one embodiment, the resource orchestrator 912 can include a software design infrastructure (SDI) management entity for the data center 900. The resource orchestrator 912 can include hardware, software, or a combination thereof.
[0098] In at least one embodiment, as in Fig. As shown in Figure 9, the framework layer 920 can contain a job scheduler 928, a configuration manager 934, a resource manager 936, and / or a distributed file system 938. The framework layer 920 can contain a framework that supports the software 932 of the software layer 930 and / or application(s) 942 of the application layer 940. The software 932 or the application(s) 942 can each contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 920 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can utilize a distributed file system 938 for processing large amounts of data (e.g., "Big Data"), but is not limited to it.In at least one embodiment, the job scheduler 928 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 900. The configuration manager 934 can be capable of configuring different layers, such as the software layer 930 and the framework layer 920, which contains Spark and the distributed file system 938, to support the processing of large amounts of data. The resource manager 936 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 938 and the job scheduler 928. In at least one embodiment, the clustered or grouped compute resources can include the grouped compute resource 914 on the infrastructure layer 910 of the data center.The Resource Manager 936 can coordinate with the Resource Orchestrator 912 to manage these allocated or assigned computing resources.
[0099] In at least one embodiment, the software contained in software layer 930 may include software 932 that is used by at least sections of the node CRs 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of framework layer 920. One or more types of software may include, but are not limited to, web page search software, email virus scanning software, database software, and streaming video content software.
[0100] In at least one embodiment, the application(s) 942 contained in the application layer 940 may contain one or more types of applications used by at least sections of the nodes CRs 916(1)-916(N), the grouped compute resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0101] In at least one embodiment, a configuration manager 934, resource manager 936, and resource orchestrator 912 can implement any number and type of self-modifying actions based on at least any set and type of data acquired in any technically feasible manner. Self-modifying actions can relieve a data center operator of the data center 900 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0102] The Data Center 900 may contain tools, services, software, or other resources to update / train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model(s) may be updated / trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to the Data Center 900.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 900 by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.
[0103] In at least one embodiment, the data center can use 900 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to update / train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0104] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the computing device(s) 800. Fig. 8. For example, each device may contain similar components, features, and / or functionality to the computing device(s) 800. Furthermore, if backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 900, an example of which is given herein with reference to Fig. 9 is described in more detail.
[0105] The components of a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can contain one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0106] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein can be implemented on any number of client devices with reference to a server.
[0107] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework, such as one that uses a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to that.
[0108] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server, the core server(s) may offload at least some functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0109] The client device(s) may include at least some of the components, features, and functions described herein with respect to Fig.The exemplary computing device(s) described in Section 800 may include 800. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied.
[0110] The disclosure of this application also contains the following numbered clauses: Clause 1. One or more processors comprising the following: one or more circuits for the following: Identifying a single frame of a video stream; Determine that the single image contains damaged or lost data; and Generating a corrected single image by adding a applies an error concealment function that is selected based at least on the position of the damaged or lost data in the single image and a region of interest in the single image. Clause 2. The one or more processors according to Clause 1, wherein the one or more circuits shall select the error concealment function at least on the basis of the position of the damaged or lost data located within the region of interest in the single image. Clause 3. The one or more processors according to Clause 1 or Clause 2, wherein the one or more circuits shall produce the corrected single image by applying the error concealment function to at least one section of the region of interest. Clause 4. The one or more processors according to any of the preceding clauses, wherein the one or more circuits shall apply the error concealment function by providing the single image as input to a machine learning model. Clause 5. The one or more processors according to any of the preceding clauses, wherein the one or more circuits are used for the following: Receiving an encoded bitstream of the video stream; and Generating the single image by decoding the encoded bitstream, where decoding the single image indicates the position of the damaged or lost data. Clause 6. The one or more processors according to any of the preceding clauses, wherein the error concealment function is a first error concealment function from a plurality of error concealment functions. Clause 7. The one or more processors according to any of the preceding clauses, wherein the one or more circuits are used for: determining that an object is detected in a predetermined number of prior frames in the video stream; and Selecting the first error masking function from the plurality of error masking functions based at least on the object detected in the predetermined number of previous frames, wherein the first error masking function uses a larger amount of computing resources compared to a second error masking function from the plurality of error masking functions. Clause 8. The one or more processors according to any of the preceding clauses, wherein the one or more circuits are used for: generating the corrected frame by applying the first error concealment function to at least one section of the frame in which the object is estimated to appear. Clause 9. The one or more processors according to any of the preceding clauses, wherein the one or more circuits are used for the following: Determine that the damaged or lost data is located outside the region of interest; and Selecting the first error cover function from the multitude of error cover functions based at least on the damaged or lost data located outside the region of interest, wherein the first error cover function uses a smaller amount of computing resources compared to a second error cover function from the multitude of error cover functions. Clause 10. The one or more processors according to any of the preceding clauses, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for carrying out simulation processes; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning processes; a system implemented using an edge device; a system that is implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a Visual Language Model (VLM); a system for generating synthetic data; a system that embodies one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 11. System, which includes the following: one or more processors for the following: Receiving a request to process a video stream comprising a large number of individual frames; Applying an error concealment function to at least one frame from the plurality of frames, wherein the error concealment function is selected at least on the basis of a position of damaged or lost data in the at least one frame and a region of interest in which the at least one frame is selected; and Executing a machine learning model that uses at least one single image as input for the request. Clause 12. System according to Clause 11, wherein the one or more processors are used for the following: Determine that at least one single image contains damaged or lost data by using a decoding process. Clause 13. System according to Clause 11 or Clause 12, wherein the machine learning model generates a clue to an object in which at least one frame is generated, and wherein the one or more processors are used for the following: Selecting a second error masking function for at least one second frame from the multitude of frames, based at least on the clue of the object in the at least one frame. Clause 14. The system according to one of clauses 11-13, wherein the second error concealment function for the at least one second frame is further selected on the basis of an expected position of the object in the at least one second frame. Clause 15. The system according to any of Clauses 11-14, wherein the position of the damage or data loss is a macroblock position or a slice position of the at least one single image. Clause 16. System according to any of Clauses 11-15, wherein the one or more processors are used for the following: Applying the error correction function by executing a second machine learning model using the at least one single image as input, wherein the second machine learning model generates replacement information for the damaged or lost data of the at least one single image. Clause 17. The one or more processors according to any of Clauses 11-16, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for carrying out simulation processes; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning processes; a system implemented using an edge device; a system that is implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a Visual Language Model (VLM); a system for generating synthetic data; a system that embodies one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. Clause 18. Procedure, which includes the following: Identifying a single frame of a video stream using one or more processors; Determine, using one or more processors, that the single image contains corrupted or lost data; and Generating a corrected frame using one or more processors by applying an error concealment function selected based on at least a position of the damaged or lost data in the frame and a region of interest in the frame. Clause 19. Procedure according to Clause 18, which further includes the following: Selecting using one or more processors of the fault concealment function, at least based on the position of the damaged or lost data located within the region of interest in the single image. Clause 20. Procedure according to Clause 18 or Clause 19, which further includes the following: Generating the corrected single image using one or more processors by applying the error concealment function to at least one section of the region of interest.
[0111] Disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract types of data. Disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. Disclosure can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected to each other via a network for communication.
[0112] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0113] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms "step" and / or "block" may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.
[0114] It is understood that the aspects and embodiments described above are only exemplary and that changes to details may be made within the scope of protection of the claims.
[0115] Each device, method and feature disclosed in the description and (where applicable) in the claims and drawings may be provided independently or in any suitable combination.
[0116] The reference numerals appearing in the claims are for illustrative purposes only and are not intended to have any limiting effect on the scope of protection of the claims.
Claims
[1] One or more processors comprising the following: one or more circuits for the following: Identifying a single frame of a video stream; Determine that the single image contains damaged or lost data; and Generating a corrected single image by adding a applies an error concealment function that is selected based at least on the position of the damaged or lost data in the single image and a region of interest in the single image. [2] The one or more processors according to claim 1, wherein the one or more circuits shall select the error concealment function at least on the basis of the position of the damaged or lost data located within the region of interest in the single image. [3] The one or more processors according to claim 1 or claim 2, wherein the one or more circuits are to produce the corrected single image by applying the error concealment function to at least one section of the region of interest. [4] The one or more processors according to any of the preceding claims, wherein the one or more circuits shall apply the error concealment function by providing the single image as input for a machine learning model. [5] The one or more processors according to any of the preceding claims, wherein the one or more circuits serve to: Receiving an encoded bitstream of the video stream; and Generating the single image by decoding the encoded bitstream, where decoding the single image indicates the position of the damaged or lost data. [6] The one or more processors according to any of the preceding claims, wherein the error concealment function is a first error concealment function from a plurality of error concealment functions. [7] The one or more processors according to claim 6, wherein the one or more circuits serve to: Determine that an object is detected in a predetermined number of previous frames in the video stream; and Selecting the first error masking function from the plurality of error masking functions based at least on the object detected in the predetermined number of previous frames, wherein the first error masking function uses a larger amount of computing resources compared to a second error masking function from the plurality of error masking functions. [8] The one or more processors according to claim 7, wherein the one or more circuits serve to: Generating the corrected single image by applying the first error concealment function to at least one section of the single image in which the object is estimated to appear. [9] The one or more processors according to any one of claims 6-8, wherein the one or more circuits serve to: Determine that the damaged or lost data is located outside the region of interest; and Selecting the first error cover function from the multitude of error cover functions based at least on the damaged or lost data located outside the region of interest, wherein the first error cover function uses a smaller amount of computing resources compared to a second error cover function from the multitude of error cover functions. [10] The one or more processors according to any of the preceding claims, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for carrying out simulation processes; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning processes; a system implemented using an edge device; a system that is implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a Visual Language Model (VLM); a system for generating synthetic data; a system that embodies one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [11] System comprising the following: one or more processors for the following: Receiving a request to process a video stream comprising a large number of individual frames; Applying an error concealment function to at least one frame from the plurality of frames, wherein the error concealment function is selected at least on the basis of a position of damaged or lost data in the at least one frame and a region of interest in which the at least one frame is selected; and Executing a machine learning model that uses at least one single image as input for the request. [12] System according to claim 11, wherein the one or more processors are used for the following: Determine that at least one single image contains damaged or lost data by using a decoding process. [13] System according to claim 11 or claim 12, wherein the machine learning model generates a reference to an object in the at least one single image, and wherein the one or more processors serve to: Selecting a second error masking function for at least one second frame from the multitude of frames, based at least on the clue of the object in the at least one frame. [14] System according to claim 13, wherein the second error concealment function for the at least one second frame is further selected on the basis of an expected position of the object in the at least one second frame. [15] System according to one of claims 11-14, wherein the position of the damage or data loss is a macroblock position or a slice position of the at least one single image. [16] System according to any one of claims 11-15, wherein the one or more processors are used for the following: Applying the error correction function by running a second machine learning model using the at least one single image as input, wherein the second machine learning model generates replacement information for the damaged or lost data of the at least one single image. [17] The one or more processors according to any one of claims 11-16, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for carrying out simulation processes; a system for performing digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing deep learning processes; a system implemented using an edge device; a system that is implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a Visual Language Model (VLM); a system for generating synthetic data; a system that embodies one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [18] Method comprising the following: Identifying a single frame of a video stream using one or more processors; Determine, using one or more processors, that the single image contains corrupted or lost data; and Generating a corrected frame using one or more processors by applying an error concealment function selected based on at least a position of the damaged or lost data in the frame and a region of interest in the frame. [19] The method of claim 18, further comprising: Selecting using one or more processors of the error concealment function, at least based on the position of the damaged or lost data located within the region of interest in the single image. [20] The method of claim 18 or 19, further comprising: Generating the corrected single image using one or more processors by applying the error concealment function to at least one section of the region of interest.