Methods, apparatuses, and media for network-based media processing
Patent Information
- Application Number
- CN202180005947.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-19
- Filing Date
- 2021-05-06
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2041-05-06
AI Technical Summary
然而,当前的NBMP设计存在如下技术问题:缺乏衡量在不同网络实体、媒体处理实体(Media Processing Entity,MPE)、源端或接收端(sink)之间所划分的工作流的质量的能力
[0012] To address one or more different technical problems, this disclosure provides, according to exemplary embodiments, technical solutions for reducing network overhead and server computational overhead while updating and delivering immersive video for one or more viewport boundaries.
Smart Images

Figure CN114600083B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 043,660, filed June 24, 2020; U.S. Provisional Application No. 63 / 087,755, filed October 5, 2020; and U.S. Application No. 17 / 233,788, filed April 19, 2021, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to associating a Network Based Media Processing (NBMP) workflow with its functional inputs and functional outputs, and signaling to notify cloud workflows of input data timeouts or input data completeness. Background Technology
[0004] The MPEG NBMP project developed a concept for processing media in the cloud. However, the current NBMP design suffers from the following technical problem: a lack of ability to measure the quality of workflows divided among different network entities, media processing entities (MPEs), and source or sink ends.
[0005] The NBMP international specification draft demonstrates the enormous potential to improve media processing efficiency, deploy media services faster and at lower costs, and provide the ability to deliver large-scale deployments by leveraging public, private, or hybrid cloud services.
[0006] Even though the NBMP specification defines idle and running states for multiple tasks and workflows, it lacks details on how these changes occur. Therefore, the following technical problems exist: a lack of defined states for each input to a task or workflow, and a lack of identification of any conditions used to change their state. Consequently, a task or workflow cannot determine whether the data in its input is complete, and there are no rules for tasks or workflows to determine when data should stop.
[0007] Furthermore, even though various applications can be run using networks and cloud platforms, the NBMP standard defines workflow descriptions to define the required processing, without providing a one-to-one mapping of workflow inputs and outputs to their functional inputs and outputs.
[0008] While current NBMP workflow descriptions provide a detailed description of the workflow (including possible Directed Acyclic Graphs, DAGs), such descriptions lack the association between workflow input and output descriptions and individual functional inputs or outputs.
[0009] Furthermore, it may be unclear which input of the workflow is associated with the input of a specific function instance, and which output of the workflow is associated with the output of a specific function instance. This is because the connection mapping only defines the connections between functions, not the connections between function instances and the workflow's inputs / outputs.
[0010] If the inputs / outputs of a workflow are distinct across all first or last functions, the workflow manager can find the correlation due to the uniqueness of each input / output. However, if a workflow includes multiple inputs / outputs with the same description (different only in their flow IDs), then identifying the correct function input / output becomes ambiguous for the workflow's inputs / outputs.
[0011] In other words, among the other technical issues mentioned above, there is also ambiguity in assigning workflow inputs and workflow outputs to specific functional inputs and functional outputs. Summary of the Invention
[0012] To address one or more different technical problems, this disclosure provides, according to exemplary embodiments, technical solutions for reducing network overhead and server computational overhead while updating and delivering immersive video for one or more viewport boundaries.
[0013] This disclosure includes a method and apparatus comprising a memory and one or more processors, the memory being configured to store computer program code, the processors being configured to access the computer program code and operate according to the instructions of the computer program code. The computer program code includes acquisition code, setting code, and determination code; the acquisition code is configured to cause the at least one processor to acquire input to at least one of a task and a workflow in an NBMP; the setting code is configured to cause the at least one processor to set a timeout for the input to at least one of the task and the workflow; the determination code is configured to cause the at least one processor to determine whether at least one of the task and the workflow has observed missing data in the input for a duration equal to the timeout, such that the determination code is further configured to cause the at least one processor to determine that other data of the input is unavailable in response to determining that at least one of the task and the workflow has observed missing data in the input for a duration equal to the timeout; and the computer program code also includes application code and processing code, the application code being configured to cause the at least one processor to update at least one application of the task and the workflow in the NBMP based on determining that other data of the input is unavailable, and the processing code being configured to cause the at least one processor to process at least one of the tasks and the workflow in the NBMP based on the update.
[0014] According to an exemplary embodiment, the determining code is further configured to cause at least one processor to determine whether the state of at least one of the tasks and workflows is set to a running state, and the determining code is further configured to cause at least one processor, when determining that the state is set to a running state, to determine whether to change the state from a running state to an idle state based on whether at least one of the tasks and workflows observes missing input data for a duration equal to the timeout.
[0015] According to an exemplary embodiment, the determining code is further configured to cause at least one processor to determine whether to associate an input of at least one of the tasks and workflows with at least one of the functional inputs and functional outputs, and the computer program code further includes association code configured to cause at least one processor to associate an input of at least one of the tasks and workflows with at least one of the functional inputs and functional outputs in response to determining that an input of at least one of the tasks and workflows is associated with at least one of the functional inputs and functional outputs.
[0016] According to an exemplary embodiment, the determining code is further configured to cause at least one processor to determine whether the input includes an indication, and the determining code is also configured to cause at least one processor to determine that other data of the input is unavailable in response to determining that the input includes an indication.
[0017] According to an exemplary embodiment, the indication includes: a complete input flag included in the input of at least one of the tasks and workflows.
[0018] According to an exemplary embodiment, the instruction is included in metadata provided along with the input of at least one of the tasks and workflows.
[0019] According to an exemplary embodiment, the acquisition code is further configured to cause at least one processor to acquire a plurality of inputs, including at least one of the inputs of a task and a workflow; the determination code is further configured to cause at least one processor to determine whether the state of at least one of the tasks and workflows is set to a running state; the determination code is further configured to cause at least one processor to determine whether all inputs of at least one of the tasks and workflows have observed missing data for all inputs for a duration equal to the timeout, and the determination code is further configured to cause at least one processor, if it is determined that the state is set to a running state, to determine whether to change the state from a running state to an idle state based on whether at least one of the tasks and workflows has observed missing data for all inputs for a duration equal to the timeout.
[0020] According to an exemplary embodiment, determining whether all inputs of at least one of the tasks and workflows are found to have missing data for a duration equal to the timeout includes: determining whether at least one of the inputs of at least one of the tasks and workflows indicates either a timeout indication or a completeness indication.
[0021] According to an exemplary embodiment, the determining code is further configured to cause at least one processor to determine whether each of the inputs of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs, and the computer program code further includes association code configured to: cause at least one processor to associate the inputs of at least one of the tasks and workflows with at least one of one or more functional inputs and functional outputs, respectively, based on the determination that each of the inputs of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs.
[0022] According to an exemplary embodiment, determining whether each of the inputs of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs includes: determining at least one of an identifier and a port name for each of the inputs of at least one of the tasks and workflows, and determining at least one of a corresponding port name and a flow identifier for at least one of a corresponding input port and a corresponding output port. Attached Figure Description
[0023] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 This is a simplified schematic diagram according to an embodiment.
[0024] Figure 2 This is a simplified schematic diagram according to an embodiment.
[0025] Figure 3 This is a simplified block diagram of the decoder according to an embodiment.
[0026] Figure 4 This is a simplified block diagram of an encoder according to an embodiment.
[0027] Figure 5 This is a simplified block diagram of an encoder according to an embodiment.
[0028] Figure 6 This is a simplified state diagram of the encoder according to an embodiment.
[0029] Figure 7 This is a simplified block diagram of the image according to an embodiment.
[0030] Figure 8 This is a simplified block diagram of an encoder according to an embodiment.
[0031] Figure 9 This is a simplified flowchart based on an embodiment.
[0032] Figure 10 This is a schematic diagram according to an embodiment. Detailed Implementation
[0033] The suggested features discussed below can be used individually or in any combination in any order. Furthermore, these embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0034] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may encode video data locally for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the encoded video data from the other terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission is common in applications such as media services.
[0035] Figure 1A second pair of terminals 101 and 104 is shown configured to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each terminal 101 and 104 can encode locally acquired video data for transmission to the other terminal via network 105. Each terminal 101 and 104 can also receive encoded video data sent by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0036] exist Figure 1 In this disclosure, terminals 101, 102, 103, and 104 may be represented as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure can be applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 refers to any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between terminals 101, 102, 103, and 104. Communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 105 may be irrelevant to the operation of this disclosure.
[0037] Figure 2 The placement of a video encoder and video decoder in a streaming environment is illustrated as an example of the application of the disclosed subject matter. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0038] The streaming system may include an acquisition subsystem 203, which may include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213. This sample stream 213 is characterized as a high-data-volume video sample stream compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video bitstream 204 is characterized as a low-data-volume encoded video bitstream compared to the sample stream and is stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. Client 212 may include video decoder 211 that decodes an incoming copy 208 of the encoded video bitstream and produces an output video sample stream 210 that can be displayed on display 209 or another display device (not depicted). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Embodiments of these standards have been described above and are further described herein.
[0039] Figure 3 This may be a functional block diagram of a video decoder 300 according to an embodiment of the present disclosure.
[0040] Receiver 302 may receive one or more encoded video sequences to be decoded by decoder 300; in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Encoded video sequences may be received from channel 301, which may be a hardware / software link to a storage device storing the encoded video data. Receiver 302 may receive encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). Receiver 302 may separate the encoded video sequences from other data. To prevent network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / resolver 304 (hereinafter referred to as the "resolver"). Buffer 303 may not be required or may be made smaller when receiver 302 receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network. To maximize usability on business packet networks such as the Internet, a buffer 303 may also be required, which can be relatively large and advantageously have an adaptive size.
[0041] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-encoded video sequence. These symbols may include information for managing the operation of the decoder 300, and potential information for controlling a display device 312 (e.g., a display), which is not part of the decoder but may be coupled to it. Control information for the display device may be fragments (not indicated) of parameter sets of Supplemental Enhancement Information (SEI messages) or Video Usability Information (VUI). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract a subset of parameters from the encoded video sequence for at least one subset of pixels in a subgroup for use in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The entropy decoder / parser can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0042] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 can receive encoded data and selectively decode specific symbols 313. In addition, the parser 304 can determine whether to provide specific symbols 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra-frame prediction unit 307, or the loop filter 311.
[0043] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 313 may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser 304 through subgroup control information parsed from the encoded video sequence. For brevity, the flow of such subgroup control information between the parser 304 and the various units described below is not described.
[0044] In addition to the functional blocks already mentioned, the decoder 300 can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.
[0045] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives from the parser 304 the quantization transform coefficients as symbols 313, as well as control information, including the transform mode used, block size, quantization factor, and quantization scaling matrix. The scaler / inverse transform unit can output a block containing sample values, which can be input into the aggregator 310.
[0046] In some cases, the output samples of the scaler / inverse transform unit 305 may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses reconstructed information extracted from the current (partially reconstructed) image 309 to generate surrounding blocks of the same size and shape as the block being reconstructed. In some cases, the aggregator 310 adds the predictive information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305 based on each sample.
[0047] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit 306 can access the reference image memory 308 to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbol 313 belonging to the block, these samples can be added by the aggregator 310 to the output of the scaler / inverse transform unit (referred to in this case as residual samples or residual signals) to generate output sample information. The motion compensation unit's retrieval of predicted samples from the address in the reference image memory can be controlled by motion vectors, and these motion vectors are available to the motion compensation unit in the form of the symbol 313, which, for example, includes X, Y, and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0048] The output samples of aggregator 310 can be employed by various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream, and these parameters can be used as symbols 313 from parser 304 in loop filter unit 311. However, video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0049] The output of the loop filter unit 311 can be a sample stream, which can be output to the display device 312 and stored in the reference image memory 557 for subsequent inter-frame image prediction.
[0050] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. Once the encoded images have been fully reconstructed and are identified as reference images (e.g., by parser 304), the current reference image 309 can become part of the reference image buffer 308, and new current image memory can be reallocated before the reconstruction of subsequent encoded images begins.
[0051] The video decoder 300 can perform decoding operations according to predetermined video compression techniques, such as those specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard (as specified in the video compression technique document or standard, particularly in its configuration file document). For compliance, the complexity of the encoded video sequence may also be required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0052] In this embodiment, receiver 302 may receive additional (redundant) data along with the encoded video. This additional data may be part of the encoded video sequence. The additional data may be used by video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0053] Figure 4 This is a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0054] Encoder 400 can receive video samples from video source 401 (which is not part of the encoder), which can capture video images that will be encoded by encoder 400.
[0055] Video source 401 can provide a source video sequence in the form of a digital video sample stream to be encoded by encoder 303. This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, video source 401 can be a storage device storing previously prepared video. In a video conferencing system, video source 401 can include a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed sequentially. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following focuses on describing samples.
[0056] According to an embodiment, encoder 400 can encode and compress images of a source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of controller 402. The controller controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily recognize other functions of controller 402 that may relate to a video encoder 400 optimized for a particular system design.
[0057] Some video encoders operate within a loop that is readily recognized by those skilled in the art as an "encoding loop." In a simplified description, an encoding loop may consist of an encoding portion including an encoder 402 (hereinafter referred to as the "source encoder") (e.g., responsible for creating symbols based on the input image to be encoded and a reference image) and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data in the same manner as the (remote) decoder creates sample data (since any compression between the symbols and the encoded video stream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream is input to a reference image memory 405. Since the decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image buffer are also bit-precisely corresponding between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values that the decoder will "see" during the prediction process. Those skilled in the art are familiar with this basic principle of reference image synchronization (and the drift that occurs, for example, due to channel errors, when synchronization cannot be maintained).
[0058] The operation of the “local” decoder 406 can be combined with, for example, the above. Figure 3 The operation is the same as that of the "remote" decoder 300 described in detail. However, please refer to another brief reference. Figure 4 When symbols are available and the entropy encoder 408 and the parser 304 are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding part of the decoder 300 (including the channel 301, receiver 302, buffer 303 and parser 304) may not be fully implemented in the local decoder 406.
[0059] It can then be observed that any decoder technique other than parsing / entropy decoding, which exists in the decoder, must also exist in the corresponding encoder in essentially the same functional form. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.
[0060] As part of its operation, the source encoder 403 can perform motion-compensated predictive coding, referencing one or more previously encoded frames from the video sequence designated as "reference frames," to predictively encode the input frame. In this way, the encoding engine 407 encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame.
[0061] The local video decoder 406 can decode encoded video data of frames that can be designated as reference frames, based on symbols created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When encoded video data can be decoded by the video decoder (… Figure 4 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process, which can be performed by the video decoder on the reference frame, and allows the reconstructed reference frame to be stored in the reference image cache 405. In this way, the encoder 400 can locally store a copy of the reconstructed reference frame that shares the same content (no transmission errors) as the reconstructed reference frame to be obtained by the remote video decoder.
[0062] Predictor 404 can perform a prediction search against encoding engine 407. That is, for a new frame to be encoded, predictor 404 can search the reference image memory 405 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. Predictor 404 can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by predictor 404, it can be determined that the input image can have prediction references obtained from multiple reference images stored in reference image memory 405.
[0063] The controller 402 can manage the encoding operations of the video encoder 403, including, for example, setting parameters and subgroup parameters for encoding video data.
[0064] The outputs of all the aforementioned functional units can be entropy encoded in the entropy encoder 408. The entropy encoder can transform the symbols generated by various functional units into an encoded video sequence by lossless compression of the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding.
[0065] Transmitter 409 may buffer the encoded video sequence created by entropy encoder 408 in preparation for transmission via communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. Transmitter 409 may combine the encoded video data from video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0066] Controller 402 manages the operation of encoder 400. During encoding, controller 405 can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types: An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.
[0067] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values for each block.
[0068] Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0069] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined by the coding assignment of the corresponding images applied to the block. For example, a block of an I-image can be non-predictively coded, or it can be predictively coded with reference to already coded blocks of the same image (spatial prediction or intra-frame prediction). A pixel block of a P-image can be predictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. A block of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.
[0070] The video encoder 400 can perform encoding operations according to predetermined video coding techniques or standards, such as those specified in ITU-T H.265 Recommendation. In operation, the video encoder 400 can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0071] In one embodiment, transmitter 409 may transmit additional data while transmitting encoded video. Source encoder 403 may include such data as part of the encoded video sequence. Additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, fragments of visual usability information (VUI) parameter sets, etc.
[0072] Figure 5A network-based media processing MPEG architecture 500 according to embodiments herein is illustrated and can be implemented relative to cloud processing, enabling the determination and utilization of the quality of workflows segmented across different network entities, MPEs, source ends, or receiver ends, thereby utilizing public cloud services, private cloud services, or hybrid cloud services as described below. The NBMP architecture 500 includes: an NBMP source 501 that provides an NBMP workflow API and workflow descriptions 511 to an NBMP workflow manager 513. The NBMP workflow manager 513 can communicate with a feature repository 503 via a feature discovery API and feature descriptions 512, and the feature repository 503 can also communicate with the NBMP source 501 via a feature discovery API and feature descriptions 502.
[0073] In addition, Figure 5 In this context, the NBMP workflow manager and the Media Processing Entity (MPE) 505 can communicate, for example, via the NBMP task API and by reporting current task status and configuration. The MPE can operate as a runtime configuration / stream / event binding entity. For instance, a media source 504 can provide one or more media streams 514 to one or more tasks 521 and 522 of the MPE 505, and these tasks can also parallelize multiple media streams 514. The MPE 505 can then transmit one or more media streams 515 to a media receiver 506.
[0074] In addition, Figure 5 In this context, the communication between NBMP source 501, NBMP workflow manager 513, function repository 503, and MPE 505 can be considered as a control flow, while the communication between media source 504, MPE 505, and media receiver 506 can be considered as a data flow.
[0075] According to an exemplary embodiment, the inputs and outputs of workflow description 511 are described by input descriptors and output descriptors, and according to an exemplary embodiment, interconnections between functional instances of the workflow defined in the connection-maparray of the processing descriptors are possible, such that the corresponding objects are shown in Tables 1 to 3 below: Table 1 - Input Descriptors
[0076] Table 2 - Output Descriptors
[0077] Table 3 - Elements of the Connectivity Mapping Array
[0078] According to the exemplary embodiment, the objects from and to are defined in Table 4: Table 4 - "From" and "To" Objects
[0079] Figure 6 A task lifecycle state diagram 600 according to an exemplary embodiment is shown. For example, the onInstantiation signal 611 may be received in the instantiation state 601. The instantiation state 601 may transition to the idle state 602 via the onTaskConfiguration signal 612, and the idle state 602 may transition back to the instantiation state 601 via the onReset signal 617. The idle state 602 may also cycle via the onTaskConfiguration signal 613, and may also transition to the running state 603 via the onStart signal 614. The running state 603 may cycle via the onTaskConfiguration signal 615, and may transition back to the idle state 602 via the onStop signal and / or the onCompletion signal 616. The running state 603 may also transition to the error state 604 via the onError signal 623, and the error state 604 may transition to the idle state 602, the instantiation state 601, and the corruption state 605 via the onErrorHandling signal 624, the onReset signal 625, and the onTermination signal 626, respectively. In addition, the running state 603, the idle state 602 and the instantiated state 601 can be transitioned to the destroyed state 605 by the onTermination signal 621, the onTermination signal 622 and the onTermination signal 611, respectively.
[0080] According to an exemplary embodiment, onStop and onCompletion signals may also be added. For example, when media data or metadata stops arriving at MPE 505, the onStop signal may indicate, for example, that task 521 and / or task 522 should transition its state from running state 603 to idle state 602. Furthermore, for example, when processing completes, such as at MPE 505, the onCompletion signal may indicate that task 521 and / or task 522 transition its state from running state 603 to idle state 602. According to an exemplary embodiment, these features are not only presented for MPE 505, but at least those state transitions (from running state 603 to / from idle state 602) may also be provided for workflow manager 513.
[0081] According to an exemplary embodiment, timeout parameters (e.g., by exemplarily adding a timeout to the input descriptor) can be used to provide and enhance NBMP input descriptors and parameters, as described in Tables 5 through 7 below: Table 5 - Input Media Parameter Objects
[0082] Table 6 - Input Metadata Parameter Object
[0083] According to the embodiments, timeout parameters can be defined as shown in Table 7: Table 7 - Input Timeout Parameters
[0084] Figure 7 An exemplary embodiment is shown, illustrating a workflow descriptor 700 with input i1 and multiple outputs o1, o2, o3, and o4; however, this number of inputs and outputs may be varied according to embodiments. Similarly, Figure 8 An exemplary embodiment of task diagram 800 is shown, which has task T0, wherein input I1 is split into multiple tasks T1, T2, T3 and another T3, and the multiple tasks have their own outputs t1, t2, t3 and t4, but this number of inputs and outputs can be varied according to embodiments. However, looking at workflow descriptor 700 and task diagram 800, there is a technical problem regarding the lack of description of the association between (o1, o2, o3, o4) and (t1, t2, t3, t4). Therefore, embodiments may refer to and rely on the information provided in Tables 8 and 9 below: Table 8 - Elements of the Link Mapping Array
[0085] According to the embodiment, the objects from and to are defined in Table 9, and may replace the definitions previously provided in Table 3.
[0086] Table 9 - Objects “From” and “To”
[0087] Figure 9 A simplified flowchart 900 according to an exemplary embodiment is shown. In S901, input information and output information are received.
[0088] In S902, instances are retrieved. For example, one or more inputs and / or outputs of any one or more tasks and workflows can be retrieved.
[0089] In S903, a mapping is established. It is possible to define the mapping regarding... Figure 7 and Figure 8 The example described illustrates the connections between functional instances; however, as a technical issue, it may be unclear which input of the workflow is associated with the input of a specific functional instance, and which output of the workflow is associated with the output of a specific functional instance. This may be because the connection mapping may only define connections between functions, not connections between functional instances and workflow inputs / outputs. If the workflow inputs / outputs are distinct across all first or last functions, the workflow manager 513 can find associations based on the uniqueness of each input / output across all inputs / outputs. However, if the workflow includes multiple inputs / outputs with the same description (different only in their flow IDs), then identifying the correct functional input / output becomes ambiguous for the workflow inputs / outputs. In other words, there will be ambiguity in assigning workflow inputs and workflow outputs to specific functional inputs and functional outputs. An example of this situation is... Figure 7 and 8 As shown in the figure, the technical solution provided by the present invention also solves this problem.
[0090] For example, in S903, the embodiment includes the workflow's inputs / outputs in the connection mapping; that is, the connection mapping should include all inputs and all outputs of the workflow. For example, a port name can be used as a function ID and a flow ID can be used as an instance. Furthermore, for example, the following text may be included in Table 8: The array of connection-mapping objects describes the media workflow DAG, that is, the connection information between different tasks in the diagram. Each element in this array represents an edge in the DAG, and each element is defined in Table 8, which can replace the definitions in Table 3.
[0091] Furthermore, according to the exemplary embodiment, Table 8 should also include all input and output ports of the General Descriptor, in the form of "from" objects and "to" objects (e.g., one "from" object for each input port and one "to" object for each output port), and for example, the "ID" and "port name" of the input / output of each corresponding workflow should be set to the "port name" and "flow ID" corresponding to the corresponding input / output port of the General Descriptor. It should be noted that for such connections, flow control parameters, colocation flags, and other parameters may be ignored, and these parameters are not included in this exemplary embodiment.
[0092] Thus, in S903, determining whether to associate an input of at least one of the tasks and workflows with at least one of the functional inputs and functional outputs involves, in response to determining that an input of at least one of the tasks and workflows is associated with at least one of the functional inputs and functional outputs, associating an input of at least one of the tasks and workflows with at least one of the functional inputs and functional outputs. This can also involve acquiring multiple inputs (including inputs of at least one of the tasks and workflows) such that determining whether each input of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs; and, based on determining whether each input of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs, associating multiple inputs of at least one of the tasks and workflows with at least one of one or more functional inputs and functional outputs, respectively. For example, according to an exemplary embodiment, in this determination process, determining whether each input of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs includes: determining at least one of the identifier and port name of each input of at least one of the tasks and workflows, and determining at least one of the corresponding port name and flow identifier of at least one of the corresponding input port and corresponding output port.
[0093] In S904, settings are made, such as setting a timeout for any input in at least one of the tasks and workflows. The features in S904 and S903 can be implemented in parallel or sequentially, and therefore, this document discloses a method for associating each input and each output of a workflow to each associated individual functional input and functional output, wherein each input or each output of the workflow is connected to one of the inputs or outputs of a functional instance / task of the workflow using a connection map, wherein each input and each output can be explicitly identified, embodying a technical solution to the aforementioned technical problem.
[0094] In S905, operations such as running one or more tasks are performed. However, as an additional technical issue, a task or workflow may not be able to determine whether the data in its input is complete. Therefore, statements in the NBMP specification such as "onStop, when media data or metadata stops arriving, the task should transition its state from the running state to the idle state" may be technically inadequate because the task or workflow may lack any rules for determining whether data has stopped.
[0095] Therefore, according to the exemplary embodiment, there is a technical solution to the problem of task or workflow implementation when input stops arriving, by defining two additional parameters for each input: "Timeout": a time interval that indicates whether the input data is considered complete if no data is received in the input, and "Complete": an input flag that, if set to "true", means that no further data has arrived in the input and the input data can be considered complete. Thus, in S904, the timeout parameter can be set by the NBMP source for workflow inputs, or by the workflow manager for task inputs. The workflow or task can then observe the input, and once no data arrives within the duration of the Timeout, it can be concluded that data has stopped arriving.
[0096] In S906, it is considered whether one or more flags have been received. If not, the process can continue checking such flags. Therefore, in S906, it is possible to consider whether a flag exists for input in the case of complete input. When the media source or connection task has no data, this flag is set to "true" so that the workflow or task knows it will not receive any additional data for that input. Therefore, in Table 6 above, the technical advantages of the added NBMP input descriptors and parameters are enhanced by using timeout parameters. Thus, a complete flag for input can exist, so that for each media or metadata input, each function can define an input parameter, and the parameter "complete flag" can be defined as indicating that if it is set to "true," the input is complete, i.e., no additional data has been received at that input. In S907, it is considered whether an idle state exists. If no idle state exists, the process can continue checking for this state. In S908, it is considered whether to change the state. If not, the process can continue checking whether to change the state. In S909, it is considered whether to terminate the process. If not, the process can continue running. In S910, the processing ends, and then proceeds to S901 to await further reception of any of the multiple inputs and multiple outputs. Therefore, during processing, it is determined whether at least one of the tasks and workflows has observed missing input data for a duration equal to the timeout; in response to determining that at least one of the tasks and workflows has observed missing input data for a duration equal to the timeout, it is determined that other input data is unavailable; it is determined whether the state of at least one of the tasks and workflows is set to a running state; and if it is determined that the state is set to a running state, it is determined whether to change the state from the running state to an idle state based on whether at least one of the tasks and workflows has observed missing input data for a duration equal to the timeout. It can also be determined whether the input includes an indication, and in response to determining that the input includes an indication, it is determined that other input data is unavailable. Such an indication may include a complete input flag included with the input of at least one of the tasks and workflows, or it may be provided as metadata included with the input of at least one of the tasks and workflows.
[0097] In addition, for example Figure 6The state transitions between states can be further enhanced based on timeout parameters and completion flags determined at either point in S906 and S907, for example, at S908 considering whether the following exists: when all media data or metadata stops arriving (by observing timeout values or completion flags set to "true" in all inputs, or any combination of both, and completing the processing of the received inputs), "onStop" indicates that the state should transition from the running state to the idle state, etc.; and when the processing is complete (by observing timeout values or completion flags set to "true" in all inputs, or any combination of both, and completing the processing of the received inputs), "onCompletion" indicates that the state should transition from the running state to the idle state, etc. According to an exemplary embodiment, this representation can be applied to either tasks or workflows. Thus, regarding the acquisition of multiple inputs, including at least one of the task and workflow inputs, at S901, it is also possible to determine, in response to either S906 or S907 or S908, whether to set the state of at least one of the task and workflow to a running state, such as additionally determining whether missing data is observed in all inputs of at least one of the task and workflow (e.g., all inputs whose duration at S907 is equal to the timeout); and if it is determined at S908 that the state is set to a running state, it is determined whether to change the state from a running state to an idle state based on whether missing data is observed in at least one of the task and workflow inputs within a duration equal to the timeout of all inputs, based on the result of either S906 or S907. For example, by determining whether at least one of the inputs of at least one of the task and workflow indicates either a timeout indication at S907 and / or a complete indication at S906, it can be determined whether missing data is observed in all inputs of at least one of the task and workflow inputs within a duration equal to the timeout.
[0098] Therefore, these features offer technical advantages by providing the following methods: a method for setting a timeout for each input of a task or workflow, such that if the task or workflow does not observe any data in its corresponding input for a duration equal to the timeout, the task or workflow can infer that no more data is available for that input; a method for signaling no additional input data available for media inputs, metadata inputs, or any other inputs using a complete input flag, such that a data sending entity can notify the task or workflow that no data will be sent to that input, and thus infer that no more data will arrive for the task or workflow; and a method for changing the state of a task or workflow from a running state to an idle state based on the timeouts of all inputs and all complete flags of all inputs, and if all inputs of the task or workflow are in a timeout or complete state, then the task or workflow should change its state from a running state to an idle state, which would otherwise be unavailable, taking into account technical problems (e.g., at least for NBMP implementations), which are advantageously solved as described in the embodiments herein.
[0099] Therefore, based on any update operation described in S906 to S909 above, at S905, based on the determination that other input data is unavailable, an update can be applied to at least one of the tasks and workflows in NBMP so that the processing of at least one of the tasks and workflows in NBMP is implemented based on the update.
[0100] The techniques described above can be implemented as computer software using computer-readable instructions and can be physically stored in one or more computer-readable media or implemented by one or more hardware processors with a specific configuration. For example, Figure 10 A computer system 1000 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0101] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by the computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.
[0102] The instructions can be executed on various types of computers or their components, such as personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0103] Figure 10The components of the computer system 1000 shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependencies or requirements relating to any one or a combination of components shown in the exemplary embodiments of the computer system 1000.
[0104] Computer system 1000 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movement), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). Human-machine interface devices may also be used to acquire certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0105] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard 1001, mouse 1002, touchpad 1003, touch screen 1010, joystick 1005, microphone 1006, scanner 1008, and camera 1007.
[0106] Computer system 1000 may also include certain human-machine interface output devices. Such human-machine interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback of touch screen 1010 or joystick 1005, but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speaker 1009, headphones (not depicted)), visual output devices (e.g., screen 1010 including CRT screen, LCD screen, plasma screen, OLED screen, each screen may or may not have touch screen input functionality, each screen may or may not have tactile feedback functionality - some of these screens are capable of outputting two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0107] The computer system 1000 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD / RW 1020 with media such as CD / DVD 1011, thumb drives 1022, removable hard disk drives or solid-state drives 1023, conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0108] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0109] Computer system 1000 may also include an interface 1099 providing access to one or more communication networks 1098. Network 1098 may be, for example, a wireless network, a wired network, or an optical network. Network 1098 may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of network 1098 include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks 1098 typically require external network interface adapters (e.g., USB ports of computer system 1000) to connect to certain general-purpose data ports or peripheral buses (1050 and 1051); as described below, other network interfaces are typically integrated into the core of computer system 1000 by connecting to the system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). Computer system 1000 can communicate with other entities using any of these networks 1098. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as using a local area network (LAN) or wide area network (WAN) to connect to other computer systems. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.
[0110] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel 1040 of the computer system 1000.
[0111] The core 1040 may include one or more central processing units (CPUs) 1041, graphics processing units (GPUs) 1042, graphics adapters 1017, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 1043, hardware accelerators 1044 for certain tasks, etc. These devices, along with read-only memory (ROM) 1045, random access memory 1046, and internal mass storage 1047 such as internal non-user-accessible hard disk drives (SD drives), SSDs, etc., may be connected via a system bus 1048. In some computer systems, the system bus 1048 may be accessed via one or more physical connectors to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus 1048 or connected to the core's system bus 1048 via a peripheral bus 1051. Peripheral bus architectures include PCI, USB, etc.
[0112] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in read-only memory (ROM) 1045 or random access memory (RAM) 1046. Transient data can also be stored in RAM 1046, while permanent data can be stored, for example, in internal mass storage 1047. Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs 1041, GPUs 1042, mass storage 1047, ROM 1045, RAM 1046, etc.
[0113] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0114] As a non-limiting example, a computer system having architecture 1000, particularly a computer system with kernel 1040, can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software contained in one or more tangible computer-readable media. Such computer-readable media may be media associated with user-accessible mass storage as described above, and some non-transitory memory of kernel 1040, such as internal mass storage 1047 or ROM 1045. Software implementing the various embodiments of this disclosure may be stored in such devices and executed by kernel 1040. Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software may cause kernel 1040, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in random access memory RAM 1046 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system 1000 may be functionalized by logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1044), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0115] While several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
Claims
1. A method for network-based media processing (NBMP), the method comprising: Obtain input from at least one of the tasks and workflows in the NBMP; Determine at least one of the identifier and port name of each input of at least one of the inputs of the task and the workflow, and determine at least one of the corresponding port name and flow identifier of at least one of the corresponding input port and the corresponding output port; Based on determining whether each of the inputs of at least one of the tasks and workflows is associated with at least one of one or more functional inputs and functional outputs, the inputs of at least one of the tasks and workflows are associated with at least one of one or more functional inputs and functional outputs, respectively. Set a timeout for the input of at least one of the tasks and the workflow; Determine whether at least one of the tasks and workflows is found to have missing data in the input within a duration equal to the timeout; In response to determining that at least one of the tasks and the workflow observes missing data in the input for a duration equal to the timeout, it is determined that other data in the input is unavailable; Based on the determination that other data for the input is unavailable, update the task in the NBMP and the at least one application in the workflow; as well as Based on the update, process at least one of the tasks in the NBMP and the workflow.
2. The method according to claim 1, further comprising: Determine whether the status of at least one of the tasks and workflows is set to a running state; as well as If the state is determined to be set to the running state, based on whether at least one of the tasks and the workflow observes missing data in the input data within the duration equal to the timeout, it is determined whether to change the state from the running state to the idle state.
3. The method according to claim 1, further comprising: Determine whether the input includes an instruction; and In response to determining that the input includes the indication, it is determined that other data for the input is unavailable.
4. The method according to claim 3, in, The indication includes: a complete input flag included in the input of at least one of the tasks and the workflow.
5. The method according to claim 3, in, The instruction is included in the metadata provided along with the input of at least one of the tasks and workflows.
6. The method according to claim 3, further comprising: Acquire multiple inputs, the multiple inputs including the inputs of at least one of the task and the workflow; Determine whether the status of at least one of the tasks and workflows is set to a running state; Determine whether missing data is observed in all inputs of at least one of the tasks and workflows within a duration equal to the timeout; as well as If the state is determined to be set to the running state, based on whether at least one of the tasks and workflows observes missing data for all inputs for a duration equal to the timeout, it is determined whether to change the state from the running state to the idle state.
7. The method according to claim 6, in, Determining whether missing data is observed in all inputs of the task and the workflow within the duration equal to the timeout includes: determining whether at least one of the inputs of the task and the workflow indicates either a timeout indication or a completeness indication.
8. An apparatus for network-based media processing (NBMP), the apparatus comprising: At least one memory is configured to store computer program code; At least one processor is configured to access and operate in accordance with the instructions of the computer program code, the computer program code comprising: The code is configured to cause the at least one processor to receive input from at least one of the tasks and workflows in the NBMP; Setting code, configured to cause the at least one processor to set a timeout for the input of the task and at least one of the workflows; and The determination code is configured to cause the at least one processor to determine whether at least one of the tasks and workflows has observed missing data in the input for a duration equal to the timeout. The determining code is further configured to cause the at least one processor to determine that other data of the input is unavailable in response to determining that at least one of the task and the workflow has observed missing data in the input for a duration equal to the timeout. The determining code is further configured to cause at least one processor to determine at least one of the identifier and port name of each input of the at least one of the task and the workflow, and to determine at least one of the corresponding port name and stream identifier of at least one of the corresponding input port and the corresponding output port; The computer program code further includes association code, which is configured to associate the inputs of the task and the workflow with at least one of the one or more functional inputs and functional outputs, respectively, based on determining whether each input of the inputs of the at least one in the task and the workflow is associated with at least one of one or more functional inputs and functional outputs. The computer program code further includes application code configured to: based on the determination that other data from the input is unavailable, cause the at least one processor to update the task in the NBMP and the at least one application in the workflow; and The computer program code further includes processing code configured to, based on the update, cause the at least one processor to process the task in the NBMP and the at least one in the workflow.
9. The apparatus according to claim 8, in, The determining code is further configured to cause the at least one processor to determine whether the state of the task and at least one of the workflows is set to a running state, and The determining code is further configured to enable the at least one processor, upon determining that the state is set to the running state, to determine whether to change the state from the running state to the idle state based on whether the at least one of the tasks and the workflow observes missing data in the input data within the duration equal to the timeout.
10. The apparatus according to claim 8, in, The determining code is further configured to cause the at least one processor to determine whether the input includes an indication, and The determining code is further configured to, in response to determining that the input includes the indication, cause the at least one processor to determine that the other data of the input is unavailable.
11. The apparatus according to claim 10, in, The indication includes: a complete input flag included in the input of at least one of the tasks and the workflow.
12. The apparatus according to claim 10, in, The instruction is included in the metadata provided along with the input of at least one of the tasks and workflows.
13. The apparatus according to claim 10, in, The acquisition code is also configured to cause the at least one processor to acquire multiple inputs, the multiple inputs including the inputs of the task and the at least one of the workflows; The determining code is further configured to cause the at least one processor to determine whether the state of the task and the workflow at least one is set to a running state; The determining code is further configured to cause the at least one processor to determine whether all inputs of the task and the at least one in the workflow are found to have missing data for all inputs within a duration equal to the timeout. The determining code is further configured to enable the at least one processor, upon determining that the state is set to the running state, to determine whether to change the state from the running state to the idle state based on whether the at least one of the tasks and the workflow observes missing data in the inputs for all inputs within a duration equal to the timeout.
14. The apparatus according to claim 13, in, Determining whether missing data is observed in all inputs of the task and the workflow within the duration equal to the timeout includes: determining whether at least one of the inputs of the task and the workflow indicates either a timeout indication or a completeness indication.
15. A non-transitory computer-readable medium storing a program that causes a computer to perform a network-based media processing (NBMP) process, the process comprising the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Task timeouts based on input data characteristics
US9430280B1