Method, apparatus, and computer program for dividing and merging multi-dimensional media data into multi-dimensional media segments

By segmenting and processing multi-dimensional media streams into parallel sub-streams with metadata for ordering, the method addresses the limitations of current NBMP frameworks in handling multi-dimensional media, achieving efficient parallel processing and merging.

JP7697037B2Active Publication Date: 2025-06-23TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023560809
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-01
Filing Date
2022-12-13
Publication Date
2025-06-23
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

Current network-based media processing (NBMP) frameworks do not support multi-dimensional segment metadata, limiting their ability to efficiently split, process, and merge multi-dimensional media data in parallel.

Method used

The proposed method segments a multi-dimensional media stream into parallel sub-streams, each with segment metadata for ordering, allowing for parallel processing and subsequent merging into a single stream using the segment metadata.

Benefits of technology

This approach enables efficient parallel processing of multi-dimensional media data, ensuring that segments are properly ordered and merged, thereby enhancing the processing capabilities of NBMP frameworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697037000005
    Figure 0007697037000005
  • Figure 0007697037000006
    Figure 0007697037000006
  • Figure 0007697037000007
    Figure 0007697037000007
Patent Text Reader

Abstract

The method includes the steps of segmenting a multidimensional media stream into a plurality of segments of multidimensional media in a multidimensional space; splitting the segmented multidimensional media stream into a plurality of sub-streams that can be processed in parallel, each of the plurality of sub-streams including segment metadata used to order segments within each sub-stream; processing each of the plurality of sub-streams in parallel; and merging the plurality of sub-streams into a single stream using segment metadata carried in an output segment, the single stream including the ordered segments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims priority to U.S. Patent Application No. 63 / 298,536, filed on January 11, 2022, and U.S. Patent Application No. 18 / 073,048, filed on December 1, 2022, and the disclosures of these are hereby incorporated by reference in their entireties.

[0002] Technical field The present disclosure provides an extension to the one-dimensional NBMP splitter and merger functions to support splitting and merging multi-dimensional media data into segments that can each be processed independently.

Background Art

[0003] A network-based media processing (NBMP) framework defines an interface that includes both a data format and an application programming interface (API) between entities connected via a digital network for media processing. The NBMP standard defines a set of tools for the independent processing of one-dimensional media segments. This framework enables the dynamic creation of media processing pipelines and provides access to processed media data and metadata in real-time or a delayed manner. Networks and cloud platforms are used to run various applications. When media is multi-dimensional, it may be necessary to segment media data into multiple dimensions. If parallel processing is required, such media data needs to be split, processed in parallel paths, and merged again. Current NBMP splitter and merger functions do not support multi-dimensional segment metadata.

Summary of the Invention

[0004] The following presents a simplified summary of one or more embodiments of the present disclosure to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments, nor is it intended to identify key or critical elements of all embodiments or to delineate the scope of any or all embodiments. The sole purpose of this summary is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description that follows.

[0005] According to some embodiments, a method is provided that is executed by at least one processor. The method includes segmenting a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space. The method further includes dividing the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream. The method further includes processing each of the plurality of sub-streams in parallel and merging the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, where the single stream includes ordered segments.

[0006] According to some embodiments, the apparatus includes at least one memory configured to store program code and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes segmentation code configured to cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space. The program code further includes splitting code configured to cause the at least one processor to split the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order the segments within each sub-stream. The program code further includes processing code configured to cause the at least one processor to process each of the plurality of sub-streams in parallel. The program code further includes merge code configured to cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried to the output segments, where the single stream includes ordered segments.

[0007] According to some embodiments, when executed by at least one processor, a non-transitory computer-readable storage medium stores instructions that cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space. The instructions further cause the at least one processor to divide the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order the segments within each sub-stream. The instructions further cause the at least one processor to process each of the plurality of sub-streams in parallel using the segment metadata carried in the multi-dimensional media segments. The instructions further cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, where the single stream includes ordered segments.

[0008] Additional embodiments are described in the following description, and are apparent in part from the following description, and / or may be learned by practice of the presented embodiments of the disclosure.

Brief Description of the Drawings

[0009] The above and other aspects, features, and aspects of the embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings.

[0010]

Figure 1

[0011]

Figure 2

[0012]

Figure 3

[0013]

Figure 4

[0014]

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0015] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0016] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementation to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from practice of the implementation. Further, one or more features or components of one embodiment may be incorporated into another embodiment (or one or more features of another embodiment), or may be combined with another embodiment. Further, in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be (at least partially) executed simultaneously, and the order of one or more operations may be switched.

[0017] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Thus, although the operation and behavior of the systems and / or methods have been described herein without reference to specific software code, it is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.

[0018] Even if certain combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0019] Any element, act, or instruction used in this specification should not be construed as essential or indispensable unless explicitly stated otherwise. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more". When only one item is intended, the term "one" or similar language is used. Also, as used in this specification, terms such as "has", "have", "having", "include", "including", etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless explicitly stated otherwise. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0020] Throughout this specification, references to "one embodiment", "an embodiment", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the solution. Thus, references throughout this specification to "in one embodiment", "in an embodiment", and similar language may all refer to the same embodiment, but not necessarily so.

[0021] Furthermore, the features, advantages, and characteristics described in this disclosure may be combined in any suitable manner in one or more embodiments. One of ordinary skill in the art will recognize, in light of the description herein, that the disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of the disclosure may be recognized in a particular embodiment.

[0022] FIG. 1 is a diagram of an environment 100 in which the methods, apparatuses, and systems described herein according to an embodiment can be implemented. As shown in FIG. 1, the environment 100 may include a user device 110, a platform 120, and a network 130. The devices in the environment 100 may be interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0023] The user device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with the platform 120. For example, the user device 110 may include a computing device (such as a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (such as a smartphone, a wireless phone, etc.), a wearable device (such as a pair of smart glasses or a smartwatch), or a similar device. In some embodiments, the user device 110 may receive information from the platform 120 and / or transmit information to the platform 120.

[0024] The platform 120 includes one or more devices as described elsewhere herein. In some implementations, the platform 120 may include a cloud server or a group of cloud servers. In some implementations, the platform 120 may be designed modularly such that software components can be swapped in or out according to specific needs. Thus, the platform 120 can be easily and / or quickly reconfigured for different uses.

[0025] In some implementations, as shown, platform 120 can be hosted in a cloud computing environment 122. In particular, while the implementations described herein will be described as having platform 120 hosted within cloud computing environment 122, in some implementations, platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment), or may be partially cloud-based.

[0026] Cloud computing environment 122 includes an environment that hosts platform 120. Cloud computing environment 122 can provide services such as computing, software, data access, storage, etc., without requiring knowledge of the physical location and configuration of the system and / or device hosting platform 120 by an end user (e.g., user device 110). As shown, cloud computing environment 122 can include a group of computing resources 124 (collectively referred to as “(plural) computing resources 124” and individually referred to as “(singular) computing resource 124”).

[0027] Computing resources 124 can include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, computing resources 124 can host platform 120. Cloud resources can include compute instances operating within computing resources 124, storage devices provided within computing resources 124, data transfer devices provided by computing resources 124, and the like. In some implementations, computing resources 124 can communicate with other computing resources 124 via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0028] As further shown in FIG. 1, computing resource 124 includes a group of cloud resources such as one or more applications (“APPs”) 124-1, one or more virtual machines (“VMs”) 124-2, virtualized storage (“VSs”) 124-3, one or more hypervisors (“HYPs”) 124-4, and the like.

[0029] Application 124-1 includes one or more software applications that can be provided to and / or accessed by user device 110 and / or platform 120. Application 124-1 can eliminate the need to install and execute software applications on user device 110. For example, application 124-1 can include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, an application 124-1 can send information to / receive information from one or more other applications 124-1 via virtual machine 124-2.

[0030] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine depending on its use and correspondence to any physical machine by the virtual machine 124-2. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (“OS”). A process virtual machine can execute a single program and support a single process. In some embodiments, virtual machine 124-2 can operate on behalf of a user (e.g., user device 110) and manage the infrastructure of cloud computing environment 122 such as data management, synchronization, or long-term data transfer.

[0031] The virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resource 124. In some implementations, within the context of the storage system, the types of virtualization may include block virtualization and file virtualization. Block virtualization may refer to the abstraction (or separation) of logical storage from physical storage such that the storage system can be accessed regardless of the physical storage or heterogeneous structure. This separation may enable the administrator of the storage system to have flexibility in how the administrator manages the storage for the end user. File virtualization may eliminate the dependency between the data accessed at the file level and the location where the file is physically stored. This may enable optimization of the performance of storage usage, server integration, and / or seamless file migration.

[0032] The hypervisor 124-4 may provide hardware virtualization technology that enables multiple operating systems (e.g., "guest operating systems") to run simultaneously on a host computer such as the computing resource 124. The hypervisor 124-4 may present a virtual operating platform to the guest operating systems and may manage the execution of the guest operating systems. Multiple instances of various operating systems may share the virtualized hardware resources.

[0033] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, an optical fiber-based network, etc. and / or a combination of these or other types of networks.

[0034] The number and arrangement of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently than those shown in FIG. 1. Further, two or more of the devices shown in FIG. 1 may be implemented within a single device, or a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of environment 100 may perform one or more functions described as being performed by another set of devices of environment 100.

[0035] FIG. 2 is a block diagram of exemplary components of one or more of the devices of FIG. 1. Device 200 may correspond to user device 110 and / or platform 120. As shown in FIG. 2, device 200 may include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a communication interface 270.

[0036] Bus 210 includes components that enable communication among the components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) for storing information and / or instructions for use by processor 220.

[0037] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240, together with a corresponding drive, can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium.

[0038] Input component 250 includes components (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone) that enable device 200 to receive information, such as via user input. Additionally or alternatively, input component 250 can include sensors (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator) for sensing information. Output component 260 includes components (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs)) that provide output information from device 200.

[0039] The communication interface 270 includes transceiver-like components (such as a transceiver and / or separate receivers and transmitters) that enable the device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of a wired and wireless connection. The communication interface 270 may enable the device 200 to receive information from and / or provide information to other devices. For example, the communication interface 270 may include an Ethernet® interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0040] The device 200 may execute one or more of the processes described herein. The device 200 may execute these processes in response to a processor 220 that executes software instructions stored by a non-transitory computer-readable medium such as the memory 230 and / or the storage component 240. The computer-readable medium is defined herein as a non-transitory memory device. The memory device includes a memory space within a single physical storage device or a memory space that spans multiple physical storage devices.

[0041] The software instructions may be read into the memory 230 and / or the storage component 240 from another computer-readable medium or from another device via the communication interface 270. When executed, the software instructions stored in the memory 230 and / or the storage component 240 may cause the processor 220 to execute one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions for executing one or more of the processes described herein. Accordingly, the embodiments described herein are not limited to a particular combination of hardware circuitry and software.

[0042] The number and arrangement of the components shown in FIG. 2 are provided as an example. In practice, device 200 may include additional components, fewer components, different components, or components arranged differently than those shown in FIG. 2. Additionally or alternatively, a set of components (e.g., one or more components) of device 200 may perform one or more functions described as being performed by another set of components of device 200.

[0043] In some embodiments, an NBMP system 300 is provided. Referring to FIG. 3, NBMP system 300 includes an NBMP source 310, an NBMP workflow manager 320, a function repository 330, one or more media processing entities 350, a media source 360, and a media sink 370.

[0044] NBMP source 310 may receive instructions from a third-party entity, communicate with NBMP workflow manager 320 via an NBMP workflow, and communicate with function repository 330 via function discovery API 391. For example, NBMP source 310 may send a workflow description document (WDD) to NBMP workflow manager 320 and read a function description of functions stored in function repository 330. These functions may be media processing functions stored in the memory of function repository 330, such as functions for media decoding, feature point extraction, camera parameter extraction, projection methods, seam information extraction, blending, post-processing, and encoding. NBMP source 310 may comprise or be implemented by at least one processor and a memory storing code configured to cause the at least one processor to perform the functions of NBMP source 310.

[0045] The NBMP source 310 may request the NBMP workflow manager 320 to create a workflow by sending a workflow description document that includes tasks 352 to be executed by one or more media processing entities 350, where the workflow description document may include several descriptors, each of which may have several parameters.

[0046] For example, the NBMP source 310 may select a function stored in the function repository 330 and send a workflow description document to the NBMP workflow manager 320, which includes various descriptors about description details such as input and output data, required functions, and workflow requirements. The workflow description document may include a set of task descriptions and an input and output connection map of the tasks 352 to be executed by one or more of the media processing entities 350. When the NBMP workflow manager 320 receives such information from the NBMP source 310, the NBMP workflow manager 320 may create a workflow by instantiating tasks based on the names of the functions and connecting the tasks according to the connection map.

[0047] Alternatively or additionally, the NBMP source 310 may request the NBMP workflow manager 320 to create a workflow by using a set of keywords. For example, the NBMP source 310 may send a workflow description document that may include a set of keywords that the NBMP workflow manager 320 may use to find appropriate functions stored in the function repository 330 to the NBMP workflow manager 320. When the NBMP workflow manager 320 receives such information from the NBMP source 310, the NBMP workflow manager 320 creates a workflow by searching for appropriate functions using the keywords that may be specified in the Processing Descriptor of the workflow description document, provisions the tasks using other descriptors in the workflow description document, and connects them to create a workflow.

[0048] The NBMP workflow manager 320 may communicate with the function repository 330 via a function discovery API 393, which may be the same or a different API as the function discovery API 391, and may communicate with one or more of the media processing entities 350 via an API 394 (e.g., an NBMP task API). The NBMP workflow manager 320 may comprise or be implemented by at least one processor and a memory storing code configured to cause the at least one processor to perform the functions of the NBMP workflow manager 320.

[0049] The NBMP workflow manager 320 may use the API 394 to set up, configure, manage, and monitor one or more tasks 352 of a workflow that can be executed by one or more media processing entities 350. In some embodiments, the NBMP workflow manager 320 may use the API 394 to update and destroy tasks 352. To configure, manage, and monitor tasks 352 of a workflow, the NBMP workflow manager 320 may send messages, such as requests, to one or more of the media processing entities 350, where each message has a number of descriptors and each descriptor has a number of parameters. The tasks 352 may each include a media processing function 354 and configurations 353 for the media processing function 354.

[0050] In some embodiments, after receiving a workflow description document from the NBMP source 310 that does not include a list of tasks (e.g., includes a list of keywords instead of a list of tasks), the NBMP workflow manager 320 selects tasks based on the task descriptions in the workflow description document and searches the function repository 330 via the function discovery API 393 to find an appropriate function to execute as task 352 for the current workflow. For example, the NBMP workflow manager 320 may select tasks based on keywords provided in the workflow description document. After an appropriate function is identified by using the set of keywords or task descriptions provided by the NBMP source 310, the NBMP workflow manager 320 may configure the selected tasks in the workflow by using the API 394. For example, the NBMP workflow manager 320 may extract configuration data from the information received from the NBMP source and configure task 352 based on that configuration data.

[0051] One or more media processing entities 350 may be configured to receive media content from the media source 360 and process the media content according to a workflow that includes task 352 created by the NBMP workflow manager 320 and output the processed media content to the media sink 370. Each of the one or more media processing entities 350 may comprise or be implemented by at least one processor and a memory storing code configured to cause the at least one processor to perform the functions of the media processing entity 350.

[0052] Media source 360 may include a memory for storing media and may be integrated with or separate from NBMP source 310. In some embodiments, NBMP workflow manager 320 may notify NBMP source 310 when a workflow is ready, and media source 360 may transmit media content to one or more of media processing entities 350 based on the notification that the workflow is ready.

[0053] Media sink 370 may include or be implemented by at least one processor and at least one display configured to display media processed by one or more media processing entities 350.

[0054] As described above, messages from NBMP source 310 to NBMP workflow manager 320 (e.g., a workflow description document for requesting creation of a workflow) and messages from NBMP workflow manager 320 to one or more media processing entities 350 (e.g., for executing a workflow) may include several descriptors, and each descriptor may have several parameters. In some examples, communication between any components of NBMP system 300 using an API may include several descriptors, and each descriptor may have several parameters.

[0055] Embodiments may relate to a method for identifying and signaling nonessential inputs, outputs, and tasks in a workflow executed on a cloud platform.

[0056] The essential outputs of a workflow can be the outputs of the workflow that must generate data for the workflow to be considered to be operating properly. The essential inputs of a workflow can be the inputs that must be processed for the workflow to create the essential outputs of the workflow. A properly operating workflow can be a workflow that processes all essential inputs and generates all essential outputs. The essential tasks of a workflow can be the tasks necessary for proper operation and can be the tasks necessary to process the data required for a properly operating workflow. For example, an essential task can be a task that processes an essential input and / or generates an essential output. In some embodiments, an essential input can be an input required for an essential task to operate, and an essential output can be an output required as an essential input for the essential task or an output required as an output for the overall workflow. A non-essential input can be an input not required by the workflow to generate the essential outputs of the workflow. For example, a workflow may generate all essential outputs even if none of the non-essential inputs are provided, provided that all of the essential inputs are provided. A non-essential task can be a task included in the workflow but not an essential task. For example, a non-essential task can process a non-essential input and generate a non-essential output.

[0057] The NBMP standard defines splitter and merger function templates. FIG. 4 shows a diagram of workflow 400 corresponding to an example of this process that uses NBMP splitter / merger functions for parallel processing of segments. In FIG. 4, task T410 can be transformed into n instances of task T that are executed in parallel.

[0058] In some embodiments, the media stream is continuous. The splitter 420 may convert the media stream into N media sub-streams. Each sub-stream may be processed by an instance of T430A - C, and then the sub-streams may be interleaved with each other in the merger 440 to produce an output 450 equivalent to the task T410 output stream. The 1:N splitter and N:1 merger functions act at the segment boundaries. Each segment may have metadata of start, duration, and length, or a start code and a sequence number associated therewith. Since the segments are independent, the sub-streams are independent of each other in terms of being processed by the task T410. The tasks T0,..., T N-1 do not need to process the segments simultaneously. Since the segments and sub-streams can be independent, each instance of the task T410 may execute at its own speed. However, the current NBMP standard only deals with 1-D segmentation of media data.

[0059] The current NBMP standard defines the following format for segment metadata. Each segment may meet the following requirements: A. A continuous set of samples B. The maximum duration of D in the scale of the time scale T. Here, D and T are configuration parameters.

[0060] Each segment may use one of the following metadata: 1. Timing metadata: A. The start time s at the time scale t B. The time scale t = T C. The duration d at the time scale t D. The length l (bytes) 2. Or sequence metadata: A. A start code that is the same and unique in all other segments B. An ascending sequence number

[0061] In both cases, the media is segmented only in one dimension (e.g., typically time). However, media signals tend to be multi-dimensional.

[0062] This disclosure extends the splitter and merger functions to support the splitting and merging of segments with multi-dimensional metadata.

[0063] The following definitions may be used: - C is an M-dimensional vector [c0, c1,..., c M-1 , where the element ci has index i, and where index i + 1 is nested within index i, meaning that a 1 increment of the vector's index i is considered a larger increment than any increment of indices i + 1, i + 2,..., M - 1 where 0 ≦ i < M. - A multi-dimensional segment of dimension M may be defined as a segment representing information about samples in a space starting at point S = [s0, s1,..., s M-1 and length D = [d0, d1,..., d M-1 , where s i and d i are non-negative integers. If non-integer values are required, the vector T = [t0, t1,..., t M-1 represents the scale factor t i for dimension i, where the actual starting point and length in dimension i are s i / t i and d i / t i respectively, where t i is a positive integer.

[0064] In an embodiment, all segments of a media stream may have one of the following metadata: 1. Segment position data: a. Scaling vector T = [t0, t1,..., t M-11. Scale factors for S and D b. Unit t i For each index s i The start vector representing the start point of the media segment in the M - dimensional space having it S =[s0, s1,..., s M-1 c. Unit t i For each index d i The length vector representing the hyper - space covered by the media segment in the M - dimensional space having it D =[d0, d1,..., d M-1 d. The size of segment L in bytes 2. Segment sequence data: a. The sequence vector representing the sequence of media segments in the M - dimensional space having each index n i =[n0, n1,..., n n = M-1 b. The start code C, a unique code at which all segments start and which is not repeated in the middle of any segment. c. The size of segment L in bytes

[0065] The splitter function may split the multi - dimensional input into multiple multi - dimensional outputs and shall have the following requirements: ● A FIFO buffer for one input and N outputs. Here, N is a configuration parameter for the number of splits. ● Input: ○ Each input segment shall each meet the following requirements: ■ Scale T =[(T0, T1,..., T M-1 at the scale of D =[D0, D1,..., D M-1 's maximum size. Here D and and T are set in the configuration of the function. ■ Contain one of the following metadata and constraints: ● Location metadata:​​​ ○ Scale vector t = [t0, t1,..., t M-1 Start position at s = [s0, s1,..., s M-1 . Here, the number s i is scaled at t i . ○ Scale t = T ○ Scale t Length at d = [d0, d1,..., d M-1 ○ Byte size L (bytes) ● Or sequence metadata: ○ An identical and unique start code in all other segments ○ A sequence vector that increases with the order of the segments n = [n0, n1,..., n M-1 ■ Samples that do not overlap with other input segments ○ Assume that the input segments are in ascending order. ● Output: ○ The media stream in all output buffers is always composed of zero or more output segments. Each output segment shall meet the following requirements: ■ Exactly identical to only one input segment ■ Contain one of the following metadata and constraints: ● Location metadata: ○ Scale vector t = [t0, t1,..., t M-1 Start position at s = [s0, s1,..., s M-1 . Here, the number s i is scaled at t i . ○ Scale t = T ○ Scale t Length at d = [d0, d1,..., d M-1 ​​​ ○ Byte size L (bytes) ● Or sequence metadata: ○ An identical and unique start code in all other segments ○ A sequence vector that increases with the order of the segments n = [n0, n1,..., n M-1 ○ The set of all output segments of all output buffers shall cover the entire processed input segments together (e.g., input samples are not excluded from the set of output segments). ● The splitting process occurs as follows: Each of the N consecutive segments g0, g1,..., g in the input is moved to one of the output buffers O0, O1,..., O in this specific order (by moving the i-th segment g N-1 to the output O i ). i N-1 ● Each output buffer is sorted in ascending order of segments.

[0066] ​​​If there is a common header, the splitter shall repeat it in all single outputs. The media format may require the use of the same byte sequence at the start of the stream, regardless of any parameter changes in the content represented by the stream. This sequence of fixed, unchanging bytes is called the common header of that media format. The common header has zero length. In the splitter, it is necessary to add the common header to all single outputs in order to maintain the compliance of the output media stream with the specific media format to which the input stream conforms. Since the common header does not describe the length of the media, the segment containing the common header has zero length. Therefore, in the case of timing metadata, if the first segment of any stream has zero length, that segment is considered the common header of that stream, and the splitter can use this property to identify the common header segment. If the common header does not exist in the input (for example, the first segment has non-zero length), the splitter does not need to add any common header to any of the outputs, and the first segment of the input is treated the same as any other media segment within that input stream. In the case of sequence metadata, if the repeat-header flag is set in the configuration, the segment with sequence number 0 is the common header segment. 1. For temporal segments, m = M = 1, and d and t indicate the duration and time scale respectively. 2. The function can support the maximum dimension of the input signal, that is, it does not support the splitting of inputs with dimensions higher than a specific value. This value is defined as the contact configuration parameter.

[0067] The following table describes templates that support the multi-dimensional splitter function.

[0068]

Table 1

[0069]

Table 2

[0070] The merger function shall merge multiple multi-dimensional inputs into a single multi-dimensional output and have the following requirements: ● An FIFO buffer with N inputs and 1 output. N is a configuration parameter for the number of merges. ● Input: ○ Each input segment shall meet the following requirements: ■ Scale T =[(T0,T1,...,T M-1 in units of D =[D0,D1,...,D M-1 maximum size. Here D and T are set in the function configuration. ■ Include one of the following metadata and constraints: ● Location metadata: ○ Scale vector t =[t0,t1,...,t M-1 start position s =[s0,s1,...,s M-1 . Here, the number s i is scaled at t i . ○ Scale t = T ○ Scale t length at d =[d0,d1,...,d M-1 ○ Byte size L (bytes) ● Or sequence metadata: ○ The same and unique start code in all other segments ○ A sequence vector that increases with the order of the segments n =[n0,n1,...,n M-1 ​ ■Samples that do not overlap with any other input segments ○Assume that the input segments are in ascending order. ●Output: ○The media stream in the output buffer is always composed of zero or more output segments. All output segments shall meet the following requirements: ■Exactly the same as only one input segment ■Contain one of the following metadata and constraints: ●Position metadata: ○Scale vector t =[t0,t1,...,t M-1 starting position in s =[s0,s1,...,s M-1 . Here, the number s i is scaled at t i . ○Scale t = T ○Scale t length in d =[d0,d1,...,d M-1 ○Byte size L (bytes) ●Or sequence metadata: ○Start code that is the same and unique in all other segments ○Sequence vector that increases with the order of the segments n =[n0,n1,...,n M-1 ○The set of all output segments shall together cover all the processed input segments of all the inputs (e.g., the samples of the inputs are not excluded from the set of output segments). ●The merge process occurs as follows: One segment from all the inputs I0, I1,..., I N-1 is moved to the output buffer in that order, and then the output buffer is sorted in ascending order of the segments.

[0071] ​​​If there is a common header, the merger shall add it only once in a single output. The media format may require the use of the same byte sequence at the start of those streams, regardless of any parameter changes in the content represented by the streams. This sequence of fixed, unchanging bytes may be called the common header of that media format. The common header may be of zero length. In the merger, it may be necessary to add the common header only once to a single output in order to maintain the compliance of the output media stream with the specific media format to which the input streams conform. Since the common header does not describe the length of the media, a segment that includes the common header may be of zero length. Thus, in the case of timing metadata, if the first segment of any stream is of zero length, that segment may be regarded as the common header of that stream, and the merger may use this property to identify the common header segment. If the common header does not exist in the input (for example, the first segment is of non-zero length), the merger does not need to add any common header to any of the outputs, and the first segment of the input is treated the same as any other media segment within that input stream. In the case of sequence metadata, if the repeat header flag is set in the configuration, the segment with sequence number 0 is the common header segment. 1. For time segments, m = M = 1, and d and t indicate the duration and time scale respectively. 2. The function may support the maximum dimension of the input signal (for example, it does not support splitting of inputs with dimensions higher than a specific value). This value may be defined as a junction configuration parameter.

[0072] The following table describes a template that supports the multi-dimensional merger function.

[0073]

Table 3

[0074]

Table 4

[0075] Figure 5 is a flowchart of an exemplary process 500 for splitting and merging a multi - dimensional media stream. In some embodiments, one or more process blocks of Figure 5 may be performed by any of the above - described elements such as, for example, the NBMP system 300 or any element included therein such as the NBMP workflow manager 320.

[0076] As shown in Figure 5, process 500 includes segmenting a multi - dimensional media stream into a plurality of segments of multi - dimensional media in a multi - dimensional space (operation 510).

[0077] As further shown in Figure 5, process 500 may include splitting the segmented multi - dimensional media stream into a plurality of parallel sub - streams (operation 520).

[0078] As further shown in Figure 5, process 500 may include processing each of the plurality of sub - streams in parallel using multi - dimensional metadata carried with each multi - dimensional media segment (operation 530).

[0079] As further shown in Figure 5, process 500 may include merging the plurality of sub - streams into a single stream using multi - dimensional metadata carried in the output segment (operation 540).

[0080] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementation to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from practice of the implementation.

[0081] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed in this specification is an example of an exemplary approach. It is understood that, based on design preferences, the specific order or hierarchy of blocks in a process / flowchart may be rearranged. Additionally, some blocks may be combined or omitted. The accompanying method claims present the elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.

[0082] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible technical detail integration level. Additionally, one or more of the above-described components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to perform an operation.

[0083] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards, mechanically encoded devices such as raised structures within grooves having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transitory signal per se, such as, for example, a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0084] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0085] The computer-readable program code / instructions for performing the operations can be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection to an external computer may be configured (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to perform the aspects or operations and may personalize the electronic circuit.

[0086] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions for implementing the manner of functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0087] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0088] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. The methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in the drawings. In some alternative implementations, the functions represented by the blocks may occur in a different order than shown in the drawings. For example, two blocks shown in succession may actually be executed simultaneously or substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified function or operation, or by a combination of special-purpose hardware and computer instructions.

[0089] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods have been described herein without reference to specific software code, it being understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.

Claims

1. A method executed by at least one processor, comprising: segmenting a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; dividing the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, each of the plurality of sub-streams including segment metadata used to order segments within each sub-stream; processing each of the plurality of sub-streams in parallel; merging the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, the single stream including ordered segments; and further comprising representing the plurality of segments of the multi-dimensional media using respective sequence vectors, the sequence vectors including the segment metadata, the segment metadata including either (1) a start vector, a length vector, and a scaling vector, or (2) a start code. The method.

2. Each of the plurality of segments of the multi-dimensional media includes segment position metadata or segment sequence metadata. The method according to claim 1.

3. The start code is the same and unique in all of the plurality of segments of the multi-dimensional media. The method according to claim 1.

4. A method executed by at least one processor, comprising: segmenting a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; The step of dividing the segmented multi-dimensional media stream into a plurality of sub-streams capable of being processed in parallel, wherein each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream. The step of processing each of the plurality of sub-streams in parallel. The step of merging the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, wherein the single stream includes ordered segments. Including Before merging the plurality of sub-streams, further including the step of ordering the plurality of segments of the multi-dimensional media in ascending order of the sequence vector for each sub-stream. Method. A method executed by at least one processor, comprising: The step of segmenting a multi-dimensional media stream into a plurality of segments of the multi-dimensional media in a multi-dimensional space. The step of dividing the segmented multi-dimensional media stream into a plurality of sub-streams capable of being processed in parallel, wherein each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream. The step of processing each of the plurality of sub-streams in parallel. The step of merging the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, wherein the single stream includes ordered segments. Including After merging the plurality of sub-streams, further including the step of ordering the plurality of segments of the multi-dimensional media in ascending order of the sequence vector for each sub-stream. Method. **Claim 6** A method executed by at least one processor, comprising: segmenting a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; dividing the segmented multi-dimensional media stream into a plurality of sub-streams capable of being processed in parallel, each of the plurality of sub-streams including segment metadata used to order segments within each sub-stream; processing each of the plurality of sub-streams in parallel; merging the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, the single stream including ordered segments; and the plurality of segments of multi-dimensional media within the plurality of sub-streams include a plurality of input segments and a plurality of output segments, each of the plurality of output segments being exactly the same as one of the plurality of input segments within the plurality of sub-streams. A method. **Claim 7** An apparatus, comprising: at least one memory configured to store program code; at least one processor configured to read the program code and operate as instructed by the program code; and the program code includes: segmentation code configured to cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; Split code configured to cause the at least one processor to split the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream, and the split code; Processing code configured to cause the at least one processor to process each of the plurality of sub-streams in parallel; Merge code configured to cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, where the single stream includes ordered segments, and the merge code; including; The program code further includes representation code configured to cause the at least one processor to represent a plurality of segments of the multi-dimensional media using respective sequence vectors, the sequence vectors including the segment metadata, and the segment metadata including one of (1) a start vector, a length vector, and a scaling vector, and (2) a start code. **Claim 8** Each of the plurality of segments of the multi-dimensional media includes segment position metadata or segment sequence metadata; The apparatus according to claim 7. **Claim 9** The start code is the same and unique in all of the plurality of segments of the multi-dimensional media; The apparatus according to claim 7. **Claim 10** An apparatus comprising: At least one memory configured to store program code; At least one processor configured to read the program code and operate as instructed by the program code; comprising, wherein the program code segmentation code configured to cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; split code configured to cause the at least one processor to split the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, wherein each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream; processing code configured to cause the at least one processor to process each of the plurality of sub-streams in parallel; merge code configured to cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, wherein the single stream includes ordered segments; including wherein the program code further includes first ordering code configured to cause the at least one processor to order, for each sub-stream, the plurality of segments of the multi-dimensional media in ascending order of a sequence vector, before merging the plurality of sub-streams. device. **Claim 11**: A device comprising at least one memory configured to store program code; at least one processor configured to read the program code and operate as instructed by the program code; comprising, wherein the program code segmentation code configured to cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; Split code configured to cause the at least one processor to split the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream, and the split code; Processing code configured to cause the at least one processor to process each of the plurality of sub-streams in parallel; Merge code configured to cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried to the output segment, where the single stream includes ordered segments, and the merge code; Including; The program code further includes second ordering code configured to cause the at least one processor to order, in ascending order of a sequence vector, the plurality of segments of the multi-dimensional media for each sub-stream after merging the plurality of sub-streams. Device. At least one memory configured to store program code; At least one processor configured to read the program code and operate as instructed by the program code; Comprising, the program code is Segmentation code configured to cause the at least one processor to segment a multi-dimensional media stream into a plurality of segments of multi-dimensional media in a multi-dimensional space; A split code configured to cause the at least one processor to split the segmented multi-dimensional media stream into a plurality of sub-streams that can be processed in parallel, where each of the plurality of sub-streams includes segment metadata used to order segments within each sub-stream. Processing code configured to cause the at least one processor to process each of the plurality of sub-streams in parallel. A merge code configured to cause the at least one processor to merge the plurality of sub-streams into a single stream using the segment metadata carried in the output segments, where the single stream includes ordered segments. Including The plurality of segments of the multi-dimensional media within the plurality of sub-streams include a plurality of input segments and a plurality of output segments, and each of the plurality of output segments is exactly the same as one of the plurality of input segments within the plurality of sub-streams. Device.

13. A computer program that, when executed by at least one processor, causes the at least one processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Initial view angle control and presentation method and system based on three-dimensional point cloud

    CN112150603A

  • Method and apparatus for stateless parallel processing of tasks and workflows

    WO2021061785A1