PERSONALIZED AGREGATION OF VOLUMETRIC VIDEOS

The aggregated volumetric video system addresses limitations by enabling user-defined rules for combining and personalizing content from multiple sources, enhancing viewer interaction and immersion.

DE112023005271T5Pending Publication Date: 2026-04-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2023-11-24
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing volumetric video technologies limit viewers to a fixed perspective and lack personalization options, restricting the ability to combine and customize content from multiple sources.

Method used

An aggregated volumetric video system that allows users to specify aggregation rules for combining and personalizing volumetric videos from multiple sources, incorporating object recognition and metadata generation to create a single, customizable experience.

Benefits of technology

Enables viewers to freely navigate and personalize volumetric video content, enhancing immersion and flexibility by allowing selection and customization of objects and attributes across multiple video sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method, a computer system, and a computer program product are provided. The computer-implemented method includes selecting, using a first attribute of a first object, the first object in a first volumetric video. The method also includes selecting, using a second attribute of a second object, the second object in a second volumetric video, wherein the first and second attributes satisfy an aggregation rule. The method further includes generating an aggregated volumetric video from the first and second volumetric videos, wherein the generation of the aggregated video involves simultaneously rendering the first and second objects in the aggregated volumetric video based on the aggregation rule.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] The present invention relates generally to the processing of volumetric video. In particular, the present invention relates to a method, a system, and a computer program product for the personalized aggregation of volumetric videos.

[0002] Immersive technologies continue to gain popularity, especially as consumer devices like VR headsets continue to fall in price. The growth in consumer adoption of such technologies has led to growth in content, as content creators have responded to consumer interest in new immersive experiences. Furthermore, advances in video recording technology have enabled content creators to increase the level of immersion in their content.

[0003] Immersive video content is generally filmed using multiple cameras from different angles simultaneously, or with a single camera from multiple positions and angles. One example of immersive filmmaking is 360° filming. The resulting 360° video content is generally created by a computer by stitching together a number of images with a limited field of view, captured simultaneously, to form a complete sphere of still or video images within which a person can stand. This is most easily viewed by a person using a VR headset, but it can also be viewed in other ways. It is also limited in that the individual perspective within scenes is fixed relative to the images themselves. In other words, a viewer can only view such scenes from a position chosen by the filmmaker, and their movement within the scene is thus restricted.Viewers can look around in all directions, but they cannot move from the physical camera position. Additionally, conventional 360° video content sacrifices depth information and volumetric content, as it is essentially a sphere with a viewer at its center and images mounted along the inner wall of that sphere. There are no objects within the scene that have any shape other than this spherical wall. This further reduces the immersiveness of the experience, as it limits the viewer's perspective.

[0004] Volumetric video differs from 360° video in that it also utilizes depth information, which is captured using photogrammetry or depth sensors (e.g., light field arrays, LiDAR). This information allows volumetric video to capture both the images of the scene and the three-dimensional parameters of the objects within it. For example, a chair within a given volumetric video scene can have both a shape (e.g., a three-dimensional geometric form corresponding to that of the chair) and superimposed images to create the impression that it is made of wood, metal, plastic, or another material. Therefore, a viewer can generally move freely within the scene in volumetric video, overcoming the movement limitations of conventional two-dimensional or 360° video recording techniques.Content created using these techniques is often referred to as volumetric video, 6-degrees-of-freedom (6DoF), light field, or free-view video. SUMMARY

[0005] The exemplary embodiments offer personalized aggregation of volumetric videos. One embodiment features selection, using a first attribute of a first object, of the first object in a first volumetric video. The embodiment also features selection, using a second attribute of a second object, of the second object in a second volumetric video, wherein the first and second attributes satisfy an aggregation rule. The embodiment further features the generation of an aggregated volumetric video from the first and second volumetric videos, wherein the generation of the aggregated video includes simultaneous rendering of the first and second objects in the aggregated volumetric video based on the aggregation rule.Further embodiments of this aspect include corresponding computer systems, devices and computer program products recorded on one or more computer-readable storage media, each configured to perform the steps of the embodiment.

[0006] One embodiment comprises a computer-usable computer program product. The computer program product includes a computer-readable storage medium and program instructions stored on the storage medium.

[0007] One embodiment comprises a computer system. The computer system comprises a processor, computer-readable memory, and a computer-readable storage medium, as well as program instructions that are stored on the storage medium and executed by the processor via the memory. BRIEF DESCRIPTION OF THE FIGURES

[0008] The novel features considered characteristic of the invention are set forth in the accompanying claims. The invention itself, as well as a preferred embodiment, further objectives and advantages, are best understood with reference to the following detailed description of exemplary embodiments, when read in conjunction with the accompanying figures, wherein: Fig. Figure 1 shows a block diagram of a computer environment according to an exemplary embodiment; Fig. Figure 2 shows a functional block diagram of an exemplary volumetric video module according to an exemplary embodiment; Fig. Figure 3 shows a functional block diagram of an exemplary video processing environment according to an exemplary embodiment; Fig. Figure 4 shows an aggregation process according to an exemplary embodiment; Fig. Figure 5 shows a functional block diagram of a volumetric video module according to an exemplary embodiment; Fig. Figure 6 shows a functional block diagram of a volumetric video module according to an exemplary embodiment; Fig. Figure 7 shows a flowchart of an exemplary process for aggregating volumetric videos according to an exemplary embodiment; Fig. Figure 8 shows a flowchart of an exemplary process for aggregating volumetric videos using an optimization technique according to an exemplary embodiment; and Fig. Figure 9 shows a flowchart of an exemplary process for aggregating volumetric videos using a security technique according to an exemplary embodiment. DETAILED DESCRIPTION

[0009] Volumetric video differs from 360° video in that it contains depth information. This additional depth information fundamentally changes how content is consumed. When viewing a scene in 360° video format, the viewer is fixed to a single position and, from that point of view, has up to three degrees of freedom, corresponding to rotations around each of the three orthogonal axes of a Cartesian coordinate system: roll (tilting the head left or right), pitch (tilting the head forward or backward), and yaw (rotating the head left or right).

[0010] In contrast, the depth information in volumetric video frees the viewer from a fixed point of view, allowing up to six degrees of freedom, corresponding to rotations around and translations along each of the three orthogonal axes of a Cartesian coordinate system: roll, pitch, and yaw as previously described, plus elevation (moving up or down), pan (moving left or right), and forward / backward pan. This allows the viewer to move freely within a volumetric video scene and view objects from different angles and positions. This additional freedom of movement significantly enhances the immersion of content in volumetric video compared to previous technologies.

[0011] One drawback of current volumetric video technology is that viewers are still limited to seeing the volumetric video scene only as it was originally created. That is, viewers have no control over the content of the volumetric video. For example, if a user is remotely participating in a conference taking place simultaneously in two or more locations, each with volumetric videos available for remote participants, the user is limited to viewing only one of the two volumetric videos at a time. In other words, the user cannot choose an option that provides a volumetric video containing elements of interest from each of the source volumetric videos.

[0012] Disclosed embodiments address these and other limitations of conventional volumetric video systems by providing an aggregated volumetric video that is an aggregation of two or more source volumetric videos. Disclosed embodiments also allow for the personalization of the aggregated volumetric video. For example, in some embodiments, a user can specify one or more aggregation rules.

[0013] Aggregation rules are primarily instructions and / or preferences that control the rendering of two or more source volumetric videos into a single aggregated volumetric video. Aggregation rules can specify which objects from the source volumetric videos should or should not be displayed in the aggregated volumetric video (e.g., "show the audience from the first source instead of the audience from the second source"). Aggregation rules can also specify attributes of objects from the source volumetric videos that should be displayed, not displayed, or modified in the aggregated volumetric video (e.g., "change the color of the audience seats from blue to black; obscure company logos on the audience's shirts and hats").

[0014] Aggregation rules can have rules of varying specificity. For example, aggregation rules can specify certain attributes, specific objects, or classes of objects, with the classes of objects varying in their specificity. As a more specific, non-restrictive example, a user creating rules for a personalized, aggregated volumetric video might have source volumetric videos from multiple auto shows. In this scenario, rules of varying specificity could include a rule for a broad class of objects, such as a preference for sedans over other vehicle types; a more specific rule might specify particular makes and models of sedans; and a rule for a specific object might indicate a famous car in one of the volumetric videos.

[0015] In some embodiments, video aggregation is performed by a volumetric video module. In some of these embodiments, the volumetric video module analyzes the digital volumetric video data from two or more source volumetric videos. In some of these embodiments, this analysis results in the identification of one or more objects in each of the source volumetric videos.

[0016] In some embodiments, the analysis generates metadata for one or more of the identified objects. In some embodiments, the metadata for an identified object includes data representing one or more attributes of the identified object. In various embodiments, the number and type of attributes vary depending on implementation decisions, constraints, and / or preferences. Non-restrictive examples of attribute types include appearance attributes, classification attributes, and / or source attributes. In exemplary embodiments, the appearance attributes of an identified object may include size, shape, and / or color attributes of the identified object.

[0017] In exemplary embodiments, classification attributes of an identified object can comprise one or more classes or categories assigned to the identified object. In some embodiments, a classification attribute includes a class (or classes) predicted by known object recognition and / or classification techniques.

[0018] For example, in some embodiments, object recognition techniques incorporate machine learning processes that use trained machine learning models to predict categories of objects in an image. It should be noted that, for the purposes of this disclosure, "images" also include images representing frames of a video. In some embodiments, the classification is performed using known semantic image segmentation techniques that assign sections or segments of a volumetric image to a corresponding class of what the section or segment of the volumetric image represents.

[0019] In some embodiments, object recognition incorporates techniques that combine classification with localization to determine the positions of classified objects in an image. Thus, in some embodiments, object recognition is performed using known techniques that determine which objects are present in an image and indicate where the objects are located within the image. In some embodiments, object recognition includes known instance segmentation techniques that distinguish between separate objects of the same class within an image.

[0020] In some embodiments, classification attributes can have multiple classes representing different degrees of specificity. For example, if an identified object is a multi-passenger van, the identified object might have "vehicle" as its first classification attribute, "multi-passenger vehicle" as its second classification attribute, and "van" as its third classification attribute.

[0021] In exemplary embodiments, source attributes may contain information that identifies the source or sources of the volumetric video. A non-limiting example: Suppose an aggregated volumetric video is created for a conference taking place simultaneously at two geographically distant locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video recorded at the first location and a second volumetric video recorded at the second location. Thus, a first identified object captured in the first volumetric video can have a source attribute indicating the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video can have a source attribute indicating the second location or the second volumetric video as the source of the second identified object.

[0022] The set of size and shape attributes can contain a single attribute or two or more. A size and shape attribute is a feature, property, or other characteristic of an object that is associated with size measurements and / or shapes of one or more parts of the object. A part of an object can be a section, segment, or component of the object. A part of an object can also be the entire object. A size measurement is any measured quantity, such as height, length, and width. A size measurement can also include the weight and volume of an object. The set of size and shape attributes can contain only size-related attributes without shape attributes. In another embodiment, the set of size and shape attributes can contain only shape-related attributes without size-related attributes.In another embodiment, the set of size and shape attributes includes both size-related and shape-related attributes.

[0023] An embodiment can be implemented as a software application. The application implementing an embodiment can be configured as a modification in an existing manufacturing system, as a separate application operated in conjunction with an existing manufacturing system, as a standalone system, or a combination thereof.

[0024] One embodiment monitors system status data for an indication of a system error that causes the system kernel to enter a halt state. The status data can vary depending on the system type (e.g., operating system and hardware), but generally includes system error messages, error codes, log entries, or other data representing a system error. The system error itself can also vary depending on the system type, but generally includes kernel errors, kernel panics, halt errors, or similar errors that result in a partial or complete loss of kernel functionality.

[0025] A system error that results in a partial or complete loss of kernel functionality is generally recognized by the system as a condition requiring a system reboot for recovery. In most system types, the loss of kernel functionality typically necessitates a reboot for recovery. Therefore, such errors are examples of failures that meet a reboot condition.

[0026] In exemplary embodiments, when a system error is detected that meets a restart condition (e.g., a system error that leads to the loss of kernel functionality), debug data is temporarily stored in a protected memory area. Since at least some kernel functionality has already been lost at this point, this embodiment involves creating and storing a copy of the debug data without kernel assistance. In some embodiments, for example, an exception handler creates and stores the copy of the debug data. The debug data can vary depending on the system type (e.g., operating system and hardware), but generally includes data from processor memory (e.g., registers and cache), logs, and / or trace arrays.

[0027] Because the kernel is in a suspended state while the debug data is being generated, the kernel functionality for processing the debug data is unavailable. Therefore, the kernel is not available at this time to filter sensitive information from the debug data. For this reason, the debug data is stored in protected memory, where it can be retained during a reboot and subsequently processed.

[0028] After the reboot, a debug device, which is untrusted by the restored system, connects via an I / O port using a trusted protocol. The untrusted device requests debug data to help determine the cause of the system error. Because the debug data is stored in protected memory, the untrusted device cannot access it directly. Instead, the untrusted instance must request the debug data from a secure debug module.

[0029] In the illustrated embodiment, the secure debug module receives the request for debug data. In response to the request, the secure debug module analyzes the debug data using a data cleansing module with a sensitive data detection / cleaning method that identifies and removes sensitive data from the debug data.

[0030] In some implementations, the data cleansing module identifies sensitive data according to an audit policy. An audit policy is a set of preferences, rules, and / or criteria that protect sensitive data within the debug data. For example, an audit policy might define "sensitive objects" as files or objects that contain certain keywords (such as "confidential" or "privileged") and / or are associated with specific keywords (e.g., in metadata) or specific flags (e.g., in metadata that designates a document or email as personal, confidential, etc.). An audit policy might also specify rules for handling sensitive objects. For example, an audit policy might require an auditor to approve the transfer of potentially sensitive objects from protected storage to the untrusted instance.Thus, in some implementations, the data protection process includes one or more safeguards in the form of data cleansing to prevent sensitive data containing debug data from reaching the untrusted instance.

[0031] In the illustrated embodiment, a window module detects whether sensitive data was processed during a time window in which the system error occurred, for example, by reviewing system logs. In some embodiments, the window module reviews the system logs using an audit policy that includes a set of preferences, rules, and / or criteria by which the window module detects sensitive data in the debug data. For example, an audit policy might define "sensitive objects" as files or objects that contain certain keywords (e.g., "confidential" or "privileged") and / or are associated with certain keywords (e.g., in metadata) or certain flags (e.g., in metadata that mark a document or email as personal, confidential, etc.). An audit policy might also specify rules for handling sensitive objects.For example, an audit policy might require an auditor to approve the transfer of potentially sensitive objects from protected storage to the untrusted instance. Thus, in some implementations, the data protection process includes one or more safeguards, such as time window analysis, to prevent sensitive data, including debug data, from reaching the untrusted instance.

[0032] In some implementations, the data cleansing module identifies sensitive data according to an audit policy. An audit policy is a set of preferences, rules, and / or criteria that protect sensitive data within the debug data. For example, an audit policy might define "sensitive objects" as files or objects that contain certain keywords (such as "confidential" or "privileged") and / or are associated with specific keywords (e.g., in metadata) or specific flags (e.g., in metadata that designates a document or email as personal, confidential, etc.). An audit policy might also specify rules for handling sensitive objects. For example, an audit policy might require an auditor to approve the transfer of potentially sensitive objects from protected storage to the untrusted instance.Thus, in some implementations, the data protection process incorporates one or more safeguards in the form of data encryption to prevent sensitive data containing debug data from reaching the untrusted instance.

[0033] For the sake of clarity and without limitation, the exemplary embodiments are described using some example configurations. Based on this disclosure, those skilled in the art can conceive numerous variations, adaptations, and modifications of a described configuration to achieve a described purpose, and these are also provided for within the scope of the exemplary embodiments.

[0034] Furthermore, simplified diagrams of data processing environments are used in the figures and exemplary embodiments. In an actual computer environment, additional structures or components may exist that are not shown or described here, or that may differ from the structures or components shown but perform a similar function without going beyond the scope of the exemplary embodiments.

[0035] The exemplary embodiments are described only as examples with regard to certain actual or hypothetical components. Specific variations of these and other similar artifacts are not intended to limit the invention. Any suitable variation of these and other similar artifacts can be selected within the scope of the exemplary embodiments.

[0036] The examples in this disclosure serve only for clarity of description and are not intended to limit the exemplary embodiments. All advantages listed here are merely examples and are not intended to limit the exemplary embodiments. Additional or other advantages may be realized through specific exemplary embodiments. Furthermore, a particular exemplary embodiment may have some, all, or none of the advantages listed above.

[0037] Furthermore, the exemplary embodiments can be implemented with respect to any data type, data source, or access to a data source via a data network. Any type of data storage device can provide the data for an embodiment of the invention, either locally on a data processing system or via a data network, within the scope of the invention. If an embodiment using a mobile device is described, any data storage device suitable for use with the mobile device can provide the data for such an embodiment, either locally on the mobile device or via a data network, within the scope of the exemplary embodiments.

[0038] The exemplary embodiments are described using specific code, computer-readable storage media, high-level functions, designs, architectures, protocols, layouts, circuit diagrams, and tools, but these descriptions are not limiting. Furthermore, in some cases, the exemplary embodiments are described using specific software, tools, and data processing environments, but only as examples to clarify the description. The exemplary embodiments can be used in conjunction with other comparable or similarly suitable structures, systems, applications, or architectures. For example, other comparable mobile devices, structures, systems, applications, or architectures can be used in connection with such an embodiment of the invention.An exemplary embodiment can be implemented in hardware, software, or a combination thereof.

[0039] The examples in this disclosure serve only for clarity of description and are not limiting to the exemplary embodiments. Additional data, operations, actions, tasks, activities, and manipulations will be conceivable from this disclosure and are provided for within the scope of the exemplary embodiments.

[0040] Various aspects of the present disclosure are described by descriptive text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic in computer program product (CPP) execution forms. With regard to flowcharts, depending on the technology used, the operations may be performed in a different order than depicted in the respective flowchart. For example, again depending on the technology used, two consecutive operations in the flowchart may be performed in reverse order, as a single integrated step, simultaneously, or at least partially overlapping in time.

[0041] An embodiment as a computer program product (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any set of one or more storage media (also called “mediums”) that are collectively contained in a set of one or more storage devices and collectively contain machine-readable code comprising instructions and / or data for performing computer operations according to a given CPP claim. A “storage device” is any tangible device capable of retaining and storing instructions for use by a computer processor. Without limitation, the computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor-based storage medium, a mechanical storage medium, or any suitable combination of the foregoing media.Some known storage devices containing these media include: floppy disks, hard disks, main memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static RAM (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disks, mechanically coded devices (such as punched cards or indentations / protrusions on a main surface of a disk), or any suitable combination of the aforementioned media. A computer-readable storage medium within the meaning of this disclosure is not to be understood as storage in the form of volatile signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses traveling through a fiber optic cable, electrical signals transmitted through a cable, and / or other transmission media.As experts know, data is occasionally moved during the normal operation of a storage device, e.g., during access, defragmentation, or cleanup, but this does not make the storage device volatile, as the data is not volatile as long as it is stored.

[0042] Fig. Figure 1 shows a block diagram of a computer environment 100. The computer environment 100 includes an example of an environment for executing at least part of the computer code involved in carrying out the inventive methods, such as an enhanced volumetric video module 200 that provides personalized aggregation of volumetric videos. In addition to the volumetric video module 200, the computer environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, a user device 103, a server 104, a public cloud 105, and a private cloud 106.In this embodiment, the computer 101 comprises: the processor set 0 (including processing circuit 120 and cache 121), the communication fabric 111, the volatile memory 112, the persistent memory 113 (including the operating system 122 and volumetric video module 200, as specified above), the peripheral device set 114 (including the user interface device set 123, memory 124, and IoT sensor set 125), and the network module 115. Server 104 comprises the remote database 130. Public cloud 105 comprises the gateway 140, the cloud orchestration module 141, the physical host system 142, the virtual machine 143, and the container 144.

[0043] A COMPUTER 101 can take the form of a desktop computer, laptop, tablet computer, smartphone, smartwatch or other portable computer, mainframe, quantum computer, or any other form of computer or mobile device capable of running a program, accessing a network, or querying a database, such as the remote database 130. As is well known in computer technology, and depending on the technology, the execution of a computer-implemented procedure can be distributed across multiple computers and / or locations. On the other hand, for the sake of simplicity, this description of the computer environment 100 focuses on a single computer, specifically Computer 101. Computer 101 may reside in a cloud, even if it is located in Fig. Computer 101 is not represented in a cloud. Conversely, computer 101 does not have to be in a cloud unless explicitly stated.

[0044] A processor set 110 comprises one or more computer processors of any type currently known or to be developed in the future. The processing circuitry 120 can be distributed across multiple packages, such as multiple coordinate integrated circuit chips. The processing circuitry 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the processor chip package and is typically used for data or code that needs to be quickly available to the threads or cores of the processor set 110. Cache memories are typically divided into multiple levels, depending on their relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located off-chip. In some computing environments, the processor set 110 may be designed for working with qubits and performing quantum computations.

[0045] Computer-readable program instructions are typically loaded onto Computer 101 to execute a sequence of operations by Processor Set 110 of Computer 101, thereby effecting a computer-implemented method such that the instructions executed in this manner instantiate the methods specified in flowcharts and / or descriptive texts of the computer-implemented methods in this document (collectively referred to as "the inventive methods"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as Cache 121 and the other storage media described below. The program instructions and associated data are retrieved by Processor Set 110 to control and direct the execution of the inventive methods.In computer environment 100, at least some of the instructions for carrying out the inventive methods in the volumetric video module 200 can be stored in persistent memory 113.

[0046] A COMMUNICATION FABRIC 111 is a signal transmission path that allows the various components of Computer 101 to communicate with each other. Typically, this fabric consists of switches and electrically conductive paths, such as buses, bridges, physical input / output ports, and the like. Other types of signal transmission paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0047] A volatile memory 112 is any type of volatile memory currently known or that may be developed in the future. Examples include dynamic RAM or static RAM types. Typically, volatile memory 112 is characterized by random access; however, this is not required unless explicitly stated. In computer 101, volatile memory 112 is contained in a single block that is internal to computer 101; however, it may alternatively or additionally be distributed across multiple packages and / or located external to computer 101.

[0048] A persistent memory 113 is any form of non-volatile memory for computers that is currently known or will be developed in the future. The non-volatile nature of this memory means that the stored data is retained regardless of whether the computer 101 and / or the persistent memory 113 are powered on. A persistent memory 113 can be read-only memory (ROM); however, typically at least a portion of the persistent memory allows data to be written, erased, and overwritten. Some known forms of persistent memory are magnetic disks and solid-state storage devices. The operating system 122 can take various forms, such as different well-known proprietary operating systems or open-source Portable Operating System Interface-like operating systems that use a kernel.The code contained in the volumetric video module 200 typically includes at least some part of the computer code involved in carrying out the inventive processes.

[0049] A peripheral device set 114 comprises a majority of the peripheral devices of the computer 101. Data communication links between the peripheral devices and the other components of the computer 101 can be implemented in various ways, such as Bluetooth connections, near-field communication (NFC) connections, wired connections (e.g., USB-type cables), plug-in connections (e.g., SD cards), connections via local area networks, and even connections via wide area networks such as the internet. In various embodiments, the user interface device set 123 can include components such as a screen, speakers, microphone, wearable devices (such as glasses and smartwatches), keyboard, mouse, printer, touchpad, game controller, and haptic devices. The memory 124 is external storage, such as an external hard drive, or plug-in storage, such as an SD card. The memory 124 can be persistent and / or volatile.In some embodiments, the memory 124 can take the form of a quantum computer storage device to store data in the form of qubits. In embodiments where computer 101 requires a large storage capacity (e.g., when computer 101 stores and manages a large database locally), this storage can be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer and another a motion detector.

[0050] A network module 115 is a collection of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers over the wide area network 102. The network module 115 may include hardware such as modems or Wi-Fi signal transceivers, software for packetizing and / or unpacking data for transmission over communication networks, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of the network module 115 are performed on the same physical hardware device. In other embodiments (for example, embodiments that utilize software-defined networking (SDN)), the control and forwarding functions of the network module 115 are performed on physically separate devices, so that the control functions manage several different network hardware devices.Computer-readable program instructions for carrying out the inventive methods can typically be downloaded from an external computer or external storage device to computer 101 via a network card or network interface included in network module 115.

[0051] WIDE WAN (WAN) 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any computer data communication technology currently known or hereafter developed. In some embodiments, the WAN 102 may be replaced and / or supplemented by local area networks (LANs) designed for communication between devices within a local area, such as a Wi-Fi network. The WAN and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.

[0052] A USER DEVICE (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of a company operating Computer 101) and can take any of the forms described above in connection with Computer 101. The User Device 103 typically receives helpful and useful data from the operations of Computer 101. For example, in a hypothetical case where Computer 101 is designed to provide a recommendation to an end-user, this recommendation would typically be communicated from the Network Module 115 of Computer 101 to the User Device 103 via the Wide Area Network 102. Thus, the User Device 103 can display or otherwise present the recommendation to an end-user. In some embodiments, the User Device 103 may be a client device, such as a thin client, heavy client, mainframe, desktop computer, etc.

[0053] A SERVER 104 is any computer system that provides at least some data and / or functionality to Computer 101. Server 104 can be controlled and used by the same organization that operates Computer 101. Server 104 represents the machine(s) that collect and store helpful and useful data for use by other computers, such as Computer 101. For example, in a hypothetical case where Computer 101 is designed and programmed to provide a recommendation based on historical data, this historical data can be delivered from Server 104 to Computer 101 from remote database 130.

[0054] A PUBLIC CLOUD 105 is any computing system usable by multiple organizations that provides on-demand computing resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without requiring the user to actively manage the underlying infrastructure. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computing hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented through virtual computing environments running on various computers that comprise the physical host system 142, which is the universe of physical computers in and / or for the public cloud 105.Virtual Computing Environments (VCEs) typically take the form of virtual machines from Virtual Machine 143 and / or containers from Container 144. It is understood that these VCEs can be stored as images and moved between different physical host systems, either as images or after VCE instantiation. The Cloud Orchestration Module 141 manages the transfer and storage of images, starts new instances of VCEs, and manages active instances of VCE deployments. The Gateway 140 is the collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate over the Wide Area Network 102.

[0055] Further explanation of virtualized computing environments (VCEs) is required. VCEs can be stored as "images." A new active instance of the VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system virtualization. This refers to an operating system feature where the kernel allows the existence of multiple isolated user-space instances, called containers. These isolated user-space instances generally behave like real computers from the perspective of the programs running within them. A computer program running on a conventional operating system can utilize all of that computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities.However, programs running in a container can only use the container's contents and assigned devices, a property known as containerization.

[0056] A private cloud (106) is similar to a public cloud (105), except that the computing resources are available only to a single organization. While the private cloud (106) is depicted as connected to the wide area network (102), in other configurations, a private cloud may be completely isolated from the internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public), often implemented by different providers. Each of the multiple clouds remains a separate and independent entity, but the larger hybrid cloud architecture is held together by standardized or proprietary technology that enables orchestration, management, and / or data / application portability between the multiple sub-clouds.In this embodiment, the public cloud 105 and the private cloud 106 are both part of a larger hybrid cloud.

[0057] Metered service: Cloud systems automatically control and optimize resource usage by employing a metering function at an abstraction level appropriate for the service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, reported, and billed, providing transparency for both the service provider and the consumer.

[0058] Fig. Figure 2 shows a functional block diagram of an exemplary volumetric video module 200 according to an exemplary embodiment. In the illustrated embodiment, the volumetric video module 200 receives a plurality of volumetric videos 206A-206C and aggregates two or more of them into an aggregated volumetric video 208.

[0059] In the illustrated embodiment, the volumetric video module 200 includes an aggregation engine 202 and aggregation rule(s) 204, which are stored on a computer-readable storage medium. In alternative embodiments, the volumetric video module 200 may have some or all of the functionality described herein, but distributed differently across one or more modules. In some embodiments, the functionality described herein is distributed across multiple systems, which may include combinations of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.

[0060] In the illustrated embodiment, the aggregation engine 202 aggregates individual volumetric videos, such as two or more of the volumetric videos 206A-206C, and creates a single aggregated volumetric video 208, also referred to here as corridor 208. In some embodiments, the volumetric videos 206A-206C are recorded simultaneously from different physical locations. For example, volumetric video 206A may be a live volumetric video broadcast from Los Angeles, volumetric video 206B a live volumetric video broadcast from Chicago, and volumetric video 206C a live volumetric video broadcast from New York for a conference taking place at these three locations. In some embodiments, the volumetric videos 206A-206C are recorded at different times from different physical locations.

[0061] Aggregation rules 204 are primarily instructions and / or preferences that control the rendering of two or more source volumetric videos into a single aggregated volumetric video. Aggregation rules 204 can include specifying objects from the source volumetric videos to be displayed or not displayed in the aggregated volumetric video (e.g., "show the audience from the first source instead of the audience from the second source"). Aggregation rules 204 can also include specifying attributes of objects from the source volumetric videos to be displayed, not displayed, or modified in the aggregated volumetric video (e.g., "change the color of the audience seats from blue to black; obscure company logos on the audience's shirts and hats").

[0062] Aggregation rules 204 can contain rules of varying specificity. For example, aggregation rules 204 can include specifying certain attributes, specific objects, or classes of objects, with the classes of objects varying in their specificity. As a more specific, non-restrictive example, a user creating rules 204 for a personalized, aggregated volumetric video might have source volumetric videos from multiple auto shows. In this scenario, rules 204 of varying specificity could include a rule for a broad class of objects, such as a preference for sedans over other vehicle types; a more specific rule might specify particular makes and models of sedans; and a rule for a specific object might indicate a famous car in one of the volumetric videos.

[0063] In some embodiments, the aggregation engine 202 performs video aggregation. In some of these embodiments, the aggregation engine 202 analyzes the digital volumetric video data from two or more source volumetric videos. In some of these embodiments, this analysis results in the identification of one or more objects in each of the source volumetric videos.

[0064] In some embodiments, the aggregation engine 202 generates metadata for one or more of the identified objects. In some embodiments, the metadata for an identified object includes data representing one or more attributes of the identified object. In different embodiments, the number and type of attributes vary depending on implementation decisions, constraints, and / or preferences. Non-restrictive examples of attribute types include appearance attributes, classification attributes, and / or source attributes. In exemplary embodiments, appearance attributes of an identified object may include size, shape, and / or color attributes of the identified object.

[0065] In exemplary embodiments, classification attributes of an identified object may have one or more classes or categories that are assigned to the identified object. In some embodiments, a classification attribute has a class (or classes) that are predicted using known object recognition and / or classification techniques.

[0066] In some embodiments, for example, the aggregation engine 202 includes an image classifier 210 that performs object recognition techniques, which in turn include machine learning processes that use trained machine learning models to predict categories of objects in an image. It should be noted that, for the purposes of this disclosure, "images" also include images representing frames of a video. In some embodiments, the classification is performed using known semantic image segmentation techniques that assign sections or segments of a volumetric image to a corresponding class of what the section or segment of the volumetric image represents.

[0067] In some embodiments, the aggregation engine 202 performs object recognition using techniques that combine classification with localization to determine the positions of classified objects in an image. Thus, in some embodiments, the aggregation engine 202 performs object recognition using known techniques that determine which objects are present in an image and indicate where the objects are located within the image. In some embodiments, the aggregation engine 202 uses object recognition techniques that incorporate known instance segmentation techniques to distinguish between separate objects of the same class within an image.

[0068] In some embodiments, classification attributes can comprise multiple classes representing different degrees of specificity. For example, if an identified object is a multi-passenger van, the identified object might have "vehicle" as its first classification attribute, "multi-passenger vehicle" as its second classification attribute, and "van" as its third classification attribute.

[0069] In exemplary embodiments, source attributes can contain information that identifies the source or sources of the volumetric video. A non-limiting example: Suppose an aggregated volumetric video is created for a conference taking place simultaneously at two geographically distant locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video recorded at the first location and a second volumetric video recorded at the second location.Thus, a first identified object captured in the first volumetric video can contain a source attribute that specifies the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video can contain a source attribute that specifies the second location or the second volumetric video as the source of the second identified object.

[0070] The set of size and shape attributes can contain a single attribute or two or more. A size and shape attribute is a feature, property, or other characteristic of an object that is associated with size measurements and / or shapes of one or more parts of the object. A part of an object can be a section, segment, or component of the object. A part of an object can also encompass the entire object. A size measurement is any measured quantity, such as height, length, and width. A size measurement can also include the weight and volume of an object. The set of size and shape attributes can contain only size-related attributes without shape attributes. In another embodiment, the set of size and shape attributes can contain only shape-related attributes without size-related attributes.In another embodiment, the set of size and shape attributes includes both size-related and shape-related attributes.

[0071] Fig. Figure 3 shows a functional block diagram of an exemplary video processing environment 300 according to an exemplary embodiment. In the illustrated embodiment, the video processing environment 300 comprises the server 306A, the server 306B, and the server 310, each of which contains a volumetric video module 200. Fig. 1 and Fig. 2 can be shown.

[0072] In the illustrated embodiment, the volumetric videos 302A-302D are examples of source volumetric videos. In the illustrated embodiment, the volumetric videos 302A-302D are each created with a plurality of cameras 304. Alternatively, one or more of the volumetric videos 302A-302D can be created with a single camera that records from multiple positions and angles.

[0073] The embodiment in Fig. Figure 3 clarifies that disclosed embodiments, in addition to aggregating source videos (such as volumetric videos 302A-302D), can also aggregate volumetric videos. The illustrated embodiment also allows the aggregation of more than two source volumetric videos.

[0074] In the illustrated embodiment, server 306A aggregates volumetric video 302A and volumetric video 302B to form aggregated volumetric video 308A, and server 306B aggregates volumetric video 302C and volumetric video 302D to form aggregated volumetric video 308B. Server 310 then aggregates the aggregated volumetric video 308A and the aggregated volumetric video 308B to form aggregated volumetric video 312. In an alternative embodiment, an aggregated volumetric video can be created using the source two-dimensional (2D) video and depth information from cameras 304 instead of volumetric video. For example, in some embodiments of the server 306A, the aggregated volumetric video 308A is generated using 2D video and depth information from the cameras 304, which are associated with volumetric video 302A and volumetric video 302B.

[0075] In another embodiment, an aggregated volumetric video can be further aggregated with another source volumetric video. For example, if nothing from volumetric video 302D is needed for the aggregated volumetric video 312, server 310 would aggregate the aggregated volumetric video 308A with volumetric video 302C to form aggregated volumetric video 312.

[0076] Fig. Figure 4 shows an aggregation process 400 according to an exemplary embodiment. In one embodiment, the aggregation process 400 is initiated from the volumetric video module 200. Fig. 1 and Fig. 2 carried out.

[0077] The embodiment in Fig. Figure 4 illustrates that alternative embodiments, in addition to aggregating two volumetric videos (source or aggregated) simultaneously, can also aggregate three or more volumetric videos (source or aggregated) simultaneously. In the illustrated embodiment, user rules have been set to aggregate volumetric video 402A and volumetric video 402B with volumetric video 402C, with volumetric video 402C being used primarily, but the audience in frame 404 from volumetric video 402A and the audience in frame 408 from volumetric video 402B are displayed instead of the audience at insertion positions 406 and 410 from volumetric video 402C.

[0078] Fig. Figure 5 shows a functional block diagram of a volumetric video module 502 according to an exemplary embodiment. The volumetric video module 502 is related to the volumetric video module 200 from Fig. 1 and Fig. 2 similarly, with the difference that the aggregation engine 504 and rendering engine 510 of the volumetric video module 502 generate volumetric videos and / or aggregated volumetric videos of an aggregation hierarchy according to user navigation inputs.

[0079] For example, in the illustrated embodiment, a hierarchy of volumetric videos comprises volumetric video 516A, volumetric video 516B, and aggregated volumetric video 516C. At time T1, a user 514 uses a headset 512 to view volumetric video 516A. The headset 512 has or is connected to the volumetric video module 502.

[0080] The aggregation engine 504 has user input and one or more source volumetric videos 206A-206C. The aggregation engine 504 also has access to aggregation rule(s) 506 and an aggregation hierarchy 508. The description of the aggregation rule(s) 204 from Fig. 2 applies equally to aggregation rule(s) 506. The aggregation hierarchy 508 can be a playlist that specifies which volumetric videos are to be presented and in what order they are to be presented. In some embodiments, a user 514 creates the aggregation hierarchy 508. In some embodiments, the volumetric video module 502 automatically generates the aggregation hierarchy 508 according to user rules or preferences (e.g., play source volumetric videos in order from nearest to furthest recording location, then aggregated).

[0081] In the illustrated embodiment, the headset 512 can have controls that allow easy forward / backward navigation in the aggregation hierarchy 508. These controls send user input to the aggregation engine 504. The aggregation engine 504 checks the aggregation hierarchy 508 and then instructs the rendering engine 510 to transmit the corresponding volumetric video. In the illustrated embodiment, the aggregation hierarchy 508 includes the volumetric video 516A, followed by the volumetric video 516B, followed by the aggregated volumetric video 516C, so that after selecting "Next" from time T1, the user 514 next sees the volumetric video 516B at time T2, and upon selecting "Next" again, sees the aggregated volumetric video 516C at time T3. The user can also navigate backward, e.g., B. to return from aggregated volumetric video 516C to volumetric video 516B, etc.

[0082] Fig. Figure 6 shows a functional block diagram of a volumetric video module 602 according to an exemplary embodiment. The volumetric video module 602 is related to the volumetric video module 200 from Fig. 1 and Fig. 2 similar, with the difference that the volumetric video module 602 includes a multiview module 610 which uses preview images 609 to create a multiview selection image 618A.

[0083] A multiview selection screen 618A comprises a display of two or more available volumetric videos, which the user can use as a menu to select a volumetric video. For example, in the illustrated embodiment, the volumetric video module 602 receives the source volumetric videos 206A-206C and generates a fourth volumetric video, which is an aggregation of the three source volumetric videos 206A-206C. The volumetric video module 602 thus has four possible volumetric videos from which the user 614 can select. The user, viewing via the headset 616, is presented with a multiview selection screen 618A, which is generated by the multiview module 610 using the preview images 609 of each of the four available volumetric videos.The multiview selection screen 618A allows user 614 to select one of the available volumetric videos by choosing the corresponding preview image in the multiview selection screen 618A. For example, if the user selects the upper left field, this selection is transmitted as user input to the aggregation engine 604. The aggregation engine 604 then provides the volumetric video data for the selected volumetric video to the rendering engine 612. The rendering engine 612 then renders the volumetric video (in this example, volumetric video 618B was selected) on user 614's headset 616.

[0084] Fig. Figure 7 shows a flowchart of an exemplary process 700 for aggregating volumetric videos according to an exemplary embodiment. In a particular embodiment, the volumetric video module 200 performs the following actions: Fig. 2, the volumetric video module 502 from Fig. 5 or the volumetric video module 602 from Fig. 6 the process 700 out.

[0085] In the illustrated embodiment, the process in block 702 first identifies source volumetric videos associated with an aggregation rule. Next, in block 704, the process determines whether the aggregation rule applies to specific objects within the video. If so, in block 706, the process segments the video data of the identified source volumetric videos. Subsequently, in block 708, the process extracts the object specified by the aggregation rule from the segmented video data, for example, to insert the object into the aggregated volumetric video generated in block 714. However, the process first checks whether there are any further objects specified by the rule (block 710) and any further source volumetric videos to be processed (block 712). Once all source volumetric videos have been checked for objects specified by the aggregation rules, the process proceeds to block 714.In block 714, the process generates the aggregated volumetric video according to the aggregation rules, i.e., including the specified objects extracted from the individual source volumetric videos.

[0086] Fig. Figure 8 shows a flowchart of an exemplary process 800 for aggregating volumetric videos using an optimization technique according to an exemplary embodiment. In a particular embodiment, the volumetric video module 200 performs Fig. 2, the volumetric video module 502 from Fig. 5 or the volumetric video module 602 from Fig. 6 the process 800 out.

[0087] In the illustrated embodiment, the process in block 802 retrieves an aggregation rule for a user requesting an aggregated volumetric video. The process then compares the requesting user's aggregation rule with a set of aggregation rules from other users in block 804. More specifically, the process searches for a matching rule from another user for whom an aggregated video is already being processed or generated. If a matching rule is found in block 806, the process performs an optimization technique, providing the requesting user with the aggregated video corresponding to the matching rule. This prevents redundant aggregation processing and reduces the system load. If no matching rule is found in block 806, the process generates an aggregated volumetric video according to the requesting user's aggregation rule, for example, following process 700. Fig. 7.

[0088] Fig. Figure 9 shows a flowchart of an exemplary process 900 for aggregating volumetric videos using a safety technique according to an exemplary embodiment. In a particular embodiment, the volumetric video module 200 performs Fig. 2, the volumetric video module 502 from Fig. 5 or the volumetric video module 602 from Fig. 6 the process 900 out.

[0089] In the illustrated embodiment, the process in block 902 identifies source volumetric videos associated with an aggregated video request from a requesting user. The process then compares an authorization rule associated with the requesting user with access rules for each of the identified source videos in block 904. In block 906, the process determines whether the requesting user is authorized to view the identified source videos. If the user is not authorized to view all identified source videos, the request is rejected in block 910. Otherwise, the process in block 908 generates an aggregated volumetric video according to the requesting user's aggregation rule, for example, according to process 700. Fig. 7.

[0090] The following definitions and abbreviations are to be used for the interpretation of the claims and the description. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements but may also include other elements not expressly listed or inherent to this type.

[0091] Furthermore, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" need not necessarily be construed as preferable or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are to be understood as including any whole number greater than or equal to one, i.e., one, two, three, four, etc. The terms "a multitude" are to be understood as including any whole number greater than or equal to two, i.e., two, three, four, five, etc. The term "connection" may include an indirect "connection" and a direct "connection."

[0092] References in the description to "an embodiment," "an exemplary embodiment," etc., mean that the described embodiment may contain a particular feature, structure, or property, but not every embodiment necessarily contains that particular feature, structure, or property. Such formulations also do not necessarily refer to the same embodiment. Furthermore, if a particular feature, structure, or property is described in connection with an embodiment, it is assumed that it is within the knowledge of a person skilled in the art to realize such a feature, structure, or property in connection with other embodiments as well, regardless of whether this is explicitly described.

[0093] The terms "approximately", "essentially", "about" and their variations are to be understood as encompassing the measurement error in determining the respective quantity based on the equipment available at the time of application. For example, "approximately" may encompass a range of ±8%, 5%, or 2% of a given value.

[0094] The description of the various embodiments of the present invention is provided for illustrative purposes only and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used has been chosen to best explain the principles of the embodiments, their practical application, or the technical improvements over technologies available on the market, or to enable those skilled in the art to understand the described embodiments.

[0095] The description of the various embodiments of the present invention is provided for illustrative purposes only and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used has been chosen to best explain the principles of the embodiments, their practical application, or the technical improvements over technologies available on the market, or to enable those skilled in the art to understand the described embodiments.

[0096] Thus, a computer-implemented method, system or device, and computer program product in the exemplary embodiments are provided for managing participation in online communities and other related features, functions, or operations. Where an embodiment or part thereof is described in connection with a device type, the computer-implemented method, system or device, computer program product, or part thereof is adapted or configured for use with a suitable and comparable embodiment of that device type.

[0097] When an embodiment is described as being implemented in an application, the provision of the application in a Software-as-a-Service (SaaS) model is provided within the framework of the exemplary embodiments. In a SaaS model, the capability of the application implementing the embodiment is provided to the user by running the application in a cloud infrastructure. The user can access the application with a variety of client devices via a thin-client interface such as a web browser (e.g., web-based email) or other lightweight client applications. The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage of the cloud infrastructure. In some cases, the user may not even manage or control the capabilities of the SaaS application.In some other cases, the SaaS implementation of the application may allow a possible exception for user-specific application configuration settings.

[0098] The present invention can be a system, a method, and / or a computer program product at any possible level of technical integration. The computer program product can comprise a computer-readable storage medium (or media) on which computer-readable program instructions are stored that cause a processor to execute aspects of the present invention.

[0099] The computer-readable program instructions described herein can be downloaded to the respective computer / processing devices from a computer-readable storage medium or to an external computer or storage device via a network, such as the internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network card or network interface in each computer / processing device receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computer / processing device.

[0100] Computer-readable program instructions for performing operations of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status data, configuration data for integrated circuits, or source code or object code written in one or more programming languages, including an object-oriented programming language such as Smalltalk, C++, or similar, and procedural programming languages ​​such as C or similar. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter scenario, the remote computer can be connected to the user's computer via any network, including a local area network (LAN) or a wide area network (WAN), or the connection can be established to an external computer (e.g., via the internet with an internet service provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute the computer-readable program instructions by using state information from the computer-readable program instructions to personalize the electronic circuitry to implement aspects of the present invention.

[0101] Aspects of the present invention are described herein with reference to flowchart representations and / or block diagrams of processes, devices (systems), and computer program products according to embodiments of the invention. It is understood that each block of the flowchart representations and / or block diagrams, and combinations of blocks in the flowchart representations and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processor of a general or special computer or other programmable data processing apparatus to create a machine such that the instructions executed by the processor of the computer or other programmable data processing apparatus provide means for implementing the functions / actions specified by the flowchart and / or block diagram block or blocks.These computer-readable program instructions can also be stored in a computer-readable storage medium that can instruct a computer, a programmable data processing apparatus and / or other devices to function in a certain manner, such that the computer-readable storage medium with instructions stored therein comprises an object of manufacture containing instructions for implementing the functional / action aspect specified by the flowchart and / or block diagram block or blocks.

[0103] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to execute a sequence of operational steps on the computer, other programmable apparatus, or other device, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified by the flowchart and / or block diagram block or blocks.

[0104] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this context, each block in the flowchart or block diagram can represent a module, segment, or part of instructions that includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions represented in the blocks may occur in a different order. For example, two consecutive blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved.It is also noted that each block of the block diagrams and / or flowchart representations and combinations of blocks in the block diagrams and / or flowchart representations can be implemented by special hardware-based systems that perform the specified functions or actions, or execute combinations of special hardware and computer instructions.

[0105] Embodiments of the present invention can also be provided as part of a service relationship with a client company, a non-profit organization, a government agency, an internal organizational structure, or the like. Aspects of these embodiments may include configuring a computer system to perform and providing software, hardware, and web services that implement some or all of the methods described herein. Aspects of these embodiments may also include analyzing the client's operations, generating recommendations based on the analysis, building systems that implement parts of the recommendations, integrating the systems into existing processes and infrastructures, measuring system usage, allocating costs to system users, and billing for system usage.Although the above embodiments of the present invention have each been described by stating their individual advantages, the present invention is not limited to any particular combination thereof. On the contrary, such embodiments can also be combined with one another in any number and at any time, in accordance with the intended provision of the present invention, without losing their advantageous effects.

Claims

[1] Computer-implemented method comprising: Select, using a first attribute of a first object, the first object in a first volumetric video; Selecting, using a second attribute of a second object, the second object in a second volumetric video, where the first attribute and the second attribute satisfy an aggregation rule; and Generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video involves simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule. [2] Method according to claim 1, wherein the selection of the first object in the first volumetric video is carried out using an instance segmentation process, wherein the instance segmentation process comprises classifying a first part of the first volumetric video as a representation of the first object with the first attribute. [3] The method of claim 2, wherein the instance segmentation process further comprises: Extracting an image segment from a frame of the first volumetric video; Classifying the extracted image segment using a trained, machine learning-based image classifier, wherein the image classifier outputs a segment classification in response to the extracted image segment; Determine that the segment classification output by the image classifier is associated with the first object with the first attribute; Specify, in response to determining that the segment classification is associated with the first object with the first attribute, the extracted image segment as a representation of at least one part of the first object, such that the first part of the first volumetric video contains the extracted image segment. [4] Method according to claim 2, further comprising: Extracting the first part of the first volumetric video from frames of the first volumetric video; and Inserting the extracted first part of the first volumetric video into frames of a third volumetric video. [5] Method according to claim 4, further comprising: Extracting a second part of the second volumetric video from frames of the second volumetric video, where the second part represents the second object with the second attribute; and Inserting the extracted second part of the second volumetric video into frames of a third volumetric video. [6] Method according to claim 1, further comprising: Transferring the aggregated volumetric video to a user device; Capturing user input that displays a selection of the first volumetric video; and Switching, in response to user input, from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device. [7] Method according to claim 1, further comprising: Transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; Capturing user input that displays a selection of the preview image; and Transmitted to the user device in response to user input of the aggregated volumetric video. [8] Method according to claim 1, further comprising: Capturing user input that displays a selection of the first volumetric video; and Determine whether the user is authorized to view the first volumetric video by comparing an authorization rule associated with the user with an access rule associated with the first volumetric video. where the generation of the aggregated volumetric video occurs in response to the determination that the user is authorized to view the first volumetric video. [9] Method according to claim 1, further comprising: Comparing a first set of aggregation rules associated with a first user with a second set of aggregation rules associated with a second user; Determine that the first set of aggregation rules is identical to the second set of aggregation rules, where both the first set of aggregation rules and the second set of aggregation rules contain said aggregation rule; and Transmitted, in response to the detection that the first set of aggregation rules matches the second set of aggregation rules, of the aggregated volumetric video to the first user and to the second user. [10] Method according to claim 9, further comprising: Specify, in response to the detection, that the first set of aggregation rules matches the second set of aggregation rules of the first user and the second user for joint aggregation processing. [11] Method according to claim 10, further comprising: Transmitted, in response to the specification of the first user and the second user for joint aggregation processing, the aggregated volumetric video is sent to a first user device associated with the first user and to a second user device associated with the second user. [12] Computer program product comprising one or more computer-readable storage media and program instructions stored thereon, wherein the program instructions are executable by a processor to cause the processor to perform the following: Select, using a first attribute of a first object, the first object in a first volumetric video; Selecting, using a second attribute of a second object, the second object in a second volumetric video, where the first attribute and the second attribute satisfy an aggregation rule; and Generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video involves simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule. [13] Computer program product according to claim 12, wherein the stored program instructions are stored in a computer-readable storage device in a data processing system and wherein the stored program instructions are transmitted over a network from a remote data processing system. [14] Computer program product according to claim 12, wherein the stored program instructions are stored in a computer-readable storage device in a server data processing system and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer-readable storage device associated with the remote data processing system, further comprising: Program instructions for measuring the usage of the program instructions associated with the request; and Program instructions for generating an invoice based on measured usage. [15] Computer program product according to claim 12, wherein the operations further comprise: Transferring the aggregated volumetric video to a user device; Capturing user input that displays a selection of the first volumetric video; and Switching, in response to user input, from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device. [16] Computer program product according to claim 12, wherein the operations further comprise: Transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; Capturing user input that displays a selection of the preview image; and Transmitted to the user device in response to user input of the aggregated volumetric video. [17] Computer program product according to claim 12, wherein the operations further comprise: Capturing user input that displays a selection of the first volumetric video; and Determine whether the user is authorized to view the first volumetric video by comparing an authorization rule associated with the user with an access rule associated with the first volumetric video. where the generation of the aggregated volumetric video occurs in response to the determination that the user is authorized to view the first volumetric video. [18] Computer system comprising a processor and one or more computer-readable storage media, and program instructions stored thereon, wherein the program instructions are executable by the processor to cause the processor to perform the following: Select, using a first attribute of a first object, the first object in a first volumetric video; Selecting, using a second attribute of a second object, the second object in a second volumetric video, where the first attribute and the second attribute satisfy an aggregation rule; and Generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video involves simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule. [19] Computer system according to claim 18, wherein the operations further comprise: Transferring the aggregated volumetric video to a user device; Capturing user input that displays a selection of the first volumetric video; and Switching, in response to user input, from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device. [20] Computer system according to claim 18, wherein the operations further comprise: Transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; Capturing user input that displays a selection of the preview image; and Transmitted to the user device in response to user input of the aggregated volumetric video.