Personalized Aggregation of Volumetric Video
The aggregated volumetric video system addresses limitations of fixed viewpoints by combining multiple source videos with user-defined rules, allowing for enhanced viewer interaction and immersion.
Patent Information
- Application Number
- JP2025530735
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-11-24
- Publication Date
- 2026-01-06
AI Technical Summary
Existing volumetric video systems limit viewers to a single created viewpoint, restricting their freedom of movement and inability to select content from multiple source videos.
An aggregated volumetric video system that combines multiple source videos based on user-defined aggregation rules, allowing for personalized rendering of objects and attributes, enabling six degrees of freedom and enhanced viewer interaction.
Enables viewers to freely navigate and personalize the content within a volumetric scene, overcoming limitations of fixed viewpoints and enhancing immersion.
Smart Images

Figure 2026500117000001_ABST
Abstract
Description
[Background technology]
[0001] The present invention relates generally to volumetric video processing, and more particularly to a method, system, and computer program for personalized aggregation of volumetric video.
[0002] Immersive technologies continue to gain popularity, particularly associated consumer devices such as virtual reality (VR) headsets, and prices continue to fall. Growing consumer acceptance of such technologies has led to growth in content as content creators respond to consumer interest in new immersive experiences. Additionally, advancements in video capture technology are enabling content creators to increase the "immersive" nature of their content.
[0003] Immersive video content is generally captured using multiple cameras simultaneously from different angles, or using a single camera from multiple locations and angles. One example of immersive filming is so-called 360-degree filming. The resulting 360-degree video content generally has a limited field of view, but is created computer-generated by stitching together multiple simultaneously captured images to form an entire sphere of still or video images within which an individual may be positioned. It is most easily viewed by an individual with a VR headset, but can also be viewed in other ways. It is also somewhat limited because the individual's viewpoint in the scene is fixed relative to the image itself. In other words, the viewer can only view such a scene from a position selected by the filmmaker, thereby limiting movement within the scene. While the viewer can look around in all directions, they cannot move from the location of the physical camera. Additionally, traditional 360-degree video content sacrifices depth and volumetric content because it is effectively a sphere with the viewer at the center and photographs attached along the interior walls of the sphere. There are no objects in the scene that have shapes other than the walls of this sphere, which further reduces the immersiveness of the experience by limiting the viewer's experience.
[0004] Volumetric video is distinguished from 360° video in that volumetric video also captures depth information using photogrammetry or depth-of-field sensors (e.g., light field arrays, LIDAR). This information results in volumetric video capturing both an image of a scene and the overall three-dimensional parameters of objects in the scene. So, for example, a chair in a given volumetric video scene may have both a shape (e.g., a three-dimensional geometry corresponding to that of the chair) and an image superimposed on it to create the impression that it is made from wood, metal, or plastic, or whatever. Thus, in volumetric video, the viewer can generally move freely within the scene, overcoming the movement constraints of traditional two-dimensional or 360° videography techniques. Content produced using these techniques is often referred to as volumetric, six-degrees-of-freedom (6DoF), light field, or free-viewpoint video. Summary of the Invention
[0005] An exemplary embodiment provides personalized aggregation of volumetric videos. One embodiment comprises selecting a first object in a first volumetric video using a first attribute of the first object. This embodiment also comprises selecting a second object in a second volumetric video using a second attribute of the second object, where the first attribute and the second attribute satisfy an aggregation rule. This embodiment also comprises generating an aggregated volumetric video from the first volumetric video and the second volumetric video, where generating the aggregated video comprises simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of an embodiment.
[0006] One embodiment includes a computer usable program product, the computer usable program product including a computer readable storage medium and program instructions stored on the storage medium.
[0007] One embodiment includes a computer system comprising a processor, a computer-readable memory, and a computer-readable storage medium, with program instructions stored on the storage medium for execution by the processor via the memory. [Brief explanation of the drawings]
[0008] The novel features believed characteristic of this invention are set forth in the appended claims. However, the invention itself, together with its preferred mode of use, further objects and advantages thereof, will best be understood by reference to the following detailed description of illustrative embodiments when read in connection with the accompanying drawings.
[0009] [Figure 1] FIG. 1 is a block diagram of a computing environment according to an illustrative embodiment.
[0010] [Figure 2] FIG. 2 is a functional block diagram of an exemplary volumetric video module according to an exemplary embodiment.
[0011] [Figure 3] FIG. 1 is a functional block diagram of an exemplary video processing environment according to an exemplary embodiment.
[0012] [Figure 4] FIG. 1 illustrates an aggregation process according to an exemplary embodiment.
[0013] [Figure 5] FIG. 2 is a functional block diagram of a volumetric video module according to an exemplary embodiment.
[0014] [Figure 6] FIG. 2 is a functional block diagram of a volumetric video module according to an exemplary embodiment.
[0015] [Figure 7] 1 is a flowchart of an example process for aggregating volumetric video, according to an example embodiment.
[0016] [Figure 8] 1 is a flowchart of an example process for aggregating volumetric videos using aggregation optimization techniques, according to an example embodiment.
[0017] [Figure 9] 1 is a flowchart of an example process for aggregating volumetric video using security techniques, according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] Volumetric video is distinguished from 360° video in that volumetric video includes depth information. This added depth information dramatically changes the way content can be consumed. When viewing a scene in a 360° video format, the viewer is locked into a single position, and from that single vantage point, the viewer may have up to three degrees of freedom, corresponding to rotation about each of the three orthogonal axes of a Cartesian coordinate system: roll (tilting the head left or right), pitch (tilting the head forward or backward), and yaw (rotating the head left or right).
[0019] In contrast, the depth information included with volumetric video frees the viewer from locked viewpoints and allows the viewer to have up to six degrees of freedom, corresponding to rotation about and translation along each of the three orthogonal axes of a Cartesian coordinate system, i.e., elevating (moving up or down), strafing (moving left or right), and surging (moving forward or backward), in addition to roll, pitch, and yaw as described above. As a result, the viewer can move freely around in a volumetric video scene and observe objects from multiple angles and viewpoints. This added freedom of movement significantly enhances the immersive nature of content delivered in volumetric video formats compared to earlier technologies.
[0020] However, a limitation of currently existing volumetric videos is that viewers are still limited to viewing only the volumetric video scene that was originally created. That is, viewers have no control over the content of the volumetric video. For example, if a user remotely participates in a conference that is taking place simultaneously at two or more locations, from which each volumetric video is made available to remote participants, the user is limited to viewing only one of the two volumetric videos at a time. In other words, the user cannot select the option of providing a volumetric video that includes an element of interest from each of the source volumetric videos.
[0021] The disclosed embodiments address these and other limitations of conventional volumetric video systems by providing an aggregated volumetric video that is an aggregation of two or more source volumetric videos. The disclosed embodiments also provide for personalization of the aggregated volumetric video. For example, in some embodiments, a user specifies one or more aggregation rules.
[0022] Aggregation rules are primarily instructions and / or preferences that guide the rendering of two or more source volumetric videos into a single aggregated volumetric video. Aggregation rules may include designating objects from the source volumetric videos to show or not show in the aggregated volumetric video (e.g., showing an audience from a first source instead of an audience from a second source). Aggregation rules may include designating attributes of objects from the source volumetric videos to show or not show or to modify in the aggregated volumetric video (e.g., changing the hue of audience chairs from blue to black; blurring company logos on audience members' shirts and hats).
[0023] The aggregation rules may include rules of different specificity. For example, the aggregation rules may include designation of a particular attribute, a particular object, or a class of objects, where the classes of objects may differ in terms of specificity. As a more specific, non-limiting example, a user creating rules for personalized aggregated volumetric videos may have source volumetric videos from multiple car shows; in this scenario, rules of different specificity may include rules for broad classes of objects, such as a preference for viewing sedans over other types of vehicles; a rule for a more specific class of object may specify a particular make and model of a sedan; or a rule for a specific object may specify a famous car on display in one of the volumetric videos.
[0024] In some embodiments, the video aggregation is performed by a volumetric video module. In some such embodiments, the volumetric video module analyzes digital volumetric video data from two or more source volumetric videos. In some such embodiments, this analysis results in the identification of one or more objects in each of the source volumetric videos.
[0025] In some embodiments, the analysis generates metadata for one or more of the identified objects. In some embodiments, the metadata for the identified objects includes data representing one or more attributes of the identified objects. In various embodiments, the number and types of attributes will vary depending on implementation decisions, constraints, and / or preferences. Non-limiting examples of types of attributes include appearance, classification, and / or source attributes. In an exemplary embodiment, appearance attributes of the identified objects may include size, shape, and / or hue attributes of the identified objects.
[0026] In example embodiments, the classification attributes of an identified object may include one or more classes or categories associated with the identified object. In some embodiments, the classification attributes include a class (or classes) predicted using known object classification and / or detection techniques.
[0027] For example, in some embodiments, the object classification technique includes a machine learning process that uses a trained machine learning model to predict categories of objects in an image. It will be understood that "images," as referred to herein, include images that are frames of video. In some embodiments, the classification is performed using known semantic image segmentation techniques that classify portions or segments of a volumetric image with the corresponding classes that the portions or segments of the volumetric image represent.
[0028] In some embodiments, object detection includes techniques that combine classification with localization techniques to determine the location of classified objects in an image. Thus, in some embodiments, object detection is performed using known techniques to determine what objects are in an image and specify where the objects are positioned within the image. In some embodiments, object detection includes known instance segmentation techniques to distinguish between distinct objects of the same class in an image.
[0029] In some embodiments, the classification attribute may include multiple classes representing different degrees of specificity. For example, assuming the identified object is a multi-passenger van, in this example, the identified object may include "vehicle" as a first classification attribute, "multi-passenger vehicle" as a second classification attribute, and "van" as a third classification attribute.
[0030] In an exemplary embodiment, the source attribute may include information identifying one or more sources of the volumetric video. As a non-limiting example, assume that an aggregated volumetric video is being generated for a conference occurring simultaneously at first and second geographically separated locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video captured at the first location and a second volumetric video captured at the second location. Thus, a first identified object captured in the first volumetric video may include a source attribute indicating the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video may include a source attribute indicating the second location or the second volumetric video as the source of the second identified object.
[0031] The set of size and shape attributes may include only a single attribute, as well as two or more attributes. Size and shape attributes are characteristics, features, or other features of the shape of an object and / or one or more portions of an object that are associated with a size measurement. A portion of an object may be a portion, section, or component of an object. A portion of an object may include the entire object. A size measurement is any measure of size, such as, but not limited to, height, length, and width. A size measurement may include the weight and volume of an object. The set of size and shape attributes may include only size-related attributes without any shape-related attributes. In another embodiment, the set of size and shape attributes may include only shape-related attributes without any size-related attributes. In yet another embodiment, the set of size and shape attributes includes both size-related attributes and shape-related attributes.
[0032] An embodiment may be implemented as a software application. An application implementing an embodiment may be configured as a modification to an existing manufacturing system, as a separate application operating in conjunction with an existing manufacturing system, as a stand-alone system, or some combination thereof.
[0033] One embodiment monitors system status data for indications of system failures that cause the system kernel to transition to a halted state. The status data may vary depending on the type of system (e.g., operating system and hardware), but may generally include system error messages, error codes, log entries, or other data representative of system errors. The system errors may also vary depending on the type of system, but may generally include kernel errors, kernel panics, stop errors, or the like that cause partial or complete loss of kernel functionality.
[0034] A system failure that causes a partial or complete loss of kernel functionality is generally recognized by the system as a condition that requires a system reboot to restore. Loss of kernel functionality typically requires a reboot to restore in most types of systems. Thus, such errors are examples of errors that meet a reboot condition.
[0035] In exemplary embodiments, upon detection of a system error that satisfies a restart condition (e.g., a system error that causes loss of kernel functionality), debug data is temporarily saved to a protected section of memory. Because at least a portion of the kernel functionality will have been lost at this point, embodiments include generating and saving a copy of the debug data without kernel assistance. For example, in some embodiments, an exception handler generates and saves a copy of the debug data. The debug data may vary depending on the type of system (e.g., operating system and hardware), but may generally include data from processor memory (e.g., registers and caches), logs, and / or trace arrays.
[0036] Since the kernel is in a halted state while the debug data is being generated, this means that kernel functions for processing the debug data are not available. As a result, the kernel is not available at this time to filter sensitive information from the debug data. For this reason, the debug data is stored in a protected section of memory where it is preserved across reboots and can be processed thereafter.
[0037] After the reboot occurs, an untrusted debug device is connected to the restored system via an I / O port using a trusted protocol. The untrusted device issues a request for debug data that will be used to attempt to determine the reason for the system failure. Because the debug data is stored in protected memory, the untrusted device is not able to access the debug data directly. Instead, the untrusted entity must request the debug data from the secure debug module.
[0038] In the illustrated embodiment, the secure debug module receives a request for debug data, and in response to the request, the secure debug module uses the data sanitization module to analyze the debug data using a sensitive data detection / sanitization process that detects and removes sensitive data in the debug data.
[0039] In some embodiments, the data sanitization module detects sensitive data according to an audit policy. An audit policy is a set of preferences, rules, and / or criteria for protecting sensitive data in debug data. For example, an audit policy may define a “sensitive object” as a file or object that contains a particular keyword (e.g., “secret” or “privileged”) and / or is associated with a particular keyword (e.g., in metadata) or a particular flag (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy may further specify rules for handling sensitive objects. As an example, an audit policy may require that a reviewer approve the transfer of any potentially sensitive object from protected memory to an untrusted entity. Therefore, in some embodiments, the data protection process includes one or more protective measures in the form of data sanitization to prevent sensitive data from being leaked in debug data sent to an untrusted entity.
[0040] In the illustrated embodiment, the window module detects whether sensitive data was processed during a time window in which a system error occurred, for example, by examining a system log. For example, in some embodiments, the window module examines the system log using an audit policy that includes a set of preferences, rules, and / or criteria that the window module uses to identify sensitive data in the debug data. For example, the audit policy may define a “sensitive object” as a file or object that contains a particular keyword (e.g., “secret” or “privileged”) and / or is associated with a particular keyword (e.g., in metadata) or a particular flag (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy may further specify rules for handling sensitive objects. As an example, the audit policy may require that a reviewer approve the transfer of any potentially sensitive object from protected memory to an untrusted entity. Therefore, in some embodiments, the data protection process includes one or more protective measures in the form of time window analysis to prevent sensitive data from being leaked in debug data sent to an untrusted entity.
[0041] In some embodiments, the data sanitization module detects sensitive data according to an audit policy. An audit policy is a set of preferences, rules, and / or criteria for protecting sensitive data in debug data. For example, an audit policy may define a “sensitive object” as a file or object that contains a particular keyword (e.g., “secret” or “privileged”) and / or is associated with a particular keyword (e.g., in metadata) or a particular flag (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy may further specify rules for handling sensitive objects. As an example, an audit policy may require that a reviewer approve the transfer of any potentially sensitive object from protected memory to an untrusted entity. Therefore, in some embodiments, the data protection process includes one or more safeguards in the form of data encryption to prevent sensitive data from being leaked in debug data sent to an untrusted entity.
[0042] For clarity of explanation, and without implying any limitations, the exemplary embodiments are described using several example configurations. From this disclosure, one skilled in the art will recognize many variations, adaptations, and modifications of the described configurations to achieve the described objectives, which are contemplated within the scope of the exemplary embodiments.
[0043] Furthermore, a simplified illustration of a data processing environment is used in the figures and exemplary embodiments. In an actual computing environment, additional structures or components not shown or described herein, or structures or components for similar functionality but different from those described herein, may be present without departing from the scope of the exemplary embodiments.
[0044] Furthermore, the exemplary embodiments are described with reference to specific actual or hypothetical components, merely as examples. Any particular manifestation of these and other similar artifacts is not intended to limit the invention. Any suitable manifestation of these and other similar artifacts may be selected within the scope of the exemplary embodiments.
[0045] The examples in this disclosure are used only for clarity of explanation and are not limited to exemplary embodiments. Any advantages listed herein are examples only and are not intended to limit the exemplary embodiments. Additional or different advantages may be realized by particular exemplary embodiments. Furthermore, particular exemplary embodiments may include some, all, or none of the above-listed advantages.
[0046] Furthermore, exemplary embodiments may be implemented with respect to any type of data, data source, or access to a data source via a data network. Any type of data storage device may provide data to an embodiment of the present invention, either locally at a data processing system or via a data network, within the scope of the present invention. Where an embodiment is described using a mobile device, any type of data storage device suitable for use with a mobile device may provide data to such an embodiment, either locally at the mobile device or via a data network, within the scope of exemplary embodiments.
[0047] The exemplary embodiments are described using specific code, computer-readable storage media, high-level features, designs, architectures, protocols, layouts, diagrams, and tools as examples only, and are not limited to the exemplary embodiments. Furthermore, for clarity of explanation, the exemplary embodiments are described in some instances using specific software, tools, and data processing environments as examples only. The exemplary embodiments may be used in conjunction with other equivalent or similarly purposed structures, systems, applications, or architectures. For example, other equivalent mobile devices, structures, systems, applications, or architectures thereof may be used in conjunction with such embodiments of the present invention within the scope of the present invention. The exemplary embodiments may be implemented in hardware, software, or a combination thereof.
[0048] The examples in this disclosure are used for clarity of explanation only and are not intended to be limiting of the exemplary embodiments. Additional data, operations, actions, tasks, activities, and manipulations are recognized from this disclosure and are contemplated within the scope of the exemplary embodiments.
[0049] Various aspects of the present disclosure are described through text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.
[0050] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated through wires, and / or other transmission media.As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but the above does not make the storage device temporary, as the data is not temporary while it is stored.
[0051] Referring to Figure 1, this figure illustrates a block diagram of a computing environment 100. The computing environment 100 includes an example environment for the execution of at least a portion of the computer code involved in performing the methods of the present invention, such as an improved volumetric video module 200 that provides personalized aggregation of volumetric video. In addition to the volumetric video module 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, computer 101 includes a set of processors 110 (including processing circuitry 120 and cache 121), a communications fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and volumetric video module 200 as identified above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0052] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the art of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or among multiple locations. While in this presentation of computing environment 100, to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not depicted in the cloud of FIG. 1 , it may be located in a cloud. However, computer 101 is not required to reside within a cloud except to any extent that may be expressly indicated.
[0053] Processor set 110 includes one or more computer processors of any type now known or later developed. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple coordinated integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computing.
[0054] Computer-readable program instructions are typically loaded onto computer 101 to cause processor set 110 of computer 101 to perform a series of operational steps, thereby realizing a computer-implemented method, whereby the instructions so executed instantiate the method specified in the flowcharts and / or descriptions of the computer-implemented method contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for performing the methods of the present invention may be stored in volumetric video module 200 in persistent storage 113.
[0055] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive pathways, such as those that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication pathways may be used, such as fiber optic and / or wireless communication pathways.
[0056] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic type random access memory (RAM) or static type RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0057] Persistent storage 113 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data remains regardless of whether power is supplied to computer 101 and / or to persistent storage 113 directly. Persistent storage 113 may be read-only memory (ROM), but typically at least a portion of persistent storage allows data to be written, data to be deleted, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 122 may take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems employing a kernel. The code contained in volumetric video module 200 typically includes at least a portion of the computer code involved in performing the methods of the present invention.
[0058] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 may be implemented in various ways, such as Bluetooth connections, Near-Field Communication (NFC) connections, connections made by cable (such as a universal serial bus (USB)-type cable), insertion-type connections (e.g., a secure digital (SD) card), connections made over a local area communication network, and even connections made over a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage 124 may be external storage, such as an external hard drive, or insertable storage, such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (e.g., computer 101 stores and manages large databases locally), this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0059] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for implementing the methods of the present invention can be downloaded to computer 101 from an external computer or external storage device, typically through a network adapter card or network interface included in network module 115.
[0060] WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances by any now known or later developed technology for communicating computer data. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.
[0061] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and may take any of the forms discussed above in connection with computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to the end user, the recommendations would typically be communicated from network module 115 of computer 101 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be a client device such as a thin client, a heavy client, a mainframe computer, a desktop computer, etc.
[0062] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0063] Public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer functionality, particularly data storage (cloud storage) and computing power, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 105's computing resources is performed by computer hardware and / or software in cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments (VCEs) running on various computers that comprise host physical machine set 142, the universe of physical computers in and / or available to public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It is understood that these VCEs may be stored as images and can be transferred among and between various physical machine hosts, either as images or after instantiation of the VCEs. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.
[0064] Some further description of a virtualized computing environment (VCE) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from the image. Two well-known types of VCE are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for the existence of multiple isolated user space instances called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running in them. A computer program running on a normal operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices assigned to the container; this feature is known as containerization.
[0065] A private cloud 106 is similar to a public cloud 105, except that the computing resources are available only for use by a single enterprise. While the private cloud 106 is shown as communicating with the WAN 102, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often each implemented by a different vendor. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0066] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, reported, and billed, providing transparency to both providers and consumers of utilized services.
[0067] 2, which illustrates a functional block diagram of an exemplary volumetric video module 200 according to an exemplary embodiment. In the illustrated embodiment, the volumetric video module 200 receives multiple volumetric videos 206A-206C and aggregates two or more of them into an aggregated volumetric video 208.
[0068] In the illustrated embodiment, volumetric video module 200 includes aggregation engine 202 and aggregation rules 204 stored on a computer-readable storage medium. In alternative embodiments, volumetric video module 200 may include some or all of the functionality described herein, but grouped differently into one or more modules. In some embodiments, the functionality described herein is distributed among multiple systems, which may include a combination of software and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0069] In the illustrated embodiment, aggregation engine 202 aggregates individual volumetric videos, such as two or more of volumetric videos 206A-206C, to create a single aggregated volumetric video 208, also referred to herein as corridor 208. In some embodiments, volumetric videos 206A-206C are captured simultaneously from different physical locations. For example, volumetric video 206A may be a live broadcast in volumetric video format from Los Angeles, volumetric video 206B may be a live broadcast in volumetric video format from Chicago, and volumetric video 206C may be a live broadcast in volumetric video format from New York, for a conference hosted in these three locations. In some embodiments, volumetric videos 206A-206C are captured at different times from different physical locations.
[0070] The aggregation rules 204 are primarily instructions and / or preferences that guide the rendering of two or more source volumetric videos into a single aggregated volumetric video. The aggregation rules 204 may include designating objects from the source volumetric videos to show or not show in the aggregated volumetric video (e.g., show an audience from a first source instead of an audience from a second source). The aggregation rules 204 may include designating attributes of objects from the source volumetric videos to show or not show, or to modify in the aggregated volumetric video (e.g., change the hue of audience chairs from blue to black; blur company logos on audience members' shirts and hats).
[0071] The aggregation rules 204 may include rules of differing specificity. For example, the aggregation rules 204 may include designations of particular attributes, particular objects, or classes of objects, where the classes of objects may differ in terms of specificity. As a more specific, non-limiting example, a user creating rules 204 for personalized aggregated volumetric videos may have source volumetric videos from multiple car shows; in this scenario, the rules 204 of differing specificity may include rules for broad classes of objects, such as a preference for viewing sedans over other types of vehicles; a rule for a more specific class of object may specify a particular make and model of a sedan; or a rule for a particular object may specify a famous car on display in one of the volumetric videos.
[0072] In some embodiments, aggregation engine 202 performs video aggregation. In some such embodiments, aggregation engine 202 analyzes digital volumetric video data from two or more source volumetric videos. In some such embodiments, this analysis results in the identification of one or more objects in each of the source volumetric videos.
[0073] In some embodiments, aggregation engine 202 generates metadata for one or more of the identified objects. In some embodiments, the metadata for the identified objects includes data representing one or more attributes of the identified objects. In various embodiments, the number and types of attributes will vary depending on implementation decisions, constraints, and / or preferences. Non-limiting examples of types of attributes include appearance, classification, and / or source attributes. In an exemplary embodiment, appearance attributes of the identified objects may include size, shape, and / or hue attributes of the identified objects.
[0074] In example embodiments, the classification attributes of an identified object may include one or more classes or categories associated with the identified object. In some embodiments, the classification attributes include a class (or classes) predicted using known object classification and / or detection techniques.
[0075] For example, in some embodiments, aggregation engine 202 includes an image classifier 210 that performs object classification techniques that include a machine learning process that uses a trained machine learning model to predict categories of objects in an image. It will be understood that "images," as referred to herein, include images that are frames of video. In some embodiments, classification is performed using known semantic image segmentation techniques that classify portions or segments of a volumetric image with the corresponding classes that the portions or segments of the volumetric image represent.
[0076] In some embodiments, aggregation engine 202 performs object detection using techniques that combine classification with localization techniques to determine the location of classified objects in an image. Thus, in some embodiments, aggregation engine 202 performs object detection using known techniques to determine what objects are in an image and specify where the objects are positioned within the image. In some embodiments, aggregation engine 202 uses object detection techniques that include known instance segmentation techniques to distinguish between distinct objects of the same class in an image.
[0077] In some embodiments, the classification attribute may include multiple classes representing different degrees of specificity. For example, assuming the identified object is a multi-passenger van, in this example, the identified object may include "vehicle" as a first classification attribute, "multi-passenger vehicle" as a second classification attribute, and "van" as a third classification attribute.
[0078] In an exemplary embodiment, the source attribute may include information identifying one or more sources of the volumetric video. As a non-limiting example, assume that an aggregated volumetric video is being generated for a conference occurring simultaneously at first and second geographically separated locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video captured at the first location and a second volumetric video captured at the second location. Thus, a first identified object captured in the first volumetric video may include a source attribute indicating the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video may include a source attribute indicating the second location or the second volumetric video as the source of the second identified object.
[0079] The set of size and shape attributes may include only a single attribute, as well as two or more attributes. Size and shape attributes are characteristics, features, or other features of the shape of an object and / or one or more portions of an object that are associated with a size measurement. A portion of an object may be a portion, section, or component of an object. A portion of an object may include the entire object. A size measurement is any measure of size, such as, but not limited to, height, length, and width. A size measurement may include weight and volume of an object. The set of size and shape attributes may include only size-related attributes without any shape-related attributes. In another embodiment, the set of size and shape attributes may include only shape-related attributes without any size-related attributes. In yet another embodiment, the set of size and shape attributes includes both size-related attributes and shape-related attributes.
[0080] 3, which illustrates a functional block diagram of an exemplary video processing environment 300 according to an exemplary embodiment. In the illustrated embodiment, video processing environment 300 includes server 306A, server 306B, and server 310, each of which may include volumetric video module 200 of FIGS. 1 and 2.
[0081] In the illustrated embodiment, volumetric videos 302A-302D are examples of source volumetric videos. In the illustrated embodiment, volumetric videos 302A-302D are each created using multiple cameras 304. Alternatively, one or more of volumetric videos 302A-302D may be created using a single camera recording from multiple locations and angles.
[0082] 3 illustrates that in addition to aggregating source videos (such as volumetric video 302A-volumetric video 302D), the disclosed embodiments may also aggregate aggregated volumetric videos. The illustrated embodiments also enable aggregating more than two source volumetric videos.
[0083] In the illustrated embodiment, server 306A aggregates volumetric video 302A and volumetric video 302B into aggregated volumetric video 308A, and server 306B aggregates volumetric video 302C and volumetric video 302D into aggregated volumetric video 308B. Server 310 then aggregates aggregated volumetric video 308A and aggregated volumetric video 308B into aggregated volumetric video 312. In alternative embodiments, the aggregated volumetric video may be aggregated using source two-dimensional (2D) video and depth information from camera 304 rather than volumetric video. For example, in some embodiments, server 306A generates aggregated volumetric video 308A using 2D video and depth information from the camera 304 associated with volumetric video 302A and from the camera 304 associated with volumetric video 302B.
[0084] In another embodiment, the aggregated volumetric video may be further aggregated with another source volumetric video. For example, if nothing from volumetric video 302D is needed for aggregated volumetric video 312, server 310 will aggregate aggregated volumetric video 308A with volumetric video 302C into aggregated volumetric video 312.
[0085] 4, which illustrates an aggregation process 400 according to an example embodiment. In one embodiment, the aggregation process 400 is performed by the volumetric video module 200 of FIGS.
[0086] 4 illustrates that in addition to aggregating two volumetric videos (source or aggregated) at a time, alternative embodiments may aggregate three or more volumetric videos (source or aggregated) at a time. In the illustrated embodiment, the user rules specify that volumetric video 402A and volumetric video 402B are aggregated with volumetric video 402C, primarily using volumetric video 402C, but showing the audience at image portion 404 from volumetric video 402A and the audience at image portion 408 from volumetric video 402B instead of the audience at insertion location 406 and insertion location 410 of volumetric video 402C.
[0087] 5, which illustrates a functional block diagram of a volumetric video module 502 according to an example embodiment. The volumetric video module 502 is similar to the volumetric video module 200 of FIGS. 1 and 2, except that an aggregation engine 504 and a rendering engine 510 of the volumetric video module 502 generate an aggregation hierarchy of volumetric videos and / or aggregated volumetric videos according to user navigation input.
[0088] For example, in the illustrated embodiment, the hierarchy of volumetric videos includes volumetric video 516A, volumetric video 516B, and aggregated volumetric video 516C. User 514 is using headset 512 to view volumetric video 516A at time T1. Headset 512 includes or is in communication with volumetric video module 502.
[0089] The aggregation engine 504 receives user input and one or more source volumetric videos 206A-206C. The aggregation engine 504 also has access to aggregation rules 506 and an aggregation hierarchy 508. The description of aggregation rules 204 in FIG. 2 applies equally to aggregation rules 506. The aggregation hierarchy 508 may be a playlist that specifies which volumetric videos to present and the order in which they should be presented. In some embodiments, a user 514 creates the aggregation hierarchy 508. In some embodiments, the volumetric video module 502 automatically generates the aggregation hierarchy 508 according to user-specified rules or preferences (e.g., play the source volumetric videos in order from closest to furthest recording location, then play the aggregated ones).
[0090] In the illustrated embodiment, the headset 512 may include controls that allow simple next / previous navigation of the aggregation hierarchy 508. These controls send user input to the aggregation engine 504, which checks the aggregation hierarchy 508 and then instructs the rendering engine 510 to send the appropriate volumetric video. Thus, in the illustrated embodiment, the aggregation hierarchy 508 includes volumetric video 516A, followed by volumetric video 516B, followed by aggregated volumetric video 516C, such that a user 514 selecting "Next" from time T1 would then view volumetric video 516B at time T2; if the user selects "Next" again, the user would view aggregated volumetric video 516C at time T3. The user may also navigate back, e.g., transitioning from aggregated volumetric video 516C back to volumetric video 516B, and so on.
[0091] 6, which illustrates a functional block diagram of a volumetric video module 602 according to an example embodiment. The volumetric video module 602 is similar to the volumetric video module 200 of FIGS. 1 and 2, except that the volumetric video module 602 includes a multi-view module 610 that uses a preview image 609 to generate a multi-view image 618A.
[0092] The multiview image 618A includes a display of two or more available volumetric videos that can be used as a menu by the user to select a volumetric video to view. For example, in the illustrated embodiment, the volumetric video module 602 receives the source volumetric videos 206A-206C and generates a fourth volumetric video that is an aggregation of the three source volumetric videos 206A-206C. Thus, the volumetric video module 602 has four possible volumetric videos from which the user 614 can select. The user viewing through use of the headset 616 is presented with a multiview image 618A created by the multiview module 610 using preview images 609 from each of the four available volumetric videos. The multiview image 618A allows the user 614 to select one of the available volumetric videos by selecting the corresponding preview image in the multiview image 618A. For example, if the user selects the top left box, this selection is sent to the aggregation engine 604 as user input. The aggregation engine 604 then provides the volumetric video data for the selected volumetric video to the rendering engine 612. The rendering engine 612 then renders the volumetric video (in this example, volumetric video 618B was selected) in the headset 616 of the user 614.
[0093] 7, which illustrates a flowchart of an example process 700 for aggregating volumetric video, according to an example embodiment. In particular embodiments, process 700 is performed by volumetric video module 200 of FIG. 2, volumetric video module 502 of FIG. 5, or volumetric video module 602 of FIG. 6.
[0094] In the illustrated embodiment, the process then identifies a volumetric video source associated with the aggregation rule at block 702. Then, at block 704, the process determines whether the aggregation rule applies to any particular object in the video. If so, then, at block 706, the process segments the video data of the identified volumetric video source. Then, at block 708, the process extracts the object specified by the aggregation rule from the segmented video data, e.g., so that the object can be included in the aggregated volumetric video generated at block 714. First, however, the process checks for other objects specified by the rule (block 710) and for additional volumetric video sources to process (block 712). Once all volumetric video sources have been processed for objects specified by the aggregation rule, the process continues at block 714. At block 714, the process generates an aggregated volumetric video according to the aggregation rule, i.e., to include the specified object extracted from the individual volumetric video sources.
[0095] 8, which illustrates a flowchart of an example process 800 for aggregating volumetric video using an aggregation optimization technique, according to an example embodiment. In particular embodiments, process 800 is performed by volumetric video module 200 of FIG. 2, volumetric video module 502 of FIG. 5, or volumetric video module 602 of FIG. 6.
[0096] In the illustrated embodiment, at block 802, the process retrieves aggregation rules for a user requesting an aggregated volumetric video. Then, at block 804, the process compares the requesting user's aggregation rules with sets of aggregation rules for other users. Specifically, the process searches for matching aggregation rules for other users for whom aggregated videos have already been processed or generated. Then, at block 806, if a matching rule is found, the process performs an optimization technique to provide the requesting user with the aggregated video corresponding to the matching rule. This prevents redundant aggregation processing, thereby reducing the workload of the involved systems. If no matching rule is found at block 806, the process generates the aggregated volumetric video according to the requesting user's aggregation rules, for example, according to process 700 of FIG. 7.
[0097] 9, a flowchart of an example process 900 for aggregating volumetric video using security techniques is shown, in accordance with an example embodiment. In particular embodiments, process 900 is performed by volumetric video module 200 of FIG. 2, volumetric video module 502 of FIG. 5, or volumetric video module 602 of FIG. 6.
[0098] In the illustrated embodiment, at block 902, the process identifies volumetric video sources associated with a request for aggregated video from a requesting user. Then, at block 904, the process compares the permission rules associated with the requesting user with the access rules for each of the identified video sources. At block 906, the process determines whether the requesting user is authorized to view the identified video sources. If the user is not authorized to view all of the identified video sources, at block 910, the request is denied. Otherwise, at block 908, the process generates an aggregated volumetric video according to the requesting user's aggregation rules, for example, according to process 700 of FIG. 7.
[0099] The following definitions and abbreviations are to be used for interpreting the claims and the specification. As used herein, the terms "comprises," "comprising," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or device that includes a list of elements is not necessarily limited to only those elements and may include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or device.
[0100] Additionally, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "plurality" is understood to include any integer greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" can include indirect and direct "connections."
[0101] References herein to "one embodiment," "an embodiment," "an exemplary embodiment," etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but that all embodiments may or may not include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with one embodiment, it is believed to be within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described.
[0102] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measurement of a particular quantity based on equipment available at the time of the filing of this application. For example, "about" can include a range of ±8%, or 5%, or 2% of a given value.
[0103] The description of various embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications, or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments described herein.
[0104] The description of various embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications, or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments described herein.
[0105] Thus, computer-implemented methods, systems, or apparatus, and computer program products are provided in exemplary embodiments for managing participation in online communities and other related features, functions, or operations. Where an embodiment, or portions thereof, are described with respect to a certain type of device, the computer-implemented method, system, or apparatus, computer program product, or portions thereof, is adapted or configured for use with suitable and equivalent designations of that type of device.
[0106] While an embodiment is described as being implemented in an application, delivery as an application in a Software as a Service (SaaS) model is contemplated within the scope of the exemplary embodiment. In the SaaS model, the capabilities of an application implementing an embodiment are provided to users by running the application in a cloud infrastructure. Users can access the application using a variety of client devices through a thin-client interface, such as a web browser (e.g., web-based email) or other lightweight client application. Users do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage of the cloud infrastructure. In some cases, users may not even need to manage or control the capabilities of the SaaS application. In other cases, a SaaS implementation of an application may allow for the possible exception of limited user-specific application configuration settings.
[0107] The present invention may be a system, method, and / or computer program product at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0108] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computing / processing device for storage.
[0109] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, or the like, and procedural programming languages such as the "C" programming language or similar. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions to personalize the electronic circuitry by utilizing state information of the computer-readable program instructions to perform aspects of the present invention.
[0110] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0111] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium, whereby the instructions can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, whereby the computer-readable storage medium having stored thereon instructions comprises an article of manufacture including instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0112] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0113] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions, that implements the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.
[0114] Embodiments of the present invention may be delivered as part of a service engagement with a client company, nonprofit organization, government agency, internal organizational structure, or the like. Aspects of these embodiments may include configuring a computer system to perform and deploying software, hardware, and web services that implement some or all of the methods described herein. Aspects of these embodiments may also include analyzing client behavior, making recommendations in response to the analysis, building a system that implements portions of the recommendations, integrating the system into existing processes and infrastructure, metering system usage, allocating costs to users of the system, and billing for system usage. Although the above embodiments of the present invention have each been described by describing their respective individual advantages, the present invention is not limited to any particular combination thereof. On the contrary, such embodiments may be combined in any manner and number in accordance with the intended deployment of the present invention without losing their beneficial effects.
Claims
1. selecting a first object in a first volumetric video using a first attribute of the first object; selecting a second object in a second volumetric video using a second attribute of the second object, where the first attribute and the second attribute satisfy an aggregation rule; and generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video comprises simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule. A computer-implemented method comprising:
2. 2. The method of claim 1 , wherein selecting the first object in the first volumetric video comprises using an instance segmentation process, the instance segmentation process including classifying a first portion of the first volumetric video as representing the first object having the first attribute.
3. The instance segmentation process is extracting image segments from frames of the first volumetric video; classifying the extracted image segments using a trained machine learning based image classifier, wherein the image classifier outputs a segment classification in response to receiving the extracted image segments; determining that the segment classification output from the image classifier is associated with the first object having the first attribute; In response to determining that the segment classification is associated with the first object having the first attribute, designating the extracted image segment as a depiction of at least a portion of the first object such that the first portion of the first volumetric video includes the extracted image segment. The method of claim 2 further comprising:
4. extracting the first portion of the first volumetric video from a frame of the first volumetric video; and inserting the so extracted first portion of the first volumetric video into a frame of a third volumetric video. The method of claim 2 further comprising:
5. extracting a second portion of the second volumetric video from a frame of the second volumetric video, wherein the second portion represents the second object having the second attribute; and inserting said so extracted second portion of said second volumetric video into a frame of a third volumetric video. The method of claim 4 further comprising:
6. transmitting the aggregated volumetric video to a user device; detecting a user input indicating a selection of the first volumetric video; and transitioning from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device in response to the user input. The method of claim 1 further comprising:
7. transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; detecting a user input indicating a selection of the preview image; and transmitting the aggregated volumetric video to the user device in response to the user input. The method of claim 1 further comprising:
8. detecting a user input from a user indicating a selection of the first volumetric video; and determining whether the user is authorized to view the first volumetric video by comparing permission rules associated with the user with access rules associated with the first volumetric video.
10. The method of claim 1, further comprising: generating the aggregated volumetric video in response to determining that the user is authorized to view the first volumetric video.
9. comparing a first set of aggregation rules associated with the first user to a second set of aggregation rules associated with the second user; detecting that the first set of aggregation rules matches the second set of aggregation rules, where the first set of aggregation rules and the second set of aggregation rules both include the aggregation rule; and transmitting the aggregated volumetric video to the first user and the second user in response to detecting that the first set of aggregation rules matches the second set of aggregation rules. The method of claim 1 further comprising:
10. 10. The method of claim 9, further comprising designating the first user and the second user for joint aggregation processing in response to detecting that the first set of aggregation rules matches the second set of aggregation rules.
11. 11. The method of claim 10, further comprising, in response to designating the first user and the second user for joint aggregation, transmitting the aggregated volumetric video to a first user device associated with the first user and a second user device associated with the second user.
12. 1. A computer program product comprising one or more computer-readable storage media and program instructions collectively stored on the one or more computer-readable storage media, the program instructions causing a processor to: selecting a first object in a first volumetric video using a first attribute of the first object; selecting a second object in a second volumetric video using a second attribute of the second object, where the first attribute and the second attribute satisfy an aggregation rule; and generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video includes simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule; a computer program product executable by the processor to cause the processor to perform operations having the steps:
13. 13. The computer program product of claim 12, wherein the stored program instructions are stored in a computer-readable storage device in a data processing system, and the stored program instructions are transferred over a network from a remote data processing system.
14. the stored program instructions are stored in a computer readable storage device at a server data processing system, and the stored program instructions are downloaded for use in a computer readable storage device associated with a remote data processing system in response to a request over a network to the remote data processing system; program instructions for metering usage of the program instructions associated with the request; and program instructions for generating a bill based on said metered usage The computer program product of claim 12 further comprising:
15. The operation is transmitting the aggregated volumetric video to a user device; detecting a user input indicating a selection of the first volumetric video; and transitioning from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device in response to the user input.
13. The computer program product of claim 12, further comprising:
16. The operation is transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; detecting a user input indicating a selection of the preview image; and transmitting the aggregated volumetric video to the user device in response to the user input.
13. The computer program product of claim 12, further comprising:
17. The operation is detecting a user input from a user indicating a selection of the first volumetric video; and determining whether the user is authorized to view the first volumetric video by comparing permission rules associated with the user with access rules associated with the first volumetric video; 13. The computer program product of claim 12, further comprising: generating the aggregated volumetric video in response to determining that the user is authorized to view the first volumetric video.
18. 1. A computer system comprising a processor and one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media, the program instructions causing the processor to: selecting a first object in a first volumetric video using a first attribute of the first object; selecting a second object in a second volumetric video using a second attribute of the second object, where the first attribute and the second attribute satisfy an aggregation rule; and generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video includes simultaneously rendering the first object and the second object in the aggregated volumetric video based on the aggregation rule; 20. A computer system executable by the processor to perform operations comprising:
19. The operation is transmitting the aggregated volumetric video to a user device; detecting a user input indicating a selection of the first volumetric video; and transitioning from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device in response to the user input.
20. The computer system of claim 18, further comprising:
20. The operation is transmitting a multiview selection image to a user device, wherein the multiview selection image includes a preview image associated with the aggregated volumetric video; detecting a user input indicating a selection of the preview image; and transmitting the aggregated volumetric video to the user device in response to the user input.
20. The computer system of claim 18, further comprising: