Personalized aggregation of volumetric video
By introducing aggregation engine and aggregation rules into the volume video system, identifying and rendering objects of multiple source volume videos, the problem that users cannot select and control video content is solved, and the user's viewing freedom and immersion are improved.
Patent Information
- Application Number
- CN202380084783.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-11-24
- Publication Date
- 2025-07-18
Smart Images

Figure CN120345236A_ABST
Abstract
Description
Background Art
[0001] The present invention generally relates to volumetric video processing. More specifically, the present invention relates to methods, systems, and computer programs for personalized aggregated volumetric video.
[0002] Immersive technologies continue to gain popularity, especially as related consumer devices such as virtual reality (VR) headsets continue to drop in price. The growth in consumer adoption of this technology has led to growth in content, as content creators have responded to consumer interest in new immersive experiences. Additionally, advancements in video capture technology have allowed content creators to increase the "immersiveness" of their content.
[0003] Immersive video content is typically captured using multiple cameras from different angles simultaneously, or using one camera from multiple positions and angles. An example of immersive shooting is so-called 360° shooting. The resulting 360° video content is typically created by a computer by stitching together multiple images with a limited field of view but captured simultaneously to form a whole sphere of a still or video image where an individual can stand. It is most easily viewed by an individual with a VR headset, but can also be viewed by other means. It is also somewhat limited because the individual perspectives within the scene are fixed relative to the image itself. In other words, a viewer can only view such a scene from the positions chosen by the filmmaker, thus limiting movement within the scene. The viewer can look around in all directions, but they cannot move away from the position of the physical camera. Additionally, traditional 360° video content sacrifices depth and volumetric content because it is effectively a sphere with the observer at the center of the sphere and the pictures posted along the inner wall of the sphere. There are no objects within the scene that have a shape other than that spherical wall. This further reduces the level of immersion of the experience by restricting the viewer's experience.
[0004] Volumetric video differs from 360° video in that volumetric video also uses photogrammetry or depth sensors (e.g., light field arrays, LIDAR) to capture depth information. This information results in a volumetric video capturing an image of the scene and the overall three-dimensional parameters of the objects within the scene. Thus, for example, a chair within a given volumetric video scene can have both a shape (e.g., a three-dimensional geometry corresponding to the shape of a chair) and an image superimposed on it to create the impression that it is made of wood, or metal, or plastic, or any possible scenario. Thus, in volumetric video, a viewer can generally move freely within the scene, overcoming the movement limitations of traditional two-dimensional or 360° video shooting techniques. The content produced using these techniques is typically referred to as volumetric, six degrees of freedom (6DOF), light field, or free viewpoint video. Summary of the Invention
[0005] Exemplary embodiments provide for personalized aggregation of volumetric videos. An embodiment includes selecting a first object in a first volumetric video using a first attribute of a first object. The embodiment also includes selecting a second object in a second volumetric video using a second attribute of a second object, where the first attribute and the second attribute satisfy an aggregation rule. The embodiment also includes generating an aggregated volumetric video from the first volumetric video and the second volumetric video, where generating the aggregated video includes rendering the first object and the second object simultaneously in the aggregated volumetric video based on the aggregation rule. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the embodiments.
[0006] An embodiment includes a computer-usable program product. The computer-usable program product includes a computer-readable storage medium and program instructions stored on the storage medium.
[0007] An embodiment includes a computer system. The computer system includes a processor, a computer-readable memory, a computer-readable storage medium, and program instructions stored on the storage medium for execution by the processor via the memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Novel features that are considered characteristic of the invention are set forth in the appended claims. However, the invention itself, as well as its preferred mode of use, further objectives and advantages, will best be understood by reference to the following detailed description of illustrative embodiments in conjunction with the accompanying drawings, in which:
[0009] Figure 1 A block diagram of a computing environment according to an illustrative embodiment is described;
[0010] Figure 2 A functional block diagram of an exemplary volumetric video module according to an illustrative embodiment is described;
[0011] Figure 3 A functional block diagram of an exemplary video processing environment according to an illustrative embodiment is described;
[0012] Figure 4 An aggregation process according to an illustrative embodiment is described;
[0013] Figure 5 A functional block diagram of a volumetric video module according to an illustrative embodiment is described;
[0014] Figure 6 A functional block diagram of a volumetric video module according to an illustrative embodiment is described;
[0015] Figure 7 A flowchart of an example process for aggregating volumetric videos according to an illustrative embodiment is described;
[0016] Figure 8 A flowchart depicting an example process for aggregating volumetric video using aggregation optimization techniques according to an illustrative embodiment; and
[0017] Figure 9 A flowchart depicting an example process for aggregating volumetric video using security techniques according to an illustrative embodiment. DETAILED DESCRIPTION
[0018] The difference between volumetric video and 360° video is that volumetric video includes depth information. This added depth information significantly changes the way the content can be consumed. When viewing a scene in 360° video format, the viewer is locked to a single position, and from that single vantage point, the viewer can have up to three degrees of freedom corresponding to rotation about each of the three orthogonal axes of the Cartesian coordinate system: roll (tilting the head left or right), pitch (tilting the head forward or backward), and yaw (turning the head left or right).
[0019] In contrast, the depth information included in volumetric video liberates the viewer from the locked vantage point and allows the viewer to have up to six degrees of freedom, which correspond to rotation about each of the three orthogonal axes of the Cartesian coordinate system and translation along each of the three orthogonal axes of the Cartesian coordinate system: roll, pitch, and yaw as described above, plus lift (moving up or down), pan (moving left or right), and dolly (moving forward or backward). As a result, the viewer can freely move around in the volumetric video scene and observe objects from multiple angles and vantage points. Compared to earlier technologies, this increased freedom of movement significantly increases the immersive nature of the content provided in volumetric video format.
[0020] However, a limitation of current volumetric video is that the viewer is still restricted to only viewing the volumetric video scene as originally created. That is, the viewer cannot control the content of the volumetric video. For example, if a user is remotely participating in a meeting that is taking place simultaneously at two or more locations, where corresponding volumetric videos are made available to the remote participants from the two or more locations, the user is restricted to viewing only one of the two volumetric videos at a time. In other words, the user cannot select the option of having a volumetric video that includes elements of interest from each source volumetric video.
[0021] The disclosed embodiments address these and other limitations of traditional volumetric video systems by providing an aggregated volumetric video that is an aggregation of two or more source volumetric videos. The disclosed embodiments also provide for personalization of the aggregated volumetric video. For example, in some embodiments, the user specifies one or more aggregation rules.
[0022] Aggregation rules are primarily instructions and / or preferences that guide the presentation of two or more source volumetric videos as a single aggregated volumetric video. Aggregation rules can include the specification of objects from the source volumetric videos to be shown or not shown in the aggregated volumetric video (e.g., showing the audience from a first source but not the audience from a second source). Aggregation rules can include the specification of attributes of objects from the source volumetric videos to be shown or not shown or modified in the aggregated volumetric video (e.g., changing the hue of the audience seats from blue to black; blurring the company logo on the shirts and hats of audience members).
[0023] Aggregation rules can include rules of different levels of specificity. For example, aggregation rules can include the specification of specific attributes, specific objects, or object classes, where the object classes can vary in terms of specificity. As a more specific non-limiting example, a user creating rules for a personalized aggregated volumetric video can have source volumetric videos from multiple car shows, and in that scenario, rules of varying specificity can include rules for broad classes of objects, such as a preference for watching sedans compared to other types of vehicles; rules for more specific object classes can specify the particular make and model of the sedan; and rules for a specific object can specify a famous car shown in one of the volumetric videos.
[0024] In some embodiments, video aggregation is performed by a volumetric video module. In some such embodiments, the volumetric video module analyzes digital volumetric video data from two or more source volumetric videos. In some such embodiments, this analysis results in the identification of one or more objects in each source volumetric video.
[0025] In some embodiments, the analysis generates metadata for one or more of the identified objects. In some embodiments, the metadata for the identified objects includes data representing one or more attributes of the identified objects. In various embodiments, the number and type of attributes will vary according to implementation decisions, constraints, and / or preferences. Non-limiting examples of attribute types include appearance, classification, and / or source attributes. In an exemplary embodiment, the appearance attributes of the identified object can include the size, shape, and / or hue attributes of the identified object.
[0026] In an exemplary embodiment, the classification attributes of the identified object can include one or more categories or types associated with the identified object. In some embodiments, the classification attributes include the class (or classes) predicted using known object classification and / or detection techniques.
[0027] For example, in some embodiments, object classification techniques include a machine learning process that uses a trained machine learning model to predict the class of an object in an image. It should be understood that the "image" referred to herein includes an image as a video frame. In some embodiments, classification is performed using known semantic image segmentation techniques that classify portions or segments of a volumetric image using the corresponding classes represented by the portions or segments of the volumetric image.
[0028] In some embodiments, object detection includes techniques that combine classification with localization techniques to determine the location of the classified objects in an image. Thus, in some embodiments, object detection is performed using known techniques that determine what objects are in the image and specify the location of the objects in the image. In some embodiments, object detection includes known instance segmentation techniques that distinguish between individual objects of the same class in an image.
[0029] In some embodiments, classification attributes may include multiple classes representing degrees of particularity of variation. For example, assume that the identified object is a multi-passenger van. In this example, the identified object may include "vehicle" as a first classification attribute, "multi-passenger vehicle" as a second classification attribute, and "van" as a third classification attribute.
[0030] In an exemplary embodiment, source attributes may include information identifying one or more sources of a volumetric video. As a non-limiting example, assume that an aggregated volumetric video is being generated for a meeting occurring simultaneously at geographically distant first and second locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video captured at the first location and a second volumetric video captured at the second location. Thus, a first identified object captured in the first volumetric video may include a source attribute representing the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video may include a source attribute representing the second location or the second volumetric video as the source of the second identified object.
[0031] The set of size and shape attributes can include only a single attribute, as well as two or more attributes. Size and shape attributes are characteristics, features, or other properties of an object that are associated with the size measurements and / or shape of one or more parts of the object. A part of an object can be a portion, section, or component of the object. A part of an object can also include the entire object. Size measurement is the measurement of any dimension, such as but not limited to height, length, and width. Size measurement can also include the weight and volume of an object. The set of size and shape attributes can include only size-related attributes and not include any shape-related attributes. In another embodiment, the set of size and shape attributes can include only shape-related attributes and not include any size-related attributes. In yet another embodiment, the set of size and shape attributes includes size-related attributes and shape-related attributes.
[0032] Embodiments can be implemented as software applications. The application programs implementing the embodiments can be configured as modifications in existing manufacturing systems, as separate application programs operating in conjunction with existing manufacturing systems, as stand-alone systems, or some combination thereof.
[0033] Embodiments monitor system state data to obtain indications of system failures that cause the system kernel to enter a suspended state. The state data can vary depending on the type of system (e.g., operating system and hardware), but generally can include system error messages, error codes, log entries, or other data indicating system errors. System errors can also vary depending on the type of system, but generally can include kernel errors, stop errors, etc. that cause a partial or complete loss of kernel functionality.
[0034] System failures that cause a partial or complete loss of kernel functionality are generally recognized by the system as conditions that require a restart of the system in order to recover. The loss of kernel functionality generally requires a restart to recover in most types of systems. Thus, such errors are examples of errors that meet the restart condition.
[0035] In an illustrative embodiment, when a system error that meets the restart condition is detected (e.g., a system error that causes a loss of kernel functionality), debug data is temporarily saved to a protected section of memory. Since at least a portion of the kernel functionality will be lost at this time, the embodiments include generating and saving a copy of the debug data without the assistance of the kernel. For example, in some embodiments, an exception handler generates and saves a copy of the debug data. The debug data can vary depending on the type of system (e.g., operating system and hardware), but generally can include data from processor memory (e.g., registers and caches), logs, and / or trace arrays.
[0036] Since the kernel is in a paused state when generating debug data, this means that kernel functions for processing debug data are unavailable. As a result, the kernel is not available at this time to filter sensitive information from the debug data. For this reason, the debug data is stored in a protected section of the memory, where it can be retained during a restart and processed thereafter.
[0037] After a restart occurs, the debug device, as a non-trusted device to the restored system, is connected via an I / O port using a trusted protocol. The non-trusted device issues a request for the debug data that will be used to attempt to determine the cause of a system failure. Since the debug data is stored in protected memory, the untrusted device cannot directly access the debug data. Instead, the untrusted entity must request the debug data from the secure debug module.
[0038] In the illustrated embodiment, the secure debug module receives a request for the debug data. In response to the request, the secure debug module uses a data sanitization module to analyze the debug data using a sensitive data detection / sanitization process that detects and removes sensitive data from the debug data.
[0039] In some embodiments, the data sanitization module detects sensitive data according to an audit policy. An audit policy is a set of preferences, rules, and / or criteria for protecting sensitive data in the debug data. For example, the audit policy may define "sensitive objects" as files or objects that contain specific keywords (e.g., "confidential" or "privileged") and / or are associated with specific keywords (e.g., in metadata) or specific flags (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy can also specify rules for handling sensitive objects. As an example, the audit policy may require an auditor to approve any potential transfer of sensitive objects from protected memory to an untrusted entity. Thus, in some embodiments, the data protection process includes one or more protection measures in the form of data sanitization to prevent sensitive data from being leaked along with the debug data sent to an untrusted entity.
[0040] In the illustrated embodiment, the window module detects whether sensitive data is being processed during a time window in which a system error occurs, for example, by looking at system logs. For example, in some embodiments, the window module uses an audit policy to review the system logs, and the audit policy includes a set of preferences, rules, and / or criteria that the window module uses to identify sensitive data in the debug data. For example, the audit policy may define a "sensitive object" as a file or object that contains a specific keyword (e.g., "confidential" or "privileged") and / or is associated with a specific keyword (e.g., in metadata) or a specific flag (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy can also specify rules for handling sensitive objects. As an example, the audit policy may require an auditor to approve the transfer of any potentially sensitive object from a protected memory to an untrusted entity. Thus, in some embodiments, the data protection process includes one or more protection measures in the form of time window analysis to prevent sensitive data from being leaked together with the debug data sent to an untrusted entity.
[0041] In some embodiments, the data sanitization module detects sensitive data according to an audit policy. The audit policy is a set of preferences, rules, and / or criteria for protecting sensitive data in the debug data. For example, the audit policy may define a "sensitive object" as a file or object that contains a specific keyword (e.g., "confidential" or "privileged") and / or is associated with a specific keyword (e.g., in metadata) or a specific flag (e.g., in metadata that identifies a document or email as personal, confidential, etc.). The audit policy can also specify rules for handling sensitive objects. As an example, the audit policy may require an auditor to approve the transfer of any potentially sensitive object from a protected memory to an untrusted entity. Thus, in some embodiments, the data protection process includes one or more protection measures in the form of data encryption to prevent sensitive data from being leaked together with the debug data sent to an untrusted entity.
[0042] For the sake of clarity of description and without implying any limitation thereto, some example configurations are used to describe the illustrative embodiments. According to the present disclosure, those of ordinary skill in the art will be able to conceive of many variations, adaptations, and modifications of the configurations for achieving the stated purposes, and these are all considered to be within the scope of the exemplary embodiments.
[0043] In addition, simplified diagrams of a data processing environment are used in the drawings and the illustrative embodiments. In an actual computing environment, there may be additional structures or components that are not shown or described herein, or structures or components that are different from those shown but are used for functions similar to those described herein, without departing from the scope of the illustrative embodiments.
[0044] In addition, the description of illustrative embodiments with respect to specific actual or hypothetical components is for example only. Any specific representation of these and other similar artifacts is not intended to limit the invention. Any suitable representation of these and other similar artifacts may be selected within the scope of the exemplary embodiments.
[0045] The examples in this disclosure are for clarity of description only and are not limiting to the illustrative embodiments. Any advantages listed herein are only examples and are not intended to limit the illustrative embodiments. Additional or different advantages may be achieved through specific illustrative embodiments. Further, a particular illustrative embodiment may have some, all, or none of the advantages listed above.
[0046] In addition, the illustrative embodiments may be implemented with respect to any type of data, data source, or access to a data source over a data network. Within the scope of the present invention, any type of data storage device may provide data locally at a data processing system or to an embodiment of the present invention over a data network. In the case where embodiments are described using a mobile device, within the scope of the illustrative embodiments, any type of data storage device suitable for use with a mobile device may provide data locally at the mobile device or to such an embodiment over a data network.
[0047] The use of specific codes, computer-readable storage media, advanced features, designs, architectures, protocols, layouts, diagrams, and tools to describe the illustrative embodiments is for example only and is not a limitation of the illustrative embodiments. Further, for clarity of description, in some instances specific software, tools, and data processing environments are used only as examples to describe the illustrative embodiments. The illustrative embodiments may be used in conjunction with other equivalent or similar purpose structures, systems, applications, or architectures. For example, within the scope of the present invention, other comparable mobile devices, structures, systems, applications, or their architectures may be used in conjunction with such embodiments of the present invention. The illustrative embodiments may be implemented in hardware, software, or a combination thereof.
[0048] The examples in this disclosure are for clarity of description only and are not limiting to the illustrative embodiments. Additional data, operations, actions, tasks, activities, and manipulations may be contemplated from this disclosure, and these additional data, operations, actions, tasks, activities, and manipulations may be envisioned within the scope of the illustrative embodiments.
[0049] Aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from the order shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0050] The term "computer program product embodiment" ("CPP embodiment" or "CPP") as used in the present disclosure describes any collection of one or more storage media (also referred to as "media") jointly included in a set of one or more storage devices, the set of one or more storage devices jointly including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions used by a computer processor. By way of non-limitation, computer-readable storage media can be electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punched cards or pits / lands formed in the major surface of a disk, or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in the present disclosure, should not be construed to store in the form of a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals transmitted through wires and / or other transmission media. As will be understood by those skilled in the art, data is typically moved at certain incidental points in time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not render the storage device transitory because the data is not transitory when it is stored.
[0051] Reference Figure 1, which depicts a block diagram of a computing environment 100. The computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the methods of the present invention, such as an improved volumetric video module 200 that provides personalized aggregation of volumetric video. In addition to the volumetric video module 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a set of processors 110 (including processing circuitry 120 and a cache 121), a communication fabric 111, volatile memory 112, persistent storage 113 (including an operating system 122 and the volumetric video module 200, as described above), a set of peripheral devices 114 (including a set of user interface (UI) devices 123, a storage device 124, and a set of Internet of Things (IoT) sensors 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0052] The computer 101 can take the form of a desktop computer, a laptop computer, a tablet computer, a smart phone, a smart watch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of running programs, accessing a network, or querying a database such as the remote database 130. As is well known in the computer art and depending on the technology, the performance of computer-implemented methods can be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of the computing environment 100, the detailed discussion focuses on a single computer, specifically the computer 101, to keep the presentation as simple as possible. The computer 101 can be located in the cloud, even though Figure 1 it is not shown in the cloud. On the other hand, the computer 101 does not need to be in the cloud unless it can be positively indicated to any extent.
[0053] The set of processors 110 includes one or more computer processors of any type now known or later developed. The processing circuitry 120 may be distributed across multiple packages, such as multiple cooperative integrated circuit chips. The processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. The cache 121 is a memory located within the processor chip package and is generally used for data or code that should be made available for rapid access by threads or cores running on the set of processors 110. Caches are generally organized into multiple levels based on their relative proximity to the processing circuitry. Alternatively, some or all of the caches of the set of processors may be located "off-chip". In some computing environments, the set of processors 110 may be designed to work with qubits and perform quantum computing.
[0054] Computer-readable program instructions are typically loaded onto the computer 101 so that the set of processors 110 of the computer 101 executes a series of operational steps to implement a computer-implemented method such that the instructions so executed will instantiate the method specified in the flowchart and / or the narrative description of the computer-implemented method included in this document (collectively referred to as "the method of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 110 to control and direct the execution of the method of the present invention. In the computing environment 100, at least some of the instructions for performing the method of the present invention may be stored in the volumetric video module 200 in the persistent storage device 113.
[0055] The communication structure 111 is a signal conduction path that allows the various components of the computer 101 to communicate with each other. Generally, this structure consists of switches and conductive paths, such as switches and conductive paths that make up a bus, a bridge, a physical input / output port, etc. Other types of signal communication paths may be used, such as fiber optic communication paths and / or wireless communication paths.
[0056] The volatile memory 112 is any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Generally, the volatile memory 112 is characterized by random access, but this is not required unless specifically stated. In the computer 101, the volatile memory 112 is located within a single package and inside the computer 101, but, alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located externally relative to the computer 101.
[0057] The persistent storage device 113 is any form of non-volatile storage for a computer that is known now or developed in the future. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 can be a read-only memory (ROM), but typically at least a portion of the permanent memory allows for the writing of data, the deletion of data, and the re-writing of data. Some common forms of persistent storage devices include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems or operating systems of the open-source portable operating system interface type that employ a kernel. The code included in the volumetric video module 200 typically includes at least some of the computer code involved in performing the methods of the present invention.
[0058] The peripheral device collection 114 includes the peripheral device collection of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as via a Bluetooth connection, a near-field communication (NFC) connection, a connection made by a cable (such as a universal serial bus (USB)-type cable), a plug-in connection (e.g., a Secure Digital (SD) card), a connection made via a local communication network, and even a connection made via a wide-area network such as the Internet. In various embodiments, the UI device collection 123 can include components such as a display screen, speakers, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. The storage device 124 can be permanent and / or volatile. In some embodiments, the storage device 124 can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 needs to have a large amount of storage (e.g., in the case where the computer 101 locally stores and manages a large database), the storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor collection 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0059] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 can include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for performing the methods of the present invention can generally be downloaded to the computer 101 from an external computer or an external storage device via a network adapter or network interface included in the network module 115.
[0060] The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances via any technology known now or developed in the future for transmitting computer data. In some embodiments, the WAN 102 can be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN generally includes computer hardware such as copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0061] The end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating the computer 101) and can take any form discussed above in connection with the computer 101. The EUD 103 typically receives useful and available data from the operation of the computer 101. For example, in the hypothetical case where the computer 101 is designed to provide recommendations to the end user, the recommendations will typically be transmitted from the network module 115 of the computer 101 to the EUD 103 via the WAN 102. In this way, the EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.
[0062] The remote server 104 is any computer system that provides at least some data and / or functionality to the computer 101. The remote server 104 can be controlled and used by the same entity that operates the computer 101. The remote server 104 represents a machine that collects and stores useful and available data used by other computers such as the computer 101. For example, in the hypothetical case where the computer 101 is designed and programmed to provide recommendations based on historical data, then that historical data can be provided to the computer 101 from the remote database 130 of the remote server 104.
[0063] The public cloud 105 is any computer system that can be used by multiple entities, which provides on-demand availability of computer system resources and / or other computer capabilities (notably data storage (cloud storage) and computing power) without direct active management by the user. Cloud computing typically exploits the sharing of resources to achieve consistency and economy of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers of a set of host physical machines 142, which is the universe of physical computers within and / or available for the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from a virtual machine group 143 and / or containers from a container group 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after instantiation of the VCE. The cloud orchestration module 141 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiations of the VCE deployment. The gateway 140 is a collection of computer software, hardware, and firmware that allows the public cloud 105 to communicate via the WAN 102.
[0064] Some further explanations of the virtualized computing environment (VCE) will now be provided. The VCE can be stored as an "image". New active instances of the VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for the existence of multiple isolated user space instances, called containers. From the perspective of the programs running within them, these isolated user space instances typically behave as actual computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices allocated to the container, which is a feature known as containerization.
[0065] A private cloud 106 is similar to a public cloud 105, except that computing resources are only available to a single enterprise. Although the private cloud 106 is depicted as communicating with the WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) that are typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable coordination, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0066] Measurable services: The cloud system automatically controls and optimizes resource usage by leveraging metering capabilities at a certain level of abstraction appropriate for the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, reported, and invoiced, providing transparency to both the provider and consumer of the services used.
[0067] Reference Figure 2 , which depicts a functional block diagram of an exemplary volumetric video module 200 according to an illustrative embodiment. In the illustrated embodiment, the volumetric video module 200 receives multiple volumetric videos 206A–206C and aggregates two or more of them into an aggregated volumetric video 208.
[0068] In the illustrated embodiment, the volumetric video module 200 includes an aggregation engine 202 and aggregation rules 204 stored in a computer-readable storage medium. In alternative embodiments, the volumetric video module 200 may include some or all of the functions described herein, but grouped differently into one or more modules. In some embodiments, the functions described herein are distributed across multiple systems, which may include a combination of software- and / or hardware-based systems, such as application-specific integrated circuits (ASICs), computer programs, or smartphone applications.
[0069] In the illustrated embodiment, the aggregation engine 202 aggregates individual volumetric videos, such as two or more of volumetric videos 206A - 206C, and creates a single aggregated volumetric video 208, also referred to herein as a corridor 208. In some embodiments, the volumetric videos 206A - 206C are captured simultaneously from different physical locations. For example, volumetric video 206A can be a live broadcast in volumetric video format from Los Angeles, volumetric video 206B can be a live broadcast in volumetric video format from Chicago, and volumetric video 206C can be a live broadcast in volumetric video format from New York, for a meeting hosted at these three locations. In some embodiments, the volumetric videos 206A - 206C are captured from different physical locations at different times.
[0070] Aggregation rules 204 are primarily instructions and / or preferences that guide the presentation of two or more source volumetric videos as a single aggregated volumetric video. Aggregation rules 204 can include the specification of objects from the source volumetric videos to be displayed or not displayed in the aggregated volumetric video (e.g., display the audience from the first source rather than the audience from the second source). Aggregation rules 204 can include the specification of attributes of objects from the source volumetric videos to be displayed or not displayed or modified in the aggregated volumetric video (e.g., change the hue of the audience chairs from blue to black; blur the company logos on the shirts and hats of audience members).
[0071] Aggregation rules 204 can include rules of different levels of specificity. For example, aggregation rules 204 can include the specification of specific attributes, specific objects, or object classes, where the object classes can vary in terms of specificity. As a more specific non - limiting example, a user creating rules 204 for a personalized aggregated volumetric video can have source volumetric videos from multiple car shows, and in this scenario, rules 204 of different specificities can include rules for broad object classes, such as a preference for watching sedans compared to other types of vehicles; rules for more specific object classes can specify the particular make and model of the sedan; and rules for a specific object can specify a famous car shown in one of the volumetric videos.
[0072] In some embodiments, the aggregation engine 202 performs video aggregation. In some such embodiments, the aggregation engine 202 analyzes digital volumetric video data from two or more source volumetric videos. In some such embodiments, this analysis results in the identification of one or more objects in each source volumetric video.
[0073] In some embodiments, the aggregation engine 202 generates metadata for one or more of the identified objects. In some embodiments, the metadata for the identified objects includes data representing one or more attributes of the identified objects. In various embodiments, the number and type of attributes will vary according to implementation decisions, constraints, and / or preferences. Non-limiting examples of attribute types include appearance, classification, and / or source attributes. In an exemplary embodiment, the appearance attributes of the identified object may include the size, shape, and / or hue attributes of the identified object.
[0074] In an exemplary embodiment, the classification attributes of the identified object may include one or more categories or classes associated with the identified object. In some embodiments, the classification attributes include the class (or classes) predicted using known object classification and / or detection techniques.
[0075] For example, in some embodiments, the aggregation engine 202 includes an image classifier 210 that performs object classification techniques including a machine learning process that uses a trained machine learning model to predict the class of an object in an image. It should be understood that the "image" referred to here includes an image as a video frame. In some embodiments, classification is performed using known semantic image segmentation techniques that classify parts or segments of a volumetric image using the corresponding classes represented by the parts or segments of the volumetric image.
[0076] In some embodiments, the aggregation engine 202 performs object detection using techniques that combine classification with localization techniques to determine the location of the classified objects in the image. Thus, in some embodiments, the aggregation engine 202 performs object detection using known techniques that determine what objects are in the image and specify where the objects are located in the image. In some embodiments, the aggregation engine 202 uses object detection techniques including known instance segmentation techniques that distinguish between individual objects of the same class in an image.
[0077] In some embodiments, the classification attributes may include multiple classes representing the degree of particularity variation. For example, assuming the identified object is a multi-passenger van, in this example, the identified object may include "vehicle" as the first classification attribute, "multi-passenger vehicle" as the second classification attribute, and "van" as the third classification attribute.
[0078] In an exemplary embodiment, a source attribute may include information identifying one or more sources of a volumetric video. As a non-limiting example, assume that an aggregated volumetric video is being generated for a meeting occurring simultaneously at geographically distant first and second locations. In this example, the aggregated volumetric video is an aggregation of a first volumetric video captured at the first location and a second volumetric video captured at the second location. Thus, a first identified object captured in the first volumetric video may include a source attribute indicating the first location or the first volumetric video as the source of the first identified object, while a second identified object captured in the second volumetric video may include a source attribute indicating the second location or the second volumetric video as the source of the second identified object.
[0079] The set of size and shape attributes may include only a single attribute, as well as two or more attributes. Size and shape attributes are characteristics, features, or other properties of an object that are associated with the dimensional measurements and / or shape of one or more parts of the object. A part of an object may be a portion, section, or component of the object. A part of an object may also include the entire object. Dimensional measurement is the measurement of any dimension, such as but not limited to height, length, and width. Dimensional measurement may also include the weight and volume of an object. The set of size and shape attributes may include only size-related attributes and not include any shape-related attributes. In another embodiment, the set of size and shape attributes may include only shape-related attributes and not include any size-related attributes. In yet another embodiment, the set of size and shape attributes includes both size-related attributes and shape-related attributes.
[0080] Reference Figure 3 , the figure depicts a functional block diagram of an exemplary video processing environment 300 in accordance with an illustrative embodiment. In the illustrated embodiment, video processing environment 300 includes servers 306A, 306B, and 310, each of which may include Figure 1 and 2 a volumetric video module 200.
[0081] In the illustrated embodiment, volumetric videos 302A - 302D are examples of source volumetric videos. In the illustrated embodiment, each volumetric video 302A - 302D is created using multiple cameras 304. Alternatively, one or more volumetric videos 302A - 302D may be created using a single camera recording from multiple locations and angles.
[0082] Figure 3 The embodiments in
[0083] In the illustrated embodiment, server 306A aggregates volumetric video 302A and volumetric video 302B into aggregated volumetric video 308A, and server 306B aggregates volumetric video 302C and volumetric video 302D into aggregated volumetric video 308B. Then, server 310 aggregates aggregated volumetric video 308A and aggregated volumetric video 308B into aggregated volumetric video 312. In an alternative embodiment, source two-dimensional (2D) video and depth information from camera 304 may be used instead of volumetric video to aggregate the aggregated volumetric video. For example, in some embodiments, server 306A uses 2D video and depth information from camera 304 associated with volumetric video 302A and from camera 304 associated with volumetric video 302B to generate aggregated volumetric video 308A.
[0084] In another embodiment, the aggregated volumetric video may be further aggregated with another source volumetric video. For example, if nothing from volumetric video 302D is needed for aggregated volumetric video 312, server 310 aggregates aggregated volumetric video 308A with volumetric video 302C into aggregated volumetric video 312.
[0085] Reference Figure 4 , which depicts an aggregation process 400 according to an illustrative embodiment. In one embodiment, aggregation process 400 is performed by the Figure 1 and 2 volumetric video module 200.
[0086] Figure 4 The embodiments in
[0087] show that, in addition to aggregating two volumetric videos (source or aggregated) at a time, alternative embodiments may aggregate three or more volumetric videos (source or aggregated) at a time. In the illustrated embodiment, the user rule specifies that volumetric video 402C is primarily used to aggregate volumetric video 402A and volumetric video 402B with volumetric video 402C, but the viewers in image portion 404 from volumetric video 402A and the viewers in image portion 408 from volumetric video 402B are shown in place of the viewers in insertion positions 406 and 410 of volumetric video 402C.
[0087] Reference Figure 5 , which depicts a functional block diagram of a volumetric video module 502 according to an illustrative embodiment. Volumetric video module 502 is similar to the Figure 1 and 2 volumetric video module 200, except that the aggregation engine 504 and rendering engine 510 of volumetric video module 502 generate a hierarchical volumetric video and / or an aggregated volumetric video based on user navigation input.
[0088] For example, in the illustrated embodiment, the hierarchy of volumetric videos includes volumetric video 516A, volumetric video 516B, and aggregated volumetric video 516C. User 514 uses the head-mounted headset 512 to view volumetric video 516A at time T1. The head-mounted headset 512 includes or communicates with a volumetric video module 502.
[0089] The aggregation engine 504 receives user input and one or more source volumetric videos 206A - 206C. The aggregation engine 504 may also access aggregation rules 506 and an aggregation hierarchy 508. Figure 2 The description of the (multiple) aggregation rules 204 applies equally to the (multiple) aggregation rules 506. The aggregation hierarchy 508 can be a playlist that specifies which volumetric videos are to be presented and the order in which they should be presented. In some embodiments, user 514 creates the aggregation hierarchy 508. In some embodiments, the volumetric video module 502 automatically generates the aggregation hierarchy 508 according to user-specified rules or preferences (e.g., play the source volumetric videos in order from the most recent recording location to the furthest, and then play the aggregated one).
[0090] In the illustrated embodiment, the head-mounted headset 512 may include controls that allow for simple next / previous navigation of the aggregation hierarchy 508. These controls send user input to the aggregation engine 504. The aggregation engine 504 examines the aggregation hierarchy 508 and then instructs the rendering engine 510 to send the appropriate volumetric video. Thus, in the illustrated embodiment, the aggregation hierarchy 508 includes volumetric video 516A, followed by volumetric video 516B, followed by aggregated volumetric video 516C, such that user 514 who selects "next" at time T1 will next see volumetric video 516B at time T2, and if the user selects "next" again, the user will see aggregated volumetric video 516C at time T3. The user can also navigate back, e.g., from aggregated volumetric video 516C to volumetric video 516B, and so on.
[0091] Reference Figure 6 to, which depicts a functional block diagram of a volumetric video module 602 according to an illustrative embodiment. The volumetric video module 602 is similar to Figure 1 and 2 's volumetric video module 200, except that the volumetric video module 602 includes a multi-view module 610 that uses preview images 609 to generate multi-view images 618A.
[0092] The multi-view image 618A includes a display of two or more available volumetric videos that a user can use as a menu to select a volumetric video to view. For example, in the illustrated embodiment, the volumetric video module 602 receives source volumetric videos 206A - 206C and generates a fourth volumetric video that is an aggregation of the three source volumetric videos 206A - 206C. Thus, the volumetric video module 602 has four possible volumetric videos from which user 614 can choose. The user viewing through the head-mounted headset 616 is presented with the multi-view image 618A, which is created by the multi-view module 610 using preview images 609 from each of the four available volumetric videos. The multi-view image 618A allows user 614 to select one of the available volumetric videos by selecting the corresponding preview image in the multi-view image 618A. For example, if the user selects the upper left box, that selection is sent as user input to the aggregation engine 604. The aggregation engine 604 then provides the selected volumetric video data to the rendering engine 612. The rendering engine 612 then presents the volumetric video (in this example, volumetric video 618B is selected) at the user's 614 head-mounted headset 616.
[0093] Reference Figure 7 , which depicts a flowchart of an example process 700 for aggregating volumetric videos according to an illustrative embodiment. In a particular embodiment, Figure 2 the volumetric video module 200 of Figure 5 the volumetric video module 502 of Figure 6 the volumetric video module 602 of
[0094] In the illustrated embodiment, next, at block 702, the process identifies volumetric video sources associated with aggregation rules. Next, at block 704, the process determines whether the aggregation rules apply to any particular objects in the videos. If so, at block 706, the process segments the video data of the identified volumetric video sources. Next, at block 708, the process extracts the objects specified by the aggregation rules from the segmented video data such that the objects can be included in the aggregated volumetric video generated at block 714. However, first, the process checks for other objects specified by the rules (block 710) and additional volumetric video sources to be processed (block 712). Once all volumetric video sources have been processed for the objects specified by the aggregation rules, the process continues to block 714. At block 714, the process generates an aggregated volumetric video according to the aggregation rule(s), i.e., to include the specified objects extracted from the individual volumetric video sources.
[0095] Reference Figure 8 , which depicts a flowchart of an example process 800 for aggregating volumetric videos using aggregation optimization techniques according to an illustrative embodiment. In a particular embodiment,Figure 2 The volumetric video module 200, Figure 5 the volumetric video module 502, or Figure 6 the volumetric video module 602 executes process 800.
[0096] In the illustrated embodiment, at block 802, the process retrieves the aggregation rules for the user requesting the aggregated volumetric video. Next, at block 804, the process compares the aggregation rules of the requesting user with the set of aggregation rules of other users. Specifically, the process searches for matching aggregation rules of another user for which an aggregated video has been processed or generated. Next, at block 806, if a matching rule is found, the process performs an optimization technique where the process provides the requesting user with the aggregated video corresponding to the matching rule. This prevents redundant aggregation processing, thereby reducing the workload of the involved system. If no matching rule is found at block 806, the process generates an aggregated volumetric video according to the aggregation rules of the requesting user, e.g., according to Figure 7 process 700.
[0097] Referring to Figure 9 , the figure depicts a flowchart of an example process 900 for aggregating volumetric video using security techniques according to an illustrative embodiment. In a particular embodiment, Figure 2 the volumetric video module 200, Figure 5 the volumetric video module 502, or Figure 6 the volumetric video module 602 executes process 900.
[0098] In the illustrated embodiment, at block 902, the process identifies the volumetric video sources associated with the aggregation video request from the requesting user. Next, at block 904, the process compares the permission rules associated with the requesting user with the access rules for each of the identified video sources. At block 906, the process determines whether the requesting user is authorized to view the identified video sources. If the user is not authorized to view all of the identified video sources, the request is rejected at block 910. Otherwise, at block 908, the process generates an aggregated volumetric video according to the aggregation rules of the requesting user, e.g., according to Figure 7 process 700.
[0099] The following definitions and abbreviations are used to interpret the claims and the specification. As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” “contains” or “containing,” or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or apparatus.
[0100] Additionally, the term “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or design described herein as “exemplary” is not necessarily to be construed as more preferred or advantageous than other embodiments or designs. The terms “at least one” and “one or more” are to be understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term “plurality” is to be understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term “connected” may include indirect “connection” and direct “connection.”
[0101] References in the specification to “one embodiment,” “an embodiment,” “example embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but each embodiment may or may not include that particular feature, structure, or characteristic. Moreover, these phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is within the knowledge of one of ordinary skill in the art to effect such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described.
[0102] The terms “about,” “substantially,” “approximately,” and variations thereof are intended to include the degree of error associated with a measurement of a particular quantity based on the equipment available at the time of filing the present application. For example, “about” may include a range of ±8% or 5% or 2% of a given value.
[0103] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application, or a technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments described herein.
[0104] Descriptions of various embodiments of the present invention have been given for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments described herein.
[0105] Accordingly, in an illustrative embodiment, a computer-implemented method, system, or apparatus, and a computer program product are provided for managing participation in an online community and other related features, functions, or operations. In the case where an embodiment is described with respect to a type of device, the computer-implemented method, system, or apparatus, computer program product, or a portion thereof is adapted or configured to be used with the appropriate and comparable performance of that type of device.
[0106] In cases where an embodiment is described as being implemented in an application, within the scope of the illustrative embodiments, delivery of the application in a software as a service (SaaS) model can be envisioned. In the SaaS model, the ability to implement the application of the embodiment is provided to users by executing the application in a cloud infrastructure. Users can access the application using various client devices through a thin client interface such as a web browser (e.g., web-based email) or other lightweight client applications. The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage of the cloud infrastructure. In some cases, the user may not even manage or control the capabilities of the SaaS application. In some other cases, the SaaS implementation of the application program may allow for possible exceptions to limited user-specific application program configuration settings.
[0107] The present invention can be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0108] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0109] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider through the Internet). In some embodiments, in order to perform aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0110] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0111] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture that includes instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0112] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0113] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0114] Embodiments of the present invention may also be delivered as part of a service engagement with a customer company, a non-profit organization, a government entity, an internal organizational structure, and the like. Aspects of these embodiments may include configuring a computer system to execute, and deploying software, hardware, and web services that implement some or all of the methods described herein. Aspects of these embodiments may also include analyzing a customer's operations, creating recommendations in response to the analysis, building a system that implements portions of the recommendations, integrating the system into existing processes and infrastructure, metering the use of the system, allocating costs to users of the system, and billing for use of the system. Although the above embodiments of the present invention have been described by stating their respective advantages, the present invention is not limited to their specific combinations. Instead, these embodiments may also be combined in any manner and in any number according to the intended deployment of the present invention without losing their beneficial effects.
Claims
1. A computer-implemented method, comprising: selecting the first object in the first volumetric video using a first attribute of the first object; selecting the second object in the second volumetric video using a second attribute of the second object, wherein the first attribute and the second attribute satisfy an aggregation rule; and generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video includes presenting the first object and the second object simultaneously in the aggregated volumetric video based on the aggregation rule.
2. The method according to claim 1, wherein, Selecting the first object in the first volumetric video includes using instance segmentation processing, wherein the instance segmentation processing includes classifying a first portion of the first volumetric video as representing the first object having the first attribute.
3. The method according to claim 2, wherein, The instance segmentation processing further includes: extracting image patches from frames of the first volumetric video; classifying the extracted image patches using a trained machine learning-based image classifier, wherein the image classifier outputs a patch classification in response to receiving the extracted image patches; determining that the patch classification output from the image classifier is associated with the first object having the first attribute; responsive to determining that the patch classification is associated with the first object having the first attribute, designating the extracted image patches as depictions of at least some of the first object such that the first portion of the first volumetric video includes the extracted image patches.
4. The method according to claim 2, further comprising: extracting the first portion of the first volumetric video from frames of the first volumetric video; and inserting the thus-extracted first portion of the first volumetric video into frames of a third volumetric video.
5. The method according to claim 4, further comprising: extracting a second portion of the second volumetric video from frames of the second volumetric video, wherein the second portion represents the second object having the second attribute; and inserting the thus-extracted second portion of the second volumetric video into frames of a third volumetric video.
6. The method according to claim 1, further comprising: transmitting the aggregated volumetric video to a user device; detecting a user input indicating a selection of the first volumetric video; and responsive to the user input, converting from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device.
7. The method according to claim 1, further comprising: transmitting a multi-view selection image to a user device, wherein the multi-view selection image includes a preview image associated with the aggregated volumetric video; detecting a user input indicating a selection of the preview image; and responsive to the user input, transmitting the aggregated volumetric video to the user device.
8. The method according to claim 1, further comprising: detecting a user input from a user indicating a selection of the first volumetric video; and Determining whether the user is authorized to view the first volumetric video by comparing the license rules associated with the user and the access rules associated with the first volumetric video, wherein generating the aggregated volumetric video is in response to determining that the user is authorized to view the first volumetric video.
9. The method according to claim 1, further comprising: Comparing a first set of aggregation rules associated with a first user and a second set of aggregation rules associated with a second user; Detecting that the first set of aggregation rules matches the second set of aggregation rules, wherein both the first set of aggregation rules and the second set of aggregation rules include the aggregation rules; and In response to detecting that the first set of aggregation rules matches the second set of aggregation rules, transmitting the aggregated volumetric video to the first user and the second user.
10. The method according to claim 9, further comprising: In response to detecting that the first set of aggregation rules matches the second set of aggregation rules, designating the first user and the second user for co-aggregation processing.
11. The method according to claim 10, further comprising: In response to designating the first user and the second user for co-aggregation processing, transmitting the aggregated volumetric video to a first user device associated with the first user and a second user device associated with the second user.
12. A computer program product, comprising one or more computer-readable storage media and program instructions jointly stored on the one or more computer-readable storage media, the program instructions being executable by a processor to cause the processor to perform operations, the operations including: Selecting the first object in the first volumetric video using a first attribute of the first object; Selecting the second object in the second volumetric video using a second attribute of the second object, wherein the first attribute and the second attribute satisfy an aggregation rule; and Generating an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video includes presenting the first object and the second object simultaneously in the aggregated volumetric video based on the aggregation rule.
13. The computer program product according to claim 12, wherein, The stored program instructions are stored in a computer-readable storage device in a data processing system, and wherein the stored program instructions are transmitted from a remote data processing system via a network.
14. The computer program product according to claim 12, wherein, The stored program instructions are stored in a computer-readable storage device in a server data processing system, and wherein the stored program instructions are downloaded to a remote data processing system in response to a request via a network for a computer-readable storage device associated with the remote data processing system, the computer program product further comprising: Program instructions for metering the use of the program instructions associated with the request; and Program instructions for generating an invoice based on the metered use.
15. The computer program product according to claim 12, the operations further comprising: Transmitting the aggregated volumetric video to a user device; Detect a user input indicating selection of the first volumetric video; and In response to the user input, convert from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device.
16. The computer program product according to claim 12, wherein the operation further comprises: Transmit a multi-view selection image to the user device, wherein the multi-view selection image includes a preview image associated with the aggregated volumetric video; Detect a user input indicating selection of the preview image; and In response to the user input, transmit the aggregated volumetric video to the user device.
17. The computer program product according to claim 12, wherein the operation further comprises: Detect a user input from the user indicating selection of the first volumetric video; and Determine whether the user is authorized to view the first volumetric video by comparing a permission rule associated with the user and an access rule associated with the first volumetric video, wherein generating the aggregated volumetric video is in response to determining that the user is authorized to view the first volumetric video.
18. A computer system, comprising a processor and one or more computer-readable storage media, and program instructions jointly stored on the one or more computer-readable storage media, the program instructions being executable by the processor to cause the processor to perform operations, the operations comprising: Select the first object in the first volumetric video using a first attribute of the first object; Select the second object in the second volumetric video using a second attribute of the second object, wherein the first attribute and the second attribute satisfy an aggregation rule; and Generate an aggregated volumetric video from the first volumetric video and the second volumetric video, wherein generating the aggregated video comprises presenting the first object and the second object simultaneously in the aggregated volumetric video based on the aggregation rule.
19. The computer system according to claim 18, wherein the operation further comprises: Transmit the aggregated volumetric video to the user device; Detect a user input indicating selection of the first volumetric video; and In response to the user input, convert from transmitting the aggregated volumetric video to the user device to transmitting the first volumetric video to the user device.
20. The computer system according to claim 18, wherein the operation further comprises: Transmit a multi-view selection image to the user device, wherein the multi-view selection image includes a preview image associated with the aggregated volumetric video; Detect a user input indicating selection of the preview image; and In response to the user input, transmit the aggregated volumetric video to the user device.
Citation Information
Cited By
Personalized aggregation of volumetric videos
US12505670B2