Object-level access control for multimedia content
By segmenting multimedia content into object-level data streams and applying encryption and digital signatures, the system addresses the lack of granularity in existing systems, providing secure and efficient access control and verification.
Patent Information
- Application Number
- PCT/CA2025/050768
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Existing multimedia management systems lack the granularity to provide selective access and verification of specific elements within multimedia content, failing to address complex privacy and access control needs, particularly in scenarios involving sensitive information and time-based access changes.
The system segments multimedia content into object-level data streams, encrypts each stream with specific keys, and applies digital signatures for authentication, enabling granular access control and verification of individual segments.
Enables precise control over access to specific elements within multimedia content, ensuring privacy and authenticity, facilitating secure and efficient sharing across diverse domains.
Smart Images

Figure CA2025050768_04122025_PF_FP_ABST
Abstract
Description
OBJECT-LEVEL ACCESS CONTROL FOR MULTIMEDIACONTENTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of United States patent application no. 63 / 654,772, filed on May 31 , 2024, and entitled “OBJECT-LEVEL ACCESS CONTROL FOR MULTIMEDIA CONTENT”, the entirety of which is hereby incorporated by reference herein.TECHNICAL FIELD
[0002] The present disclosure relates to systems and methods for managing multimedia content and in particular to systems and methods for managing multimedia content through media segmentation, permission control and verification at objectlevel.BACKGROUND
[0003] The widespread sharing of multimedia content, including video and audio files, has revolutionized communication and information dissemination. However, this ease of sharing has also brought forth significant challenges related to privacy, security, and content authenticity. The emergence of sophisticated content manipulation techniques, such as deep fakes, has further exacerbated these concerns, making it increasingly difficult to distinguish genuine content from fabricated or altered versions. In particular, a similar level of advancement in multimedia management approaches in response to the increased burden and challenges for multimedia control and security have yet to be observed.
[0004] Existing solutions for protecting multimedia content often rely on filelevel encryption and file-level digital signatures. While these methods offer some degree of security and integrity protection, they lack the granularity required to address the complex privacy and access control needs of today's digital landscape. Current multimedia management systems and methods are limited to file-level and media stream-level encryption, which treat the entire content as a single unit, making it impossible to selectively share or restrict access to specific elements within the media. Similarly, traditional digital signatures verify the integrity of the entire file butdo not provide mechanisms for verifying the authenticity of individual objects or segments within the content.
[0005] Current file-level encryption and signature schemes, which are typically used for multimedia management, not only lack the granularity to address complex access control requirements but also fail to provide mechanisms for controlling access based on specific time intervals. This is crucial in scenarios where access privileges might change over time or where certain segments of the multimedia content need to be restricted to specific timeframes. Moreover, the lack of object-level verification mechanisms makes it difficult to establish the authenticity and integrity of specific portions of the media, which is particularly important in scenarios involving legal evidence or journalistic integrity.
[0006] Accordingly, systems and methods for improved multimedia content management and in particular object-level access and verification of multimedia content remain highly desirable.SUMMARY
[0007] In accordance with one aspect of the present disclosure, a method of managing access to multimedia content is disclosed, the method comprising: receiving a multimedia file; generating at least one data stream on a basis of at least one object of interest in the multimedia file such that each data stream of the at least one data stream is a segment of the multimedia file that comprises and corresponds to a respective object of interest of the of at least one object of interest; generating at least one encryption key, each encryption key of the at least one encryption key corresponding to a respective data stream of the at least one data stream; encrypting the at least one data stream with a corresponding encryption key of the at least one encryption key, the encrypted at least one data stream configured to be decrypted by the corresponding encryption key; assigning a digital signature to each data stream of the at least one data stream; and generating a media container by packaging the at least one data stream, the at least one encryption key, and the digital signature.
[0008] In some aspects, the method further comprises: encrypting a timebased segment of one of the at least one data stream with one of at least one time-specific encryption key, the at least one time-specific encryption key generated as one or more of the at least one encryption keys.
[0009] In some aspects, the method further comprises: encrypting the at least one encryption key with a first public key; and storing the at least one encryption key in a secure repository.
[0010] In some aspects, the method further comprises: generating at least one hash value, each hash value of the at least one hash value corresponding to a respective data steam of the at least one data stream; and packaging the at least one hash value into the media container, each hash value of the at least one hash value assigned the digital signature.
[0011] In some aspects, the method further comprises: generating an overall hash value corresponding to the multimedia file; and packaging the overall hash value into the media container, the overall hash value assigned the digital signature.
[0012] In some aspects, the method further comprises: identifying the at least one object of interest based on input from a user; and tracking the at least one object of interest to generate the at least one data stream.
[0013] In some aspects, the at least one object of interest is identified by the user using natural language queries, selection of a particular area, instance segmentation, or combinations thereof; the at least one object of interest is visualbased, audio-based, text-based, or combinations thereof; the visual-based at least one object of interest is tracked using a Kalman filter algorithm, an optical flow algorithm, an Al algorithm, or combinations thereof; and the audio-based at least one object of interest is tracked using audio source separation, speaker diarization, audio fingerprinting, feature matching, or combinations thereof.
[0014] In some aspects, the multimedia file is segmented using a Contrastive Language-Image Pre-training (CLIP) model or a Segment Anything Model (SAM).
[0015] In some aspects, the method further comprises: encoding the at least one data stream, wherein the at least one data stream is video-based, audio-based, text-based, or combinations thereof; the video-based at least one data streamencoded using H.264 or HEVC codec; the audio-based at least one data stream encoded using AAC or MP3 codec; and the at least one data stream comprising one or more mask images for separating the respective object of interest from background, and the one or more mask images are encoded using Run-Length Encoding (RLE), bit-plane encoding, PNG compression, GIF compression, H.264 codec, HEVC codec, or combinations thereof.
[0016] In some aspects, the video-based at least one data stream is processed by binary mask generation, alpha channel encoding, color key filling, or combinations thereof to isolate the respective object of interest; the audio-based at least one data stream is processed by silence detection and removal, audio segmentation, background noise preservation, or combinations thereof to isolate the respective object of interest.
[0017] In some aspects, the method further comprises: embedding, in a header of the media container, metadata for synchronizing the at least one data stream, the metadata comprising: frame number, decode timestamp, presentation timestamp, object identifier, spatial information, object timestamp, or combinations thereof; and synchronizing the at least one data stream.
[0018] In some aspects, the secure repository is a key-value database, a hardware security module, or a blockchain storage.
[0019] In some aspects, the at least one data stream is encrypted using Advanced Encryption Standard (AES) with a key length of 256 bits.
[0020] In some aspects, the at least one encryption key is encrypted using an asymmetric encryption algorithm and wherein the asymmetric encryption algorithm is Rivest-Shamir-Adleman or Elliptic Curve Cryptography.
[0021] In some aspects, a format of the media container is MKV, MP4, or AVI.
[0022] In some aspects, each data stream of the at least one data stream is assigned a unique identifier, codec information, timestamps, language codes, or combinations thereof; a header of the media container comprises: format version, file size, duration, list of the plurality of data streams, or combinations thereof; and-metadata of the media container comprises the at least one encryption key and digital signature.
[0023] In some aspects, the method further comprises: receiving a request for access of multimedia content; determining at least one authorized data stream from the at least one data stream; retrieving at least one authorized encryption key from the at least one encryption key, each authorized encryption key of the at least one authorized encryption key corresponding to a respective authorized data stream of the at least one authorized data stream; and generating an authorized media container by packaging the at least one authorized encryption key into the media container.
[0024] In some aspects, the method further comprises: authenticating the request; and decrypting the at least one authorized encryption key using a first private key.
[0025] In some aspects, the method further comprises: encrypting the at least one authorized encryption key with a second public key; and storing the at least one authorized encryption key in the secure repository, each of the at least one authorized encryption key configured to decrypt a respective authorized data stream of the at least one authorized data stream for access; and the at least one authorized encryption key configured to be decrypted by a second private key.
[0026] In some aspects, the method further comprises: decrypting at the at least one authorized encryption key with a second private key; and decrypting the at least one authorized data stream with the at least one authorized encryption key.
[0027] In some aspects, the method further comprises: in response to a request to access the encrypted time-based segment, decrypting the time-based segment of the one of the at least one data stream with a corresponding one of the at least one time-specific encryption key; and outputting the decrypted time-based segment to a video display or audio speaker.
[0028] In some aspects, the method further comprises: receiving a content verification request corresponding to at least one data stream for verification from theat least one data stream; and returning at least one verification hash value for comparison with at least one corresponding current hash value of the at least one data stream for verification, each verification hash value of the at least one verification hash value corresponding to a respective data stream of the at least one data stream for verification.
[0029] In some aspects, the method further comprises: receiving a verification digital signature; assigning the verification signature to the at least one verification hash value; and packaging the verification digital signature into the media container.
[0030] In some aspects, the method further comprises: displaying the at least one data stream for user playback and management.
[0031] In accordance with another aspect of the present disclosure, a method of managing access to multimedia content is disclosed, the method comprising: receiving a multimedia file; generating at least one data stream for each of at least one object of interest in the multimedia file such that each data stream of the at least one data stream is a segment of the multimedia file that comprises and corresponds to a respective representation of a respective object of interest of the at least one object of interest; generating at least one encryption key, each encryption key of the at least one encryption key corresponding to a respective data stream of the at least one data stream; encrypting each of the at least one data stream with a corresponding encryption key of the at least one encryption key, each encrypted data stream configured to be decrypted by the corresponding encryption key; receiving at least one digital signature for the at least one data stream; and generating a media container by packaging the at least one data stream and the at least one digital signature.
[0032] In some aspects, a plurality of data streams are generated for each of the at least one object of interest.
[0033] In some aspects, the method further comprises: generating at least one encryption key identifier, each being a unique identifier corresponding to a respective encryption key; and packaging the at least one encryption key identifier in the media container.
[0034] In some aspects, the method further comprises: storing the at least one encryption key in a key management service secure repository; each of the at least one encryption key is retrievable for decrypting the respective data stream from the key management service using a respective encryption key identifier and corresponding authentication data.
[0035] In some aspects, the method further comprises: embedding a representation table as metadata in the media container file for each of the at least one data stream; the representation table comprises: an identifier of the respective object of interest and a type of the respective representation.
[0036] In some aspects, the representation table further comprises a respective encryption key identifier.
[0037] In some aspects, the representation table further comprises codec information, a location pointer, language information, resolution information, or combinations thereof.
[0038] In some aspects, the method further comprises: embedding one or more privacy directives as metadata in the media container file for each of the plurality of data streams; the one or more privacy directives comprise information for accessing the respective data stream.
[0039] In some aspects, the one or more privacy directives comprise: consent status data, default representation for rendering data, retention policy data, geographical policy data, purpose policy data, or combinations thereof.
[0040] In some aspects, each respective representation is a high quality representation, a low resolution representation, an de-identified representation, a text description representation, a sensor data representation, or an embedding representation.
[0041] In some aspects, the de-identified representation is generated by modifying the respective object of interest using at least one artificial intelligence model to generate a representation of the respective object of interest lacking features permitting identification of the respective object of interest.
[0042] In some aspects, the at least one artificial intelligence model is configured to modify features in the respective according to a predetermined policy or threshold by performing feature identification and feature alteration.
[0043] In some aspects, the method further comprises: evaluating the modified object of interest based on one or more privacy policies; and storing records of the evaluating as metadata in the media container file.
[0044] In some aspects, the method further comprises: generating one or more root digests for the at least one data stream and corresponding metadata.
[0045] In some aspects, each of the at least one digital signature is assigned to a respective root digest for verification.
[0046] In some aspects, each root digest comprises a plurality of leaf nodes, each corresponding to: one of the at least one data stream, one of the at least one object of interest, or a representation table comprising metadata for the respective representation.
[0047] In some aspects, metadata of the media container file is accessible in a bitstream using a supplemental-enhancement-information (SEI) message or from the container box.
[0048] In some aspects, the method further comprises: encrypting a timebased segment of one of the at least one data stream with one of at least one timespecific encryption key, the at least one time-specific encryption key is generated as one or more of the at least one encryption key.
[0049] In some aspects, the time-based segment data stream is implemented using a Self-Decodable Access Point (SDAP) codec.
[0050] In some aspects, the method further comprises: generating a streaming manifest comprising: data of the at least one object of interest; data of each respective representation; and the plurality of encryption key identifiers.
[0051] In some aspects, the streaming manifest is configured to be decoded for streaming one or more of the at least one data stream based on data in the streaming manifest and an authorization of the at least one data stream.
[0052] In some aspects, the media container comprises metadata corresponding to revocation conditions for one or more of the at least one data stream.
[0053] In some aspects, the method further comprises: identifying the at least one object of interest based on input from a user; and tracking the at least one object of interest to generate the at least one data stream.
[0054] In some aspects, the at least one object of interest is identified by the user using natural language queries, selection of a particular area, instance segmentation, or combinations thereof; the at least one object of interest is visualbased, audio-based, text-based, or combinations thereof; the visual-based at least one object of interest is tracked using a Kalman filter algorithm, an optical flow algorithm, an Al algorithm, or combinations thereof; and the audio-based at least one object of interest is tracked using audio source separation, speaker diarization, audio fingerprinting, feature matching, or combinations thereof.
[0055] In some aspects, the method further comprises: encoding the at least one data stream, the at least one data stream is video-based, audio-based, textbased, or combinations thereof; the video-based at least one data stream is encoded using H.264 or HEVC codec; the audio-based at least one data stream is encoded using AAC or MP3 codec; and the at least one data stream comprises one or more mask images for separating the respective object of interest from background, and the one or more mask images are encoded using Run-Length Encoding (RLE), bitplane encoding, PNG compression, GIF compression, H.264 codec, HEVC codec, or combinations thereof.
[0056] In some aspects, the video-based at least one data stream is processed by binary mask generation, alpha channel encoding, color key filling, or combinations thereof to isolate the respective object of interest; and the audio-based at least one data stream is processed by silence detection and removal, audio segmentation,background noise preservation, or combinations thereof to isolate the respective object of interest.
[0057] In some aspects, the method further comprises: embedding, in a header of the media container, metadata for synchronizing the at least one data stream, the metadata comprising: frame number, decode timestamp, presentation timestamp, object identifier, spatial information, object timestamp, or combinations thereof; and synchronizing the at least one data stream.
[0058] In some aspects, the method further comprises: receiving a request for access of multimedia content; authenticating the request for access; determining at least one authorized data stream from the at least one data stream; retrieving at least one authorized encryption key identifier from the at least one encryption key identifier, each authorized encryption key identifier corresponding to a respective authorized data stream; retrieving at least one authorized encryption key corresponding to the at least one authorized encryption key identifier from the at least one encryption key from the key management service; and decrypting the at least one authorized data stream using the at least one authorized encryption key.
[0059] In some aspects, the key management service is configured to return the at least one authorized encryption key based on the authenticating and the at least one authorized encryption key identifier.
[0060] In some aspects, the key management service is configured to query the media container file for a revocation condition for the at least one authorized encryption key.
[0061] In some aspects, the at least one authorized encryption key is retrieved as a wrapped encryption key wrapped by the key management service, and the method further comprises: unwrapping the at least one authorized encryption key in a local channel session or for a predetermined maximum duration.
[0062] In some aspects, the method further comprises: rendering the at least one authorized data stream for viewing according to metadata corresponding to the at least one authorized data stream.
[0063] In some aspects, the method further comprises: receiving a content verification request corresponding for the plurality of data streams; verifying the at least one digital signature; recomputing one or more verification root digests, each corresponding to a respective root digest and recomputed using data corresponding to the respective root digest; verifying the at least one data stream by comparing the one or more verification root digests to the one or more root digests; and assigning a verification signature for each root digest.
[0064] In accordance with another aspect of the present disclosure, a method for managing access to multimedia content is disclosed, comprising: receiving a request for access of multimedia content, the multimedia content comprising a plurality of data streams, each data stream of the plurality of data streams is a segment of the multimedia content that comprises and corresponds to a respective representation of a respective object of interest in the multimedia content; authenticating the request for access, the request for accessing identifying one or more first data streams for access from the at least one data stream; retrieving one or more encryption key identifiers, each corresponding to a respective one of the first data streams; retrieving one or more encryption keys corresponding to the one or more first data streams from a key management service in response to authentication by the key management service based on the authenticating of the request for access and the one or more encryption key identifiers; and decrypting the one or more first data streams using the one or more encryption keys in response to the request for access.
[0065] In accordance with another aspect of the present disclosure, a method of managing access to multimedia content is disclosed, the method comprising: receiving a multimedia file; generating a plurality of data streams for each of at least one object of interest in the multimedia file such that each data stream of the plurality of data streams is a segment of the multimedia file that comprises and corresponds to a respective representation of a respective object of interest of the at least one object of interest; generating a plurality of encryption keys, each encryption key of the plurality of encryption keys corresponding to a respective data stream of the plurality of data streams; encrypting each of the plurality of data streams with a correspondingencryption key of the plurality of encryption keys, each encrypted data stream configured to be decrypted by the corresponding encryption key; generating a plurality of encryption key identifiers, each being a unique identifier corresponding to a respective encryption key; and generating a media container by packaging the plurality of data streams and the plurality of encryption key identifiers.
[0066] In some aspects, the method further comprises: generating a root digest for the plurality of data streams and corresponding metadata; and assigning a digital signature to the root digest for verification.
[0067] In some aspects, the plurality of encryption keys are stored in a key management service secure repository; and wherein each of the plurality of encryption keys is retrievable for decrypting the respective data stream from the key management service using a respective encryption key identifier and corresponding authentication data.
[0068] In accordance with another aspect of the present disclosure, a method of managing access to multimedia content is disclosed, the method comprising: requesting access, at an access control server, of one or more of a plurality of encrypted data streams each corresponding to a segment of a multimedia file and a representation of a respective object of interest; receiving authorization to the one or more encrypted data streams from the access control server; retrieving one or more encryption key identifiers from the one or more encrypted data streams, each encryption key identifier corresponding to a respective encrypted data stream; requesting one or more encryption keys corresponding to the one or more encryption key identifiers from a key management server based on the authorization and using the one or more encryption keys; receiving the one or more encryption keys; and decrypting the one or more encrypted data streams using the one or more encryption keys.
[0069] In accordance with another aspect of the present disclosure, a system for managing access to multimedia content is disclosed, the system comprising: one or more processing units configured to perform the method of any one of the above aspects.
[0070] In accordance with another aspect of the present disclosure, a non- transitory computer-readable medium having computer readable instructions stored thereon is disclosed, which, when executed by at least one processor, causes the at least one processor to perform the method of any one of the above aspects.BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Further features and advantages of the present disclosure will become apparent from the following detailed description, taken in combination with the appended drawings, in which:FIG. 1 depicts a representation of a system for managing multimedia content, according to an example embodiment.FIG. 2A depicts generation of a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIG. 2B depicts a method of generating a plurality of streams corresponding to identified objects of interest, according to an example embodiment.FIG. 3 depicts a method of generating a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIGs. 4A and 4B depict allocation of access to media content in a multimedia container file for providing segmented access and verification of media content, according to example embodiments.FIGs. 5A and 5B depict methods of allocating access of media content in a multimedia container file for providing segmented access and verification of media content, according to example embodiments.FIG. 6 depicts verification of media content in a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIGs. 7 A and 7B depict methods of verifying media content in a multimedia container file for providing segmented access and verification of media content, according to example embodiments.FIG. 8A depicts a sequence diagram of generating a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIG. 8B depicts a sequence diagram of allocating access of media content in a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIG. 8C depicts a sequence diagram of verifying media content in a multimedia container file for providing segmented access and verification of media content, according to an example embodiment.FIGs. 9A and 9B depict segmentation of source media into a plurality of media streams, according to example embodiments.FIGs. 10A and 10B depict interfaces of a system for managing multimedia content, according to example embodiments.FIGs. 11 A and 11 B depict segmentation approaches for a plurality of media streams, according to example embodiments.FIGs. 12A and 12B depict generated representations of original objects, according to example embodiments.
[0072] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0073] Growing concerns surrounding privacy and content authenticity necessitate the development of more sophisticated access control and verification mechanisms for multimedia content. Object-level access control offers a promising solution by enabling granular control over individual elements within multimediacontent. Specifically, the multimedia content may be segmented according to objects of interest in the multimedia content. This approach allows content owners to define specific permissions for different users, ensuring that only authorized individuals have access to relevant portions of the content. Additionally, robust content authentication methods, such as digital signatures for individual objects or segments, are essential for establishing trust and verifying the integrity and origin of multimedia content. The present disclosure aims to provide a flexible, secure, and trusted framework for addressing these concerns by enabling content sharing across diverse domains, such as smart cities, law enforcement, insurance, journalism, and beyond. Accordingly, in accordance with the present disclosure, it is possible to strike a balance between protecting privacy and enabling effective collaboration and information sharing among authorized parties.
[0074] In particular, the widespread sharing of multimedia content has introduced significant challenges related to privacy, security, and content authenticity. Sophisticated content manipulation techniques, such as deep fakes, exacerbate these concerns by making it difficult to distinguish genuine content from fabricated or altered versions. Current multimedia management approaches have not adequately addressed these increasing challenges in multimedia control and security.
[0075] The limitations of current solutions are particularly evident in scenarios where sensitive information or the privacy of individuals is involved. Multimedia management is often restricted to the manual dissection of a video stream into multiple files by time, which is a non-ideal solution for a number of reasons. Specifically, it can be very difficult to manage access to parts or sections of a single multimedia content which may be particularly challenging as unrelated individuals or objects can have overlapping presence within the multimedia content. Using an example of a video captured by a surveillance camera for a traffic accident, law enforcement agencies may require access to the full video for their investigation, while insurance companies might only need specific portions relevant to their claims process. Additionally, individuals involved in the accident should have access to their own image or video segments but not necessarily the entire footage, and that these individuals should not have information pertaining to unrelated parties, even if thoseparties appear in the video at the same time. Furthermore, it is crucial to protect the privacy of bystanders or unrelated individuals who may have been captured in the video inadvertently. However, the options for providing customized streams or files based on relevant portions of the multimedia content as a whole are limited and can be very cost and labour prohibitive.
[0076] For example, existing solutions for protecting multimedia content often rely on file-level encryption and file-level digital signatures. While these methods offer some degree of security and integrity protection, they lack the granularity required to address the complex privacy and access control needs of today's digital landscape. Current multimedia management systems and methods are limited to file-level and media stream-level encryption, which treat the entire content as a single unit, making it impossible to selectively share or restrict access to specific elements within the media. Similarly, traditional digital signatures verify the integrity of the entire file but do not provide mechanisms for verifying the authenticity of individual objects or segments within the content.
[0077] Existing file-level encryption and signature schemes, which are typically used for multimedia management, not only lack the granularity to address complex access control requirements but also fail to provide mechanisms for controlling access based on specific time intervals. This can be crucial in scenarios where access privileges might change over time or where certain segments of the multimedia content need to be restricted to specific timeframes. Moreover, the lack of object-level verification mechanisms makes it difficult to establish the authenticity and integrity of specific portions of the media, which is particularly important in scenarios involving legal evidence or journalistic integrity. Furthermore, existing systems typically do not support the efficient generation, storage, and cryptographically controlled access to multiple, semantically distinct representations of the same underlying object within a multimedia file, nor do they typically integrate rich, privacy-aware metadata directly into the multimedia codec design to make the content self-describing and selfenforcing regarding object access and privacy directives. Consequently, delivering such granular, privacy-respecting content efficiently via adaptive streaming protocols also remains a significant challenge.
[0078] Specifically, existing multimedia management limitations are evident in scenarios involving sensitive information, individual privacy, or the need for granular control over different facets of the same object. Traditional management methodologies often rely on coarse-grained operations, like manually dissecting video into multiple time-based files, which fails to address semantic complexity. This makes managing access to specific parts or semantic elements of a single recording difficult, especially with overlapping, unrelated individuals or objects. For instance, in a traffic accident video, law enforcement might need complete, unaltered recordings with highest-fidelity representations of all objects. Insurance companies might only require specific representations of involved vehicles and contextual objects, not identifiable features of bystanders. Involved individuals should ideally access representations pertinent to themselves and their property (e.g., their vehicle's visual data, their spoken statements as a textual transcript), but not necessarily all footage or high- fidelity representations of other parties or uninvolved bystanders, whose privacy is paramount. Further, existing systems offer limited and cumbersome options for such customized, multi-stakeholder access to different semantic layers or representations of an event. Manual redaction, versioning, and distributing distinct files for each view are labor-intensive, error-prone, costly, and challenges data integrity, auditability, and synchronized understanding. These methods do not support a single, authoritative source file with multiple, independently secured, verifiable representations of its constituent objects of interest.
[0079] Correspondingly, the present disclosure relates to novel systems and methods for granular access control, content authentication, and secure sharing of multimedia content, as to help address the limitations of existing file-level protection mechanisms. By employing deep learning for object segmentation and creating separate encrypted streams, the present disclosure can enable selective access and privacy protection at the object level. Further, it is possible to provide digital signatures for ensuring content integrity and authenticity, from the owner’s and / or users’ (e.g. witnesses) signatures to help provide a robust verification process. In particular, the present disclosure relates to systems and methods for managing multimedia content, and more particularly to systems and methods for managing multimedia content through object-level media segmentation, permission control, and verification,including the management of multiple distinct semantic representations per object of interest, codec-integrated privacy-aware metadata, and adaptive streaming of said object representations.
[0080] In particular, the present disclosure can address limitations in existing systems and methods by providing granular access control, robust content authentication, and secure sharing of multimedia content at the level of individual objects of interest and their distinct representations. Preferred embodiments may employ an object-centric codec with integrated artificial intelligence (Al) for object identification and multi-representation generation. By creating separate, independently encrypted data stream representations for each facet of an object of interest (e.g., high-fidelity visual, de-identified visual, textual description, sensorbased data), the present disclosure can enable selective access and privacy protection with unprecedented granularity. Furthermore, integrating privacy-aware metadata directly within the media container or codec bitstream, coupled with hierarchical integrity mechanisms (e.g., signed Merkle trees covering content representations and metadata), can allow verifiable authenticity and adherence to privacy directives. This approach can facilitate digital signatures for the integrity of specific object layer representations and their governing policies (e.g., not just for an entire file) originating from the content owner and potentially augmented by third-party attestations.
[0081] The systems and methods can comprise a number of components and functionalities. Object identification and segmentation may be provided, in which deep learning artificial intelligence (Al) models can be employed to identify and segment objects within media content based on natural language queries or automated object detection. Object tracking can also be applied to generate a discrete stream focused on each object instance. Stream encryption may also be provided, in which each object or group of object streams can be individually encrypted using symmetric and / or asymmetric encryption algorithms to ensure secure storage and transmission. Access control may also be provided, in which a flexible access control mechanism can allow users to define permissions for individual object streams for granting or restricting access as needed. Key management and secure sharing schemes canalso be implemented to allow keys to be provided to authorized users to permit decrypting and accessing of individual streams. Playback synchronization of the media can also be performed by decrypting the streams the user has keys to and combining the streams. Digital signature capabilities may also be provided, in which for each object stream, a unique identifier, such as a hash, can be calculated. The identifier can be digitally signed to confirm the integrity of the content and verify the ownership of the media. Additionally, authorized users can provide their own digital signatures for further verification and authentication.
[0082] By enabling object-level segmentation, encryption, access control and verification, the present disclosure provides a flexible and secure framework for sharing multimedia content while protecting privacy. The present disclosure may further allow for time-based access control, enabling the definition of permissions for specific object streams within designated time intervals as described further herein.
[0083] The present disclosure may provide systems and methods for generating a multimedia container file for managing multimedia content. A number of objects of interest may be identified and tracked within a multimedia file. These objects of interest may be isolated and tracked to generate a plurality of streams, each of which are segmented from the multimedia file and corresponding to a particular object of interest. An encryption key is generated for each of the streams, which is used to encrypt the corresponding stream. The encryption keys can also be encrypted using a public key of the owner of the multimedia file and stored in secured storage. A hash value can also be generated for verification purposes for each of the streams. The owner can additionally digitally sign each of the streams or the generated hash values. The final media container file is generated by packaging the streams, the encryption keys, the signatures (e.g. by including the digitally signed content), and the hash values. The owner may be then provided with the media container file, which they can distribute as desired with access to particular streams or for verification purposes.
[0084] The present disclosure may also provide systems and methods for providing access to streams within the media container file. An authorized user, for example the owner of the media container file, can select a stream or streams in themedia container file to be authorized to another user. The encryption keys corresponding to the streams to be authorized can be retrieved from the plurality of encryption keys that were stored and encrypted using a public key of the other user, which may also be stored in the secure storage. The media container file may be updated to include the encryption keys for the other user. Alternatively, a new media container file may be generated using the streams, the encryption keys for the owner and / or other user, the signatures, and the hash values. This media container file can be provided to the other user. The other user can use their private key to decrypt the decryption keys, which they can use to decrypt the streams that they are authorized to access.
[0085] The present disclosure may also provide systems and methods for verifying content in the media container file. A user having access to one or more streams within a media container file can, at the request of the owner for example, verify that the streams they are accessing have not been modified by comparing the packaged hash values of the streams with the current hash values of the streams. The user can additionally sign each of the verified streams or the verified hash values. The media container file may be updated to include the verification signatures. Alternatively, a new media container file may be generated using the streams, the encryption keys, the hash values, and the signatures of the owner and / or other user. This media container file can be provided to the owner.
[0086] By combining advanced deep learning models with sophisticated encryption and digital signature techniques in multimedia content management, the present disclosure can provide a secure, flexible framework for object-level access control and privacy protection. Further, users may be empowered to exercise precise control over specific elements within video and audio files to safeguard the content's security, integrity, and authenticity. By enabling fine-grained access control coupled with privacy safeguards at the object level, the systems and methods of the present disclosure may have a wide range of applications including enhancing corporate communications, enabling more effective remote collaboration, securing sensitive medical data, and enriching educational resources. Moreover, it can be beneficial to implement the present disclosure in the media and entertainment sectors to create amore equitable digital ecosystem for creators and consumers alike by protecting consumer privacy and copyright holder works from illegitimate access and tampering.
[0087] As described further herein, the present disclosure can comprise a plurality of components and functionalities for object-centric multimedia management with integrated security and privacy, such as:1. Object Identification and Segmentation Module: Al models (e.g., general segmentation or vision models) can be used to identify, classify, and segment semantic objects of interest within input multimedia, guided by user input (e.g., natural language queries, interactive selection) or operating autonomously.2. Multi-Representation Generation Module: For each identified object of interest, the system can generate multiple distinct data stream representations (e.g., high-quality visual, de-identified visual, low-resolution visual, textual descriptions, sensor-based streams, feature embeddings).3. Object-Centric Codec and Packaging Module: An object-aware and privacy- aware codec can manage encoding of these multiple representations (e.g., potentially as distinct object layers / tracks) and embeds integrated privacy- aware metadata (e.g., including Object IDs, Representation Tables, key representations or IDs (KI Ds), Actionable Privacy Directives). It can package these elements with a hierarchical integrity signature into a self-describing media container.4. Representation-Specific Encryption Module: The system can encrypt each distinct data stream representation with its unique, representation-specific content-encryption key (CEK).5. Key Management Service or System (KMS) and Access Control Module (ACM): The system can manage the CEK lifecycle, including secure key storage, zero-knowledge key referencing (e.g., where the media container holds KI Ds only), secure key delivery, and key revocation. The Access Control module can determine user authorization for specific object representations based on policies and credentials, issuing tokens for KMS interaction.6. Integrity Verification Module: The system can enable verification of the hierarchical integrity digest to detect tampering of representations or governing metadata, and supports associating additional verification signatures.7. Adaptive Streaming Module: The system can generates manifests (e.g., DASH MPD, HLS M3U8) describing available object representations and security requirements, enabling users to stream only authorized representations.8. Compliant Client Playback Module: The system can authenticate users; retrieve keys from KMS; parse manifests; request / decrypt segments; verify integrity; respect privacy directives; and composite representations for synchronized playback.9. Secure Hardware Appliance: The system can be implemented as an edge device or server incorporating an HSM and GPU / FPGA acceleration to execute modules in real-time.
[0088] Accordingly, resulting from growing concerns about multimedia privacy and content authenticity necessitate more sophisticated access control and verification, the present disclosure can enhance object-level access control by managing multiple, semantically distinct representations per object of interest, enabling granular control over individual elements and their forms within multimedia content. Specifically, content can be segmented by objects of interest. For each identified object of interest, multiple representations (e.g., high-fidelity, de-identified, textual, sensor-based) may be generated, each independently encrypted and associated with specific, embedded access policies and privacy directives. This allows content owners to define precise user permissions, ensuring authorized access only to relevant representations of relevant objects of interest. Furthermore, robust content authentication, such as digital signatures over a hierarchical integrity digest covering content representations and their governing metadata can be utilized to establish trust and verify integrity. The present disclosure can therefore provide a flexible, secure framework for facilitating content sharing across diverse domains (e.g., smart cities, law enforcement, and journalism) while balancing privacy protection with effective collaboration among authorized parties.
[0089] By enabling object-level segmentation into multiple distinct semantic representations -each with its own encryption - and by providing for granular access control, integrated privacy directives, and comprehensive integrity verification, the present disclosure can offer a flexible, robust, and secure framework for sharing multimedia content while protecting privacy and ensuring authenticity. This framework can support basic access control and advanced features such as time-segment specific encryption for representations, performance-bound Al de-identification, and verifiable key revocation for "right-to-be-forgotten" compliance.
[0090] A particularly useful example application of the present disclosure can be for the processing of video footage from a traffic camera that has captured a vehicle accident. The video can be automatically split into streams corresponding to each vehicle and person. Each party involved in the accident may be provided the access to decrypt only the streams relevant to them. Specifically, an individual would have access to the stream tracking their own vehicle but not those of others. Integrity verification would allow all parties to confirm the video of their vehicle is unmodified, and to digitally sign the streams to attest that it is indeed authentic footage of their vehicle. Insurance companies and police can be granted access to only the streams and portions of the video relevant to their investigation. Other individuals captured in the periphery can have their streams fully encrypted to protect their privacy.
[0091] Furthermore, the present disclosure can have a degree of flexibility in design and could be adapted for emerging trends and technologies in key management and secure content sharing. For example, the systems and methods of the present disclosure can leverage traditional cryptographic techniques or integrate decentralized solutions such as blockchain to meet the ever-changing landscape of digital security and privacy.
[0092] Embodiments are described below, by way of example only, with reference to FIGs. 1-12B. It would be appreciated that various algorithms and / or models are described with respect to the figures and that the specific implementations of these models would be known to a person skilled in the art in view of the present disclosure.
[0093] FIG. 1 depicts a representation of a system for managing multimedia content. The system comprises one or more servers 108 communicatively coupled to a communication network 106 (e.g. the internet). The implementation of servers 108 is not restrictive and servers 108 may be a physical, cloud-based, or a hybrid thereof, for example. A first user 102, second user 126, and third user 128 can communicate with servers 108 over the communication network 106 via user devices, for example a computer 104 as shown in FIG. 1. The present system can accommodate further users or systems / entities in place of users, as needed. The user devices are not restricted to those expressly shown and may be any suitable device known in the art such as smart phones and tablets. Servers 108 may provide a graphical user interface (GUI) on the user devices of users 102, 126, 128 for ease of communication and operation control by the users 102, 126, 128. The implementation of the GUI is not restrictive and may be, for example, a mobile / computer application or a web page.
[0094] In accordance with the present disclosure, the servers 108 are configured to receive a multimedia file 122 from the first user 102. That is, the first user 102 can be considered as the “owner” of the multimedia file 122, which may be, for example, a video file. The first user 102 may be interested in providing portions or segments of the multimedia file 122 to other users. For example, the multimedia file 122 may be a video file of a traffic accident. The first user 102 may wish to provide video stream(s) of the two cars that have collided to their insurance company. However, they may wish to redact or censor the footage of other vehicles and bystanders not involved in the collision for the privacy purposes in the video stream(s) provided to the insurance company. The first user 102 may also wish to provide a video stream to a witness to verify the sequence of events. However, they would like such a video stream to include the likeness of the witness in addition to the vehicles involved in the collision with all other vehicles and bystanders redacted or censored. Further, the insurance company would only need to access the first video stream including the collided vehicles and not the second video stream including the witness, and vice versa. Accordingly, the first user 102 may request the servers 108 to generate these segmented data streams that contain only a portion of the original footage, the access of which could be individually controlled. Other uses and purposes for segmenting the multimedia file are also possible.
[0095] In accordance with the present disclosure and as further described herein, the systems and methods provided by the servers 108 can be configured to provide selective access to parts of the multimedia file 122. Specifically, the servers 108 can generate a media container file 124 configured to provide granular or selective access to portions of media content within the multimedia file 122. The first user 102 may provide one or more criteria for generating one or more segmented data streams from the multimedia file 122. Using the previously referenced example, the first user 102 may wish to generate a data stream of their own vehicle, which was involved in a crash, without any other vehicles or bystanders appearing in the data stream. They may also wish to generate a second and third data stream respectively corresponding to a second and third vehicle involved in the accident as well as a fourth data stream of a witness to the accident, all of which without other vehicles and bystanders not expressly selected for inclusion in the streams. As such, the user may prompt the servers 108 to generate a data stream corresponding to their own vehicle, a data stream corresponding to the second vehicle, a data stream corresponding to a third vehicle, and a data stream corresponding to the witness. That is, the first user 102 may provide, identify or select objects of interest to the server 108, which may be vehicles or witnesses in the above example, based upon which the data streams are to be generated. The servers 108 may use one or more algorithms to allow the user to make the selection of the objects of interest. For example, the servers 108 may be configured to allow the selection of objects of interest via natural language selection or brushstroke selection (i.e. selection of a particular area or object using pointers). Alternatively, the servers 108 may automatically detect objects of interest in the multimedia file 122 and generate a list of objects of interest for selection by the first user 102. The selection process of objects of interest will be described further herein.
[0096] The servers can identify the objects of interest in the multimedia file 122 and isolate the objects of interest from the other elements in the multimedia file 122. By tracking the objects of interest through the multimedia file 122, the servers 108 can generate data streams containing the respective object of interest where each of the generated data streams is a part or segment of the multimedia file 122. Further, the generated data streams may be processed, synchronized and encoded by the server 108. It is also possible to generate a hash value (e.g. a unique identifier) for each ofthe data streams and the overall data stream (e.g. all data streams), which can be used for verification purposes. The hash values may be stored for safekeeping. The first user 102 can also provide their digital signature to the servers 108, which may be used to sign the hash values and / or the data streams to ensure authenticity. The servers 108 generate an encryption key for each of the generated data streams and the overall data stream, which may be used to encrypt each of the generated data streams and the overall data stream. The servers 108 can also encrypt the generated encryption keys using a public key of the first user 102. The public key is a part of the public / private key pair, which can be used for encryption and decryption. The servers can store the encryption keys for safeguarding. By packaging the generated data steams, the encryption keys, and the digital signature (e.g. the digitally signed content), the servers 108 can generate a media container file 124 which comprises encrypted data streams 124a for access control. It should be noted that the hash values may also be packaged into the media container file 124. The systems and method for generating the media container file 124 will be described further herein. The media container file 124 can be provided to the first user 102 and / or stored in the servers 108. Further details with respect to object-based stream generation and encryption is described further herein.
[0097] In accordance with the present disclosure, a second user 126 may wish to access the media content of the multimedia file 122. Accordingly, the second user 126 may request access of media content (e.g. access of one or more encrypted data streams 124a) from the first user 102 directly or through the servers 108. If the first user 102 would like to grant access of the media content to the second user 126, for example, by allowing the second user to access one or more encrypted data streams 124a, the servers 108 can authenticate the first user 102 to retrieve the encryption keys (if previously stored) and decrypt the encryption keys (if previously encrypted) with their private key. The servers 108 can process the request to access one or more encrypted data streams 124a and select encryption keys that correspond to encrypted data streams 124a that are authorized to be accessed by the second user 126. The servers 108 can request the second user 126 to provide their public key, which may be used by the servers 108 to encrypt the selected encryption keys. The selected encryption keys can be stored for safeguarding. The selected encryption keys canalso be provided to the second user 126. The servers 108 may generate an updated media container file that also comprises the selected encryption keys. The media container file 124 can be provided to the second user 126 such that they may access the one or more encrypted data streams 124a that they are authorized to access by decrypting the one or more encrypted data streams 124a with the selected encryption keys, which may be decrypted by their private key, if required. Further details and alternative implementations for the second user 126 to access the encrypted data is described further herein.
[0098] Following the previous example, the second user 126 may be the insurance company, which is requesting access to video of the accident. The first user 102 can provide the insurance company with access to the first, second, and third streams in the media container file respectively corresponding to their own vehicle, the second vehicle, and the third vehicle involved in the collision. It should be noted that the first, second, and third streams only comprise footage of the first user’s vehicle, the second vehicle, and the third vehicle, respectively, without footage of other vehicles or bystanders. Further, the insurance company would not be able to access the fourth stream comprising footage of the witness. In some embodiments, the first user 102 configures the servers 108 to provide pre-set authorization for select entities such that the select entities may undergo the authorization process as set forth above without intervention by the first user. It would be appreciated that other uses of the disclosed process are also possible. The systems and methods for providing access to one or more encrypted data streams 124a will be described in further detail herein.
[0099] In accordance with the present disclosure, a third user 128 (e.g. a “witness”) may wish to verify the media content in the multimedia file 122, for example, at the request of the first user 102. The second and third user are not mutually exclusive. The first user 102 may directly or through the servers 108 request the third user 128 to verify media content of the one or more encrypted data streams 124a containing portions of media content of the multimedia file 122. The third user may request to verify one or more encrypted data streams 124a to the servers 108. Accordingly, the third user 128 can be provided access to the one or more encrypteddata streams 124a as described above. The generated hash values are provided to the third user 128, for example, by the servers 108 through the media container file 124 or after being retrieved from storage (if stored). The servers 108 may calculate the current hash values of the one or more encrypted data streams 124a for comparison with the provided hash values. If the hash values match, the matching values confirm that the content of the one or more encrypted data streams 124a is authentic and accordingly has not been altered or tampered with. The third user 128 may also provide their digital signature to the servers 108, which may be used to sign the hash values and / or the one or more encrypted data streams 124a to confirm authenticity. The provided digital signature and associated hash values may be stored for safekeeping. The servers 108 may also generate an updated media container file that also comprises the third user’s signature. The media container file 124 can be provided to the first user 102 as proof of authenticity. Further details and alternative implementations for the third user 128 to verify the encrypted data is described further herein.
[0100] Following the previous example, the third user 126 may be a witness to the traffic accident. The first user 102 may request that the third user 126 verify that the events as shown in the data stream(s) are accurate. Accordingly, the witness may be provided access to the fourth stream comprising footage of themselves in the accident. The witness may verify that the events shown in the fourth stream are accurate and provide their digital signature as proof. The first user 102 may then be able to use the media container file 124 as evidence for the events of the traffic accident. An unaffiliated third party may also be able to verify any or all of the encrypted data streams 124a in the media container file 124 following the disclosed process and that other uses of the disclosed process are also possible. The systems and methods for verifying the one or more encrypted data streams 124a will be described in further detail herein.
[0101] In a particular implementation, the servers 108 each comprise a CPU 110, a non-transitory computer-readable memory 112, a non-volatile storage 114, an input / output interface 116, and graphical processing units (“GPU”) 118. The non- transitory computer-readable memory 112 comprises computer-executableinstructions stored thereon at runtime which, when executed by the CPU 110, configure the server to perform the above described processes of generating the media container file 124, providing access to one or more encrypted data streams 124a, and providing verifying the one or more encrypted data streams 124a. The nonvolatile storage 114 has stored on it computer-executable instructions that are loaded into the non-transitory computer-readable memory 112 at runtime. The non-transitory computer-readable memory 112 can also have stored thereon one or more machine learning models for performing object segmentation and representation generation. The input / output interface 116 allows the server to communicate with one or more external devices such the computer 104 (e.g. via network 106). The non-transitory computer-readable memory 112 may also comprise a secure repository 120 for the storage of the encryption keys, hash values, and signatures, as disclosed above. The secure repository 120 will be described in further detail herein. The GPU 118 may be used to control a display and may be used process the multimedia file 122 and to generate the individual data streams comprising the objects of interest as described above (e.g., using the machine learning models) and in further detail herein. In some aspects, the secure repository may be stored at a separate server. The computing components of the separate server are similar to those shown for servers 108. It would be appreciated that the CPU 110 and GPU 118 may be one or more processors or microprocessors, which are examples of suitable processing units, which may additional or alternatively comprise an artificial intelligence accelerator, programmable logic controller, a microcontroller (which comprises both a processing unit and a non-transitory computer readable medium), Al accelerator, neural processing unit (NPU), or system-on-a-chip (SoC). As an alternative to an implementation that relies on processor-executed computer program code, a hardware-based implementation may be used. For example, an application-specific integrated circuit (ASIC), field programmable gate array (FPGA), or other suitable type of hardware implementation may be used as an alternative to or to supplement an implementation that relies primarily on a processor executing computer program code stored on a computer medium.
[0102] FIG. 2A depicts generation of a multimedia container file for providing segmented access and verification of media content. A user 202 may upload amultimedia file 204 for processing by the systems and methods of the present disclosure. A plurality of data streams may be generated from the multimedia file 204, which may be viewed and managed during playback once the generation of the multimedia container file is complete. The multimedia file 204 may be a video file and may be in a format known to a person skilled in the art including but not limited to MP4, MOV, AVI, WebM or MKV. Note that while the term “data stream(s)’ is used herein, each data stream (or a plurality of data streams together) can be encoded, rendered, and streamed, for example over an access control server. Additionally, each data stream can comprise a representation of an object of interest, as described further herein. That is, each data stream can be considered to be an object or data layer, which can be encoded and packaged into a media container or rendered for streaming.
[0103] As depicted in FIG. 2A, the multimedia file 204 can comprise a header 206, metadata 208, as well as a video stream 210, an audio stream 212, and a subtitle or text stream 214 corresponding to the media content. The header 206 comprises information with regard to the multimedia file 204, such as format, size, and duration. The metadata 208 comprises information with regard to the media content in the multimedia file 204. The systems and methods of the present disclosure may support various means for receiving the multimedia file 204 and can employ secure data transfer protocols. Specifically, the supported communication may include: web interface communication, API calls, or a dedicated messaging system. The multimedia file 204 can be safeguarded by TLS / SSL encryption protocols and may be uploaded to a server or cloud infrastructure from a user device (e.g. from the device storage or cloud-based storage connected to the user device). It should be noted that the user may be required to be authenticated through conventional means prior to uploading the multimedia file 204. A confirmation that the upload of the multimedia file 205 is completed may be output if the multimedia file 204 is successfully received and free of errors and file corruption.
[0104] The systems and methods of the present disclosure may provide various means for objects of interest in the multimedia file 204 to be selected and / or identified (216). The user 202 may wish to generate a number of data streams basedon the objects of interest such that each data stream is based on a particular object of interest. The generated data streams may be individually managed with regard to access permissions. The user 202 may interact with the system through various methods to identify objects of interest including but not limited to natural language queries, brushstroke selection, automatic instance segmentation, or combinations thereof. For example, objects of interest can identified either by Al or user-guided selection . An example of natural language queries may allow the user202 to describe the objects of interest to be identified in plain language, such as providing prompts of "blue car" or "person with a backpack" to specify the objects of interest. In some embodiments, deep learning models may be utilized to process the queries and identify corresponding objects of interest within the multimedia content 204. An example of brushstroke selection may allow the user 202 to directly select objects such as the face or body of a person or vehicle within the video frames of the multimedia file 204 (e.g. in the video stream 212) through a user interface using a brushstroke tool, which can provide a more intuitive and precise method for the identification of objects of interest. An example of automatic instance segmentation may include the use of automatic instance segmentation algorithms to identify and segment all discernible objects of interest within the video frames of the multimedia file 204. For objects of interest in the audio stream 212, separate sources can be automatically identified and separated. This option may be useful for scenarios where comprehensive objects of interest identification is desired, such as for audio separation where all audio sources are automatically divided into different tracks, allowing the user 202 to select which tracks should be encrypted or signed, as described further herein.
[0105] In accordance with the present disclosure, the selected / identified objects of interest may be segmented or extracted (218), for example, by separating it from the background. That is, the multimedia file 204 may be segmented using a media segmentation module, for example. Deep learning models for object identification and separation may be used to separate / segment objects of interest in both the video stream 210 and the audio stream 212. The deep learning models can be trained on extensive datasets containing image-text-audio data, to enable accurate recognition and association of visual and auditory concepts with textual descriptions(e.g. the concepts / descriptors input during objects of interest selection). Accordingly, the present disclosure can accurately identify and locate objects within the video based on the user's queries. Models, including but not limited to: CLIP (Contrastive Language-Image Pre-training), SAM (Segment-Anything-Model), or a combination thereof can be employed for detection and segmentation of objects of interest in the video stream 210. CLIP may refer to a neural network architecture that has been trained on a dataset of image-text pairs for the understanding and association of visual concepts with their corresponding textual descriptions. In some aspects, binary masks may be generated in which pixels corresponding to the object of interest can be marked as “foreground” (e.g. with a value of 1 ) and the remaining pixels can be marked as “background” (e.g. with a value of 0). For example, a mask may be applied or associated with each frame of the video stream 210. Techniques including but not limited to source separation, speaker diarization models, or a combination thereof can be used to identify audio-based objects of interest (e.g. distinct sounds and speakers) within the audio stream 212. Source separation models may be used to identify distinct sounds (e.g. from the object of interest) and output estimates of the individual source sounds. Speaker diarization models can segment the audio or the audio stream 212 into speaker-homogeneous regions and label each segment with a speaker (e.g. as an object of interest) identity. A data stream is generated for each of the selected objects of interest.
[0106] In accordance with the present disclosure, the objects of interest may be tracked (220) once identified. Object-tracking algorithms may be used to follow the movement of visual-based objects of interest throughout the video sequence (e.g. the video stream 210), which can generate a number of individual “tracks” corresponding to the number of objects of interest. Various tracking algorithms can be utilized, including but not limited to: Kalman filter, optical flow, deep learning trackers, or a combination thereof. Kalman filters can predict the future position of objects of interest based on their past movement and current measurements, which can provide robust tracking even in cases of temporary occlusion or noise. Optical flow can analyze the movement of pixels between consecutive frames to estimate the motion of objects of interest, which can enable accurate tracking of complex movements. Newly developed deep learning trackers / models can learn complex object appearances andmotion patterns for objects of interest, which can provide highly accurate and robust tracking performance. Further, masks may be generated for each track for the duration of the track.
[0107] It should be noted that tracking is performed differently for audio stream212, which includes audio-based objects of interest that cannot be tracked using movement. Due to the dynamic nature of sound and the potential overlap of multiple sound sources, audio-based objects of interest may be tracked with a focus on maintaining the association between identified sound sources or speakers and their corresponding audio segments throughout the audio stream. Various techniques can be used to track audio-based objects of interest in the audio stream 212, including but not limited to: audio source separation, speaker diarization, audio finger printing, feature matching, or a combination thereof. Audio source separation algorithms may be applied to isolate and extract individual sound sources (e.g. objects of interest) as they evolve or change over time. This can allow the creation of separate audio streams for each identified sound for independent access control and analysis, as described further herein. Speaker diarization techniques may be used for audio stream 212 that contains speech. Speaker diarization may be employed to segment the audio stream 212 based on speaker (e.g. object of interest) identity. As speakers take turns or speak concurrently, changes in speaker activity can be tracked and segmentation can be adjusted accordingly, which can ensure that the utterances of each speaker (e.g. with speech from each speaker being an object of interest) are accurately associated with their respective audio stream. Matching and correlation techniques, such as audio fingerprinting or feature matching, can be used to track specific sounds or audio events, considered as objects of interest, across the audio stream 212, which can allow the association between an identified sound (e.g. object of interest) and its corresponding audio segments to be maintained even if the sound's (e.g. object of interest) characteristics change slightly over time or if the sound is interrupted by other sounds.
[0108] The application of various techniques for object of interest tracking can ensure that access control and signature verification are applied consistently and accurately (as described further herein) to the relevant audio objects of interestthroughout the entire audio stream 212, regardless of their temporal dynamics or potential overlap with other sound sources. Further, the specific choice of object of interest identification and tracking methods can be adapted and customized based on the application requirements and the type of multimedia file 204 being processed to ensure flexibility and optimal performance.
[0109] Note that a plurality of different data representations can be generated for each object of interest (221 ), for example as a plurality of data streams, as described further herein.
[0110] In accordance with the present disclosure, a separate video or audio stream may be generated for each identified object of interest (222), containing only the information related to the respective object of interest. For example, the tracking of the respective object of interest can enable the creation of separate data streams, each focusing on a unique object of interest. As an example, if the video stream 210 contains two different blue cars, two distinct streams can be generated, one for each car, thereby enabling fine-grained access control at the object level. That is, the video segmentation process results in the creation of multiple discrete individual data streams, with each stream representing a specific object of interest. This stream generation and encoding process is described herein with reference to FIG. 2B, which depicts the process for generating a plurality of streams corresponding to identified objects of interest.
[0111] In some embodiments, a plurality of data streams can be generated for each (visual) object of interest. That is, for each visual object of interest, a plurality of data streams having a different representation of the object of interest may be generated. As used herein, an "object of interest" can refer to a conceptually distinct entity within multimedia content deemed relevant for separate identification, tracking, and management, such as a specific person, vehicle, spoken phrase, or on-screen textual element. Each set of coded data (e.g., pixels, audio samples, text segments, or associated spatial / temporal metadata like segmentation masks) corresponding to such an identified "object of interest" within the multimedia bitstream or container can be referred to as an "Object Layer." Each "object layer" can be uniquely identified (e.g., by an object ID: “ObjectlD_X”) and can be the primary unit for which distinctsemantic representations are generated and access control is applied. Note that each object layer can correspond to a particular data stream (e.g. a video stream). The term "ObjectX" as used herein, corresponds to a generic placeholder to refer to an exemplary instance of an "Object of Interest".
[0112] As depicted in FIG. 2B, the objects of interest may be isolated (250). For visual-based objects of interest in the video stream 210, the relevant pixel data corresponding to the objects of interest from each video frame of the video stream 210 may be extracted. For example, the objects of interest can be isolated by applying the generated mask (described previously), utilizing the alpha channel information, identifying pixels based on the chosen color key filling, or combinations thereof. For audio-based objects of interest in the audio stream 212, the corresponding audio segments may be isolated based on the output of audio source separation and speaker diarization models, which can ensure that only the relevant audio data associated with the identified sound or speaker (e.g. identified as objects of interest) is included in the stream to be generated. The original audio track (e.g. audio stream 211 ) may be preserved and securely stored alongside the separated tracks, which can ensure that no loss of audio information takes place. That is, elements that are not the object of interest can be removed, redacted, or censored from the streams to prevent information with regard to these elements from becoming available in association with the object of interest.
[0113] As depicted in FIG. 2B, empty space handling may be performed (252). For the video stream 210, techniques including but not limited to: mask generation, alpha channel encoding, identifying color key filling, or combinations thereof may be performed. Mask generation can refer to the creation of a binary mask for each video frame in the video stream 210 to generate a data stream, as described previously. The generated masks may be stored in separate video streams that are generated, with each frame of the mask stream representing the corresponding mask image for the video frame. This can allow for playback options that involve displaying only the masks without the original video content. The mask streams can be associated with a corresponding video data stream to isolate the object of interest, if needed. An alpha channel can be added to each video frame of the video stream 210 to generate a datastream, representing the transparency of each pixel. Pixels belonging to the objects of interest can have an alpha value of 255 (e.g. fully opaque), while empty spaces can have an alpha value of 0 (e.g. fully transparent). This technique can allow for smooth compositing of the object of interest onto any background. The empty space surrounding the object can be filled with a specific color chosen as the “color key”. This color key can be used during playback to identify and replace the empty areas with the desired background or transparency.
[0114] For the audio stream 212, techniques including but not limited to: silence detection and removal, audio segmentation and timestamping, or combinations thereof may be performed. It should be noted that some audio streams comprise objects of interest that produce sounds intermittently. As such, periods of silence or background noise may be identified and removed to reduce the overall size of the generated audio stream and optimizes storage and transmission efficiency. Various audio processing techniques which can be employed for silence detection includes but is not limited to: energy thresholding, where silence / noise can be detected based on the audio signal's (e.g. of the object of interest) energy / signal falling below a predefined threshold; zero-crossing rate analysis, where silence / noise can be identified by analyzing the rate at which the audio signal (e.g. of the object of interest) crosses the zero amplitude level; spectral analysis, where silence / noise can be detected by analyzing the frequency content of the audio signal (e.g. of the object of interest) and identifying segments (in the audio stream 212) with minimal spectral energy; or combinations thereof. The audio stream 212 may also be segmented into distinct segments based on the presence or absence of sound of one or more objects of interest. Timestamps may be associated with each segment to indicate the start and end times of the corresponding audio object of interest. The timestamps may be stored within the metadata 208 (of the multimedia file 204 or the individual data stream) or the header 206 and can be used to synchronize the audio playback with the corresponding video and mask streams, as described further herein.
[0115] In some embodiments, The synchronization metadata can comprise, for each data stream or object of interest:1. Frame numbers, sample indices, or other discrete-time markers.2. Decode timestamps (DTS) and presentation timestamps (PTS) for each frame or sample, ensuring precise alignment of all visual, audio, textual, and sensorbased representations linked to the object of interest.3. The object of interest pertaining to the data stream.4. Timestamped Spatial Extent Time Series data, for example sequences of bounding boxes, segmentation masks, or 3-D coordinates which are aligned with the PTS / DTS of the corresponding visual frames for accurate spatial overlay and interaction analysis.5. Object Appearance Timestamps or Temporal Extent data, for example indicating start and end PTS / DTS for an object’s presence or for specific events / states in its timeline.Accordingly, a playback client can uses this metadata to align and composite authorised, selected representations, ensuring a coherent and temporally accurate experience
[0116] By performing the above described processes, objects of interests within the multimedia file 204 can be isolated and tracked to generate the individual data streams. As depicted in FIG. 2B, the data streams may be encoded (254). In particular, individual video (data) streams are generated for visual-based objects of interest from the video stream 210 and individual audio (data) streams are generated for audio-based objects of interest from the audio stream 212. In some embodiments, the individual video streams comprising isolated video data, along with the chosen empty space representation (mask, alpha channel, or color key), can be encoded using any suitable video codec, for example H.264, HEVC or, MKV, which can compress the data for efficient storage and transmission. In some embodiments, the individual audio streams comprising isolated audio segments can be encoded using any appropriate audio codec, for example AAC, WAV, or MP3, which can reduce the data size while preserving audio quality. Variable bit rate encoding can also be utilized for the encoding of individual audio streams. In some embodiments, the mask images / streams used to represent the empty space around each object can be encoded using various efficient methods and approaches to minimize storagerequirements and transmission overhead including but not limited to: lossless image compression, lossless image formats (e.g. PNG or GIF with lossless compression and wide compatibility with various systems), video codecs with intra-frame coding, specialized mask encoding techniques (e.g. specialized encoding techniques for binary images), or combinations thereof. For example, techniques like Run-Length Encoding (RLE) or bit-plane encoding can be used to compress the mask data without any loss of information to help ensure accurate reconstruction of object shapes (of objects of interest) during playback. It should also be noted that video codecs like H.264 or HEVC, which can be used for video stream encoding, can also be utilized in intra-frame coding mode to compress individual mask images efficiently. The described encoding process may also be adapted based on factors such as desired compression ratio, computational complexity, and compatibility with playback systems. For embodiments, where different representations are generated for each object of interest, all representations can be encoded as individual data streams.
[0117] As depicted in FIG. 2B, the generated and / or encoded data streams may be synchronized (256). This process can ensure that individual data streams are synchronized during playback, regardless of the media type (e.g. audio, video, text, or images). In some aspects, all of the data streams may begin playback at the same time. If one of the video data streams is missing content in certain frames due to the isolation of the object of interest (e.g. masks used to cover the segments of the stream), empty frames can be inserted to maintain consistent timing across all of the video data streams. Similarly, silent segments can be added to audio tracks if one of the audio data streams includes specific sounds or speakers that have been removed. This process can ensure basic synchronization and is most effective when all streams have matching frame rates. In some aspects, metadata containing timing information and object details for each of the individual video and audio data stream may be generated. The generated metadata can be embedded in the header 206 or a separate associated file. Examples of generated metadata can be one or more of: frame numbers for identifying the sequence of frames within each individual data stream; decode timestamps (DTS) for specifying the precise timing for decoding each frame in each individual data stream; presentation timestamps (PTS) for indicating the timing for presenting each frame during playback; object IDs for identifying eachobject of interest within the individual data streams; spatial information for providing details about the location of the corresponding object of interest within each frame (e.g. bounding box coordinates); and object appearance / disappearance timestamps for marking the specific timeframes when the corresponding object of interest appear or disappear within the individual data streams. In some embodiments, a playback engine may be used to utilize the generated metadata to align and synchronize the streams accurately. In particular, techniques such as buffering (e.g. temporarily storing data to compensate for variations in network speed or processing time); timecode matching (e.g. aligning the individual data streams based on their respective timestamps); playback speed adjustment (e.g. fine-tuning the playback speed of individual dada streams to ensure they remain synchronized); or a combination thereof may be used. Further details with regard to the generated metadata is described further herein.
[0118] It should be noted that the synchronization approaches may be adapted depending on the media type. For example, for the synchronization of video data streams, frame numbers, DTS, and PTS may be preferred to ensure smooth and continuous playback of video content. Similarly, for the synchronization of audio data streams, aligning waveforms, adjusting playback speed to match durations, or synchronizing based on timestamps may be preferred to ensure that sounds and speech segments are presented in the correct order and timing. Further, text or images can be associated with specific video frames or timestamps in the data stream and the presentation timing can be adjusted based on the overall flow and context of the multimedia content. These synchronization approaches can help ensure a seamless and coherent multimedia presentation, even when dealing with multiple data streams of different media types and varying durations.
[0119] As depicted in FIG. 2B, consideration may be also given to error handling (258). This may be done, for example, to address potential uncertainties or inaccuracies in object identification, tracking, or segmentation. In some embodiments, confidence scores can be used, in which deep learning models and tracking algorithms may be configured to provide confidence scores associated with their outputs. These scores can be used to identify and flag potential errors, allowing forfurther review or manual correction. In some aspects, user feedback can be used, in which the system can allow users (e.g. the user 202) to provide feedback on the accuracy of the identification and segmentation of their selected objects of interest. This feedback can be used to improve the performance of the deep learning models and refine the generated masks. In some embodiments, manual correction may be used, in which users can be provided with tools to manually adjust boundaries of objects of interest, correct misidentified objects of interest, or refine segmentation masks. This can help ensure that the final output (i.e. the generated individual data streams) accurately reflects the intended access control and redaction requirements. In some embodiments, redundancy and fallback mechanisms may be used, in which redundant identification or tracking methods can be employed to increase robustness and mitigate the impact of errors from any single method. Additionally, fallback mechanisms can be implemented to ensure that access control and privacy protection are maintained even in cases of significant errors or uncertainties. Finally, at 260, the corresponding data stream(s) are output, which can be then encrypted.
[0120] As described above, a primary data stream can be generated for an identified object of interest (e.g., "ObjectX," such as Person A or Vehicle B). In some embodiments, a plurality of distinct data stream comprising unique representations for each uniquely identified object of interest can be generated. This multirepresentation approach can address varied use cases, user authorizations, privacy constraints, and resource limitations that necessitate different views or data forms for the same underlying object. Instead of a monolithic stream, the system can generate a structured set of alternative or complementary data streams, each encapsulating a specific representation of an object of interest.
[0121] Note that each data stream, including each distinct data representation / data stream for a particular object of interest can be independently manageable and associated with the primary object of interest. Each object of interest can be identified by an object identifier or ID. As described herein, each distinct representation of a particular object of interest (e.g., each of a plurality of data streams for an object of interest) can be encrypted with its own unique cryptographic key for granular, context- aware access control. The generation of these multiple representations can occurconcurrently with, or subsequent to, initial object identification and tracking. Specific representations generated for an object of interest may be determined by system configuration, content-owner policies, the object's nature, or real-time analysis of its context and sensitivity.
[0122] Example data streams generated for an object of interest is described below. Each data stream contains a representation of the corresponding object of interest. Note that multiple types of data streams can be generated for each object of interest. The user may choose the type(s) of data streams to be generated for each / all of the objects of interest. Alternatively, all possible data streams may be generated.High Quality Representation
[0123] This type of data stream can preserve the maximum possible visual fidelity of the object of interest as captured in the original multimedia. That is, the representation of the object of interest for this type of data stream can be substantially the same as it appears in the original multimedia, for example having substantially the same resolution and / or bit rate.
[0124] For a visual object of interest, this type of data stream can be generated by segmenting pixels corresponding to the object of interest from video frames, as described above. For example, Al-driven semantic segmentation can be used to create precise masks. Pixels within these masks can be extracted. Minimal or high- quality compression (e.g., intra-frame modes of codecs like H.265 / HEVC or AV1 with high bitrate settings) can then be applied to the segmented pixel data. The background outside the object mask may be rendered transparent (e.g., via an alpha channel), filled with a neutral color, or retained if the stream is a full-frame crop around the object of interest. Note that the audio corresponding to the object of interest, for example spatially or semantically linked original audio (e.g., speech from a person object) can also be segmented, included, and compressed using a high-fidelity audio codec (e.g., AAC, Opus) and associated with the generated data stream (e.g., via the object ID).
[0125] This type of data stream can be valuable for users having complete authorization of the multimedia content, for example requiring complete, unaltered visual / audio detail for primary evidence, detailed analysis, or archival.De-identified Representation
[0126] This type of data stream can generate a representation of the (visual) object of interest that is visually modified to obscure or replace personally identifiable information (PH) or other sensitive features, while attempting to retain contextual relevance and naturalness, offering a more advanced solution than simple blurring or pixelation.
[0127] One or more Al models can first detect specific sensitive features within a segmented object of interest or regions therein (e.g., faces via face detectors, license plates via ANPR pre-processing, logos / text via OCR / object detectors). Subsequently, the one or more Al models can modify the object of the interest visually for the data stream by generating an modified representation thereof (within the data stream). For example, human faces can be altered using techniques including Generative Adversarial Networks (GANs) trained for face anonymization (e.g., anonymization-focused GANs or face attribute swapping techniques) to replace original faces with synthetic, non-existent ones, or subtly alter key facial features to prevent identification while maintaining a human-like appearance. The transformation level can be a parameter. As another example, license plates can be modified via simple blurring or Al inpainting to replace the plate characters with plausible but incorrect ones (e.g., alternate numbers), or to blend the plate area with surrounding vehicle texture. As another example, for other PH such as names on badges, Optical Character Recognition (OCR) can be used to identify the sensitive text, followed by targeted masking or replacement of the identified text with generic text.
[0128] As such, pixel data for the object of interest or portions thereof can be modified for this type of data stream, which aims to make re-identification from this type of data streams alone extremely difficult or impossible. This type of data stream may be provided to users for or requiring privacy protection (e.g., GDPR, CCPAcompliance), particular where the data stream is shared with users lacking explicit consent for identifiable features, or for public release of certain footage.
[0129] In some embodiments, the de-identified data streams are required to be demonstrably resistant to automated re-identification according to a policy-defined numerical threshold (e.g., set by the user or relevant regulatory requirements). Note that various transformation techniques (e.g., GAN, diffusion, classical blurring) are possible for generating this type of data streams.
[0130] As an example, a set of privacy policies may be applied for the generation of the de-identified data streams. A policy-driven privacy envelope can include profile configuration parameters according to target privacy metrics. The configuration can include (1 ) a metric for a generic privacy meter, (2) a metric for a maximum allowed probability of human object of interest (e.g., a face) verifier matching a de-identified person to its original identity at a set False Positive Rate (FPR), and (3) a metric for a maximum allowed character-recognition accuracy on an anonymized text-based object of interest at a stated FPR. Note that policy can reference any benchmark dataset and verifier, as documented within the policy, and may be set for each data stream.
[0131] As another example, a default set of thresholds may be used according to regional requirements, where the user or the system can select one or more of the policies the data stream should comply with. Specifically, EU GDPR can require metric (2) to be < 1 % FPR = 1 x 10"3and metric (3) to be < 2% @ FPR = 1 x 10"4; US HI PAA can require metric (2) to be < 5% @ FPR = 1 x 10"3and metric (3) to be < 5% @ FPR = 1 x 10"3; and Canadian PIPEDA can require metric (2) to be < 2% @ FPR = 1 x 10"3and metric (3) to be < 3% @ FPR = 1 x 10"4
[0132] The Al models for generating this type of data stream may include specific face recognition (e.g., ArcFace-ResNet100) and plate recognition (e.g., OpenALPR-YOLOv5) models trained on standard datasets (e.g., Celeb-1 M / LFW, public ANPR corpus).
[0133] In some embodiments, the model(s) may be retrained, updated or substituted at regular intervals (e.g. quarterly), where the model benchmarks can bere-run after every model retraining to ensure model compliance. Each data stream can be generated with a signed tuple corresponding to the information of the model. The model information (e.g., model version, benchmark name / type, threshold, measured value, evaluation date, etc.) can be stored in the media container (e.g., as metadata) and incorporated as a Merkle tree leaf (described further herein). Benchmarking can be done to ensure that the data stream is compliant, for example, if a measured value exceeds its threshold, the data stream is flagged, and packaging of the data stream can be halted until a conformant representation is produced.
[0134] In some embodiments, users can use an internal validation corpus for private reporting; note that only standard benchmark results may be embedded for cross-deployment comparability.Low Resolution Representation
[0135] This type of data stream can provide a reduced data footprint representation of the object of interest by sacrificing detail for computational efficiency. That is, the low resolution representation can be a lower resolution or bit rate data stream in comparison to the original data stream.
[0136] This type of data stream can be generated from the High quality representation data streams by methods such as:Spatial Downsampling: Reducing pixel dimensions of the data stream (e.g., from a 1080p object crop to 360p).Temporal Downsampling: Reducing frame rate of the data stream (e.g., from 30fps to 10fps).Aggressive Compression: Re-encoding the data stream at a much lower bitrate using standard video codecs.Thumbnail / lcon Generation: Creating a single representative static image or a short animated icon of the object for use in the data stream.
[0137] This type of data stream can be provided to users requiring low- bandwidth streaming, for quick III previews, overview timeline generation, or for users with highly restricted access needing only a coarse indication of object presence and movement.Text Description Representation
[0138] This type of data stream can replace the object of interest with textual data describing the object o interest, its attributes, actions, or associated speech.
[0139] For visual objects of interest, the data stream can be generated using Al-driven video / image captioning models (e.g., Transformer-based or similar architectures) to analyze the unaltered data stream of corresponding to the object of interest (e.g., high quality representation or a pre-processed version) to generate descriptive sentences (e.g., "Person A wearing red jacket walks towards door"). Further, object-based Al detection models can provide attributes (e.g., color, type) of the object of interest, which are structured into text.
[0140] From audio objects of interests, or if the object of interest has associated audio, Automatic Speech Recognition (ASR) engines (e.g., large language modelbased ASR) can be used to convert spoken words into a time-stamped transcript and used to generate a data stream or by including the transcript in the data stream. This transcript may be further processed for speaker diarization if multiple speakers are associated with the object or its vicinity. Note that the generated text can be formatted with timestamps (e.g., WebVTT, SRT, custom XML / JSON) for synchronization with other representations.
[0141] This type of data stream can be provided to user with accessibility needs(e.g., for visually impaired users), for quick summarization, object-level keyword searchability, or providing context when visual access is restricted.Sensor Data Representation
[0142] Some multimedia can include synchronized data from non-visual sensors (e.g., Lidar, Radar, Thermal). This type of data stream can utilizerepresentations which capture sensor data specifically pertaining to the object of interest.
[0143] To segment the object of interest or data corresponding thereto, the subset of sensor readings (e.g., Lidar point cloud cluster, radar signature, thermal blob) spatially and temporally corresponding to the visually identified data of interest can be identified then segmented. Example segmentation methods include:Sensor Fusion Pre-processing: Calibrating and aligning coordinate systems of different sensors (e.g., camera and Lidar);Cross-Modal Association: Using (visual) object of interest's 2D bounding box or segmentation mask from the visual stream to project its spatial extent into the 3D Lidar space (or other sensor space) at the corresponding timestamp; andExtracting sensor data (e.g., Lidar points) falling within this projected 3D volume.
[0144] The segmented sensor data can be stored / generated as a time- synchronized stream (e.g., a sequence of Lidar point clouds, time-series radar signatures).
[0145] This type of data stream can be useful to provide complementary information about the object of interest’s physical properties, spatial presence, or state not apparent from visual data alone (e.g., precise Lidar distance / shape, thermal temperature, radar velocity). This type of data stream can be particularly useful for autonomous systems, forensics, and detailed environmental analysis.Feature Embedding Representation
[0146] This type of data stream can replace the object of interest (visual or audio) with a highly compact, numerical vector representation of the object of interest. This type of data stream can be useful in capturing the essential discriminative characteristics of the object of interest without storing raw visual or audio data.
[0147] To generate data streams for visual data of objects, the segmented visual data of the object of interest (e.g., from a high quality representation) can beprocessed by a pre-trained deep learning model (e.g., a Convolutional Neural Network or Vision Transformer) where the final vector output is used for stream generation. For example, activation values from an intermediate or final layer (pre-classification) can be extracted as the feature vector (embedding). For tasks such as face recognition, specialized models (e.g., ArcFace, FaceNet) can generate highly discriminative facial embeddings.
[0148] To generate data streams for audio data of objects, audio embeddings can be generated as the output of models (e.g., VGGish, TRILL) or custom networks trained on audio tasks processing the data of object, capturing characteristics of sound events, speaker identity, or speech content.
[0149] These vector / embedding representations can be of a fixed dimensionality (e.g., 128, 256, 512 dimensions) and stored as a time-series in the data stream.
[0150] This type of data stream can enable privacy-preserving analytics. In particular, these data streams can be used in operations including object reidentification across video segments, similarity searching (e.g., "find similar objects"), or anomaly detection can be performed on these embeddings without accessing potentially sensitive raw pixel / audio data, which can be crucial for large-scale analysis while respecting privacy.
[0151] Various advantages are possible by having multiple representations for objects of interest. For example, semantic linkage and co-management of these distinct representations and data streams under a single object identifier corresponding to a particular object of interest is possible. Each representation can be explicitly defined as a particular facet or version of the same underlying object of interest. Further, each data stream or each type of representation can be provided with its only type of encryption as described below, which allow for highly granular access control. For example, a user may be authorized to view all data representations or data streams associated with an object of interest. Alternatively, the user may only be authorized to view a certain type of representation for the object of interest, for example the user may be allowed to view de-identified representationand the textual description representation but denied access to high quality representation and sensor data representation. These representations can be managed within a single, coherent framework, with their availability and characteristics detailed in a Representation Table as described further below. Notably, their integrity, along with that of associated metadata, can be protected by a hierarchical integrity signature, also described further below.
[0152] Referring back to FIG. 2A, the generated individual data streams are encrypted (224). In particular, a stream encryption and key management module may be employed to encrypt the generated individual data streams and manage the encryption keys (described further herein). The stream encryption and key management module may interact with a secure repository 226 as well as one or more access control modules to ensure confidentiality and controlled access to multimedia content. In accordance with the present disclosure, a unique encryption key 230 is generated for each of the generated individual data streams containing the respective object of interest. The encryption key 230 may be a symmetric encryption key and may be of default or user-defined length. By generating an encryption key 230 for each of the individual data streams, if one of the encryption keys is compromised, the security of other individual data streams are not affected, which can enhance overall data protection. Further, separate encryption keys can be generated for different time segments within a stream. That is, it is possible to generate one or more time-specific encryption keys, each of which corresponding to a particular time segment of one of the individual data streams, which may or may not overlap. This feature can allow for additional granular access control and time-based permissions in that the user 202 can control access to a particular timed portion of one or more individual data streams. In accordance with the present disclosure, each of the generated individual data streams is encrypted using the corresponding encryption key 230 that was generated. That is, all of the individual data streams including video, audio, text, and mask streams are encrypted. In particular, the encryption algorithm may be a symmetric encryption algorithm such as Advanced Encryption Standard (AES) with a key length of 256 bits. In some aspects, the user may be able to choose the algorithm for encryption and / or the key length on a provided user interface. Further, the encryptionalgorithm and key lengths may be adjusted based on specific security requirements and compliance standards.
[0153] In some embodiments, for time segment based encryption, when a new time-specific encryption key, K_E(t0, tx) , is introduced for a segment [t0, tx) of an object of interest, the first encrypted sample (e.g., video frame, audio packet) in that interval can be a Self-Decodable Access Point (SDAP). An SDAP is a frame or sample decodable without reference to any earlier frames / samples potentially encrypted with a different key. The specific type of frame or sample constituting an SDAP depends on the multimedia codec used for the representation, as detailed below:AVC I H.264: Instantaneous Decoder Refresh (IDR) picture,HEVC I H.265: Clean Random Access (CRA) picture, IDR picture,AV1 : Keyframe (specifically, a non-delta keyframe),VVC I H.266: IDR_W_RADL picture, IDR_N_LP picture, CRA_NUT picture,Opus (audio): Any independently decodable Opus packet, andAAC (audio): Any AAC frame not relying on prior frame temporal prediction (e.g., all frames).
[0154] The encryption keys 230 may be stored in a secure repository 226, as depicted in FIG. 2A. The secure repository 226 may be managed by the system of the present disclosure and can provide a protected and isolated environment for key management to reduce the risk of unauthorized access or key leakage. An access control module or similar process may be implemented to interact with the secure repository 226 to manage and control the retrieval and distribution of the decryption keys 230 to authorized users based on user instructions and / or pre-set policies. This can ensure that only authorized individuals can access the content of the individual data streams. In some embodiments, the secure repository 226 is implemented as a Key Management Service (KMS), which manages access to the decryption keys 230. In some aspects, the secure repository 226 is implemented as a secure key-value databases. These databases can be implemented with various access controlmechanisms and encryption features and can offer a software-based option for secure storage. In some aspects, the secure repository 226 is implemented as hardware security modules (HSMs). The HSMs can provide protection against physical and logical attacks and can be suitable for high-security environments. In some aspects, the secure repository 226 is implemented as a blockchain-based storage, which can provide distributed and tamper-proof storage to enhance transparency and security in key management. It should be noted that the secure repository 226 can securely store encryption keys 230 to prevent unauthorized access and ensuring data confidentiality, enforce access control policies (e.g. default policies or set by the user 202) by only releasing decryption keys 230 to authorized users or systems, and facilitate key management operations such as key generation, rotation, and revocation to help ensure the long-term security and integrity of the encryption keys 230. In particular, the secure repository 222 can store the specific content-encryption keys as the decryption keys 230, which are each linked or associated with a corresponding encryption key representation (e.g., encryption key ID), stored in the metadata of the data stream.
[0155] In some embodiments, the key management service is configured for secure storage, management, and delivery of representation-specific contentencryption keys (CEKs) or master keys from which CEKs are derived. It can utilize a highly secure repository for sensitive key material. This secure repository is configured for robust protection against unauthorized access and tampering. Implementations can comprise the below aspects.1. A hardened key-value database for cryptographic key storage, with strong access controls based on authenticated KMS processes, encryption-at-rest using master KMS keys, and immutable audit logging of key management operations.2. One or more Hardware Security Modules (HSMs) compliant with industry standards (e.g., FIPS 140-2 Level 3 or higher), used to protect KMS master keys and perform critical cryptographic operations like CEK generation, CEK wrapping for secure client delivery, and signing of audit or revocation data within a tamper-resistant boundary.3. A blockchain-based distributed ledger for managing key-associated metadata, such as encryption key representations and IDs (KI Ds), key revocation status, authorization policies, or key lifecycle audit trails. This can enhance tamperevidence and transparency for key metadata management, while sensitive cryptographic key material itself may remain off-chain within HSMs or encrypted databases.The KMS can ensure the confidentiality, integrity, and availability of cryptographic keys stored therein.
[0156] In accordance with the present disclosure, the encryption keys 230 may be formatted prior to storage. For example, the encryption keys 230 may be formatted according to an encryption standard determined by the user 202 or a default encryption standard. The encryption keys 230 may be also organized and identified with relevant metadata, such as the corresponding object ID and stream type (e.g. video, audio, mask, text). By associating relevant metadata to each of the encryption keys 230, the encryption keys 230 can be quickly associated with the correct multimedia file 204, data stream, object of interest in the data stream, and user 202.
[0157] In accordance with the present disclosure, the encryption keys 230 may be encrypted. The encryption process may be asymmetric. In some aspects, each encryption key for each data stream is individually encrypted using a public key of the user 202. The asymmetric encryption process can ensure that only the user 202, possessing the corresponding private key, can decrypt the encryption keys 230 and therefore access the media content in the encrypted data streams unless the encryption keys 230 are shared by the user 202. The asymmetric encryption algorithm for encrypting the encryption keys 230 may be RSA (Rivest-Shamir-Adleman) or ECC (Elliptic Curve Cryptography).
[0158] It should be noted that the encrypted encryption keys may also be stored in the secure repository 226. In particular, the user 202 may be provided with the option to store the encrypted encryption keys in the secure repository 226. For example, if the user 202 would like other users to access one or more of the encrypted data streams, they may wish to store the encrypted encryption keys in the securerepository 226. The encrypted encryption keys may be stored with the corresponding metadata, as described previously. Other users may be able to request access to one or more encrypted data streams, for example, via the access control module. In that case, if the request is authorized, the stored encryption keys may be retrieved and distributed to the other users for use in accessing the authorized data streams, for example, via the access control module. In some embodiments, no encryption keys may be stored in the secure repository 226. For example, the user 202 may wish to retain sole control over the encrypted encryption keys and therefore maintain sole access over all of the encrypted data streams. Accordingly, the encrypted encryption keys may be written into the header 206 instead such that other users would not have access to the encryption keys, for example, through the access control module. The encrypted encryption keys may be written into the header 206 with the corresponding metadata, as described previously. The encrypted encryption keys and / or the corresponding metadata may also be written into the header 206 and stored in the secure repository 226.
[0159] In accordance with the present disclosure, generation and storage of digital signatures and / or corresponding hash values may be implemented to improve the integrity and authenticity of the multimedia content. In some aspects, a hash value 228 (e.g. string of characters) may be generated for each of the individual data streams (e.g. all of the video, audio, text, and mask streams). Specifically, a unique cryptographic hash value 228 may be computed for each generated individual data stream using a cryptographic hash function. The cryptographic hash function may be a secure hash function such as SHA-256. The hash values 228 may be useful in verifying the authenticity of the media content (e.g. the individual data streams). For example, when accessed, a user can compute a current hash value of the accessed media content (e.g. the individual data streams), the current hash value can be compared to the previously generated hash value (e.g. generated at the time of data stream generation) to confirm that the hash value is unchanged. If the hash value remains unchanged, it serves as proof that the media content has not been altered since the previous hash value was generated. In some embodiments, a hash value (i.e. an overall hash value) can be generated for the entirety of the media content or the overall media content (e.g. multimedia file 204). In some embodiments, a hashvalue (i.e. an overall hash value) may also be generated for all of the individual data streams and associated metadata. The generation of a hash value for all of the media content can ensure that the authenticity of the entirety of the media content can be validated. Any / all of above described hash values 228 may be stored in the secure repository 226 for security purposes. Further, by storing the hash values 228 in the secure repository 226, other users may be able to request access to the hash values 228 for use in verifying the authenticity of one or more data streams, for example, via the access control module. The stored hash values 228 may be provided to the other users if authorized, for example, via the access control module, for use in comparison with hash values that are subsequently generated.
[0160] Further, digital signatures may also be provided to ensure data authenticity. In some embodiments, the user 202 may digitally sign (e.g. provide digital signatures for) each of the individual data streams and / or overall media content, for example, by signing each of the generated hash values 228 and / or the hash value of the entirety of the media content. The digital signatures may be provided using a private key of the user 202. A digital signature algorithm including but not limited to RSA and ECDSA may be used to generate the digital signature, which can bind or associate the hash value to the owner's identity to help ensure authenticity. The digital signatures may also be stored in the secure repository 226, for example, in association with the corresponding hash values 228. It should be noted that the digital signatures and their corresponding hash values 228 may be written into the header 206 such the validation and verification of the media content may be performed. It should be noted that by writing the signatures and their corresponding hash values 228 into the header 206, verification can be performed regardless of whether the encryption keys 230 are stored in the secure repository 226 or not.
[0161] In accordance with the present disclosure, a media container file 232 (e.g. a multimedia container file) is generated. In particular, the encrypted data streams corresponding to the objects of interest, the encrypted encryption keys, the digital signatures, and the hash values 228 may be packaged into the media container file 232. The media container file 232 may be useful in improving the portability, organization, and secure distribution of the multimedia content while maintainingcompatibility and data integrity. As shown in FIG. 2A, the media containerfile 232 may comprise a header 206 and metadata 208. The header 206 and metadata 208 may remain unchanged from those of the multimedia file 204 or may be updated to reflect the information and contents of the media container file 232. For example, the header 206 may be updated to include the encrypted encryption keys, the digital signatures, and / or the hash values (238). A new header 206 may be generated, the header 206 comprising information including: format version, file size, duration, the list of included data streams, or a combination thereof. In some embodiments, one or more of the encrypted encryption keys, the digital signatures, and the hash values may be embedded in the metadata section of the media container file 232. The embedding process may be performed using standard cryptographic protocols. For example, AES-256 may be used for symmetric encryption and RSA or ECDSA may be used asymmetric encryption which can help ensure the confidentiality and authenticity of the sensitive data such as the encryption keys and has values. The structure and organization of the metadata within the container header can be made adhere to the specifications of the chosen container format and any relevant encryption or digital signature standards, which can promote interoperability and compatibility with various playback and processing tools.
[0162] The header 206 of the media container file 232 can comprise information pertaining to the file format version and compatibility information; the file size; the duration of the media; and top level index to object of interest descriptors.
[0163] In some embodiments, the media container file 232 comprises a root digest for securing contents in the media container file 232 as well as associated metadata 208. The root digest may itself by signed by the user 202 and maybe stored with the metadata 208 in the header. For example, the root digest can be a single, digitally signed root digest for cryptographically binding all critical payload elements and governing metadata. The root digest can be configured to cover every normative byte of encrypted representation payloads and all governing metadata elements (e.g., object definitions, representation tables, KIDs, privacy directives, container structural information); to produce a collision-resistant root digest from these elements; and to authenticate the root digest using a strong digital signature scheme. In a particularembodiment, the root digest can be implemented as a hierarchical hash structure, such as a Merkle tree. However, other functionally equivalent structures (e.g., nested hash-chains, cryptographic accumulators, Post-Quantum Cryptography (PQC) Merkle variants) are also possible. In some embodiments, a root digest can be generated for each data stream or each object of interest and their corresponding metadata.
[0164] Accordingly, the root digest can also a verifier to detect unauthorized changes to: (a) any encrypted object representation bitstream, or (b) the metadata controlling its access, decryption, or processing (e.g., including key IDs, privacy directives, and structural information)by processing the root digest.
[0165] An example Merkle tree implementation of the root digest is shown below, comprising a plurality of leaf nodes each corresponding to a hash of a distinct, canonicalized portion of content or metadata. Example leaf classes / nodes is shown below with the corresponding data.1 . Each encrypted distinct data stream I data representation covered by the root digest (e.g., a node for each individual data stream).2. Each row (or a canonical serialization of the entire table) of a Representation Table for each object (described below), for example including Representation ID, Representation Type, Codec Format Information, Location Pointer, Encryption Key ID Representation, etc.3. Object of interest fields for each object of interest, for example including Object ID, Object Attributes such as Semantic Type and Sensitivity Level, and Actionable Privacy Directives.4. Container structural metadata such as file header information, track map / layout, Ke Locator Box URI, audit records (described above), and revocation log head hash.
[0166] To construct the Merkle tree, suitable hash functions, canonicalization, and suitable tree structures can be used, for example as outlined below.Hash Function: The hash function for leaf and internal node digests can be SHA-256. However, other hash algorithms are possible as well. A Hash algorithm ID (e.g., an OID or standardized enumeration) can be stored per leaf or associated with the Merkle root, referencing the specific algorithm used (e.g., SHA-256, SHA-3, PQC hash algorithm). This can enable different algorithms to be adopted without altering the fundamental tree structure or signature format.Canonicalization: Before hashing, textual metadata (e.g., JSON / XML portions of integrated metadata) can be serialized into a canonical byte representation (e.g., RFC 8785 Canonical JSON). Binary data blocks or container boxes can be hashed in their raw byte order per their specifications, ensuring consistent hash values.Tree Construction: Child node digests (or leaf data) can be concatenated in a predefined order (e.g., lexicographically based on canonical serialization, ora predefined structural order) and then hashed using the specified has function to form parent nodes, recursively, until a single Merkle Root hash is obtained. The Merkle tree version or specific construction rules (e.g., Tree Version) may also be recorded.
[0167] Each root digest or a tuple comprising the root digest (e.g., Merkle Root, Hash algorithm ID, Tree Version) can be digitally signed by the user 202, packager, or other trusted entity (e.g., using ECDSA-P-256). The signature can be time-stamped (e.g., via an RFC 3161 compliant Time-Stamp Authority) for nonrepudiation regarding signing time. The resulting signed blob (e.g., containing the Merkle root and signature) can embedded in a dedicated location within the media container or in metadata 208, such as a proprietary ISO-BMFF uuid box (e.g., named SIGEN_INTEGRITY_SIG) or an similar structure (e.g., Matroska SimpleTag).
[0168] Note that any alteration to an encrypted data stream, a Representation Table entry, a privacy flag, a structural header, or any other data covered by a leaf node will alter the root digest (e.g., change at least one leaf digest). This change can propagate up the Merkle tree, resulting in a different Merkle Root candidate that will not match the signed original Merkle Root. During verification, this mismatch causes overall integrity verification to fail, thereby delivering an end-to-end tamper-evident guarantee for the entire package. Various hash algorithms, signature schemes, orhierarchical hashing structures (e.g., different tree balancing or ordering rules) can be used for implementation.
[0169] In some embodiments, the encryption keys 230 themselves are not stored in the media container file 232. That is, the encryption keys 230 cannot be directly decrypted from the media container file 232. Instead, a representation of the keys can be included in the metadata (238) of the media container file. The representation can comprise data for identifying a storage location (e.g., secure repository 226) as well as an identifier for each key. In particular, the representation of the keys in the media container file 232 can comprise:Encryption-Key Identifier (KID): A value or identifier for each key (e.g., an Encryption Key ID Representation, store in a Representation Table described below) that uniquely denotes the specific content-encryption key (CEK) used for that representation. Note that the KID is not the CEK itself.Key Locator: Metadata, such as a URI for a Key Management Service (e.g., secure repository 226), which can be used by an authorized client to obtain the corresponding CEK.Consequently, possessing the media container file alone cannot yield any plaintext CEK, achieving a "zero-knowledge distribution" of keys.
[0170] In some embodiments, non-limiting implementation for signalling KIDs and Key Locators within the container includes:The KID can be implemented as a 128-bit RFC-4122 UUID, which can be stored in a standard field (e.g., the default_KID field of an ISO-BMFF tenc box associated with the representation's track / layer) or as an Encryption Key ID Representation field within an integrated metadata Representation Table in the media container file 232 (described below).The Key Locator can be implemented as a URI placed in a proprietary uuid box (e.g., named KeyLocatorBox) within the media container file. Alternatively, other formats such as Matroska tags, or signalled in adaptive streaming manifests such as DASH or HLS, can be used to implement the Key Locator.
[0171] To associate a time-segment data stream and its corresponding encryption key ID, the media container file 232 can explicitly signal the association, as shown below.ISO-BMFF based containers (e.g., MP4): The encryption key representation (e.g., a key ID) can be written to a description box such as Sample Group Description Box (e.g., grouping_type 'seig') where it is referenced by a Sample To Group Box (sbgp) that maps contiguous sample ranges (by duration / count) to the appropriate sample group, thus linking time segment data stream to its decryption key ID.HLS Manifests: An EXT-X-KEY tag can be placed in the media playlist of the container file immediately before the time segment data stream (e.g., .ts file, fMP4 segment) where the first sample (e.g., time-segment data stream) is the SDAP and which can be encrypted with the key specified in said tag. The key format attribute (e.g., KEYFORMAT="urn:uuid:<kid>" within EXT-X-KEY) can signal the corresponding encryption key ID.DASH Manifests: Content Protection descriptors within Adaptation Set or Representation elements can specify the encryption key ID (e.g., cenc:default_KID). For per-segment / period key changes, multiple Content Protection descriptors with different encryption key IDs and key acquisition information can be aligned with segment boundaries starting with SDAPs.
[0172] In some embodiments, the encoder settings or live streaming constraints can limit frequent insertion of full SDAPs (which can be larger). Accordingly, the system can employ gradual decoder refresh (GDR) mechanisms (e.g., GDR slices in HEVC; coded-metadata refresh in AV1 ). In such cases, encryption key rotation can be deferred until the GDR window fully completes and a truly independent decoding state is achieved, ensuring no inter-segment prediction dependency across key boundaries.
[0173] In some embodiments, preprocessing by the packaging system for generating the media container file 232 can detect no suitable SDAP at a desired keyrotation boundary within a pre-encoded representation. Accordingly, the packaging system can insert the suitable SDAP. This can be achieved by partially re-encoding asmall segment around the boundary to create an SDAP, or by splicing in a short, independently coded segment (e.g., an IDR-only GOP) at the boundary, before applying the time-segment specific encryption. This measure can ensure deterministic playback for compliant clients across key-change boundaries.
[0174] This SDAP alignment strategy for time-segment encryption can prevent decode failures arising from inter-frame prediction dependencies crossing a keychange boundary. It can also enable fine-grained key rotation for enhanced security (e.g., limiting content encrypted with a single key), more responsive revocation (as keys for older segments can become unavailable), and precise enforcement of timebased access rights.
[0175] In particular, the codec of the media container file 232 can include functionalities for object identification, multi-representation generation, encryption logic, and metadata management directly. This codec can be configured to be both object and privacy aware, in contrast from the traditional frame-based or simple stream-based processing.
[0176] Unlike conventional codecs which are focused on pixel / sample compression across undifferentiated frames, the codec of media container file 232 can comprise the below functionalities: a) Internally perform or tightly integrate with Al modules for real-time or near real-time semantic object identification and tracking within input multimedia. b) Segment multimedia content into distinct object layers. For example, a video frame processed by this codec can be structured (e.g., physically in the bitstream or conceptually) as a composite of a Base / Background Layer (static / global scene elements) and one or more object layers, each corresponding to an identified object of interest or representation thereof. c) Encode multiple distinct data stream representations for each object of interest (as described above), and employ efficient strategies for each representation type. This can include scalable coding techniques (e.g., spatial, temporal, quality, semantic scalability per object layer) or encoding each representation as a distinct, timed sub-track for the object of interest. In some embodiments, encoding includes managing motion vectors, transformations, and predictive coding elements on a per-object- basis, and per-representation-basis to optimize compression while maintaining individual stream integrity. d) Embed comprehensive, structured, privacy-aware metadata directly into its generated bitstream or into tightly coupled container header structures, thereby making the multimedia file self-describing and facilitating autonomous discovery and policy enforcement.
[0177] In particular, the codec can generate and embed the metadata block 208. The metadata 208 can enable ease of discovery, granular access control, and privacy management for the media content in the media container file 232. For each uniquely identified object of interest within the multimedia, the metadata 208 can include a plurality of information, as described below.
[0178] Object ID: Each data stream (e.g., Object Layer) can be assigned an object identifier for the corresponding object of interest (e.g., ObjectlD_X) that is unique at least within the current media container file. The object ID can explicitly indicates if the data stream may be used for cross-container correlation. Examples of the identifier scheme which include:Container-local: A UUID v4 that can be generated at packaging, ensuring uniqueness only within the container for maximum privacy (e.g., unrelated IDs for the same real- world entity in different files).Namespace-global: A UUIDv5 can be derived from a content-owner-chosen ID (e.g., namespace_uuid) and a stable hasd (e.g., raw_entity_hash) of the real-world entity (e.g., plate number, biometric template), for uses like federated analytics or continuous training pipelines requiring cross-file entity linkage.Scope Signalling: A global ID flag (is_global_id flag) (e.g., one bit) can be stored in the media container file 232 or a bitstream SEI message (e.g., sigen_object_layer SEI). This ID can declare the scope of the object of interest, allowing downstream systems to determine if cross-file correlation is permissible.Integrity Binding: Both the object of interest ID and sigaling scope flag can be included in the root digest.Policy Enforcement: A policy ID can be included to identify additional rules introduced at deployments (e.g., "global IDs allowed only on private networks"). This policy ID can be implemented via a KMS, as described further herein.
[0179] Object Attributes: For each object of interest, a set of descriptive characteristics can be provided, which can include:Semantic Type: An enumerated type indicating the object layer's classification (e.g., "Person," "Vehicle," "Face," "LicensePlate," "Building," "TextRegion," "AudioEvent_Speech," "AudioEvent_Alarm").Sensitivity Level: An enumerated type indicating the sensitivity of the object data (e.g., "PII_High," "PII_Low," "CommerciallySensitive," "Public," "InternalUseOnly").Temporal Extent: Data defining the start time and end time (or duration) of the object of interest's presence within the multimedia content.Spatial Extent: A time-series of data (e.g., sequences of bounding boxes, segmentation masks, or 3D coordinates) defining the object of interest's spatial location and shape over time (e.g., over its presence in the multimedia content). These data can be a compact representation or a pointer to a separate metadata stream.Confidence Score: A score generated during project identification / tracking indicating the confidence in the object of interest's detection and classification.
[0180] Object Relationships: In some embodiments, metadata 208 can define relationships between different object of interest instances. Examples are provided below:Hierarchical Relationship: Data which identifies a hierarchical relationship between objects of interest (e.g., ObjectlD_Y (License Plate) is a child of ObjectlD_X (Vehicle)).Interaction Relationship: data which identifies an interactional relationship between objects of interest (e.g., ObjectlD_A (Person) interacts with ObjectlD_B (a held object) during a specific time interval).
[0181] Representation Table: For each object of interest or each data stream, a structured index lists may be generated, including all available distinct data stream representations generated for the object of interest. Each entry can corresponds to one representation or data stream and comprising details thereof, using data as described below.Representation ID: An unique identifier for the specific data representation or data stream for the specific object of interest (e.g., "ObjectX_Repr_HQ_Visual," "ObjectX_Repr_AI_Deidentified_Visual").Representation Type: A standardized enumeration indicating the representation's nature, corresponding to the type of data stream / data representation (e.g., high quality representation, de-identified representation, etc.).Codec Format Info: Data on the codec, profile, level, and parameters used to encode the data stream (e.g., "HEVC Mainl O L5.1 ," "AAC-LC 128kbps," "WebVTT," "ProprietaryLidarFormat_v2").Location Pointer: Information to locate the bitstream segments for this data stream within the media container (e.g., track ID, layer ID in scalable coding, byte offsets, segment index for streaming).Encryption Key ID Representation (KID): An identifier uniquely linking the data stream to a specific content-encryption key (as described above and further below), which can used to retrieve or derive the correct decryption key for accessing the data stream.PTS / DTS timestamps (explicit or implicit via stts, ctts), for example if the data stream is a time segment data stream..Associated Metadata: Data comprising pointers to further metadata specific to this data stream (e.g., language of a textual description, resolution of a visual representation).
[0182] Actionable Privacy Directives: This data can be associated with an object ID or representation ID and can comprise embedded flags, rules, or policy references for guiding data handling, access, and lifecycle. Examples are provided below:Consent Status Flag: An enumerated value identifying the current access authorization consent for a data stream or object of interest (e.g., CONSENT OBTAINED, CONSENT EXPLICITLY DENIED, CONSENT PENDING, CONSENT NOT APPLICABLE, CONSENT WITHDRAWN).Default Representation On No Consent ID: Specifies the representation ID (e.g., an data stream / data representation type of the object of interest) that a compliant playback system should render if consent for higher-fidelity representations is not affirmed for the current user / context.Retention Policy Tag: A tag linking to a data retention policy or rule (e.g., "RETAIN_30_DAYS," "LEGAL_HOLD") applicable to the object of interest or its data representations.Right To Be Forgotten Marker: A Boolean flag indicating if the object of interest or data stream (e.g., containing PH) is subject to erasure requests (e.g., under GDPR). If true, this may point to a data stream / data representation (e.g., a de-identified one) that could be retained post-anonymization. This marker can also be coupled with revocation mechanisms as described further herein.Purpose Limitation Tags: Tags indicating approved purposes for processing or viewing the object of interest's data or specific data representations / data streams (e.g., "SECURITY_MONITORING_ONLY," "ANALYTICS_ANONYMIZED").Geofencing Policy ID: An identifier for a policy restricting access / playback based on geographic location.
[0183] In some embodiments, for every data stream or every object of interest, the codec or multimedia container can have embed metadata that: 1) enumerates its set of semantically distinct representations (e.g., reprjd, repr_type); 2) binds each representation to a distinct content-encryption-key identifier (KID) (e.g., enc_kid); and 3) carries privacy-action directives (e.g., PII_MASK, RIGHT_TO_ERASE). This metadata can itself be integrity-protected by the package-wide signing scheme, as described above. Example signal mechanisms are provided below:HEVC / H.265: Mapping each data stream to Region-based Processing Units (RPUs) and carrying proprietary SEI messages (e.g., sigen_object_layer) in the same access unit.VVC / H.266: Treating each object of interest as an independent sub-picture; for example, a data stream SEI (e.g., NAL-unit type 62) can reference the SubpicID and lists the representation ID and key ID.MPEG-I Immersive Video (MIV): Assigning each object of interest to a distinct patch ID within an atlas and attaching metadata in an atlas-level SEI.OMAF / 3600streaming: Exposing the object of interest as a viewport region; propagating its metadata to a DASH manifest via Essential Property elements keyed by the object’s UUID.
[0184] In some embodiments, the underlying codec or container can lack native region signalling. Accordingly, the same metadata fields can be embedded in a proprietary ISO-BMFF uuid box (e.g., SIGEN_OBJECT_LAYER_BOX) or an equivalent Matroska / WebM construct, aligned with the samples constituting the data streams.
[0185] Advantageously, existing ISO / IEC tools typically describe where a region resides or how to stream different resolutions. In contrast, the present disclosure can bind, at object and data stream granularity, three dimensions-semantic representation, per-representation cryptography, and machine-readable privacy directives-into a single, integrity-protected metadata unit (metadata 208). Specifically, existing HEVC, WC, MIV, and OMAF profiles do not natively provide this tri-partitebinding. Note that the system can be implemented using any suitable codec syntax element, container box, or manifest extension for fulfilling this functional requirement.
[0186] The above-described codec can ensure the included metadata 208 is structured and embedded for discovery by compliant client devices (e.g., media players, editing software, analytics platforms) or server-side systems. In some embodiments, metadata 208 resides in the file header as depicted in FIG. 2A, corresponding to a dedicated "moov" box (in MP4-family containers), or equivalent structures in other formats (e.g., MKV segment information).
[0187] Upon opening or processing the multimedia container 232, a compliant system can first parse the metadata 208, allowing it to:1 . Discover all semantically identified objects of interest.2. Determine the full range of available distinct data stream representations for each object of interest.3. Understand the encoding format and location of each data stream / data representation's bitstream.4. Identify the encryption status and key requirements (e.g., via Encryption Key ID or Representation) for each data stream.5. Become aware of actionable privacy directives associated with objects of interest or the corresponding data streams.
[0188] Accordingly, the multimedia container 232 can significantly reduce reliance on external databases or sidecar files for managing object-level information, access rights, and privacy policies. A compliant playback client, for example, can uses the metadata 208 with authorized decryption keys to decide which data streams to request, decrypt, and render, while automatically adhering to embedded privacy directives (e.g., defaulting to an de-identified data stream if consent for a high-quality view is not affirmed). The codec itself, during encoding, can also use the metadata 208 to correctly package data streams, data representations, and associated security information.
[0189] Note that, as described above, a unique, representation-specific content-encryption key (CEK) can be generated for each distinct data representation or data stream. That is, each object of interest may be associated with a plurality of encryption keys, each corresponding to a particular data stream for the object of interest. For example, for a particular object of interest:The high quality representation data stream can be encrypted with a key labelled as CEK_ObjectX_HQ_Visual.The de-identified representation data stream can be encrypted with a key labelled as CEK_ObjectX_AI-Deid_Visual.The textual description representation data stream can be encrypted with a key labelled as CEK_ObjectX_Textual_ASR.The sensor data representation data stream can be encrypted with a key labelled as CEK_ObjectX_Lidar.
[0190] This per-representation encryption can enable highly granular access control, as a user's authorization determines which specific CEKs (or means to obtain them) they receive.
[0191] The format of the media container file 232 is not restrictive and may be any suitable media container format based on compatibility requirements and desired features. Examples of acceptable formats include MKV, MP4, and AVI, but other formats are possible as well. The format may be adapted based on specific features required (e.g. support for multiple audio tracks, subtitles, chapters, or advanced metadata options).
[0192] In accordance with the present disclosure, the encrypted individual data streams are integrated into the media container file 232. As depicted in FIG. 2A, the individual data streams can include any one or more of: one or more video data streams 240 segmented and generated from the video stream 210, one or more audio data streams 244 segmented and generated from the audio stream 212, one or more text streams 246 segmented and generated from the subtitle stream 214, as well as one or more mask streams 242, each of which may correspond and / or be associatedwith a respective video data stream 240 for isolating the corresponding identified object of interest, as described previously. Various other data streams (as described above) corresponding to the objects of interest can also be included. It should be noted that the one or more video data streams 240, one or more audio data streams 244, and / or one or more text streams 246 may also respectively include the video stream 210, audio stream 212, and / or subtitle stream 214. The video stream 210, audio stream 212, and / or subtitle stream 214 may be encrypted as described previously. Encryption keys, hash values, and / or digital signatures for the video stream 210, audio stream 212, and / or subtitle stream 214 may also be included in the media container file 232 in accordance with the above described processes. It should also be noted that the generation of the one or more text streams 246 by segmenting the subtitle stream 214 and subsequent encoding may be performed in a similar manner to the one or more video data streams 240 and one or more audio data streams 244. For example, the subtitle stream 214 may be divided into different subtitle tracks or segmented based on the speaker of the subtitle text.
[0193] In some embodiments, to ensure the integrity of the container file during transmission and storage, error detection and correction mechanisms may be employed for error handling and data integrity checking. For example, cyclic redundancy checks (CRC) or error-correcting codes (ECC) may be used to detect and correct any detected errors that have occurred due to data corruption or transmission issues, which can help ensure that the media container file 232 remains intact and the multimedia content (i.e. the data streams) can be reliably accessed and played back. In some embodiments, the finalization of the media container file 232 includes verifying the integrity of the media container file 232 using error detection mechanisms and / or adding any necessary metadata or structural elements required by the container format. The finalization process can help ensure that all data is written correctly and the file structure is consistent with the chosen container format’s specifications.
[0194] FIG. 3 depicts a method of generating a multimedia container file for providing segmented access and verification of media content, it should be noted that although no explicit references are made, various processes described herein withregard to FIG. 3 may correspond to those described with regard to FIGs. 2A and 2B. It should be noted that boxes shown in dashes refer to optional processes.
[0195] Initially, a multimedia file is received (302). The multimedia file may be in any conventional format and may be uploaded by an user or otherwise transferred from the user who intends to manage the access to the media content of the multimedia file. The user may be required to authenticate prior to uploading the file. An indication may be output if the multimedia file is successfully received.
[0196] In accordance with the present disclosure, one or more objects of interest may be selected (304). The objects of interest may be used as a basis for generating one or more data streams, the access of which may be managed. In other words, the user may be able to manage the access to the media content by limiting the access to one or more objects of interest. It should be noted that the objects of interest may be visual, audio, or text based. The user may be prompted to indicate one or more objects of interest based upon which one or more data streams are to be generated. Once the selection is made and received, the multimedia file is segmented (306). In particular, the multimedia file is segmented into one or more data streams. The selected objects of interest may be identified and isolated. Specifically, each of the selected objects of interest may be tracked (308) within the multimedia file to isolate each of the objects of interest from the other elements in the multimedia file, as described previously. For example, masks may be generated forframes containing visual objects of interest and visual objects of interest may be isolated from background noise.
[0197] For each object of interest, a corresponding data stream is generated (310), which may only contain the object of interest. The format of the data stream is determined by the object of interest. For example, a visual object of interest may be used to generate a video data stream (e.g. by extracting the pixels in each frame) and an audio object of interest may be used to generate an audio data stream. Each data stream may be processed. For example, masks may be associated with the video data stream and non-object of interest areas in video may be censored. Audio data streams may be processed to remove silence and add timestamps for synchronization. In some embodiments, multiple data streams can be generated foran object of interest, each corresponding to a different data representations for the object of interest. The generated data streams may be encoded in a format corresponding to the type of the data stream (312). It should be noted that each of the video, audio, mask, and text data streams is encoded. For example, video data streams can be encoded using a video codec, audio data streams can be encoded with an audio codec, and mask data streams can be encoded as images. The data streams may also be synchronized (314) to facilitate smooth and logical playback where the data streams are in sync with one another and the original media content. For example, the data streams can be made to start concurrently with empty frames / silence inserted for missing or redacted content or by aligning the streams using metadata of the data streams.
[0198] For each data stream, a corresponding encryption key is generated (316). The encryption keys may be symmetric keys and are unique to each data stream. It should be noted that encryption keys can also be generated for segments of particular data streams on a time basis. For example, a particular data stream may have two corresponding encryption keys, one of which corresponding to the first half of the data stream while the second of which corresponding to the second half of the data stream.
[0199] Each of the data streams is encrypted using the corresponding encryption key or keys (318). In particular, the encryption algorithm may be a symmetric algorithm such that the corresponding encryption key or keys can be used to decrypt the corresponding data stream or segment of the data stream, as the case may be.
[0200] Each of the encryption keys may be encrypted (320) using a public key of the user. An asymmetric encryption algorithm can be used such that only the user, having the corresponding private key, can decrypt the encrypted (symmetric) encryption keys and access the encrypted data streams. However, in the future, the user may also decrypt the encrypted encryption keys for use by other users which they may use to decrypt the data streams authorized by the user. The encrypted encryption keys may be stored in a secure repository (322) for future access. For example, an access control module may retrieve the stored encryption keys toauthorized users. Encryption key data corresponding to encryption key representation can be generated for each encryption key, (e.g., as encryption key ID), the data can be stored as metadata in the container for users to access the corresponding data stream. The encryption keys themselves (e.g., CEK) can be stored in the secure repository, in particular in a key management system for encryption access and management. In particular, the encryption keys can be transmitted to and stored at the key management server. In response, the key management system can provide a corresponding key identifier for each encryption key as the encryption key representation, which can be stored in the media containerfile as metadata. Forfuture access, the key management service can authenticate a key request (e.g., presenting the key identifier(s) for a specific encryption key(s)) and authenticate the request to grant access to the encryption key.
[0201] Hash values may be generated for each of the data streams (324). In some embodiments, a hash value may also be generated for the complete media content (e.g. all of the data streams), which can be compared with hash values generated using the same algorithm at a later time to verify authenticity. The user can provide their own digital signature for authenticity verification (326). In particular, the user can use their private key to provide their digital signature for or in association with each of the data streams. Specifically, the user can digitally sign each of the generated hash values. The digital signatures and / or the hash values can be stored in the secure repository. In some embodiment, one or more root digests each capturing some or all stream data and metadata can be generated, which can be digitally signed (individually) and then package for content verification.
[0202] In accordance with the present disclosure, a media container file (e.g. a multimedia container file) is generated (328). The media container file is generated by packaging the encrypted data streams into the same container file. Relevant header and metadata information is also included in the media container file. In some embodiments, one or more of the encrypted encryption keys, the digital signatures (e.g. the digitally signed content), and the hash values are also packaged into the media container file, for example, in the header or as metadata. The media container file may be used to manage access to the data streams contained therein. The usermay authorize the access of one or more data streams or segments of data streams of their choice to other user(s).
[0203] The systems and methods of the present disclosure can allow the owner of a media container file (e.g. generated as described above) to define granular access control policies for individual data streams, which only contain their respective objects of interest. Permissions may be assigned based on user roles, groups, or other criteria for flexible and customized access control. FIGs. 4A and 4B depict allocation of access to media content in a multimedia container file for providing segmented access and verification of media content. In particular, FIG. 4A depicts a first embodiment of the media access framework and FIG. 4B depicts a preferred embodiment of the media access framework comprising an Access Control Server (ACS) and a Key Management Service (KMS) for Zero-Knowledge key delivery based on Key Identifiers (KI Ds) within the container and client-side token-based authorization.
[0204] Referring first to FIG. 4A, A user 402 may wish to access media content owned or controlled by an owner 404. In particular, the user 402 may be interested in accessing objects of interest in the media content and therefore the corresponding one or more particular data streams. The media content may be a part of a media container file 412. The media container file 412 (e.g. a multimedia container file) comprises a header 414 and metadata 416, and may also comprise the encrypted keys of the owner 418. The media container file 412 also comprises a plurality of encrypted data streams. For example, the media container file 412 may include one or more encrypted video streams 420, one or more encrypted mask streams 422, one or more encrypted audio streams 424, and / or one or more encrypted text streams 426. The user 402 may request the access of one or more of the encrypted data streams in the media container file 412.
[0205] Accordingly, the user 402 may make a request for access of one or more data streams to the owner 404. In some aspects, the systems and methods of the present disclosure may provide an avenue of communication between the user 402 and the owner 404. For example, a web interface, API call, or a dedicated messaging system may be implemented to establish a communication channel between the user402 and the owner 404. In some aspects, the user 402 may specify the desired level of access, which may be a part of the request for access. For example, the user 402 may identify the particular data streams, the types of content, and / or which of the objects of interest they wish to access. Specifically, the user 402 can request access to specific video or audio data streams, mask data streams, or other relevant information associated with the multimedia content, for example, in the media container file.
[0206] If the owner 404 approves the access request of the user 402, the owner 404 may be authenticated. The authentication (e.g. using owner credentials or a secure authentication mechanism) can ensure that only the legitimate owner can initiate an authorization process. The owner 404 may retrieve all encrypted encryption keys 410 associated with the multimedia container file that are stored in a secure repository 406 (as described previously), which may be configured to securely store hash values 408 and encrypted encryption keys 410. For example, the owner 404 may interact with the secure repository 406 directly or through an access control module (not depicted) to retrieve the encrypted encryption keys 410. Once retrieved, the owner 404 can use their private key to decrypt the encrypted encryption keys 410. It should be noted that the encrypted encryption keys 410 were previously encrypted with the public key of the owner, as described previously. In some embodiments, only encrypted encryption keys 410 that correspond to the data streams that the user 402 is authorized to access are retrieved and decrypted. In some embodiments, the owner 404 may have established pre-set policies that allow the user 402 to access the one or more data streams. In some embodiments, the access request of the user 402 may be evaluated against the pre-set policies of the owner 404 (e.g. by the access control module) which may include access control policies. For example, the access control policies can be used to define which specific data stream(s) are permitted to be accessed by each user or group entities. By evaluating the user request against the access control policies (e.g. if policies indicate that the requested data streams are allowed to be accessed), the access request may be granted (e.g. by the access control module).
[0207] In accordance with the present disclosure, the user 402 may be prompted to provide their public key. The user 402 may transmit their public key (e.g. to the access control module), for example, through a protected communication channel to ensure the confidentiality and integrity of the user public key during transmission. The encrypted encryption keys 410 (e.g. the authorized encryption keys) corresponding to the encrypted data streams that the user is permitted to access (e.g. authorized data streams) may be selected (if not selected previously), for example, by the access control module. Each of selected encryption keys can be encrypted (428) by the public key of the user 402. The encryption algorithm may be an asymmetric encryption and can follow substantially the same process as the encryption of the encryption keys by the public key of the owner, as described previously. The selected encryption keys can be delivered to the user 402, for example after encryption and by the access control module, through a protected communication channel. In some embodiments, the selected encryption keys may also be stored in the secure repository 406 after encryption. This can help maintain a secure record of which encryption keys have been shared with each user. Further, future access to the same one or more data streams by the user 402 can be more easily facilitated through direct retrieval from the secure repository 406.
[0208] The user 402 can decrypt the received encryption keys using their private key. The user 402 may also store the decrypted encryption keys for future access of the authorized one or more data streams. A second media container file 412a is provided to the user 402 to access one or more of the authorized data streams. In some aspects, a second media container file 412a is generated, for example, by updating the media container file 412 and provided to user 402. The second media container file 412a may also be substantially the same as the original media container file 412. In some embodiments, the second media container file 412a may be updated to include the encrypted encryption keys 432 selected for the user 402 (e.g. for accessing the one or more authorized data streams) in the header or metadata information of the second media container file 412a, as depicted in FIG. 4A. In that case, the user 402 may obtain the encrypted encryption keys that are selected from the second media container file 412a and may decrypt these encrypted encryption keys using their private key.
[0209] By using the decrypted encryption keys, the user 402 can decrypt the authorized data streams within the second media container file 412a (e.g. a multimedia container file) by decrypting the corresponding encrypted data stream. It should be noted that the user 402 may only decrypt data streams for which they have the encryption key. That is, only the data streams which the owner 404 has authorized and thereby provided the encryption keys for use by the user 402 may be accessed by the user 402. Even if the user 402 had access to the non-authorized encryption keys (for example, from the encrypted keys of the owner 418 in the second media container file 412a), they would be unable to access the data streams that correspond to the non-authorized encryption keys as the user 402 would be unable to decrypt the non-authorized encryption keys. For example, the user 402 may have requested and been granted access to an authorized video stream 420a, authorized mask stream 422a, authorized audio stream 424a, and authorized text stream 426a. Specifically, the user 402 would have the decrypted encryption keys corresponding to the authorized video stream 420a, authorized mask stream 422a, authorized audio stream 424a, and authorized text stream 426a, which they may use to decrypt the authorized streams to access the authorized streams. In some aspects, the user 402 may have requested and / or been granted access to one or more segments of one or more data streams (if the one or more data streams have been encrypted in segments, as described previously). Accordingly, an analogous process can be performed to access the segment(s) of the one or more data streams.
[0210] It should be noted that the exchange of keys as described in the present disclosure may be implemented under secure key exchange protocols such as Diffie- Hellman to establish shared encryption keys between authorized parties and enable secure communication and access to the data streams. Further, blockchain technology may be utilized to implement decentralized access control mechanisms, where permissions are stored and managed on a distributed ledger to provide transparency and immutability.
[0211] Referring now to FIG. 4B, a user may access media content, for example included in the media container file 412 through the use of an access control server 430. The access control sever 430 is coupled to the secure repository 406,implemented as a key management service (KMS). The access control sever 430 may have stored on it the media container file 412. The user can access or retrieve the media container file 412 from the access control server 430. The access control sever 430 can also provide the user 402 with access credentials, for example authentication tokens, for accessing content (e.g., data streams) in the media container file 412. Each authentication credentials may correspond to a particular data stream. The access control sever 430 can grant the credentials based on default permissions / policies or those set by the owner of the media container file 412. The user 402 can also request access to content in the media container file 412 from the owner through the access control sever 430.
[0212] The user 402 can communicate with the key management service 406, for example through the access control sever 430. In particular, the communication can be a mutually authenticated, forward-secret channel (e.g., TLS 1.3 with an ephemeral Diffie-Hellman key exchange). The user 402 can present key representations 418 (e.g., encryption key ID) from the media containerfile 412 forthe desired data streams to the key management service 406, along with the credentials (e.g., a JWT or SAML assertion from the access control server) received from the access control server 430 during authentication.
[0213] The key management service 406 can verify the request from the key representations and the credentials. Policy checks may also be applied (e.g., role, geo-fence, time window, revocation list). If access for the data streams is granted (e.g., the credentials are approved), the requested encryption keys stored in the key management service 406 and corresponding to the key representations and desired data streams is returned to the user 402. In some embodiments, the key management service 406 wraps the encryption keys (e.g., with AES Key-Wrap per RFC 3394) using a per-session secret derived from the secure channel handshake (e.g., via a TLS exporter).
[0214] The user 402 can then decrypt the data streams that they are authorized to access, for example by locally unwrapping the encryption keys and decrypting the data streams in the media container file 412.
[0215] In some embodiments, the key management service 406 can also implement revocation mechanisms. For example, deleting or flagging a encryption key ID within the key management service 406 renders the corresponding representation in every distributed copy of the container undecipherable by compliant clients that must fetch the key. Non-compliant users and playback clients that have cached an earlier key would lose access once that key's validity period, if any, expires.
[0216] In some embodiments, offline access to the data streams may be permitted. In particular, for scenarios requiring intermittent connectivity for a trusted client, the media container file can carry (and be transmitted to the user with) one or more data stream-specific encryption keys that have been each wrapped using an asymmetric public key (e.g., an Elliptic Curve public key) specific to the target device or user. Each wrapped encryption key can be set to have a short expiry interval (e.g., two hours). After expiry, the client must reacquire a fresh key from the key management service 406. Therefore, limited offline usability with the overall zeroknowledge security for long-term access is possible. Such an asymmetrically wrapped key can be stored alongside the key representation (e.g., encryption key ID) and any Key Locator metadata within the container.
[0217] In some embodiments, a data stream or data streams associated with a particular object of interest can be flagged for erasure or access revocation (e.g., via a Right To Be Forgotten Marker). That is, under certain conditions (e.g., time period expiry) no compliant user will be able to decrypt the flagged data stream(s), even if copies of the media container file 412 are already distributed. In particular, the disclosed system can:1. break the cryptographic link between the ciphertext of the revoked representation and any usable decryption key for future sessions;2. allow a verifier (e.g., an auditor) to prove that a revocation event occurred for a given key or object; and3. limit the lifetime of keys that may have been cached on trusted devices to enforce timely revocation.
[0218] In some embodiments, time-bound licensing (Epoch Keys) can be used to implement key revocation. This can be implemented using a wrapping strategy such that each representation or data stream-specific content-encryption key (CEK), denoted K_E(object uuid, representation ID), can be wrapped by the KMS under an Epoch Public Key, PK_epoch,n, before being returned to the user 402. Further, the KMS can rotate the PK_epoch,n (e.g., generating a new key pair PK_epoch,n+1 I SK_epoch,n+1) at configurable intervals (e.g., a default of 24 hours) such that the prior keys are unusable after rotation. In some embodiments, a flag (e.g., must check revocation flag) or an embedded license expiry in the delivered wrapped key (or in metadata like the SIGEN_OBJECT_LAYER_BOX) can require the user 402 to relicense (e.g., request a newly wrapped CEK) from the KMS before or upon an epoch boundary. In some embodiments, revocation actions can be performed. For example, if the a right to be forgotten marker is true for an object of interest, or access is otherwise revoked, the KMS can omit the CEKs corresponding to the object of interest from the set it is willing to wrap with the next (PK_epoch,n+1 ) and subsequent epoch keys. Any previously delivered CEK wrapped with the previous epoch thus becomes unusable by the user after the epoch ends, and therefore re-requesting is necessary.
[0219] In some embodiments, online status check (OCSP-Like) can be implemented for key revocation. In particular, the KMS can expose a lightweight, secure API (e.g., a RESTful endpoint over HTTPS) for querying encryption keys that returns a status (e.g., valid, revoked, expired, unknown) for any queried encryption key ID. Further, a playback gate can be utilized. The user402, before each decryption session or periodically, can be required to query the status of encryption keys (e.g., by encryption key ID) prior to accessing the data streams. The playback client then refuses playback or decryption if an encryption key’s status is not valid. This check can be mandated by metadata flags within the container. Note that the KMS can query the data streams via the media container file and / or the access control server for the status of the encryption key ID for revocation.
[0220] In some embodiments, attribute-based encryption (offline compatible) can be implemented for key revocation, lin particular, for deployments requiring robustoffline revocation, Ciphertext-Policy Attribute-Based Encryption (CP-ABE) may be used, as described below.Encryption Policy: An object of interest's ciphertext can be encrypted under a CP-ABE policy defined by the content owner / system, for example setting the key to expire after a certain time (e.g., policy := object_uuid_X AND (user_consent_status_for_X = true) AND (current_time < key_expiry_timestamp)).Attribute Issuance: The KMS can act as an attribute authority, issuing cryptographic attributes to users, allowing the key to expire after a certain time (e.g., user_consent_status_for_X = true, key_expiry_timestamp = [timestamp]). The user 402 can use these attributes to derive the corresponding decryption key.Revocation: The KMS can revoke or expire user attributes. For example, if access to data streams for an object of interest is withdrawn, the KMS can change the user's consent or access attribute to the object of interest to false. If the policy then evaluates to false based on the user's current attributes, the playback client cannot derive the decryption key, even with a local ciphertext copy.
[0221] In some embodiments, every key revocation action (e.g., flagging a KID as revoked, expiring an attribute, omitting from an epoch key catalog) can be executed within a Hardware Security Module (HSM) ora secure logging component of the KMS. This can yield a signed audit tuple (e.g., (object uuid, representation ID revoked, revocation reason, timestamp, signature of KMS / HSM)). These tuples can form an append-only revocation log. The head hash of this log (or hashes of significant revocation events) can be added (e.g., at set intervals) as a leaf in the Merkle tree of the media container file 412, making revocation evidence itself tamper-evident and verifiable.
[0222] Note that where intermittent connectivity is required, decrypted encryption keys can be cached by the user 402 inside a Trusted Execution Environment (TEE) or similar secure storage. A cache lifetime for such keys can be set by the owner or KMS, At_max (e.g., default of 2 hours and configurable by policy / license), which can be enforced by the TEE’s secure clock or by the deliveredlicense terms. After expiry, the user 402 is required to re-authenticate with the KMS and re-acquire or re-validate keys, thereby enforcing intervening revocation events.
[0223] By enabling encryption keys to be revoked, the present disclosure can enable right-to-be-forgotten requests and other revocation needs. This can ensure that, after revocation, no standards-compliant user can decrypt the data streams for the revoked encryption keys, regardless of how widely the media container file has been distributed.
[0224] In some embodiments, the KMS 406 can establish a secure key delivery channel for providing the encryption keys 410 to the user 402. In particular, the KMS can be configured to enable the below aspects:1. Confidentiality & Integrity: The encryption keys cannot be read or altered by unauthorized parties in transit.2. Mutual Authentication: Both the user 402 and KMS positively identify each other before key material exchange.3. Forward Secrecy: Session keys protecting the exchange can be derived from ephemeral secrets, so compromise of long-term private keys (KMS or user 402) does not expose past exchanged encryption keys.4. Token Binding: Any authorization token (e.g., JWT, SAML assertion) presented by the user 402 to the KMS can be cryptographically tied to the specific secure channel over which it is presented, preventing token replay by an attacker on a different connection.5. Short-Lived Credentials: Delivered encryption keys, or tokens / licenses authorizing their use, can be configured to expire within a bounded interval, forcing periodic re-authentication and enabling timely revocation enforcement.
[0225] In some embodiments, the KMS can implement the below aspects:1. Protocol Choice: Use TLS 1 .3 (or later) for the user-KMS channel. TLS 1 .3 can inherently provide forward secrecy (via ephemeral Diffie-Hellman, e.g., x25519or secp256r1 ) and ensure confidentiality / integrity. Mutual authentication can be achieved via client and server digital certificates.2. Token Binding: The user can provide a hash of a TLS Exporter value (e.g., tls_exporter_hash = TLS-Exporter("EXPORTER-SIGEN-KEY-REQUEST", context_string, 32)) within the authorization token (as described above) sent to the KMS. Upon receiving the token over the TLS channel, the KMS can recompute this hash from its live TLS transcript and the same context string. A mismatch would abort the request for the encryption key, binding the token to that specific TLS session.3. Key Delivery: After successful token validation and policy checks, the KMS can wrap the requested encryption key using an Authenticated Encryption with Associated Data (AEAD) algorithm (e.g., AES-GCM) or a dedicated key-wrap algorithm (e.g., AES Key-Wrap per RFC 3394). The wrapping key can be derived from a secret established during the TLS 1 .3 handshake (e.g., another TLS Exporter value or a Key Derivation Function such as HKDF on the TLS session master secret). The wrapped encryption key can be sent to the user over the TLS channel; where the KMS can discard the per-session wrapping key post-use.4. Lifetime Control: The unwrapped encryption key can be marked as valid for a specific duration, At_max (e.g., two hours), for example set by KMS license response or embedded policy. This lifetime can be enforced within the user’s secure key store or Trusted Execution Environment (TEE). On expiry, the user is required to re-establish a fresh, conformant channel to re-acquire / re-validate the encryption key.Alternative secure transport implementations, such as QUIC over TLS 1.3 can be used as well.
[0226] Note that the encryption keys require protection when stored (e.g., for offline use) or when transmitted from the KMS to the user. This protection can be achieved by cryptographically wrapping each encryption key, so that plaintext keys are never exposed. For transmission, the encryption key can be wrapped with asymmetric key derived from the mutually authenticated, forward-secret session (e.g., a TLS 1 .3 exporter key), as described above. Wrapping algorithms such as AES-Key- Wrap (RFC 3394) or AES-GCM may be used. In some embodiments, offline playback can be enabled using encryption keys embedded in the media container file and wrapped with the client’s public key using an asymmetric scheme such as ECIES or RSA-OAEP. The corresponding private key, securely stored on the device (e.g., in a TEE), would be required to unwrap the encryption key.
[0227] In some embodiments, the user 402 can request and access data streams directly through the access control server 430. That is, the user 402 can, through the access control server 430, request and receive the encryption keys the decrypt the requested data streams. Further, the access control server 430 can provide playback functionalities to allow the decrypted data streams to be played, for example through streaming.
[0228] In particular, for streaming delivery, the disclosed systems and methods can support adaptive streaming (e.g., via protocols like MPEG-DASH or HLS). This adaptive streaming can be implemented through the access control server 430 and can be configured to be aware of, and leverage, the object of interest’s data streams having the multi-representation structure. Such an approach can enable efficient, scalable, and secure delivery of tailored multimedia experiences to the user 402 over varying network conditions, while strictly adhering to the granular access control policies and embedded privacy directives.
[0229] In some embodiments, the access control server 430 can comprise a multimedia server, content preparation module, or Content Delivery Network (CDN) edge server which generates streaming manifests (e.g., MPEG-DASH Media Presentation Description - MPD; HLS M3U8 playlists). The manifests can be constructed by parsing and interpreting the codec-integrated or container-embedded privacy-aware metadata (e.g., metadata 415). The manifest generation process can comprise:1. Identifying distinct data streams associated with each object of interest (each uniquely identified by an identifier of a particular object of interest) or primaryobject of interest streams as defined by the object-centric codec and described in the integrated metadata.2. For each identified object of interest, consulting its Representation Table within the integrated metadata to ascertain available distinct data stream representations (e.g., high quality representation data stream, de-identified representation data stream, etc.).3. Structuring the manifest to explicitly represent these objects of interest and their available representations, including information for client-side selection, decryption, and playback. Examples can include:MPEG-DASH: This can involve defining Period elements. Within a Period, an Adaptation Set may correspond to an object of interest or a logical grouping of its representations. Each distinct semantic representation (e.g., high quality representation, de-identified representation, etc.) is then a separate Representation element within its Adaptation Set, including bandwidth, segment information (e.g., Segment Template, Segment List), and critical metadata-derived elements such as:A unique tag for the specific Representation Type (e.g., via EssentialProperty or SupplementalProperty with a schemeldUri like "urn:ourpatent:repr_type" and a value like "De-identified representation");The encryption key representation (e.g. encryption key ID) for decryption, which can be signalled via standard Content Protection descriptors (e.g., for CENC, with cenc:default_KID set to the encryption key ID from the Representation Table). Multiple Content Protection descriptors can support different DRM systems or key acquisition protocols; andCustom descriptors reflecting actionable privacy directives (e.g., <EssentialProperty schemeldUri="urn:ourpatent:privacy_directive" value="default_no_consent" / >) to enable policy-aware client decisions.HLS: The master M3U8 playlist can list variant streams (EXT-X-STREAM-INF tags) corresponding to quality levels or object of interest representation combinations. Media playlists contain segments, data streams corresponding to representations andencryption key IDs can be signalled using custom HLS tags (e.g., EXT-X-OBJECT- LAYER-ID, EXT-X-REPRESENTATION-TYPE, EXT-X-KEYFORMAT with a URI pointing to a Key Locator) or standard EXT-X-KEY tags with encryption key IDs in their URI attribute (e.g., KEYFORMAT="urn:uuid:[KID]"), associated with specific representation segments. Privacy directives can also be signalled via custom tags.4. Ensuring segment information (e.g., URLs, byte ranges, durations, SDAP alignment) for each data stream / data representation is accurately included, allowing independent client fetching and decoding.
[0230] The integrity of these generated manifests can be ensured, for example, by including their hash as a leaf in the root digest or through separate digital signatures.
[0231] FIGs. 5A and 5B depict methods of allocating access of media content in a multimedia container file for providing segmented access and verification of media content. It should be noted that although no explicit references are made, various processes described herein with regard to FIGs. 5A and 5B may correspond to those described with regard to FIGs. 4A and 4B. It should be noted that boxes shown in dashes refer to optional processes. In particular, FIG. 5A depicts a first embodiment of the media access method and FIG. 5B depicts a preferred embodiment of the media access method through the use of an Access Control Server (ACS) and a Key Management Service (KMS) for Zero-Knowledge key delivery based on Key Identifiers (KI Ds) within the container and client-side token-based authorization.
[0232] Initially, an access request for one or more data streams in a media container file may be received (504). Specifically, each of the data streams may correspond to a respective object of interest. For example, a second user may request access to the data streams directly (e.g. through an GUI or API) or to a first user that is the owner of the media container file (e.g. a multimedia container file), who may make the access request on the behalf of the second user. The access request may also specify the level of access, such as the specific data streams that the second user would like to access. The user may also request to access segment(s) of data streams.
[0233] The first user may be authenticated (506) and may retrieve all of the symmetric encryption keys (508) corresponding to the media container file from a secure repository. The encryption keys may be encrypted symmetric keys. In some embodiments, the first user may choose to only retrieve symmetric key(s) corresponding to the data streams (or segment(s) thereof) that they intend to authorize access of (e.g. the data streams requested by the second user). The retrieved encryption keys may be decrypted by the private key of the first user (510).
[0234] The access request may be evaluated to determine if the second user is authorized to access the requested data streams (512) or segment(s) of the data streams as well as which data streams the second user is permitted to access, if any. For example, the access request may be evaluated against an access policy defined by the first user which outlines the data streams that are authorized for access by each other. In some aspects, the first user may determine if and which data streams (and / or segment(s) thereof) the second user is allowed to access. The encryption keys corresponding the data streams (and / or segment(s) thereof) that the second user is allowed to access may be selected (514) and encrypted with a public key of the second user (516). The encrypted keys may be transmitted to the second user or stored (518). The encrypted keys may be stored in the secure repository for future access and / or in the header or metadata of the media container file.
[0235] The media container file comprising all of the data streams may be updated or generated and provided to the second user (520). The second user can use their private key to decrypt the encrypted keys corresponding to the data streams (and / or segment(s) thereof) that they are authorized to access. By using the decrypted keys, the second user can access the corresponding data streams (and / or segment(s) thereof) they are authorized to access (522). However, they would be unable to decrypt any other decryption keys and therefore unable to decrypt or access any data streams they are not authorized to access.
[0236] Referring now FIG. 5B, another method for accessing media content is depicted, comprising analogous steps to those shown in FIG. 5A. Differences to FIG. 5A are highlighted below for brevity.
[0237] At 502, the user requests access to one or more data streams, for example in a multimedia container file through an access control server (ACS). At 504, the user authenticates themselves to the access control server. The access control server can evaluate user’s identity and authenticate the user based on stored policies, and subsequently determine data streams the user is authorized to access based on their identity and policies. The user can be authenticated by using robust mechanisms (e.g., username / password, multi-factor authentication, OAuth 2.0, SAML).
[0238] In particular, the ACS can evaluate the authenticated user's identity, roles, group memberships, and contextual information against pre-defined access control policies. These policies can specify which objects of interest and types of representations (e.g., by specific Representation IDs) the user is permitted to access. The ACS can also consult Actionable Privacy Directives. If authorization is granted for at least some representations, the ACS can issue a cryptographically signed authorization token (e.g., JWT) to the user (e.g., the user device). This token can securely encapsulate the user's identity, the scope of granted access (e.g., authorized objects of interest, data streams, data representations, encryption keys, etc.), token lifetime, and other relevant data, which can be provided to the key management service.
[0239] In some embodiments, the access control server can provide the user with the media container file. The user can then, at 514, select the desired data streams for access. The access control server can determine the corresponding encryption key representations via the corresponding metadata. In a further embodiment, the user can select one or more data streams for each of a plurality of objects of interest based on privacy policies. For example, the user may select (or may only be able to select) data streams with de-identified representations for access.
[0240] At 507, the user authenticates to a key management service (e.g., via the access control server) using their credentials (e.g., issued by the access control server). The key management service can verify the user request, which comprises the encryption key representations for the desired encryption keys, and verify the user based on authorization token, policies, and encryption ID revocation status. Onceauthenticated, the key management service, based on the user’s identity, roles, permissions, and request context, can provide the user with the secure set of encryption keys for accessing the data streams (or tokens / licenses allowing CEK derivation) at 508. Note that the request can comprise object of interest, Representation Type or Representation ID, and encryption key ID.
[0241] In particular, to verify the user, the key management service can verify the authorization token's (issued by the ACS to the user) authenticity, integrity, expiry, and its binding to the secure channel; extract authorized encryption key IDs and permission scopes; as well as perform policy checks including consulting an internal revocation list for requested encryption key IDs. If verified, the (plaintext) encryption keys associated with the encryption key IDs are retrieved and returned to the user.
[0242] In some embodiments, the key management service can wrap the encryption keys using Session-Derived Key prior to transmission to the user, where the user can unwrap the encryption key locally.
[0243] In some embodiments, the data streams can be streamed over the access control server. Accordingly, the user can fetch the adaptive streaming manifest (MPD or M3U8 master playlist) from the content server or CDN of the access control server, which parses the manifest to understand available objects of interest, their data representations / data streams, bandwidth requirements, segment locations, associated key IDs, and any signalled privacy directives or representation type information. Based on user selection or default presentation policies, the playback client can use information from the parsed manifest and the media container's integrated metadata (e.g., Representation Table ) to identify the encryption key IDs for each desired and potentially authorized representation.
[0244] At 509, the data streams can be decrypted using the received encryption keys. For each relevant object of interest in the manifest, the playback client can cross-reference available manifest representations with the set of authorized encryption keys and associated data. The actionable privacy directives from the manifest (derived from embedded metadata) can be reviewed for each object of interest and its representations. For example, if a Consent Status Flag for an objectof interest indicates no consent for its high quality representation data stream, the playback client (even if possessing a key for other contexts) consults the existing policy (e.g., Default Representation On No Consent ID). If this points to an deidentified representation data stream for which the user holds a key, that version is selected. Additionally, based on the data stream mapping, authorization check, privacy directive evaluation, and other factors (e.g., user preferences, device capabilities), the playback client can establish a playback plan, for example by selecting one authorized and appropriate data representation / data stream per object of interest for rendering / processing.
[0245] In some embodiments, the playback client can adaptively request media segments from the server / CDN. For each object of interest in its playback plan, segments can be requested only from the available representations. Standard adaptive bitrate (ABR) switching can occur within the chosen representation if multiple bitrates / qualities are offered, based on network conditions and buffer status. Segment requests can align with SDAPs if per-segment key rotation is active.
[0246] At 511 , the data streams are rendered for viewing by the user (522). The access control server can decode and composite the individually fetched, decrypted data streams (which may be diverse in type / fidelity, including visual layers, synchronized audio, textual overlays, rendered sensor data) onto a base layer or presentation canvas. This can create the final, coherent, policy-compliant multimedia view. If the user lacks keys for an essential representation, or if all authorized representations are unsuitable due to prevailing privacy directives, the corresponding object(s) of interest and the corresponding data stream(s) may be omitted, replaced by a placeholder, or indicated as "data not available / authorized," per playback policy.
[0247] In some embodiments, the different types of data streams can be decoded as described below.1. Decrypted visual representations (e.g., high quality representation, low quality representations, etc.) are decoded by a video decoder and output to a display.2. Decrypted audio representations are decoded by an audio decoder and output to speakers or headphones.3. Decrypted textual representations (e.g., subtitles) are parsed and rendered as captions or subtitles, or provided to accessibility services, and synchronized with visual / audio playback using embedded synchronization metadata.4. Decrypted sensor-based representations (e.g., Lidar point clouds) may be visualized (e.g., in 2D / 3D, overlaid on visual data) or used for further processing.
[0248] Advantageously, this adaptive streaming mechanism can be distinct from conventional adaptive bitrate streaming. While incorporating bitrate adaptation for selected individual representations, the present disclosure can also dynamically assemble a multimedia experience from granular, cryptographically enforced access rights to different semantic representations of individual objects of interest. The stream can adapt not only to network bandwidth but, critically, also to user permissions, object-specific privacy rules embedded in the media, and the semantic nature of the content delivered for each constituent object of interest. This can ensure only specifically authorized data for each representation is transmitted to and decrypted by the playback client, providing bandwidth efficiency and robust enforcement of nuanced access control policies, all while respecting privacy directives managed directly within the multimedia content.
[0249] Accordingly, it is possible to use the disclosed systems and methods to leverage object-centric metadata and key management, ensuring that in a streaming context, access is granted with fine-grained precision. The present disclosure can dynamically enforce authorization and privacy considerations at the point of key delivery and segment request, adhering to the principle of least privilege and embedded directives, even for complex multi-representation content.
[0250] The systems and methods of the present disclosure can allow a user to verify the authenticity and integrity of the media content in a media container file. FIG. 6 depict verification of media content in a multimedia container file for providing segmented access and verification of media content, which can be implemented via a access control server.
[0251] An owner 602 of a media container file 612 can request a witness 604 to verify the contents inside the media container file 612 (e.g. a multimedia container file). The process for generating the media container file 612 is as described previously. The witness 604 may initiate a request to verify the contents of the media container file 612, for example, through a GUI or API call. The witness 604 may have been granted access to some or all of the data streams in the media container file 612 following the processes as described previously. In this case, it is assumed that the owner 602 has granted the witness 604 access to all of the data streams for verification. However, the verification of selected data streams and / or segments of data streams are possible as well. The process for accessing the data streams are as described above.
[0252] The witness 604 may access (606) the authorized data streams for verification (e.g. verification streams) in the media container file 612. In some embodiments, the media container file 612 can comprise a header 614, metadata 616, and one or more encrypted data streams. The data streams may include encrypted video streams 620, encrypted mask streams 622, encrypted audio streams 624, and encrypted text streams 626. In some embodiments, The media container file 612 can also comprise the encrypted keys of the owner 618 configured to be used to decrypt each of the encrypted data streams and the encrypted keys of the witness 632 configured to be used to decrypt the encrypted data streams that the witness 604 is authorized to access. In some embodiments, the media container file 612 have stored therein encryption key representations rather than the keys themselves. In a particular implementation, the witness 604 is authorized to access all of the data streams.
[0253] In accordance with the present disclosure, the stored hash values 632 (as described previously) may be retrieved for verification. The stored hash values 632 may be retrieved from a secure repository 630 or from the media container file 612 (e.g. from the header 614 of the media container file 612), depending on where the hash values 532 are stored. The secure repository 630 may be configured to store the hash values 632, the encryption keys 634, and / or the digital signatures 636 previously signed by the owner 602. It should be noted that the digital signatures 636may also be retrieved from the secure repository 630 or from the media container file 612 (e.g. from the header 614 of the media container file 612), depending on where the hash values 532 are stored. The retrieved hash values may be presented to the witness 604 for verification. In some embodiments, only the hash values corresponding to the data streams that the witness 604 is authorized to access are retrieved and presented or only the values corresponding to the data streams for which verification is requested are retrieved and presented.
[0254] The witness 604 may calculate a current hash value for each of the accessible data stream or streams to be verified. In particular, the current hash value for each data stream can be calculated using the same cryptographic hash function as used to obtain the stored hash values. The witness 604 may then verify the authenticity of the media content (608) by comparing each of calculated current hash values with each of the stored hash values for the same data stream. If the values match, it serves as confirmation that the contents of the corresponding data stream have not been altered or tampered with since the stored hash values 632 were generated.
[0255] In some embodiments, the metadata 616 can comprise one or more root digests, corresponding to the media container file or data streams(s) therein, as described above. Accordingly, the witness 604, can, through the access control server, parse the media container to extract all data elements originally included as Merkle tree leaves (e.g., encrypted representations, metadata tables, object attributes, privacy directives, structural info) in each root digest. The leaf digests can be computed over the extracted bytes using the hash algorithm specified and canonicalization rules. The Merkle tree can then be reconstructed from the recomputed leaf digests according to Tree Version rules to obtain a Merkle Root candidate for each root digest.
[0256] For each root digest, the embedded signed blob containing the original root digest, the hash algorithm (or ID thereof), the Tree Version and the digital signature provided by the owner on these data can be retrieved. The owner’s digital signal for the root digest can be first verified, for example using the witness’ public key. If the signature is valid, the original root digest can be compared with therecomputed root digest (Merkle Root candidate) to verify that no alterations have been made.
[0257] In some embodiments, The witness 604 may also provide their digital signature as confirmation of authenticity (610). The digital signature may be signed in association with the data streams for which the content has been verified. For example, the witness 604 can digitally sign the stored hash values of the data streams that have been verified, although signing the current hash values is also possible. The digital signature may be signed using a private key of the witness 604, for example, using a digital signature algorithm such as RSA or ECDSA. Further, the digital signature may be stored (628) in association with the corresponding hash values in the secure repository 630. In some embodiments, each root digest can be digitally signed by the witness 604 and saved to the media container file, as described below.
[0258] In some embodiments, upon verification, an attestation log comprising records of the verification process, the verified data streams, the corresponding metadata, and the root digest may be generated (640). In addition to the witness signed root digest, the attestation log can also be signed by the witness 604.
[0259] In some embodiments, a second media container file 612a (e.g. a multimedia container file) is provided to the user 602 and / or witness 604 as confirmation of authentication. In some embodiments, a second media container file 412a is generated, for example, by updating the media container file 612, and provided to the user 602 and / or witness 604. In some embodiments, the digital signature may be packaged (e.g. the digitally signed content) into the second media container file 612a as metadata 632a. As depicted in FIG. 6, digital signature may be packaged with the encrypted keys of the witness 632. The second media container file 612a may also be provided to an external party as additional proof that reflects the added verification.
[0260] FIGs. 7A and 7B depict methods of verifying media content in a multimedia container file for providing segmented access and verification of media content, it should be noted that although no explicit references are made, various processes described herein with regard to FIGs. 7A and 7B may correspond to thosedescribed with regard to FIG. 6. It should be noted that boxes shown in dashes refer to optional processes. In particular, FIG. 7A depicts a first embodiment of the media verification framework and FIG. 7B depicts a preferred embodiment of the media verification framework comprising an Access Control Server (ACS) and a Key Management Service (KMS) for Zero-Knowledge key delivery based on Key Identifiers (KI Ds) within the container and client-side token-based authorization.
[0261] Referring first to FIG. 7A, Initially, a request for verification of one or more or all data streams in a media container file (e.g. a multimedia container file) may be processed (702). For example, the owner of the media container file may request a witness to verify the data streams. The request can be made through a GUI or API. The witness may access the data streams to be verified (704), if needed.
[0262] Hash values that correspond to the data streams to be verified can be retrieved from storage and provided to the witness (706). The stored hash values may be generated previously for the data streams to be verified. In some embodiments, the stored hash values are retrieved from a secure repository. In some embodiments, the stored hash values are retrieved from the media container file. The current hash values of the data streams to be verified can be generated and provided to the witness. The witness can compare the current hash value of each of the data stream(s) to be verified with the corresponding stored hash value to verify if the data stream(s) has been modified since the stored hash value was generated (708). That is, if the hash values are consistent with each other, the authenticity of the data stream(s) can be verified and confirmed.
[0263] Once the authenticity of the data stream(s) to be verified is confirmed, the witness can provide their digital signature (710). In particular, the witness may use their private key to digitally sign the stored hash values, which can serve as further proof of authenticity and as verification that the data streams have not been altered since the stored hash values were generated. The witness signatures can also be stored in the secure repository, for example, with signed hash values to serve as proof (712).
[0264] The media container file comprising all of the data streams may be updated or generated and provided to the owner (714). In particular, the witness signatures (e.g. the digitally signed content) may also be packaged into the media container file (e.g. a multimedia container file). For example, the witness signatures may be included in the media container file with the associated stored hash values as metadata. The media container file can serve as proof of verification or confirmation of authenticity.
[0265] Referring now to FIG. 7B, another method for verifying media content is depicted, comprising analogous steps to those shown in FIG. 7A. Difference to FIG. 7A are highlighted below for brevity.
[0266] In particular, at 705, the witness can retrieve metadata from the media container file for verification. Specifically, the one or more original root digests and the corresponding hash algorithm ID for identifying the hashing algorithm used, tree version, and owner’s signature can be retrieved. The data elements (e.g., data stream, metadata corresponding to the data streams etc.) corresponding to each root digest can be retrieved or extracted from the media container file. At 707, each root digest can be recomputed using the hash algorithm specified by the algorithm ID and the extracted data element to reconstruct the root digest. The owner’s signature can also be verified. By comparing each reconstructed root digest with the original root digest, the corresponding contents of the media container file can be verified (708).
[0267] The witness can provide their digital signature at 710. In particular, the witness can digitally sign each original root digest, where the signature can be associated with an ID of the witness as well as a timestamp of the signing time. The witness signature can be packaged into the media container file (at 714) in a number of suitable manners. The signature may be directly embedded or inserted into the media container file (e.g., as metadata) or appended to an audit log track in the media container file. In some embodiments, the signature can be recorded in an external secure edge, for example managed by a key management service.
[0268] In some embodiments, an attestation log can be generated at 716.
[0269] In accordance with the present disclosure, the systems and the methods of the present disclosure may be implemented with a number of components, each of which may perform a number of specialized functions. FIGs. 8A-8C depict sequence diagrams of generating a multimedia container file for providing segmented access and verification of media content, the authorization of access for the data streams within the multimedia container file, and the verification of the data streams within the multimedia container file showing various components and processes, it should be noted that although no explicit references are made, various processes described herein with regard to FIGs. 8A-8C may corresponds to those described with regard to FIGs. 1-7B. Note that FIGs. 8A-8C are provided as an overview of the interaction between various users and modules of the disclosed systems according to example embodiments only. In particular, FIGs. 8A-8C generally correspond to the embodiments of the present disclosure as shown in FIGs. 4A, 5A, and 7A and certain embodiments in FIGs. 2A, 2B, and 3.
[0270] Referring to FIG. 8A, a first user 802 can upload or transfer a multimedia or media file to a media engine 808 for segmentation (814). Specifically, the first user 802 may wish to generate one or more encrypted data streams for access management. The media engine 808 may prompt the first user 802 to make a selection of objects of interest for use in data stream generation. The first user 802 can select or otherwise identify a number of objects of interest (816) to the media engine 808, for example, through a GUI of their user device. It should be noted that the objects of interest may be visual, audio, or text based.
[0271] The media engine 808 can segment the received media file into a number of data streams (818) based on the selection of objections of interest and additional criteria set by the user. The media engine may be configured for the processing of the media content in the media file and include functionalities for object of interest identification / tracking, media file segmentation, and data stream generation. Each segmented data stream can comprise one or more objects of interest, for example, as defined / selected by the user, where other elements are censored or redacted. The identified objects of interest are tracked (820) by the media engine 808 to extract the objects of interest from the remaining elements in the originalmedia file. The media engine 808 can generate a number of data streams (822) based on the objects of interest selected by the first user 802. The data streams may also be generated based on criteria set by the first user 802 where they could specify the content (e.g. objects of interest) to be included with each stream. The media engine 808 can also perform additional processes for data stream generation. For example, the media engine 808 may synchronize the generated data streams. The media engine 808 may also include data with regard to each of the data streams as metadata. The generated data streams can be encoded using a suitable method based on the type of data stream by the media engine 808 and forwarded to a file engine 810.
[0272] The file engine 810 may be configured to perform various processes related to the management of the data streams. For example, the file engine 810 may include functionalities for data stream encryption, data stream packaging and media container file management. The file engine 810 can generate an encryption key (824) for each of the data streams generated by and received from the media engine 808. The encryption keys are unique. The user may also define the format of the key and / or the encryption criteria. In some aspects, the file engine 810 can generate a encryption key for a particular time segment of a data stream. Accordingly, after encryption, only the particular time segment of the data stream can be decrypted and accessed. The file engine 810 can encrypt each of the data streams (826) using the corresponding encryption key. In particular, each of the data streams can be encrypted by the file engine 810 using a symmetric algorithm. The file engine 810 can also calculate or generate a hash value (828) for each of the data streams for content verification, which can be done using a suitable hash function. The file engine 810 can also generate an overall hash value corresponding to all the media content (e.g. all of the data streams).
[0273] The file engine 810 can prompt the first user 802 to provide their digital signature (830) to digitally sign the data streams, for example, through an implemented GUI. The first user 802 can use their private key to digitally sign each of the generated hash values (832) for verification. The digital signature(s) (e.g. the signed hash values) are returned (834) to the file engine 810. The file engine 810 canforward sensitive information (836) to a repository 812. For example, the sensitive information may be information that needs to be securely protected, such as the digital signature, the hash values, and / or the encryption keys.
[0274] The repository 812 may be configured to manage sensitive data. In particular, the repository 812 may perform encryption and storage functionalities to secure and protect sensitive data. The repository 812 may store the digital signature and the hash values for retrieval in the future. The repository 812 can encrypt each of the received encryption keys (838) with a public key of the first user 802, which may be requested and received. The repository 812 can encrypt the encryption keys using an asymmetric encryption algorithm. The repository can also store the encrypted encryption keys (840) for safeguarding and future retrieval. In some aspects, the encrypted encryption keys may be returned (842) to the file engine 810.
[0275] The file engine 810 may generate a multimedia / media container file (844) by packaging the generated data streams. The file engine 810 can also update the metadata and header of the multimedia container file as needed. For example, the file engine 804 may specify the number of data streams and the content of each data stream. In some aspects, the file engine 810 can include the signed (or unsigned) hash values in the metadata of the multimedia container file for verification. In some aspects, the file engine 810 can include the encrypted encryption keys in the metadata of the multimedia container file such that the encrypted encryption keys can be accessed by the first user 802 from the multimedia container file. The file engine 810 can return the generated multimedia container file (846) to the first user 802, which can be used by the first user 802 to manage the access of media content through authorization of the data streams in the multimedia container file.
[0276] Referring to FIG. 8B, a second user 804 may request media content access (848) from the first user 802. For example, the second user 804 may request to access one or more encrypted data streams in the multimedia container file corresponding to objects of interest that the second user 804 would like to view. In some embodiments, the second user 804 may request to access a time segment of a particular data stream, which may be processed in the same manner as described herein. A communication channel may be established between the first and secondusers, for example, via a GUI, to allow the request to be viewed and authorized. If the first user 802 intends to authorize the request, they can request the repository 812 to return all of the stored encryption keys (850). The repository 812 may request the first user 802 to authenticate themselves to ensure the security of the encryption keys.
[0277] The repository 812 can return all of the encryption keys (852) to the first user 802. The first user 802 can decrypt the encryption keys using their private key such that the encryption keys can be made available. The first user 802 can select encryption keys that are authorized (854) for use by the second user 804. For example, the first user 802 may authorize access to one or more of the data streams requested by the second user 804 and select the encryption keys corresponding to the authorized data streams. In some aspects, the repository 812 can evaluate the access request of the second user 804 and select the corresponding encryption keys. In some embodiments, the repository 812 can evaluate the request against rules defined by the first user 802 to determine which encryption keys are to be selected, if any.
[0278] The second user 804 may be prompted to provide their public key (856), which they can transmit to the repository 812. The repository 812 can encrypt (858) each of the selected encryption keys with the public key of the second user 804, which can be performed using an asymmetric algorithm. The repository can also store the selected keys (860) that have been encrypted using the public key of the second user 804 for safeguarding and future retrieval.
[0279] The selected keys that have been encrypted using the public key of the second user may be returned (862) by the repository 812 to the file engine 810. The file engine 810 can generate a new multimedia / media container file (864) by updating the header or metadata of the original multimedia container file to include the selected keys that have been encrypted using the public key of the second user 804. The updated multimedia container file can be returned (866) to the second user 804. The second user 804 can retrieve the selected keys from the repository 812 or the updated multimedia container file. By using their private key, the second user 804 can decrypt the selected keys and use the decrypted keys to access the one or more data streams (868) they are authorized to access.
[0280] Referring to FIG. 8C, the first user 802 may wish to have a third user 806 verify the authenticity of the media content in the multimedia container file (870), for example, through an implemented GUI. The third user 806 can request and access the multimedia container file (872), particularly the data streams that the first user 802 would like to verify and are contained therein, as described previously.
[0281] The file engine 810 can generate and provide the current hash values of the data streams that need to be verified to the third user 806. The repository 812 can return the stored hash values to the third user 806. In some embodiments, the stored hash values can be retrieved from the multimedia container file. The user can verify that the data streams (for verification) have not been altered (874) by comparing the current hash values with the stored hash values and confirm that the hash values have not changed. It should be noted that the current and stored hash values are calculated using the same hashing function. The third user 806 can digitally sign the stored hash values using their private key as proof of their authenticity. The digital signature(s) (e.g. the third user’s signed hash values) are sent (876) to the repository 812. The repository 812 can store the hash values signed by the third user (878) as proof against tampering.
[0282] The repository 812 can also forward the hash values signed by the third user (880) to the file engine 810. The file engine 810 may also prompt and receive the hash values signed by the third user directly from the third user 806. A verified multimedia / media container file can be generated (882) by updating the multimedia container file. The file engine 810 can update the multimedia container file to include the hash values signed by the third user in the header or metadata of the multimedia container file. The verified multimedia container file can be returned (884) from the file engine 810 to the first user 802 such that the first user 802 has proof that that the media content in the multimedia container file is authentic.
[0283] Advantageously, the present disclosure's modular design (e.g., KMS, object-aware codec, adaptive streaming components) can offer significant design flexibility and adaptability foremerging trends in key management, cryptography (e.g., post-quantum algorithms), Al models for object analysis / representation generation, and secure content sharing. For example, the systems and methods may integratepost-quantum cryptography or decentralised audit ledgers such as blockchain for managing aspects of identity, policy, or audit trails, addressing the evolving digital security and privacy landscape.
[0284] FIGs. 9A-11 C relate to an example application of the present disclosure in processing video footage from a traffic incident between Vehicle A and Vehicle B. An object-centric codec can identify relevant objects of interest (e.g., Vehicle A, Driver A, Vehicle B, Bystander D, License Plate A). For each object of interest, multiple representations corresponding to distinct data streams can be generated (e.g., high quality representation, de-identified representations of faces and plates, textual transcript of statements, Lidar data of vehicle positions / damage, etc.). Access for different stakeholders could be:1 . Law Enforcement: May receive encryption keys for high quality representation data streams of involved vehicles / persons, full audio transcripts, and Lidar data for accident reconstruction.2. Driver A: May receive encryption keys for high quality representation of their own vehicle and their own de-identified representations, plus de-identified representations of Vehicle B, and textual descriptions of official interactions; but not encryption keys for identifiable visuals of Driver B or Bystander D.3. Insurance Company (Driver A): May receive encryption keys for de-identified representations of Vehicle A and B, Lidar data, and textual incident descriptions; but not identifiable visuals of persons unless consent / policy allows.4. Public News Report: May receive encryption keys only for de-identified representations of the scene (with all persons / plates anonymized) or a textual summary.5. Bystander D: Their data (e.g., visual representation) remains encrypted with non-distributed keys, or only the de-identified representations are made available under strict controls, respecting privacy unless legally mandated.The hierarchical integrity signature can also allow parties involved to verify received representations and governing metadata. Key revocation can be used if, for example, consent for a representation is withdrawn.
[0285] FIG. 9A and 9B depict segmentation of source media into a plurality of media streams. In particular, FIG. 9A depicts the segmentation of an example video stream in which a traffic accident was captured. A plurality of processed streams are generated, each of which can be tailored for distinct stakeholders while maintaining individual privacy.
[0286] The source media 902 comprises a video sequence illustrating a traffic incident initiated by a pedestrian crossing at a green light. As depicted in FIG. 9A, a car and a truck collide in the closest lane to the camera, observed by nearby pedestrians and an oncoming vehicle. After the incident, the witnesses leave, and the drivers engage in a dispute. The source media 902 may be processed (904) using the systems and methods of the present disclosure as described previously to generate five data streams. A first data stream 906 may include the complete video without any alterations. The first data stream 906 may be provided to law enforcement and depicts the entire event without any omissions.
[0287] The remaining streams may include various masks and / or filters as described previously to redact or censor select areas. A second data stream 908 may be from the perspective of a witness and provided to the witness. The second data stream 908 may include the vehicles involved in the accident and the witness, which can be identified as the objects of interest, which are then isolated and tracked to generate the second data stream 908. As shown in FIG. 9A, elements that are not included as the objects of interest may be censored to protect their privacy. A third data stream 910 may include the vehicles involved in the accident and the drivers of the vehicle. The vehicles involved in the accident and the drivers thereof can be identified as the objects of interest, which are then isolated and tracked to generate the third data stream 910. A fourth data stream 912 may include all vehicles and their movements, which can be provided to the insurance company as it includes all vehicle movements that may be relevant for the claim. All vehicles can be identified as the objects of interest, which are then isolated and tracked to generate the fourth datastream 912. A fifth data stream 914 may be a time based segment stream in which only the portion of the video relevant to the accident is included. The fifth data stream 914 may be useful for stakeholders requiring a concise review of the incident, such as accident reconstruction experts or emergency response coordinators.
[0288] It should be noted that in at least some embodiments the above described entities would not have access to any other data streams. Specifically, the data streams may be encrypted as described preciously with access to each stream provided to an authorized entity as described above. The entities may decrypt the corresponding data stream to access the corresponding media content. For example, the insurance company would be able to view the fourth data stream 912 but would be unable to view any other data streams. As such, the insurance company would only be able to see the movement of the vehicles, as included in the fourth data stream 912. As another example, the stakeholder entities would only be able to view the fifth data stream 914 but none of the other data streams. Further, the stakeholder entities would be only able to view the limited segment / duration of the data stream in which the accident occurs but would not be able to view the events that take place before or after the accident.
[0289] FIG. 9B is substantially analogous to FIG. 9A. However, in FIG. 9B, a sixth data stream 916 corresponding to a data stream comprising de-identified representations of the identified objects of interest is shown.
[0290] In accordance with the present disclosure, a playback engine may be provided to implement the processes and functionalities of the systems and methods as described with regard to FIGs 1-9B. In particular, FIGs. 10A and 10B depict example user interfaces of such a playback engine for implementing the systems and methods of managing multimedia content.
[0291] Referring first to FIG. 10A, The user interface (e.g. GUI) may be implemented as a webpage or application. A playback panel 1002 may be provided for the users, including standard playback functionalities to allow a user to control and view the playback of any of the original media streams or generated data streams. The user interface may also include a toolbar or taskbar 1012 configured to providegeneral settings control for the user. For example, options may be provided for opening / saving / exporting / uploading data streams, verifying user credentials (e.g. login), and general user interface / playback engine settings.
[0292] The user interface comprises a manual object detection panel 1004 to allow the user to select and identify objects of interests. For example, the manual object detection panel 1004 can include an brushstroke selection option 1006 for selecting objects of interest that appear in the playback panel, as well as a text prompt option 1008 to allow the user to enter text queries for selecting objects of interest. There may also be a class option 1010 to limit the automatic detection of objects of interest and / or the selection of objects of interest to a particular type (e.g. vehicle or pedestrian).
[0293] The selected objects of interest can be tracked to generate a number of data streams. For example, a number of video data streams may be generated, each corresponding to a particular object of interest 1014 that was detected. A mask stream 1016 may be generated for each of the video data streams, which the user can choose to selectively apply or remove to any particular video stream for playback shown in the playback panel 1002, the results of which may be previewed as well. The user can also detect and remove any available audio data streams 1018 that have been generated as needed. Briefly referring to FIG. 10B, substantially analogous to FIG. 10A, de-identified representations 1028 for identified objects of interest can be selected by the user for data stream generation.
[0294] The user interface can also comprise a timeline panel 1020 which includes timeline tools 1022 such as segment, undo / redo, and delete. The timeline panel 1020 can display a breakdown of the video's sequence, including the source tracks, icons and colored markings that can indicate the presence of distinct objects of interest at designated intervals. For example, the timeline panel can display the video track 1024 as well as the subtitle and audio track 1026 of the current media stream, to allow the media stream to be edited.
[0295] The user interface can also present information about each object of interest or data representations thereof, the representations that the user is authorizedto view, an integrity status (if verification is performed), as well as permission controls such as requesting higher level access, sharing data streams, initiation verification, etc.
[0296] It should be noted that the playback engine may be configured to perform stream assembly and access control, as described previously. For example, upon initiating playback through the user interface, the playback engine can retrieve the user's available decryption keys (e.g. from the secure repository or local storage), which can be used decrypt the authorized data streams within the media container file. In some aspects, the playback engine can also selectively assemble the authorized video, audio, text, and mask data streams for display to the user, where the displayed stream can be determined on the data streams authorized to be accessed by the user.
[0297] In some aspects, the playback engine can be configured to retrieve the stored hash values and digital signatures for the authorized data streams from the multimedia container file, as described previously. The playback engine may also be conjured to calculate the hash values for the decrypted streams using the same cryptographic hash function as was used for the stored hash values. In some embodiments, the playback engine can verify the user's digital signatures using their public key to provide assurance of the content's authenticity and origin. In some embodiments, the playback engine may determine that witness digital signatures are present in one or more data streams and validate witness digital signatures using the public key of the witness to confirming verification and content trustworthiness. In some embodiments, the user interface may display visual indicators or messages that represent verification status of the data streams. For example, icons or text labels to signify owner-signed, user-verified, or unverified content may be displayed.
[0298] In some embodiments, the playback engine is configured to synchronize the data streams, for example, by utilizing embedded synchronization metadata, as described previously. The synchronization metadata can include timestamps, frame numbers, and object IDs for synchronizing the playback of different data streams. This can ensures that video, audio, subtitles, and mask data streams are aligned correctly and presented coherently to maintain the temporal and spatial relationships betweenobjects of interest. In some embodiments, the playback engine may be configured to utilize techniques to represent redacted area. For example, the playback engine may decrypt one or more corresponding mask stream and apply the binary masks in the mask streams to each video frame, revealing the authorized objects and concealing the redacted areas. Other techniques such as alpha channel composition, color key replacement and zero padding may also be used to represent redacted areas. In some embodiments, the playback engine may omit data streams that a user is not authorized to access. Accordingly, the objects of interest in the omitted data streams are not shown to the particular user. In some embodiments, the playback engine can replace the redacted content with alternative representations to provide context or maintain visual coherence. For example, the redacted content may be redacted or censored by blurring / pixelation, using placeholder images, or with silhouettes.
[0299] In some embodiments, the playback engine is configured to provide the user with various playback functionalities, for example through the user interface depicted in FIG. 10A. For example, options such as play, pause, stop, seek, volume adjustment, and stream selection may be included for the displayed media stream. The playback engine may also provide options for how redacted content may be represented. The playback engine can also provide closed captions, audio descriptions, and adjustable playback speeds to enhance accessibility.
[0300] FIGs. 11A and 11 B depict example segmentation approachs for a plurality of media streams. As depicted, audio stream 1102 comprises the audio track of the entire sequence of events shown in the original media file. The audio track may be segmented into audio data streams 1104. For example, the sounds specific to the events depicted in the media file, such as the accident and the arguments, can be extracted as objects of interest to form a first audio data stream 1104a while non- relevant sounds such as background noise form a second audio data stream 1104b. Similarly, video stream 1106 can comprise the unedited video footage / track of the media file, which can be segmented into a plurality of video data streams 1108. Each of the video data streams 1108 may correspond to a particular person or vehicle that can be considered an object of interest. The individual video data streams of 1108 may be shown in the order of appearance in the original video stream. For example,a first video data stream 1108a depicts a pedestrian that appears in the video stream 1104 from the beginning and is shown first. The duration of the individual data streams can correspond to the appearance of appears for the respective object of interest. For example, a video data stream 1108b includes one of the vehicles involved in the accident and has a length that is shorter than the original video stream 1106. A background video data stream without any object of interest may also be segmented. Mask data streams 1112 can be generated by applying masks which redact parts of the video of the corresponding video data streams 1108. For example, each of the mask data streams 1112 corresponds to a respective video data stream 1108. As depicted in FIG. 11 A, the object of interest that appears in the stream is redacted or censored. For example, mask data streams 1112a and 1112b corresponding to video data streams 1108a and 1108b show the pedestrian and vehicle being redacted, respectively. It should be noted that the mask data streams 1112 may be applied in combination with one or more of the video data streams 1108 to provide a video stream where one or more objects of interest are redacted. Text streams 1114 can also be provided, which may include subtitles or other text information. For example, the text streams 1114 can include a transcript 1116 of the argument between the parties involved in the accident. It should be noted that the depicted data streams have yet to be encrypted and may be arranged and processed as required to generate a multimedia container.
[0301] Referring now to FIG. 11 B, which is substantially analogous to FIG. 11 A, a plurality of data streams 1118 comprising de-identified representations of the objects of interest is also shown. The de-identified representations can be used to generate one or more video streams where the original object(s) of interest is replaced with contextually relevant representation(s) that cannot be used to identify the original object(s). For example, a video data stream 1120a can be generated where the original person is replaced by an Al altered person.
[0302] FIGs. 12A and 12B depict example comparisons of objects of interest prior and after de-identification. In particular, 1202a and 1204a depict two objects of interest as they originally appear and 1202b and 1204b depict the same two objects of interest after de-identification.
[0303] It would be appreciated by one of ordinary skill in the art that the system and components shown in the figures may include components not shown in the drawings. For simplicity and clarity of the illustration, elements in the figures are not necessarily to scale and are only schematic. It will be apparent to persons skilled in the art that a number of variations and modifications can be made without departing from the scope of the invention as described herein.
[0304] It is contemplated that any part of any aspect or embodiment discussed in this specification can be implemented or combined with any part of any other aspect or embodiment discussed in this specification, so long as such those parts are not mutually exclusive with each other.
[0305] It should be recognized that features and aspects of the various examples provided above can be combined into further examples that also fall within the scope of the present disclosure.
[0306] When used in this specification and claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps or integers are included. The terms are not to be interpreted to exclude the presence of other features, steps or components. Additionally, the term "connect" and variants of it such as "connected", "connects", and "connecting" as used in this description are intended to include indirect and direct connections unless otherwise indicated. For example, if a first device is connected to a second device, that coupling may be through a direct connection or through an indirect connection via other devices and connections. Similarly, if the first device is communicatively connected to the second device, communication may be through a direct connection or through an indirect connection via other devices and connections. Further, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0307] The embodiments have been described above with reference to flow, sequence, and block diagrams of methods, apparatuses, systems, and computer program products. In this regard, the depicted flow, sequence, and block diagrams illustrate the architecture, functionality, and operation of implementations of variousembodiments. For instance, each block of the flow and block diagrams and operation in the sequence diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified action(s). In some alternative embodiments, the action(s) noted in that block or operation may occur out of the order noted in those figures. For example, two blocks or operations shown in succession may, in some embodiments, be executed substantially concurrently, or the blocks or operations may sometimes be executed in the reverse order, depending upon the functionality involved. Some specific examples of the foregoing have been noted above butthose noted examples are not necessarily the only examples. Each block of the flow and block diagrams and operation of the sequence diagrams, and combinations of those blocks and operations, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0308] Use of language such as "at least one of X, Y, and Z," "at least one of X, Y, or Z," "at least one or more of X, Y, and Z," "at least one or more of X, Y, and / or Z," or "at least one of X, Y, and / or Z," is intended to be inclusive of both a single item (e.g., just X, or just Y, or just Z) and multiple items (e.g., {X and Y}, {X and Z}, {Y and Z}, or {X, Y, and Z}). The phrase "at least one of" and similar phrases are not intended to convey a requirement that each possible item must be present, although each possible item may be present.
[0309] The invention may also broadly consist in the parts, elements, steps, examples and / or features referred to or indicated in the specification individually or collectively in any and all combinations of two or more said parts, elements, steps, examples and / or features. In particular, one or more features in any of the embodiments described herein may be combined with one or more features from any other embodiment(s) described herein.
Claims
CLAIMS:1 . A method of managing access to multimedia content, the method comprising: receiving a multimedia file; generating at least one data stream for each of at least one object of interest in the multimedia file such that each data stream of the at least one data stream is a segment of the multimedia file that comprises and corresponds to a respective representation of a respective object of interest of the at least one object of interest; generating at least one encryption key, each encryption key of the at least one encryption key corresponding to a respective data stream of the at least one data stream; encrypting each of the at least one data stream with a corresponding encryption key of the at least one encryption key, each encrypted data stream configured to be decrypted by the corresponding encryption key; receiving at least one digital signature for the at least one data stream; and generating a media container by packaging the at least one data stream and the at least one digital signature.
2. The method of claim 1 , wherein a plurality of data streams are generated for each of the at least one object of interest.
3. The method of claim 1 , further comprising: generating at least one encryption key identifier, each being a unique identifier corresponding to a respective encryption key; and packaging the at least one encryption key identifier in the media container.
4. The method of claim 3, further comprising: storing the at least one encryption key in a key management service secure repository;wherein each of the at least one encryption key is retrievable for decrypting the respective data stream from the key management service using a respective encryption key identifier and corresponding authentication data.
5. The method of claim 1 , further comprising: embedding a representation table as metadata in the media container file for each of the at least one data stream; wherein the representation table comprises: an identifier of the respective object of interest and a type of the respective representation.
6. The method of claim 5, wherein the representation table further comprises a respective encryption key identifier.
7. The method of claim 5, wherein the representation table further comprises codec information, a location pointer, language information, resolution information, or combinations thereof.
8. The method of claim 1 , further comprising: embedding one or more privacy directives as metadata in the media container file for each of the plurality of data streams; wherein the one or more privacy directives comprise information for accessing the respective data stream.
9. The method of claim 8, wherein the one or more privacy directives comprise: consent status data, default representation for rendering data, retention policy data, geographical policy data, purpose policy data, or combinations thereof.
10. The method of claim 1 , wherein each respective representation is a high quality representation, a low resolution representation, an de-identified representation, a text description representation, a sensor data representation, or an embedding representation.
11. The method of claim 10, wherein the de-identified representation is generated by modifying the respective object of interest using at least one artificial intelligence model to generate a representation of the respective object of interest lacking features permitting identification of the respective object of interest.
12. The method of claim 11 , wherein the at least one artificial intelligence model is configured to modify features in the respective according to a predetermined policy or threshold by performing feature identification and feature alteration.
13. The method of claim 11 , further comprising: evaluating the modified object of interest based on one or more privacy policies; and storing records of the evaluating as metadata in the media container file.
14. The method of claim 1 , further comprising: generating one or more root digests for the at least one data stream and corresponding metadata.
15. The method of claim 14, wherein each of the at least one digital signature is assigned to a respective root digest for verification.
16. The method of claim 14, wherein each root digest comprises a plurality of leaf nodes, each corresponding to: one of the at least one data stream, one of the at least one object of interest, or a representation table comprising metadata for the respective representation.
17. The method of claim 1 , wherein metadata of the media container file is accessible in a bitstream using a supplemental-enhancement-information (SEI) message or from the container box.
18. The method of claim 1 , further comprising: encrypting a time-based segment of one of the at least one data stream with one of at least one time-specific encryption key,wherein the at least one time-specific encryption key is generated as one or more of the at least one encryption key.
19. The method of claim 18, wherein the time-based segment data stream is implemented using a Self-Decodable Access Point (SDAP) codec.
20. The method of claim 1 , further comprising: generating a streaming manifest comprising: data of the at least one object of interest; data of each respective representation; and the plurality of encryption key identifiers.
21. The method of claim 20, wherein the streaming manifest is configured to be decoded for streaming one or more of the at least one data stream based on data in the streaming manifest and an authorization of the at least one data stream.
22. The method of claim 1 , wherein the media container comprises metadata corresponding to revocation conditions for one or more of the at least one data stream.
23. The method of claim 1 , further comprising: identifying the at least one object of interest based on input from a user; and tracking the at least one object of interest to generate the at least one data stream.
24. The method of claim 23, wherein the at least one object of interest is identified by the user using natural language queries, selection of a particular area, instance segmentation, or combinations thereof; wherein the at least one object of interest is visual-based, audio-based, textbased, or combinations thereof;wherein the visual-based at least one object of interest is tracked using a Kalman filter algorithm, an optical flow algorithm, an Al algorithm, or combinations thereof; and wherein the audio-based at least one object of interest is tracked using audio source separation, speaker diarization, audio fingerprinting, feature matching, or combinations thereof.
25. The method of claim 1 , further comprising: encoding the at least one data stream, wherein the at least one data stream is video-based, audio-based, text-based, or combinations thereof; wherein the video-based at least one data stream is encoded using H.264 or HEVC codec; wherein the audio-based at least one data stream is encoded using AAC or MP3 codec; and wherein the at least one data stream comprises one or more mask images for separating the respective object of interest from background, and the one or more mask images are encoded using Run-Length Encoding (RLE), bit-plane encoding, PNG compression, GIF compression, H.264 codec, HEVC codec, or combinations thereof.
26. The method of claim 25, wherein the video-based at least one data stream is processed by binary mask generation, alpha channel encoding, color key filling, or combinations thereof to isolate the respective object of interest; and wherein the audio-based at least one data stream is processed by silence detection and removal, audio segmentation, background noise preservation, or combinations thereof to isolate the respective object of interest.
27. The method of claim 1 , further comprising:embedding, in a header of the media container, metadata for synchronizing the at least one data stream, the metadata comprising: frame number, decode timestamp, presentation timestamp, object identifier, spatial information, object timestamp, or combinations thereof; and synchronizing the at least one data stream.
28. The method of claim 4, further comprising: receiving a request for access of multimedia content; authenticating the request for access; determining at least one authorized data stream from the at least one data stream; retrieving at least one authorized encryption key identifier from the at least one encryption key identifier, each authorized encryption key identifier corresponding to a respective authorized data stream; retrieving at least one authorized encryption key corresponding to the at least one authorized encryption key identifier from the at least one encryption key from the key management service; and decrypting the at least one authorized data stream using the at least one authorized encryption key.
29. The method of claim 28, wherein the key management service is configured to return the at least one authorized encryption key based on the authenticating and the at least one authorized encryption key identifier.
30. The method of claim 28, wherein the key management service is configured to query the media container file for a revocation condition for the at least one authorized encryption key.31 . The method of claim 28, wherein the at least one authorized encryption key is retrieved as a wrapped encryption key wrapped by the key management service, andwherein the method further comprises: unwrapping the at least one authorized encryption key in a local channel session or for a predetermined maximum duration.
32. The method of claim 28, further comprising: rendering the at least one authorized data stream for viewing according to metadata corresponding to the at least one authorized data stream.
33. The method of claim 14, further comprising: receiving a content verification request corresponding for the plurality of data streams; verifying the at least one digital signature; recomputing one or more verification root digests, each corresponding to a respective root digest and recomputed using data corresponding to the respective root digest; verifying the at least one data stream by comparing the one or more verification root digests to the one or more root digests; and assigning a verification signature for each root digest.
34. A method for managing access to multimedia content, comprising: receiving a request for access of multimedia content, the multimedia content comprising a plurality of data streams, each data stream of the plurality of data streams is a segment of the multimedia content that comprises and corresponds to a respective representation of a respective object of interest in the multimedia content; authenticating the request for access, the request for accessing identifying one or more first data streams for access from the at least one data stream; retrieving one or more encryption key identifiers, each corresponding to a respective one of the first data streams;retrieving one or more encryption keys corresponding to the one or more first data streams from a key management service in response to authentication by the key management service based on the authenticating of the request for access and the one or more encryption key identifiers; and decrypting the one or more first data streams using the one or more encryption keys in response to the request for access.
35. A method of managing access to multimedia content, the method comprising: receiving a multimedia file; generating a plurality of data streams for each of at least one object of interest in the multimedia file such that each data stream of the plurality of data streams is a segment of the multimedia file that comprises and corresponds to a respective representation of a respective object of interest of the at least one object of interest; generating a plurality of encryption keys, each encryption key of the plurality of encryption keys corresponding to a respective data stream of the plurality of data streams; encrypting each of the plurality of data streams with a corresponding encryption key of the plurality of encryption keys, each encrypted data stream configured to be decrypted by the corresponding encryption key; generating a plurality of encryption key identifiers, each being a unique identifier corresponding to a respective encryption key; and generating a media container by packaging the plurality of data streams and the plurality of encryption key identifiers.
36. The method of claim 35, further comprising: generating a root digest for the plurality of data streams and corresponding metadata; and assigning a digital signature to the root digest for verification.
37. The method of 35, wherein the plurality of encryption keys are stored in a key management service secure repository; and wherein each of the plurality ofencryption keys is retrievable for decrypting the respective data stream from the key management service using a respective encryption key identifier and corresponding authentication data.
38. A method of managing access to multimedia content, the method comprising: requesting access, at an access control server, of one or more of a plurality of encrypted data streams each corresponding to a segment of a multimedia file and a representation of a respective object of interest; receiving authorization to the one or more encrypted data streams from the access control server; retrieving one or more encryption key identifiers from the one or more encrypted data streams, each encryption key identifier corresponding to a respective encrypted data stream; requesting one or more encryption keys corresponding to the one or more encryption key identifiers from a key management server based on the authorization and using the one or more encryption keys; receiving the one or more encryption keys; and decrypting the one or more encrypted data streams using the one or more encryption keys.
39. A system for managing access to multimedia content, the system comprising: one or more processing units configured to perform the method of any one of claims 1 to 38.
40. A non-transitory computer-readable medium having computer readable instructions stored thereon, which, when executed by at least one processor, causes the at least one processor to perform the method of any one of claims 1 to 38.
Citation Information
Patent Citations
Multiple stakeholder secure memory partitioning and access control
US20080109662A1
Method and system for transmitting end-user access information for multimedia content
US20090110059A1
System for controlling access and distribution of digital property
US20140123218A1
Cited By
Video transmission method and device for bridge detection, equipment and storage medium
CN121418550A