System and method for real-time artificial intelligence-based video compression and decompression

AI-based video compression and decompression systems with metadata encoding and parallel processing address bandwidth inefficiencies by enhancing compression efficiency and quality, enabling effective object detection and scale restoration.

WO2026009040A1PCT designated stage Publication Date: 2026-01-08FOUR DROBOTICS CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/051498
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-30
Filing Date
2025-02-13
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Video data transmission over networks with limited bandwidth, such as wireless networks, is inefficient due to high bandwidth consumption, and existing video compression and decompression technologies do not effectively utilize artificial intelligence for enhanced efficiency.

Method used

Implementing artificial intelligence-based video compression and decompression systems that utilize metadata encoding of compression parameters and parallel processing to enhance efficiency, utilizing trained models for object detection and pixel estimation in compressed frames.

Benefits of technology

The system reduces bandwidth usage by efficiently compressing and decompressing video data, enabling effective object detection and restoration to the original scale, improving video transmission quality and reducing network load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025051498_08012026_PF_FP_ABST
    Figure IB2025051498_08012026_PF_FP_ABST
Patent Text Reader

Abstract

What is disclosed is: A method for video reception comprising: receiving compressed video frames and metadata associated with the compressed video frames. The metadata comprises information related to selected objects of interest, and information related to a pre-compression video scale. The metadata is extracted from the video frames, and the above information is obtained. The method comprises performing a first sequence and a second sequence in parallel. The first sequence comprises: inputting the compressed video frames and the information related to the selected objects of interest to a first trained model, detecting the one or more objects of interest in the compressed frame, and generating the one or more objects of interest in the pre-compression video scale. The second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels. A plurality of new frames is generated based on outputs from the first and second sequences.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR REAL-TIME ARTIFICIAL INTELLIGENCE-BASED VIDEO COMPRESSION AND DECOMPRESSIONFIELD OF THE INVENTION

[0001] The present disclosure relates to video compression and decompression, in particular artificial intelligence and machine learning based video compression and decompression.BRIEF SUMMARY

[0002] A destination subsystem for video reception, wherein: the destination subsystem is communicatively coupled to a source subsystem via a network; the destination subsystem comprises: a reception module communicatively coupled to the network, a metadata extraction module, a destination storage module, wherein: the destination storage module stores one or more trained models, and a video decompression module, wherein: the reception module, destination storage module, video decompression module and metadata extraction module are communicatively coupled to each other via one or more destination interconnections; the reception module receives, from the source subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames, wherein: the metadata comprises: information related to selected objects of interest, and information related to a pre-compression video scale; the metadata extraction module: extracts the associated metadata, and obtains, based on the extracted metadata: the information related to the selected objects of interest, and the information related to the pre-compression video scale, the video decompression module performs a first sequence and a second sequence in parallel, wherein: the first sequence comprises: inputting the compressed video frames and the information related to the selected objects of interest to a first trained model obtained from the destination storage module, detecting, by the first trained model, one or more of the selected objects of interest in the compressed frame, and based on the detecting, generating the one or more of the selected objects of interest in the pre-compression video scale using the information related to the pre-compression video scale, the second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale, and the video decompression module generates a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0003] A method for video reception comprising: receiving, by a destination subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames, wherein: the associated metadata comprises: information related to one or more selected objects of interest, and information related to a pre-compression video scale; extracting, by the destination subsystem, the associated metadata; based on the extracted metadata, obtaining: the information related to the one or more selected objects of interest, and the information related to the pre-compression video scale; performing a first sequence and a second sequence in parallel, wherein: the first sequence comprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the selected one or more objects of interest in the pre-compression video scale using the information related to the pre-compression video scale, and the second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale; and generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0004] A method for video transmission comprising: capturing, by a source subsystem, a plurality of video frames at a pre-compression video scale; selecting one or more objects of interest from the captured plurality of video frames; generating information related to the objects of interest; compressing the captured plurality of video frames, further wherein: information related to the pre-compression video scale is stored; generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the pre-compression video scale; and transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destination subsystem.

[0005] A method for video transmission and reception comprising: capturing, by a source subsystem, a plurality of video frames at a pre-compression video scale; selecting one or more objects of interest from the captured plurality of video frames; generating information related to the objects of interest; compressing the captured plurality of video frames, further wherein: information related to the pre-compression video scale is stored; generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the pre-compression video scale; transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destinationsubsystem; receiving, by the destination subsystem, the compressed video frames and the associated metadata; extracting, by the destination subsystem, the associated metadata from the video frames; based on the extracting, obtaining: the information related to the selected one or more objects of interest, and the information related to the pre-compression video scale; performing a first sequence and a second sequence in parallel, wherein: the first sequence comprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting one or more of the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the one or more detected objects of interest in the pre-compression scale using the information related to the pre-compression video scale, and the second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale; and generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0006] The foregoing and additional aspects and embodiments of the present disclosure will be apparent to those of ordinary skill in the art in view of the detailed description of various embodiments and / or aspects, which is made with reference to the drawings, a brief description of which is provided next.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The foregoing and other advantages of the disclosure will become apparent upon reading the following detailed description and upon reference to the drawings.

[0008] FIG. 1 illustrates an example embodiment of a system which uses Al to perform video compression and decompression.

[0009] FIG. 2 illustrates an example embodiment of an Al subsystem.

[0010] FIG. 3 illustrates an example embodiment of a source subsystem.

[0011] FIG. 4 illustrates an example embodiment of a destination subsystem.

[0012] FIG. 5 illustrates an example embodiment of a process of training an Al model.

[0013] FIG. 6 illustrates an example embodiment of a process for video compression and transmission.

[0014] FIG. 7 shows an example embodiment of a process for video reception and decompression.

[0015] While the present disclosure is susceptible to various modifications and alternative forms, specific embodiments or implementations have been shown by way of example in the drawings and will be described in detail herein. It should be understood, however, that the disclosure is not intended to be limited to the particular forms disclosed. Rather, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of an invention as defined by the appended claims.DETAILED DESCRIPTION

[0016] Video data consumes relatively large amounts of bandwidth during transmission over networks, which may be problematic for networks having limited bandwidth such as wireless networks. Therefore, video compression prior to transmission is useful to reduce bandwidth usage for video transmission. In such systems, video decompression is then employed at the receiving end to enable a viewer to be able to view and use the information contained in the original video.

[0017] Artificial intelligence (Al) offers much potential to enhance video compression and decompression. In particular, parallel processing using different Al techniques can be used to enhance the efficiency of Al-based video decompression.

[0018] While the description below refers to Al, one of ordinary skill in the art would know that machine learning (ML) is a sub-field of Al, and therefore ML techniques, models and algorithms can be used where appropriate.

[0019] Compression parameters can be encoded into the metadata transmitted along with the compressed video. These compression parameters can then be extracted from the metadata at the receiving end and used in the parallel processing during decompression of the compressed video.

[0020] Patent Cooperation Treaty (PCT) Application Publication WO2023 / 084258, to Giataganas et al, filed 10 November 2021 and published on 19 May 2023, hereinafter referred to as “Giataganas ‘258”; and PCT Application Publication WO2023 / 084259, to Giataganas et al, filed 10 November 2021 and published on 19 May 2023, hereinafter referred to as “Giataganas ‘259”, concern compression of video of surgical procedures. While Giataganas ‘258 contemplates the use of metadata encoding and different Al techniques including Generative Adversarial Networks (GAN), Giataganas ‘258 and Giataganas ‘259 do not contemplate the use of parallel operations in decompression of the videos.

[0021] PCT Application Publication WO2023 / 183378, to Webb et al, filed 22 March 2023 and published on 28 September 2023 and hereinafter referred to as “Webb”, concerns end-to-endimage and video compression of hologram images. While Webb discloses the use of Al techniques such as GAN in video decompression, Webb does not contemplate the use of metadata encoding for video decompression.

[0022] Similarly, US Patent Application Publication 2024 / 0022759 to Asif et al, filed 30 July 2023 and published on 18 January 2024, and hereinafter referred to as “Asif’, discusses the use of GAN to identify salient features. However, like Webb, Asif does not contemplate the use of metadata encoding for video decompression.

[0023] A system and method which: performs Al-based video compression, utilizes metadata encoding of compression parameters, and uses parallel processing to perform Al-based video decompression efficiently, to overcome the shortcomings of the prior art is described below.

[0024] FIG. 1 shows an example embodiment of a system 100 which overcomes the shortcomings of the prior art. Al subsystem 101, source subsystem 103 and destination subsystem 105 are communicatively coupled to each other by one or more networks 107. In keeping with the above discussion with regard to ML being a sub-field of Al, one of ordinary skill in the art would understand that in some embodiments, Al subsystem 101 implements ML techniques, models and algorithms where appropriate. In some of these embodiments, Al subsystem 101 is an ML subsystem.

[0025] One or more networks 107 can be implemented using a variety of networking and communications technologies. In some embodiments, one or more networks 107 is implemented using wired technologies such as Firewire, Universal Serial Bus (USB), Ethernet and optical networks. In some embodiments, one or more networks 107 are implemented using wireless technologies such as WiFi, BLUETOOTH®, NFC, 3G, LIE and 5G. In some embodiments, one or more networks 107 are implemented using satellite communications links. In some embodiments, the communication technologies stated above include, for example, technologies related to a local area network (LAN), a campus area network, a controller area network (CAN), or a metropolitan area network (MAN). In yet other embodiments, one or more networks 107 are implemented using terrestrial communications links. In some embodiments, one or more networks 107 comprise at least one public network. In some embodiments, one or more networks 107 comprise at least one private network. In some embodiments, one or more networks 107 compriseone or more subnetworks. In some of these embodiments, some of the subnetworks are private. In some of these embodiments, some of the subnetworks are public. In some embodiments, communications within one or more networks 107 are encrypted.

[0026] Al subsystem 101 plays a variety of roles. These include but are not limited to: Storing and analyzing videos; and Implementation of Al-related operations related to videos as will be explained below.

[0027] An example embodiment of Al subsystem 101 is shown in FIG. 2. Al subsystem 101 comprises communications subsystem 201, compression / decompression (hereinafter referred to as “codec”) processing subsystems 203-1 to 203-N, Al interconnection 205 and database 207.

[0028] Communications subsystem 201 is coupled to one or more networks 107. Communications subsystem 201 receives information from, and transmits information to one or more networks 107. Communications subsystem 201 communicates using the communications and networking protocols and techniques that one or more networks 107 utilizes. Communications subsystem 201 receives information from one or more networks 107 within, for example, incoming signals 250; and transmits information to one or more networks 107 within, for example, outgoing signals 260. Examples of information within incoming signals 250 are data for training and testing such as historical data sets or user supplied commands, inputs and parameters which are used for operation of the Al subsystem 101. Examples of information within outgoing signals 260 are, Al model related information, parameters associated with trained Al models, results of training and results of testing.

[0029] Al interconnection 205 communicatively couples the various components of Al subsystem 101 to each other. In some embodiments, Al interconnection 205 is implemented using, for example, network technologies known to those in the art. These include, for example, satellite networks, wireless networks, wired networks, Ethernet networks, local area networks, metropolitan area networks and optical networks. In some embodiments, Al interconnection 205 comprises one or more subnetworks. In another embodiment, Al interconnection 205 comprises other technologies to connect multiple components to each other including, for example, buses, coaxial cables, and USB connections.

[0030] Codec processing subsystems 203-1 to 203-N perform the functionalities necessary to perform Al video compression and decompression using one or more programs and algorithms. These algorithms and programs are stored in, for example:• database 207, or• within codec processing subsystems 203-1 to 203 -N.

[0031] Examples of Al operations performed by codec processing subsystems 203-1 to 203 -N comprise:• Performance of Al -related operations such as: o pre-processing of video data sets prior to performing Al operations such as training and testing, o model training using training video data sets, o model testing using testing video data sets, o selecting appropriate models to use, o performance evaluation of different Al models, and o post-processing of video data sets;• In some embodiments, performing functions necessary to enable searching of database 207 such as implementation of appropriate data search algorithms;

[0032] Various implementations are possible for Al subsystem 101 and its components. In some embodiments, Al subsystem 101 is implemented using a cloud-based approach. In other embodiments, Al subsystem 101 is implemented across one or more facilities, where each of the components are located in different facilities and Al interconnection 205 is then a network-based connection. In further embodiments, Al subsystem 101 is implemented within a single server or computer. In yet other embodiments, Al subsystem 101 is implemented in software. In other embodiments, Al subsystem 101 is implemented using a combination of software and hardware. In yet other embodiments, Al subsystem 101 is hosted by a cloud services provider such as AMAZON® Web Services.

[0033] Databases 207 stores information and data for use by Al subsystem 101. This includes, for example:• one or more algorithms and programs necessary to perform Al -related operations and other operations, and• data needed for codec processing subsystems 203-1 to 203 -N to perform operations and functions. This comprises, for example: o training video data sets; o testing video data sets;o video data related to previous operations; o search histories; and o previous search results.

[0034] In some embodiments, database 207 further comprises a database server. The database server receives one or more commands from, for example, codec processing subsystems 203-1 to 203 -N and communication subsystem 201, and translates these commands into appropriate database language commands to retrieve and store data into databases 207. In one embodiment, database 207 is implemented using one or more database languages known to those of ordinary skill in the art, including, for example, Structured Query Language (SQL). In a further embodiment, database 207 stores data for a plurality of sources and destinations. Then, there may be a need to keep the set of data related to each source or destination separate from the data relating to the other sources or destinations. In some embodiments, database 207 is partitioned so that data related to each source or destination is separate from the other sources or destinations. In a further embodiment, when data is entered into database 207, associated metadata is added so as to make it more easily searchable. In a further embodiment, the associated metadata comprises one or more tags. In yet another embodiment, database 207 presents an interface to enable the entering of search queries. In yet another embodiment, database 207 requires authentication prior to entry and retrieval of data.

[0035] Videos are recorded at source subsystem 103 and compressed prior to transmission over one or more networks 107 to either destination subsystem 105 or Al subsystem 101. Source subsystem 103 can be located in, for example, buildings, vehicles, facilities and other locations where video data is recorded.

[0036] An example embodiment of source subsystem 103 is shown in FIG. 3. In FIG. 3, source subsystem 103 comprises: video capture module 301, video pre-processing module 303, source storage module 305, video compression module 307, metadata generation module 309, transmission module 311, source interconnection 313, andsource power supply module 315.

[0037] Video capture module 301 plays the role of capturing video using appropriate devices known to those of ordinary skill in the art comprising, for example, video cameras or imaging sensors. In some embodiments, video capture module 301 converts analog video to digital video format. In some embodiments video capture module 301 is implemented using a combination of hardware and software. In some embodiments, video capture module 301 is implemented using hardware.

[0038] Video pre-processing module 303 plays the role of pre-processing videos prior to compression. In some embodiments, video pre-processing module 303 is implemented using a combination of hardware and software. In some embodiments, video pre-processing module 303 is implemented using hardware.

[0039] Source storage module 305 plays the role of storing data and programs for operation of source subsystem 103. Examples of data and programs stored are: video frames that have been captured by video capture module 301 ; pre-processing algorithms used by, for example, video capture module 301 or video preprocessing module 303; trained Al models used in the operation of the components of source subsystem 103; parameters used by the pre-processing algorithms; video frames which have undergone pre-processing; compression algorithms used by video compression module 307; parameters used by the compression algorithms; compressed video frames; programs and algorithms used by metadata generation module 309; parameters to operate programs and algorithms used by metadata generation module 309; metadata generated by metadata generation module 309; and data and programs used in the operation of transmission module 311.In some embodiments, source storage module 305 is implemented using storage devices and technologies known to those of ordinary skill in the art.

[0040] Source storage module 305 can be implemented in a variety of ways. In some embodiments, source storage module 305 is implemented in a distributed fashion. In someembodiments, source storage module 305 is implemented as a plurality of portions communicatively coupled to each other via, for example, interconnections 313.

[0041] Then, in some of these embodiments, one or more of the plurality of portions is associated with each of the: video capture module 301, video pre-processing module 303, video compression module 307, metadata generation module 309, and transmission module 311.

[0042] In some of these embodiments, the one or more of the plurality of portions are integrated into each of the modules mentioned above.

[0043] In yet other embodiments, there is no source storage module 305. Instead, each of the modules mentioned above comprise storage sub-modules to store data used by each module.

[0044] Video compression module 307 plays the role of compressing videos using a video compression algorithm known to those of ordinary skill in the art. An example of such a video compression algorithm is the H.264 suite. As explained previously, in some embodiments, the algorithms and parameters to operate these algorithms are stored in source storage module 305. In yet other embodiments, the algorithms and parameters to operate these algorithms are stored in video compression module 307. In some embodiments, video compression module 307 is implemented using a combination of hardware and software. In some embodiments, video capture module is implemented using hardware.

[0045] Metadata generation module 309 plays the role of generating metadata for insertion into compressed video frames. As explained previously, in some embodiments, the programs and algorithms for metadata generation, and parameters to operate these programs and algorithms, are stored in source storage module 305. In yet other embodiments, the programs and algorithms and parameters to operate these programs and algorithms are stored in metadata generation module 309. In some embodiments, metadata generation module 309 is implemented using a combination of hardware and software. In some embodiments, metadata generation module 309 is implemented using hardware.

[0046] Transmission module 311 plays the role of transmitting compressed video frames and associated metadata over one or more networks 107 to either destination subsystem 105 or Alsubsystem 101. Transmission module 311 implements protocols and techniques necessary to enable transmission of videos and data over one or more networks 107. Similar to the other examples above, in some embodiments the data and programs used in the operation of transmission module 311 are stored in source storage module 305. In yet other embodiments, the data and programs used in the operation of transmission module 311 are stored in transmission module 311 itself.

[0047] Source interconnection 313 plays a similar role to Al interconnection 205 in FIG. 2, that is, source interconnection 313 communicatively couples the various components of source subsystem 103 to each other. In some embodiments, source interconnection 313 is implemented using, for example, network technologies known to those in the art. These include, for example, satellite networks, wireless networks, wired networks, Ethernet networks, local area networks, metropolitan area networks and optical networks. In some embodiments, source interconnection 313 comprises one or more subnetworks. In another embodiment, source interconnection 313 comprises other technologies to connect multiple components to each other including, for example, buses, coaxial cables and USB connections.

[0048] Source power supply module 315 supplies power to the various components of source subsystem 103 to enable their operation. In some embodiments, source power supply module 315 comprises one or more batteries implemented using battery technology known to those of ordinary skill in the art. In some embodiments, source power supply module 315 comprises renewable power sources such as solar cells.

[0049] Destination subsystem 105 receives videos transmitted from source subsystem 103 over one or more networks 107. Destination subsystem 105 can be located in, for example, buildings, control centres, or control rooms.

[0050] FIG. 4 shows an example embodiment of destination subsystem 105. In FIG. 4, destination subsystem 105 comprises: reception module 401 , metadata extraction module 403, destination storage module 405, video decompression module 407, video viewing module 409, destination interconnection 411, anddestination power supply module 413.

[0051] Reception module 401 plays the role of receiving the compressed videos and associated metadata transmitted over one or more networks 107 from source subsystem 103. Reception module 401 implements protocols and techniques necessary to enable reception of videos and data from one or more networks 107. Similar to the other examples above for the source subsystem 103, in some embodiments the data and programs necessary for the operation of reception module 401 are stored in destination storage module 405. In yet other embodiments, the data and programs necessary for the operation of reception module 401 are stored in reception module 401 itself.

[0052] Metadata extraction module 403 plays the role of extracting metadata which has been inserted into compressed video frames. Similar to the other examples above for the source subsystem 103, in some embodiments the data and programs necessary for the operation of metadata extraction module 403 are stored in destination storage module 405. In yet other embodiments, the data and programs necessary for the operation of metadata extraction module 403 are stored in metadata extraction module 403 itself. In some embodiments, metadata extraction module 403 is implemented using a combination of hardware and software. In some embodiments, metadata extraction module 403 is implemented using hardware.

[0053] Destination storage module 405 plays the role of storing data and programs for operation of destination subsystem 105. Examples of data and programs stored are: data and programs used in the operation of reception module 401 ; compressed video frames and associated metadata received by reception module 401 ; metadata extracted by metadata extraction module 403; information obtained from the extracted metadata; trained Al models used in the operation of the components of destination subsystem 105; and video frames generated by video decompression module 407.

[0054] In some embodiments, destination storage module 405 is implemented using storage devices and technologies known to those of ordinary skill in the art.

[0055] Similar to the source storage module 305, destination storage module 405 can be implemented in a variety of ways. In some embodiments, destination storage module 405 is implemented in a distributed fashion. In some embodiments, destination storage module 405 isimplemented as a plurality of portions communicatively coupled to each other via interconnections 411, and one or more of the plurality of portions is associated with each of the: reception module 401 , metadata extraction module 403, video decompression module 407, and video viewing module 409.

[0056] In some of these embodiments, the one or more of the plurality of portions are integrated into each of the modules mentioned above.

[0057] In yet other embodiments, there is no destination storage module 405. Instead, each of the modules mentioned above comprise storage sub-modules to store data used by each module.

[0058] Video decompression module 407 plays the role of decompressing transmitted compressed videos. In some embodiments, video decompression module 407 is implemented using a combination of hardware and software. In some embodiments, video decompression module 407 is implemented using hardware.

[0059] Video viewing module 409 allows viewers to play back decompressed videos. In some embodiments, video viewing module 409 comprises one or devices and equipment necessary to enable video playback, such as displays, screens and projectors. In some embodiments, video viewing module 409 comprises, for example, devices such as computers and tablets.

[0060] Destination interconnection 411 plays a similar role to Al interconnection 205 in FIG. 2 and source interconnection 313 in FIG. 3, that is, destination interconnection 411 communicatively couples the various components of destination subsystem 103 to each other. In some embodiments, destination interconnection 411 is implemented using, for example, network technologies known to those in the art. These include, for example, satellite networks, wireless networks, wired networks, Ethernet networks, local area networks, metropolitan area networks and optical networks. In some embodiments, destination interconnection 411 comprises one or more subnetworks. In another embodiment, destination interconnection 411 comprises other technologies to connect multiple components to each other including, for example, buses, coaxial cables and USB connections.

[0061] Destination power supply module 413 supplies power to reception module 401, metadata extraction module 403, destination storage module 405, video decompression module 407 and video viewing module 409. In some embodiments, destination power supply module 413comprises one or more batteries implemented using battery technology known to those of ordinary skill in the art. In some embodiments, destination power supply module 413 comprises renewable power sources such as solar cells.

[0062] In some embodiments, Al subsystem 101 is implemented as part of the destination subsystem 105.

[0063] As explained above, there are embodiments where components of Al subsystem 101, transmission subsystem 103 and destination subsystem 105 are implemented in hardware. In some embodiments, these hardware implementations are field programmable gate array (FPGA)-based. In yet other embodiments, some of these hardware implementations comprise the use of at least one processor.

[0064] FIG. 5 shows an example embodiment of a process of training an Al model for use in system 100, performed by Al subsystem 101. In step 501, pre-processing of a data set stored in database 207 prior to performing Al operations such as training and testing is performed by codec processing subsystems 203-1 to 203-N. The data sets which are stored in database 207 are obtained from, for example, historical data sets. Examples of historical data sets comprise, for example:Compressed videos,Decompressed videos,Metadata associated with videos, andUncompressed videos.

[0065] These historical data sets are, for example, those stored in at least one of source storage module 305 of source subsystem 103, and destination storage module 405 of destination subsystem 105.

[0066] Examples of pre-processing of data sets comprise: performing appropriate normalization operations on the received data sets; performing one or more data cleaning operations, such as: o filtering out outliers, o data scrubbing, o data validation, o data transformation, and o removing duplicate data, as well as irrelevant data; performing one or more labelling operations; andsplitting the data set into a training data set and a testing data set.

[0067] In some embodiments, the one or more labelling operations are performed on the data set using, for example, labels stored in database 207.

[0068] In some embodiments, as explained above, pre-processing of data sets comprises splitting the data set into a training data set and a testing data set. Specifically, different proportions of the data set are assigned to training and testing. For example, in some embodiments, 80% of the data set is assigned to the training data set while the remaining 20% is assigned to the testing data set. Techniques to assign data to training and testing are known to those of ordinary skill in the art and will not be discussed here. As would be known to one of ordinary skill in the art, in the embodiments where the pre-processing comprises one or more labelling operations, the splitting of the data set into a training data set and a testing data set occurs either before or after the one or more labelling operations are performed.

[0069] Then, as is known to those of ordinary skill in the art, the testing data set is used solely for model testing, while the training data is used to train the model. The testing and training data sets are then stored in, for example, database 207.

[0070] One of ordinary skill in the art would understand that different models can be used. Examples of models include but are not limited to: k-nearest neighbours, decision trees, naive Bayes, random forest, gradient boosting, logistic regression, support vector machine, neural networks, and- GANs.

[0071] In some embodiments, the models comprise one or more real-time object detection algorithms such as You Only Look Once (YOLO) as described in Redmon J, Diwala S, Girshick R, Farhadi A. “You only look once: Unified, real-time object detection”, Proceedings of the IEEE conference on computer vision and pattern recognition 2016 (pp. 779-788).

[0072] In step 503, after the data is pre-processed and split into training and testing data sets, model training is performed by codec processing subsystems 203-1 to 203 -N using the labelled features in the training data set. Techniques to perform model training are known to those of ordinary skill in the art and will not be discussed here. The parameters necessary to implement the model training techniques are retrieved from, for example, database 207. In some embodiments the results of the training operations are stored in, for example, database 207.

[0073] In step 505, after the model is trained, the model is tested by codec processing subsystems 203-1 to 203-N using the testing data set stored in database 207. In some embodiments, as part of step 505, one or more model performance metrics are calculated. Examples of model performance metrics comprise:Accuracy,- Fl,Sensitivity, andSpecificity.

[0074] The calculated one or more measures of model performance metrics are stored in, for example, database 207.

[0075] Based on these calculated measures, in step 507 a confidence score is calculated by codec processing subsystems 203-1 to 203-N. In some embodiments, this confidence score is stored in, for example, database 207.

[0076] The calculated confidence score is compared against a threshold by codec processing subsystems 203-1 to 203-N in step 509. When the threshold is met, the model and all parameters for usage of the model are then uploaded via Al interconnection 205, communications subsystem 201 and one or more networks 107 to, for example, one or more of source subsystem 103 and destination subsystem 105 in step 511. In some embodiments, the model and all parameters for usage of the model are uploaded to one or more of the source storage module 305 of source subsystem 103, and destination storage module 405 of destination subsystem 105, in step 511.

[0077] When the threshold is not met, then in step 513 one or more adjustments are made by codec processing subsystems 203-1 to 203-N to improve the confidence score. Examples of such actions comprise:Changing the model used,Changing the nature of the labelling, andChanging one of more parameters associated with the model training.

[0078] After the one or more adjustments are completed, the process returns to step 503, where: training (step 503), testing (step 505), calculation of the confidence score (step 507), and comparison against the threshold (step 509), are repeated until the threshold confidence score is met. Then, the model and all associated parameters are uploaded to the destination storage module 405 of destination subsystem 105 in step 511.

[0079] FIG. 6 shows an example embodiment of a process for video compression and transmission. In step 601, video frames are captured by video capture module 301 of source subsystem 103 and stored in source storage module 305 via source interconnections 313.

[0080] In step 603, one or more pre-processing operations are performed on the stored captured video frames. In the embodiments where there is a separate video pre-processing module 303: the stored captured video frames are retrieved from source storage module 305 via source interconnections 313 by video pre-processing module 303. Video pre-processing module 303 then performs at least one pre-processing operation on the stored captured video frames. In some embodiments, the at least one pre-processing operation comprises using a trained Al model retrieved from source storage module 305 to select objects of interest from the retrieved video frames, and then generating information related to the selected objects of interest. The information related to the selected objects of interest comprise, for example, features of the selected objects of interest. In some embodiments, the trained Al model used comprises one or more real-time object detection algorithms, as previously explained. In some embodiments, the one or more real-time object detection algorithms are used to perform the selecting. In some embodiments, the at least one pre-processing operation comprises the trained Al model detecting redundant information in the stored captured video frames, and removing the detected redundant information.

[0081] This information is stored in source storage module 305 via source interconnections 313. In other embodiments, step 603 is performed by video capture module 301. The video frames which have undergone pre-processing are stored, for example, in source storage module 305.

[0082] In step 605, video compression module 307 then retrieves pre-processed video frames from source storage module 305. The retrieved video frames are then compressed by video compression module 307 using a video compression algorithm retrieved from, for example, source storage module 305. The parameters used by the video compression algorithm are also retrieved from, for example, source storage module 305. In an example embodiment, video compression is performed using, for example, the H.264 video compression technique. In some embodiments, the video compression comprises down-scaling the video from a higher-resolution pre-compression to a lower-resolution post-compression. In some embodiments, this comprises down-scaling a video from a higher resolution such as a 1920 x 1080 resolution, an 8K resolution, or a 4K resolution, to a 640 x 480 resolution post-compression. The pre-compression scale information is then stored, for example, in source storage module 305, along with the compressed video frames.

[0083] In step 607, metadata generation module 309 retrieves: the information related to the selected objects of interest, the pre-compression video scale information, the compression algorithm used, the compressed video frames, and other data, which one of ordinary skill in the art would know is used in the generation of metadata. from source storage module 305. Metadata generation module 309 then generates metadata associated with the compressed video frames based on one or more of: the information related to the selected objects of interest, the pre-compression video scale information, the compressed video frames, and the other data, as explained above. In some embodiments, the information related to the selected objects of interest and the pre-compression video scale information is encoded into the metadata. The generated metadata is stored in, for example, source storage module 305.

[0084] In step 609, transmission module 311 then retrieves the compressed video frames together with the associated metadata from source storage module 305, wherein the metadata associated with each compressed frame is attached to that compressed video frame. Transmission module 311 then transmits the compressed video frames together with the associated metadataover one or more networks 107 to the destination subsystem 105. As explained previously, in some embodiments, transmission module 311 uses data and programs stored in source storage module 305 to perform this transmission.

[0085] FIG. 7 shows an example embodiment of a process for video reception and decompression at destination subsystem 105, using the model trained using the process previously described in FIG. 5. In step 701, reception module 401 receives the compressed video frames and the attached associated metadata from one or more networks 107 and stores the received compressed video frames in destination storage module 405.

[0086] In step 703, metadata extraction module 403: retrieves the stored compressed video frames from destination storage module 405, extracts the associated attached metadata from each video frame, obtains: o the information related to the selected objects of interest, o the pre-compression video scale information, and o the other data described above in step 607. from the extracted metadata, and stores this obtained information in destination storage module 405.

[0087] Sequence 705 comprising steps 707, 709 and 711; and sequence 725 comprising step 727; are performed in parallel to enable improved efficiency of operation.

[0088] Referring to sequence 705: In step 707, video decompression module 407 retrieves a trained Al model from destination storage module 405. The compressed video frames and information related to selected objects of interest, such as features of the selected objects of interest are retrieved from destination storage module 405 and input to the trained model by video decompression module 407.

[0089] In step 709, the selected objects of interest are detected in the compressed frame by video decompression module 407 using the trained Al model. In some embodiments, this comprises generating one or more bounding boxes of objects of interest.

[0090] In step 711, video decompression module 407 retrieves a trained Al model from destination storage module 405, uses this retrieved trained model to take the bounding boxes of objects of interest created in step 709 as an input, and generate objects in the original precompression video scale. As described above, the pre-compression video scale information wasobtained from the metadata in step 703. In some embodiments, this trained model is a GAN-based model. The output from sequence 705 is stored, for example, in destination storage module 405.

[0091] Referring to sequence 725: In step 727 video decompression module 407 iterates through every pixel on the frame and estimates the corresponding pixels based on the destination scale. Techniques to iterate and estimate pixels are known to those of ordinary skill in the art. In some embodiments, the iteration and estimation operations are performed using a trained Al model retrieved from destination storage module 405. The iterated and estimated pixels are then stored in, for example destination storage module. The output from sequence 725 is stored, for example, in destination storage module 405.

[0092] In step 729, the stored outputs from sequences 705 and 725 are retrieved. New video frames are generated by video decompression module 407 based on the retrieved outputs, and stored in destination storage module 405. In some embodiments, step 729 comprises video decompression module 407 performing post-processing operations on the generated video frames. The generated video frames are then stored, for example, in destination storage module 405.

[0093] In step 731, the video frames generated in step 729 are made available to video viewing module 409 for viewing.

[0094] In one example embodiment, a destination subsystem for video reception. The destination subsystem is communicatively coupled to a source subsystem via a network. The destination subsystem comprises: a reception module communicatively coupled to the network, a metadata extraction module, a destination storage module which stores one or more trained models, and a video decompression module. The reception module, destination storage module, video decompression module and metadata extraction module are communicatively coupled to each other via one or more destination interconnections. The reception module receives, from the source subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames. The metadata comprises: information related to selected objects of interest, and information related to a pre-compression video scale. The metadata extraction module extracts the associated metadata and obtains, based on the extracted metadata: the information related to the selected objects of interest, and the information related to the precompression video scale. The video decompression module performs a first sequence and a second sequence in parallel. The first sequence comprises: inputting the compressed video frames and the information related to the selected objects of interest to a first trained model obtained from thedestination storage module, detecting, by the first trained model, one or more of the selected objects of interest in the compressed frame, and based on the detecting, generating the one or more of the selected objects of interest in the pre-compression video scale using the information related to the pre-compression video scale. The second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale. The video decompression module generates a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0095] In one or more of the above examples, the generating of the one or more of the selected objects of interest in the pre-compression video scale is performed using a second trained model.

[0096] In one or more of the above examples, the second trained model is a GAN-based model.

[0097] In one or more of the above examples, the iterating and estimating are performed using a second trained model.

[0098] In one or more of the above examples, the generating of the plurality of new frames comprises one or more post-processing steps.

[0099] In one or more of the above examples, the destination subsystem comprises a video viewing module communicatively coupled to the reception module, destination storage module, video decompression module and metadata decoding module via the one or more destination interconnections. The generated plurality of new frames is made available to the video viewing module.

[0100] In one or more of the above examples, the destination subsystem is communicatively coupled to an Al subsystem via the network; and the Al subsystem comprises: one or more codec processing subsystems, a database, and a communications subsystem coupled to the network. The one or more codec processing subsystems, the database and the communications subsystem are coupled to each other via one or more Al interconnections.

[0101] In one or more of the above examples, the codec processing subsystems performs pre-processing of a data set stored in the database. The data set is split into a training data set and a testing data set, and one or more labelling operations are performed on the training data set and the testing data set using one or more labels stored in the database. The codec processing subsystem uses the training data set to perform training of one or more models. The codec processingsubsystem uses the one or more labels to perform the training. The codec processing subsystem tests the one or more trained models using the testing data set. The testing comprises calculating one or more model performance metrics. The codec processing subsystem calculates one or more confidence scores based on the calculated one or more performance metrics. The codec processing subsystem compares the calculated one or more confidence scores against one or more thresholds. Based on the comparing, the codec processing subsystem either: performs one or more adjustments, or uploads the one or more trained models and all parameters for usage of the one or trained models to the destination storage module.

[0102] In one example embodiment, a method for video reception. The method comprises: receiving, by a destination subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames. The associated metadata comprises: information related to one or more selected objects of interest, and information related to a precompression video scale. The method comprises extracting, by the destination subsystem, the associated metadata. Based on the extracted metadata: the information related to the one or more selected objects of interest, and the information related to the pre-compression video scale are obtained. The method comprises performing a first sequence and a second sequence in parallel. The first sequence comprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the selected one or more objects of interest in the pre-compression video scale using the information related to the pre-compression video scale. The second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the precompression video scale. The method comprises generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0103] In one or more of the above examples, the detecting of the selected one or more objects of interest is performed using the first trained model; and the detecting comprises generating one or more bounding boxes.

[0104] In one or more of the above examples, a second trained model uses the generated one or more bounding boxes as an input to perform the generating of the selected one or more objects of interest in the pre-compression video scale.

[0105] In one example embodiment, a method for video transmission. The method comprises capturing, by a source subsystem, a plurality of video frames at a pre-compression video scale. The method further comprises selecting one or more objects of interest from the captured plurality of video frames. The method further comprises generating information related to the objects of interest. The method further comprises compressing the captured plurality of video frames. Information related to the pre-compression video scale is stored. The method further comprises generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the precompression video scale. The method further comprises transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destination subsystem.

[0106] In one or more of the above examples, the generating of the metadata comprises encoding the information related to the selected objects of interest, and the information related to the pre-compression video scale, into the metadata.

[0107] In one or more of the above examples, the generating of the metadata is further based on the compressed video frames, and other data.

[0108] In one example embodiment, a method for video transmission and reception. The method comprises: capturing, by a source subsystem, a plurality of video frames at a precompression video scale. The method further comprises selecting one or more objects of interest from the captured plurality of video frames. The method further comprises generating information related to the objects of interest. The method further comprises compressing the captured plurality of video frames, further wherein: information related to the pre-compression video scale is stored. The method further comprises generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the pre-compression video scale. The method further comprises transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destination subsystem. The method further comprises receiving, by the destination subsystem, the compressed video frames and the associated metadata. The method further comprises extracting, by the destination subsystem, the associated metadata from the video frames. The method further comprises obtaining, based on the extracting, the information related to the selected one or more objects of interest, and the information related to the pre-compression video scale. The method further comprises performing a first sequence and a second sequence in parallel. The first sequencecomprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting one or more of the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the one or more detected objects of interest in the pre-compression scale using the information related to the pre-compression video scale. The second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale. The method further comprises generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

[0109] In one or more of the above examples, the generating of the metadata comprises encoding the information related to the selected objects of interest, and the information related to the pre-compression video scale, into the metadata.

[0110] In one or more of the above examples, the generating of the metadata is further based on the compressed video frames, and other data.

[0111] In one or more of the above examples, the detecting of the selected one or more objects of interest is performed using the first trained model; and the detecting comprises generating one or more bounding boxes.

[0112] In one or more of the above examples, the generating of the selected one or more objects of interest in the pre-compression video scale is performed using a second trained model. The second trained model uses the generated one or more bounding boxes as an input to perform the generating of the selected one or more objects of interest in the pre-compression video scale.

[0113] Although the algorithms described above including those with reference to the foregoing flow charts have been described separately, it should be understood that any two or more of the algorithms disclosed herein can be combined in any combination. Any of the methods, algorithms, implementations, or procedures described herein can include machine-readable instructions for execution by: (a) a processor, (b) a controller, and / or (c) any other suitable processing device. Any algorithm, software, or method disclosed herein can be embodied in software stored on a non-transitory tangible medium such as, for example, a flash memory, a CD- ROM, a floppy disk, a hard drive, a digital versatile disk (DVD), or other memory devices, but persons of ordinary skill in the art will readily appreciate that the entire algorithm and / or parts thereof could alternatively be executed by a device other than a controller and / or embodied in firmware or dedicated hardware in a well-known manner, for example, it may be implemented byan application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable logic device (FPLD), discrete logic, field programmable gate array (FPGAs). Also, some or all of the machine-readable instructions represented in any flowchart depicted herein can be implemented manually as opposed to automatically by a controller, processor, or similar computing device or machine. Further, although specific algorithms are described with reference to flowcharts depicted herein, persons of ordinary skill in the art will readily appreciate that many other methods of implementing the example machine readable instructions may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the blocks described may be changed, eliminated, or combined.

[0114] It should be noted that the algorithms illustrated and discussed herein as having various modules which perform particular functions and interact with one another. It should be understood that these modules are merely segregated based on their function for the sake of description and represent computer hardware and / or executable software code which is stored on a computer-readable medium for execution on appropriate computing hardware. The various functions of the different modules and units can be combined or segregated as hardware and / or software stored on a non-transitory computer-readable medium as above as modules in any manner, and can be used separately or in combination.

[0115] While particular implementations and applications of the present disclosure have been illustrated and described, it is to be understood that the present disclosure is not limited to the precise construction and compositions disclosed herein and that various modifications, changes, and variations can be apparent from the foregoing descriptions without departing from the spirit and scope of an invention as defined in the appended claims.

Claims

WHAT IS CLAIMED IS:

1. A destination subsystem for video reception, wherein: the destination subsystem is communicatively coupled to a source subsystem via a network; the destination subsystem comprises: a reception module communicatively coupled to the network, a metadata extraction module, a destination storage module, wherein: the destination storage module stores one or more trained models, and a video decompression module, wherein: the reception module, destination storage module, video decompression module and metadata extraction module are communicatively coupled to each other via one or more destination interconnections; the reception module receives, from the source subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames, wherein: the metadata comprises: information related to selected objects of interest, and information related to a pre-compression video scale; the metadata extraction module: extracts the associated metadata, and obtains, based on the extracted metadata: the information related to the selected objects of interest, and the information related to the pre-compression video scale,the video decompression module performs a first sequence and a second sequence in parallel, wherein: the first sequence comprises: inputting the compressed video frames and the information related to the selected objects of interest to a first trained model obtained from the destination storage module, detecting, by the first trained model, one or more of the selected objects of interest in the compressed frame, and based on the detecting, generating the one or more of the selected objects of interest in the pre-compression video scale using the information related to the precompression video scale, the second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale, and the video decompression module generates a plurality of new frames based on an output from the first sequence and an output from the second sequence.

2. The system of claim 1 , wherein the generating of the one or more of the selected objects of interest in the pre-compression video scale is performed using a second trained model.

3. The system of claim 2, wherein the second trained model is a generative adversarial network (GAN)-based model.

4. The system of claim 1, wherein the iterating and estimating are performed using a second trained model.

5. The system of claim 1, wherein the generating of the plurality of new frames comprises one or more post-processing steps.

6. The system of claim 1 , wherein: the destination subsystem comprises a video viewing module communicatively coupled to the reception module, destination storage module, video decompression module and metadata decoding module via the one or more destination interconnections; and the generated plurality of new frames is made available to the video viewing module.

7. The system of claim 1 , wherein: the destination subsystem is communicatively coupled to an Al subsystem via the network; and the Al subsystem comprises: one or more codec processing subsystems, a database, and a communications subsystem coupled to the network, wherein the one or more codec processing subsystems, the database and the communications subsystem are coupled to each other via one or more Al interconnections.

8. The system of claim 7, wherein: the codec processing subsystems performs pre-processing of a data set stored in the database, further wherein: the data set is split into a training data set and a testing data set, and one or more labelling operations are performed on the training data set and the testing data set using one or more labels stored in the database;the codec processing subsystem uses the training data set to perform training of one or more models, further wherein: the codec processing subsystem uses the one or more labels to perform the training; the codec processing subsystem tests the one or more trained models using the testing data set, further wherein: the testing comprises calculating one or more model performance metrics; the codec processing subsystem calculates one or more confidence scores based on the calculated one or more performance metrics; the codec processing subsystem compares the calculated one or more confidence scores against one or more thresholds; and based on the comparing, the codec processing subsystem either: performs one or more adjustments, or uploads the one or more trained models and all parameters for usage of the one or trained models to the destination storage module.

9. A method for video reception comprising: receiving, by a destination subsystem, a plurality of compressed video frames and metadata associated with the plurality of compressed video frames, wherein: the associated metadata comprises: information related to one or more selected objects of interest, and information related to a pre-compression video scale; extracting, by the destination subsystem, the associated metadata; based on the extracted metadata, obtaining: the information related to the one or more selected objects of interest, and the information related to the pre-compression video scale; performing a first sequence and a second sequence in parallel, wherein:the first sequence comprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the selected one or more objects of interest in the pre-compression video scale using the information related to the precompression video scale, and the second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale; and generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

10. The method of claim 9, wherein the detecting of the selected one or more objects of interest is performed using the first trained model; and the detecting comprises generating one or more bounding boxes.

11. The method of claim 10, wherein: a second trained model uses the generated one or more bounding boxes as an input to perform the generating of the selected one or more objects of interest in the pre-compression video scale.

12. A method for video transmission comprising:capturing, by a source subsystem, a plurality of video frames at a pre-compression video scale; selecting one or more objects of interest from the captured plurality of video frames; generating information related to the objects of interest; compressing the captured plurality of video frames, further wherein: information related to the pre-compression video scale is stored; generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the pre-compression video scale; and transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destination subsystem.

13. The method of claim 12, further wherein: the generating of the metadata comprises encoding: the information related to the selected objects of interest, and the information related to the pre- compress ion video scale, into the metadata.

14. The method of claim 12, wherein: the generating of the metadata is further based on: the compressed video frames, and other data.

15. A method for video transmission and reception comprising:capturing, by a source subsystem, a plurality of video frames at a pre-compression video scale; selecting one or more objects of interest from the captured plurality of video frames; generating information related to the objects of interest; compressing the captured plurality of video frames, further wherein: information related to the pre-compression video scale is stored; generating metadata associated with the compressed video frames based on: the information related to the selected objects of interest, and the information related to the pre-compression video scale; transmitting, from the source subsystem, the compressed video frames and the associated metadata to a destination subsystem; receiving, by the destination subsystem, the compressed video frames and the associated metadata; extracting, by the destination subsystem, the associated metadata from the video frames; based on the extracting, obtaining: the information related to the selected one or more objects of interest, and the information related to the pre-compression video scale; performing a first sequence and a second sequence in parallel, wherein: the first sequence comprises: inputting the compressed video frames and the information related to the selected one or more objects of interest to a first trained model, detecting one or more of the selected one or more objects of interest in the compressed frame, and based on the detecting, generating the one or more detected objects of interest in the pre-compression scale using the information related to the precompression video scale, andthe second sequence comprises: iterating through every pixel on the compressed frames, and estimating the corresponding pixels based on the pre-compression video scale; and generating a plurality of new frames based on an output from the first sequence and an output from the second sequence.

16. The method of claim 15, further wherein: the generating of the metadata comprises encoding: the information related to the objects of interest, and the information related to the pre- compress ion video scale, into the metadata.

17. The method of claim 15, wherein: the generating of the metadata is further based on: the compressed video frames, and other data.

18. The method of claim 15, wherein the detecting of the selected one or more objects of interest is performed using the first trained model; and the detecting comprises generating one or more bounding boxes.

19. The method of claim 18, wherein:the generating of the one or more detected objects of interest in the pre-compression video scale is performed using a second trained model; and the second trained model uses the generated one or more bounding boxes as an input to perform the generating of the selected one or more objects of interest in the pre-compression video scale.

Citation Information

Patent Citations

  • Video compression using neural networks

    US20210329306A1

  • Maintaining fixed sizes for target objects in frames

    US20210365707A1

  • Event / object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

    US20220059132A1

  • Video encoding through non-saliency compression for live streaming of high definition videos in low-bandwidth transmission

    US20240022759A1