Context-adaptive lossless coding of event frames

The context-adaptive lossless compression method using a deep-ternary-tree context-tree model efficiently encodes event frames, addressing inefficiencies in existing methods by replacing young models with mature ones, achieving superior coding performance and reduced runtime.

WO2025214577A1PCT designated stage Publication Date: 2025-10-16HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/059533
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing methods for encoding event frames from event cameras are inefficient, complex, or dependent on the performance of existing video codecs, and do not effectively handle high-resolution event data with sparse distributions.

Method used

A context-adaptive lossless compression method using a deep-ternary-tree context-tree model and a novel model search procedure to encode event frames efficiently, replacing young context models with mature models for improved coding efficiency.

Benefits of technology

Achieves improved coding performance of 34.34% compared to FLIF and 6.95% compared to prior solutions, with a significantly reduced encoding runtime, making it suitable for high-resolution event data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024059533_16102025_PF_FP_ABST
    Figure EP2024059533_16102025_PF_FP_ABST
Patent Text Reader

Abstract

A lossless compression method for encoding event camera frame sequences (1), wherein each event frame (8) of the event frame sequence (1) is traversed line by line trinary symbols (2) at a current position (10) representing event polarity are encoded using a probability distribution (15) corresponding to a context-tree-leaf (14) of a context-tree-model (13) selected using a deep trinary tree (12) of the causal context (11) of the current position (10) based on previously encoded trinary symbols (2). The method can further comprise a model search procedure for the encoding of the current trinary symbol (2) by determining whether a given context-tree-model (13) is young or mature based on previously encoded symbol history and replacing each young context-tree-leaf model (22) by the closest mature context-tree-leaf model (21).
Need to check novelty before this filing date? Find Prior Art

Description

[0001]CONTEXT-ADAPTIVE LOSSLESS CODING OF EVENT FRAMES TECHNICAL FIELD The disclosure relates in general to the field of digital image processing, and more specifically to a novel lossless compressionmethod for encoding synchronous or asynchronous event data obtained from event cameras.BACKGROUNDA new type of sensor, called event camera, was recently developed based on the new biomimetic technologies proposed in the neuromorphic engineering domain. The event camera is bio-inspired by the human brain as each pixel operates separately to mimic the biological neural system and performs simple tasks with a small energy consumption. More exactly, in contrast toconventional camera where all are operating simultaneously, the event camera sensor introduces a novel designed where eachpixel detects and reports independently only the changes (increased or decease) of the incoming light intensity above a threshold or remains silent otherwise. During operation, each pixel stores a reference brightness level, and continuously compares it to the current brightness level.If the difference in brightness exceeds a threshold, that pixel resets its reference level and generates an event: a discrete packetthat contains the pixel address, timestamp, and polarity (increase or decrease) of a brightness change, or an instantaneous measurement of the illumination level. Thus, event cameras output an asynchronous stream of events triggered by changes in scene illumination. As a result, only information about variable objects is included within the event stream delivered by a pixel event sensor, and there is no information about homogeneous surfaces or motionless backgrounds. In particular, event cameras provide the possibility of a very high temporal resolution as asynchronous events can be triggeredat the smallest timestamp distance of 10^^ s, which is equivalent to achieving a frame rate of up to 10^ frames per second(fps). Due to these advantages, event cameras can be efficiently applied in the technical fields of object recognition, autonomousvehicles, and robotics, among others. Two types of sensors currently available on the market are dynamic vision sensors (DVS), which capture sequences of asynchronous events; and dynamic and active-pixel vision sensors (DAVIS), which add a second camera to the DVS sensor, an active pixel sensor (APS) for capturing greyscale or color (RGB) frames. In certain circumstances, such as textured scenes with rapid motion, millions of events are generated per second. In order to process such busy scenes, existing event processes can require massive parallel computations. In the literature, the event data compression problem remains understudied. Only a few coding solutions are proposed to either encode asynchronous event sequences, or synchronous event frame (EF) sequences. In one prior art solution, the asynchronous event sequences are proposed to be encoded lossless by exploiting the spatial and temporal characteristics of the event location information. This proposed method is designed using three key strategies that are used to exploit the spatial and temporal characteristics of the spike location information for compression, including the adaptive macro-cube partitioning structure, the address-prior mode and the time-prior mode. The event information can be partitioned into multiple macro-cubes in temporal, each of which has the full spatial resolution of the pixel array. A macro-cube is then spatially split into small event-cubes. The encoding process for the event-cube consists of the event location coding, which isdone using the address-prior (AP) mode and time-prior (TP) mode, and the event polarity coding, which utilizes the previouscoding information as the reference to predict the imminent polarity. Finally, the computed residual error of the AP mode, TP mode, and polarity coding are fed into an adaptive context-based entropy coder. In the AP mode, the location histogram mapand location histogram counts are encoded. The location histogram map indicates wheatear an event exists or not, while thelocation histogram counts saves the number of events that occurred for each pixel. In the TP mode, the offsets of the spatial location to a central point is encoded using an entropy coder. The disadvantage of this proposed solution is that it employs complex coding techniques, and it is designed to deal only with raw asynchronous data, and thus not suitable to encode Event Frames. In another prior art solution, a so-called Time Aggregation-based Lossless Video Encoding for Neuromorphic vision sensor data (TALVEN) algorithm is proposed, where an event-accumulation process is employed to generate EFs by concatenatingthe positive and negative polarity EF counts and form an EF sequence which is then compressed by the known High EfficiencyVideo Coding (HEVC) standard. This method is designed to process DVS data by aggregating the events at fixed time intervals. In this case, the EFs are generated by recording the location histogram count, i.e., recording the number of event count at eachpixel. Four different spatio-temporal intervals are studied: 1, 2, 5, and 10 ms. The values are constrained by the number ofevents that can occur in the corresponding interval. The event sequence is divided into two separate frames, one associated topositive increase of luminance intensity and one for decrease of luminance, i.e., according to the vent polarity. The two framesare then merged in one single superframe composed of the “positive polarity” frame on the left and the “negative polarity” frame on the right. Finally, the superframes are encoded using the HEVC codec. The disadvantage of this proposed solution is that the method’s performance depends on the HEVC performance whileencoding the superframes. In the case of higher resolution sensors, e.g. 640 x 480, or 1280 x 720 pixels, the captured eventsare even more sparsely disturbed inside the frames. Therefore, the superframes are not efficiently encoded by HEVC, as the video coded is designed to encode a different type of information. Hence, one major disadvantage is that the method’s performance depends on the performance of the selected video codec. In a prior art application of the inventor, an efficient context-based lossless image codec for encoding event camera frames was introduced. The asynchronous event sequence acquired in a time-window generates an Event Frame (EF) using an event accumulation process, where the pixel-polarity is set as the sign of the event-pixel polarity sum. This results in a more efficient way to store the event data, where up to eight EFs are represented as a pair of an event map image (EMI), containing the spatial information, and a vector, containing the polarity information. The codec encodes the EMI using: (i) a binary map, which signals the positions where at least one event occurs in the EFs; (ii) the number of events for each signaled position; and (iii) their EF index. Template context modelling and adaptive Markov modelling are employed to encode these three types of data.The disadvantage of this proposed solution is that it only provides a performance-oriented lossless compression codec designedto encode a sequence of up to 8 event frames. In a further prior art solution, a low-complexity–oriented lossless compression codec was proposed for encoding event cameraframes and for integration in Event Signal Processing (ESP) chips which uses a fixed-length representation. In another solution,a novel low-complexity memory-efficient fixed-length representation of EFs was proposed using multi-level lookup tables, which is suitable for hardware implementation in very-low-power chips. In another solution, a coding solution was proposedfor encoding asynchronous event sequence and for integration in ESP chips. These methods were all designed to encode longerevent frame sequences. Considering the increasingly high frame rates that the event camera can achieve, there is a growing need for a method or system that provides a solution for efficient encoding of event data received from event cameras, and in particular for a solution that is suitable for lossless compression of event camera frame sequences. SUMMARYAccordingly, described herein are systems and methods of efficient event frame encoding for event cameras. Such systems andmethods can accumulate events received within a time period from an event camera in the form of an event stream, convert the asynchronous events into event frames, and efficiently encode the event frames in a lossless format for further processing by other applications. The foregoing and other objects are achieved by the features of the independent claims which disclose a novel method of encoding an event frame sequence using arithmetic coding. A first contribution proposes the use of a deep-ternary-tree of the current pixel position context as the context-tree model selector. The arithmetic codec encodes each trinary symbol using the probability distribution of the associated context-tree-leaf model. Another contribution proposes a novel context design based on several frames, where the context order controls the codec's complexity. Another contribution proposes a model search procedure to replace the context-tree prune-and-encode strategy by searching for the closest “mature” context model betweenlower-order context-tree models.According to a first aspect, there is provided a method for encoding an event frame sequence, each event frame comprising a trinary symbol at each pixel location representing event polarity by encoding the event frame sequence using arithmetic coding, wherein each event frame of the event frame sequence is traversed line by line and each trinary symbol at a current position is encoded using a probability distribution corresponding to a context-tree-leaf of a context-tree-model selected using a deeptrinary tree of the causal context of the current position based on previously encoded trinary symbols.This results in an efficient, context-adaptive compression method for lossless encoding of event frame sequences. Using deeptrinary trees for context modelling selection enables that the model's probability distribution encodes the current trinary symbol using a small codelength. The experimental evaluation shows that the proposed method provides an improved coding performance of 34.34% and a smaller runtime of up to 5.18x compared with state-of-the-art lossless image codec FLIF and, respectively, 6.95% and 14.42x compared with some prior solutions. In a possible implementation form of the first aspect the method further comprises the step of generating the event frame sequence by obtaining an asynchronous event stream, such as from an event camera, the event stream comprising asynchronousevents detected at pixel locations with associated polarity indicating variations in light intensity, and converting theasynchronous event stream into synchronous event frames by dividing the event stream into time intervals, and converting thedetected events at each pixel location over a given time interval to generate a trinary symbol at each pixel location of acorresponding event frame. In an embodiment the trinary symbols at each pixel location are generated by summing the polarities of detected events at eachpixel location over a given time interval.In a further possible implementation form of the first aspect the method comprises a context-tree modeling step to create a set of context-tree-models comprising context-tree-leafs, each context-tree-model comprising a specific symbol distribution modeled by the coding history of trinary symbols that occurred in the past in the causal context associated with the respective context-tree-model. In a further possible implementation form of the first aspect the causal context of a trinary symbol at the current position is based on previously encoded trinary symbols from neighboring pixel positions in the sequence of event frames, taking intoaccount both the current event frame and a predefined number of previous event frames. This results in a novel context designbased on several event frames, where the context order governs the codec's complexity for improved coding efficiency. In an embodiment the predefined number of previous event frames is 5.In an embodiment the deep trinary tree has a maximum tree-depth of 26 levels.In another embodiment the deep trinary tree has a maximum tree-depth of 20 levels.In a further possible implementation form of the first aspect, after encoding the trinary symbol, the respective context-tree- model model is updated to learn from the past. In a further possible implementation form of the first aspect, in order for the context-tree-model to retain only most relevant past information, an update of model frequency counts is performed if a trinary symbol frequency sum of a context-tree-model reaches a context frequency sum threshold: ∑f^ = τ^wherein the context frequency sum threshold is determined τ^ depending on the resolution of the event frames.In an embodiment said threshold is τ^ = 2^^of trinary symbols. In a further possible implementation form of the first aspect, the deep trinary tree is created by incorporating a number of trinarysymbols collected from the neighborhood of the current position, defining the context length, wherein up to six closest causalneighbors are selected from the current event frame and up to five closest causal neighbors are selected from previous event frames.In a further possible implementation form of the first aspect, the causal neighbors are selected based on a maximum distancecriterion, with a maximum distance of 2 pixels in the current event frame and a maximum distance of 1 pixel in previous eventframes.In a further possible implementation form of the first aspect, the causal neighbors are selected from current and previous eventframes according to at least one of a template context and a predefined context order.In a further possible implementation form of the first aspect, the method comprises a model search procedure by determiningwhether a given context-tree-model is young or mature based on previously encoded symbol history available for computing aprobability distribution, and for the encoding of the current trinary symbol, replacing each young context-tree-leaf model by the closest mature context-tree-leaf model, wherein parent nodes of a mature context-tree-leaf model are obtained by merging a set of neighboring young context-tree-leaf models.This model search procedure is proposed to replace the time-consuming context-tree prune-and-encode strategy by searchingfor the closest “mature” context model between lower-order context-tree models, i.e., a “young” context-tree-leaf model isreplaced by a “mature” context model, having an improved coding efficiency (“wisdom”) as the probability distribution iscomputed over a larger number of symbols thanks to lower-order context. In a further possible implementation form of the first aspect, a maturity of a given context-tree-model is determined according to a summed context count: wherein f^represents a minimum context count, and N^represents model number, and wherein κ represents tree depth of thecontext-tree-model.In a further possible implementation form of the first aspect, the maximum leaf-ancestor tree-depth ζ^ is limited with respectto the tree depth of the context-tree-model ζ^ < κ to limit complexity.In an embodiment the minimum context count is f^ = 32. In an embodiment, at each tree depth the counts of all subtrees of a node are added to compute the frequencies of the probability distribution associated with the respective node to reduce memory use. In a further possible implementation form of the first aspect the trinary symbols at each pixel location of a corresponding event frame are selected as "0", "1", or "2" based on the accumulated polarity p, initializing with “1” and accumulating all events triggered at the same pixel position by summing their polarity and set the trinary symbols based on the sign of the summed polarity as "0" for negative polarity p=-1, “1” for no event p=0, and “2” for positive polarity p=1. In an embodiment the probability distribution for each trinary symbol is computed using a Laplace estimator and fed to anarithmetic coder to encode the respective trinary symbol.According to a second aspect, there is provided a computer-based system comprising a processor coupled to a storage device;wherein the storage device comprises instructions that, when executed by the processor, cause the computer-based system toencode the sequence of event frames according to the method of any one of the possible implementation forms of the first aspect. The resulting system is surprisingly effective for encoding a plurality of event frames in an event frame sequence. According to a third aspect, there is provided a non-transitory computer readable medium having stored thereon programinstructions that, when executed by a processor, cause the processor to perform the method according to any one of the possibleimplementation forms of the first aspect.These and other aspects will be apparent from the embodiment(s) described below. BRIEF DESCRIPTION OF THE DRAWINGSIn the following detailed portion of the present disclosure, the aspects, embodiments and implementations will be explained inmore detail with reference to the example embodiments shown in the drawings, in which:Fig. 1 illustrates a method for encoding an event frame sequence in accordance with an example of the embodiments of thedisclosure;Fig. 2 illustrates converting an event stream into a plurality of event frames in accordance with an example of the embodimentsof the disclosure;Fig. 3 illustrates trinary symbols at each pixel location of a corresponding event frame in accordance with an example of theembodiments of the disclosure; Fig. 4 illustrates the causal context of a current position of an event frame based on previously encoded trinary symbols inaccordance with another example of the embodiments of the disclosure;Fig.5 illustrates the equations for context-tree-leaf model index selection based on context in accordance with another example of the embodiments of the disclosure;Fig. 6 illustrates a trinary tree of a given symbol length and selection of an associated model based on a given context inaccordance with another example of the embodiments of the disclosure; Fig.7 illustrates the equations of determining probability distribution for encoding a trinary symbol at a current position of anevent frame and updating the model in accordance with another example of the embodiments of the disclosure;Fig.8 illustrates a flow chart of updating the model frequency counts based on ternary symbol frequency sum in accordancewith another example of the embodiments of the disclosure;Fig. 9 illustrates a flow chart of a mature context model search procedure in accordance with another example of theembodiments of the disclosure; Fig. 10 illustrates a method of obtaining parent nodes of a mature context-tree-leaf model by merging a set of neighboringyoung context-tree-leaf models in accordance with an example of the embodiments of the disclosure;Figs. 11-13 illustrate equations for merging a set of neighboring young context-tree-leaf models in accordance with an exampleof the embodiments of the disclosure; Fig.14 illustrates merging a set of neighboring young context-tree-leaf models of a given trinary tree of a given symbol lengtha given context in accordance with an example of the embodiments of the disclosure;Fig. 15 shows average lossless compression results of methods in accordance with example embodiments of the disclosure fordifferent context-tree depths;Fig. 16 shows average lossless compression results of methods in accordance with example embodiments of the disclosurecompared to prior art methods; and Fig. 17 shows a table of average lossless compression rate and runtime results of methods in accordance with example embodiments of the disclosure compared to prior art methods. DETAILED DESCRIPTIONIn the following detailed description, numerous specific details are set forth by way of examples in order to provide a thoroughunderstanding of the relevant disclosure. However, it should be apparent to those skilled in the art that the present disclosuremay be practiced without such details. In other instances, well known methods, procedures, systems, and / or components havebeen described at a relatively high-level, without detail, in order to avoid unnecessarily obscuring aspects of the present disclosure. In the disclosure, some terms such as “ternary” and “trinary” which are understood in the technical field to refer to the same feature or quality are used interchangeably. The proposed method is referred in the following disclosure as Context-Adaptive Deep-Trinary-Tree, or CADeTT in short.Fig. 1 illustrates a method for encoding an event frame sequence 1 in accordance with an example of the embodiments of thedisclosure. Each event frame 8 of an event frame sequence 1 comprises a trinary symbol 2 at each pixel location 5 representingevent polarity 6, as shown in Fig. 2 and 3. The trinary symbol 2 is encoded at a current pixel position 10 by employing arithmeticcoding, where context-tree modelling based on deep-ternary-trees 12 is used to compute the symbol probability distribution 15. Hence, CADeTT uses context-tree modelling to learn to compute an efficient ternary symbol probability distribution basedon the prior encoded (decoded) symbol distribution of the context 11 of the current pixel position 10. In particular, each event frame 8 of the event frame sequence 1 is traversed line by line and each trinary symbol 2 at a current position 10 is encoded using a probability distribution 15 corresponding to a context-tree-leaf 14 of a context-tree-model 13 selected using a deep trinary tree 12 of the causal context 11 of the current position 10 based on previously encoded trinarysymbols 2, as will be explained in more detail in the disclosure below.The disclosure is organized as follows. Section A.1 presents the EF generation process. Section A.2 describes the proposed context-tree modelling process. Section A.3 describes the proposed “mature” context model search procedure. A.1 EVENT FRAME SEQUENCE GENERATIONIn some applications, the input data is sometimes received as an asynchronous event stream 3 which is then processed toproduce an output. However, in general, the event data is usually preferred as a synchronous image input, where the sequenceof asynchronous events is divided into spatio-temporal neighborhoods (i.e., time-volumes) of a time interval 9 resulting in synchronous event frames 8. Fig. 2 illustrates converting such an event stream 3 received from an event camera sensor of resolution W x H over a timeperiod comprising asynchronous events 4. An asynchronous event 4 is triggered at pixel location, (^, ^), ∀^ ∈ [1,  ^], ^ ∈[1,  ^], timestamp ^, with polarity ^, where ^ = +1 (or − 1) signals an increase (decrease) in light intensity. An event 4 isdenoted as ^ = (^,  ^,  ^,  ^). An asynchronous event sequence 1 collects all ^ events 4 triggered over time period ^ as ^^=.Event camera encapsulates all information using a 64 bits per event (bpev) representation.Due to large bitrate requirements and since image processing applications prefer to consume the event data as frames, anaccumulation process can be employed to transform the 1-D event sequence to a 2-D EF sequence using a time-window Δtused to control the EF sparsity and event information loss. The process starts by dividing ^^ into spatio-temporal volumes ofsize ^ × ^ × Δ, denoted ^^, as ^^ = In an exemplary embodiment, the sum-accumulation processis employed, where each generates an EF 8, denoted ^^ , as follows: initialize ^^ with ``1" and then all events 4 triggered atthe same pixel position 5, (^,  ^), are accumulated by summing their polarity, ^ = ∑ ^^ , and setting ^^(^,  ^), as a trinarysymbol 2 ^ ∈ {0,   1,   2}, where: ``0" signals ^ < 0 (^ = −1), ``1'' signals ^ = 0, and ``2" signals ^ > 0 (^ = 1), see Figure3. Hence, ^^ generates the synchronous EF sequence 1 ℱ = {^^} ^^^^^ , stored in memory using as raw data using a 2 bits perpixel (bpp) representation. A.2 CONTEXT-TREE MODELLING USING DEEP-TRINARY-TREES CADeTT employs a context-tree modelling process that consists in using available causal neighboring symbols 18, from the current event frame 16 and previous event frames 17, to create a causal context 11 with the role of dividing the set of already known (encoded / decoded) symbols 2 into a set of models 13. Each model 13 has a specific symbol distribution modelled bythe coding history of the symbols 2 that occurred in the past in the specific context 11 associated with the model 13. Since theEF 8 has a trinary symbol raw representation, each increase of the context length order further divides the set of known trinarysymbols 2 into a set of models 3 × larger. E.g., a zero-order model uses no past information, therefore, a single model is usedto compute the symbol probability distribution feed to the arithmetic coder. CONTEXT SELECTION In the following disclosure ^ represents the number of trinary symbols 2 collected from the current pixel neighbourhood, i.e.,the context length (tree-depth). Figure 4 illustrates a proposed exemplary method for current pixel context 11 selection fromthe neighborhood, where up to six closest causal neighbors 18 are selected from current event frame 16, ^^ , and up to fivecausal neighbors 18 are selected from previous event frames 17, ^^^^ , ^^^^ , ^^^^ . In an embodiment the predefined number ofprevious event frames 17 for context selection is limited at 5.The context order 20 in which the neighbors 18 are used to create the context-tree plays an important role. An exemplarycontext order 20 as illustrated in Fig.4 only selects the neighbors found at a position at a maximum distance of 2 pixels in thecurrent event frame 16 ^^ and 1 pixel in previous event frames 17.TRINARY-CONTEXT-TREEA trinary-context-tree 12 depicts all the possible symbol combinations that divide the set of known symbols into ^^ = 3^models, each associated with a tree-leaf. In the following disclosure ∈ {1,  2, … ,  3^} denotes the ^-th model at treedepth ^, where context ^^ of ^ trinary symbol length, denoted ^^ = [^^ ^^  … ^^], selects the modelin Fig.5: Figure 6 depicts the deep trinary tree 12 of ^ = 5 and how the causal context 11 ^^ = [1 1 1 1 0] selects the associated context-tree-model 13 ^^^^^. In an embodiment the deep trinary tree has a maximum tree-depth of 26 levels. In another embodiment the deep trinary tree has a maximum tree-depth of 20 levels. PROBABILITY DISTRIBUTIONIn the following disclosure ^^^^ ^^ denotes the frequency count and ^^^^^ denotes the probability of trinary symbol ^. In anexemplary embodiment, a Laplace estimator can compute the probability that ^^(^,  ^) is ^ as shown in Figure 7: The probability distribution can then be fed to the arithmetic coder to encode ^^(^,  ^). Next, the model ^^^is first updated by ^^ ^ ^increasing ^ ^^^^(^, ^)as ^^^^^ ← ^^^^^+ 1, so that the model can learn from the past. Furthermore, due to the arithmetic coder’s limited 16-bit precision, a second update can also be performed: when the ternary ymbol frequency sum ∑^^^s ^ (∑^^^^^^ ^^ ) reaches a threshold ^^ of ternary symbols, the model frequency count can be updateds f^ ←^^ ^where ^^ =∑^^ denotes probability distribution 15; or as ^^ ← ^^^^^a ^^^ ^^^^ ^^ + 1^ ≫ 1, where ``≫" denotes the shift-to-right operator, so the model can retain only most relevant past information.In an embodiment ^^ = 2^^. However, experiments show that ^^ = 2^^ provides improved results.A.3 “MATURE” CONTEXT MODEL SEARCH It can be noted that the compression of the first several trinary symbols 2 in each model 13 is inefficient as only a few symbols were triggered in the corresponding context 11, i.e., the model 13 is ``too young" as not enough information from symbol history is available to compute an efficient probability distribution 15. The deeper the trinary tree 12, the larger the number of contexts 11 and the number of symbols 2 that are needed to train the models 13. In prior art literature, to increase the coding gains, a prune-and-encode strategy is used. First, the image / data is traversed to collect the set of counts of every context tree note. Next, codelength estimations are computed for each node, and the context-tree is pruned to obtain the optimal context-tree by removing the child nodes that do not perform better than the parent node.Finally, the optimal context-tree is first encoded, and then the image / data using the optimal context-tree. Such a procedure istime-consuming and must be applied several times in the case of EF sequences 1. The coding performance depends on tree-depth and does not always provide the best performance due to additional optimal context-tree cost. In the method according to the examples in the disclosure, no pruning process is applied as only the set of models at maximum tree-depth, associated with the tree-leaves, are stored in memory. In the proposed method as shown in Fig.9, “young” context-tree-leaf models 22 are replaced with a “mature” context-tree-leaf model 21 as parent nodes 23 are obtained by merging the setof neighboring context-tree-leaf models 21. The coding efficiency is improved as the model's probability distribution 15 is computed over a more relevant list of symbols. Similar to tree pruning, the coding performance drops if the model obtained byreducing the context length ^ with ^ steps is ``too old" as too many symbols are collected. Although not shown in the figures, the present disclosure also extends to a computer-based system comprising an event cameraas described in the examples before, configured to record an event stream 3; and at least a processor coupled to a storage deviceand configured to convert the event stream 3 into a number of event frames 8 as described before. This storage device mayfurther comprise instructions that, when executed by the processor, cause the computer-based system to execute a methodaccording to any one of the above examples. The processor accordingly may process images and / or data relating to one or more functions described in the present disclosure. In some embodiments, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction-set processor (ASIP), a graphics processing unit (GPU), a physics processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), acontroller, a microcontroller unit, a reduced instruction-set computer (RISC), a microprocessor, or the like, or any combinationthereof. In some embodiments the processor may be further configured to control a display device of the computer-based system (not shown) to display any of the raw or encoded data received from the event camera 1. The display may include a liquid crystal display (LCD), a light emitting diode (LED)-based display, a flat panel display or curved screen, a rollable or flexible display panel, a cathode ray tube (CRT), or a combination thereof. In some embodiments the processor may be further configured to control an input device (not shown) for receiving a user input. An input device may be a keyboard, a touch screen, a mouse, a remote controller, a wearable device, or the like, or a combination thereof. The input device may include a keyboard, a touch screen (e.g., with haptics or tactile feedback, etc.), a speech input, an eye tracking input, a brain monitoring system, or any other comparable input mechanism. The input information received through the input device may be communicated to the processor for further processing. Another type ofthe input device may include a cursor control device, such as a mouse, a trackball, or cursor direction keys to communicatedirection information and command selections to, for example, the processor and to control cursor movement on a display device. The storage device can be configured to store data directly obtained from the event camera 1 and any further camera; and / or processed data from the processor. In some embodiments, the storage device may store images received from the respective cameras and / or processed images received from the processor with different formats including, for example, bmp, jpg, png, tiff, gif, pcx, tga, exif, fpx, svg, psd, cdr, pcd, dxf, ufo, eps, ai, raw, WMF, or the like, or any combination thereof. In some embodiments, the storage device may store algorithms to be applied in the processor, such as an encoding algorithm as described in any of the examples above. In some embodiments, the storage device may include a mass storage, a removable storage, a volatile read-and-write memory, a read-only memory (ROM), or the like, or any combination thereof. Exemplary mass storage may include a magnetic disk, an optical disk, a solid-state drive, etc. Although also not illustrated explicitly, the storage device and the processor can also be implemented as part of a server that isin data connection with the event camera 1 or a client device (such as an autonomous vehicle) that is connected to the eventcamera 1 through a network, wherein the client device can send event streams over the network to the server. The server maythen run the encoding algorithms as explained before. In some embodiments, the network may be any type of a wired or wirelessnetwork, or a combination thereof. Merely by way of example, the network may include a cable network, a wire line network,an optical fiber network, a telecommunication network, an intranet, an Internet, a local area network (LAN), a wide areanetwork (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), a public telephone switched network (PSTN), a Bluetooth network, a ZigBee network, a near field communication (NFC) network, or the like, or any combination thereof. The various aspects and implementations have been described in conjunction with various embodiments herein. However, other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed subject-matter, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measured cannot be used to advantage. A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. The reference signs used in the claims shall not be construed as limiting the scope. EXPERIMENTAL EVALUATION The experimental evaluation was performed over the ETH_Training dataset [M. Gehrig, W. Aarents, D. Gehrig and D. Scaramuzza, “DSEC: A Stereo Event Camera Dataset for Driving Scenarios,” IEEE Trans. Robot. Autom, vol.6, no.3, pp.4947-4954, Jul. 2021], called herein DSEC, of 82 event sequences, where ^ × ^ = 640 × 480 and the EF sequences weregenerated using Δ = 1^^. Due to memory requirements and time constraints, a maximum of 20,000 EFs were generated,corresponding to the first 20 seconds in the sequence.CADeTT-M (CADeTT) denotes the version with (without) the “mature" model search, with different context-tree-depth. E.g.,CADeTT−^17 employs CADeTT with ^ = 17; CADeTT-M−^20 employs CADeTT with “mature" model search and ^ =20, It was noted that CADeTT-M provides notable coding gains from ^ = 14.The following state-of-the-art methods are introduced for comparison: (1) the HEVC standard [G. J. Sullivan, J. Ohm, W. Han, and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEEE Trans. Circuits Syst. Video Technol., vol.22, no.12, pp.1649–1668, Dec.2012.] using the FFmpeg implementation [online source: http: / / ffmpeg.org / documentation.html]; (2) the VVC standard [B. Bross, J. Chen, J. -R. Ohm, G. J. Sullivan and Y. -K. Wang, “Developments in international video coding standardization after AVC with an overview of versatile video coding (VVC),” Proc. IEEE, vol.109, no.9, pp.1463- 1493, Sep.2021.] using the VVC Test Model [Available: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware VTM]; (3) the CALIC image codec [X. Wu and N. Memon, “Context-based, adaptive, lossless image coding,” IEEE Trans. Commun., vol.45, no.4, pp.437–444, Apr.1997.], and (4) the FLIF image codec [J. Sneyers and P. Wuille, “FLIF: Free lossless image format based on MANIAC compression,” Proc. IEEE Int. Conf. Image Process., Phoenix, AZ, USA, Sep.2016, pp.66–70.], [Available: https: / / flif.info].The methods encoded an intensity-like frame sequence, where each frame ^ merged five consecutive EFs as ^ = ∑^ ^^^ ^^3^. Moreover, the CADeTT results were compared with prior work from the inventor: (a) [I. Schiopu and R.C. Bilcu, “Lossless Compression of Event Camera Frames,” IEEE Signal Processing Letters, vol.29, pp. 1779-1783, 2022], where the proposed performance-oriented EF codec, LCEFC, encodes up to 8 EFs; and (b) [I. Schiopu and R.C. Bilcu, “Low-Complexity Lossless Coding for Memory-Efficient Representation of Event CameraFrames,” IEEE Sensors Letters, vol. 6, no. 11, pp. 1-4, Nov. 2022, Art no. 7004704], where the MER codec provides randomaccess to any group of pixels, e.g., 8 × 8 × 8, and the SAFE codec encodes a large number of EFs.The compression results were compared using the compression ratio (CR), defined as the ratio between raw data size (2 bit perpixel representation) and compressed size. Encoding runtime is reported as milliseconds per EF (^^ / ^^).In the experiments, the proposed method was tested under different configurations: by variating the context-tree-depth between0 and 26, and by imposing a maximum EF sequence length of 8 EF, 100 EF, and ^^^ available EFs.Figure 15 shows the average lossless compression results over the DSEC for CADeTT and CADeTT-M. When ^^^ availableEFs are encoded, the experiments show that: (i) after ^ = 20 the runtime performance is strongly affected by the large memorysize required to store the context model; (ii) after ^ = 17 the coding gains of CADeTT become neglectable; and (iii)CADeTT−^0 provides the fastest runtime of around 4 ^^ / ^^ or around 80 seconds per 20,000 EF sequence.When 100 or 8 EFs are encoded together, the experiments show that the proposed method is still able to provide good coding gains. Note that the use of a smaller context-tree-depth is recommended as not enough trinary symbols are encoded so that most of the contexts can compute efficient probability distributions. One can note that CADeTT offers the possibility to controlthe complexity using the context-tree-depth ^.Figures 16 and 17 show the average lossless compression results over DSEC. One can note that CADeTT−^0 has a similarperformance as SAFE (b), the prior low-complexity solution of the inventor, CADeTT−^17 provides the best bitrate-runtimetrade-off and CADeTT-M−^20 an improved coding performance of 34.34% compared with FLIF and 6.95% comparedwith LCECF (a), the prior performance-oriented solution of the inventor, and a runtime performance of 5.18 × smallercompared with FLIF and 14.42 × smaller compared with LCECF.CONCLUSIONS The proposed method according to the examples of the disclosure introduces an efficient context-adaptive lossless compression method for encoding event frame sequences 1. CADeTT proposes the use of a deep-ternary-tree 12 of the current pixel position context 11 as the context-tree model selector. CADeTT-M proposes a novel model search procedure for searching for the closest “mature” context model with improved coding efficiency with the probability distribution computed over a morerelevant list of symbols. The experimental evaluation shows an improved average performance of around 34.34% comparedwith FLIF and around 6.95% compared with the prior performance-oriented solution of the same inventor, while providing a much smaller encoding runtime. ACRONYMS AND ABBREVIATIONS ACRONYM FULL NAME EXPLANATIONCADeTT Context-Adaptive Deep-Trinary-Tree Proposed methodContext-Adaptive Deep-Trinary-Tree with mature CADeTT-M Proposed method model search procedure HEVC High Efficiency Video Coding Video Coding StandardVVC Versatile Video Coding Video Coding StandardCR Compression Ratio Compression metricDVS Dynamic Vision Sensor Event sensorsDAVIS Dynamic and Active-pixel Vision Sensor Type of Event SensorAPS Active Pixel Sensor Grayscale sensorESP chip Event Signal Processing chip Low-power event-processing chipEF Event Frame -Random Access -CALIC Context Adaptive Lossless Image Codec State-of-the-art methodFLIF Free Lossless Image Format State-of-the-art method

Claims

CLAIMS 1. A method for encoding an event frame sequence (1), each event frame (8) comprising a trinary symbol (2) at each pixel location (5) representing event polarity (6) by: encoding the event frame sequence (1) using arithmetic coding,wherein each event frame (8) of the event frame sequence (1) is traversed line by line and each trinary symbol (2) at a current position (10) is encoded using a probability distribution (15) corresponding to a context-tree-leaf (14) of a context-tree-model (13) selected using a deep trinary tree (12) of the causal context (11) of the current position (10) based on previously encoded trinary symbols (2).

2. The method of claim 1, further comprising the step of generating the event frame sequence (1) by: obtaining an asynchronous event stream (3), the event stream (3) comprising asynchronous events (4) detected at pixel locations (5) with associated polarity (6) indicating variations in light intensity, and converting the asynchronous event stream (3) into synchronous event frames (8) by dividing the event stream (3) into time intervals (9) and converting the detected events (4) at each pixel location (5) over a given time interval (9) togenerate a trinary symbol (2) at each pixel location (5) of a corresponding event frame.

3. The method of any one of claims 1 or 2, wherein the method comprises a context-tree modeling step to create a set of context- tree-models comprising context-tree-leafs (14), each context-tree-model (13) comprising a specific symbol distribution modeled by the coding history of trinary symbols (2) that occurred in the past in the causal context (11) associated with the respective context-tree-model (13).

4. The method of any one of claims 1 to 3, wherein the causal context (11) of a trinary symbol (2) at the current position (10) is based on previously encoded trinary symbols (2) from neighboring pixel positions in the sequence of event frames (8), taking into account both the current event frame (16) and a predefined number of previous event frames (17).

5. The method of any one of claims 1 to 4, wherein after encoding the trinary symbol (2), the respective context-tree-model (13) model is updated to learn from the past.

6. The method of any one of claims 1 to 5, wherein in order for the context-tree-model (13) to retain only most relevant pastinformation, an update of model frequency counts is performed if a trinary symbol (2) frequency sum of a context-tree-model (13) reaches a context frequency sum threshold: ∑^^ = ^^wherein the context frequency sum threshold is determined ^^ depending on the resolution of the event frames (8).

7. The method of any one of claims 1 to 6, wherein the deep trinary tree (12) is created by incorporating a number of trinary symbols (2) collected from the neighborhood of the current position (10), defining the context length, wherein up to six closest causal neighbors (18) are selected from the current event frame (16) and up to five closest causal neighbors (18) are selected from previous event frames (17).

8. The method of claim 7, wherein the causal neighbors (18) are selected based on a maximum distance criterion, with a maximum distance of 2 pixels in the current event frame (16) and a maximum distance of 1 pixel in previous event frames (17).

9. The method of any one of claims 7 or 8, wherein the causal neighbors (18) are selected from current and previous event frames (17) according to at least one of a template context (19) and a predefined context order (20).

10. The method of any one of claims 1 to 9, wherein the method comprises: a model search procedure by determining whether a given context-tree-model (13) is young or mature based on previously encoded symbol history available for computing a probability distribution (15), and for the encoding of the current trinary symbol (2), replacing each young context-tree-leaf model (22) by the closestmature context-tree-leaf model (21),wherein parent nodes (23) of a mature context-tree-leaf model (21) are obtained by merging a set of neighboring young context-tree-leaf models (22).

11. The method of claim 10, wherein maturity of a given context-tree-model (13) is determined according to a summed context count:wherein ^^ represents a minimum context count, and ^^ represents model number, and wherein ^ represents tree depth of thecontext-tree-model (13).

12. The method of any one of claim 10 or 11, wherein the maximum leaf-ancestor tree-depth ^^is limited with respect to thetree depth of the context-tree-model (13) ^^ < ^ to limit complexity.

13. The method of any one of claims 1 to 12, wherein the trinary symbols (2) at each pixel location (5) of a corresponding event frame (8) are selected as "0", "1", or "2" based on the accumulated polarity (6) p, initializing with “1” and accumulating all events (4) triggered at the same pixel position by summing their polarity (6) and set the trinary symbols (2) based on the sign of the summed polarity (6) as "0" for negative polarity (6) p=-1, “1” for no event p=0, and “2” for positive polarity (6) p=1.

14. A computer-based system comprising a processor coupled to a storage device; wherein the storage device comprises instructions that, when executed by the processor, cause the computer-based system to encode the sequence of event frames(8) according to the method of any one of claims 1 to 13.

15. A non-transitory computer readable medium having stored thereon program instructions that, when executed by a processor,cause the processor to perform the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Context-based lossless image compression for event camera

    WO2023160789A1