Event encoding method

By using run encoding in event vision sensors to compress event data, the problem of low compression efficiency in the prior art is solved, and more efficient data compression and processing is achieved.

CN120201205APending Publication Date: 2025-06-24OMNIVISION TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378932.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing event data encoding methods are not compressed efficiently, and they fail to fully utilize the sparseness and repetition of event data, resulting in high pressure on data transmission, storage and processing.

Method used

Run coding is used to compress and encode the polarity and/or timestamp of the event. Based on the event frame data format, the sparsity and repetition of the event data are used to further improve the compression efficiency.

Benefits of technology

It significantly reduces the amount of data, reduces the pressure on data transmission, storage and processing, and improves compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201205A_ABST
    Figure CN120201205A_ABST
Patent Text Reader

Abstract

The invention provides an event coding method, which comprises the following steps that: an event vision sensor acquires a plurality of original events at different positions, and each event comprises a polarity and a timestamp; the polarity is one of a positive event with increased brightness, a negative event with decreased brightness or a no event with brightness change smaller than a threshold value; and carrying out compression coding on the polarity and / or the timestamp by adopting run-length coding. According to the method, the event data is compressed based on the event frame data format and run length coding, the compression efficiency is further improved by utilizing the sparsity (a lot of data is 0) and repeatability of the event data, and the data volume is remarkably reduced, so that the pressure of data transmission, storage and processing is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of event signal compression encoding processing of event vision sensors, and particularly relates to an event encoding method. Background Art

[0002] With the development of society, the image acquisition field has higher and higher requirements for the quality of acquired images. In the field of motion image acquisition, an event vision sensor that collects data based on "event signals" can be used. The imaging device of the event vision sensor (EVS) has been widely used because of its fast data acquisition and low power consumption. The photodiode in the event vision sensor converts the sensed optical signal into an electrical signal and amplifies and outputs the change in the electrical signal caused by the change in brightness. When the change in the electrical signal is greater than a certain threshold, this change is defined as an "event signal", and only when the event signal occurs, that is, when the pixel senses a change in brightness, the sensor will generate a pulse and perform data communication.

[0003] The EVS can achieve high-speed data capture, thus generating a large amount of event data. If the original event data is not compressed and encoded, it will pose a great challenge to the transmission, storage, and processing of data.

[0004] Existing event data encoding will adopt some lossless encoding methods in the form of stream data or frame data. The compression efficiency of these solutions is still not high enough, and the characteristics of event data are not fully utilized to explore the potential of compression. Summary of the Invention

[0005] The purpose of the present invention is to provide an event encoding method, which uses run-length encoding to compress and encode the polarity and / or timestamp of events, compresses event data based on the event frame data format and run-length encoding, and further improves the compression efficiency by utilizing the sparsity (many data are 0) and repeatability of event data, significantly reducing the amount of data, thereby reducing the pressure on data transmission, storage, and processing.

[0006] The present invention provides an event encoding method, including:

[0007] An event vision sensor acquires a plurality of original events at different positions, and each of the events includes a polarity and a timestamp; the polarity is one of a positive event with increasing brightness, a negative event with decreasing brightness, or a non-event with a brightness change less than a threshold;

[0008] Use run-length encoding to compress and encode the polarity and / or the timestamp.

[0009] Further, the polarity is represented by two binary digits, and the non-event is represented by 00;

[0010] The positive event is represented by 11, and the negative event is represented by 10 or 01;

[0011] Alternatively, the negative event is represented by 11, and the positive event is represented by 10 or 01.

[0012] Further, all events within a preset first time window or reaching a first event quantity threshold are packaged together as an event frame; one event frame contains the positive event, the negative event, and the no-event;

[0013] The event frame includes a two-dimensional array of multiple rows and columns composed of two-bit binary numbers of the polarity of each event; wherein, the two-bit binary number of the polarity of one event in the same row serves as one column; four adjacent events form an event block in different combination ways, and one event block is one byte; after the event blocks are arranged in a certain scanning order, the event frame is divided by byte and forms a two-dimensional array after scanning and sorting.

[0014] Further, when the number of columns of the two-dimensional array is an integer multiple of 4, the total number of bytes in each row is 1 / 4 of the number of columns; when the number of columns is not an integer multiple of 4, one or more 00s are added at the end of the row or one or more binary numbers with unused polarity are added at the end of the row to make each row include an integer number of bytes.

[0015] Further, all positive events within a preset second time window or reaching a second event quantity threshold are packaged together as a positive event frame; one positive event frame includes the positive event and the no-event;

[0016] All negative events within a preset third time window or reaching a third event quantity threshold are packaged together as a negative event frame; one negative event frame includes the negative event and the no-event.

[0017] Further, each event in the positive event frame is represented by 1-bit binary number, the positive event is represented by 1; the no-event in the positive event frame is represented by 0; eight consecutive events in the positive event frame form a byte; a first marker bit is added to the header of the encoded file of the positive event frame;

[0018] Each event in the negative event frame is also represented by 1-bit binary number, the negative event is represented by 1; the no-event in the negative event frame is represented by 0; eight consecutive events in the negative event frame form a byte; a second marker bit is added to the header of the encoded file of the negative event frame;

[0019] Wherein, the first flag bit is 0 and the second flag bit is 1; or the first flag bit is 1 and the second flag bit is 0;

[0020] Both the positive event frame and the negative event frame are compressed using the run-length encoding.

[0021] Further, any all-zero row concentration area that appears in any one of the two-dimensional array after scan sorting, the positive event frame combined in bytes, and the negative event frame combined in bytes in the event frame can be compressed using the all-zero row run-length encoding; the all-zero row concentration area is distributed in 1 row or consecutive n rows, and all the binary numbers in each row are 0;

[0022] The run-length encoding of the all-zero row concentration area includes a first header information and first data. The first header information is set as an identification bit representing the all-zero row, and the value of the first data is n - 1.

[0023] Further, four adjacent events form an event block in different combination ways. The different combination ways specifically include: any one combination of four events located in 1 row and consecutive 4 columns, four events located in 2 rows and 2 columns, and four events located in 1 column and consecutive 4 rows.

[0024] Further, block run-length encoding is used for compression, and different event blocks are encoded in different scan orders.

[0025] Further, within the same event frame, an event block array is formed with the event blocks as units. The event block array is p rows and m columns, including:

[0026] The m column event blocks in the first row are K11, K12, K13... K1m in sequence;

[0027] The m column event blocks in the second row are K21, K22, K23... K2m in sequence;

[0028] The m column event blocks in the third row are K31, K32, K33... K3m in sequence;

[0029] The m column event blocks in the last row of p rows are Kp1, Kp2, Kp3... Kpm in sequence;

[0030] The event block array composed of each type of event block can be scanned in any of the following ways:

[0031] A: Scan from the beginning to the end of the row in the order of K11, K12, K21, K22, K13, K14, K23, K24 in sequence;

[0032] B: Scan from the beginning to the end of the row in the order of K11, K21, K12, K22, K13, K23 in sequence;

[0033] C: Scan from the beginning to the end of the row in the order of K11, K21, K31, K12, K22, K32, K13, K23, K33 in sequence;

[0034] D: Scan row by row from the first row to the last row;

[0035] E: Scan column by column from the first column to the last column.

[0036] Furthermore, in at least two or more event frames, event blocks at the same or similar spatial positions in each of the event frames are combined together and scanned in different scanning manners.

[0037] Furthermore, in at least two or more of the said event frames,

[0038] The first event frame includes event blocks 11, 12... 1n distributed in positions 11, 12... 1n one by one from left to right in sequence;

[0039] The second event frame includes event blocks 21, 22... 2n distributed in positions 21, 22... 2n one by one from left to right in sequence;

[0040] Perform scanning and encoding in the scanning order of event block 11, event block 21, event block 12, event block 22... event block 1n, event block 2n; or, perform scanning and encoding in a staggered up-and-down manner in the scanning order of event block 11, event block 22, event block 13, event block 24;

[0041] The third event frame includes event blocks 31, 32... 3n distributed in positions 31, 32... 3n one by one from left to right in sequence;

[0042] Perform scanning and encoding in the scanning order of event block 11, event block 21, event block 31, event block 12, event block 22, event block 32... event block 1n, event block 2n, event block 3n;

[0043] Among them, the corresponding position distributions of the first event frame, the second event frame, and the third event frame are the same, and the number of rows and columns of the corresponding event blocks is also the same.

[0044] Furthermore, in the two-dimensional array after the scanning sorting, in the first region, there are a consecutive repeated first values, and the corresponding run-length encoding includes a second header information and a second data. The second header information is set as an identification bit representing a, and the value of the second data is the first value; the first value is any integer in [0 - 255]; and

[0045] In the two-dimensional array after scanning and sorting, among c consecutive data in the second region, any two adjacent data are not repeated. The corresponding run-length encoding includes third header information and third data. The third header information is set as an identification bit representing c, and the third data takes the original c data without compression.

[0046] Further, a complete event frame is divided into a top frame and a bottom frame. The encoding of each row in the top frame starts with reserved first header information as a start flag, and the encoding of each row in the top frame ends with reserved second header information as an end flag;

[0047] The encoding of each row in the bottom frame starts with reserved third header information as a start flag, and the encoding of each row in the bottom frame ends with reserved fourth header information as an end flag.

[0048] Further, reserved fifth header information is used as a first-in-first-out almost full identification bit in both the top frame and the bottom frame; the data in the row with the first-in-first-out almost full identification bit is prone to errors or loss, and the correctness of the data in this row needs to be checked or the data in this row needs to be directly deleted;

[0049] Reserved sixth header information is used as an empty data identification bit to be filled in both the top frame and the bottom frame.

[0050] Further, reserved seventh header information is used to represent which frame of the image frame of the image sensor; reserved eighth header information is used to represent which frame of the event frame of the event vision sensor; to obtain image data and event data for the same scene at the same moment, and to strictly align the image data and the event data.

[0051] Further, in the two-dimensional array after scanning and sorting, the identification bit of the header information representing the number of consecutive repeated data is set as an integer in the range of (-127 to -1); the identification bit of the header information representing the number of consecutive non-repeated data is set as an integer in the range of (0 to 115); the identification bit of the reserved header information for special purposes is set as -128 and an integer in the range of (116 to 127).

[0052] Further, the event frame is not the original event frame, but a processed event frame formed by processing the event through at least one of accumulation, counting, weighted average, and linear interpolation;

[0053] The bit width of each processed event in the processed event frame is correspondingly extended to 8 bits or a higher bit width;

[0054] Alternatively, the latest or oldest events of each pixel within a period of time are taken to form a final event frame. At this time, the bit width of a single said event can still use 2 bits.

[0055] Furthermore, each said event is represented by four elements (x, y, p, t), where x and y represent spatial coordinates, p represents the polarity of the event as 1 or -1, and t represents the timestamp of the event;

[0056] Run - length encoding is used to compress - encode the timestamp. Specifically, when the timestamps of d events corresponding to d consecutive x - coordinates with equal y - coordinates are all equal, the corresponding run - length encoding includes header information representing d and timestamp data.

[0057] Furthermore, if the original event data is not output in frame format, the (x, y, p, t) stream format is converted into frame format for encoding. After conversion into frame format, the spatial coordinates do not need to be represented, and only (p, t) needs to be represented;

[0058] The compression encoding supports both the sensor being directly represented in event - frame format and being converted from other formats into the event - frame format; after having the event frame, event encoding is performed;

[0059] The method of building the event frame includes: scanning all the events. If the positions of multiple events are different, they can be placed in the same said event frame; if a new event appears at the same position, it is placed in the next event frame.

[0060] Furthermore, the run - length encoding is combined with at least one of entropy encoding and dictionary encoding for compression encoding.

[0061] Compared with the prior art, the present invention has the following beneficial effects:

[0062] The present invention provides an event encoding method, including: an event vision sensor acquires a plurality of original events at different positions, and each event includes a polarity and a timestamp; the polarity is one of a positive event with increasing brightness, a negative event with decreasing brightness, or a non - event with a brightness change less than a threshold; run - length encoding is used to compress - encode the polarity and / or the timestamp. The present invention is based on the event - frame data format and run - length encoding to compress event data, and utilizes the sparsity (many data are 0) and repeatability of event data to further improve the compression efficiency, significantly reduce the data volume, and thus relieve the pressure of data transmission, storage, and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic flowchart of an event encoding method according to an embodiment of the present invention.

[0064] Figure 2Schematic diagram of combining four adjacent events in the event encoding method of the present invention to form a byte.

[0065] Figures 3 to 6 Schematic diagram of the first to fourth different event blocks in the event encoding method of the present invention arranged in different scanning orders.

[0066] Figure 7 Comparison diagram of the final compression effects of different scanning orders of different event blocks in the event encoding method of the present invention.

[0067] Figure 8 Schematic diagram of the reserved header in the top frame TOP in the event encoding method of the present invention.

[0068] Figure 9 Schematic diagram of the reserved header in the bottom frame BOTTOM in the event encoding method of the present invention.

[0069] Figure 10 Schematic diagram of the reserved headers corresponding to the image sensor and the event vision sensor respectively in the event encoding method of the present invention.

[0070] Figure 11 Schematic diagram of the combined application of the reserved headers corresponding to the image sensor and the event vision sensor respectively in the event encoding method of the present invention.

[0071] Figure 12 Schematic diagram of using run - length encoding to compress and encode timestamps in the event encoding method of the present invention. Detailed implementation manners

[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non - precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.

[0073] For ease of description, some embodiments of the present application may use spatial relative terms such as "above", "below", "top", "bottom", etc. to describe the relationship between one element or component and another (or other) element or component as shown in the drawings of the embodiments. It should be understood that in addition to the orientations described in the drawings, spatial relative terms are also intended to include different orientations during the use or operation of the device. For example, if the device in the drawing is flipped, the element or component described as "below" or "beneath" other elements or components will subsequently be positioned "above" or "on top of" other elements or components. The terms "first", "second", etc. in the following text are used to distinguish between similar elements and are not necessarily used to describe a specific order or time sequence. It is to be understood that these terms may be replaced where appropriate.

[0074] An embodiment of the present invention provides an event encoding method, as Figure 1 shown, including:

[0075] Step S1, an event vision sensor acquires a plurality of original events at different positions, and each event includes a polarity and a timestamp; the polarity is one of a positive event with increasing brightness, a negative event with decreasing brightness, or a non-event with a brightness change less than a threshold;

[0076] Step S2, perform run-length encoding on the polarity and / or the timestamp for compression encoding.

[0077] The following details each step of the event encoding method of the embodiment of the present invention with reference to the accompanying drawings.

[0078] Step S1, an event vision sensor acquires a plurality of original events at different positions, and each event includes a polarity and a timestamp. Different from a traditional camera, an event vision sensor (event camera) does not capture images of the entire scene at a fixed frequency, but only records these changes instantaneously (usually at the microsecond level) when the pixel brightness in the scene changes. Each "event" is actually a very simplified data point, including the timestamp when the change occurs, the pixel position, and the direction of the brightness change (increase or decrease, that is, a positive event or a negative event). Therefore, an event camera can encode the motion in the scene very efficiently, especially in a high dynamic range or fast motion scene.

[0079] The polarity is one of a positive event with increasing brightness, a negative event with decreasing brightness, or a non-event with a brightness change less than a threshold; that is, each EVS pixel can only be in one of a positive event, a negative event, and a non-event. In one example, the polarity can be represented by two binary digits, the non-event is represented by 00; the positive event is represented by 11, and the negative event is represented by 10 or 01; or, the negative event is represented by 11, and the positive event is represented by 10 or 01.

[0080] Since the event camera outputs an asynchronous event stream without a fixed frame rate, a first time window or a first event quantity threshold is defined as needed, and all events within this first time window or reaching the first event quantity threshold are "packaged" together for processing. Such a "package" can be regarded as an "event frame", although it is essentially different from the frame of a traditional camera. This process is sometimes called "time surface" construction or "event accumulation", aiming to facilitate subsequent vision processing algorithms or networks to process these event data in a way similar to processing image frames. In this example, an event frame contains positive events, negative events, and non-events; the positive events and negative events are merged into one frame for run-length encoding.

[0081] As Figure 2As shown, the event frame includes a multi-row and multi-column two-dimensional array composed of two-bit binary numbers representing the polarity of each event; among them, the two-bit binary numbers representing the polarity of an event in the same row form a column; four adjacent events form an event block in different combination ways, and an event block is one byte; after the event blocks are arranged in a certain scanning order, the event frame is divided into units of bytes, and a two-dimensional array after scanning and sorting is formed. Figure 2 It shows that every four consecutive events in the same row are combined into a first event block; each row scans the first event block one by one from left to right.

[0082] When the number of columns of the two-dimensional array is an integer multiple of 4, the total number of bytes in each row is 1 / 4 of the number of columns; when the number of columns is not an integer multiple of 4, one or more 00s are added at the end of the row or one binary number with unused polarity is added at the end of the row so that each row includes an integer number of bytes.

[0083] In another example, the positive events and negative events can also be separated to form a frame each, which are the positive event frame and the negative event frame respectively. Specifically, all positive events within a preset second time window or reaching the second event quantity threshold are packed together as a positive event frame; a positive event frame includes positive events and no events. All negative events within a preset third time window or reaching the third event quantity threshold are packed together as a negative event frame; a negative event frame includes negative events and no events. Each event in the positive event frame is represented by a 1-bit binary number, and a positive event is represented by 1; no event in the positive event frame is represented by 0; every 8 consecutive events in the positive event frame are combined to form a byte; a first marker bit is added to the head of the encoded file of the positive event frame; each event in the negative event frame is also represented by a 1-bit binary number, and a negative event is represented by 1; no event in the negative event frame is represented by 0; every 8 consecutive events in the negative event frame are combined to form a byte; a second marker bit is added to the head of the encoded file of the negative event frame; among them, the first marker bit is 0 and the second marker bit is 1; or the first marker bit is 1 and the second marker bit is 0; both the positive event frame and the negative event frame are compressed using run-length encoding.

[0084] Step S2: Use run-length encoding to compress and encode the polarity and / or timestamp. The present invention is based on the event frame data format and run-length encoding to compress event data, and utilizes the sparsity (many data are 0) and repeatability of event data to further improve the compression efficiency and reduce the data volume. The idea of run-length encoding is to achieve compression by recording consecutive repeated data values and their repetition times. For example, the string "AAAABBBCCD" can be encoded as "4A3B2C1D". Each encoding pair consists of a head and data (HEAD+DATA).

[0085] Exemplarily, a specific allocation example of header information is provided. A part of the header (-127 to -1) is used to represent the pattern of data repetition (i.e., |HEAD| + 1 consecutive repeated data), another part of the header (0 to 115) is used to represent the pattern of different consecutive data (i.e., HEAD + 1 different data), and there is a reserved part of the header (-128, 116 to 127) for special purposes. One byte has 8 bits and can represent 256 numbers. Generally, the numbers represented by one byte are represented by unsigned 0 to 255. To clearly distinguish the pattern of data repetition and the pattern of different consecutive data, it is agreed that negative numbers (-127 to -1) are used to represent the pattern of data repetition, and non-negative numbers (0 to 115) are used to represent the pattern of different consecutive data. In this way, by looking at the positive or negative sign, it is possible to clearly know the data pattern. Using signed -128 to 127 to represent the header information with the same 256 numbers, with 0 corresponding to -128, is more intuitive. The specific numerical setting of the header information is not restricted and can be adjusted according to the situation.

[0086] For any one of the two-dimensional array after scanning and sorting in the event frame, the positive event frame combined in bytes, and the negative event frame combined in bytes, if there is a region with all-zero rows concentrated, run-length encoding for all-zero rows can be used for compression; the region with all-zero rows concentrated is distributed in 1 row or consecutive n rows, and each row has all binary numbers as 0; the run-length encoding of the region with all-zero rows concentrated includes the first header information and the first data. The first header information is set as the identification bit representing all-zero rows, and the value of the first data is n - 1. The first header information can occupy one byte or two or more (including two) bytes; the first data can occupy one byte or two or more (including two) bytes. For the situation where there may be a single row or consecutive multiple rows with all zeros in the event frame, the run-length encoding for all-zero rows of the present invention can be used for compression. For example, a row contains 2048 EVS pixels, and the value of each pixel is 0. After packing four consecutive events, the number of all-zero bytes is 2048 / 4 = 512 bytes. If conventional run-length encoding is used, a total of 8 bytes are required. At this time, a reserved header can be allocated for the all-zero row. The first header information is set as the identification bit representing 0, and the first header information is, for example, -128, and then followed by a value representing the number of all-zero rows. Generally, the number of all-zero rows starts counting from 0, and the actual number of all-zero rows is this value plus 1. When there is only one row with all zeros, the encoding is (-128, 0); when there are five consecutive rows with all zeros, the encoding is (-128, 4).

[0087] Block run-length encoding can be used for compression, and different event blocks are encoded according to different scanning orders.

[0088] Within the same event frame, four adjacent events form an event block in different combinations. The different combinations specifically include any one of the following: four events located in one row and four consecutive columns (1x4), four events located in two rows and two columns (2x2), and four events located in one column and four consecutive rows (4x1).

[0089] Within the same event frame, an event block array is formed with event blocks as units. The event block array is p rows and m columns, including:

[0090] The m column event blocks in the first row are K11, K12, K13... K1m in sequence;

[0091] The m column event blocks in the second row are K21, K22, K23... K2m in sequence;

[0092] The m column event blocks in the third row are K31, K32, K33... K3m in sequence;

[0093] The m column event blocks in the last row, the pth row, are Kp1, Kp2, Kp3... Kpm in sequence;

[0094] The event block array composed of each type of event block can be scanned in any of the following ways:

[0095] A: Sequentially scan from the beginning to the end of the row in the order of K11, K12, K21, K22, K13, K14, K23, K24;

[0096] B: Sequentially scan from the beginning to the end of the row in the order of K11, K21, K12, K22, K13, K23;

[0097] C: Sequentially scan from the beginning to the end of the row in the order of K11, K21, K31, K12, K22, K32, K13, K23, K33;

[0098] D: Scan row by row from the first row to the last row;

[0099] E: Scan column by column from the first column to the last column.

[0100] Figures 3 to 6 Shows some coding embodiments formed by combining different event blocks and different scanning orders (other combinations can also be derived, not limited). Here, x represents the basic event block (four events just form one byte), and xyzk represents the scanning order of the event block. Run-length encoding is performed on consecutive bytes according to the given scanning order.

[0101] The repeatability of event data within the same event frame belongs to two-dimensional coding in space. Exemplarily, the first type: such as Figure 3As shown, every four consecutive events in the same row are combined into a first event block (1x4); each row is scanned one by one from left to right for the first event block. The second method: As Figure 4 shown, in two adjacent rows, two adjacent events in the upper row and two events in the directly lower row are combined into a second event block (2x2); every two adjacent rows are scanned one by one from left to right for the second event block. The third method: As Figure 5 shown, in two adjacent rows, every four adjacent events in the upper row or every four adjacent events in the directly lower row are combined into a third event block (2x4); in the column composed of the third event blocks, each column is scanned column by column from top to bottom for the third event block. The fourth method: As Figure 6 shown, in four adjacent rows, two adjacent events in the first row and two events in the directly lower second row, or two adjacent events in the third row and two events in the directly lower fourth row are combined into a fourth event block (4x2); in the column composed of the fourth event blocks, each column is scanned column by column from top to bottom for the fourth event block. The fifth method: In four adjacent rows, four events in the same column are combined into a fifth event block; each column is scanned column by column from top to bottom for the fifth event block.

[0102] Figure 7 shows the compression ratios of run-length encoding for the same given event frames under different event blocks and different scanning orders. The vertical axis represents the compression ratio, which is the amount of data after compression / the amount of data before compression; the horizontal axis represents the number of frames. Generally Figure 6 in the fourth type (4x2) of , that is, for the basic event block of 2x2, a better compression ratio can be obtained by scanning four consecutive rows in the form of 2x2 event blocks first from top to bottom and then returning to the top of the event block in the next column and scanning downwards (the lower the compression ratio, the better the compression efficiency). Figure 7 In , m represents the specific average compression ratio value.

[0103] The present invention also considers the time dimension and adopts run - length encoding or other encoding between multiple event frames. Among at least two or more event frames, event blocks at the same or similar spatial positions in each event frame are combined together and scanned in different scanning manners. The probability that the polarities or timestamps of event blocks at the same or similar spatial positions are relatively close or the same is high. Exemplarily, among at least two or more event frames, the first event frame includes event blocks 11, 12... 1n distributed one - to - one from left to right at positions 11, 12... 1n respectively; the second event frame includes event blocks 21, 22... 2n distributed one - to - one from left to right at positions 21, 22... 2n respectively; scanning and encoding are performed in the scanning order of event blocks 11, 21, 12, 22... 1n, 2n; or, scanning and encoding are performed in a staggered up - and - down scanning order in the scanning order of event blocks 11, 22, 13, 24. The third event frame includes event blocks 31, 32... 3n distributed one - to - one from left to right at positions 31, 32... 3n respectively; scanning and encoding are performed in the scanning order of event blocks 11, 21, 31, 12, 22, 32... 1n, 2n, 3n; wherein, the corresponding position distributions of the first event frame, the second event frame, and the third event frame are the same, and the number of rows and columns of the corresponding event blocks are also the same.

[0104] A specific allocation example of a header information, where a part of the header (-127 to -1) is used to represent the pattern of data repetition (i.e., |HEAD| + 1 consecutive repeated data). In the two - dimensional array after scanning and sorting, if there are a consecutive repeated first values in the first region, the corresponding run - length encoding includes a second header information and a second data. The second header information is set as the identification bit representing a, for example, the second header information is -(a - 1), and the value of the second data is the first value; the first value is any integer in [0 - 255]; both the second header information and the second data each occupy one byte. Exemplarily, if there are 9 consecutive repeated 120s in the first region, the corresponding second header information is -8, that is, the identification bit representing 9 is -8; the value of the corresponding second data is 120. The second header information can occupy one byte or two or more (including two) bytes; the second data can occupy one byte or two or more (including two) bytes.

[0105] A specific allocation example of a header information. Another part of the header (0 to 115) is used to represent different patterns of continuous data (i.e., HEAD + 1 different data). In the scanned and sorted two-dimensional array, among the continuous c data in the second region, any two adjacent data are not repeated, that is, c different data. The corresponding run-length encoding includes a third header information and a third data. The third header information is set as an identification bit representing c. For example, the third header information is c - 1, and the third data takes the original c data without compression.

[0106] Other important information may be embedded in the encoded events. To facilitate the distinction from the conventional run-length encoding, reserved headers are used for special marking. Figures 8 to 11 Shows the usage scenarios of different reserved headers. In a specific allocation example of a header information, the identification bits of the reserved header information for special purposes are set as integers within the range of -128 and (116 to 127) for example. The allocation and use of the reserved headers are relatively flexible, as long as the given allocation rules are followed for use.

[0107] Such as Figure 8 and Figure 9 As shown, a complete event frame is divided into a top frame TOP and a bottom frame BOTTOM. Figure 8 In it, the encoding of each row in the top frame TOP starts with a reserved first header information (such as 124) as the start flag, and the encoding of each row in the top frame ends with a reserved second header information (such as 125) as the end flag. Figure 9 In it, the encoding of each row in the bottom frame BOTTOM starts with a reserved third header information (such as 122) as the start flag, and the encoding of each row in the bottom frame ends with a reserved fourth header information (such as 123) as the end flag. In both the top frame and the bottom frame, a reserved fifth header information (such as 126) is used as the first-in first-out (FIFO) almost full identification bit; the data of the row with the first-in first-out almost full identification bit is prone to errors or loss, and the correctness of the data in this row needs to be checked or the data in this row can be directly deleted. In both the top frame and the bottom frame, a reserved sixth header information (such as 127) is used as the identification bit for empty data to be filled.

[0108] Such as Figure 10 and Figure 11 As shown, a reserved seventh header information (such as 121) is used to represent the ID of the image frame of the image sensor (CIS); a reserved eighth header information (such as 120) is used to represent the ID of the event frame of the event vision sensor (EVS); to obtain the image data of the image sensor and the event data of the event vision sensor for the same scene at the same moment, so that the image data and the event data are strictly aligned.

[0109] In addition to the typical embodiments mentioned in the above inventions, there are other different embodiments, which are specifically reflected in aspects such as the representation of events / event frames, the organization and scanning order of event blocks, the encoding of other relevant data besides event polarity, the fusion of run-length encoding and other encoding methods, etc. The embodiments may be flexible combinations of the above aspects.

[0110] The representation of an event frame is not limited to being split into a Top frame and a Bottom frame. It can also be a complete frame. The event frame may not be the original event frame either, but a processed event frame formed by processing the events through at least one of accumulation, counting, weighted average, and linear interpolation; the bit width of each processed event in the processed event frame is correspondingly extended to 8 bits or a higher bit width (usually an integer byte); or, the latest or oldest events of each pixel within a period of time are taken to form the final event frame, and at this time, the bit width of a single event can still be 2 bits. For example, in an original event frame, 8 original events occur within 10 ms. The polarity of each original event is used as the minimum object for compression encoding. Event accumulation processing includes, for example, taking the sum of the polarity values of each of the 8 original events as the minimum object for compression encoding, and the 8 original events constitute a processed event in the processed event frame. Event counting processing includes, for example, taking the number of original events occurring within every 10 ms as the minimum object for compression encoding, and the value (number) of the first processed event in the processed event frame is 8.

[0111] The event frame may also be converted from other formats, such as Figure 12 shown. Each row in the table is an event, and each event is represented by four elements (x, y, p, t), where x and y represent spatial coordinates, p represents the polarity of the event (1 / -1), and t represents the timestamp of the event. For this case, the polarity and timestamp of the events need to be represented in the format of a frame according to the spatial coordinates, and the polarity and timestamp at positions where no event occurs are both 0.

[0112] Run-length encoding is used to compress the timestamp. Specifically, when the timestamps of d events corresponding to d consecutive x coordinates with equal y coordinates are all equal, the corresponding run-length encoding includes header information representing d and the timestamp data. For example Figure 12 in, the 4 timestamps of 899877 ns, the 7 timestamps of 899879 ns, and the 5 timestamps of 899880 ns can all be compressed using run-length encoding including header information and timestamp data.

[0113] If the original event data is not output in a frame format, the (x, y, p, t) stream format is converted into a frame format for encoding. After conversion into the frame format, the spatial coordinates do not need to be represented, and only (p, t) needs to be represented; the compression encoding supports both the direct representation of the sensor as an event frame format and the conversion from other formats to the event frame format; after the event frame is obtained, event encoding is performed; the method for constructing the event frame includes: scanning all the events. If the positions of multiple events are different, they can be placed in the same event frame; if a new event appears at the same position, it is placed in the next event frame.

[0114] The present invention can perform compression encoding by using run-length encoding in combination with at least one of entropy encoding and dictionary encoding. In addition to run-length encoding, other encoding methods can be further applied after run-length encoding to improve the compression efficiency; the present invention is a lossless encoding. Run-length encoding is based on the statistics of the continuous data distribution (independent of the probability model), and entropy encoding depends on the probability model. These two types of encoding are independent of each other. Entropy encoding specifically includes Huffman encoding, arithmetic encoding, interval encoding, and context-adaptive binary arithmetic encoding. Huffman encoding belongs to a type of entropy encoding.

[0115] The present invention will be applied to products related to event vision sensors (EVS), including independent EVS products and hybrid sensor products that integrate traditional CIS (CMOS image sensor) and EVS. These products can be used in multiple fields such as mobile phones, vehicles, security, and industry.

[0116] In summary, the present invention provides an event encoding method, including: an event vision sensor acquires a plurality of original events at different positions, and each event includes a polarity and a timestamp; the polarity is one of a positive event with increasing brightness, a negative event with decreasing brightness, or a non-event with a brightness change less than a threshold; run-length encoding is used to perform compression encoding on the polarity and / or the timestamp. The present invention is based on the event frame data format and run-length encoding to compress event data, and utilizes the sparsity (many data are 0) and repeatability of event data to further improve the compression efficiency, significantly reduce the data volume, and thus relieve the pressure of data transmission, storage, and processing.

[0117] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the methods disclosed in the embodiments, since they correspond to the devices disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0118] The above description is only a description of the preferred embodiments of the present invention, and does not limit any scope of the rights of the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and decorations made to the above embodiments according to the technical essence of the present invention without departing from the technical solution of the present invention all fall within the protection scope of the technical solution of the present invention.

Claims

1. An event coding method, characterized in that: include: The event vision sensor acquires a plurality of raw events at different locations, each of the events including polarity and time stamp; The polarity is one of a positive event of brightness increase, a negative event of brightness decrease, or no event of brightness change less than a threshold value; Run-length coding is used to compress and encode the polarity and / or the time stamp.

2. The event encoding method according to claim 1, characterized in that: The polarity is represented by two binary digits, and the no event is represented by 00; The positive event is represented by 11, and the negative event is represented by 10 or 01; Alternatively, the negative event is represented by 11, and the positive event is represented by 10 or 01.

3. The event encoding method according to claim 2, characterized in that: Packing all the events within a preset first time window or reaching a first event quantity threshold together as an event frame; One of the event frames includes the positive event, the negative event and the no event; The event frame includes a two-dimensional array of multiple rows and columns composed of two-bit binary numbers of the polarity of each event; wherein, the two-bit binary number of the polarity of an event in the same row is a column; four adjacent events form an event block in different combinations, and an event block is a byte; after the event blocks are arranged in a certain scanning order, the event frame is divided into bytes to form a two-dimensional array after scanning sorting.

4. The event encoding method according to claim 3, characterized in that: When the number of columns of the two-dimensional array is an integer multiple of 4, the total number of bytes in each row is 1 / 4 of the number of columns; when the number of columns is not an integer multiple of 4, one or more 00s are added to the end of the row or one or more unused binary numbers of the polarity are added to the end of the row so that each row includes an integer number of bytes.

5. The event encoding method according to claim 1, characterized in that: Packing all the positive events within a preset second time window or reaching a second event quantity threshold together as a positive event frame; one positive event frame includes the positive event and the no event; Packing all the negative events within a preset third time window or reaching a third event quantity threshold together as a negative event frame; One negative event frame includes the negative event and the no event.

6. The event encoding method according to claim 5, characterized in that: Each event in the forward event frame is represented by a 1-bit binary number, the forward event is represented by 1; no event in the forward event frame is represented by 0; eight consecutive events in the forward event frame form a byte; the header of the encoded file of the forward event frame is added with a first marker bit; Each event in the negative event frame is also represented by a 1-bit binary number, the negative event is represented by 1; no event in the negative event frame is represented by 0; a combination of 8 consecutive events in the negative event frame constitutes a byte; the header of the encoded file of the negative event frame is added with a second mark bit; The first mark bit is 0, and the second mark bit is 1; or the first mark bit is 1, and the second mark bit is 0; The positive event frame and the negative event frame are both compressed by using the run-length coding.

7. The event coding method according to claim 3 or 6, characterized in that: The all-0 row concentrated area in any one of the two-dimensional array after scan sorting in the event frame, the positive event frame after combination in byte units, and the negative event frame after combination in byte units can be compressed by using all-0 row run-length coding; the all-0 row concentrated area is distributed in 1 row or n consecutive rows, and the binary numbers in each row are all 0; The run-length encoding of the area where all zero rows are concentrated includes first header information and first data, the first header information is set to an identification bit representing an all-0 row, and the value of the first data is n-1.

8. The event encoding method according to claim 3, characterized in that: The four adjacent events form an event block in different combinations, and the different combinations specifically include: any combination of the four events located in 1 row and 4 consecutive columns, the four events located in 2 rows and 2 columns, and the four events located in 1 column and 4 consecutive rows.

9. The event encoding method according to claim 3, characterized in that: Block run-length coding is used for compression, and different event blocks are encoded in different scanning orders.

10. The event encoding method according to claim 8, characterized in that: In the same event frame, an event block array is formed with the event blocks as units. The event block array has p rows and m columns, including: The event blocks in the first row and m columns are K11, K12, K13…K1m; The event blocks in the second row and m columns are K21, K22, K23…K2m; The event blocks in the third row and m columns are K31, K32, K33…K3m; The event blocks of the last row p and m columns are Kp1, Kp2, Kp3…Kpm; The event block array composed of each event block is scanned in any of the following ways: A: Scan from the beginning to the end of the line in the order of K11, K12, K21, K22, K13, K14, K23, K24; B: Scan from the beginning to the end of the line in the order of K11, K21, K12, K22, K13, and K23; C: Scan from the beginning to the end of the line in the order of K11, K21, K31, K12, K22, K32, K13, K23, K33; D: Scan the entire line from the first line to the last line line by line; E: Scan the entire column from the first column to the last column one by one.

11. The event encoding method according to claim 3, characterized in that: In at least two or more of the event frames, event blocks at the same or similar spatial positions in each of the event frames are combined together and scanned in different scanning modes.

12. The event encoding method according to claim 11, characterized in that: In at least two or more of the event frames, The first event frame includes event blocks 11, event blocks 12, ... event blocks 1n which are distributed one by one at positions 11, 12, ... 1n from left to right; The second event frame includes event blocks 21, event blocks 22, ... event blocks 2n which are distributed one by one at positions 21, 22, ... 2n from left to right; Scanning and encoding are performed in the scanning order of event block 11, event block 21, event block 12, event block 22 ... event block 1n, event block 2n; or, scanning and encoding are performed in an up-and-down staggered manner in the scanning order of event block 11, event block 22, event block 13, event block 24; The third event frame includes event blocks 31, event blocks 32, ... event blocks 3n which are distributed one by one at positions 31, 32, ... 3n from left to right; Scanning and encoding are performed in the scanning order of event block 11, event block 21, event block 31, event block 12, event block 22, event block 32 ... event block 1n, event block 2n, event block 3n; The corresponding positions of the first event frame, the second event frame and the third event frame are distributed in the same manner, and the numbers of rows and columns of corresponding event blocks are also the same.

13. The event encoding method according to claim 3, characterized in that: In the two-dimensional array after the scan sorting, there are a first values ​​that are continuously repeated in the first area, and the corresponding run-length encoding includes second header information and second data, the second header information is set to represent the identification bits of a, and the value of the second data is the first value; the first value is any integer in [0-255]; and In the two-dimensional array after the scan sorting, there are c consecutive data in the second area and any two adjacent data are not repeated. The corresponding run-length encoding includes third header information and third data. The third header information is set to an identification bit representing c data, and the third data takes the original c data without compression.

14. The event encoding method according to claim 3, characterized in that: A complete event frame is divided into a top frame and a bottom frame, the encoding of each line in the top frame starts with the reserved first header information as a start mark, and the encoding of each line in the top frame ends with the reserved second header information as an end mark; The encoding of each line in the bottom frame starts with the reserved third header information as a start mark, and the encoding of each line in the bottom frame ends with the reserved fourth header information as an end mark.

15. The event encoding method according to claim 14, characterized in that: The top frame and the bottom frame both use the reserved fifth header information as a FIFO almost full flag; the data of the row with the FIFO almost full flag is prone to error or loss, and the correctness of the row data needs to be checked or the row data needs to be directly deleted; The top frame and the bottom frame both use the reserved sixth header information as the empty data identification bit to be filled.

16. The event encoding method according to claim 14, characterized in that: The reserved seventh header information is used to indicate the image frame number of the image sensor; the reserved eighth header information is used to indicate the event frame number of the event vision sensor; To obtain image data and event data for the same scene at the same time, so that the image data and the event data are strictly aligned.

17. The event encoding method according to claim 14, characterized in that: In the two-dimensional array after the scan and sorting, the identification bit of the header information representing the number of consecutive repeated data is set to an integer in the range of (-127 to -1); the identification bit of the header information representing the number of consecutive non-repeating data is set to an integer in the range of (0 to 115); the identification bit of the reserved header information representing special purposes is set to an integer in the range of -128 and (116 to 127).

18. The event encoding method according to claim 3, characterized in that: The event frame is not an original event frame, but a processed event frame formed by processing the event by at least one of accumulation, counting, weighted averaging and linear interpolation; The bit width of each processing event in the processing event frame is correspondingly expanded to 8 bits or higher; Alternatively, the latest or oldest event of each pixel within a period of time is taken to form a final event frame, and at this time, the bit width of a single event can still use 2 bits.

19. The event encoding method according to claim 1, characterized in that: Each event is represented by four elements (x, y, p, t), where x and y represent spatial coordinates, p represents the polarity of the event as 1 or -1, and t represents the timestamp of the event; The timestamp is compressed and encoded using run-length coding, specifically including: when the y coordinates are equal and the timestamps of d events corresponding to consecutive d x coordinates are all repeated and equal, the corresponding run-length coding includes header information representing d identification bits and timestamp data.

20. The event encoding method according to claim 19, characterized in that: If the original event data is not output in a frame format, the (x, y, p, t) stream format is converted into a frame format for encoding. After conversion into a frame format, the spatial coordinates do not need to be represented, and only (p, t) is required; The compression encoding supports both direct representation of the sensor into the event frame format and conversion from other formats into the event frame format; After the event frame is obtained, event coding is performed; The method of constructing the event frame includes: scanning all the events, and if the positions of multiple events are different, they can be placed in the same event frame; If a new event occurs at the same location, it is placed in the next event frame.

21. The event encoding method according to claim 1, characterized in that: The run-length coding is combined with at least one of entropy coding and dictionary coding to perform compression coding.