A method and system for compact binary serialization of native type data and zero-copy read-write

CN122817136APending Publication Date: 2026-09-25LINGSHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610993984.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0029]1、本发明通过面向原生类型的紧凑编码方式,减少类描述、字段名、对象包装信息等冗余内容;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817136A_ABST
    Figure CN122817136A_ABST
Patent Text Reader

Abstract

The application discloses a kind of compact binary serialization and zero-copy read-write method and system of native type data, it is related to computer data processing and data storage technical field, specifically includes: according to the field order, byte sequence and coding analysis rule determined by pre-set structure description, signed integer or long integer is executed positive and negative interlaced mapping and variable-length integer coding, float type is written original binary bit, according to variable-length coding, write array length and continuously write element original byte to native array, converge into compact binary data stream;Data stream is written in byte buffer according to field order, record effective data length and form data frame;When receiving submission instruction, data frame is written in channel or file buffer in batches;When reading data frame, restore native scalar according to coding analysis rule;For the native array that element is continuously written according to fixed byte width, create native type buffer view of shared storage area, read array element through view and output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer data processing and data storage technology, specifically to a compact binary serialization and zero-copy read / write method and system for native data types. Background Technology

[0002] In scenarios involving data transmission, file storage, network communication, and cache persistence, basic data such as integers, long integers, floating-point numbers, and native arrays frequently need to be serialized, written, and read between memory, files, and network channels. Existing common serialization methods often process object structures, typically carrying class descriptions, field names, object headers, wrapper class information, text delimiters, or other format identifiers when saving or transmitting data. For large amounts of similar native data, these additional elements enter the data stream along with the actual numerical data, easily reducing the proportion of effective data and increasing storage space usage and network transmission burden.

[0003] In terms of read and write processing, traditional byte stream methods typically input and output bytes one at a time. Raw data types need to be split into multiple bytes before writing and then reassembled into the corresponding type after reading. When processing large amounts of data continuously, frequent small-granularity reads and writes increase the number of system calls. At the same time, raw arrays often need to be copied into temporary byte arrays during deserialization and then converted into arrays or collections of objects of the corresponding type, resulting in additional memory usage, data copying, and object creation overhead.

[0004] Furthermore, some serialization frameworks rely on structure description files, reflection, dynamically generated code, or runtime type adaptation mechanisms. For data streams containing only native scalars and native arrays, these mechanisms may introduce additional parsing steps. Different modules, platforms, or languages ​​also need to consistently handle field order, byte order, array length, and message boundaries; without stable boundary recognition and type resolution conventions, the reading end is prone to problems such as parsing position offsets, length judgment errors, or inconsistent type restoration.

[0005] Therefore, this invention proposes a compact binary serialization and zero-copy read / write method and system for native data types. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the present invention aims to provide a compact binary serialization and zero-copy read / write method and system for native data types.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A compact binary serialization and zero-copy read / write method for primitive data types, characterized by comprising:

[0009] Obtain raw type data, a pre-defined structure description, and a directly addressable byte buffer, wherein the raw type data includes raw scalars and raw arrays that are not wrapped by objects;

[0010] Based on the pre-defined structure description, determine the field order, byte order and encoding parsing rules of the original type data. Perform positive and negative interleaving mapping and variable-length integer encoding on signed integers or long integers. Write the original binary bits on floating-point types. Write the array length on the original array according to variable-length encoding and write the original bytes of the elements continuously. Aggregate the encoding results into a compact binary data stream.

[0011] Write a compact binary data stream into a byte buffer in field order, and record the effective data length to form a data frame;

[0012] When a commit instruction for a byte buffer is received, data frames are written in batches to a file channel, network channel, or memory-mapped file buffer.

[0013] When reading a data frame, the original scalar is restored according to the encoding parsing rules. For an original array whose elements are written continuously with a fixed byte width, an original type buffer view that matches the array element type and shares the storage area is created in the sub-interval where the element data is located. The array elements are read through the view, and the restored original scalar and the original type buffer view used to represent the original array are output.

[0014] Furthermore, for signed integers or long integers, perform positive and negative interleaving mapping and variable-length integer encoding, including: mapping the value to be encoded to a non-negative integer according to the sign and magnitude of the value; extracting seven bits of valid data each time to form an encoded byte starting from the low bit of the non-negative integer; and determining the continuation flag bit of the encoded byte based on whether there is any unencoded valid data after extraction.

[0015] Further, determining the continuation flag of the encoded byte includes: determining the encoding termination position based on the remaining value after extracting seven valid data bits each time; when the remaining value is zero, setting the continuation flag of the current encoded byte to the end state, so that the number of encoded bytes of integers does not exceed five and the number of encoded bytes of long integers does not exceed ten.

[0016] Furthermore, writing raw binary bits to floating-point numbers includes: determining a data width of four bytes or eight bytes based on the type of the floating-point number; obtaining the standard floating-point number binary bits corresponding to the floating-point number; and writing the binary bits into the byte buffer according to the byte order determined by the preset structure description.

[0017] Furthermore, writing the array length into the native array using variable-length encoding and continuously writing the original bytes of the elements includes: generating variable-length encoding of the array length based on the number of elements contained in the native array; calculating the element data length based on the number of elements and the fixed byte width corresponding to the array elements; allocating a continuous byte interval after the variable-length encoding of the array length, and writing the original bytes of each array element into the continuous byte interval according to the array index order.

[0018] Furthermore, the directly addressable byte buffer can be a direct memory byte buffer or a memory-mapped file buffer; during the writing of a compact binary data stream, the read / write position, valid data limit position, and total capacity of the byte buffer are maintained, and the valid data length is calculated based on the start read / write position and end write position of the data frame.

[0019] Furthermore, before writing the encoded result, the remaining capacity of the byte buffer is calculated based on the total capacity and the current read / write position; the remaining capacity is compared with the byte length of the encoded result to be written; when the remaining capacity is less than the byte length, the expanded buffer capacity is determined based on the byte length, or the process is switched to another byte buffer or mapped file segment.

[0020] Furthermore, writing data frames in batches to a file channel, network channel, or memory-mapped file buffer includes: determining a continuous valid interval corresponding to one or more complete data frames based on the start position and valid data length of each data frame; using the continuous valid interval as the data range for a batch write; and updating or resetting the read / write position of the byte buffer after confirming that the continuous valid interval has been written.

[0021] Furthermore, the native type buffer view shares the same underlying storage area with the sub-interval where the element data is located. This ensures that when reading array elements through the native type buffer view, the element data is not copied to the intermediate byte array. Moreover, when the sub-interval where the element data is located is writable, modifications to array elements through the native type buffer view are synchronized to the sub-interval where the element data is located.

[0022] A compact binary serialization and zero-copy read / write system for primitive data types includes:

[0023] The data acquisition module is used to acquire raw type data, preset structure descriptions, and directly addressable byte buffers. The raw type data includes raw scalars and raw arrays that are not wrapped by objects.

[0024] The data encoding module is used to determine the field order, byte order and encoding parsing rules of the original data type according to the preset structure description. It performs positive and negative interleaving mapping and variable length integer encoding on signed integers or long integers, writes the original binary bits on floating-point types, writes the array length on the original array according to the variable length encoding and writes the original bytes of the elements continuously, and gathers the encoding results into a compact binary data stream.

[0025] The data frame forming module is used to write a compact binary data stream into a byte buffer in field order and record the effective data length to form a data frame.

[0026] The data writing module is used to write data frames in batches to a file channel, network channel, or memory-mapped file buffer when a commit instruction for a byte buffer is received.

[0027] The data restoration module is used to restore the original scalar according to the encoding parsing rules when reading data frames. For original arrays whose elements are written continuously with a fixed byte width, it creates an original type buffer view that matches the array element type and shares the storage area in the sub-interval where the element data is located. The array elements are read through the view, and the restored original scalar and the original type buffer view used to represent the original array are output.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. This invention reduces redundant content such as class descriptions, field names, and object wrapper information by using a compact coding method oriented towards native types;

[0030] 2. This invention enables small integers to occupy fewer bytes by performing positive and negative interleaving mapping and variable-length integer encoding on signed integers. By writing floating-point numbers to the original binary bits, it avoids the space and computational overhead caused by text conversion. By fronting the array length and laying out the original bytes of the elements continuously, the array boundaries can be quickly determined.

[0031] 3. This invention reduces fine-grained I / O operations by using directly addressable byte buffers, valid data frame length, and batch write mechanisms; and reduces intermediate copying and object creation overhead when reading native arrays by creating native type buffer views of shared storage areas in array element sub-intervals. Attached Figure Description

[0032] Figure 1 A flowchart illustrating a compact binary serialization and zero-copy read / write method for a primitive data type;

[0033] Figure 2 A flowchart illustrating the process of batch writing data frames;

[0034] Figure 3This is a schematic diagram of the architecture of a compact binary serialization and zero-copy read / write system for a primitive data type. Detailed Implementation

[0035] Example 1: Refer to Figure 1 A compact binary serialization and zero-copy read / write method for primitive data types, comprising:

[0036] A compact binary serialization and zero-copy read / write method for primitive data types can be applied to scenarios such as Java big data transmission, local disk storage, network communication, memory cache serialization, file persistence, and high-speed data exchange. The executor can be a serialization component running on a server, gateway, terminal device, or data processing node, or an encoding / decoding program embedded in a communication framework, caching framework, or file storage framework.

[0037] The input objects in this embodiment include primitive type data, a preset structure description, and a directly addressable byte buffer. The primitive type data includes unwrapped primitive scalars and primitive arrays, such as integers, long integers, single-precision floating-point numbers, double-precision floating-point numbers, bytes, and arrays of the above types. Specifically, primitive data includes byte, boolean, short, char, int, long, float, double, byte arrays, int arrays, long arrays, float arrays, and double arrays. The output includes the data frame formed during the encoding stage, and the primitive scalars restored during the reading stage, as well as a view of the primitive type buffer that can be used to read elements of the primitive array.

[0038] S1. Obtain raw data type, pre-defined structure description, and directly addressable byte buffer.

[0039] It receives raw data type from the upper-layer business and obtains the preset structure description corresponding to that raw data type. The preset structure description is used to enable the encoding and decoding ends to process the same data frame according to a consistent field order and parsing rules without depending on the field names or class descriptions in the data frame.

[0040] In this embodiment, the preset structure description is represented by a structure description table, configuration object, interface parameter set, or a description object generated at compile time. The preset structure description includes field sequence number, field type, whether the field is an array, array element type, array encoding mode, byte order, integer encoding method, floating-point writing method, frame header format, and exception checking strategy. The array encoding mode includes fixed-width continuous encoding mode and variable-length element encoding mode; only when the array encoding mode is fixed-width continuous encoding mode, the reading end creates a native type buffer view. The data frame includes a frame header and a frame body. The frame header is located before the frame body and includes at least the magic number, format version, structure description identifier, frame body length, and flag bits. The structure description identifier is used to indicate the preset structure description used by both the encoding and decoding ends; the frame body length is used to determine the parsing boundary of the current data frame; the flag bits are used to indicate byte order, whether it is read-only, whether it contains a check value, or whether it is a variable frame.

[0041] For example, the predefined structure description specifies that the first field is an int scalar, the second field is a long scalar, and the third field is a double array. The encoding end writes the corresponding compact binary data sequentially according to the order of these fields, and the decoding end reads the data in the same order, without needing to write field names, field type names, or object wrapper information into the data frame.

[0042] Directly addressable byte buffers are used to carry encoded, compact binary data streams. These byte buffers can be direct memory byte buffers or memory-mapped file buffers. Direct memory byte buffers are suitable for network transmissions, local caching, or medium-sized batch writes; memory-mapped file buffers are suitable for large file writes, file segmentation mapping, or scenarios requiring direct access to file content.

[0043] In this embodiment, zero-copy read / write means that when reading a raw array, the contiguous byte range containing the array elements is not copied to an intermediate byte array or wrapper object collection. Instead, a raw type buffer view matching the array element type is created based on the contiguous byte range, and the array elements are accessed through this raw type buffer view. For the write process of file channels or network channels, this embodiment reduces intermediate buffer copying and fine-grained write operations through direct memory byte buffers, memory-mapped file buffers, and batch write mechanisms, but does not limit the absence of any data copying within the operating system, disk controller, or network device.

[0044] The byte buffer maintains read / write positions, valid data limit positions, and total capacity. For ease of explanation, let's assume the byte buffer is... The starting write position of the data frame is The current write position is The total capacity of the buffer is The effective data length of the data frame is Current remaining capacity It can be represented as The aforementioned state variables are used to subsequently determine whether there is available space for writing, calculate the effective length of the data frame, and determine the continuous valid interval for batch writing.

[0045] When raw data types are missing, predefined structure descriptions are absent, or predefined structure descriptions do not match the input data, the data acquisition module can refuse to enter the encoding process and return an exception status to the upper layer indicating missing structure description, mismatched field types, or invalid input data. In some embodiments, if the upper-layer business allows empty arrays, the array length can be recorded as zero; if the upper-layer business allows null references, the null value handling method can be agreed upon by the predefined structure description. When null value handling is not enabled, null references can be treated as invalid input.

[0046] S2. Perform compact binary encoding on the native type data according to the predefined structure description.

[0047] The field order, byte order, and encoding / parsing rules for each field are determined based on the predefined structure description, and each field is processed sequentially according to the field order. Different encoding methods are used for different types of raw data to reduce redundant bytes while maintaining parsability.

[0048] For signed integers or signed long integers, an alternating positive and negative mapping is performed first. The alternating positive and negative mapping is used to map signed integers to non-negative integers, so that positive and negative numbers with smaller absolute values ​​correspond to smaller non-negative integers, thus adapting to subsequent variable-length encoding.

[0049] Let the value to be encoded be The non-negative integers after mapping are The bit width is In a specific embodiment, the mapping is performed using the following relationship: .in, For integers, we can use 32; for long integers, we can use 64. Indicates left shift, Indicates a signed right shift. This represents a bitwise XOR operation. The above formula is an optional implementation; any implementation that can map the value to be encoded to a non-negative integer based on the sign and magnitude of the value is acceptable.

[0050] After the mapping is completed, for non-negative integers Perform the shortest variable-length integer encoding. From Starting from the least significant bit, seven bits of valid data are extracted each time to form a coded byte. If there is still unencoded valid data after extraction, the continue flag of that coded byte is set to continue; if the remaining value is zero, the continue flag of that coded byte is set to end. A coded byte can consist of the lower seven valid data bits and one continue flag, where the continue flag is located in the highest bit of the coded byte. When the continue flag is in the continue state, it indicates that there are still coded bytes to follow; when the continue flag is in the end state, it indicates that the variable-length integer encoding of the current field has ended.

[0051] In some embodiments, the number of encoded bytes for integers does not exceed five, and the number of encoded bytes for long integers does not exceed ten. Encoding of the current value ends immediately when the remaining value is zero to avoid writing redundant subsequent encoded bytes, thus forming the shortest variable-length integer encoding. If, during the encoding process, it is found that more than the maximum number of encoded bytes for the corresponding type is required to complete the encoding, it can be determined that the data to be encoded or the type configuration is abnormal, and the encoding of the current field is stopped.

[0052] For floating-point numbers, the data width is determined based on the specific floating-point type. Single-precision floating-point numbers correspond to four bytes, and double-precision floating-point numbers correspond to eight bytes. The standard floating-point binary bits are obtained, and written to the compact binary data stream according to the byte order determined by the preset structure description. This process does not convert the floating-point number to a string, nor does it write field names or object wrapper information. If little-endian byte order is specified in the preset structure description, the least significant byte is written first; if big-endian byte order is specified, the most significant byte is written first. Using the same byte order at both the encoding and decoding ends ensures that the floating-point number is correctly restored.

[0053] Furthermore, for byte type data, a single byte of the original value is written during encoding; for boolean type data, a single byte is used to represent the logical value, where 0 represents false and 1 represents true; for short and char type data, the corresponding original bytes are written according to the byte order determined by the predefined structure description, with a data width of two bytes. For the raw arrays of the above types, the variable-length encoding of the number of array elements is written first, and then the original bytes of each array element are written continuously according to the fixed byte width of the corresponding element.

[0054] For raw arrays, first generate a variable-length encoding for the array length based on the number of elements in the raw array. Let the number of array elements be... The array elements are of type T, and the element width is a fixed number of bytes. The length of the array elements is... It can be represented as: .in, This indicates the byte width of a single array element; for example, byte is 1 byte, int and float are 4 bytes, and long and double are 8 bytes. Array length. This indicates the number of array elements, not the total number of bytes occupied by the array. The decoding end can then determine the number of elements based on this information. and Calculate the boundaries of element data.

[0055] After variable-length encoding of the array, a contiguous byte range is allocated, and the raw bytes of each array element are written into this range in array index order. No delimiters, field names, or object wrapper information are written between array elements. For arrays where elements are represented with a fixed byte width, this contiguous byte range can be located as the sub-range containing the element data during the read phase, and further used to create a native type buffer view.

[0056] When the array length is zero, after writing the variable-length encoding representing zero length, the original bytes of the elements are not written. If the number of array elements... Fixed byte width of element If the product exceeds the maximum allowed data frame length in the current implementation, the current data frame encoding will be terminated and an array length out-of-bounds exception will be output. If the preset structure description indicates that a field is a scalar but the actual input is an array, or a field is an array but the actual input is a scalar, then it can be determined that the field attributes do not match.

[0057] Through the above processing, the encoding results of different primitive types are aggregated into a compact binary data stream according to the field order in the preset structure description.

[0058] S3. Write the compact binary data stream into a byte buffer and record the effective data length to form a data frame.

[0059] A compact binary data stream is written to a directly addressable byte buffer in field-by-field order. At the start of writing, the current read / write position is recorded as the starting position for writing the data frame. Each time an encoded byte is written, the current read / write position is... Subsequently, after all fields corresponding to a data frame have been written, the effective data length is calculated based on the end and start writing positions. ,Right now From the byte buffer Start, length is A continuous region constitutes a complete data frame.

[0060] A data frame is a binary data segment with a defined start position and valid data length. Data frames serve as the basic unit for subsequent batch writes and are also used by the decoding end to determine the parsing boundaries of the frame. In some embodiments, a data frame can be represented by a data frame description record, which includes the frame start position, valid data length, frame state, and target write position. The frame state can include states such as writing, committable, writing out, and committed, to prevent incomplete half-frames of data from being written out in batches.

[0061] Before writing the encoded result, the data frame forming module calculates the remaining capacity R based on the total capacity C and the current read / write position p. When R is less than the length of the bytes to be written for the encoded result, the data frame forming module can determine the expanded buffer capacity based on the length of the bytes to be written, or switch to another byte buffer or mapped file segment. In some embodiments, the expanded capacity may not be less than the sum of the current capacity and the length of the bytes to be written; in other embodiments, a doubling expansion method or segmented switching according to a preset maximum capacity is used. All of the above expansion and switching methods are based on the principle of ensuring that the data frame can be written completely and subsequently located.

[0062] If the byte buffer is read-only, the mapped file segment is not writable, expansion fails, or switching to a new buffer is impossible, the data frame forming module can stop writing the current data frame and mark it as invalid or a write failure. For data that has been written but has not yet formed a complete data frame, it can be rolled back to... Alternatively, record the state of abnormal frames to prevent them from being written out as valid data frames later.

[0063] S4. When a commit instruction for a byte buffer is received, write the data frames in batches to the file channel, network channel, or memory-mapped file buffer.

[0064] Listen for or receive commit commands. Commit commands can be generated by the upper-layer business explicitly calling the commit interface, or triggered by events such as the completion of an encoding task, the end of a batch processing cycle, the need to refresh the cache, or the completion of a write process. Upon receiving a commit command, select complete data frames in a commit-ready state from the byte buffer, and determine one or more consecutive valid intervals corresponding to complete data frames based on the start position and valid data length of each data frame.

[0065] Reference Figure 2 Let the starting position of the first data frame to be submitted be... The end position of the last consecutive data frame is Then the continuous effective interval can be represented as The data writing module uses this continuous valid interval as the data range for a single batch write. If multiple data frames are consecutively arranged in the byte buffer, they can be merged into a single batch write; if there are incomplete data frames or invalid regions in the middle, only the first continuous complete interval is selected for writing to avoid writing incomplete frames to the target channel.

[0066] When the target is a file channel, the data writing module writes a continuous valid range to the local file or persistent file. During writing, the file write offset can be recorded, allowing the data frame to be located within the file using its starting offset and valid length. When the target is a network channel, the data writing module can send a continuous valid range as a network message payload and confirm the number of bytes sent based on the write result returned by the network channel. When the target is a memory-mapped file buffer, the data writing module determines the mapped write area based on the current write position of the memory-mapped file buffer and the length of the data to be written, and writes the data frame to that mapped write area.

[0067] In some embodiments, when the remaining capacity of the current memory-mapped file buffer is less than the effective data length of the data frame to be written, a new mapped file segment is established based on the next file write offset, and subsequent data frames are written to the new mapped file segment. Each mapped file segment may have a segment start file offset, segment capacity, intra-segment write position, and segment status. The actual storage location of the data frame in the target file can be determined by the segment start file offset and the intra-segment write position.

[0068] After confirming that a continuous valid range has been written, update or reset the read / write position of the byte buffer. For a fully committed buffer, the read / write position can be reset to the beginning position to reuse the buffer; for a buffer with uncommitted data frames, the uncommitted data frames can be retained and only the committed areas can be marked as reusable. If a partial write occurs on a file channel or network channel, record the written position and the remaining unwritten range, and continue writing during subsequent commits or retries to avoid duplicate writes or missed data. If the target channel is unavailable, the data frame can be kept in a committable state, and a write failure status can be returned to the upper layer.

[0069] S5. Read the data frame and restore the original type data according to the encoding parsing rules.

[0070] Data frames are retrieved from byte buffers, file channel read results, network receive buffers, or memory-mapped file buffers. The parsing boundaries are determined based on the start position and valid data length of the data frame, and each field is read sequentially according to the field order and encoding parsing rules in the preset structure description. Since the encoding and decoding ends use the same preset structure description, the decoding end does not need to read the field names or class descriptions from the data frame to determine the type and reading method of each field.

[0071] For signed integers or signed long integers, the variable-length integer code is read byte by byte starting from the current read position. For each byte read, the lower seven bits are taken as valid data, and the continuation flag is used to determine whether to continue reading subsequent bytes. When a byte with the continuation flag set to "end" is read, the already read valid data bits are combined into a mapped non-negative integer u, and then a reverse mapping with alternating positive and negative bits is performed on u to obtain the original signed integer or long integer. In some embodiments, the sign can be determined based on the least significant bit of u, and the original value can be obtained from the right shift result. If the integer does not reach the end state within five bytes, or the long integer does not reach the end state within ten bytes, the field is considered incompletely encoded or the data frame is corrupted, and parsing of the current data frame is stopped.

[0072] For floating-point types, the system determines whether to read four or eight bytes of raw binary data based on the preset structure description, and then restores them to single-precision or double-precision floating-point values ​​according to the byte order in the preset structure description. If the number of readable bytes remaining in the current field is insufficient to match the corresponding data width, the data frame boundary is deemed abnormal, and parsing of the current data frame is stopped.

[0073] For a raw array, first read the variable-length encoding of the array length to obtain the number of array elements, n. Then, determine the array element type T according to the predefined structure description, and calculate the element data length D = n × w(T) based on the fixed byte width w(T) corresponding to element type T. Let the starting position of the array element data be... The end position of the array element data is .Will Compare with the valid data boundaries of the data frame. When Confirmation is made when the data frame's valid data boundaries are not exceeded. The sub-interval containing the array element data; when If the data exceeds the valid data boundary of the data frame, the array length or data frame boundary is determined to be abnormal, and the parsing of the current data frame is stopped.

[0074] For native arrays whose elements are written continuously in fixed-byte widths, a native type buffer view matching the array element type is created in the sub-interval containing the element data. This view shares the same underlying storage area as the sub-interval containing the element data and interprets the underlying bytes according to the byte order in the predefined structure description. For int arrays, an integer buffer view can be created; for long arrays, a long integer buffer view can be created; for float arrays, a floating-point buffer view can be created; for double arrays, a double-precision floating-point buffer view can be created; and for byte arrays, a byte buffer view can be used directly.

[0075] Read the first through the primitive type buffer view When there are multiple array elements, the starting position of the array element is used. Array index and element fixed byte width Locate the corresponding byte position, its element offset is Therefore, when reading array elements, it is not necessary to first copy the element data to an intermediate byte array, nor is it necessary to generate wrapper objects element by element. In embodiments where the sub-interval containing the element data is writable, when array elements are modified through a primitive type buffer view, the modification result is synchronized to the sub-interval containing the element data; in read-only buffers or read-only memory-mapped file buffers, this view is only used for reading and does not perform write modifications.

[0076] After the read operation is complete, the output includes the restored native scalar and a native type buffer view representing the native array. For native scalars, the output value can be the corresponding integer, long integer, floating-point, or double-precision floating-point value. For native arrays, the output can be a native type buffer view or the array read result provided to the upper layer based on this view. In scenarios requiring compatibility with ordinary array interfaces, the upper-layer business logic can construct an array object based on the view; in zero-copy read / write scenarios, it is preferable to directly use the native type buffer view to access array elements.

[0077] During parsing, length verification, boundary verification, and encoding integrity verification can also be performed. Length verification is used to determine if the array length is within the allowed range; boundary verification is used to determine if the read position of each field exceeds the valid length of the data frame; encoding integrity verification is used to determine if the variable-length integer encoding ends within the maximum allowed number of encoded bytes for the corresponding type. When any verification fails, parsing the current data frame stops, and an exception status is output. In network reception scenarios, the exception data frame can be discarded and the system can wait for the next complete data frame; in file reading scenarios, the file offset and valid length of the exception data frame can be recorded for subsequent location.

[0078] Example 2: Refer to Figure 3 A compact binary serialization and zero-copy read / write system for primitive data types, comprising:

[0079] The data acquisition module is used to acquire raw type data, preset structure descriptions, and directly addressable byte buffers. The raw type data includes raw scalars and raw arrays that are not wrapped by objects.

[0080] The data encoding module is used to determine the field order, byte order and encoding parsing rules of the original data type according to the preset structure description. It performs positive and negative interleaving mapping and shortest variable length integer encoding on signed integers or long integers, writes the original binary bits on floating-point types, writes the array length on the original array according to the variable length encoding and writes the original bytes of the elements continuously, and gathers the encoding results into a compact binary data stream.

[0081] The data frame forming module is used to write a compact binary data stream into a byte buffer in field order and record the effective data length to form a data frame.

[0082] The data writing module is used to write data frames in batches to a file channel, network channel, or memory-mapped file buffer when a commit instruction for a byte buffer is received.

[0083] The data restoration module is used to restore the original scalar according to the encoding parsing rules when reading data frames. For original arrays whose elements are written continuously with a fixed byte width, it creates an original type buffer view that matches the array element type and shares the storage area in the sub-interval where the element data is located. The array elements are read through the view, and the restored original scalar and the original type buffer view used to represent the original array are output.

[0084] The modules mentioned above can interact with each other through memory objects, buffer references, file channel references, or network channel references. When implemented within the same process, the modules can share the same directly addressable byte buffer; in distributed or cross-process scenarios, the data frames output by the data writing module can be read by the data restoration module of another process or another device.

[0085] Through the above steps, this embodiment can encode unwrapped raw scalars and raw arrays into a compact binary data stream, forming a data frame with an effective data length in a directly addressable byte buffer; upon submission, it writes to the target channel in batches as complete data frames; upon reading, it restores the raw scalar according to a preset structure description, and creates a raw type buffer view of a shared storage area on the sub-interval where the element data of the raw array is located. Thus, while reducing redundant serialization information, it lowers the overhead of intermediate copying and object creation during the raw array reading process, improving the processing efficiency of raw type data in storage, network communication, and cache persistence scenarios.

[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A compact binary serialization and zero-copy read / write method for primitive data types, characterized in that, include: Obtain raw type data, a pre-defined structure description, and a directly addressable byte buffer. The raw type data includes raw scalars and raw arrays that are not wrapped by objects. Based on the pre-defined structure description, determine the field order, byte order and encoding parsing rules of the original type data. Perform positive and negative interleaving mapping and variable-length integer encoding on signed integers or long integers. Write the original binary bits on floating-point types. Write the array length on the original array according to variable-length encoding and write the original bytes of the elements continuously. Aggregate the encoding results into a compact binary data stream. Write a compact binary data stream into a byte buffer in field order, and record the effective data length to form a data frame; When a commit instruction for a byte buffer is received, data frames are written in batches to a file channel, network channel, or memory-mapped file buffer. When reading a data frame, the original scalar is restored according to the encoding parsing rules. For an original array whose elements are written continuously with a fixed byte width, an original type buffer view that matches the array element type and shares the storage area is created in the sub-interval where the element data is located. The array elements are read through the view, and the restored original scalar and the original type buffer view used to represent the original array are output.

2. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, Perform positive and negative interleaving mapping and variable-length integer encoding on signed integers or long integers, including: mapping the value to be encoded to a non-negative integer according to the sign and magnitude of the value; extracting seven bits of valid data each time to form an encoded byte starting from the least significant bit of the non-negative integer; and determining the continuation flag bit of the encoded byte based on whether there is any unencoded valid data after extraction.

3. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 2, characterized in that, Determine the continuation flag of the encoded byte, including: determining the encoding termination position based on the remaining value after extracting seven valid data bits each time; when the remaining value is zero, set the continuation flag of the current encoded byte to the end state, so that the number of encoded bytes of integers does not exceed five and the number of encoded bytes of long integers does not exceed ten.

4. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, Writing raw binary bits to a floating-point number includes: determining a data width of four bytes or eight bytes based on the floating-point type; obtaining the standard floating-point number binary bits corresponding to the floating-point number; and writing the binary bits into the byte buffer according to the byte order determined by the preset structure description.

5. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, Write the array length into the native array using variable-length encoding and write the original bytes of the elements consecutively. This includes: generating variable-length encoding for the array length based on the number of elements in the native array; calculating the element data length based on the number of elements and the fixed byte width corresponding to the array elements; allocating a consecutive byte interval after the variable-length encoding of the array length, and writing the original bytes of each array element into the consecutive byte interval according to the array index order.

6. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, The directly addressable byte buffer can be a direct memory byte buffer or a memory-mapped file buffer; during the writing of a compact binary data stream, the read / write position, valid data limit position, and total capacity of the byte buffer are maintained, and the valid data length is calculated based on the start read / write position and end write position of the data frame.

7. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 6, characterized in that, Before writing the encoded result, calculate the remaining capacity of the byte buffer based on the total capacity and the current read / write position; compare the remaining capacity with the byte length of the encoded result to be written; when the remaining capacity is less than the byte length, determine the expanded buffer capacity based on the byte length, or switch to another byte buffer or mapped file segment.

8. The compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, Batch writing data frames to a file channel, network channel, or memory-mapped file buffer includes: determining one or more consecutive valid intervals corresponding to complete data frames based on the start position and valid data length of each data frame; using the consecutive valid intervals as the data range for a batch write; and updating or resetting the read / write position of the byte buffer after confirming that the consecutive valid intervals have been written.

9. A compact binary serialization and zero-copy read / write method for primitive data types according to claim 1, characterized in that, The native type buffer view shares the same underlying storage area with the sub-interval where the element data is located. This means that when reading array elements through the native type buffer view, the element data is not copied to the intermediate byte array. Furthermore, when the sub-interval where the element data is located is writable, modifications to array elements through the native type buffer view are synchronized to the sub-interval where the element data is located.

10. A compact binary serialization and zero-copy read / write system for primitive data types, applied in the compact binary serialization and zero-copy read / write method for primitive data types as described in claims 1-9, characterized in that, include: The data acquisition module is used to acquire raw type data, preset structure descriptions, and directly addressable byte buffers. The raw type data includes raw scalars and raw arrays that are not wrapped by objects. The data encoding module is used to determine the field order, byte order and encoding parsing rules of the original data type according to the preset structure description. It performs positive and negative interleaving mapping and variable length integer encoding on signed integers or long integers, writes the original binary bits on floating-point types, writes the array length on the original array according to the variable length encoding and writes the original bytes of the elements continuously, and gathers the encoding results into a compact binary data stream. The data frame forming module is used to write a compact binary data stream into a byte buffer in field order and record the effective data length to form a data frame. The data writing module is used to write data frames in batches to a file channel, network channel, or memory-mapped file buffer when a commit instruction for a byte buffer is received. The data restoration module is used to restore the original scalar according to the encoding parsing rules when reading data frames. For original arrays whose elements are written continuously with a fixed byte width, it creates an original type buffer view that matches the array element type and shares the storage area in the sub-interval where the element data is located. The array elements are read through the view, and the restored original scalar and the original type buffer view used to represent the original array are output.