Graph database decimal index coding and decoding method based on key value storage
By decomposing decimal values into symbols, weights, and number arrays, and employing a compact binary encoding method, the efficiency and sorting problems of high-precision decimal data storage and retrieval in graph database systems are solved, achieving efficient storage and retrieval, and making it suitable for analysis tasks in financial transactions and social networks.
Patent Information
- Application Number
- CN202511367845.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-24
AI Technical Summary
When processing high-precision decimal data, existing graph database systems suffer from low storage efficiency, insufficient comparison performance, flawed negative number sorting logic, and imperfect special value handling, making it difficult to meet the needs of complex analytical tasks.
A key-value-based graph database decimal index encoding method is adopted, which decomposes decimal values into symbols, weights, and numeric arrays. A compact binary encoding method is used, and a set of auxiliary encoding parameters is dynamically generated. Escape characters and padding characters are used to avoid conflicts. Pruning operations optimize storage space and ensure the correct sorting of negative weights and special values.
It significantly improves the indexing efficiency and query performance of high-precision decimal data, ensuring mathematical correctness and query performance. It is suitable for large-scale network data analysis scenarios, such as amount sorting in financial transaction graphs and weight filtering in social network graphs.
Smart Images

Figure CN120849671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of graph database management technology, and in particular to a method for encoding and decoding decimal indexes in graph databases based on key-value storage. Background Art
[0002] In modern database systems, graph databases are widely used in scenarios such as social networks, financial transactions, and knowledge graphs due to their efficient modeling capabilities for complex relational networks. These applications typically involve the storage and retrieval of high-precision decimal data (e.g., monetary calculations, geographic coordinate analysis, scientific computing), which places stringent demands on numerical accuracy, storage efficiency, and query performance.
[0003] However, existing graph database systems face the following technical bottlenecks when processing high-precision decimal data:
[0004] 1. Low storage efficiency. In existing technologies, high-precision decimal data is typically stored using string or binary encoding. While string storage provides a clear representation of numerical values, redundant characters (such as leading zeros and decimal points) consume additional storage space. Binary encoding schemes (such as IEEE 754 floating-point numbers) are difficult to reliably handle high-precision requirements and are prone to precision loss when processing values outside the standard range. These methods significantly increase storage overhead and limit system scalability in large-scale dataset scenarios.
[0005] 2. Insufficient comparison performance. In key-value storage systems, comparing decimal data based on strings requires parsing the numerical value character by character, resulting in a time complexity of O(n), where n is the number of digits. This is inefficient in frequent range queries or sorting operations. For example, in financial trading scenarios where it is necessary to quickly filter transaction amounts within a specific range, existing methods struggle to meet real-time requirements.
[0006] 3. Defects in negative number sorting logic. Existing encoding schemes do not adhere to the lexicographical order of negative numbers as they do to their mathematical numerical order. For example, the string encoding "-100" is lexicographically greater than "-2," while the actual numerical value "-100 < -2," leading to incorrect sorting results. Such issues can cause serious errors in social network relationship weight analysis or financial risk assessment.
[0007] 4. Inadequate handling of special values. Existing technologies lack a unified standard for supporting special values (such as NaN, positive infinity, and negative infinity). For example, in knowledge graphs, the handling of missing or outlier values (such as NaN) may lead to inconsistent query results or even system anomalies due to different encoding methods. Furthermore, the sorting logic for special values is not clearly defined, further increasing the complexity of data management.
[0008] In summary, existing graph database systems have significant shortcomings in storing and querying high-precision decimal data. There is an urgent need for an indexing scheme that can balance storage efficiency, comparison performance, negative number sorting logic, and special value handling to meet the needs of complex analysis tasks. Summary of the Invention
[0009] The purpose of this invention is to provide a method for encoding and decoding decimal indexes in graph databases based on key-value storage, so as to solve the problems of low storage efficiency, insufficient comparison performance, defects in negative number sorting logic, and imperfect handling of special values in existing graph database systems when processing high-precision decimal data.
[0010] To achieve the above objectives, this application adopts the following technical solution:
[0011] This application discloses a decimal index encoding method for graph databases based on key-value storage, comprising the following steps:
[0012] Receive the value to be encoded and its sorting direction parameter;
[0013] The numerical value to be encoded is normalized, and the sign, weight, and number array of the normalized numerical value are extracted, wherein the elements in the number array represent multiple digits with a predetermined base;
[0014] Based on the sorting direction parameter and symbol, a set of encoding auxiliary parameters is dynamically generated, which includes sorting mask, escape character and padding character;
[0015] Based on the set of encoding auxiliary parameters, the symbols, weights, and number arrays are converted into a binary encoded stream;
[0016] Add a data termination marker consisting of the escape characters to the end of the binary encoded stream;
[0017] The output binary encoded stream is used for ordered storage and retrieval of data.
[0018] Preferably, the normalization process for the value to be encoded includes:
[0019] Call the trimming method to remove leading or trailing zeros from the value to be encoded.
[0020] Preferably, the method further includes extracting the number of elements of the normalized value to be encoded;
[0021] Before dynamically generating the encoding auxiliary parameter set based on the sorting direction parameter and the symbol, the following steps are also included:
[0022] In response to the fact that the number of elements is zero, its weight is adjusted to the preset minimum weight;
[0023] In response to the symbol representing a negative value, the weight is negated.
[0024] Preferably, the step of dynamically generating an encoding auxiliary parameter set based on the sorting direction parameter and the symbol includes:
[0025] If the sorting direction is ascending, then the initial sorting mask, escape character, and padding character are respectively the first preset sorting mask, the first preset escape character, and the first preset padding character;
[0026] If the sorting direction is descending, then the initial sorting mask, escape character, and padding character are respectively the second preset mask, the second preset escape character, and the second preset padding character;
[0027] If the symbol represents a negative value, then the determined sorting mask, escape character, and padding character are inverted bit by bit.
[0028] Preferably, the step of converting the symbols, weights, and number arrays into a binary encoded stream includes:
[0029] The symbols and weights are encoded into binary data respectively, and the high-order bits are inverted to obtain the first byte sequence and the second byte sequence;
[0030] Iterate through the array of numbers, and for each number element, convert it to big-endian byte order with the most significant byte first and the least significant byte last.
[0031] If the value to be encoded is negative, the converted digital element is XORed with the sorting mask, and the result of the XOR operation is split into high-order bytes and low-order bytes.
[0032] When the generated high-order byte or low-order byte is the same as the escape character, the padding character is added after it.
[0033] Preferably, the data termination marker is a sequence of two consecutive escape characters.
[0034] A method for decoding decimal indexes in a graph database based on key-value storage, comprising:
[0035] Receive encoded data streams and their sorting direction parameters;
[0036] The symbols and weights are parsed from the encoded data stream;
[0037] Based on the sorting direction parameters and symbols, a set of decoding auxiliary parameters is dynamically generated, which includes sorting mask, escape characters, and padding characters.
[0038] The data termination marker in the encoded data stream is located based on the escape character;
[0039] Traverse the data segments in the encoded data stream before the data termination marker, and reconstruct the digital array according to the decoding auxiliary parameter set, wherein the elements in the digital array represent multi-digit numbers with a predetermined base;
[0040] Based on the symbols, weights, and the restored array of numbers, the original numerical values are reconstructed.
[0041] Preferably, the step of dynamically generating a decoding auxiliary parameter set based on the sorting direction parameter and the symbol includes:
[0042] If the sorting direction is ascending, then the initial sorting mask, escape character, and padding character are respectively the first preset mask, the first preset escape character, and the first preset padding character;
[0043] If the sorting direction is descending, then the initial sorting mask, escape character, and padding character are respectively the second preset mask, the second preset escape character, and the second preset padding character;
[0044] If the symbol represents a negative value, then the determined sorting mask, escape character, and padding character are inverted bit by bit.
[0045] Preferably, locating the data termination marker in the encoded data stream based on the escape character includes:
[0046] Search for two consecutive escape characters in the encoded data stream as data termination markers.
[0047] Preferably, the restored digital array includes:
[0048] Read two bytes at a time as a processing unit;
[0049] If the read byte is equal to an escape character, skip the next byte;
[0050] Combine the two bytes into big-endian byte order;
[0051] If the symbol represents a positive value, then the big-endian byte order is directly converted to the original number element;
[0052] If the symbol represents a negative value, then the big-endian byte order is XORed with the sorting mask, and the result is converted into the original number element.
[0053] Preferably, the method further includes:
[0054] When the restored numerical array is empty, the weight is adjusted to the preset minimum weight value.
[0055] Preferably, after resolving the symbols and weights, the following is also included:
[0056] The symbols and weights are inverted by taking the high-order bits, and the weights after taking the high-order bits are converted into the original numerical elements;
[0057] If the sign of the inverted high-order bits represents a negative value, then the inverse operation is performed on the converted weights.
[0058] This application has the following beneficial effects:
[0059] This application achieves efficient storage and retrieval by decomposing decimal values into signs, weights, and numeric arrays, and employing a compact binary encoding method. Specifically, negative weights are inverted to ensure the mathematical correctness of the sorting; specific sign flags support the handling of NaN, positive infinity, and negative infinity; escape characters and null characters are used to avoid conflicts; and storage space is optimized through pruning operations. This significantly improves the indexing efficiency and query performance of high-precision decimal data in graph databases, making it suitable for large-scale network data analysis scenarios, such as amount sorting in financial transaction graphs or weight filtering in social network graphs. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is a flowchart of a decimal index encoding method for a graph database based on key-value storage, provided in an embodiment of this application.
[0062] Figure 2 This is a flowchart illustrating the encoding of symbols, weights, and numeric arrays provided in the embodiments of this application;
[0063] Figure 3 This is a flowchart of a decimal index decoding method for a graph database based on key-value storage, provided in an embodiment of this application.
[0064] Figure 4 This is a flowchart of restoring a digital array provided in an embodiment of this application;
[0065] Figure 5 This is a schematic diagram of an electronic device that implements a decimal index encoding and decoding method for a graph database based on key-value storage, as provided in an embodiment of this application. Detailed Implementation
[0066] To make the technical solution of this application clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The terms "first," "second," etc., in the claims and specification of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate. This is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0067] Glossary:
[0068] Decimal data type: A high-precision numeric data type that can represent fractions and decimals with arbitrary precision, used to store decimal values in databases.
[0069] Key-value storage: A data storage paradigm where data is stored in the form of key-value pairs, commonly used in NoSQL databases (such as graph databases).
[0070] NumericDigit: A 16-bit unsigned integer, storing a number unit (0-9999) in a base of 10000, corresponding to a maximum of 4 decimal digits.
[0071] Weight: Represents the position of the first non-zero digit in a decimal value relative to the decimal point, used to determine the magnitude and precision of the value. In decimal representation, weight is defined as the position index of digit grouping. Specifically, with the decimal point as the reference, the integer part (to the left of the decimal point) is divided into groups of four digits from right to left, with the rightmost group (containing the units digit) having a weight of 0, and each group to the left having a weight increasing by 1. The fractional part (to the right of the decimal point) is divided into groups of four digits from left to right, with the leftmost group (containing the tenths digit) having a weight of -1, and each group to the right having a weight decreasing by 1. In this embodiment, the position of the first non-zero digit relative to the decimal point is taken as the weight of the entire value. For example, in the value 12345678.9012, the first non-zero digit is 1, located in the ten-millions place, so the weight of 12345678.9012 is recorded as 1.
[0072] Sign: Indicates the positive or negative sign of a decimal value (such as POSITIVE for non-negative numbers and NEGATIVE for negative numbers) or a special state (such as NaN for non-numeric values, POS_INF for positive infinity, and NEG_INF for negative infinity).
[0073] Digit arrays: A list of multiple NumericDigits arranged in order, which are then combined to form a complete high-precision value.
[0074] Number of elements (Ndigits): The number of NumericDigits actually used in the digits array determines the precision and size of the value.
[0075] Precision: The total number of significant digits in a decimal value, ignoring leading zeros. For example, 0.00456 has 3 significant digits.
[0076] Scale: The number of digits after the decimal point in a decimal value.
[0077] OrderMask: Specifies the sorting field and its direction using the binary pattern of a bitmask.
[0078] Escape character: A special character used to avoid conflicts during encoding.
[0079] Null character: Represents a specific value or separator in the encoding.
[0080] In this application, the high-precision decimal data type is implemented based on the Decimal class to support numerical operations of arbitrary precision. To ensure consistency and lossless precision in data processing, all external input data, regardless of its source format (string, integer, or floating-point), must be normalized to Decimal objects before encoding. A Decimal object is a complete data structure containing attributes such as sign, digit array (Digits), weight (Weight), and number of elements (Ndigits), as shown in the example below:
[0081] For the value 123.45, sign=POSITIVE, ndigits=2, weight=0, scale=2, digits=[0123,4500];
[0082] For the value -0.0006789, sign=NEGATIVE, ndigits=2, weight=-1, scale=7, digits=[0006,7890];
[0083] For the value 123456789, sign=POSITIVE, ndigits=3, weight=2, scale=0, digits=[0001,2345,6789];
[0084] For the special value NaN, sign=NaN, ndigits=0, weight=0, scale=0, digits=[] (empty array);
[0085] For the special value positive infinity, sign=POS_INF, ndigits=0, weight=0, scale=0, digits=[] (empty array);
[0086] For the special value negative infinity, sign=NEG_INF, ndigits=0, weight=0, scale=0, digits=[] (empty array).
[0087] Example 1
[0088] like Figure 1 As shown, this embodiment of the disclosure provides a decimal index encoding method for graph databases based on key-value storage, including steps S110-S160.
[0089] S110, Receive the value to be encoded and its sorting direction parameters.
[0090] In the encoding method proposed in this embodiment, the initial step includes the encoder receiving two core parameters from the calling end: the numerical object to be encoded and its corresponding sorting direction parameter. The numerical object to be encoded is a pre-constructed high-precision decimal Decimal object, which contains structured attribute information such as sign, number array, weight, and number of elements. The sorting direction parameter is a Boolean variable `isDesc`, whose value explicitly specifies the sorting rule to be followed during encoding; that is, when the value is false, it indicates an ascending order rule, and when the value is true, it indicates a descending order rule. By receiving and parsing the above input parameters, the foundation for subsequent numerical conversion and byte order processing is laid, thereby ensuring that the encoding operation is performed under clear logical goals and rule constraints.
[0091] S120. Normalize the numerical value to be encoded and extract the sign, weight and number array of the normalized numerical value to be encoded, wherein the elements in the number array represent multiple digits with a predetermined base.
[0092] Furthermore, the numerical values to be encoded are normalized, including:
[0093] Call the trimming method to remove leading or trailing zeros from the value to be encoded.
[0094] After obtaining the Decimal numerical object to be encoded and the sorting direction parameter isDesc, if there are redundant 0s in the Decimal numerical object, the trim method is first called to normalize the numerical value to be encoded. The purpose of this step is to eliminate redundant information in the numerical representation. By removing meaningless leading or trailing zeros, the unity and simplicity of the data format are ensured, thereby effectively reducing the data volume during subsequent storage and transmission.
[0095] After the normalization process is completed, the structured attribute information of the numerical value to be encoded, such as Sign, Digits, Weight, and Ndigits, is extracted. In the embodiments of the present disclosure, the elements in Digits decompose the numerical value using base 10000 (ten-thousand进制), that is, each array element corresponds to four decimal digits, thereby achieving an efficient compressed representation of large numbers.
[0096] It should be noted here that in the embodiments of the present disclosure, Sign determines the sorting priority, NEG_INF < NEGATIVE < POSITIVE < POS_INF < NAN, ensuring the correct order of special values and positive and negative numbers.
[0097] S130. Dynamically generate an encoding auxiliary parameter set based on the sorting direction parameter and the sign. The encoding auxiliary parameter set includes a sorting mask, an escape character, and a padding character.
[0098] Furthermore, it also includes extracting the number of elements of the normalized numerical value to be encoded;
[0099] Before dynamically generating the encoding auxiliary parameter set based on the sorting direction parameter and the sign, it also includes:
[0100] In response to the number of elements being zero, adjust its weight to a preset minimum weight;
[0101] In response to the numerical value to be encoded being negative, perform an inversion operation on its weight.
[0102] After extracting the various attribute information of the numerical value to be encoded, first determine whether the number of its elements Ndigits is zero. If the numerical value is zero, that is, Ndigits = 0, then adjust its weight Weight to a preset minimum weight value MIN_WEIGHT. In the embodiments of the present disclosure, this minimum weight value is set to -32768. This operation ensures that a zero value can be assigned a unified and recognizable extreme weight during the sorting process, so as to be located in the correct position in the sorting sequence.
[0103] Next, it is determined whether the sign indicates a negative value. If the value is negative, its current weight is inverted by multiplying by -1. This step ensures that negative numbers with larger absolute values (i.e., smaller negative values) receive smaller weights, while negative numbers with smaller absolute values (i.e., larger negative values) receive larger weights. This guarantees that the values within the negative range are correctly arranged from smallest to largest when sorted by weight, ultimately achieving overall order for all values (negative, zero, and positive).
[0104] Furthermore, based on the sorting direction parameter and the sign, a set of encoding auxiliary parameters is dynamically generated, including:
[0105] If the sorting direction is ascending, the initial sort mask, escape character, and padding character are respectively the first preset sort mask, the first preset escape character, and the first preset padding character;
[0106] If the sorting direction is descending, the initial sorting mask, escape character, and padding character are the second preset mask, the second preset escape character, and the second preset padding character, respectively.
[0107] If the symbol represents a negative value, then the determined sort mask, escape character, and padding character are inverted bit by bit.
[0108] After completing the normalization and weight adjustment of the values to be encoded, the corresponding encoding auxiliary parameters are configured according to the values of the sorting direction parameter isDesc and the symbol Sign. Specifically, the boolean value of the sorting direction parameter isDesc is first determined. If isDesc is false, indicating a request for ascending order, the first preset sorting mask ASC_MASK (0x0000) is selected as the initial sorting mask, the first preset escape character ASC_ESCAPE (0x00) is selected as the initial escape character, and the first preset padding character ASC_NULL (0x01) is selected as the initial padding character. It should be noted that in this embodiment, the null character (i.e., the NULL character) is selected as the padding character to fill specific fields during encoding to ensure format alignment and correct parsing. If isDesc is true, indicating a request for descending order, the second preset sorting mask DESC_MASK (0xFFFF) is selected as the initial sorting mask, the second preset escape character DESC_ESCAPE (0xFF) is selected as the initial escape character, and the second preset padding character DESC_NULL (0xFE) is selected as the initial padding character. Similarly, the padding character here is also a null character. Choosing a different sort mask is fundamental to all subsequent encoding operations; it determines the tone of the byte sequence to ensure that the final lexicographical order conforms to the expected sorting direction.
[0109] After obtaining the initial set of auxiliary encoding parameters, the sign of the value to be encoded is further examined. If the sign indicates that the value is negative, then the currently selected sorting mask, escape character, and padding character are all bitwise inverted (i.e., all binary bits change from 0 to 1 and from 1 to 0). This step is a key mechanism to ensure that negative numbers obtain a self-consistent representation in sorting encoding that is the opposite of positive numbers. By inverting, a complementary encoding space is constructed for negative numbers, so that regardless of the sorting direction, negative numbers can maintain their mathematical magnitude in the final encoded sequence.
[0110] S140. Based on the encoding auxiliary parameter set, convert the symbols, weights, and number arrays into a binary encoded stream.
[0111] Furthermore, if Figure 2 As shown, converting the symbols, weights, and numeric arrays into a binary encoded stream includes:
[0112] S210. Encode the symbols and weights into binary data respectively, and invert the high bits to obtain the first byte sequence and the second byte sequence;
[0113] S220. Traverse the array of numbers. For each number element, convert it to big-endian byte order with the most significant byte first and the least significant byte last.
[0114] S230. If the value to be encoded is negative, the converted digital element is XORed with the sorting mask, and the result of the XOR operation is split into high-order byte and low-order byte.
[0115] S240. When the generated high-order byte or low-order byte is the same as the escape character, add a padding character after it.
[0116] After determining the set of auxiliary encoding parameters, the encoding process begins. First, the sign and weight are encoded as 16-bit unsigned integers, resulting in the first byte sequence and the second byte sequence. Then, the encoded results are added to the beginning of the encoded stream in order (first byte sequence then second byte sequence) to ensure that the sign byte is compared first during sorting. That is, entries with positive signs must be placed after entries with negative signs, thus ensuring that negative numbers always appear before positive numbers during sorting.
[0117] Then, iterate through each element in the number array and perform the following operations: Convert the current number element to a 16-bit unsigned integer representation in big-endian byte order. Big-endian byte order means that the bytes of a number are arranged in the order of the most significant byte first and the least significant byte last. For example, the hexadecimal representation of decimal 2 is 0x0002, which occupies 2 bytes. When represented in big-endian byte order, the most significant byte 0x00 comes first and the least significant byte 0x02 comes last, resulting in [0x00, 0x02]. Since 2 is a positive number, there is no need to perform a bitwise XOR operation between its big-endian byte order and a specific sorting mask (typically 0x0000). However, if it is a negative number, a bitwise XOR operation between the big-endian byte order and the sorting mask is performed to generate an intermediate value. Then, the XOR result, i.e., the intermediate value, is split into a high 8-bit byte byte1 and a low 8-bit byte byte2. For positive numbers, after converting them to big-endian byte order, the high 8-bit byte and the low 8-bit byte already exist, so no splitting is needed. Append byte1 to the encoding stream in sequence. If byte1 is equal to an escape character (such as 0x00), add a padding character (such as 0xFF, represented as cnull) after byte1 to avoid parsing conflicts. Append byte2 to the encoding stream in sequence. If byte2 is equal to an escape character, add cnull after byte2.
[0118] This step ensures that the encoded byte sequence of each number does not conflict with escape characters during transmission, while the additional order of byte1 and byte2 (low byte first, high byte last) is consistent with the original big-endian byte order representation, providing a reliable data structure for subsequent sorting.
[0119] S150. Add a data termination marker consisting of escape characters to the end of the binary encoded stream.
[0120] After the numeric array is encoded, two consecutive escape characters (such as a non-negative terminating marker [0x00, 0x00] or a negative terminating marker [0xFF, 0xFF]) are added to the end of the binary encoded stream to form a clear delimiter sequence. In this embodiment, the two consecutive escape characters are used as fixed delimiters to precisely mark the end position of the numeric array encoded portion. During parsing, this sequence can be scanned to distinguish the numeric array from subsequent data, avoiding parsing errors caused by confusion between data content and boundaries.
[0121] S160: Output a binary encoded stream for ordered data storage and retrieval.
[0122] The output binary encoded stream is meticulously designed to support efficient and ordered storage and accurate retrieval. Its structured encoding mechanism ensures that data can be directly sorted by byte sequence after storage without additional parsing or conversion, while providing clear boundary markers and logical hierarchies for retrieval operations.
[0123] In one embodiment, the above encoding process is explained using the ascending order of the positive numbers 12.34 and 123.45 as an example.
[0124] Convert 12.34 to a Decimal object, resulting in: sign=POSITIVE, weight=0, digits=[12,3400], ndigits=2, scale=2.
[0125] Convert 123.45 to a Decimal object, resulting in: sign=POSITIVE, weight=0, digits=[123,4500], ndigits=2, scale=2.
[0126] In this embodiment, isDesc is false, and 12.34 and 123.45 are both positive numbers. Therefore, in the encoding auxiliary parameter set, orderMask=0x0000, escape=0x00, and cnull=0x01.
[0127] For 12.34, its symbol POSITIVE = 0x0062, inverted by the most significant bit → [0x80, 0x62]. Here, 0x0062 is the hexadecimal code of POSITIVE itself. Converting it to binary gives 0000 0000 0110 0010. Then, by inverting the most significant bit, converting the 0 in the most significant bit to 1, we get 1000 0000 0110. 0010, its corresponding hexadecimal value is 0x8062 = [0x80, 0x62]; weight = 0, inverting the high-order bits gives [0x80, 0x00]; the number array digits = [12, 3400] → [0x00, 0x01, 0x0c, 0x0D, 0x48]. Here, first generate the big-endian byte order [0x00, 0x0C], then escape 0x00 to get [0x00, 0x01, 0x0c], then... 3400 is encoded in big-endian to get [0x0D, 0x48], and finally combined to get [12, 3400] → [0x00, 0x01, 0x0c, 0x0D, 0x48]; the separator is [0x00, 0x00]; all the above contents are combined in a fixed order to get the complete binary encoding result [0x80, 0x62, 0x80, 0x00, 0x00, 0x01, 0x0C, 0x0D, 0x48, 0x00, 0x00].
[0128] For 123.45, its symbol POSITIVE = 0x0062, high-order bits inverted → [0x80, 0x62]; weight = 0, high-order bits inverted → [0x80, 0x00]; digit array digits = [123, 4500] → [0x00, 0x01, 0x7B, 0x11, 0x94]; separator is [0x00, 0x00]; the complete binary encoding result is [0x80, 0x62, 0x80, 0x00, 0x00, 0x01, 0x7B, 0x11, 0x94, 0x00, 0x00]. The calculation logic is the same as for 12.34, and will not be repeated here.
[0129] Since the first 4 bytes (sign and weight) are the same, comparing the numeric parts, and since 0x0C < 0x7B, 12.34 < 123.45, which is in ascending order.
[0130] In another embodiment, the above encoding process is explained using the descending order of negative values -12.34 and -123.45 as an example.
[0131] Convert -12.34 to a Decimal object, resulting in: sign=NEGATIVE, weight=0, digits=[12,3400], ndigits=2, scale=2.
[0132] Convert -123.45 to a Decimal object, resulting in: sign=NEGATIVE, weight=0, digits=[123,4500], ndigits=2, scale=2.
[0133] For -12.34, its sign is NEGATIVE = 0x0020 (high-order bit inverted) → [0x80, 0x20], where 0x0020 is the encoding of NEGATIVE itself; weight = 0 (high-order bit inverted) → [0x80, 0x00]; digits = [12, 3400] → [0xFF, 0xFE, 0xF3, 0xF2, 0xB7]. First, a 12-bit big-endian byte order [0x00, 0x0C] is generated, then it is XORed with the sort mask 0xFFFF to obtain [0xFF, 0xFD]. 0xFF is escaped to obtain [0xFF, 0xFD]. [xFE,0xFD], then big-endian encoding of 3400 is obtained as [0x0D,0x48], and then bitwise XOR operation is performed with sorting mask 0xFFFF to obtain [0xF2,0xB7], and finally combined to obtain [12,3400]→[0xFF,0xFE,0xFD,0xF2,0xB7]; the separator is [0xFF,0xFF]; the complete binary encoding result is [0x80,0x20,0x80,0x00,0xFF,0xFE,0xFD,0xF2,0xB7,0xFF,0xFF].
[0134] For -123.45, its sign NEGATIVE = 0x0020, high-order bits inverted → [0x80, 0x20]; weight = 0, high-order bits inverted → [0x80, 0x00]; digit array digits = [123, 4500] → [0xFF, 0xFE, 0x84, 0xEE, 0x6B]; separator is [0xFF, 0xFF]; the complete binary encoding result is [0x80, 0x20, 0x80, 0x00, 0xFF, 0xFE, 0x84, 0xEE, 0x6B, 0xFF, 0xFF]. The calculation logic is the same as for -12.34, and will not be repeated here.
[0135] Since the first 4 bytes (sign and weight) are the same, comparing the numeric parts, and since 0x84 < 0xF2, -123.45 < -12.34, which is in ascending order.
[0136] In another embodiment, the above encoding process is explained using the ascending order of the special values NEG_INF, 0, and POS_INF as an example.
[0137] Convert NEG_INF to a Decimal object to get: sign=NEG_INF(0x01), digits=[], ndigits=0, weight=0.
[0138] Converting 0 to a Decimal object gives: sign = POSITIVE, digits = [], ndigits = 0, weight = MIN_WEIGHT.
[0139] Converting POS_INF to a Decimal object gives: sign = POS_INF(0xA3), digits = [], ndigits = 0, weight = 0.
[0140] The complete binary encoding result of NEG_INF is [0x80, 0x01, 0x7F, 0x00, 0x00, 0x00].
[0141] The complete binary encoding result of 0 is [0x80, 0x62, 0x7F, 0x00, 0x00, 0x00].
[0142] The complete binary encoding result of POS_INF is [0x80, 0xA3, 0x7F, 0x00, 0x00, 0x00].
[0143] Comparing the signs, since 0x01 < 0x62 < 0xA3, that is, NEG_INF < 0 < POS_INF, which conforms to ascending order.
[0144] The encoding method provided by the embodiments of the present disclosure decomposes the decimal value into a sign, a weight, and an array of digits, and adopts a compact binary encoding method to achieve efficient storage and query. Specifically, taking the negation of the negative weight ensures the mathematical correctness of sorting; handling NaN, positive infinity, and negative infinity is supported through specific sign flags; escape characters and null characters are also used to avoid conflicts, and the storage space is optimized through pruning operations. It significantly improves the indexing efficiency and query performance of high-precision decimal data in the graph database and is applicable to analysis scenarios of large-scale network data, such as sorting amounts in financial transaction graphs or filtering weights in social network graphs.
[0145] Embodiment 2
[0146] As Figure 3 shown, the embodiments of the present disclosure provide a decimal index decoding method for a graph database based on key-value storage, which is used to decode the data generated according to the encoding method described in Embodiment 1, including steps S210 - S260.
[0147] S310. Receive the encoded data stream and its sorting direction parameter.
[0148] In the decoding method proposed in this embodiment, the initial step includes receiving two core parameters passed from the calling end: the encoded data stream and its corresponding sorting direction parameter. The encoded data stream is a standardized binary sequence generated after encoding the Decimal object according to the encoding method provided in Embodiment 1. Its structure strictly follows a predefined format, including a symbol sequence, weight information, data termination markers, and necessary auxiliary parameters. The sorting direction parameter is completely consistent with the rules used in the encoding stage (e.g., both specify that the symbols are arranged in ascending or descending order), which ensures that the decoder can directly reuse the sorting logic during encoding. Through this parameter synchronization mechanism, the decoder can efficiently reconstruct the original Decimal object without additional calculation or negotiation. This design not only eliminates the risk of parameter inconsistency between the encoding and decoding ends but also significantly reduces the system implementation complexity and ensures high reliability of the decoding process.
[0149] S320. Parse the symbols and weights from the encoded data stream.
[0150] In Example 1, the first and second byte sequences generated by Sign and Weight are added to the beginning of the encoded stream in a fixed order (i.e., the first byte sequence first, then the second byte sequence). This design ensures that the decoder can directly obtain key information through a simple sequential reading mechanism—when the decoder starts parsing from the beginning of the encoded stream, it first reads and parses the preceding byte sequence (i.e., the first byte sequence) to recover the symbol information, and then parses the subsequent byte sequence (i.e., the second byte sequence) to recover the weight information. This allows for the complete extraction of symbols and weights without scanning or processing the rest of the encoded stream. The parsing process includes inverting the highest bit of both the read symbol and weight, converting the inverted weight into its original numerical element, and if the inverted symbol represents a negative value, performing an inverse operation on the converted weight to restore the original weight. This structured sequence arrangement not only simplifies the encoding / decoding logic but also significantly improves the efficiency and reliability of information acquisition.
[0151] S330. Based on the sorting direction parameter and symbol, dynamically generate a decoding auxiliary parameter set, which includes sorting mask, escape character and padding character.
[0152] In the encoding stage described in Example 1, the encoder generates unique encoding auxiliary parameters based on the sorting direction and the symbol. In the decoding stage, the method for determining the decoding auxiliary parameter set is completely consistent with the generation logic of the encoding auxiliary parameter set in Example 1. Specifically, the decoder directly calls the same generation logic using the received sorting direction parameter and the parsed symbol, without requiring additional input or parameter negotiation. This deterministic mechanism based on sorting and symbols ensures that the parameter set generation logic at both ends of the encoding and decoding process is completely consistent, eliminating the need for additional synchronization steps, thereby achieving efficient and unambiguous parameter reuse.
[0153] S340. Locate the data termination marker in the encoded data stream based on the escape character.
[0154] A data termination marker consists of two consecutive escape characters, its core advantage being to avoid conflicts with single escape characters in the original data. In the encoded data stream, the decoder uses a sequential scanning mechanism to compare bytes one by one, starting from the third byte sequence at the beginning. When two consecutive identical escape characters are detected (e.g., "0x00, 0x00" in the byte sequence), it is determined to be a data termination marker. At this point, all bytes before the data termination marker (excluding the data termination marker itself, the first byte sequence, and the second byte sequence) constitute the complete encoded stream of the original numeric array. This mechanism efficiently locates data boundaries through simple linear scanning, without the need for additional parsing or predefined lengths, significantly simplifying the boundary handling logic of the data stream.
[0155] S350. Traverse the data segments in the encoded data stream before the data termination marker, and reconstruct the digital array according to the decoding auxiliary parameter set, wherein the elements in the digital array represent multi-digit numbers with a predetermined base.
[0156] Furthermore, such as Figure 4 As shown, restore the number array, including:
[0157] S410: Read two bytes at a time as a processing unit;
[0158] S420. When the read byte is equal to an escape character, skip the next byte;
[0159] S430. Combine the two bytes into big-endian byte order;
[0160] S440. If the sign represents a positive value, then directly convert the big-endian byte order to the original number element;
[0161] S450. If the sign represents a negative value, then perform an XOR operation between the big-endian byte order and the sorting mask, and convert the result into the original number element.
[0162] The original number array can be obtained by restoring the encoded stream of the located number array. In this embodiment, restoring the original number array requires step-by-step execution according to specific byte sequence parsing rules. First, two bytes are read sequentially as a basic processing unit. If the currently read byte is equal to a preset escape character, the byte immediately following it is automatically ignored to ensure that the escape sequence does not affect the actual data parsing. Then, the two valid bytes are combined in order to form a complete big-endian byte order, which is based on hexadecimal representation. Next, it is determined whether the parsed symbol represents a negative value. If not, the big-endian byte order is directly converted into the original number element. If so, the big-endian byte order is XORed with the sorting mask, and the result is then converted into the original number element.
[0163] S360 reconstructs the original numerical value based on the symbols, weights, and the restored array of numbers.
[0164] After the numerical array is restored, the original values are gradually reconstructed by integrating the symbols, weights, and the restored numerical array. This process requires strict assurance that the order and mapping relationships of each data component are correct, and finally, a numerical result that conforms to the original semantics is generated through arithmetic composition.
[0165] The embodiments disclosed herein enable efficient decimal data indexing in the key-value store of a graph database, while maintaining the correctness of mathematical sorting and support for special values.
[0166] Example 3
[0167] like Figure 5 As shown, this embodiment of the present disclosure provides an electronic device, including a memory 501 and a processor 502. The memory 501 is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor 502 to implement the above-described method for encoding and decoding a decimal index of a graph database based on key-value storage.
[0168] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the electronic device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0169] A computer-readable storage medium storing a computer program that, when executed by a computer, implements the above-described method for encoding and decoding a decimal index of a graph database based on key-value storage.
[0170] For example, a computer program can be divided into one or more modules / units. One or more modules / units are stored in memory 501 and executed by processor 502. Data I / O interface transmission is completed by input interface 505 and output interface 506 to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions. The instruction segments are used to describe the execution process of the computer program in the computer device.
[0171] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a memory 501 and a processor 502. Those skilled in the art will understand that this embodiment is merely an example of a computer device and does not constitute a limitation on the computer device. It may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include an input device 507, a network access device, a bus, etc.
[0172] Processor 502 can be a Central Processing Unit (CPU), or other general-purpose processor 502, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processor 502 can be a microprocessor 502, or any conventional processor 502, etc.
[0173] The memory 501 can be an internal storage unit of a computer device, such as a hard drive or memory. The memory 501 can also be an external storage device of a computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 501 can include both internal and external storage units. The memory 501 is used to store computer programs and other programs and data required by the computer device. The memory 501 can also be used for temporary storage in the output device 508. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM) 503, random access memory (RAM) 504, discs, or optical discs.
[0174] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A decimal index encoding method for graph databases based on key-value storage, characterized in that, Includes the following steps: Receive the value to be encoded and its sorting direction parameter; The numerical value to be encoded is normalized, and the sign, weight, and number array of the normalized numerical value are extracted, wherein the elements in the number array represent multiple digits with a predetermined base; Based on the sorting direction parameter and symbol, a set of encoding auxiliary parameters is dynamically generated, which includes sorting mask, escape character and padding character; Based on the set of encoding auxiliary parameters, the symbols, weights, and number arrays are converted into a binary encoded stream; Add a data termination marker consisting of the escape characters to the end of the binary encoded stream; The output binary encoded stream is used for ordered storage and retrieval of data.
2. The decimal index encoding method for a graph database based on key-value storage according to claim 1, characterized in that, The normalization process for the value to be encoded includes: Call the trimming method to remove leading or trailing zeros from the value to be encoded.
3. The decimal index encoding method for a graph database based on key-value storage according to claim 1, characterized in that, The method also includes extracting the number of elements in the normalized numerical value to be encoded; Before dynamically generating the encoding auxiliary parameter set based on the sorting direction parameter and the symbol, the following steps are also included: In response to the fact that the number of elements is zero, its weight is adjusted to the preset minimum weight; In response to the symbol representing a negative value, the weight is negated.
4. The decimal index encoding method for a graph database based on key-value storage according to claim 3, characterized in that, The dynamic generation of an encoding auxiliary parameter set based on the sorting direction parameter and the symbol includes: If the sorting direction is ascending, then the initial sorting mask, escape character, and padding character are respectively the first preset sorting mask, the first preset escape character, and the first preset padding character; If the sorting direction is descending, then the initial sorting mask, escape character, and padding character are respectively the second preset mask, the second preset escape character, and the second preset padding character; If the symbol represents a negative value, then the determined sorting mask, escape character, and padding character are inverted bit by bit.
5. The decimal index encoding method for a graph database based on key-value storage according to claim 1, characterized in that, The step of converting the symbols, weights, and number arrays into a binary encoded stream includes: The symbols and weights are encoded into binary data respectively, and the high-order bits are inverted to obtain the first byte sequence and the second byte sequence; Iterate through the array of numbers, and for each number element, convert it to big-endian byte order with the most significant byte first and the least significant byte last. If the value to be encoded is negative, the converted digital element is XORed with the sorting mask, and the result of the XOR operation is split into high-order bytes and low-order bytes. When the generated high-order byte or low-order byte is the same as the escape character, the padding character is added after it.
6. The decimal index encoding method for a graph database based on key-value storage according to claim 1, characterized in that, The data termination marker is a sequence of two consecutive escape characters.
7. A method for decoding decimal indexes in a graph database based on key-value storage, characterized in that, include: Receive encoded data streams and their sorting direction parameters; The symbols and weights are parsed from the encoded data stream; Based on the sorting direction parameters and symbols, a set of decoding auxiliary parameters is dynamically generated, which includes sorting mask, escape characters, and padding characters. The data termination marker in the encoded data stream is located based on the escape character; Traverse the data segments in the encoded data stream before the data termination marker, and reconstruct the digital array according to the decoding auxiliary parameter set, wherein the elements in the digital array represent multiple digits with a predetermined base; Based on the symbols, weights, and the restored array of numbers, the original numerical values are reconstructed.
8. A method for decoding a decimal index in a graph database based on key-value storage according to claim 7, characterized in that, The dynamic generation of a decoding auxiliary parameter set based on the sorting direction parameter and the symbol includes: If the sorting direction is ascending, then the initial sorting mask, escape character, and padding character are respectively the first preset mask, the first preset escape character, and the first preset padding character; If the sorting direction is descending, then the initial sorting mask, escape character, and padding character are respectively the second preset mask, the second preset escape character, and the second preset padding character; If the symbol represents a negative value, then the determined sorting mask, escape character, and padding character are inverted bit by bit.
9. A method for decoding a decimal index in a graph database based on key-value storage according to claim 7, characterized in that, The step of locating the data termination marker in the encoded data stream based on the escape character includes: Search for two consecutive escape characters in the encoded data stream as data termination markers.
10. A method for decoding a decimal index in a graph database based on key-value storage according to claim 7, characterized in that, The restored digital array includes: Read two bytes at a time as a processing unit; If the read byte is equal to an escape character, skip the next byte; Combine the two bytes into big-endian byte order; If the symbol represents a positive value, then the big-endian byte order is directly converted to the original number element; If the symbol represents a negative value, then the big-endian byte order is XORed with the sorting mask, and the result is converted into the original number element.
11. A method for decoding a decimal index in a graph database based on key-value storage according to claim 7, characterized in that, The method further includes: When the restored numerical array is empty, the weight is adjusted to the preset minimum weight value.
12. A method for decoding a decimal index in a graph database based on key-value storage as described in claim 7, characterized in that, After resolving the symbols and weights, the following is also included: The symbols and weights are inverted by taking the high-order bits, and the weights after taking the high-order bits are converted into the original numerical elements; If the sign of the inverted high-order bits represents a negative value, then the inverse operation is performed on the converted weights.
Citation Information
Patent Citations
Index map encoding and decoding methods based on ascending and descending tuples
CN105554504A
Indexing method and device based on key value pair KV system, electronic equipment and medium
CN111241108A
Method for optimizing stored data of database
CN114048192A
Graph data processing method, apparatus and device, and computer readable storage medium
CN120277240A
K-ARY tree to binary tree conversion through complete height balanced technique
WO2013186588A2
Cited By
Cross-system acquisition and restoration method, system and equipment for symbolic variables of welding robot
CN121535789A