Data information processing method and system based on big data
By replacing the path encoding from the root node to the branch node in the Hoffman tree as the fixed-length encoding of the branch node's sequence number in the layer, the special layer is determined, and the problem of limited compression of the Hoffman encoding is solved, and more efficient compression of data information is achieved.
Patent Information
- Application Number
- CN202510704095.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The compression rate of Hoffman encoding is limited by the frequency information of the data and cannot be further improved.
By replacing the path from the root node to the branch node in the Hoffman tree as the fixed-length encoding of the branch node's sequence number at the layer where it is located, a special layer is determined, and the data is encoded at the special layer to shorten the encoding length.
It improves the compression rate of data information and reduces data storage costs.
Smart Images

Figure CN120281323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing. More specifically, the present invention relates to a data information processing method and system based on big data. Background Art
[0002] With the acceleration of digital transformation, enterprises and service providers are facing an unprecedented data deluge; from the Internet, Internet of Things to enterprise internal systems, the speed and scale of data generation are increasing exponentially; this rapidly growing digital information not only requires more powerful data centers to process and analyze this data, but also requires enterprises and service providers to have larger storage capacities.
[0003] The "4V" model of big data - large volume, variety of data types, high velocity, veracity - and data value together constitute the core characteristics of big data.
[0004] These characteristics make big data have broad application prospects in various fields, but also bring challenges in data storage, processing and analysis.
[0005] Huffman Coding has been widely used in the field of data compression due to its simple and efficient characteristics; by assigning shorter codes to data with higher frequencies and longer codes to data with lower frequencies, lossless compression of data is achieved.
[0006] By reasonably utilizing the frequency information of data, Huffman Coding can significantly reduce the storage cost of data while ensuring data integrity; therefore, the core of Huffman Coding lies in the utilization of the frequency information of data, and at the same time, the efficiency of Huffman Coding is also limited by the frequency information of data, resulting in the inability to further improve the compression ratio. Summary of the Invention
[0007] To solve the above technical problem that the efficiency of Huffman Coding is also limited by the frequency information of data, resulting in the inability to further improve the compression ratio, the present invention provides solutions in the following aspects.
[0008] In a first aspect, the present invention provides a method for processing data information based on big data, including: constructing a Huffman tree according to the frequency statistics results of all data in the data information; taking any layer that contains both leaf nodes and branch nodes in the Huffman tree as the target layer, and denoting all data whose layer where the leaf node corresponding to the previous data in the data information is the target layer as identification data; taking all data whose layer number of the row where the corresponding leaf node is located in all identification data is greater than the layer number of the target layer as target data; determining the length of the fixed-length coding of the branch nodes in the target layer according to the number of all branch nodes in the target layer; determining the reduction amount according to the length of the path coding from the root node to the branch nodes in the target layer and the length of the fixed-length coding of the branch nodes in the target layer; in response to the number of identification data being less than the product of the number of target data and the reduction amount, taking the target layer as a special layer; encoding all data through the Huffman tree and the special layer, wherein the encoding method for any target data corresponding to the special layer includes: taking the branch node where the leaf node corresponding to the target data is located in the special layer as the guiding node; determining the fixed-length coding of the guiding node according to the serial number of the guiding node in the special layer; splicing the fixed-length coding of the guiding node and the path coding from the guiding node to the leaf node corresponding to the target data to obtain the encoding result of the target data.
[0009] The present invention obtains, from the Huffman tree, a special layer where the increase in data volume is less than the decrease in data volume when the path coding from the root node to the branch node is replaced with the fixed-length coding of the serial number of the branch node in its layer, ensuring that when all data are encoded through the Huffman tree and the special layer subsequently, the compression ratio of the compression result of the data information is improved; further, in the process of encoding all data through the Huffman tree, for any target data corresponding to the special layer, the path coding from the root node to the branch node of the target data in the special layer is replaced with the fixed-length coding of the serial number of the branch node in its layer, so as to shorten the encoding length of the target data, thereby improving the compression ratio of the compression result of the data information.
[0010] Preferably, the step of determining the length of the fixed-length coding of the branch nodes in the target layer according to the number of all branch nodes in the target layer includes: denoting the number of all branch nodes in the target layer as , then the length of the fixed-length coding of the branch nodes in the target layer is equal to , represents rounding up, represents the logarithmic function with base 2.
[0011] Preferably, determining the reduction amount according to the length of the path encoding of the branch node from the root node to the target layer and the length of the fixed-length encoding of the branch node of the target layer includes: the length of the path encoding of the branch node from the root node to the target layer is equal to the layer number of the target layer minus 1; then the reduction amount is equal to the difference between the length of the path encoding of the branch node from the root node to the target layer and the length of the fixed-length encoding of the branch node of the target layer.
[0012] By calculating the reduction amount, the present invention can accurately determine whether the data volume can be reduced when encoding the target data corresponding to the target layer subsequently by replacing the path encoding from the root node to the branch node with the fixed-length encoding of the serial number of the branch node in the layer where it is located, thereby determining a special layer that can improve the compression ratio of the compression result of the data information.
[0013] Preferably, encoding all data through the Huffman tree and the special layer further includes: for any one data in the data information, if the data is the identification data corresponding to the special layer but not the target data corresponding to the special layer, taking the path encoding from the root node to the leaf node corresponding to the data as the encoding result of the data.
[0014] Preferably, encoding all data through the Huffman tree and the special layer further includes: for any one data in the data information, if the data is not the identification data corresponding to the special layer, taking the path encoding from the root node to the leaf node corresponding to the data as the encoding result of the data.
[0015] Preferably, the method further includes: for any one data in the data information: if the data is the target data corresponding to the special layer, adding an identifier 1 in front of the encoding result of the data; if the data is the identification data corresponding to the special layer but not the target data corresponding to the special layer, adding an identifier 0 in front of the encoding result of the data.
[0016] By adding identifiers, when the layer where the leaf node corresponding to the previous data is located is the special layer, it is distinguished whether the path encoding from the root node to the branch node is replaced with the fixed-length encoding of the serial number of the branch node in the layer where it is located when encoding the subsequent data, thereby ensuring the accuracy of subsequent decoding.
[0017] Preferably, obtaining the frequency statistics result of all the data includes: in the data information, counting the number of times each type of data appears, where the same data belongs to the same type of data; the frequency of each type of data refers to the ratio of the number of times each type of data appears to the total number of all data in the data information.
[0018] Preferably, the method further includes: taking the frequencies of all the data as supplementary information and storing them.
[0019] The present invention stores the frequencies of all data as supplementary information, which can ensure the correct decoding of the subsequent compression results of the data information.
[0020] Preferably, the method further includes: storing the layer numbers of all special layers as supplementary information.
[0021] The present invention stores the layer numbers of all special layers as supplementary information, which can ensure the correct decoding of the subsequent compression results of the data information.
[0022] In a second aspect, the present invention provides a data information processing system based on big data, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned data information processing method based on big data is implemented.
[0023] By adopting the above technical solution, a computer program is generated according to the above-mentioned data information processing method based on big data and stored in the memory to be loaded and executed by the processor, so as to manufacture a terminal device according to the memory and the processor, which is convenient to use.
[0024] The beneficial effects of the present invention are as follows: In the present invention, from the Huffman tree, a special layer is obtained where the increase in data volume is less than the decrease in data volume when the path encoding from the root node to the branch node is replaced with the fixed-length encoding of the serial number of the branch node in the layer where it is located. In the subsequent process of encoding all data through the Huffman tree and the special layer, for any target data corresponding to the special layer, the path encoding from the root node to the branch node of the target data in the special layer is replaced with the fixed-length encoding of the serial number of the branch node in the layer where it is located, so as to shorten the encoding length of the target data, and further improve the compression ratio of the compression result of the data information. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart schematically showing a data information processing method based on big data in the present invention; Figure 2 is a schematic diagram schematically showing the constructed Huffman tree. DETAILED DESCRIPTION OF THE INVENTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0028] An embodiment of the present invention discloses a data information processing method based on big data. Refer to Figure 1 , which includes steps S1 - S3: S1. Construct a Huffman tree according to the frequency statistics results of all data in the data information.
[0029] Huffman coding is an algorithm widely used in data compression. By assigning shorter codes to data with higher frequencies and longer codes to data with lower frequencies, lossless compression of data can be achieved.
[0030] The core idea of Huffman coding is to construct a Huffman tree (HuffmanTree) according to the frequencies of all data. The Huffman tree is a binary tree, and each leaf node in the tree represents a data and its occurrence frequency; the construction process of the Huffman tree ensures that data with higher frequencies has shorter codes, while data with lower frequencies has longer codes; in this way, Huffman coding can effectively reduce the total length of data and achieve data compression.
[0031] In order to construct a Huffman tree, it is necessary to first count the types of data in the data information and the frequency of each type of data; specifically, in the data information, count the number of times each data appears, where the same data belongs to the same type of data; correspondingly, the frequency of each type of data refers to the ratio of the number of times each data appears to the total number of all data in the data information.
[0032] Furthermore, construct a Huffman tree according to the frequencies of all data; since the Huffman tree is a binary tree, the node in the first layer of the Huffman tree is the root node, and the nodes in the remaining layers are divided into leaf nodes and branch nodes. The leaf nodes have no subtrees, while the branch nodes have subtrees.
[0033] Exemplarily, for the data information {5, 6, 1, 3, 7, 4, 5, 8, 6, 6, 6, 5, 8, 6, 6, 8, 1, 6, 6, 5, 8, 1, 5, 6, 6, 8, 6, 5, 1, 3, 7, 2, 6, 1, 3, 6, 5, 8, 1, 8}, the total number of all data in the data information is 40, and there are 8 types of data, namely: 1, 2, 3, 4, 5, 6, 7, 8. The number of times each data appears is 6, 1, 3, 1, 7, 13, 2, 7 respectively, then the frequencies of each type of data are 0.15, 0.025, 0.075, 0.025, 0.175, 0.325, 0.05, 0.175 respectively; according to the frequencies of all data, the schematic diagram of the constructed Huffman tree is as Figure 2 shown.
[0034] In addition, in order to be able to correctly decode, it is necessary to store the frequency information of each data. Therefore, the frequencies of all data are used as supplementary information and stored.
[0035] S2. Obtain all special layers in the Huffman tree.
[0036] In the process of encoding all data through the Huffman tree, if the layer where the leaf node corresponding to the previous data is located (denoted as the th layer) is above the layer where the leaf node corresponding to the next data is located (denoted as the th layer), that is , then there is a corresponding branch node in the layer where the subtree of the next data is located, which is the layer where the leaf node corresponding to the previous data is located, that is, the th layer; Exemplarily, if the previous data is 6, the layer where data 6 is located is the 3rd layer, and the next data is 1, the layer where data 1 is located is the 4th layer. The subtree where data 1 is located has a corresponding branch node in the layer where data 6 is located, that is, the 3rd layer, specifically, the branch node with a frequency of 0.325 in the 3rd layer.
[0037] Conventional Huffman coding encodes the path from the root node to the next data as the coding result of the next data. When decoding subsequently, the path on the Huffman tree is determined according to the coding result, and then the next data is determined according to the path; Exemplarily, the coding result of data 1 is 100.
[0038] Since in the case where the layer where the leaf node corresponding to the previous data is above the layer where the leaf node corresponding to the next data, there is a corresponding branch node in the layer where the subtree of the next data is located in the layer where the leaf node corresponding to the previous data is located. That is to say, as long as the previous data is known, one can know the layer where a branch node of the subtree of the next data is located according to the layer where the leaf node corresponding to the previous data is located. On this basis, only by recording the serial number of the branch node in the layer where the leaf node corresponding to the previous data is located can the position of the branch node be determined. Furthermore, according to the position of the branch node and the path coding from the branch node to the next data, the path on the Huffman tree can be determined, and then the next data can be determined according to the path.
[0039] Therefore, the path coding from the root node to the next data is split into the path coding from the root node to the branch node (the branch node corresponding to the next data in the layer where the leaf node corresponding to the previous data is located) and the path coding from the branch node to the next data. Among them, the length of the path coding from the root node to the branch node is equal to , represents the number of layers of the branch node, that is, the number of layers of the layer where the leaf node corresponding to the previous data is located; At this time, through the branch node in the The fixed-length coding of the serial numbers of the layers is used to replace the path coding from the root node to the branch node, where the branch node is in the The length of the fixed-length coding of the serial numbers of the layers is equal to , indicating the number of all branch nodes in the layer. That is to say, as long as the number of all branch nodes in the layer is satisfied, the coding length of the latter data can be shortened through the replacement operation.
[0040] Exemplarily, when the previous data is 6 and the latter data is 1, it satisfies the situation that the layer where the leaf node corresponding to the previous data is located (the 3rd layer) is above the layer where the leaf node corresponding to the latter data is located (the 4th layer). At this time, the subtree where the latter data 1 is located has a corresponding branch node in the layer where the previous data 6 is located, that is, the branch node with a frequency of 0.325 in the 3rd layer; since there is only one branch node in the 3rd layer, therefore, there is no need to encode the serial number of this branch node. That is to say, the length of the fixed-length coding of the serial number of the branch node in the 3rd layer is equal to , indicating the number of all branch nodes in the 3rd layer, and ; at this time, the number of all branch nodes in the 3rd layer is satisfied , so through the replacement operation, the coding length of the latter data can be shortened.
[0041] In summary, in the subsequent process of encoding data through the Huffman tree, if you want to shorten the coding length of the data by replacing the path coding from the root node to the branch node with the fixed-length coding of the serial number of the branch node in the layer where it is located, the data needs to meet two conditions: Condition 1: It is required that the number of all branch nodes in the layer where the leaf node corresponding to the previous data of this data is located (the layer) is satisfied ; among them, since the layer is the layer where the leaf node corresponding to the previous data is located, that is to say, Condition 1 also includes a hidden condition, that is, it is required that there must be leaf nodes in the layer.
[0042] Condition 2: It is required that the layer where the leaf node corresponding to the previous data of this data is located (the layer) is above the layer where the leaf node corresponding to this data is located (the layer), that is ; among them, since the layer where the leaf node corresponding to the previous data is located (the layer) is above the layer where the leaf node corresponding to the data is located (the layer), that is to say, condition 2 also includes a hidden condition, that is, it is required that the layer must have branch nodes.
[0043] For the hidden conditions in condition 1 and condition 2, it only involves the layer number and the number of leaf nodes and branch nodes in the layer. Therefore, it is possible to directly judge each layer of the Huffman tree to obtain the alternative layers that meet the conditions: it is required that there are branch nodes and leaf nodes in the alternative layer, and the number of all branch nodes in the alternative layer meets .
[0044] Specifically, for the constructed Huffman tree, for the layer in the Huffman tree, obtain the number of all leaf nodes and the number of all branch nodes in the layer, and use it to judge whether the layer is used as an alternative layer. The specific operation process is as follows: 1. If there are both leaf nodes and branch nodes in the layer, and the number of all branch nodes in the layer meets , use the layer as an alternative layer, indicating rounding up.
[0045] 2. If there are no leaf nodes in the layer, or there are no branch nodes in the layer, or the number of all branch nodes in the layer meets , then the layer cannot be used as an alternative layer.
[0046] Exemplarily, for the Huffman tree shown in Figure 2 , the process of obtaining all alternative layers in the Huffman tree is as follows: (1) Since there are no leaf nodes in the 1st layer and the 2nd layer, the 1st layer and the 2nd layer cannot be used as alternative layers.
[0047] (2) Since there are both leaf nodes and branch nodes in the 3rd layer, and the number of all branch nodes in the 4th layer = 1, then , so the number of all branch nodes in the 3rd layer meets , then use the 3rd layer as an alternative layer.
[0048] (3) Similarly, since there are both leaf nodes and branch nodes in the 4th, 5th, and 6th layers, and the number of all branch nodes in the 4th layer = 1, the number of all branch nodes in the 5th layer = 1, the number of all branch nodes in the 6th layer = 1, , , , so the number of all branch nodes in the 4th, 5th, and 6th layers , and all satisfy , then the 4th, 5th, and 6th layers are used as alternative layers.
[0049] (4) Since there are no branch nodes in the 7th layer, the 7th layer cannot be used as an alternative layer.
[0050] It should be noted that by screening all layers through the number of branch nodes to select alternative layers, the number of layers that need to be judged and compared when obtaining special layers subsequently can be reduced, improving the data compression efficiency.
[0051] In addition, when the layer where the leaf node corresponding to the previous data of the data is an alternative layer, an identifier is needed to identify whether the layer where the leaf node corresponding to the previous data of the data is above the layer where the leaf node corresponding to this data is located, so as to further determine whether to shorten the encoding length of the data by replacing the path encoding from the root node to the branch node with the fixed-length encoding of the serial number of the branch node in its layer; since there are only two cases: the layer where the leaf node corresponding to the previous data of the data is above the layer where the leaf node corresponding to this data is located or the layer where the leaf node corresponding to the previous data of the data is not above the layer where the leaf node corresponding to this data is located, therefore, only an identifier with a length of 1 needs to be introduced: when the layer where the leaf node corresponding to the previous data of the data is above the layer where the leaf node corresponding to this data is located, that is , the identifier is 1, when the layer where the leaf node corresponding to the previous data of the data is not above the layer where the leaf node corresponding to this data is located, that is , the identifier is 0.
[0052] Therefore, when encoding the data information subsequently, for the layer where the leaf node corresponding to the previous data is located (the The (layer) is the data of the alternative layer. If the path encoding from the root node to the branch node is replaced with the fixed-length encoding of the serial number of the branch node in its layer, an identifier with a length of 1 needs to be added, which will result in an increase in the amount of data in the encoded compression result of the final data information. Moreover, the increase in the amount of data is equal to the number of all data whose leaf node corresponding to the previous data is in the alternative layer multiplied by the length of the identifier, that is, the increase in the amount of data is equal to , is the number of all data whose leaf node corresponding to the previous data in the data information is in the alternative layer (the layer).
[0053] In addition, for all data whose leaf node corresponding to the previous data is in the alternative layer, only when the layer (the layer) where the leaf node corresponding to the previous data of the data is above the layer (the layer) where the leaf node corresponding to this data is, that is, only when, will the path encoding from the root node to the branch node of this data in the layer be replaced with the fixed-length encoding of the serial number of the branch node in its layer to shorten the encoding length of the data; since the length of the path encoding from the root node to the branch node of this data in the layer is equal to , and the length of the fixed-length encoding of the serial number of the branch node of this data in the layer is equal to , is the number of all branch nodes in the layer. Therefore, the reduction in the encoding length of this data is equal to ; In summary, when the alternative layer (the layer) is used as the special layer, the reduction in the amount of data is equal to the number of all data whose leaf node corresponding to the previous data is in the alternative layer (the layer) and that satisfy multiplied by the reduction in the encoding length of each data , that is, the reduction in the amount of data is equal to . .
[0054] In summary, if the increase in the amount of data is greater than or equal to the reduction , replacing the path encoding from the root node to the branch node with the fixed-length encoding of the serial number of the branch node in its layer cannot shorten the length of the encoded result of the data information; only when the increase in the amount of data is less than the reduction Only when the above conditions are met, can the length of the encoded result of the data information be shortened by replacing the path encoding from the root node to the branch node with the fixed-length encoding of the serial number of the branch node in its layer. Therefore, when the increase in data volume is less than the decrease, the alternative layer meeting this condition is regarded as a special layer. For the data where the layer where the leaf node corresponding to the previous data is located is a special layer, only then will we consider replacing the path encoding from the root node to the branch node with the fixed-length encoding of the serial number of the branch node in its layer.
[0055] Specifically, for the alternative layer with the layer number of , the specific operation process for determining whether to regard the alternative layer as a special layer is as follows: 1. Obtain all the data in the data information where the layer where the leaf node corresponding to the previous data is located is the th layer, and denote it as the marked data; denote the quantity of all the marked data as .
[0056] 2. Denote all the marked data that meet the condition that the layer number of their corresponding rows is greater than as the target data, and denote the quantity of all the target data as .
[0057] 3. If is less than , it indicates that when the th layer is regarded as a special layer, the increase in data volume is less than the decrease in data volume. Therefore, the th layer is regarded as a special layer, where is the quantity of all the branch nodes in the th layer.
[0058] 4. If is greater than or equal to , it indicates that when the th layer is regarded as a special layer, the increase in data volume is greater than or equal to the decrease in data volume. Therefore, the th layer cannot be regarded as a special layer.
[0059] Exemplarily, when the data information is {5, 6, 1, 3, 7, 4, 5, 8, 6, 6, 6, 5, 8, 6, 6, 8, 1, 6, 6, 5, 8, 1, 5, 6, 6, 8, 6, 5, 1, 3, 7, 2, 6, 1, 3, 6, 5, 8, 1, 8}, for the Huffman tree as shown in Figure 2 , all the alternative layers in the obtained Huffman tree include the 3rd layer, the 4th layer, the 5th layer, and the 6th layer. The process for determining whether to regard the alternative layer as a special layer is as follows: (1)For the 3rd layer, which includes 3 leaf nodes and 1 branch node, the data corresponding to the 3 leaf nodes are Data 5, Data 6, and Data 8 respectively; obtain all the data in the data information where the layer of the leaf node corresponding to the previous data is the 3rd layer, that is, all the data where the previous data is Data 5 / Data 6 / Data 8, a total of 26, and denote them as marked data. Then the quantity of all marked data = 26; among these 26 marked data, the number of target data that satisfy the layer number of the corresponding row being greater than 3 is 6. Then = 6; the 3rd layer includes 1 branch node. Then = 1. Correspondingly, = 12; since = 26, therefore, is greater than , indicating that when the 3rd layer is used as the special layer, the increase in the data volume is greater than the decrease in the data volume. Therefore, the 3rd layer cannot be used as the special layer.
[0060] (2)For the 4th layer, which includes 1 leaf node and 1 branch node, the data corresponding to the 1 leaf node is Data 1; obtain all the data in the data information where the layer of the leaf node corresponding to the previous data is the 4th layer, that is, all the data where the previous data is Data 1, a total of 6, and denote them as marked data. Then the quantity of all marked data = 6; among these 6 marked data, the number of target data that satisfy the layer number of the corresponding row being greater than 4 is 3. Then = 3; the 4th layer includes 1 branch node. Then = 1. Correspondingly, = 9; since = 6, therefore, is less than , indicating that when the 4th layer is used as the special layer, the increase in the data volume is less than the decrease in the data volume. Therefore, the 4th layer is used as the special layer.
[0061] (3)Similarly, for the 5th layer, , , , therefore, = 8, is less than , indicating that when the 5th layer is used as the special layer, the increase in the data volume is less than the decrease in the data volume. Therefore, the 5th layer is used as the special layer; for the 6th layer, , , , therefore, = 10, is less than , indicating that when the 6th layer is used as the special layer, the increase in the data volume is less than the decrease in the data volume. Therefore, the 6th layer is used as the special layer.
[0062] In addition, in order to be able to correctly decode, the layer numbers of each special layer need to be stored. Therefore, the layer numbers of all special layers are used as supplementary information and stored.
[0063] S3. Encode all data according to the Huffman tree and the special layers in the Huffman tree, obtain the compression result of the data information and store it.
[0064] Huffman coding encodes each data by the path coding from the root node of the Huffman tree to the leaf node corresponding to each data, so it is necessary to assign the encodings of 0 and 1 to the paths in the Huffman tree; therefore, in one embodiment, starting from the root node of the Huffman tree, assign the encoding "0" to the left subtree and the encoding "1" to the right subtree until reaching the leaf node; in another embodiment, starting from the root node of the Huffman tree, assign the encoding "1" to the left subtree and the encoding "0" to the right subtree until reaching the leaf node.
[0065] Further, encode all data through the Huffman tree. During the encoding process, for any one data: (1) If the data is the target data corresponding to the special layer, then use the branch node of the leaf node corresponding to the data in the special layer as the guiding node; determine the fixed-length encoding of the guiding node according to the sequence number of the guiding node in the special layer; splice the fixed-length encoding of the guiding node and the path encoding from the guiding node to the leaf node corresponding to the target data to obtain the encoding result of the target data; at the same time, add an identifier 1 in front of the encoding result of the data.
[0066] Exemplarily, when the previous data is 1 and the next data is 3, the leaf node corresponding to the previous data is in the 4th layer, and the leaf node corresponding to the next data is in the 5th layer; since the layer where the leaf node corresponding to the previous data is located, that is, the 4th layer, is a special layer, so judge the size relationship between the layer numbers of the layer where the leaf node corresponding to the previous data is located and the layer where the leaf node corresponding to the next data is located. Since the layer number of the layer where the leaf node corresponding to the previous data is 4 and the layer number of the layer where the leaf node corresponding to the next data is 5, and 5>4, so the next data is the target data corresponding to the special layer, and obtain the branch node corresponding to the subtree where the leaf node is located in the 4th layer, that is, the branch node with a frequency of 0.175 in the 4th layer, denoted as the guiding node ; since there is only one branch node in the 4th layer, it is not necessary to process the guiding node Encode according to the serial number; obtain the leading node to the leaf node Path encoding is 0, the encoding result of the next data is 0, and an identifier 1 is added in front of the encoding result of the next data. Therefore, the encoding result of data 6 is 01.
[0067] (2) If the data is the identification data corresponding to the special layer but not the target data corresponding to the special layer, use the path encoding from the root node to the leaf node corresponding to the data as the encoding result of the data; at the same time, add the identifier 0 in front of the encoding result of the data.
[0068] Exemplarily, when the previous data is 1 and the next data is 5, the leaf node corresponding to the previous data is in the 4th layer, and the leaf node corresponding to the next data is in the 3rd layer; since the leaf node corresponding to the previous data is in the 4th layer which is the special layer, so judge the leaf node corresponding to the previous data and the leaf node corresponding to the next data The size relationship of the layer numbers of the layers where they are located. Since the leaf node corresponding to the previous data is in the 4th layer, and the leaf node corresponding to the next data is in the 3rd layer, 3 < 4, so the next data is the identification data corresponding to the special layer but not the target data corresponding to the special layer. Use the path encoding from the root node to the leaf node as the encoding result of the next data, then the encoding result of data 6 is 11, and an identifier 0 is added in front of the encoding result of the next data. Therefore, the encoding result of data 6 is 011.
[0069] (3) If the data is not the identification data corresponding to the special layer, use the path encoding from the root node to the leaf node corresponding to the data as the encoding result of the data.
[0070] Exemplarily, when the previous data is 5 and the next data is 6, the leaf node corresponding to the previous data is in the 3rd layer. Since the 3rd layer is not a special layer, directly use the path encoding from the root node to the leaf node as the encoding result of the next data, then the encoding result of data 6 is 11.
[0071] Exemplarily, when the data information is {5, 6, 1, 3, 7, 4, 5, 8, 6, 6, 6, 5, 8, 6, 6, 8, 1, 6, 6, 5, 8, 1, 5, 6, 6, 8, 6, 5, 1, 3, 7, 2, 6, 1, 3, 6, 5, 8, 1, 8}, the data information is encoded by conventional Huffman coding and the method of this embodiment respectively, and the results are as follows: (1) When the data information is encoded by conventional Huffman coding, the encoding results of each data are 00, 11, 100, 1010, 10111, 101101, 00, 01, 11, 11, 11, 00, 01, 11, 11, 01, 100, 11, 11, 00, 01, 100, 00, 11, 11, 01, 11, 00, 100, 1010, 10111, 101100, 11, 100, 1010, 11, 00, 01, 100, 01 respectively. The encoding result of this data information is 0011100101010111101101000111111100011111011001111000110000111101110010010101011110110011100101011000110001, and the length of the encoding result of this data information is equal to 106.
[0072] (2) When the data information is encoded by the method of this embodiment, the encoding results of each data are 00, 11, 100, 10, 11, 11, 00, 01, 11, 11, 11, 00, 01, 11, 11, 01, 100, 011, 11, 00, 01, 100, 000, 11, 11, 01, 11, 00, 100, 10, 11, 10, 11, 100, 10, 011, 00, 01, 100, 001 respectively. The encoding result of this data information is 001110010111100011111110001111101100011110001100000111101110010010111011100100110001100001, and the length of the encoding result of this data information is equal to 90.
[0073] (3) In summary, when the data information is encoded by conventional Huffman coding, the length of the encoding result of this data information is equal to 106; when the data information is encoded by the method of this embodiment, the length of the encoding result of this data information is equal to 90; therefore, the method of this embodiment can improve the compression ratio of the compression result of the data information.
[0074] It should be noted that during the process of compressing and encoding data information through Huffman coding, when obtaining a fixed-length coding by replacing the path coding from the root node to a branch node with the sequence number of the branch node in its layer from the Huffman tree, there is a special layer where the increase in data volume is less than the decrease in data volume. Thus, when subsequently encoding all data through the Huffman tree and the special layer, for the next data that satisfies the condition that the layer where the leaf node corresponding to the previous data is located is the special layer and the layer number of the corresponding leaf node is greater than the layer number of the special layer, during encoding, the path coding from the root node to the branch node of the next data in the special layer is replaced with the fixed-length coding of the sequence number of the branch node in its layer, so as to shorten the coding length of the next data, and further improve the compression ratio of the compression result of the data information.
[0075] Furthermore, when it is necessary to view or use the data information, the compression result of the data information is decoded. The specific process is as follows: 1. Construct a Huffman tree according to the frequencies of all the data stored in the supplementary information.
[0076] 2. Decode the compression result of the data information according to the Huffman tree and the special layer in the Huffman tree. During the decoding process, mark the just-decoded data as the previous data, and mark the next data to be decoded as the next data; obtain the layer where the leaf node corresponding to the previous data is located and denote it as the th layer. The specific decoding process is as follows: 2.1. Determine whether the th layer is the special layer: If the th layer is not the special layer, directly obtain the path coding according to the remaining compression result and the Huffman tree, and obtain the leaf node according to the path coding, and take the data corresponding to the leaf node as the next data.
[0077] 2.2. If the th layer is the special layer, determine whether, during encoding, the path coding from the root node to the branch node is replaced with the fixed-length coding of the sequence number of the branch node in its layer according to whether the first binary data in the remaining compression result is 0 or 1: (1) If it is 0, it means that during encoding, the path coding from the root node to the branch node is not replaced with the fixed-length coding of the sequence number of the branch node in its layer. Directly obtain the path coding according to the remaining compression result and the Huffman tree, and obtain the leaf node according to the path coding, and take the data corresponding to the leaf node as the next data.
[0078] (2) If it is 0, it means that during encoding, the path encoding from the root node to the branch node is replaced with a fixed-length encoding of the serial number of the branch node in its layer. First, obtain the number of all branch nodes in the layer, and then obtain a binary number with a length of from the very front of the remaining compression result. Use the decimal number corresponding to this binary number as the serial number to obtain the branch node. According to the branch node and the remaining compression result, obtain the path encoding from the Huffman tree, and obtain the leaf node . Use the data corresponding to the leaf node as the subsequent data.
[0079] An embodiment of the present invention also discloses a data information processing system based on big data, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a data information processing method based on big data according to the present invention is implemented.
[0080] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.
Claims
1. A data information processing method based on big data, characterized in that Including: Construct a Huffman tree according to the frequency statistics results of all data in the data information; Take any layer in the Huffman tree that contains both leaf nodes and branch nodes as the target layer. Denote all the data whose layer where the leaf node corresponding to the previous data in the data information is the target layer as the identification data; Take all the data whose layer number of the row where the corresponding leaf node in all the identification data is greater than the layer number of the target layer as the target data; Determine the length of the fixed-length coding of the branch nodes in the target layer according to the number of all branch nodes in the target layer; Determine the reduction amount according to the length of the path coding from the root node to the branch nodes in the target layer and the length of the fixed-length coding of the branch nodes in the target layer; In response to the number of identification data being less than the product of the number of target data and the reduction amount, take the target layer as the special layer; Encode all the data through the Huffman tree and the special layer. Among them, the encoding method for any one of the target data corresponding to the special layer includes: Taking the branch node where the leaf node corresponding to the target data is located in the special layer as the guiding node; Determine the fixed-length coding of the guiding node according to the serial number of the guiding node in the special layer; Concatenate the fixed-length coding of the guiding node with the path coding from the guiding node to the leaf node corresponding to the target data to obtain the encoding result of the target data.
2. The data information processing method based on big data according to claim 1, wherein The determining the length of the fixed-length coding of the branch nodes in the target layer according to the number of all branch nodes in the target layer includes: Denote the number of all branch nodes in the target layer as , then the length of the fixed-length encoding of the branch nodes in the target layer is equal to , denotes rounding up, denotes the logarithmic function with base 2.
3. A data information processing method based on big data according to claim 1, characterized in that, The determining the reduction amount according to the length of the path coding from the root node to the branch nodes in the target layer and the length of the fixed-length coding of the branch nodes in the target layer includes: The length of the path coding from the root node to the branch nodes in the target layer is equal to the layer number of the target layer minus 1; Then the reduction amount is equal to the difference between the length of the path coding from the root node to the branch nodes in the target layer and the length of the fixed-length coding of the branch nodes in the target layer.
4. A data information processing method based on big data according to claim 1, characterized in that, The encoding all the data through the Huffman tree and the special layer further includes: For any one of the data in the data information, if the data is the identification data corresponding to the special layer but not the target data corresponding to the special layer, take the path coding from the root node to the leaf node corresponding to the data as the encoding result of the data.
5. A data information processing method based on big data according to claim 1, characterized in that The encoding all the data through the Huffman tree and the special layer further includes: For any one of the data in the data information, if the data is not the identification data corresponding to the special layer, take the path coding from the root node to the leaf node corresponding to the data as the encoding result of the data.
6. A data information processing method based on big data according to claim 1, characterized in that, The method further includes: For any one of the data in the data information: If the data is the target data corresponding to the special layer, add an identifier 1 in front of the encoding result of the data; If the data is the identification data corresponding to the special layer but not the target data corresponding to the special layer, add an identifier 0 in front of the encoding result of the data.
7. A data information processing method based on big data according to claim 1, characterized in that, Obtaining the frequency statistics results of all the data includes: In the data information, count the number of occurrences of each type of data, where the same data belongs to the same type of data; The frequency of each type of data refers to the ratio of the number of occurrences of each type of data to the total number of all data in the data information.
8. A data information processing method based on big data according to claim 1, characterized in that The method further includes: taking the frequencies of all data as supplementary information and storing them.
9. A data information processing method based on big data according to claim 1, characterized in that, The method further includes: taking the layer numbers of all special layers as supplementary information and storing them.
10. A data information processing system based on big data, characterized in that, It includes: a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a data information processing method based on big data according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Variable length encoding of compressed data
CA2388006A1
Data compression coding method and device and storage electronic equipment
CN109981111A
Data compression method and device, electronic equipment and storage medium
CN113746487A
Inspection method and system for power transformation equipment based on image recognition
CN118314481A
Variable length encoding method, variable length decoding method, storage medium, variable length encoding device, variable length decoding device, and bit stream
US20050012647A1