Huffman encoding method and encoding apparatus, chip and storage medium

WO2026201005A1PCT designated stage Publication Date: 2026-10-01THEO END (SHENZHEN) COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/086035
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure CN2026086035_01102026_PF_FP_ABST
    Figure CN2026086035_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a Huffman encoding method and encoding apparatus, a chip and a storage medium. The Huffman encoding method comprises: acquiring a character string to be encoded and associated encoding information, the encoding information comprising the length of the character string to be encoded, an alphabet to which the character string to be encoded belongs, and a code length limit for codes to be generated; scanning symbols of the character string on the basis of a preset scan length, counting the frequency of occurrence of each of the scanned symbols, and generating a symbol frequency table that corresponds to the character string to be encoded and is based on the preset scan length; according to the preset scan length and the code length limit, determining a reference frequency for frequency adjustment, so as to optimize the symbol frequency table on the basis of the reference frequency; according to the optimized symbol frequency table, generating a Huffman codebook for the character string to be encoded; and encoding, by using the Huffman codebook, the character string to be encoded, and outputting a corresponding encoding result. The present invention can adapt to different scanning progress levels and construct codebooks satisfying code length limits, thereby effectively improving Huffman encoding speed.
Need to check novelty before this filing date? Find Prior Art

Description

A Huffman coding method, coding device, chip, and storage medium

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510370509.3, filed on March 26, 2025, entitled “A Huffman Encoding Method, Encoding Device, Chip and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to the field of communication computing technology, and in particular to a Huffman coding method, coding device, chip, and storage medium. Background Technology

[0004] Huffman coding is a type of prefix coding with compression effects. It obtains a codebook based on statistical distribution by assigning shorter codewords to more frequent symbols and longer codewords to less frequent symbols. Based on this codebook, the original string can be encoded and uniquely recovered through decoding. Therefore, Huffman coding can reduce the length of the encoded string to less than the length of the original string, thus achieving compression.

[0005] In practical applications, the maximum code length in Huffman codes is generally limited. This is beneficial for preserving the codebook and for constructing compatible encoding schemes. For example, in the Deflate standard (RFC 1951), the maximum code length of Huffman codes for literal / length or distance alphabets is limited to 15. Furthermore, Huffman codes are a type of entropy coding, requiring statistical analysis of symbol frequencies to construct a codebook with better compression. However, in implementation, scanning and statistical analysis of the data to be compressed is necessary, which limits the speed of Huffman coding.

[0006] Therefore, existing encoding methods cannot complete the statistical analysis of symbol frequencies with fewer scans, or can not quickly generate codebooks that meet the length limit. Summary of the Invention

[0007] The purpose of this invention is to propose an encoding scheme that at least partially solves one of the aforementioned technical problems.

[0008] To achieve the above objectives, a first aspect of the present invention provides a Huffman encoding method, the Huffman encoding method comprising: acquiring a string to be encoded and associated encoding information, the encoding information including the length of the string to be encoded, the alphabet to which it belongs, and the restricted code length to be encoded; scanning the symbols of the string based on a preset scan length, statistically analyzing the frequency of occurrence of each symbol in the scanned symbols, and generating a symbol frequency table corresponding to the string to be encoded based on the preset scan length; determining a reference frequency for frequency adjustment based on the preset scan length and the restricted code length, and optimizing the symbol frequency table based on the reference frequency; generating a Huffman codebook for the string to be encoded based on the optimized symbol frequency table; encoding the string to be encoded using the Huffman codebook, and outputting the corresponding encoding result.

[0009] According to one embodiment of the present invention, optimizing the symbol frequency table based on the reference frequency includes: traversing the occurrence frequency of each statistically recorded symbol, comparing the occurrence frequency of the symbol with the reference frequency, and if the occurrence frequency of the symbol is less than the reference frequency, adjusting the occurrence frequency of the symbol to the reference frequency; otherwise, keeping the occurrence frequency of the symbol unchanged.

[0010] According to one embodiment of the present invention, the reference frequency is calculated using the following formula:

[0011] Where p is the reference frequency, F k+3 Here, k is the Fibonacci number, k is the restricted code length, and n is the preset scan length. The alphabet to which the string to be encoded belongs.

[0012] According to one embodiment of the present invention, generating a Huffman codebook for the string to be encoded based on the optimized symbol frequency table includes: sorting the symbol frequencies in the adjusted symbol frequency table in ascending order and recording the position information of each symbol frequency after sorting; cyclically merging the frequencies of the symbols according to the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol of the string to be encoded, so as to generate a code length table for the string to be encoded; and generating a Huffman codebook for encoding the string to be encoded based on the code length table.

[0013] According to one embodiment of the present invention, the step of cyclically merging the frequencies of the symbols based on the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol in the string to be encoded, so as to generate a code length table of the string to be encoded, includes: obtaining and merging the frequencies of the two lowest-frequency symbols in the symbol frequency table; adjusting the symbol frequency sequence based on the merged frequencies, while maintaining the ascending order of the symbol frequency sequence; incrementally adjusting and updating the merged symbol positions by adjusting the code lengths of the merged symbols; repeating the above process until the last symbol is merged to obtain the Huffman code lengths of all symbols; and adjusting the code length order of the symbols using a permutation function so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

[0014] According to one embodiment of the present invention, generating a codebook for encoding the string to be encoded based on the code length table includes: traversing the code length table and counting the occurrence frequency of each code length; determining the starting prefix of the code length based on the occurrence frequency of the code length; for each symbol of the string to be encoded, converting the starting prefix of its corresponding code length into a binary representation, and truncating the low-order bits based on the corresponding code length as the codeword of the symbol; incrementally updating the starting prefix of the corresponding code length to assign different codewords to the next symbol with that code length based on the updated starting prefix.

[0015] According to an embodiment of the present invention, the Huffman method further includes: before generating the codebook, obtaining the maximum code length in the code length table; comparing the maximum code length with the restricted code length; if the maximum code length is greater than the restricted code length, adjusting the code length table so that all code lengths in the code length table are less than or equal to the restricted code length.

[0016] According to one embodiment of the present invention, adjusting the code length table includes: obtaining two symbols in the code length table corresponding to a first preset code length, wherein the first preset code length is associated with the maximum code length in the code length table; obtaining the symbol in the code length table that is closest to a second preset code length, wherein the second preset code length is associated with the restricted code length; adjusting the code length of the obtained symbols based on the maximum code length, the code length corresponding to the symbol closest to the second preset code length, a preset code length increment depth, and a increment depth; repeating the above process to cyclically adjust the code length table until the maximum code length in the code length table is less than or equal to the restricted code length.

[0017] According to one embodiment of the present invention, the frequency of occurrence of each symbol among the scanned symbols is characterized by the following formula:

[0018] Among them, I j,iIt is an indicator variable, and this indicator variable is characterized by the following formula:

[0019] Among them, f i Let be the frequency of occurrence of symbol i, and n be the preset scan length, s j Let j be the sequence of symbols being scanned.

[0020] To achieve the above objectives, the present invention provides a Huffman encoding device, comprising: an acquisition unit for acquiring a string to be encoded and associated encoding information, the encoding information including the length of the string to be encoded, the alphabet to which it belongs, and the restricted code length to be encoded; a statistics unit for scanning the symbols of the string based on a preset scan length, counting the occurrence frequency of each symbol in the scanned symbols, and generating a symbol frequency table corresponding to the string to be encoded based on the preset scan length; an optimization unit for determining a reference frequency for frequency adjustment based on the preset scan length and the restricted code length, and optimizing the symbol frequency table based on the reference frequency; a generation unit for generating a Huffman codebook for the string to be encoded based on the optimized symbol frequency table; and an output unit for encoding the string to be encoded using the Huffman codebook and outputting the corresponding encoding result.

[0021] According to one embodiment of the present invention, the optimization unit is used to optimize the symbol frequency table based on the reference frequency, including: traversing the occurrence frequency of each statistically recorded symbol, comparing the occurrence frequency of the symbol with the reference frequency, and if the occurrence frequency of the symbol is less than the reference frequency, adjusting the occurrence frequency of the symbol to the reference frequency; otherwise, keeping the occurrence frequency of the symbol unchanged.

[0022] According to one embodiment of the present invention, the reference frequency is calculated using the following formula:

[0023] Where p is the reference frequency, F k+3 Here, k is the Fibonacci number, k is the restricted code length, and n is the preset scan length. The alphabet to which the string to be encoded belongs.

[0024] According to an embodiment of the present invention, the generation unit is configured to generate a Huffman codebook for the string to be encoded based on an optimized symbol frequency table, comprising: sorting the symbol frequencies in the adjusted symbol frequency table in ascending order and recording the position information of each symbol frequency after sorting; determining the Huffman code length of each symbol in the string to be encoded by cyclically merging the frequencies of the symbols based on the sorted symbol frequencies and the position information, so as to generate a code length table for the string to be encoded; and generating a Huffman codebook for encoding the string to be encoded based on the code length table.

[0025] According to one embodiment of the present invention, the generation unit is used to cyclically merge the frequencies of the symbols based on the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol in the string to be encoded, so as to generate a code length table of the string to be encoded, including: obtaining and merging the two symbols with the lowest frequencies in the symbol frequency table; adjusting the symbol frequencies based on the frequencies after merging the two symbols, while maintaining the ascending order of the symbol frequencies; incrementally adjusting and updating the positions of the merged symbols by adjusting the code lengths of the merged symbols; repeating the above process until merging to the last symbol to obtain the Huffman code lengths of all symbols; and adjusting the code length order of all symbols using a permutation function so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

[0026] According to an embodiment of the present invention, the generation unit is used to generate a codebook for encoding the string to be encoded based on the code length table, including: traversing the code length table and counting the occurrence frequency of each code length; determining the starting prefix of the code length based on the occurrence frequency of the code length; for each symbol of the string to be encoded, converting the starting prefix of its corresponding code length into a binary representation, and truncating the low bits based on the corresponding code length as the codeword of the symbol; and incrementally updating the starting prefix of the corresponding code length to assign different codewords to the next symbol with that code length.

[0027] According to an embodiment of the present invention, the Huffman coding apparatus further includes an adjustment unit, the adjustment unit being configured to: obtain the maximum code length in the code length table before generating the codebook; compare the maximum code length with the restricted code length; and if the maximum code length is greater than the restricted code length, adjust the code length table so that all code lengths in the code length table are less than or equal to the restricted code length.

[0028] According to an embodiment of the present invention, the adjustment unit is used to adjust the code length table, comprising: acquiring two symbols in the code length table corresponding to a first preset code length, wherein the first preset code length is associated with the maximum code length in the code length table; acquiring a symbol in the code length table corresponding to a second preset code length, wherein the second preset code length is associated with the restricted code length; adjusting the code length of the acquired symbols based on the first preset code length, the second preset code length, and preset code length increment and decrement depths; repeating the above process to cyclically adjust the code length table until the maximum code length in the code length table is less than or equal to the restricted code length.

[0029] According to one embodiment of the present invention, the frequency of occurrence of each symbol among the scanned symbols is characterized by the following formula:

[0030] Among them, I j,i It is an indicator variable, and this indicator variable is characterized by the following formula:

[0031] Among them, f i Let be the frequency of occurrence of symbol i, and n be the preset scan length, s j Let j be the sequence of symbols being scanned.

[0032] To achieve the above objectives, a third aspect of the present invention provides a chip comprising: a processor and a memory; wherein the memory stores a program capable of running on the processor, and the processor is configured to implement the Huffman coding method described in the first aspect when executing the program.

[0033] To achieve the above objectives, a fourth aspect of the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, wherein the computer instructions are used to cause the computer to execute the Huffman coding method according to the first aspect.

[0034] It is evident that the present invention has the following advantages compared to the prior art:

[0035] 1) The Huffman tree only needs to be constructed once, the coding scheme is simple, the computational complexity is low, and it is suitable for hardware implementation;

[0036] 2) It can adapt to different scanning progress and construct a codebook that meets the code length limit, which effectively improves the speed of Huffman coding;

[0037] 3) It achieves a balance between coding efficiency and performance, making the overall coding scheme more comprehensive and efficient.

[0038] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0039] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0040] Figure 1 is a schematic flowchart illustrating a Huffman coding method according to an exemplary embodiment;

[0041] Figure 2 is a schematic diagram of a Huffman tree structure according to an exemplary embodiment;

[0042] Figure 3 is a schematic diagram of another Huffman tree structure according to an exemplary embodiment;

[0043] Figure 4 is a schematic diagram of yet another Huffman tree structure according to an exemplary embodiment; and

[0044] Figure 5 is a schematic block diagram of a Huffman coding device according to an exemplary embodiment. Detailed Implementation

[0045] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0046] Before describing and illustrating the present invention, a brief explanation of the relevant concepts mentioned below will be provided to facilitate a better understanding of the invention.

[0047] The alphabet refers to the set of characters to be encoded, which is the basis of the Huffman coding algorithm.

[0048] A Huffman tree, also known as an optimal binary tree, is a binary tree with the shortest weighted path length. Its core concepts include path, path length, weight, and weighted path length. A path is the branch between two nodes in the tree, and the path length is the number of branches on the path. The weighted path length of a tree is the sum of the weighted path lengths of all leaf nodes in the tree.

[0049] Code length refers to the length of the encoded sequence obtained after encoding each character.

[0050] A codebook is actually a data structure or table that stores the correspondence between characters and Huffman codes.

[0051] In Huffman coding, it is usually necessary to statistically analyze the frequency of all symbols in the data to be compressed. This results in a large storage space required in the hardware implementation, limiting the coding speed. Therefore, this invention aims to propose a Huffman coding scheme that reduces the storage space required by the hardware by statistically analyzing partial frequencies. This not only enables the construction of a codebook that meets the length limit but also improves the coding speed and adapts to different scanning progress.

[0052] Specifically, Figure 1 is a schematic flowchart illustrating a Huffman coding method according to an exemplary embodiment. As shown in Figure 1, the coding method includes:

[0053] Step S110: Obtain the string to be encoded and its associated encoding information, wherein the encoding information includes the length of the string to be encoded, the alphabet to which it belongs, and the limited code length to be encoded.

[0054] For example, let the alphabet be... The string is S, where the alphabet contains all possible characters. This invention is based on the alphabet. A string S of length m can be represented as:

[0055] s0, s1, ..., s m-1 ,

[0056] in, To avoid loss of generality, we can denote this as a set. for

[0057] This invention requires Huffman encoding of the string S, while ensuring that the length of the Huffman codeword corresponding to each symbol in the string is less than or equal to a predetermined length k, where k > 0. This predetermined length k is the codeword limit. Without loss of generality, this invention sets the codeword limit to satisfy the following conditions: Otherwise, the number of possible leaf nodes in a Huffman tree would be less than the number of symbols, making it impossible to construct a Huffman tree and thus impossible to obtain Huffman coding.

[0058] Step S120: Scan the symbols of the string based on a preset scan length, count the frequency of each symbol in the scanned symbols, and generate a symbol frequency table corresponding to the string to be encoded based on the preset scan length.

[0059] For example, Huffman coding generally requires counting the occurrence probability of symbols to allocate different code lengths based on the probability of the symbols. In the present invention, the scanning length can be set as n, 0<n≤m. For example, the present invention can scan the first n symbols in the character string S, that is, scan s0, s1, ..., s n-1 , thereby counting the occurrence probability of each symbol among the n symbols, where n can be regarded as an adjustable parameter.

[0060] Wherein, when n=m, it is a complete scan, so the statistics on the symbol frequency of the character string S is complete. When n<m, this kind of scanning is partial scanning, so the statistics on the symbol frequency of the character string S is incomplete, thus the frequency statistics is incomplete.

[0061] The focus of the description of the present invention lies in the symbol frequency statistics under the partial scanning condition, so the preset scanning length is less than the length of the character string to be encoded, that is, n<m.

[0062] Step S130: determine a reference frequency for frequency adjustment according to the preset scanning length and the limited code length, so as to optimize the symbol frequency table based on the reference frequency.

[0063] For example, as mentioned above, when n=m, it is a complete scanning condition. The present invention does not process the counted frequency at all in this case, and this case is not the focus of the present invention. Step S130 is mainly directed to the condition of n<m, that is, the incomplete scanning condition. In this condition, there may be some symbols in the character string to be encoded that do not appear in the partial scanning statistics but appear in the subsequent unscanned character string, therefore, the occurrence frequency of some symbols may be too low in the counted symbol occurrence frequency. In view of this, embodiments of the present invention need to adjust the too low frequency. For frequency adjustment, a reference frequency needs to be determined first, and this reference frequency can be regarded as a reference value for frequency adjustment. In embodiments of the present invention, the reference frequency is determined based on the preset scanning length and the limited code length, so as to optimize and adjust the too low symbol frequency in the preliminarily counted symbol frequency table. On one hand, the symbol frequency table generated based on partial scanning can more accurately reflect the frequency distribution of symbols; on the other hand, the code length generated based on the optimized symbol frequency can meet the requirement of the limited code length, which is beneficial to subsequent encoding operations.

[0064] Step S140: generate a Huffman codebook for the character string to be encoded according to the optimized symbol frequency table.

[0065] For example, after optimizing the symbol frequency table, it is necessary to use the optimized symbol frequency table to construct a Huffman tree according to the Huffman coding algorithm, and then obtain the Huffman code length of each symbol from the Huffman tree, thus forming the Huffman codebook.

[0066] Step S150: Encode the string to be encoded using the Huffman codebook and output the corresponding compressed data.

[0067] For example, once the Huffman codebook is formed, it means that all symbols in the string to be encoded have obtained their corresponding binary codewords. Since Huffman coding is a prefix code, it can directly concatenate the codewords corresponding to symbols without needing delimiters.

[0068] In this embodiment of the invention, a partial scanning method is used to scan the string to be encoded, and the symbol frequency of the partial scan is optimized to generate a codebook that meets the code length limit for Huffman coding. Thus, this embodiment of the invention achieves Huffman coding through partial scanning, which not only improves the encoding speed but also generates codebooks that meet the code length limit. The entire encoding process only requires building the Huffman tree once, making the overall encoding scheme simple, efficient, and easy to implement in hardware.

[0069] The invention will now be described in further detail and with reference to the accompanying drawings.

[0070] In a preferred embodiment, the frequency of occurrence of each symbol among the scanned symbols is characterized by the following formula:

[0071] Among them, I j,i It is an indicator variable, and this indicator variable is characterized by the following formula:

[0072] Among them, f i Let i be the frequency of occurrence of the symbol i, and n be the preset scan length.

[0073] For example, embodiments of the present invention use an indicator variable to characterize symbol frequency, which is a special type of random variable used to indicate whether an event has occurred. When an event occurs, the indicator variable has a value of 1, and when the event does not occur, its value is 0. Specifically, let the alphabet be... It includes all possible symbols. Let the preset scan length be n, and the symbol sequence with scan length n be s0, s1, ..., s... n-1 Indicator variable I j,i Then s is used to characterize the position j in the sequence. j Is it the symbol i? If so, then Ij,i is 1; otherwise, I j,i is 0. In this way, by traversing the entire sequence and checking the symbol at each position, the total number of occurrences of each symbol in the sequence can be counted.

[0074] The indicator variable binary representation method adopted in the embodiments of the present invention is relatively convenient and easy to process in terms of frequency statistics. The number of occurrences of a symbol can be obtained only by accumulating the indicator variable corresponding to the symbol, and no complex logic or data structure is required, which effectively improves the overall efficiency of Huffman coding.

[0075] In a preferred embodiment, for the foregoing step S130, optimizing the symbol frequency table based on the reference frequency comprises:

[0076] Traversing the occurrence frequency of each counted symbol, comparing the occurrence frequency of the symbol with the reference frequency, if the occurrence frequency of the symbol is less than the reference frequency, adjusting the occurrence frequency of the symbol to the reference frequency, otherwise keeping the occurrence frequency of the symbol unchanged.

[0077] For example, let the determined reference frequency be p, and the occurrence frequency of the symbol be f i , and the adjusted frequency of the symbol is f i , if f i < p, then f i = p, otherwise f i = f i . It can be seen that after such frequency adjustment, the occurrence frequency of each symbol is adjusted to be not less than the determined reference frequency, which ensures that symbols with low original frequencies can be assigned a relatively reasonable code length in subsequent coding operations, thereby avoiding the problem of excessively long code lengths for some symbols.

[0078] In a preferred embodiment, the reference frequency is calculated by the following formula:

[0079] wherein p is the reference frequency, F k+3 is the (k+3)-th term of the Fibonacci sequence, k is the limited code length, n is the preset scanning length, is the alphabet to which the string to be encoded belongs.

[0080] For example, in the above formula represents rounding up, F k+3 is a Fibonacci number, wherein the Fibonacci number can be defined by the following formula:

[0081] F0=0, F1=1, F j =Fj-1 +F j-2 , j>1

[0082] In an embodiment of the present invention, the present invention uses the (k+3)-th term F of the Fibonacci sequence k+3 and introduces the restricted code length k and a preset scanning length n to determine a reference frequency, which is closely related to the structure of the subsequent Huffman tree and the distribution of nodes. In addition, the Fibonacci sequence can flexibly adopt different calculation methods for reference frequencies according to different restricted code lengths k. Specifically, when , the reference frequency is 1, and when , the reference frequency is

[0083] An embodiment of the present invention uses the Fibonacci sequence to calculate the reference frequency. On one hand, it can ensure that the subsequent Huffman coding satisfies the constraint of the restricted code length, and on the other hand, it ensures that all symbols have corresponding codewords, and the calculation is simple.

[0084] In a preferred embodiment, for the above step S140, generating a Huffman codebook for the string to be encoded according to the optimized symbol frequency table includes:

[0085] Step S210: sorting the symbol frequencies in the adjusted symbol frequency table in ascending order, and recording the position information of each symbol frequency after sorting;

[0086] Step S220: according to the sorted symbol frequencies and the position information, cyclically merging the frequencies of the symbols, determining the Huffman code length of each symbol of the string to be encoded, so as to generate a code length table of the string to be encoded;

[0087] Step S230: based on the code length table, generating a Huffman codebook for encoding the string to be encoded.

[0088] For example, for codebook construction, it is first necessary to sort the adjusted symbol frequencies in ascending order to obtain an ascending sequence, which can be recorded as where it is satisfied that when l<r, Since the frequencies of characters are sorted in ascending order, the frequency position corresponding to symbol i may no longer be i. Therefore, during the sorting process, it is necessary to record and save the position of symbol i after sorting. An embodiment of the present invention uses permutation π to record the position information of sorting. Specifically, π(j) = l if and only if is at position l after sorting. Obviously, the inverse mapping satisfies π -1 (l) = i.

[0089] Where the positive mapping π(j) = l represents that the symbol i is located at position l in the sorted sequence, and the inverse mapping π -1 (l) = i indicates that the original symbol corresponding to the l-th position after sorting is i. This ensures that even if the symbol position changes due to sorting, the original symbol and sorting position can still be quickly located through the forward mapping π and the inverse mapping π-1. Especially when the symbol frequency changes dynamically, the substitution π can efficiently adjust the codebook structure and optimize the codebook generation efficiency.

[0090] Once the symbol frequencies are sorted, a code length table can be generated based on the sorted frequency table. After obtaining the code length table, the corresponding codebook can be generated based on the structure of the Huffman tree.

[0091] More preferably, for step S220, based on the sorted symbol frequencies and the position information, a code length table including the Huffman code length corresponding to each symbol is generated, including:

[0092] Step S310: Obtain and merge the two symbols with the lowest frequencies in the symbol frequency table; adjust the symbol frequency sequence based on the merged frequencies, while maintaining the ascending order of the symbol frequency sequence;

[0093] Step S320: The code length of the merged symbols is adjusted incrementally and the position of the merged symbols is updated.

[0094] Step S330: Repeat the above process until the last symbol is merged to obtain the Huffman code lengths of all symbols; use the permutation function to adjust the code length order of the symbols so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

[0095] For example, in this embodiment of the invention, the two symbols with the lowest frequencies in the symbol frequency table are first obtained; then, these two symbols are merged and a new node is created accordingly, the frequency of which is the sum of the frequencies of the two symbols. The new node is then inserted into the symbol frequency table, while maintaining the ascending order of the symbol frequency table. In the process of merging symbols in the frequency table, for each merged symbol, the code length of the symbol is increased and the position of the merged symbol is updated. Finally, a permutation function is used to adjust the order of the code lengths of the symbols.

[0096] The following is a specific example to illustrate the above generation process:

[0097] 1. Initialization operation

[0098] Assume the input is an ascending array of symbol frequencies [1, 2, 3, 4]. Used to record the temporary code length of each character, initially 0, b l Frequency wA copy, used for dynamic adjustments during the merge, A l Used to record the position label of a symbol (initially the index of the symbol, such as 0, 1, 2, 3).

[0099] 2. Cyclic merging frequency

[0100] After initialization, frequencies are merged cyclically, with each iteration merging the two smallest frequencies to gradually generate the code length. The specific process is as follows:

[0101] For example, given the input frequency array mentioned above, the first loop merges the two smallest frequencies b0 = 1 and b... 1=2 The frequency t after merging is obtained. =3 .

[0102] Once the new frequency t after merging is determined, it is necessary to determine the position to be inserted into the current frequency list, such as between 3 and 4. This position is denoted as κ.

[0103] After determining the insertion position, the merged frequency t is inserted into κ, and the remaining frequencies are arranged in order, such as the merged frequency list becoming [3, 3, 4].

[0104] Add 1 to the code length of the two merged symbols, i.e., the symbols labeled 0 and 1 (because these two symbols were merged once).

[0105] Adjust the position labels of the symbols by subtraction, so that the labels reflect the new positions of the symbols after merging (for example, a symbol originally labeled 2 may become labeled 0).

[0106] 3. Repeat the merging process until completion.

[0107] The two lowest frequencies are merged in each cycle until all symbols are merged into a single unit. The code length of a symbol is determined by the number of merges; the more times a symbol is merged, the longer its code length becomes (for example, a symbol merged 3 times has a code length of 3).

[0108] For the frequency list [3, 3, 4] from the first cycle merging, a second cycle merging is performed: the two lowest frequencies 3 and 3 are merged to obtain the merged frequency t = 6. The code length of the symbols labeled 0 and 1 (corresponding to the two merged 3s) is increased by 1 to become 2, and the new frequency list is [6, 4].

[0109] Furthermore, a third loop merging is performed on the new frequency list [6, 4]: frequencies 6 and 4 are merged to obtain the merged frequency t = 10, and the symbol code length of tags 0 and 1 is increased by 1 to become 3.

[0110] 4. Position adjustment

[0111] The code length order is adjusted according to the permutation function π to ensure that the output code length table is consistent with the original symbol order, and the code length of the original symbol is [3, 3, 2, 1].

[0112] In other words, the generation process of the code length table described above is actually the process of constructing a Huffman tree. For each symbol frequency, a tree (with only a root node) is generated, and these trees form a forest. When there are more than one tree in the forest, the two trees with the lowest frequencies are taken from the forest and merged into a new tree. Specifically, a root node is constructed, and the two taken trees are designated as the left and right subtrees, respectively. The frequency of the merged tree is the sum of the frequencies of the left and right subtrees. This new tree is then reinserted into the forest, and the merging process is repeated until there is only one tree left in the forest. At this point, each symbol corresponds to a leaf node in this tree. The path from the root node to this leaf node is the codeword of that symbol. It is clear that this process focuses on the depth of the leaf nodes, not their specific positions. The depth of each leaf node and its frequency in the tree are recorded in ascending order. In each update, the depths of nodes with sorting positions of 0 and 1 are updated. Since only the merged trees experience frequency changes, as long as the merged tree is inserted into the list of remaining trees according to its new frequency, the frequency sorting will remain in ascending order. In this way, we only need to update the sorting position of the tree to which the leaf node belongs according to the insertion position.

[0113] The following section further explains the specific process of generating the code length table from the perspective of constructing a Huffman tree:

[0114] Suppose there are 4 symbols 0, 1, 2, 3, with frequencies of 3, 9, 4, and 11 respectively. Initially, each node is the root node, with depths (code lengths) of 0, 0, 0, 0, and ascending frequencies of 3, 4, 9, 11. The ascending positions of the nodes are 0, 2, 1, 3.

[0115] After initialization, frequencies are merged cyclically, with each iteration merging the two smallest frequencies to gradually generate the code length. The specific process is as follows:

[0116] For example, given the input frequency array mentioned above, in the first loop, the two smallest frequencies, 3 and 4, are merged, resulting in a new tree with a frequency of 3 + 4 = 7. The updated depth and frequencies are 1, 0, 1, 0 and 7, 9, 11, respectively, and the ascending positions of the nodes are 0, 1, 0, 2.

[0117] Following this pattern, the second merging cycle combines frequencies 7 and 9, resulting in a new tree with a frequency of 7 + 9 = 16. The updated depths and frequencies are 2, 1, 2, 0 and 11, 16, respectively, with the node's ascending position being 1, 1, 1, 0. The third merging cycle combines frequencies 11 and 16, resulting in an updated depth and frequency of 3, 2, 3, 1 and 27, with the node's ascending position being 0, 0, 0, 0. The depth table 3, 2, 3, 1 represents the output code length, and Figure 2 shows the corresponding tree.

[0118] In this embodiment of the invention, the code length is generated based on the ascending frequency of the symbols and the permutation information (position information). The code length allocation can be dynamically adjusted according to the input symbol frequency. Only one Huffman tree needs to be constructed, while ensuring consistent coding efficiency and meeting different coding requirements.

[0119] More preferably, for step 230 above, generating a codebook for encoding the string to be encoded based on the code length table includes:

[0120] Step S410: Traverse the code length table and count the occurrences of each code length;

[0121] Step S420: Determine the starting prefix of the code length based on the number of occurrences of the code length;

[0122] Step S430: For each symbol of the string to be encoded, convert the starting prefix of its corresponding code length into a binary representation, and truncate the low bits based on the corresponding code length as the codeword of the symbol;

[0123] Step S440: Incrementally update the starting prefix of the corresponding code length to assign a different codeword to the next symbol with that code length.

[0124] For example, setting array z j and y j And the variable t, where the array z j The array y is used to record the number of occurrences of each code length j. j The variable t is used to store the starting codeword for each code length j, and the variable t is used to calculate the codeword prefix.

[0125] First, iterate through the code length d of all symbols. i The number of occurrences for each code length is counted and stored in array z. j Secondly, starting from code length 1, it iterates up to the maximum code length k. The variable t is used to accumulate the number of symbols of the previous code length and shifts it left by 1 bit to make room for the codeword of the new code length. Thus, the starting codeword of each code length j is calculated and stored in array y. j In the middle. Again, based on the symbol's code length d. i , extract y jbinary low d i The bit is used as the codeword c of the current symbol. i This assigns a codeword to the current symbol. Simultaneously, it adds a prefix of that codeword length. Incremental, for example, the increment depth can be set to 1, i.e. This ensures that codewords of the same length are consecutive and unique.

[0126] In short, the embodiments of the present invention are based on the given depth of each symbol, and the codeword of each symbol can be obtained by traversing the constructed Huffman depth. That is, the starting codeword (i.e. the starting prefix) of each layer is calculated first, and then the corresponding codeword is assigned to the symbol in that layer.

[0127] For example, given the depth table (code length table) 3, 2, 3, 1, the starting codewords for layers 1, 2, and 3 are 0, 10, and 110, respectively, and the corresponding codewords for symbols 0, 1, 2, and 3 are 110, 10, 111, and 0, respectively.

[0128] In this embodiment of the invention, the generated codewords are prefix codes, meaning that no single codeword is a prefix of any other codeword. This ensures that the encoded data stream can be accurately restored to the original data without additional delimiters during decoding, avoiding ambiguity and improving decoding accuracy and efficiency. Furthermore, this embodiment of the invention can dynamically generate codebooks based on different depth tables (i.e., different symbol frequency distributions).

[0129] In a preferred embodiment, the encoding method further includes: before generating the codebook, obtaining the maximum code length in the code length table; comparing the maximum code length with the restricted code length; if the maximum code length is greater than the restricted code length, adjusting the code length table so that all code lengths in the code length table are less than or equal to the restricted code length.

[0130] For example, as previously described, one of the objectives of this invention is to ensure that the encoded code length meets the requirements of the code length restriction. In fact, in most cases, the code length table generated based on the method mentioned in the above embodiments can meet the code length restriction requirements. However, in some special cases, such as when... However, the generated code length table may still be larger than the code length limit. Therefore, this invention requires further adjustment of the code length table. Thus, in this embodiment, after generating the code length table, the code lengths of all symbols are traversed to determine if any code length exceeds the code length limit. If so, the generated code length table needs to be adjusted to ensure that all code lengths meet the code length limit requirement.

[0131] More preferably, the adjusted code length table includes:

[0132] Step S510: Obtain two symbols in the code length table that correspond to a first preset code length, wherein the first preset code length is associated with the maximum code length in the code length table;

[0133] Step S520: Obtain the symbol corresponding to the second preset code length in the code length table, wherein the second preset code length is associated with the restricted code length;

[0134] Step S530: Based on the first preset code length, the second preset code length, and the preset code length increment and decrement depth, adjust the code length of the obtained symbol.

[0135] Step S540: Repeat the above process, cyclically adjusting the code length table until the maximum code length in the code length table is less than or equal to the limit code length.

[0136] For example, in this embodiment of the invention, the first step is to query the code length table to see if there are two symbols corresponding to the first preset code length, which is associated with the maximum code length. For instance, the first preset code length is initially the maximum code length. If no two symbols corresponding to the maximum code length are found in the code length table, the maximum code length is decremented and updated for a new query. The decrement depth of the maximum code length can be set to 1. For example, if no symbols are found, the query continues to check if there are two symbols corresponding to the maximum code length minus 1, and so on. When two symbols are found, they can be denoted as the first symbol and the second symbol.

[0137] Next, the code length table is checked to see if a symbol corresponding to the second preset code length exists. This second preset code length is associated with the limit code length. For example, the second preset code length is initially the limit code length minus 1. If no symbol corresponding to the limit code length minus 1 is found in the code length table, the second preset code length is decreased and updated. The decrease depth can be set to 1, for example, the second preset code length is updated to the limit code length minus 2, and the search continues, and so on. If a symbol is found, it is recorded as the third symbol.

[0138] Therefore, after obtaining the aforementioned symbols, the present invention needs to adjust the code length of these symbols. Specifically, the code lengths of the first and third symbols are adjusted incrementally based on the code length of the third symbol, with an increment depth of 1, meaning the code lengths of the first and third symbols are adjusted to the code length of the third symbol plus 1. Additionally, the code length of the second symbol is adjusted incrementally based on its own code length, with a decrease depth of 1, meaning the code length of the second symbol is adjusted to the code length of the second symbol minus 1.

[0139] Finally, after adjusting the code length of the above symbols, the maximum code length in the adjusted code length table is retrieved again and compared with the restricted code length. If it is determined that the adjusted maximum code length is greater than the restricted code length, the above adjustment process is repeated until the maximum code length in the code length table is less than or equal to the restricted code length.

[0140] The following is a more concrete example to illustrate the above adjustment process:

[0141] Let the maximum code length in the code length table be d. max The code length is limited to k.

[0142] 1. Retrieve the largest code length d from the current code length table. max If d max If d is less than or equal to the limit code length k, then no adjustment to the code length table is needed. max When the value is greater than k, enter the adjustment loop:

[0143] 2. Search the code length table for two codes with the same length d. max If the symbol is not found, then d will be used. max Decrease by 1, search again, and so on, denoting the two found symbols as γ and θ respectively.

[0144] 3. Search for a symbol with a code length of k-1 in the code length table. If it is not found, update the code length by decreasing it and search again. Repeat this process and denote the found symbol as β.

[0145] 4. Once symbols γ, θ, and β are found, the corresponding code lengths need to be adjusted:

[0146] Among them, d γ Adjusted to d β +1, d β Adjust to d β +1, and d θ Adjust to d θ -1.

[0147] After adjusting the code length, repeat the above process until dmax is less than or equal to k.

[0148] Clearly, by further combining this with the concept of Huffman trees, the above adjustment process is actually about finding leaf nodes in the tree that are at or below level k-1. Such nodes indicate that there is space at level k or lower, so nodes above level k can be moved to that level.

[0149] For example, if we assume the code length is limited to 2, meaning the maximum depth is limited to 2, then the tree shown in Figure 3 cannot meet the code length requirement. Combining the above adjustment process, this invention finds node 0 with a depth of 1 in Figure 3, merges node 2 and node 0, and moves node 3 up one level, resulting in the tree in Figure 4. The tree shown in Figure 4 can then satisfy the requirement of a maximum depth of 2, i.e., the initial depth table is 0, 1, 2, 3, while the adjusted depth table is 2, 2, 2, 2.

[0150] In this embodiment of the invention, by adjusting the code length, it can be ensured that all code lengths can meet the code length restriction requirements.

[0151] Furthermore, based on the Huffman coding method described in the above embodiments, this invention further introduces the practical application of the method. Specifically, the Huffman coding scheme of this invention can be applied to the Deflate standard (RFC 1951). For literal / length alphabets, The code length is limited to k=15. In application, the frequencies of n symbols are first counted, based on incomplete statistics from a partial scan. Then, the counted frequencies are post-processed for adjustment and optimization, and a corresponding code length table is generated, thus producing Huffman codewords. Finally, the generated codewords are used to encode the string to be encoded. It is worth noting here that, due to F... 18 -286+1>0, therefore, no adjustment to the code length table is required.

[0152] Accordingly, based on the same inventive concept, as shown in Figure 5, the present invention provides a Huffman encoding device 500, comprising: an acquisition unit 501, configured to acquire a string to be encoded and associated encoding information, the encoding information including the length of the string to be encoded, the alphabet to which it belongs, and the restricted code length to be encoded; a statistics unit 502, configured to scan the symbols of the string based on a preset scan length, count the occurrence frequency of each symbol in the scanned symbols, and generate a symbol frequency table corresponding to the string to be encoded based on the preset scan length; an optimization unit 503, configured to determine a reference frequency for frequency adjustment based on the preset scan length and the restricted code length, and optimize the symbol frequency table based on the reference frequency; a generation unit 504, configured to generate a Huffman codebook for the string to be encoded based on the optimized symbol frequency table; and an output unit 505, configured to encode the string to be encoded using the Huffman codebook and output the corresponding encoding result.

[0153] In a preferred embodiment, the frequency of occurrence of each symbol among the scanned symbols is characterized by the following formula:

[0154] Among them, I j,i It is an indicator variable, and this indicator variable is characterized by the following formula:

[0155] Among them, f i Let i be the frequency of occurrence of the symbol i, and n be the preset scan length.

[0156] In a preferred embodiment, the optimization unit is used to optimize the symbol frequency table based on the reference frequency, including: traversing the occurrence frequency of each statistically recorded symbol, comparing the occurrence frequency of the symbol with the reference frequency, and if the occurrence frequency of the symbol is less than the reference frequency, adjusting the occurrence frequency of the symbol to the reference frequency; otherwise, keeping the occurrence frequency of the symbol unchanged.

[0157] In a preferred embodiment, the reference frequency is calculated using the following formula:

[0158] Where p is the reference frequency, F k+3 Here, k is the Fibonacci number, k is the restricted code length, and n is the preset scan length. The alphabet to which the string to be encoded belongs.

[0159] In a preferred embodiment, the generation unit is used to generate a Huffman codebook for the string to be encoded based on the optimized symbol frequency table, including: sorting the symbol frequencies in the adjusted symbol frequency table in ascending order and recording the position information of each symbol frequency after sorting; and cyclically merging the frequencies of the symbols based on the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol in the string to be encoded, so as to generate a code length table for the string to be encoded.

[0160] Based on the code length table, a Huffman codebook is generated for encoding the string to be encoded.

[0161] In a preferred embodiment, the generation unit is used to cyclically merge the frequencies of the symbols according to the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol in the string to be encoded, so as to generate a code length table of the string to be encoded. This includes: obtaining and merging the two symbols with the lowest frequencies in the symbol frequency table; adjusting the symbol frequencies based on the merged frequencies while maintaining the ascending order of the symbol frequencies; incrementally adjusting and updating the merged symbol positions based on the code lengths of the merged symbols; repeating the above process until the last symbol is merged to obtain the Huffman code lengths of all symbols; and adjusting the code length order of all symbols using a permutation function so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

[0162] In a preferred embodiment, the generation unit is used to generate a codebook for encoding the string to be encoded based on the code length table, including: traversing the code length table and counting the occurrence frequency of each code length; determining the starting prefix of the code length based on the occurrence frequency of the code length; for each symbol of the string to be encoded, converting the starting prefix of its corresponding code length into a binary representation, and truncating the low-order bits based on the corresponding code length as the codeword of the symbol; and incrementally updating the starting prefix of the corresponding code length to assign different codewords to the next symbol with that code length.

[0163] In a preferred embodiment, the Huffman coding apparatus further includes an adjustment unit, which is configured to: obtain the maximum code length in the code length table before generating the codebook; compare the maximum code length with the restricted code length; and if the maximum code length is greater than the restricted code length, adjust the code length table so that all code lengths in the code length table are less than or equal to the restricted code length.

[0164] In a preferred embodiment, the adjustment unit is used to adjust the code length table, comprising: acquiring two symbols in the code length table corresponding to a first preset code length, the first preset code length being associated with the maximum code length in the code length table; acquiring a symbol in the code length table corresponding to a second preset code length, the second preset code length being associated with the restricted code length; adjusting the code length of the acquired symbols based on the first preset code length, the second preset code length, and preset code length increment and decrement depths; repeating the above process to cyclically adjust the code length table until the maximum code length in the code length table is less than or equal to the restricted code length.

[0165] For specific implementation details and further advantages of the Huffman encoding device in the embodiments of the present invention, please refer to the above embodiments of the Huffman encoding method, which will not be elaborated further here.

[0166] Accordingly, the present invention provides a chip comprising: a processor and a memory; wherein the memory stores a program capable of running on the processor, and the processor is configured to implement the Huffman coding method described in the above embodiments when executing the program.

[0167] Accordingly, the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the Huffman coding method described in the above embodiments.

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0169] The units described in the embodiments of the present invention can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0170] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0171] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0172] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0173] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0174] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A Huffman coding method, characterized in that, The Huffman coding method includes: Obtain the string to be encoded and its associated encoding information, wherein the encoding information includes the length of the string to be encoded, the alphabet to which it belongs, and the code length limit to be encoded; The symbols of the string to be encoded are scanned based on a preset scan length, the frequency of each symbol in the scanned symbols is counted, and a symbol frequency table corresponding to the string to be encoded is generated based on the preset scan length. Based on the preset scan length and the limited code length, a reference frequency for frequency adjustment is determined, so as to optimize the symbol frequency table based on the reference frequency; Generate a Huffman codebook for the string to be encoded based on the optimized symbol frequency table; The string to be encoded is encoded using the Huffman codebook, and the corresponding encoding result is output.

2. The Huffman coding method according to claim 1, characterized in that, The optimization of the symbol frequency table based on the reference frequency includes: The frequency of each symbol is traversed and compared with the reference frequency. If the frequency of the symbol is less than the reference frequency, the frequency of the symbol is adjusted to the reference frequency; otherwise, the frequency of the symbol remains unchanged.

3. The Huffman coding method according to claim 1, characterized in that, The reference frequency is calculated using the following formula: Where p is the reference frequency, F k+3 Here, k is the Fibonacci number, k is the restricted code length, and n is the preset scan length. The alphabet to which the string to be encoded belongs.

4. The Huffman coding method according to claim 1, characterized in that, The step of generating the Huffman codebook for the string to be encoded based on the optimized symbol frequency table includes: Sort the symbol frequencies in the optimized symbol frequency table in ascending order and record the position information of each symbol frequency after sorting. Based on the sorted symbol frequencies and the position information, the frequencies of the symbols are cyclically merged to determine the Huffman code length of each symbol in the string to be encoded, so as to generate the code length table of the string to be encoded. Based on the code length table, a Huffman codebook is generated for encoding the string to be encoded.

5. The Huffman coding method according to claim 4, characterized in that, The step of cyclically merging the frequencies of the symbols based on the sorted symbol frequencies and the position information to determine the Huffman code length of each symbol in the string to be encoded, in order to generate a code length table for the string to be encoded, includes: Obtain and merge the two symbols with the lowest frequencies in the symbol frequency table; Adjust the symbol frequency sequence based on the merged frequency, while maintaining the ascending order of the symbol frequency sequence; The code length of the merged symbols is incremented and the position of the merged symbols is updated. Repeat the above process until merging to the last symbol to obtain the Huffman code lengths of all symbols; The code length order of the symbols is adjusted using a permutation function so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

6. The Huffman coding method according to claim 1, characterized in that, The step of generating a codebook for encoding the string to be encoded based on the code length table includes: Iterate through the code length table and count the occurrences of each code length; The starting prefix of the code length is determined based on the number of times the code length appears; For each symbol in the string to be encoded, the starting prefix of its corresponding code length is converted into a binary representation, and the low-order bits are truncated based on the corresponding code length as the codeword of the symbol; The starting prefix of the corresponding code length is incremented and updated so that different codewords are assigned to the next symbol with that code length based on the updated starting prefix.

7. The Huffman coding method according to claim 1, characterized in that, The Huffman coding method also includes: Before generating the codebook, obtain the maximum code length from the code length table; Compare the maximum code length and the limited code length. If the maximum code length is greater than the limited code length, adjust the code length table so that all code lengths in the code length table are less than or equal to the limited code length.

8. The Huffman coding method according to claim 7, characterized in that, The adjustment of the code length table includes: Obtain two symbols in the code length table that correspond to a first preset code length, wherein the first preset code length is associated with the maximum code length in the code length table; Obtain the symbol corresponding to the second preset code length from the code length table, where the second preset code length is associated with the restricted code length; Based on the first preset code length, the second preset code length, and the preset code length increment and decrement depths, the code length of the obtained symbols is adjusted. Repeat the above process to cyclically adjust the code length table until the maximum code length in the code length table is less than or equal to the limit code length.

9. The Huffman coding method according to claim 1, characterized in that, The frequency of occurrence of each symbol in the scanned symbols is represented by the following formula: Among them, I j,i It is an indicator variable, and this indicator variable is characterized by the following formula: Among them, f i Let be the frequency of occurrence of symbol i, and n be the preset scan length, s j Let j be the sequence of symbols being scanned.

10. A Huffman coding device, characterized in that, The Huffman encoding device includes: The acquisition unit is used to acquire the string to be encoded and the associated encoding information, wherein the encoding information includes the length of the string to be encoded, the alphabet to which it belongs, and the limited code length to be encoded; The statistics unit is used to scan the symbols of the string based on a preset scan length, count the occurrence frequency of each symbol in the scanned symbols, and generate a symbol frequency table corresponding to the string to be encoded based on the preset scan length. An optimization unit is used to determine a reference frequency for frequency adjustment based on the preset scan length and the limited code length, so as to optimize the symbol frequency table based on the reference frequency. The generation unit is used to generate a Huffman codebook for the string to be encoded based on the optimized symbol frequency table; The output unit is used to encode the string to be encoded using the Huffman codebook and output the corresponding encoding result.

11. The Huffman coding device according to claim 10, characterized in that, The optimization unit is used to optimize the symbol frequency table based on the reference frequency, including: The frequency of each symbol is traversed and compared with the reference frequency. If the frequency of the symbol is less than the reference frequency, the frequency of the symbol is adjusted to the reference frequency; otherwise, the frequency of the symbol remains unchanged.

12. The Huffman coding device according to claim 10, characterized in that, The reference frequency is calculated using the following formula: Where p is the reference frequency, F k+3 Here, k is the Fibonacci number, k is the restricted code length, and n is the preset scan length. The alphabet to which the string to be encoded belongs.

13. The Huffman coding device according to claim 10, characterized in that, The generation unit is used to generate a Huffman codebook for the string to be encoded based on the optimized symbol frequency table, including: Sort the symbol frequencies in the adjusted symbol frequency table in ascending order and record the position information of each symbol frequency after sorting. The frequency of the symbols is cyclically merged according to the sorted symbol frequency and the position information to determine the Huffman code length of each symbol in the string to be encoded, so as to generate the code length table of the string to be encoded. Based on the code length table, a Huffman codebook is generated for encoding the string to be encoded.

14. The Huffman coding device according to claim 13, characterized in that, The generation unit is used to cyclically merge the frequencies of the symbols according to the sorted symbol frequencies and the position information, and determine the Huffman code length of each symbol in the string to be encoded, so as to generate a code length table for the string to be encoded, including: Obtain and merge the two symbols with the lowest frequencies in the symbol frequency table; The symbol frequency is adjusted based on the frequency after the two symbols are merged, while maintaining the ascending order of the symbol frequencies; The code length of the merged symbols is incremented and the position of the merged symbols is updated. Repeat the above process until merging to the last symbol to obtain the Huffman code lengths of all symbols; The code length order of all symbols is adjusted using a permutation function so that the code length order of the generated code length table corresponds to the original arrangement order of the symbols.

15. The Huffman coding apparatus according to claim 10, characterized in that, The generation unit is used to generate a codebook for encoding the string to be encoded based on the code length table, including: Iterate through the code length table and count the occurrences of each code length; The starting prefix of the code length is determined based on the number of times the code length appears; For each symbol in the string to be encoded, the starting prefix of its corresponding code length is converted into a binary representation, and the low-order bits are truncated based on the corresponding code length as the codeword of the symbol; The starting prefix of the corresponding code length is incremented and updated so that different codewords are assigned to the next symbol with that code length based on the updated starting prefix.

16. The Huffman coding apparatus according to claim 10, characterized in that, The Huffman encoding device further includes an adjustment unit, which is used for: Before generating the codebook, obtain the maximum code length from the code length table; Compare the maximum code length and the limited code length. If the maximum code length is greater than the limited code length, adjust the code length table so that all code lengths in the code length table are less than or equal to the limited code length.

17. The Huffman coding apparatus according to claim 16, characterized in that, The adjustment unit is used to adjust the code length table, including: Obtain two symbols in the code length table that correspond to a first preset code length, wherein the first preset code length is associated with the maximum code length in the code length table; Obtain the symbol corresponding to the second preset code length from the code length table, where the second preset code length is associated with the restricted code length; Based on the first preset code length, the second preset code length, and the preset code length increment and decrement depths, the code length of the obtained symbols is adjusted. Repeat the above process to cyclically adjust the code length table until the maximum code length in the code length table is less than or equal to the limit code length.

18. The Huffman coding apparatus according to claim 10, characterized in that, The frequency of occurrence of each symbol in the scanned symbols is represented by the following formula: Among them, I j,i It is an indicator variable, and this indicator variable is characterized by the following formula: Among them, f i Let be the frequency of occurrence of symbol i, and n be the preset scan length, s j Let j be the sequence of symbols being scanned.

19. A chip, characterized in that, include: A processor and a memory; wherein the memory stores a program capable of running on the processor, the processor being configured to implement the Huffman coding method according to any one of claims 1-9 when executing the program.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the Huffman coding method according to any one of claims 1-9.