Semantic-based huffman encoding method, semantic-based huffman decoding method and related device
Through the semantic-based Huffman encoding method, synonymous mapping and semantic Huffman codebooks are used to solve the problem of limited encoding compression efficiency in communication technology, and more efficient encoding compression and transmission efficiency are achieved.
Patent Information
- Application Number
- PCT/CN2024/091201
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-05-06
- Publication Date
- 2025-07-31
AI Technical Summary
The existing communication technology is difficult to further improve the source encoding and compression efficiency in ultra-high-speed transmission, due to the grammatical information compression limit of classical information theory.
The semantic-based Huffman encoding method is adopted to pre-construct the synonymous map codebook and the semantic Huffman codebook, determine the encoding fields of the synonymous set, and encode them according to the order of the information sequence, and decode the synonymous set at the receiving end.
Effectively compress the average code length of the coding sequence, improve the compression efficiency during the coding process, improve the transmission efficiency of communication technology, and ensure that the transmission process is not distorted.
Smart Images

Figure CN2024091201_31072025_PF_FP_ABST
Abstract
Description
Semantic-based Huffman coding method, decoding method and related equipment Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a semantic-based Huffman encoding method, decoding method, and related equipment. Background Art
[0002] As future communications scenarios increasingly demand ultra-high-speed transmission, efficiently compressing and reliably transmitting large amounts of data has become a key research focus in current communications technology. However, because classical information theory describes information based on symbol probabilities, or grammatical information, the theoretical limit of source coding compression corresponds to the compression limit of grammatical information. At this compression limit, further improvements in source coding compression efficiency are impossible, limiting the development of communications technology.
[0003] Summary of the Invention
[0004] A first aspect of the present application provides a semantic-based Huffman coding method applied to a transmitting end, comprising: in response to receiving an information sequence containing at least one codeword sent by a source, performing synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook, and determining a synonymous set corresponding to the codeword; determining a coding field corresponding to the synonymous set according to the pre-constructed semantic Huffman codebook; sorting all coding fields in sequence according to the order of the codewords in the information sequence to obtain a coding sequence corresponding to the information sequence; and sending the coding sequence to a receiving end, so that the receiving end decodes the coding sequence.
[0005] A second aspect of the present application provides a semantic-based Huffman decoding method applied to a receiving end, comprising: in response to receiving a coding sequence sent by a transmitting end, determining a synonymous set sequence corresponding to the coding sequence based on a pre-constructed Huffman code book; extracting a source symbol as a target source symbol from each synonymous set in the synonymous set sequence according to a preset extraction method; and sorting the target source symbols in sequence according to the order of the synonymous sets in the synonymous set sequence to obtain a decoding sequence corresponding to the coding sequence.
[0006] The third aspect of the present application further provides a semantic-based Huffman coding device, comprising:
[0007] a first determining module configured to, in response to receiving an information sequence containing at least one codeword sent by a source, perform synonymous mapping on the codeword based on a pre-built synonymous mapping codebook, and determine a synonymous set corresponding to the codeword;
[0008] A second determining module is configured to determine the coding field corresponding to the synonymous set according to a pre-built semantic Huffman codebook;
[0009] an encoding module configured to sequentially sort all encoding fields according to the order of the codewords in the information sequence to obtain an encoding sequence corresponding to the information sequence; and
[0010] The sending module is configured to send the coding sequence to a receiving end, so that the receiving end decodes the coding sequence.
[0011] A fourth aspect of the present application further provides a semantic-based Huffman decoding device, comprising:
[0012] A third determining module is configured to, in response to receiving a coding sequence sent by a transmitting end, determine a synonym set sequence corresponding to the coding sequence based on a pre-constructed Huffman codebook;
[0013] an extraction module configured to extract a source symbol as a target source symbol from each synonymous set in the synonymous set sequence according to a preset extraction method; and
[0014] The decoding module is configured to sequentially sort the target source symbols according to the order of the synonymous sets in the synonymous set sequence to obtain a decoding sequence corresponding to the coding sequence.
[0015] The fifth aspect of the present application also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, wherein the processor implements the method described in the first aspect or the second aspect when executing the computer program.
[0016] The sixth aspect of the present application further provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the method as described in the first aspect or the second aspect.
[0017] The seventh aspect of the present application further provides a computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method as described in the first aspect or the second aspect.
[0018] As can be seen from the above, the semantic-based Huffman coding method, decoding method and related equipment provided by the present application, in response to receiving an information sequence containing at least one codeword sent by a source, performs synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook to determine the synonymous set corresponding to the codeword. Each codeword is associated with a synonymous set according to semantics, and the semantics of each codeword contained in the synonymous set are the same. The coding field corresponding to the synonymous set is determined according to the pre-constructed semantic Huffman codebook, and each synonymous set corresponds to a coding field, that is, the coding field corresponding to each codeword in the synonymous set during the coding process is the same. According to the order of the codewords in the information sequence, all coding fields are sorted in sequence to obtain the coding sequence corresponding to the information sequence, and the coding sequence obtained by encoding the synonymous set is used to replace the traditional Huffman coding method, which can effectively compress the average code length of the coding sequence, improve the compression efficiency in the coding process, and then improve the transmission efficiency of the coding sequence, which is conducive to further improving the efficiency of communication technology. The coded sequence is sent to a receiving end so that the receiving end decodes the coded sequence. The obtained decoded sequence retains the semantics of the information sequence, ensuring that the transmission process is not distorted. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] FIG1 is a flow chart of the semantic-based Huffman coding method according to an embodiment of the present application.
[0021] FIG2 is a schematic diagram showing the correspondence between the semantic information space and the grammatical information space according to an embodiment of the present application.
[0022] FIG3 is a schematic diagram of a process for constructing a semantic Huffman codebook according to an embodiment of the present application.
[0023] FIG4 is a schematic diagram of the process of constructing a Huffman tree according to an embodiment of the present application.
[0024] FIG5 is a schematic diagram of the structure of the Huffman tree according to an embodiment of the present application.
[0025] FIG6 is a flowchart of a semantic-based Huffman decoding method according to another embodiment of the present application.
[0026] FIG7 is a schematic diagram of the structure of a semantic-based Huffman coding device according to an embodiment of the present application.
[0027] FIG8 is a schematic structural diagram of a semantic-based Huffman encoding and decoding device according to another embodiment of the present application.
[0028] FIG9 is a schematic diagram of the hardware structure of the electronic device of the present application. DETAILED DESCRIPTION
[0029] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0030] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0031] As mentioned in the background, with the increasing demand for ultra-high-speed transmission in future communications scenarios, efficient compression and reliable transmission of large amounts of data have become key research areas in current communications technology. In 1948, Shannon proposed classical information theory. In this theory, he introduced the concept of entropy and proposed a theoretically achievable limit for information compression, specifically, reducing the average number of bits required for information transmission through coding. Since then, numerous coding schemes based on classical information theory have been designed and implemented.
[0032] Source coding is a key technology for efficient data compression. Lossless source coding aims to reduce data storage or transmission requirements without losing the original data. This coding method is commonly used in fields such as digital audio, images, and text. Entropy coding is often used for lossless source coding of discrete sources. Entropy coding is a technique that uses the probability distribution of symbol occurrence to encode data, aiming to obtain shorter codes for symbols with high probability of occurrence, thereby achieving efficient data compression. Huffman coding is a special case of entropy coding. This algorithm primarily encodes characters or symbols based on their frequency of occurrence in the data to be compressed, achieving efficient data compression. Huffman coding achieves efficient data encoding by assigning shorter codes to frequently occurring symbols and longer codes to less frequently occurring symbols. The theoretical compression limit of traditional Huffman coding is the Shannon entropy. For messages of finite length, Huffman coding can often achieve the Shannon limit in practice.
[0033] With the development of semantic communication technology, a growing number of studies have shown that, in practical communication scenarios, semantic communication methods that incorporate semantic domain information processing can achieve more efficient source compression schemes than existing source coding methods. The newly proposed semantic information theory describes the relationship between semantic information and syntactic information as a one-to-many functional mapping. The semantic source coding theorem developed based on this theory proves that semantic source compression coding can achieve more efficient source compression coding schemes than traditional source coding schemes. This demonstrates that the introduction of semantic information theory into source compression can further compress the source without compromising reconstruction quality. Therefore, semantic Huffman coding, based on semantic information theory, has higher compression efficiency than traditional Huffman coding and has the potential to become an important solution to the problem of high-efficiency source compression.
[0034] In light of this, this application proposes a semantically based Huffman encoding and decoding method. Guided by semantic information theory, the transmitting and receiving ends acquire identical synonym sets, establishing a connection between semantics and syntax through these synonym sets. A semantic Huffman codebook is constructed based on the prior probabilities of different synonym sets. At the transmitting end, a source symbol sequence is encoded using the semantic Huffman codebook. At the receiving end, the encoded symbol sequence is decoded using the semantic Huffman codebook, ultimately yielding a decoded source sequence.
[0035] It should be noted that binary coding is generally used when performing source coding, so this application mainly describes semantic Huffman coding based on binary coding. If multi-base coding is involved, simple changes can be made based on this application.
[0036] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0037] The embodiment of the present application proposes a semantic-based Huffman coding method, which is applied to a transmitting end. Referring to FIG1 , the method includes the following steps:
[0038] Step 102: In response to receiving an information sequence including at least one codeword sent by a source, perform synonymous mapping on the codeword based on a pre-built synonymous mapping codebook to determine a synonymous set corresponding to the codeword.
[0039] Specifically, a source is an entity that generates various types of information. The symbols provided by a source are uncertain and can be described by random variables and their statistical properties. Information is abstract, while a source is concrete. For example, when people converse, the human vocal system is the speech source; when people read books or newspapers, the illuminated books and newspapers themselves are text sources. Other common sources include image sources and digital sources. When the transmitter receives an information sequence from the source, it needs to perform synonym mapping on each codeword in the sequence, mapping codewords with the same meaning into a synonym set. First, the semantic information of each codeword must be determined. A codeword is one of all the source symbols corresponding to the source, as well as the semantic information of each synonym set. The semantic information of the codeword is compared with the semantic information of the synonym set. If they are identical, the codeword is mapped to the synonym set. A pre-built synonym codebook specifies a one-to-one correspondence between synonym sets and semantic information. The semantic information uniquely identifies a synonym set.
[0040] In a specific example, the information sequence sent by the source is denoted as u=[u1,…,u i …,u M ],u i Represents the i-th codeword, and the information sequence length is M. Traverse each codeword u in the information sequence i , determine the corresponding synonym set according to the synonym mapping After completing the mapping of the codewords, the synonymous set sequence is obtained The sequence length is M.
[0041] Step 104: Determine the coding field corresponding to the synonymous set according to the pre-built semantic Huffman codebook.
[0042] In this embodiment, the semantic Huffman codebook is pre-built, and the semantic Huffman codebook specifies a one-to-one correspondence between synonym sets and coding fields. After the synonym set is determined, the coding field that uniquely corresponds to the synonym set can be queried through the semantic Huffman codebook. In a specific example, the synonym set sequence U0 is traversed, and for each synonym set in U0, Obtain the coding field according to the semantic Huffman codebook mapping The mapping relationship is recorded as f:
[0043] Step 106: Sort all coding fields in sequence according to the order of the code words in the information sequence to obtain a coding sequence corresponding to the information sequence.
[0044] After determining the coding fields corresponding to the synonymous set, that is, determining the coding fields corresponding to the codewords, the coding fields are arranged in sequence according to the order of the codewords in the information sequence to obtain the coding sequence. The code sequence b is obtained by sequentially combining the codes, and the transmitter completes the encoding of the information sequence. Compared to traditional Huffman coding, which encodes codewords one by one, this embodiment encodes the synonymous set obtained by codeword mapping. This can shorten the average code length, further compress the code field, improve the compression efficiency of the code, and thus improve the transmission efficiency of the code sequence.
[0045] Step 108: Send the coded sequence to a receiving end, so that the receiving end decodes the coded sequence.
[0046] After encoding, the transmitter sends the coded sequence to the receiver, which decodes it to obtain a semantically undistorted decoded sequence. The decoding process also follows the semantic Huffman codebook. The receiver reads each code in the coded sequence one by one, traverses the code fields contained in the semantic Huffman codebook, finds the code field that matches the read code, and determines the corresponding synonym set. From this synonym set, a source symbol is extracted according to a specific rule as the decoding field corresponding to the code field. All the decoding fields are sequentially combined to obtain the decoded sequence.
[0047] Based on the above steps 102 to 108, the semantic-based Huffman coding method provided in this embodiment includes, in response to receiving an information sequence containing at least one codeword sent by a source, performing synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook to determine a synonymous set corresponding to the codeword. Each codeword is associated with a synonymous set based on semantics, and the semantics of each codeword contained in the synonymous set are the same. The coding field corresponding to the synonymous set is determined based on the pre-constructed semantic Huffman codebook, and each synonymous set corresponds to a coding field, that is, the coding field corresponding to each codeword in the synonymous set during the coding process is the same. According to the order of the codewords in the information sequence, all coding fields are sorted in sequence to obtain a coding sequence corresponding to the information sequence, and the coding sequence obtained by encoding the synonymous set is used to replace the traditional Huffman coding method, which can effectively compress the average code length of the coding sequence, improve the compression efficiency during the coding process, and then improve the transmission efficiency of the coding sequence, which is conducive to further improving the efficiency of communication technology. The coded sequence is sent to a receiving end so that the receiving end decodes the coded sequence. The obtained decoded sequence retains the semantics of the information sequence, ensuring that the transmission process is not distorted.
[0048] In some embodiments, performing synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook and determining the synonymous set corresponding to the codeword may include: based on semantic information of the codeword, querying the synonymous set corresponding to the semantic information in the synonymous mapping codebook as the synonymous set corresponding to the codeword; wherein, in the synonymous mapping codebook, the semantic information and the synonymous set have a one-to-one correspondence.
[0049] Specifically, the semantic information of the codeword is determined, the semantic information is found in the synonym mapping codebook, and the synonym set corresponding to the semantic information is used as the synonym set corresponding to the codeword. The mapping relationship is expressed as f: There is a one-to-one correspondence between semantic information and synonym sets, and all source symbols in a synonym set have the same semantic information. The method in this embodiment can quickly determine the synonym set corresponding to the codeword, thereby improving the encoding speed.
[0050] In some embodiments, before performing synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook, the above method may further include: segmenting the information sequence according to the minimum symbol syntax unit of the information source.
[0051] Specifically, before performing synonymous mapping on the codewords, the information in the information sequence needs to be segmented to obtain codewords. For example, if the source sequence is a section of English text, the minimum symbol grammatical unit is an English word, the English text is segmented into individual English words, and each English word is used as a codeword. If the source sequence is an irregular sequence of English letters, the minimum symbol grammatical unit is an English letter, the English letter sequence is segmented into individual English letters, and each English letter is used as a codeword. The method of this embodiment can effectively segment the information sequence, facilitating the subsequent encoding of the information sequence.
[0052] In some embodiments, the method of pre-constructing the synonymous mapping codebook may include: obtaining all source symbols of the source, classifying all source symbols according to the semantic information of the source symbols, combining source symbols with the same semantic information into a synonymous set, and constructing the synonymous mapping codebook based on the mapping relationship between the source symbols and the synonymous set.
[0053] Specifically, when constructing a synonym mapping codebook, it is necessary to obtain all the source symbols of the source, classify the source symbols according to the semantic information of each source symbol, and divide the source symbols with the same semantic information into a synonym set, that is, all the source symbols contained in a synonym set have the same semantic information. If a source symbol does not have other corresponding source symbols with the same semantic information, then the source symbol is divided into a synonym set separately. After all the source symbols are divided, a synonym mapping codebook is constructed based on the mapping relationship between each source symbol and the synonym set. In the embodiment of the present disclosure, a semantic knowledge base can be constructed based on an existing database, and then the wish symbols are classified based on the constructed semantic knowledge base. Specifically, for example, for text sources, a synonym dictionary based on synonyms can be constructed for classification. Specifically, different synonym construction methods can be used to construct an entropy function synonym dictionary for different sources and needs.
[0054] Figure 2 shows the correspondence between the semantic information space and the grammatical information space. As shown in Figure 2, the semantic information space contains six different types of semantic information, specifically v1, v2, v3, v4, v5, and v6. The grammatical information space contains six different synonymous subsets, specifically U1, U2, U3, U4, U5, and U6. Each synonymous subset contains several source symbols, and each synonymous subset corresponds to a piece of semantic information. Based on the semantic information, v1 is mapped to U1, v2 is mapped to U2, v3 is mapped to U3, v4 is mapped to U4, v5 is mapped to U5, and v6 is mapped to U6.
[0055] In a specific example, all source symbols are obtained from the source, and all source symbols are numbered and arranged to form a source symbol set Among them, N i Indicates the number of source symbols. Construct a synonym set sequence U, which contains multiple synonym sets i s The process of classifying all source symbols according to semantics is as follows:
[0056] 1) Initialize the synonym set sequence U to be empty and initialize the variable j = 1;
[0057] 2) For the source symbol u xi , traverse the current synonym set sequence U, if there is a synonym set in U with the same semantic information as the source symbol Then the source symbol u xi Partition into synonym sets In the middle, the traversal ends and the variable j is increased by 1;
[0058] 3) If there is no synonym set in U with the same semantic information as the source symbol, the source symbol is assigned to an empty synonym set, and this synonym set is placed at the end of the synonym set sequence U, and the variable j is incremented by 1;
[0059] 4) Determine whether j is less than N i If it is less than, return to step 2), otherwise end the process.
[0060] Through the above steps 1) to 4), the classification of all source symbols is completed, and the length of the current synonymous set sequence is recorded as Each synonym set contains at least one source symbol. Different synonym sets do not intersect with each other. The union of all synonym sets is the same as all source symbols, that is,
[0061] In some embodiments, referring to FIG3 , the method for pre-constructing the semantic Huffman codebook may include the following steps:
[0062] Step 202: Obtain the prior probability of the source symbol.
[0063] Specifically, the prior probability distribution of the source symbols can be obtained through statistical analysis. By performing statistical analysis on a large amount of data, the frequency of occurrence of each symbol in the source can be estimated. This method is suitable for situations where a large number of samples are available, such as text, images, audio, etc. By performing statistics on the data, the relative frequency of each source symbol can be obtained, thereby estimating its probability distribution. In addition, in some cases, a probability model can be used to fit the distribution of the source. For example, a Gaussian distribution, a Poisson distribution, or other statistical models are used to describe the statistical characteristics of the source. After all the source symbols are arranged, the i-th source symbol u xi The prior probability of xi ).
[0064] Step 204: Calculate the total probability of the synonymous set based on the prior probability.
[0065] Furthermore, calculating the total probability of the synonymous set based on the prior probability includes:
[0066] The sum of the prior probabilities of all the source symbols included in the synonymous set is taken as the total probability of the synonymous set.
[0067] After determining all the source symbols contained in the synonymous set, the prior probability of each source symbol is summed up to obtain the total probability of the synonymous set, that is,
[0068] Step 206: Use the synonym sets as leaf nodes and construct a Huffman tree using a Huffman tree construction method based on the total probability of the synonym sets.
[0069] The total probability of all synonymous sets is used as the leaf nodes of the Huffman tree. The two leaf nodes with the lowest total probability are selected and merged into a parent node, where the total probability of the parent node is equal to the sum of the total probabilities of the two leaf nodes. Next, the total probabilities of the parent nodes and the unmerged leaf nodes are sorted by numerical value. The two nodes with the lowest total probability are again selected and merged to form a new parent node, where the total probability of the new parent node is equal to the sum of the total probabilities of the merged two nodes. This process is repeated until the total probability of the new parent node after the merger is equal to 1. This parent node serves as the root node, forming the Huffman tree.
[0070] FIG4 shows a schematic diagram of constructing a Huffman tree. As shown in FIG4 , the synonym set sequence U contains Synonymous sets, each synonymous set is used as a leaf node of the Huffman tree. First, select the leaf nodes U1 and U2 with the smallest total probability. The total probability of node U1 is p(U1), and the total probability of node U2 is p(U2). Merge them to get the parent node Parent Node The total probability is And assign path symbols to nodes U1 and U2. The path symbol corresponding to node U1 is 0, and the path symbol corresponding to node U2 is 1. The path symbol assignment principle is that the path symbols corresponding to the left branch are all 0 (or 1), and the path symbols corresponding to the right branch are all 1 (or 0). At this time, excluding the leaf nodes U1 and U2, the two nodes with the smallest total probability are nodes U3 and U4. The total probability of node U3 is p(U3), and the total probability of node U4 is p(U4). Then merge nodes U3 and U4 to get the parent node Parent Node The total probability is If we remove the leaf nodes U1, U2, U3, and U4, the node with the smallest total probability is and Then the node and nodes Merge again, and so on, to get the final root node U T , and the root node U T The total probability is 1.
[0071] The construction process of the Huffman tree can also be described by the following process:
[0072] 1) Initialize the Huffman tree;
[0073] 2) Select two synonymous sets with the smallest total probability in the synonymous set sequence U and As a child node, i s ≠j s , the two child nodes are merged into one parent node, denoted as U t . Parent node U t The total probability is equal to Child nodes and and parent node U t Form a branch and insert the branch into the Huffman tree. The specific insertion method includes: if there is a node in the existing Huffman tree and According to the node and Merge to the root in the Huffman tree to generate the parent node U t ; If there is a node in the current Huffman tree or Then insert the non-existent node as the leaf node of the current Huffman tree, and then and Merge to the root in the Huffman tree to generate the parent node U t; If the node does not exist in the current Huffman tree and Then the node and Are inserted as the leaf nodes of the current Huffman tree, and then according to the node and Merge to the root in the Huffman tree to generate the parent node U t ;
[0074] 3) The node and Remove the node U from the synonym set sequence U. t Add to the synonym set sequence U, and reorder all synonym sets contained in the current synonym set sequence U according to the total probability value;
[0075] 4) Determine whether there is only one synonym set left in the synonym set sequence U. If not, repeat steps 2) to 3). If so, end the process and complete the construction of the Huffman tree.
[0076] Step 208: Determine the encoding field corresponding to the synonym set based on the Huffman tree.
[0077] FIG5 shows a schematic diagram of the structure of a Huffman tree. As shown in FIG5 , the Huffman tree includes three synonym sets, namely U1, U2, and U3. When determining the coding field corresponding to synonym set U1, starting from the root node of the Huffman tree and following the illustrated path to reach leaf node U1, all path symbols passed through are successively 00, that is, the coding field corresponding to synonym set U1 is 00. Similarly, when determining the coding field corresponding to synonym set U2, all path symbols passed through are successively 1, that is, the coding field corresponding to synonym set U2 is 1. When determining the coding field corresponding to synonym set U3, all path symbols passed through are successively 01, that is, the coding field corresponding to synonym set U3 is 01. According to the method of this step, the coding field corresponding to each synonym set can be obtained based on the Huffman tree.
[0078] It should be noted that the coding fields corresponding to the source symbols contained in each synonymous set are all the coding fields corresponding to the synonymous set. In this application, a Huffman tree is constructed using the total probability of the synonymous sets. During the encoding process, the synonymous sets are encoded instead of the source symbols. This can effectively reduce the average code length of the encoded sequence and improve the compression efficiency of the encoding process.
[0079] Step 210: Construct the semantic Huffman codebook based on the correspondence between the synonym sets and the coding fields.
[0080] After the Huffman tree is constructed through the above steps, the coding field of each synonym set is determined, and a mapping relationship is formed between the synonym set and the coding field, which is stored in the semantic Huffman code book, so that the coding field corresponding to each synonym set can be quickly determined based on the Huffman code book.
[0081] The present application also provides a semantic-based Huffman decoding method, which is applied to a receiving end. Referring to FIG6 , the method includes the following steps:
[0082] Step 302: In response to receiving a coding sequence sent by a transmitting end, determine a synonym set sequence corresponding to the coding sequence based on a pre-constructed Huffman codebook.
[0083] Specifically, after receiving the code sequence, each code in the code sequence is traversed, and synonym sets corresponding to the traversed codes are determined based on the Huffman codebook. All synonym sets obtained after the traversal are combined into a synonym set sequence in the determined order. The method for constructing the Huffman codebook and the synonym set in this step is the same as in the previous embodiment and will not be repeated here.
[0084] Step 304: Extract a source symbol from each synonymous set in the synonymous set sequence according to a preset extraction method as a target source symbol.
[0085] Specifically, each source symbol in a synonymous set has the same semantic information. After the synonymous set is determined, extracting a source symbol from the synonymous set according to a preset extraction method can accurately represent the semantic information encoded in the coding sequence. Preset extraction methods can include random selection, fixed selection, and maximum probability selection. Among them, fixed selection refers to selection based on a given fixed symbol; maximum probability selection refers to extraction based on the probability of each symbol in the synonymous set, for example, only selecting the symbol with the highest probability. The specific extraction method is not specifically limited in this embodiment.
[0086] Step 306: Sort the target source symbols in sequence according to the order of the synonymous sets in the synonymous set sequence to obtain a decoding sequence corresponding to the encoding sequence.
[0087] Specifically, after all target source symbols are extracted according to a preset extraction method, the target source symbols are sorted according to the order of the synonymous sets in the synonymous set sequence to obtain the corresponding decoding sequence, and the semantic information represented by the decoding sequence is the same as that of the information sequence.
[0088] In some embodiments, determining the synonym set sequence corresponding to the encoding sequence based on a pre-constructed Huffman codebook includes:
[0089] Read each code in the code sequence one by one, and determine the synonym sets corresponding to the read code fields in sequence based on the Huffman code book;
[0090] In response to the completion of reading the coding sequence, all synonymous sets are sorted according to the sequentially determined order to form the synonymous set sequence.
[0091] Specifically, each encoding symbol in the encoding sequence is read one by one. Based on the several encoding symbols that have been read, a corresponding synonymous set is determined using the Huffman codebook. Each synonymous set corresponds to a codeword, which includes several encoding symbols. For example, the codeword corresponding to the synonymous set is U, which includes three encoding symbols 001, i.e., U=001. Furthermore, for the several bits of encoding that have been read, the Huffman codebook is used to find the same encoding field as the several bits of encoding that have been read, and then a synonymous set corresponding to the encoding field is determined.
[0092] After reading is completed, the synonym sets are sorted according to the order of the previously determined synonym sets to form a synonym set sequence. During decoding, the synonym sets are determined by encoding to ensure that the decoded sequence has the same semantic information as the encoded sequence.
[0093] In a specific example, the decoding process at the receiving end can also be described by the following process:
[0094] 1) Initialize the decoding position identifier l=1, initialize the sequence number i s =1;
[0095] 2) Starting from the lth bit of the coding sequence b, according to the current code word 0 or 1 of the lth bit, start searching along the path from the root node of the Huffman tree. Each time you move a node, the decoding position identifier l is increased by 1 until a leaf node is found. The current code word decoding is ended and the path symbols 0 or 1 of the path passed are arranged and combined in order to obtain the current coding field. Determine the corresponding synonymous set based on the coding field through the Huffman code book Serial number i s Add 1;
[0096] 3) Check whether the decoding position identifier l reaches the length of the encoding sequence b. If not, return to step 2). If it reaches, set the synonym set Connect them in sequence according to the sequence number to form a synonym set sequence;
[0097] 4) Select a source symbol u from each synonymous set in the synonymous set sequence according to the preset extraction method xi , as the target source symbols, the target source symbols are connected in sequence to obtain the decoding sequence c.
[0098] The encoding and decoding process of this application is described below through specific embodiments.
[0099] This embodiment includes 6 types of source symbols, and the corresponding prior probabilities are shown in Table 1. The calculated source entropy H=2.1219 bits.
[0100] Table 1 Source prior probability table
[0101] According to the traditional Huffman coding method, the codebook can be constructed to obtain Table 2:
[0102] Table 2 Source symbol coding table
[0103] The average code length calculated from Table 2 is: L1 = 0.2×3 + 0.4×1 + 0.2×3 + 0.1×3 + 0.1×3 = 2.2 bits.
[0104] Given a synonymous mapping, synonym set U1 = {u1}, synonym set U2 = {u2,u3}, synonym set U3 = {u4,u5}. The prior probability of the synonym set is shown in Table 3:
[0105] Table 3 Prior probability table of synonymous sets
[0106] According to the Huffman tree shown in Figure 5, the encoding fields corresponding to the synonym sets are shown in Table 4:
[0107] Table 4 Coded fields corresponding to synonym sets
[0108] The average code length calculated in Table 4 is: L2 = 0.2×2 + 0.6×1 + 0.2×2 = 1.4 bits. Compared with L1, L2 significantly shortens the average code length.
[0109] For a given information sequence [u1u3u2u2u5u2u1u3u4u2], the coding sequence according to traditional Huffman coding is [0100111100110100110001], and the code length is 22 bits. The coding sequence obtained according to the semantic Huffman coding in this application is [00111011001011], and the code length is 14 bits. After decoding, the decoding sequence of traditional Huffman coding is [u1u3u2u2u5u2u1u3u4u2], and the decoding sequence after semantic Huffman coding is [u1u2u3u3u4u2u1u2u4u3]. It can be seen that semantic Huffman coding has higher compression efficiency than traditional Huffman coding, and can even break through the Shannon information entropy of the source, all thanks to the additional information brought by the synonym set. Although the decoded sequence of semantic Huffman coding does not necessarily perfectly match the information sequence sent by the source, according to the definition of synonymous sets, the two have the same semantic information, so it does not affect the final transmission decision result. The results of this embodiment show that the semantic Huffman coding method can significantly improve the compression efficiency of the source compared to traditional Huffman coding schemes, thus proving the feasibility and high effectiveness of the semantic Huffman coding proposed in this application.
[0110] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0111] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0112] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a semantic-based Huffman coding device.
[0113] 7 , the semantic-based Huffman coding apparatus is applied to a transmitting end and includes:
[0114] A first determining module 402 is configured to, in response to receiving an information sequence including at least one codeword sent by a source, perform synonymous mapping on the codeword based on a pre-built synonymous mapping codebook to determine a synonymous set corresponding to the codeword;
[0115] The second determining module 404 is configured to determine the coding field corresponding to the synonymous set according to a pre-built semantic Huffman codebook;
[0116] The encoding module 406 is configured to sequentially sort all the encoding fields according to the order of the codewords in the information sequence to obtain an encoding sequence corresponding to the information sequence;
[0117] The sending module 408 is configured to send the coded sequence to a receiving end, so that the receiving end decodes the coded sequence.
[0118] In some embodiments, the first determining module 402 is further configured to query the synonymous mapping codebook for a synonymous set corresponding to the semantic information of the codeword as the synonymous set corresponding to the codeword based on the semantic information of the codeword, wherein in the synonymous mapping codebook, the semantic information and the synonymous set have a one-to-one correspondence.
[0119] In some embodiments, before performing synonymous mapping on the codeword based on a pre-built synonymous mapping codebook, the first determining module 402 is further configured to segment the information sequence according to the minimum symbol syntax unit of the information source.
[0120] In some embodiments, a construction module is further included, which is configured to obtain all source symbols of the source, classify all source symbols according to the semantic information of the source symbols, combine source symbols with the same semantic information into a synonymous set, and construct the synonymous mapping codebook based on the mapping relationship between the source symbols and the synonymous set.
[0121] In some embodiments, the construction module is further configured to obtain the prior probability of the source symbol; calculate the total probability of the synonymous set based on the prior probability; use the synonymous set as a leaf node, and construct a Huffman tree based on the total probability of the synonymous set using a Huffman tree construction method; determine the coding field corresponding to the synonymous set based on the Huffman tree; and construct the semantic Huffman codebook based on the correspondence between the synonymous set and the coding field.
[0122] In some embodiments, the construction module is further configured to take the sum of the prior probabilities of all source symbols included in the synonymous set as the total probability of the synonymous set.
[0123] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a semantic-based Huffman coding device.
[0124] 8 , the semantic-based Huffman decoding device, applied to a receiving end, includes:
[0125] The third determining module 502 is configured to, in response to receiving a coding sequence sent by the transmitting end, determine a synonym set sequence corresponding to the coding sequence based on a pre-constructed Huffman codebook;
[0126] The extraction module 504 is configured to extract a source symbol as a target source symbol from each synonymous set in the synonymous set sequence according to a preset extraction method;
[0127] The decoding module 506 is configured to sequentially sort the target source symbols according to the order of the synonymous sets in the synonymous set sequence to obtain a decoding sequence corresponding to the coding sequence.
[0128] In some embodiments, the third determination module 502 is further configured to read each code in the code sequence one by one, and determine the synonymous sets corresponding to the read code fields in sequence based on the Huffman code book; in response to the completion of reading the code sequence, sort all the synonymous sets in the order determined in sequence to form the synonymous set sequence.
[0129] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0130] The apparatus of the above embodiment is used to implement the corresponding semantic-based Huffman encoding method or semantic-based Huffman decoding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0131] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the semantic-based Huffman encoding method or the semantic-based Huffman decoding method described in any of the above embodiments is implemented.
[0132] FIG9 shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0133] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0134] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0135] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0136] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0137] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0138] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0139] The electronic device of the above embodiment is used to implement the corresponding semantic-based Huffman encoding method or semantic-based Huffman decoding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0140] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the semantic-based Huffman encoding method or the semantic-based Huffman decoding method as described in any of the above embodiments.
[0141] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0142] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the semantic-based Huffman encoding method or semantic-based Huffman decoding method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0143] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the method described in any of the above embodiments, which has the beneficial effects of the corresponding method embodiments and will not be repeated here.
[0144] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0145] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0146] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0147] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.
Claims
1. A semantic-based Huffman coding method, applied to a sending end, includes: In response to receiving an information sequence containing at least one codeword sent by a source, perform a synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook, and determine a synonymous set corresponding to the codeword; Determine an encoding field corresponding to the synonymous set according to a pre-constructed semantic Huffman codebook; According to the order of the codewords in the information sequence, sort the encoding fields in sequence to obtain an encoding sequence corresponding to the information sequence; And Send the encoding sequence to a receiving end, so that the receiving end decodes the encoding sequence.
2. The method according to claim 1, wherein The performing a synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook and determining a synonymous set corresponding to the codeword includes: Based on the semantic information of the codeword, query in the synonymous mapping codebook for a synonymous set corresponding to the semantic information as the synonymous set corresponding to the codeword; wherein, in the synonymous mapping codebook, there is a one-to-one correspondence between the semantic information and the synonymous set.
3. The method according to claim 1, further includes: Before performing a synonymous mapping on the codeword based on a pre-constructed synonymous mapping codebook, perform a segmentation process on the information sequence according to the minimum symbol syntax unit of the source.
4. The method according to claim 1, further includes: Obtain all source symbols of the source; Classify all source symbols according to the semantic information of the source symbols; Combine source symbols with the same semantic information into a synonymous set; And Construct the synonymous mapping codebook based on the mapping relationship between the source symbols and the synonymous sets.
5. The method according to claim 4, further includes: Obtain the prior probability of the source symbols; Calculate the total probability of the synonymous set based on the prior probability; Use the synonymous set as a leaf node, and based on the total probability of the synonymous set, adopt a method for constructing a Huffman tree to construct a Huffman tree; Based on the Huffman tree, determine an encoding field corresponding to the synonymous set; And Construct the semantic Huffman codebook based on the corresponding relationship between the synonymous set and the encoding field.
6. The method according to claim 5, wherein, The calculating the total probability of the synonymous set based on the prior probability includes: Taking the sum of the prior probabilities of all source symbols included in the synonymous set as the total probability of the synonymous set.
7. A semantic-based Huffman decoding method, applied to a receiving end, includes: In response to receiving an encoding sequence sent by a sending end, determine a synonymous set sequence corresponding to the encoding sequence based on a pre-constructed Huffman codebook; Extract a source symbol as a target source symbol from each synonymous set in the synonymous set sequence according to a preset extraction method; And According to the order of the synonymous sets in the synonymous set sequence, sort the target source symbols in sequence to obtain a decoding sequence corresponding to the encoding sequence.
8. The method according to claim 7, wherein The determining a synonymous set sequence corresponding to the encoding sequence based on a pre-constructed Huffman codebook includes: Read each code in the coded sequence one by one, and sequentially determine the synonym set corresponding to the read code field based on the Huffman codebook; and In response to determining that the reading of the coded sequence is completed, sort all the synonym sets in the determined order to form the synonym set sequence.
9. A semantic-based Huffman coding device, comprising: A first determination module, configured to, in response to receiving an information sequence including at least one codeword sent by a source, perform a synonym mapping on the codeword based on a pre-constructed synonym mapping codebook, and determine the synonym set corresponding to the codeword; A second determination module, configured to determine the coding field corresponding to the synonym set according to a pre-constructed semantic Huffman codebook; A coding module, configured to sequentially sort all the coding fields according to the order of the codewords in the information sequence to obtain the coded sequence corresponding to the information sequence; And A sending module, configured to send the coded sequence to a receiving end, so that the receiving end decodes the coded sequence.
10. A semantic-based Huffman decoding device, comprising: A third determination module, configured to, in response to receiving a coded sequence sent by a sending end, determine the synonym set sequence corresponding to the coded sequence based on a pre-constructed Huffman codebook; An extraction module, configured to extract a source symbol as a target source symbol from each synonym set in the synonym set sequence according to a preset extraction method; And A decoding module, configured to sequentially sort the target source symbols in the order of the synonym sets in the synonym set sequence to obtain the decoded sequence corresponding to the coded sequence. When the processor executes the program, the method described in any one of claims 1 to 6 or claims 7 to 8 is implemented.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer instructions are used to cause a computer to execute the method described in any one of claims 1 to 6 or claims 7 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, Computer program instructions, when the computer program instructions run on a computer, cause the computer to execute the method described in any one of claims 1 to 6 or claims 7 to 8.
13. A computer program product, comprising:
Citation Information
Patent Citations
Theme word extraction method, and method and device for obtaining related digital resource by using same
CN105224521A
Compression method of data reported by terminal, sending end and receiving end
CN108737392A
Text carrier-free information hiding method based on synonym expansion and label transfer
CN112989809A
Semantic encoding method, semantic encoding device, semantic decoding method and semantic decoding device
CN115955297A
Clustering correction method based on polysemous words and synonyms
CN116384378A