Semantic coding and decoding method and device based on textrank
Through textrank-based semantic coding method and Bert-Base-Chinese model decoding technology, the problem that Chinese text semantic coding methods in the prior art cannot take into account communication effectiveness, delay and reliability, and efficient and stable semantic coding and decoding effects are achieved.
Patent Information
- Application Number
- CN202510078425.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, Chinese text semantic coding methods cannot take into account communication effectiveness, delay and reliability at the same time.
The texttrank-based semantic encoding method is adopted to encode the sent text through keyword extraction and hollowing out processing, and decode it using the Bert-Base-Chinese model to recover the text.
It realizes stable and reliable text compression, shortens the encoding and decoding delay, and improves the universality of communication.
Smart Images

Figure CN120049894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information coding, and particularly relates to a semantic coding and decoding method and device based on TextRank. Background Art
[0002] With the explosive growth of Internet application requirements, the amount of data in the network has also increased. The contradiction between the limited resources of traditional communication based on Shannon's information theory and the increasing demand for traffic has become increasingly prominent. The trend of the development of traditional communication systems is to increase the physical dimension of information transmission and change the way of resource utilization. The bottleneck of the traditional communication development model with improving bandwidth and spectrum efficiency as the core is that the current spectrum efficiency has approached the upper limit given by information theory. More and more research has begun to shift the goal of communication from the symbol level to the semantic level, saving bandwidth resources by transmitting semantics rather than complete content.
[0003] At present, scholars at home and abroad have proposed many semantic encoding and decoding strategies based on deep learning. There are mainly two semantic encoding and decoding strategies. The first strategy is that the mapping method is formulated by the author himself, and a traditional encoder is used for encoding. The second strategy directly uses a pre-trained language model to extract the semantics of the source and perform encoding. At present, most existing methods have various defects, mainly communication effectiveness, latency, and reliability. The methods using the first strategy require the researcher himself to have a certain research on the text, and do not have the universality and generality of multiple languages. Some methods using the second strategy have good reliability and latency, but insufficient effectiveness, and some methods have good reliability and effectiveness, but too long latency. That is to say, in current semantic communication, the Chinese text semantic encoding and decoding method cannot simultaneously take into account communication effectiveness, latency, and reliability. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a semantic coding and decoding method and device based on TextRank.
[0005] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] The present invention provides a semantic coding method based on TextRank, including:
[0007] Obtain the text to be sent;
[0008] Extract keywords from the text to be sent by using a keyword extraction algorithm to obtain a keyword table;
[0009] Perform blanking processing on the non-keyword words in the said sending text according to the said keyword table and the preset blanking rate, to obtain the text after blanking processing; the text after blanking processing is the text obtained by randomly removing some non-keyword words in the said sending text;
[0010] Encode the text after blanking processing to obtain a digital code sequence;
[0011] Process the digital code sequence to obtain a digital code sequence to be decoded and send it.
[0012] The present invention also provides a semantic decoding method based on textrank, including:
[0013] Receive the digital code sequence to be decoded;
[0014] Process the digital code sequence to be decoded to obtain a digital code sequence;
[0015] Perform text decoding on the digital code sequence to obtain the decoded text, wherein the decoded text is the text after blanking processing obtained by randomly removing some non-keyword words in the sending text with the preset blanking rate;
[0016] Fill in non-keyword words into the text after blanking processing to restore the semantic information of the said sending text, to obtain the restored text.
[0017] The present invention also provides a semantic encoding device based on textrank, including a processor, a communication interface, a memory and a communication bus, and the processor, the communication interface and the memory complete mutual communication through the communication bus;
[0018] The memory is used to store a computer program;
[0019] When the processor is used to execute the program stored in the memory, it realizes the steps of the above-mentioned semantic encoding method based on textrank.
[0020] The present invention also provides a semantic decoding device based on textrank, including a processor, a communication interface, a memory and a communication bus, and the processor, the communication interface and the memory complete mutual communication through the communication bus;
[0021] The memory is used to store a computer program;
[0022] When the processor is used to execute the program stored in the memory, it realizes the steps of the above-mentioned semantic decoding method based on textrank.
[0023] Compared with the prior art, the beneficial effects of the present invention:
[0024] 1) During the process of text encoding of the present invention, when performing the hollowing-out process on the transmitted text, only keywords and non-keywords are distinguished, and only the words to be removed are selected from the non-keywords according to the hollowing-out rate. Therefore, the compression of the text by the present invention is stable, reliable, and effective. In addition, during the process of text encoding of the present invention, only some algorithms are used, and no complex language model is used for processing. Therefore, there is no problem of long delay caused by the high complexity of the model structure. Therefore, the delay of the semantic encoding method of the present invention is short;
[0025] 2) During the process of text decoding of the present invention, the Bert-Base-Chinese model is used for text restoration. Since the model parameters and complexity of the Bert-Base-Chinese model are low, there is no problem of long delay caused by the high complexity of the model structure. Therefore, the delay of the semantic decoding method of the present invention is short. In addition, since the model parameters and complexity of the Bert-Base-Chinese model are low, the deployment difficulty is low, and it can be deployed on any device for decoding, so the communication universality is higher.
[0026] The following will further describe the present invention in detail with reference to the drawings and specific embodiments. Description of the Drawings
[0027] Figure 1 is a schematic flowchart of a semantic encoding method based on textrank provided by an embodiment of the present invention;
[0028] Figure 2 is a schematic flowchart of a semantic decoding method based on textrank provided by an embodiment of the present invention;
[0029] Figure 3 is a working flowchart of a semantic encoding and decoding system based on textrank provided by an embodiment of the present invention;
[0030] Figure 4 is a curve graph showing the variation of the hollowing-out delay with the character length of the text provided by an embodiment of the present invention. Detailed Embodiments
[0031] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0032] Figure 1 is a schematic flowchart of a semantic encoding method based on textrank provided by an embodiment of the present invention. This encoding method is applied to the semantic encoding end, such as Figure 1 shown, and the method includes:
[0033] S101. Obtain the text to be sent.
[0034] Here, the text to be sent can be any Chinese text, and the present invention does not limit this.
[0035] S102. Use a keyword extraction algorithm to extract the keywords in the text to be sent, and obtain a keyword table.
[0036] Here, the keyword extraction algorithm can be the textrank algorithm, and the present invention does not limit this.
[0037] S103. According to the keyword table and a preset blanking rate, perform blanking processing on the non-keywords in the text to be sent, and obtain the text after blanking processing; the text after blanking processing is the text obtained by randomly removing some non-keywords in the text to be sent.
[0038] Here, the preset blanking rate is less than or equal to 0.5.
[0039] S104. Encode the text after blanking processing to obtain a digital code sequence.
[0040] S105. Process the digital code sequence to obtain a digital code sequence to be decoded and send it.
[0041] In some embodiments, the above S102 is implemented through the following steps:
[0042] S1021. According to a preset ratio range and the total number of words in the text to be sent, obtain the number range of keywords.
[0043] Specifically, by multiplying the preset ratio range by the total number of words in the text to be sent, the number range of keywords in the text to be sent can be obtained.
[0044] Exemplarily, the preset ratio range is 15% - 30%. For example, 25%. By setting the proportion of keywords to 25%, both the compression rate and the effectiveness and stability of encoding can be taken into account.
[0045] S1022. Use the textrank keyword extraction algorithm to sort the words in the text to be sent, and obtain a sorting result; the sorting result represents the sorting of the importance of the words in the text to be sent for understanding the text to be sent.
[0046] S1023. According to the sorting result, regard the top N words with the highest importance as keywords, and obtain a keyword table; N is any integer in the number range of keywords.
[0047] Here, all the selected keywords can be stored in a table, so as to obtain a keyword table for use in subsequent blanking processing.
[0048] In some embodiments, the above S103 is implemented through the following steps:
[0049] S1031. Perform word segmentation on the sent text to obtain multiple words arranged in order.
[0050] S1032. Determine the total number M of non-keyword words to be removed according to the total number of the multiple words arranged in order and a preset deletion rate.
[0051] Specifically, by multiplying the total number of words obtained after word segmentation by the preset deletion rate, the total number M of non-keyword words to be removed can be obtained.
[0052] S1033. Randomly shuffle the arrangement order of the multiple words arranged in order to obtain multiple randomly arranged words.
[0053] S1034. According to the arrangement order of the multiple randomly arranged words, determine whether the current word is in the keyword list. If it is a keyword, continue to determine the next word. If it is not a keyword, determine whether the adjacent word of the current word in the sent text has been removed. If it has been removed, continue to determine the next word. If it has not been removed, determine whether the character length of the current word is less than or equal to a preset length. If it is greater than the preset length, continue to determine the next word. If it is less than or equal to the preset length, remove the current word. After that, count and determine whether the total number of the words that have been removed currently is less than or equal to M. If it is less than M, continue to determine the next word until M non-keyword words with character lengths less than or equal to the preset length and non-adjacent positions in the sent text are removed from the multiple randomly arranged words, obtaining the text after deletion processing.
[0054] Exemplarily, the preset length can be 2 or 3. When the preset length is 2, the quotient of the semantic similarity corresponding to the text after deletion processing divided by the compression rate is higher than the quotient of the semantic similarity corresponding to the text after deletion processing obtained when the preset length is 3. Here, the semantic similarity corresponding to the text after deletion processing is the semantic similarity between the text after deletion processing and the sent text, and the compression rate corresponding to the text after deletion processing is the quotient obtained by dividing the number of bits of the Huffman code of the text after deletion processing by the number of bits of the Huffman code of the sent text.
[0055] Here, through S1031 to S1034, some non-keyword words can be removed from the sent text, and these non-keyword words removed are some non-keyword words in the sent text with character lengths less than or equal to 2 and non-adjacent to each other. Therefore, the compression of the text by the present invention is effective and stable.
[0056] In some embodiments, the above S104 is implemented through the following steps: S1041. Use a Huffman encoder to perform Huffman source coding on the text after the hollowing-out process according to the code table, obtaining a digital code sequence. Here, the digital code sequence is a 01 code sequence.
[0057] In some embodiments, the above S105 is implemented through the following steps: After successively performing modulation, encryption, and channel coding on the digital code sequence, a digital code sequence to be decoded is obtained; the digital code sequence to be decoded is sent. The present invention does not limit the specific modulation, encryption, and channel coding methods.
[0058] The present invention also provides a flowchart of a semantic decoding method based on TextRank. This encoding method is applied to the semantic decoding end, such as Figure 2 as shown, and this method includes:
[0059] S201. Receive the digital code sequence to be decoded.
[0060] Here, the digital code sequence to be decoded is generated by using the above-mentioned semantic encoding method based on TextRank.
[0061] S202. Process the digital code sequence to be decoded to obtain a digital code sequence.
[0062] S203. Perform text decoding on the digital code sequence to obtain the decoded text. Among them, the decoded text is the text after the hollowing-out process obtained by randomly removing some non-keywords in the sent text with a preset hollowing-out rate.
[0063] S204. Fill in non-keywords into the text after the hollowing-out process to restore the semantic information of the sent text, obtaining the restored text.
[0064] In some embodiments, the above S202 is implemented through the following steps: After successively performing channel decoding, decryption, and modulation and demodulation on the digital code sequence to be decoded, a digital code sequence is obtained. The present invention does not limit the specific channel decoding, decryption, and modulation and demodulation methods.
[0065] In some embodiments, the above S203 is implemented through the following steps: Use a Huffman decoder to perform Huffman source decoding on the digital code sequence according to the code table to obtain the decoded text.
[0066] In some embodiments, the above S204 is implemented through the following steps:
[0067] S2041. Use the Bert-Base-Chinese model to generate, according to the text after the hollowing-out process, the word corresponding to each empty position in the text after the hollowing-out process and the probability of each word, where each empty position is the position of a non-keyword removed in the sent text.
[0068] Here, the Bert-Base-Chinese model processes a maximum of 512 characters each time. When the total character length of the text after blanking processing exceeds 512, the text after blanking processing can be divided into several segments, and each segment of text is input into the Bert-Base-Chinese model to obtain the words corresponding to each blank position in this segment of text and the probability of each word.
[0069] S2042. Select the word with the highest probability from the words corresponding to each blank position. When the probability of the word with the highest probability is greater than or equal to the preset probability threshold, fill the word with the highest probability into the blank position. When the probability of the word with the highest probability is less than the preset probability threshold, do not fill a word into the blank position. After processing all the blank positions in the text after blanking processing, a restored text is obtained.
[0070] Exemplarily, the preset probability threshold can be 0.3.
[0071] Here, when performing text filling, filling the word with a probability greater than or equal to 0.3 into the corresponding blank position and not filling the word into the blank position when the maximum probability of the word corresponding to the blank position is less than 0.3 can prevent the model from filling the answers with low certainty into the blanks. Therefore, the possibility of filling incorrect answers into the blank positions is reduced, making the semantic similarity between the sent text and the finally decoded restored text high and stable, that is, stable and reliable.
[0072] In some embodiments, other Chinese filling models can also be used for decoding, such as the ERNIE model, etc.
[0073] The following is through Figure 3 Describe the processes and relationships of the semantic encoding method based on textrank and the semantic decoding method based on textrank provided by the present invention. As Figure 3As shown, when the semantic encoding end needs to send text to the semantic decoding end, the semantic encoding end first preprocesses the text to be sent through keyword extraction and blanking processing to obtain the text after blanking processing. Then, Huffman encoding is performed on the text after blanking processing to obtain a digital code sequence. Then, modulation, encryption, and channel encoding processing can be performed on the digital code sequence, or only channel encoding processing can be performed on the digital code sequence to obtain the digital code sequence to be decoded. Then, the digital code sequence to be decoded is sent to the semantic decoding end through the channel. After receiving the digital code sequence to be decoded, the semantic decoding end first performs demodulation, decryption, and channel decoding processing on the digital code sequence to be decoded, or only performs channel decoding processing on the digital code sequence to be decoded to obtain a digital code sequence. Then, the semantic decoding end continues to perform Huffman decoding on the digital code sequence to obtain the decoded text. Then, filling in the blanks is performed according to the decoded text to obtain the restored text, and the restored text is the text received by the semantic decoding end.
[0074] Assume that the text to be sent by the semantic encoding end is T, and the text after blanking processing obtained by performing blanking processing on text T is The blanked part (i.e., the removed part) is T'; the restored text of the semantic decoding end is Y, and the restored effective semantic part is The ambiguous semantic part is Tq, H(.) represents semantic information, and H(Tq) are both positive values. I(X:Y) is the mutual semantic information between the semantic encoding end and the semantic decoding end. Then the relationship between these texts can be expressed by the following formula:
[0075]
[0076] Next, an example is used to specifically show the principle of the semantic-based encoding and decoding method of the present invention. Specifically, a part of the continuous text in a famous novel is selected as the text T to be sent by the semantic encoding end.
[0077] 1) The text T to be sent is specifically:
[0078] T = "Early in the morning, a warm sunbeam slipped through the crack of the curtain and cast a golden light trace on the wooden floor. I pushed open the window, and the moist spring breeze came blowing in with the fragrance of the soil, a bit cool but refreshing my spirit. Somehow, the old peach tree in the yard had already burst into bloom, and the pink and white petals were gently quivering in the morning light, like the shy blushing cheeks of a young girl. A few early-rising bees were buzzing cheerfully as they shuttled among the flowers. I noticed that on the thick branch of the peach tree, a sparrow was busily carrying dry grass in its beak. It had chosen this place to build its nest. Walking along the bluestone path towards the river, the wild grass by the roadside had already sprouted tender green shoots. The yellow little flowers of dandelions dotted among them, like stars scattered on the ground. The willows by the river had put out tender green buds, and their soft branches swayed gently in the wind, creating delicate ripples on the water surface. The river water was crystal clear, and a few small fish were swimming among the waterweeds. The sunlight refracted through the water surface, shining with glistening waves. In the distance, in the fields, the rape flowers were in full bloom, and the vast expanse of golden flowers stretched all the way to the horizon. The farmers were bending down to work in the fields, and their figures merged with this golden expanse, forming a vivid picture of spring plowing. Occasionally, an egret would fly up from the fields, and its white wings glistened in the sunlight, adding a bit of vitality to this picture scroll. The air was filled with the fresh smell of grass and soil, mixed with a faint fragrance of flowers. The spring breeze brushed against my face, a bit cool but making me feel extremely comfortable. I took a deep breath, as if trying to inhale the whole spring into my lungs. At this moment, I suddenly understood that spring is not just a season, but more a rhythm of life, a whisper of the awakening earth, and a joyous song of the growth of all things. Standing by the river, looking at this vibrant scene, a warm current welled up in my heart. Spring always comes quietly like this, yet it can make the whole world take on a brand-new look. It tells us in the gentlest way: Life is always full of hope, and beauty will always come as promised."
[0079] 2) After segmenting the sent text T, the resulting word list T n is as follows:
[0080] T n=['Early in the morning',','a ray of warm sunlight slipped into the room through the gap in the curtains,' casting a golden streak of light on the wooden floor. ','I opened the window,' and the moist spring breeze, carrying the fragrance of the soil, blew towards me,' bringing a little coolness, but making people feel refreshed.',' The old peach tree in the yard has been in full bloom since I don't know when. The pink and white petals tremble gently in the morning light, like the blushing cheeks of a young girl. A few early bees shuttle among the flowers, making a cheerful buzzing sound. ','I' noticed','on the thick branches of the peach tree,'a sparrow' was busy picking up dead grass,'it turned out'that'it'had chosen'this'place'to'build'a'nest.','Walking'towards'the'river'along'the'bluestone'path,'the'weeds'on'the'roadside'had'sprout'new'green'buds'.','I'm'not's'not'just'a'split'into'a'place'.','I'm ... The yellow dandelion flowers dotted the scene, like stars scattered on the ground. The willow trees by the river sprouted new green buds, and their soft branches swayed in the wind, creating fine ripples on the water. The river water was crystal clear,
[0081] A few small fish swam among the water plants, and the sunlight shone through the water surface. In the fields in the distance, rape flowers were in full bloom, and the golden sea of flowers stretched all the way to the horizon. Farmers bent over and worked in the fields.
[0082] ', 'Their figures blend in with this expanse of gold, forming a vivid picture of spring plowing. Occasionally, egrets fly up from the fields, and their white wings glisten in the sunlight, adding a touch of vitality to this scroll. The air is filled with the fresh scents of grass and soil, mixed with a faint floral fragrance. The spring breeze caresses the face, bringing a hint of coolness, yet making people feel extremely comfortable. I take a deep breath as if to inhale the entire spring into my lungs. At this moment, I suddenly understand that spring is not just a season, but more of a rhythm of life, a whisper of the awakening earth, and a song of the growth of all things. Standing by the river, looking at this vibrant scene, a warm current surges in my heart. Spring always comes quietly like this, yet can transform the whole world brand-new. It tells us in the gentlest way: Life is always full of hope, and beauty will always arrive as promised.']
[0083] 3) The keyword list K extracted from the word list T n is as follows:
[0084] K=['peach tree', 'new buds', 'life', 'spring breeze', 'soil', 'through', 'embellishment', 'yellow', 'sky', 'slip into', 'labor', 'shining', 'among the flowers', 'painting', 'moist', 'window', 'bending over', 'people', 'blooming', 'sea of flowers', 'whispering', 'growth', 'pink and white', 'all things', 'small fish', 'morning light', 'branches', 'gap', 'curtain', 'room', 'emerge', 'water surface', 'tremble', 'river water', 'engulf', 'yard', 'girl', 'cheek', 'shuttle', 'like', 'branches', 'selected', 'vivid', 'incomparable', 'surging'] ,'always','bees','dry grass','compose','feel','warm current','as promised','early rise','light trace','field','ripples','blooming','green grass','pushing open','breath','small flowers','way','tell','wood floor','rhythm','draw out','earth','unknown','singing','petals','extend','issue','wings','add','reflect','gold','water grass','following the wind','weeds','sparrow','spirit','rapeseed','mix','air','extract','scatter','season','attention','face','figure']
[0085] 4) When the hollowing rate is 0.25, the text after hollowing out as follows:
[0086] = "In the early morning, a ray of warm sunlight slipped into the room through the gap in the curtains, casting a golden streak of light on the wooden floor. I opened the window and the moist spring breeze, carrying the fragrance of the soil, blew in my face. It was cool, but it made me feel refreshed. The peach tree in the yard was already full of flowers. The pink and white petals trembled gently in the morning light, like a girl's blushing cheeks. A few early bees flew among the flowers, emitting K][MASK] buzzing sound[MASK] I noticed[MASK] on the thick[MASK] branch of the peach tree, a sparrow[MASK][MASK] was busy picking up[MASK] dry grass,[MASK][MASK] it chose this place[MASK][MASK]. [MASK][MASK] Walking along the bluestone road[MASK] by the river[MASK], the weeds on the roadside[MASK] have sprouted[MASK][MASK] new buds[MASK], the yellow flowers of dandelions dotted in between,[MASK] like stars scattered on the ground. [MASK][MASK]'s[MASK] [MASK]Tender green buds are sprouting, and soft branches are swaying in the wind, [MASK] leaving fine ripples on the water surface. The river water is so clear that you can see a few small fish swimming among the water plants, [MASK][MASK] reflecting [MASK][MASK] the ripples through the water surface. In the fields in the distance, rape flowers are in full bloom, and the golden sea of flowers stretches all the way to the horizon. Farmers are bending over to work in the fields, and their [MASK] figures blend into this [MASK][MASK] land. [MASK]constitutes[MASK][MASK] a vivid[MASK] picture of spring ploughing.[MASK]Occasionally[MASK] egrets fly[MASK] from the fields,[MASK][MASK] their wings shining in the sun[MASK] adding a bit of agility[MASK] to the picture.[MASK]The air[MASK] is filled with the fresh breath of grass and soil[MASK] mixed with the faint fragrance of flowers[MASK]The spring breeze[MASK][MASK] blows on your face, bringing[MASK] a little coolness[MASK] but making you feel incomparably[MASK][MASK].[MASK]Take a deep breath, [MASK][MASK] and [MASK][MASK] put all the spring into my lungs. [MASK]At this moment, [MASK]I suddenly understand [MASK] that spring is not just a [MASK][MASK] season, [MASK][MASK]it is a kind of [MASK] rhythm of life [MASK]It is the whisper of the awakening of the earth, [MASK]the song of the growth of all things. [MASK]By the river, looking at this vibrant scene [MASK]I feel a [MASK][MASK] warm current [MASK]Spring [MASK][MASK] comes [MASK] quietly [MASK][MASK] like this, [MASK] and [MASK] can [MASK]make the whole [MASK][MASK] look brand new. It tells [MASK][MASK] in the most gentle [MASK] way: life [MASK][MASK] is full of hope, [MASK][MASK] and will always arrive as promised [MASK]".
[0087] Each "[MASK]" in the hollowed-out text represents a character, and "[MASK]" or "[MASK][MASK]" in the hollowed-out text represents an empty position. The size of the hollowed-out text after Huffman coding is 3883 bits, and the size of the original sent text is 4340 bits.
[0088] 5) The recovered text Y obtained by the receiver after decoding (the blank part is caused by the probability of the corresponding word with the highest probability being less than 0.3) is as follows:
[0089] Y="In the early morning, a ray of warm sunshine slipped into the room through the gap in the curtains, casting a golden streak of light on the wooden floor. I pushed open the window, and the moist spring breeze, wrapped in the fragrance of the soil, blew in my face, with a hint of coolness, but it made people feel refreshed. The cherry tree in the yard has been full of flowers without knowing when, and the pink and white petals trembled gently in the morning light, like the blushing cheeks of a girl. A few early bees shuttled among the flowers, making a buzzing sound. I noticed that on the thick branches of the peach tree, a sparrow was busy carrying dry grass, and it chose this place to build a nest. Walking along the bluestone road to the riverside, the wild grass on the roadside has sprouted new buds, and the yellow flowers of dandelions dotted in between, like stars scattered on the ground. The tender green new buds have been pulled out, and the soft branches swayed gently in the wind, drawing fine ripples on the water surface. The river water is clear and you can see the bottom. A few small fish swim among the water plants, reflecting the ripples through the water surface. In the fields in the distance, rape flowers The golden sea of flowers is in full bloom, stretching all the way to the horizon. The farmers are bent over and working in the fields, their figures blending into the fields, forming a vivid picture of spring plowing. Occasionally, egrets fly up from the fields, their wings shining in the sun, adding a bit of agility to the picture. The air is filled with the fresh scent of grass and soil, mixed with a faint scent of flowers. The spring breeze blows across the face, bringing a little coolness, but making people feel extremely comfortable. I took a deep breath and wanted to put the whole spring into my lungs. At this moment, I suddenly realized that spring is not just a season, it is a rhythm of life, a whisper of the awakening of the earth, and a song of joy for the growth of all things. Standing by the river, looking at this vibrant scene, a warm current surged in my heart. Spring is always like this, coming quietly, so that the whole thing can be refreshed. It tells us in the most gentle way: life will always be full of hope, and spring will always come as promised. "
[0090] The advantages of the present invention are further illustrated by experiments below.
[0091] The experimental texts of the present invention are 18 excerpts from famous novels with a length of 600 to 12100 characters and 9 news excerpts from the Internet with a length of 500 to 3800 characters. Each excerpt takes 0.25, 0.3, 0.35 and 0.4 as four different hollowing rates, and 500 experiments are carried out. The reliability and validity data are shown in Table 1, where the number of Huffman coding bits of the hollowed-out text divided by the number of Huffman coding bits of the original sent text is used as the compression rate, and the semantic similarity between the encoding end and the decoding end is calculated by the STS model. The change of hollowing delay with the character length of the text is shown in Table 1. Figure 4 As shown, Figure 4 The horizontal axis unit is the number of characters, the vertical axis unit is seconds, and the fill-in-the-blank text processing speed is about 1000 words per second.
[0092] Table 1
[0093] mask rate compression rate semantic similarity semantic similarity / compression rate 0.25 0.8993 0.9526 1.0592 0.3 0.8694 0.9339 1.0742 0.35 0.8485 0.9073 1.0693 0.4 0.8259 0.8728 1.0569
[0094] When the effectiveness is reflected by the text compression ratio, the reliability is reflected by the semantic similarity between the encoding end and the decoding end, and the timeliness is reflected by the system delay, obviously, the encoding and decoding methods of the present invention have relatively high effectiveness, reliability and timeliness.
[0095] In the description of this specification, the descriptions with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0096] In the specification, the word "including" does not exclude other components or steps, and "a" or "an" does not exclude the case of multiple. Certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good effects.
[0097] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A semantic encoding method based on textrank, characterized in that: include: Get sent text; Using a keyword extraction algorithm to extract keywords from the sent text to obtain a keyword table; According to the keyword table and the preset knockout rate, the non-keywords in the sent text are knocked out to obtain a text after the knockout process; the text after the knockout process is a text obtained by randomly removing some non-keywords in the sent text; Encoding the hollowed-out text to obtain a digital code sequence; After processing the digital code sequence, a digital code sequence to be decoded is obtained and sent.
2. The semantic encoding method based on textrank according to claim 1, characterized in that: The step of performing a hollowing-out process on the non-keywords in the sent text using a preset hollowing-out rate according to the keyword table to randomly remove at least part of the non-keywords in the sent text to obtain a hollowed-out text includes: Performing word segmentation processing on the sent text to obtain a plurality of words arranged in sequence; Determine the total number M of non-keywords to be removed according to the total number of the plurality of words arranged in order and the preset knockout rate; Randomly disrupting the order of the plurality of sequentially arranged words to obtain a plurality of randomly arranged words; According to the arrangement order of the multiple randomly arranged words, determine whether the current word is in the keyword table. If it is a keyword, continue to determine the next word. If it is not a keyword, determine whether the adjacent words of the current word in the sent text are removed. If they are removed, continue to determine the next word. If they are not removed, determine whether the character length of the current word is less than or equal to a preset length. If it is greater than the preset length, continue to determine the next word. If it is less than or equal to the preset length, remove the current word. Afterwards, count and determine whether the total number of words currently removed is less than or equal to M. If it is less than M, continue to determine the next word, until M non-keywords whose character lengths are less than or equal to the preset length and whose positions in the sent text are not adjacent are removed from the multiple randomly arranged words to obtain the text after the hollowing out process.
3. The semantic encoding method based on textrank according to claim 1, characterized in that: The keyword extraction algorithm is used to extract keywords from the sent text to obtain a keyword table, including: Obtaining a quantity range of keywords according to a preset ratio range and a total quantity of words in the sent text; Using the textrank keyword extraction algorithm, the words in the sent text are sorted to obtain a sorting result; the sorting result represents the order of importance of the words in the sent text for understanding the sent text; According to the ranking results, the first N words with the highest importance are all used as keywords to obtain the keyword table; N is any integer in the number range of the keywords.
4. The semantic encoding method based on textrank according to claim 1, characterized in that: The step of encoding the hollowed-out text to obtain a digital code sequence includes: A Huffman encoder is used to perform Huffman source encoding on the hollowed-out text according to a code table to obtain the digital code sequence.
5. The semantic encoding method based on textrank according to claim 1, characterized in that: The step of processing the digital code sequence to obtain a digital code sequence to be decoded and sending the digital code sequence includes: After sequentially modulating, encrypting and channel coding the digital code sequence, a digital code sequence to be decoded is obtained; The digital code sequence to be decoded is sent.
6. A semantic decoding method based on textrank, characterized in that: include: receiving a digital code sequence to be decoded; Processing the digital code sequence to be decoded to obtain a digital code sequence; Performing text decoding on the digital code sequence to obtain a decoded text, wherein the decoded text is a hollowed-out text obtained by randomly removing some non-keywords in the sent text using a preset hollowing-out rate; Non-keywords are filled into the text after the hollowing-out process to restore the semantic information of the sent text and obtain a restored text.
7. The semantic decoding method based on textrank according to claim 6 is characterized in that: The step of filling non-keywords into the text after the hollowing-out process to restore the semantic information of the sent text and obtain the restored text includes: Using the Bert-Base-Chinese model, based on the text after the knockout processing, generate a word corresponding to each empty position in the text after the knockout processing and the probability of each word, wherein each empty position is a position where a non-keyword removed from the sent text is located; A word with the highest probability is selected from the words corresponding to each empty position, and when the probability of the word with the highest probability is greater than or equal to a preset probability threshold, the word with the highest probability is filled into the empty position, and when the probability of the word with the highest probability is less than the preset probability threshold, no word is filled into the empty position, and after processing all the empty positions in the hollowed-out text, the restored text is obtained.
8. The semantic decoding method based on textrank according to claim 6 is characterized in that: The text decoding of the digital code sequence to obtain the decoded text includes: A Huffman decoder is used to perform Huffman source decoding on the digital code sequence according to a code table to obtain the decoded text.
9. A semantic coding device based on TextRank, comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1 to 5 when executing the program stored in the memory.
10. A semantic decoding device based on TextRank, comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 6 to 8 when executing the program stored in the memory.