A Chinese text semantic compression method based on deep learning

By combining the self-attention mechanism with Bi-LSTM network, the traditional recurrent neural network model is improved, and the problems of long training time and poor semantic compression effect of traditional models are solved, achieving more efficient semantic compression effect and bandwidth resource saving.

CN114925701BActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210608472.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-05-06
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

Traditional text semantic compression models based on recurrent neural networks or convolutional neural networks have too long training time and have poor semantic compression effects, which cannot meet the requirements of semantic communication.

Method used

The semantic compression method of Chinese text based on deep learning is adopted to combine the self-attention mechanism with the Bi-LSTM network to improve the traditional recurrent neural network model, reduce the model training time and improve the semantic compression effect.

Benefits of technology

It shortens the model training time and improves the semantic compression effect, which can effectively save bandwidth resources in wireless network communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925701B_ABST
    Figure CN114925701B_ABST
Patent Text Reader

Abstract

A Chinese text semantic compression method based on deep learning can compress the semantics of an input Chinese text to the greatest extent through a system model. The invention combines the advantages of bidirectional long short-term memory network (Bi-LSTM) and self-attention mechanism (Self-Attention), greatly improving the text semantic compression effect at the sending end in the wireless communication network, effectively saving the bandwidth resources required for wireless communication transmission, thereby further improving the information processing efficiency at the receiving end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and communication, and in particular to a Chinese text semantic compression method based on deep learning. Background Art

[0002] With the vigorous development of mobile wireless communication technology in recent years, the amount of data has exploded. How to transmit a large amount of text data in the network under limited bandwidth resources is an urgent problem to be solved. Some people use text semantic compression technology to compress long texts, retain the semantics, save bandwidth resources to the maximum extent, and improve the processing efficiency of information at the receiving end. However, the traditional semantic compression model based on recurrent neural networks (RNN) or convolutional neural networks (CNN) has poor effect and cannot meet the requirements of semantic communication. In recent years, some scholars have proposed to perform text semantic compression based on the attention mechanism, which has improved the semantic compression effect compared with recurrent neural networks or convolutional neural networks. Summary of the invention

[0003] In order to solve the shortcomings of traditional semantic compression methods, such as long model training time and poor semantic compression effect, the present invention proposes a Chinese text semantic compression method based on deep learning with short model training time and good semantic compression effect. The method combines the self-attention mechanism (Self-attention) and the Bi-LSTM (Bi-directional Long Short-Term Memory) network to improve the traditional semantic compression model based on recurrent neural network, which can compress semantics to the greatest extent and can be well applied in the fields of control and communication.

[0004] The technical solution adopted by the present invention to solve its technical problem is:

[0005] A Chinese text semantic compression method based on deep learning, comprising the following steps:

[0006] 1) First, preprocess the input text as follows: normalize the sentence s to be transmitted to have n words, where the parameter n can be set by yourself; then use the Jieba Chinese word segmentation tool to remove stop words and perform word segmentation to obtain w1, w2, w3, …, w n , and then use the Word2Vec Chinese pre-training model to output each word w1,w2,w3,…,w n The corresponding word vector is represented by c1, c2, c3, …, c n Indicates that the word vector group c1, c2, c3, ..., cn Note it as C;

[0007] 2) Group the word vectors c1, c2, c3, ..., c n Input to the encoder, the encoder has the same two layers. In the first layer of the encoder, the word vector group first enters the self-attention mechanism. The calculation process is as follows:

[0008] q i =W q ×C,i∈[1,n] (1)

[0009] k i =W k ×C (2)

[0010] v i =W v ×C (3)W q ,W k ,W v is a trainable parameter matrix with a dimension of 256;

[0011] 3) For each q i (i∈[1,n]), let it be the same as every k i (i∈[1,n]) performs vector dot multiplication, for q 1 We get α 1,1 ,α 1,2 ,α 1,3 ,…,α 1,n , α 1,1 ,α 1,2 ,α 1,3 ,…,α 1,n Perform Softmax normalization operation and get in:

[0012]

[0013] Then Respectively with their corresponding v 1 ,v 2 ,v 3 ,…,v n Multiply and add the results to get vector a 1 ; Perform the above operation n times to get vector a 1 ,a 2 ,a 3 ,…,a n , the formula is as follows:

[0014]

[0015] At this point, the first self-attention mechanism operation is completed; the vector generated by the self-attention mechanism operation is called the attention vector, such as a 1 ,a 2 ,a 3 ,…,a n ;

[0016] 4) The attention vector a 1 ,a 2 ,a 3 ,…,a n Input the bidirectional long short-term memory neural network Bi-LSTM layer respectively, and obtain vector b 1 ,b 2 ,b 3 ,…,b n , dimension and a 1 ,a 2 ,a 3 ,…,a n same;

[0017] 5) Vector b 1 ,b 2 ,b 3 ,…,b n Entering the second layer of the encoder, the self-attention operation in the first layer is repeated first, and the output attention vector is then passed through the bidirectional long short-term memory neural network output vector group e 1 ,e 2 ,e 3 ,…,e n , e 1 ,e 2 ,e 3 ,…,e n Multiply by the trainable parameter matrix with dimension 256 Get vectors respectively

[0018] 6) Enter the decoder part. The decoder has two layers. In the first layer, an initial word vector with a dimension of 256 is first generated. <cls>Input to the decoder to start decoding operation;

[0019] 7) The word vector of the first target word is used as the input of the decoder for the second decoding. Similarly, the word vector of the first target word is multiplied by a square matrix with a dimension of 256. Get the corresponding vector n q ,n k ,n v Reserved for follow-up;

[0020] 8) The second target word is used as the input of the decoder for the third decoding, and the above decoding steps are repeated until all target words are output, thereby obtaining the predicted semantics.

[0021] 9) The model parameters are trained by minimizing the negative logarithmic loss function. The model parameters include matrix elements and neural network weights.

[0022] Further, in step 6), the initial word vector <cls>The self-attention mechanism is operated, and the resulting attention vector is recorded as m. The next step is to enter the Decoder-Encoder Attention layer to operate the attention mechanism. The process is as follows: multiply the vector m by a square matrix with a dimension of 256 Get vector q m , the vector q m Respectively with vector Perform a dot multiplication operation to obtain The formula is as follows:

[0023]

[0024] in, is the vector e i With square Multiply the resulting vector;

[0025] q m is the vector m and the square matrix Multiply the resulting vector;

[0026] right Perform Softmax normalization operation to obtain Then Corresponding to each Multiply and add the results to get the attention vector r 1 , vector r 1 Then pass through the feedforward neural network FFNN layer to get the vector vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The dimension with the largest probability value corresponds to the first target word.

[0027] Furthermore, in step 7), the second decoding operation is described as follows: vector n q ,n k ,n v Perform a self-attention mechanism operation on the initial word vector to obtain the attention vector h corresponding to the first target word vector, and multiply h by the trainable parameter matrix with a dimension of 256 Get vector q h , the vector q h Respectively with vector Perform a dot multiplication operation to obtain γ i (i∈[1,n]), the formula is as follows:

[0028]

[0029] in, is the vector e i With square Multiply the resulting vector;

[0030] q h is the attention vector h and the square matrix Multiply the resulting vector;

[0031] Again i (i∈[1,n]) performs Softmax normalization operation to obtain Will Corresponding to each Multiply and add the results to get the attention vector r 2 , vector r 2 After the feed-forward neural network FFNN layer, the vector vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The one with the highest probability corresponds to the second target word.

[0032] Furthermore, in step 9), the loss function is defined as:

[0033]

[0034] in, To generate the words in the standard semantic sentence for the decoder at time t The probability of , T is the total time required for the decoder to generate a semantic sentence.

[0035] The beneficial effects of the present invention are: integrating the bidirectional long short-term memory neural network with the self-attention mechanism, improving the traditional recurrent neural network model to achieve better semantic compression effect, thereby effectively saving bandwidth resources in wireless network communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the Chinese text semantic compression system model based on deep learning. It is mainly composed of two parts: encoder and decoder. The encoder includes the following parts: Self-Attention mechanism and bidirectional long short-term memory network (Bi-LSTM); the decoder includes the following parts: Self-Attention mechanism, Decoder-Encoder Attention mechanism, and feed-forward neural network (FFNN). DETAILED DESCRIPTION

[0037] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0038] Reference Figure 1 , a Chinese text semantic compression method based on deep learning, which can compress the semantics of the target text to the greatest extent. The present invention can be applied to the fields of control and communication, such as Figure 1 As shown, the text semantic compression method for this scenario includes the following steps:

[0039] 1) First, preprocess the input text as follows: normalize the sentence s to be transmitted to have n words, where the parameter n can be set by yourself; then use the Jieba Chinese word segmentation toolkit to remove stop words and perform word segmentation to obtain w1, w2, w3, …, w n , and then use the Word2Vec Chinese pre-training model to output each word w1,w2,w3,…,w n The corresponding word vector is represented by c1, c2, c3, …, c n Indicates that the word vector group c1, c2, c3, ..., c n Note it as C;

[0040] 2) Group the word vectors c1, c2, c3, ..., c n Input to the encoder, the encoder has the same two layers. In the first layer of the encoder, the word vector group first enters the self-attention mechanism, and the calculation process is as follows:

[0041] q i =W q ×C,i∈[1,n] (1)

[0042] k i =W k ×C (2)

[0043] v i =W v ×C (3)W q ,W k ,W v is a trainable parameter matrix with a dimension of 256;

[0044] 3) For each q i (i∈[1,n]), let it be the same as every k i (i∈[1,n]) performs vector dot multiplication, for q 1 We get α 1,1 ,α 1,2 ,α 1,3 ,…,α 1,n , for α 1,1 ,α 1,2 ,α 1,3 ,…,α 1,n Perform Softmax normalization operation and get in:

[0045]

[0046] Then Respectively with their corresponding v 1 ,v 2 ,v 3 ,…,v n Multiply and add the results to get vector a 1 ; Perform the above operation n times to get vector a 1 ,a 2 ,a 3 ,…,a n , the formula is as follows:

[0047]

[0048] At this point, the first self-attention mechanism operation is completed; the vector generated by the self-attention mechanism operation is called the attention vector, such as a 1 ,a 2 ,a 3 ,…,a n ;

[0049] 4) The attention vector a 1 ,a 2 ,a 3 ,…,a n Input the bidirectional long short-term memory neural network Bi-LSTM layer respectively, and obtain vector b 1 ,b 2 ,b 3 ,…,b n , dimension and a 1 ,a 2 ,a 3 ,…,a n same;

[0050] 5) Vector b 1 ,b 2 ,b 3 ,…,b n Entering the second layer of the encoder, the self-attention operation in the first layer is repeated first, and the output attention vector is then passed through the bidirectional long short-term memory neural network output vector group e 1 ,e 2 ,e 3 ,…,e n , e 1 ,e 2 ,e 3 ,…,e n Multiply by the trainable parameter matrix with dimension 256 Get vectors respectively

[0051] 6) Enter the decoder part. The decoder is also divided into two layers. In the first layer, an initial word vector with a dimension of 256 is first <cls>Input to the decoder to start decoding operation. The calculation process is as follows: Initial word vector <cls>The self-attention mechanism operation is performed as described in step 3), and the resulting attention vector is recorded as m. Next, the Decoder-Encoder Attention operation is performed by multiplying the vector m by a square matrix with a dimension of 256 Get vector q m ; The vector q m Respectively with vector Perform vector dot multiplication to get The formula is as follows:

[0052]

[0053] in, is the vector e i With square The multiplication vector, q m is the vector m and the square matrix Multiply the resulting vector;

[0054] right Perform Softmax normalization operation to obtain Then Corresponding to each Multiply and add the results to get the thought vector r 1 , vector r 1 Then the vector is obtained through the feedforward neural network FFNN vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The highest probability corresponds to the first target word.

[0055] 7) The word vector of the first target word is used as the input of the decoder for the second decoding. Similarly, the word vector of the first target word is multiplied by the trainable parameter matrix with a dimension of 256. Get the corresponding vector n q ,n k ,n v Reserved for subsequent operations; the second decoding operation is as follows: vector n q ,n k ,n v Perform a self-attention mechanism operation with the initial word vector to obtain the thought vector h corresponding to the word vector of the first target word, and multiply h by a square matrix with a dimension of 256 Get vector q h , the vector q h Respectively with vector Perform a dot multiplication operation to obtain γ i (i∈[1,n]), the formula is as follows:

[0056] in is the vector u i With square The multiplication vector, q h is the thought vector h and the square matrix Multiply the resulting vector;

[0057] Again i (i∈[1,n]) performs Softmax normalization operation to obtain Will Corresponding to each Multiply and add the results to get the thought vector r 2 , vector r 2 After the feed-forward neural network FFNN, the vector vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The highest probability corresponds to the second target word.

[0058] 8) The second target word is used as the input of the decoder for the third decoding. The above decoding steps are repeated until all target words are output, thus obtaining the predicted sentence.

[0059] 9) The model parameters can be trained by minimizing the loss function, which includes matrix elements and neural network weights. The loss function is defined as:

[0060]

[0061] in, To generate the words in the standard semantic sentence for the decoder at time t The probability of , T is the total time required for the decoder to generate a semantic sentence.

[0062] The solution of this embodiment integrates the bidirectional long short-term memory neural network with the self-attention mechanism, and improves the traditional recurrent neural network model to achieve a better semantic compression effect, thereby effectively saving bandwidth resources in wireless network communications.

[0063] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.< / cls> < / cls> < / cls> < / cls>

Claims

1. A Chinese text semantic compression method based on deep learning, characterized in that: The method comprises the following steps: 1) First, preprocess the input text as follows: normalize the sentence s to be transmitted to have n words, where the parameter n can be set by yourself; then use the Jieba Chinese word segmentation tool to remove stop words and perform word segmentation to obtain w1, w2, w3, ···, w n , and then use the Word2Vec Chinese pre-training model to output each word w1,w2,w3,···,w n The corresponding word vectors are represented by c1, c2, c3, ···, c n Indicates that the word vector group c1,c2,c3,···,c n Note it as C; 2) Group the word vectors c1, c2, c3, ···, c n Input to the encoder, the encoder has the same two layers. In the first layer of the encoder, the word vector group first enters the self-attention mechanism. The calculation process is as follows: q i =W q ×C, i∈[1,n] (1) k i =W k ×C (2) v i =W v ×C (3) Among them, W q ,W k ,W v is a trainable parameter matrix with a dimension of 256; 3) For each q i , i∈[1,n], let it be the same as every k i Perform vector dot multiplication, i∈[1,n], for q 1 We get α 1,1 ,α 1,2 ,α 1,3 ,···,α 1,n , α 1,1 ,α 1,2 ,α 1,3 ,···,α 1,n Perform Softmax normalization operation and get in: Then Respectively with their corresponding v 1 ,v 2 ,v 3 ,···,v n Multiply and add the results to get vector a 1 ; Perform the above operation n times to get vector a 1 ,a 2 ,a 3 ,···,a n , the formula is as follows: At this point, the first self-attention mechanism operation is completed; the vector generated by the self-attention mechanism operation is called the attention vector, that is, a 1 ,a 2 ,a 3 ,···,a n ; 4) The attention vector a 1 ,a 2 ,a 3 ,···,a n Input the bidirectional long short-term memory neural network Bi-LSTM layer respectively, and obtain vector b 1 ,b 2 ,b 3 ,···,b n , dimension and a 1 ,a 2 ,a 3 ,···,a n same; 5) Vector b 1 ,b 2 ,b 3 ,···,b n Entering the second layer of the encoder, the self-attention operation in the first layer is repeated first, and the output attention vector is then passed through the bidirectional long short-term memory neural network output vector group e 1 ,e 2 ,e 3 ,···,e n , e 1 ,e 2 ,e 3 ,···,e n Multiply by the trainable parameter matrix with dimension 256 Get vectors respectively i∈[1,n]; 6) Enter the decoder part. The decoder has two layers. In the first layer, an initial word vector with a dimension of 256 is first generated. <cls> Input to the decoder to start decoding operation;< / cls> 7) The word vector of the first target word is used as the input of the decoder for the second decoding. Similarly, the word vector of the first target word is multiplied by a square matrix with a dimension of 256. Get the corresponding vector n q ,n k ,n v Reserved for follow-up; 8) The second target word is used as the input of the decoder for the third decoding, and the above decoding steps are repeated until all target words are output, thereby obtaining the predicted semantics. 9) Train the model parameters by minimizing the negative logarithmic loss function, where the model parameters include matrix elements and neural network weights; In step 6), the initial word vector <cls>The self-attention mechanism is operated, and the resulting attention vector is recorded as m. The next step is to enter the Decoder-Encoder Attention layer to operate the attention mechanism. The process is as follows: multiply the vector m by a square matrix with a dimension of 256 Get vector q m , the vector q m Respectively with vector Perform a dot product operation, i∈[1,n], and get i∈[1,n], the formula is as follows:< / cls> in, is the vector e i With square The multiplication vector, q m is the vector m and the square matrix Multiply the resulting vector; right Perform Softmax normalization operation to obtain i∈[1,n], and then Corresponding to each Multiply, i∈[1,n], and add the results to get the attention vector r 1 , vector r 1 Then pass through the feedforward neural network FFNN layer to get the vector vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The dimension with the largest probability value corresponds to the first target word.

2. A Chinese text semantic compression method based on deep learning as claimed in claim 1, characterized in that: In step 7), the second decoding operation is described as follows: q ,n k ,n v Perform a self-attention mechanism operation on the initial word vector to obtain the attention vector h corresponding to the first target word vector, and multiply h by the trainable parameter matrix with a dimension of 256 Get vector q h , the vector q h Respectively with vector Perform a dot multiplication operation, i∈[1,n], and get γ i , i∈[1,n], the formula is as follows: in, is the vector e i With square The multiplication vector, q h is the attention vector h and the square matrix Multiply the resulting vector; Again i Perform Softmax normalization operation to obtain i∈[1,n], Corresponding to each Multiply, i∈[1,n], and add the results to get the attention vector r 2 , vector r 2 After the feed-forward neural network FFNN layer, the vector vector Enter the second layer of the decoder, repeat the operation of the first layer in the second layer, and finally output the probability vector through the Softmax layer. The one with the highest probability corresponds to the second target word.

3. The Chinese text semantic compression method based on deep learning as claimed in claim 1, characterized in that: In step 9), the loss function is defined as: in, To generate the words in the standard semantic sentence for the decoder at time t The probability of , T is the total time required for the decoder to generate a semantic sentence.