Information Processing Method, Apparatus, Electronic Device, and Storage Medium
Automatically process text sequences through multi-layer coding models, solving the problem of inefficiency in human work and achieving more efficient and accurate text classification.
Patent Information
- Application Number
- CN202011322555.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-11-23
AI Technical Summary
In the prior art, title compliance testing mainly relies on people's work and is inefficient.
The target text sequence is encoded using a multi-layer encoding model, the encoding results of each encoding layer are obtained, multiple enhanced semantic vectors are fused, and the fused semantic vectors are classified to obtain the normative classification results of the text sequence.
Improve the processing efficiency of text classification and the accuracy of classification results.
Smart Images

Figure CN113761931B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information processing technologies, and in particular, to an information processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In the Internet era, in order to attract users' attention, there often appear some clickbait titles, which are usually false propaganda titles that violate the Advertising Law and are fraudulent. In order to improve the accuracy of information and restrict the behavior of information publishers to create a healthy and good network environment, it is necessary to perform compliance detection on titles.
[0003] In the process of implementing the present invention, the inventors found that currently, the compliance detection of titles mainly relies on the manual operation of platform reviewers, and this method has relatively low efficiency. Summary of the Invention
[0004] Embodiments of the present invention provide an information processing method, apparatus, electronic device, and storage medium, which can improve processing efficiency and the accuracy of classification results.
[0005] In a first aspect, an embodiment of the present invention provides an information processing method, and the method includes:
[0006] Encoding a target text sequence using a multi-layer encoding model, and obtaining the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence;
[0007] Fusing the multiple enhanced semantic vectors to obtain a fused semantic vector;
[0008] Classifying the fused semantic vector to obtain a compliance classification result of the target text sequence.
[0009] In a second aspect, an embodiment of the present invention provides an information processing apparatus, and the apparatus includes:
[0010] An encoding module, configured to encode a target text sequence using a multi-layer encoding model, and obtain the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence;
[0011] A fusion module, configured to fuse the multiple enhanced semantic vectors to obtain a fused semantic vector;
[0012] A classification module, configured to classify the fused semantic vector to obtain a compliance classification result of the target text sequence.
[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the information processing method described in any one of the embodiments of the present invention is implemented.
[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the information processing method described in any one of the embodiments of the present invention is implemented.
[0015] In the embodiments of the present invention, a multi-layer coding model is used to encode a target text sequence, and the encoding results of each encoding layer in the multi-layer coding model are obtained to get multiple enhanced semantic vectors of the target text sequence. Then, the multiple enhanced semantic vectors are fused to obtain a fused semantic vector. Finally, the fused semantic vector is classified to obtain a normative classification result of the target text sequence. That is, in the embodiments of the present invention, the model can be used to automatically classify the target text sequence to obtain the normative classification result of the target text sequence, improving the processing efficiency compared with the manual operation method. In addition, since each encoding layer in the multi-layer coding model represents different semantics of the text sequence, by taking the encoding results of each encoding layer in the multi-layer coding model and fusing the multiple encoding results as the classification basis, the classification basis is made richer and the accuracy of the classification result is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of an information processing method provided by an embodiment of the present invention.
[0017] Figure 2 is another flowchart of an information processing method provided by an embodiment of the present invention.
[0018] Figure 3 is a schematic diagram of a network structure provided by an embodiment of the present invention.
[0019] Figure 4 is a schematic diagram of a structure of an encoding layer provided by an embodiment of the present invention.
[0020] Figure 5 is a schematic diagram of a structure of an information processing device provided by an embodiment of the present invention.
[0021] Figure 6 is a schematic diagram of a structure of an information processing system provided by an embodiment of the present invention.
[0022] Figure 7 is a schematic diagram of a structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0024] Since the existing text classification methods mainly rely on manual operations and have low classification efficiency, embodiments of the present invention provide an information processing method that can achieve automatic classification of texts, improve the classification efficiency, and improve the accuracy of classification results.
[0025] Figure 1 FIG. is a schematic flowchart of an information processing method provided by an embodiment of the present invention. This method can be executed by an information processing device provided by an embodiment of the present invention, and the device can be implemented in a software and / or hardware manner. In a specific embodiment, the device can be integrated in a server. The following embodiments will be described by taking the device integrated in the server as an example. Refer to Figure 1 and the method can specifically include the following steps:
[0026] Step 101, encoding a target text sequence using a multi-layer encoding model, and obtaining the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence.
[0027] In specific implementation, the target text sequence can be the title or summary of an item (such as a commodity), an article, a news, etc. The target text sequence can be uploaded by an item provider (such as a merchant), a content provider (such as a writer) to the server through a terminal, and the server performs a normative classification on the uploaded target text sequence to obtain a normative classification result of the target text sequence.
[0028] In a specific embodiment, the server may encode the target text sequence using a multi-layer encoding model, such as the Bert model. The multi-layer encoding model has multiple encoding layers, such as 4 encoding layers, 12 encoding layers, etc. Generally speaking, for the same text sequence, the semantic representations encoded by each encoding layer of the multi-layer encoding model for this text sequence are different, that is, the encoding results of different encoding layers for this text sequence are different. Generally, the higher encoding layers of the multi-layer encoding model are mainly used to encode the semantic features of the text, that is, the output of the higher encoding layers is the encoding result of the semantic features of the text sequence; the middle encoding layers of the multi-layer encoding model are mainly used to encode the syntactic features of the text, that is, the output of the middle encoding layers is the encoding result of the syntactic features of the text sequence; the lower encoding layers of the multi-layer encoding model are mainly used to encode the surface features such as the word positions of the text, that is, the output of the lower encoding layers is the encoding result of the surface features of the text sequence. Specifically, in the embodiment of the present invention, when encoding the target text sequence using the multi-layer encoding model, the encoding results of each encoding layer in the multi-layer encoding model can be obtained to obtain multiple enhanced semantic vectors of the target text sequence.
[0029] Exemplarily, the multi-layer encoding model can be a pre-trained model. For example, the Bert model can be trained first using the training data of task A, and each parameter of the Bert model is learned through task A, and the learned parameters are saved for later use; when a new task B comes, in the network parameter initialization stage, each parameter of the Bert model learned using the training data of task A can be loaded, and then the training data of task B is input into the Bert model, and the output of the Bert model is input into the classification task model for training the classification task model (such as a fully connected layer and a fully connected layer). During the training of the classification task model, each parameter of the Bert model is fine-tuned through backpropagation, that is, the parameters are better adjusted so that the Bert model is more suitable for the current task B.
[0030] In the embodiment of the present invention, since the encoding result of the target text sequence is to be used for subsequent classification tasks, the target text sequence can be encoded in the following manner:
[0031] (1) Add marker characters to the target text sequence.
[0032] That is, for ease of operation, marker characters (tokens) can be added to the target text sequence. The marker characters such as "CLS" can be added at the start or end position of the target text sequence. For example, if the target text sequence is "world-leading xx", the target text sequence with marker characters added can be "[CLS]world-leading xx".
[0033] (2) Convert the target text sequence with marker characters added into a target vector sequence.
[0034] For example, by querying a character vector table, each character in the target text sequence can be converted into a one-dimensional vector, i.e., a character vector. Combining all the character vectors together, a character vector sequence of the target text sequence is obtained. According to the position of each character in the target text sequence, a position vector of each character is obtained. Combining all the position vectors together, a position vector sequence of the target text sequence is obtained. The character vector sequence and the position vector sequence of the target text sequence are superimposed, and the superimposed vectors are normalized to obtain a target vector sequence, which is used as the input for encoding by a multi-layer encoding model. In a specific implementation, the size of the character vector sequence can be 21128*128, that is, it covers 21128 characters, and the vector dimension of each character is 128. The size of the position vector sequence can be 512*128, that is, it includes 512 position information, and the position vector dimension is 128.
[0035] (3) Encode the target vector sequence using a multi-layer encoding model.
[0036] In a specific implementation, when obtaining the encoding result of each encoding layer in the multi-layer encoding model, the encoding result of each encoding layer in the multi-layer encoding model for the marked character can be obtained, that is, the output vector at the corresponding position of the marked character, so as to obtain multiple enhanced semantic vectors of the target text sequence. That is, in the embodiments of the present invention, the encoding result of each encoding layer for the marked character can be used as the encoding result of each encoding layer for the target text sequence. The output vector at the corresponding position of the marked character can be regarded as gathering the representation of the entire target text sequence. It can be understood that compared with other existing words in the target text sequence, this marked character without obvious semantic information will more "fairly" integrate the semantic information of each word in the text.
[0037] Step 102: Fuse multiple enhanced semantic vectors to obtain a fused semantic vector.
[0038] For example, a fully connected layer (such as fully connected layer one) can be used to fuse multiple enhanced semantic vectors. For example, multiple enhanced semantic vectors can be concatenated to obtain a concatenated semantic vector, and the concatenated semantic vector is input into the fully connected layer one to perform feature fusion compression on the concatenated semantic vector by the fully connected layer one to reduce the vector dimension and facilitate subsequent binary classification. For example, the concatenated semantic vector can be fused and compressed into a 1*128 vector to obtain a fused semantic vector.
[0039] Step 103: Classify the fused semantic vector to obtain a normative classification result of the target text sequence.
[0040] For example, another fully connected layer (fully connected layer two) can be used to classify the fused semantic vector. For example, the fused semantic vector can be input into the fully connected layer two to generate a probability distribution value of the normality degree of the target text sequence by using the fully connected layer two, and it is determined whether the target text sequence is normal according to the probability distribution value.
[0041] Exemplarily, a probability threshold for compliance can be preset in advance, and the probability distribution value output by the fully connected layer two is compared with the probability threshold. If the probability distribution value is greater than the probability threshold, it is determined that the target text sequence is normal; if the probability distribution value is not greater than the probability threshold, it is determined that the target text sequence is not normal.
[0042] After the server obtains the normality classification result of the target text sequence, it can label the target text sequence according to the normality classification result. The text labels include normal and not normal. And so on, all target text sequences uploaded by the terminal can be labeled with text labels to generate a classification list of the target text sequences. When the terminal needs to query whether a certain target text sequence is normal, the server can directly query the classification list to obtain the classification result of the target text sequence.
[0043] In the above technical solution, the target text sequence is encoded by using a multi-layer encoding model, and the encoding results of each encoding layer in the multi-layer encoding model are obtained to get multiple enhanced semantic vectors of the target text sequence. Then, the multiple enhanced semantic vectors are fused to obtain a fused semantic vector. Finally, the fused semantic vector is classified to obtain the normality classification result of the target text sequence. That is, in the embodiments of the present invention, the model can be used to automatically classify the target text sequence to obtain the normality classification result of the target text sequence, which improves the processing efficiency compared with the manual operation method. In addition, since each encoding layer in the multi-layer encoding model represents different semantics of the text sequence, by taking the encoding results of each encoding layer in the multi-layer encoding model and fusing the multiple encoding results as the classification basis, the classification basis is more abundant and the accuracy of the classification result is improved.
[0044] Please refer to Figure 2 , Figure 2 which is another flow diagram of the information processing method according to the embodiments of the present invention. This method can be applied to Figure 3 the network structure shown in Figure 3 In the network structure shown in Figure 2 , after the target text sequence is converted into a target vector sequence, it enters the multi-layer encoding model. The multi-layer encoding model is composed of multiple encoding layers. The output result of each encoding layer goes outwards to the first fully connected layer, and the first fully connected layer performs feature fusion and compression. The data after fusion and compression enters the second fully connected layer, and the second fully connected layer performs classification prediction. As
[0045] Step 201, add marker characters to the target text sequence.
[0046] The target text sequence can be, for example, the title or summary of an item (such as a commodity), an article, a news item, etc. The target text sequence can be uploaded by the item provider (such as a merchant), the content provider (such as a writer) to the server through a terminal. The server performs a normative classification on the uploaded target text sequence to obtain the normative classification result of the target text sequence.
[0047] Specifically, after obtaining the target text sequence, the server can first delete the special characters in the target text sequence. Special characters such as %□, □& and other emoticon characters, and then add marker characters to the target text sequence. Marker characters such as "CLS". The marker characters can be added at the start or end position of the target text sequence. For example, if the target text sequence is "Global leader xx", the target text sequence with marker characters added can be "[CLS] Global leader xx".
[0048] Step 202, convert the target text sequence with marker characters added into a target vector sequence.
[0049] For example, by querying the character vector table, each character in the target text sequence can be converted into a one-dimensional vector, that is, a character vector. Combine all the character vectors to obtain the character vector sequence of the target text sequence; according to the position of each character in the target text sequence, obtain the position vector of each character, and combine all the position vectors to obtain the position vector sequence of the target text sequence; superimpose the character vector sequence and the position vector sequence of the target text sequence, and normalize the superimposed vectors to obtain the target vector sequence E, and use this target vector sequence E as the input for encoding by the multi-layer encoding model. In a specific implementation, the size of the character vector sequence can be 21128*128, that is, covering 21128 characters, the vector dimension of each character is 128, and the size of the position vector sequence can be 512*128, that is, including 512 position information, and the position vector dimension is 128.
[0050] Step 203, encode the target vector sequence using the multi-layer encoding model.
[0051] Specifically, a multi-layer encoding model such as the Bert model has multiple encoding layers, such as 4 encoding layers, 12 encoding layers, etc. For example, the multi-layer encoding model can be a pre-trained model. For example, the Bert model can be first trained using the training data of a preset task, and various parameters of the Bert model can be learned through the training data of the preset task, and the learned various parameters can be saved for later use. When used to implement the classification task of the embodiments of the present invention, in the network parameter initialization stage, various parameters of the Bert model learned using the training data of the preset task can be loaded, the training data of the classification task of the embodiments of the present invention can be input into the Bert model, and the output of each layer of the Bert model can be input into the classification task model of the embodiments of the present invention (such as a fully connected layer 1 and a fully connected layer 2). When training the classification task model, various parameters of the Bert model can be fine-tuned through backpropagation, so that the Bert model is more suitable for the classification task of the embodiments of the present invention, thereby improving the accuracy of the classification result.
[0052] Among them, the training data of the classification task of the embodiments of the present invention can be obtained by the following method: a large number of preset text sequences are obtained, and a text label (such as adding by means of manual annotation) is added to each preset text sequence. The text labels include standard and non-standard, and the preset text sequences with added text labels are used as the training data of the classification task.
[0053] In specific implementation, the training of the multi-layer encoding model and the classification task model can be completed in advance on a server. The training main program can be written in the Python language, and the framework of the model can be built using TensorFlow. After the model training is completed, the trained model can be saved on the disk of the server, and the saved content can include the model structure and the trained parameters in the model. After the model is trained, test data can be used to test the model, and the accuracy evaluation result of the model can be obtained according to the test result and the actual classification result.
[0054] In the embodiments of the present invention, each encoding layer of the multi-layer encoding model can be composed of a Transformer, and the Transformer can be as Figure 4 shown, mainly composed of an attention layer, an intermediate layer, and a feed-forward layer.
[0055] After training the multi-layer encoding model and the classification task model, the target vector sequence can be input into the multi-layer encoding model for encoding. In a Transformer of the multi-layer encoding model, in the attention layer, a Query (abbreviated as Q) vector (i.e., the sequence vector to be queried), a Key (abbreviated as K) vector (i.e., the index vector to be queried), and a Value (abbreviated as V) vector (i.e., the value vector of the character itself) will be created for each character in the target text sequence through linear transformation. The Q vector, K vector, and V vector can be understood as further characterizations of the characters in the target text sequence. For example, the vectors corresponding to the first position in the Q vector, K vector, and V vector can all be regarded as the characterization vectors of the "CLS" token character. In the linear transformation process, the involved formulas are as follows:
[0056] K = relu(EW n,k + b n,k )
[0057] Q = relu(EW n,q + b n,q )
[0058] V = relu(EW n,v + b n,v )
[0059] Among them, n represents the nth Transformer of the multi-layer encoding model, W n,k , b n,k , W n,q , b n,q , W n,v , b n,v are the parameters obtained by model training, W n,k , W n,q , W n,v are the corresponding weight matrices, b n,k , b n,q , b n,v are the corresponding bias vectors.
[0060] After obtaining the Q vector, K vector, and V vector for each character, Q and K T can be dot-producted, which is equivalent to calculating the correlation between pairwise characters, and then normalizing the dot-product result to calculate the correlation matrix between characters Among them, represents the vector dimension; according to the correlation matrix between characters the V vector of the character is converted into an A vector, and the A vector can be understood as a deeper-level characterization that incorporates the correlation between characters or the influence of other characters on this character.
[0061]
[0062] In the middle layer, after extracting features from the output of the attention layer again, the original feature vectors (i.e., the target vector sequence E) are superimposed, and normalization is performed to obtain the output Y of the first sub-layer n,1 , and then through feature transformation to obtain the output Y of the middle layer n,m , and the calculation formula is as follows:
[0063] Y n,1 =Normalization((AW n,1 +b n,1 )+E)
[0064] Y n,m =gelu(Y n,1 W n,m +b n,m )
[0065] Among them, m represents the middle layer of the Transformer, W n,1 , b n,1 , W n,m , b n,m are parameters obtained by model training, W n,1 , W n,m are the corresponding weight matrices, b n,1 , b n,m are the corresponding bias vectors.
[0066] In the forward propagation layer, a residual connection is added to improve the training speed of the model. Among them, (Y n,m W n,2 +b n,2 ) can be understood as the residual. Due to the introduction of the addition term, the vanishing gradient can be effectively suppressed during backpropagation. When solving the parameter variables in the matrix Y n,m , the convergence speed can be improved. After passing through the feed-forward network, the word sequence vector Y n,2 is obtained. The word sequence vector Y n,2 can be understood as the representation vector of the deep text sequence, that is, the enhanced semantic vector. The calculation formula of Y n,2 is as follows:
[0067] Y n,2 =Normalization((Y n,m W n,2 +b n,2 )+Y n,m )
[0068] Among them, W n,2 , b n,2 are parameters obtained by model training, W n,2 is the corresponding weight matrix, b n,2is the corresponding bias vector.
[0069] Step 204: Obtain the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence.
[0070] Generally speaking, for the same text sequence, the semantic representations encoded by each encoding layer of the multi-layer encoding model for this text sequence are different, that is, the encoding results of different encoding layers for this text sequence are different. Generally, the high encoding layers of the multi-layer encoding model are mainly used to encode the semantic features of the text, that is, the output of the high encoding layers is the encoding result of the semantic features of the text sequence; the middle encoding layers of the multi-layer encoding model are mainly used to encode the syntactic features of the text, that is, the output of the middle encoding layers is the encoding result of the syntactic features of the text sequence; the low encoding layers of the multi-layer encoding model are mainly used to encode the surface features such as the word positions of the text, that is, the output of the low encoding layers is the encoding result of the surface features of the text sequence. Specifically, in the embodiments of the present invention, when encoding the target text sequence using the multi-layer encoding model, the encoding results of each encoding layer in the multi-layer encoding model can be obtained to obtain multiple enhanced semantic vectors of the target text sequence.
[0071] Specifically, when marker characters are added to the target text sequence, when obtaining the encoding results of each encoding layer in the multi-layer encoding model, the encoding results of each encoding layer in the multi-layer encoding model for the marker characters can be obtained, that is, the output vectors at the corresponding positions of the marker characters, so as to obtain multiple enhanced semantic vectors of the target text sequence. That is, in the embodiments of the present invention, the encoding results of each encoding layer for the marker characters can be used as the encoding results of each encoding layer for this target text sequence, and the output vectors at the corresponding positions of the marker characters can be regarded as aggregating the representations of the entire target text sequence.
[0072] Step 205: Use the first fully connected layer to fuse the multiple enhanced semantic vectors to obtain a fused semantic vector.
[0073] For example, multiple enhanced semantic vectors can be concatenated to obtain a concatenated semantic vector, and the concatenated semantic vector is input into the first fully connected layer to perform feature fusion and compression on the concatenated semantic vector by the first fully connected layer to reduce the vector dimension and facilitate subsequent binary classification. For example, the concatenated semantic vector can be fused and compressed into a 1*128 vector to obtain a fused semantic vector.
[0074] In a specific embodiment, for example, after concatenating the encoding results of each encoding layer in the multi-layer encoding model for the marker characters to obtain a concatenated semantic vector Z, and inputting the concatenated semantic vector Z into the first fully connected layer for feature fusion and compression to obtain a fused semantic vector Y, then:
[0075] Y = f(W c Z + bc )
[0076] Among them, W c is the weight matrix of the first fully-connected layer, and b c is the bias vector of the first fully-connected layer.
[0077] Step 206: Classify the fused semantic vector using the second fully-connected layer to obtain the normalization classification result of the target text sequence.
[0078] For example, the fused semantic vector can be input into the second fully-connected layer to generate the probability distribution value of the normalization degree of the target text sequence using the second fully-connected layer, and determine whether the target text sequence is normalized according to the probability distribution value.
[0079] For instance, the fused semantic vector Y can be input into the second fully-connected layer, and the output function f uses the sigmoid function to obtain the probability distribution value of the normalization degree of the target text sequence Then:
[0080]
[0081] where W is the weight matrix of the second fully-connected layer, and b is the bias vector of the second fully-connected layer.
[0082] Exemplarily, a probability threshold for compliance can be preset, and the probability distribution value output by the second fully-connected layer is compared with the probability threshold. If the probability distribution value is greater than the probability threshold, it is determined that the target text sequence is normalized; if the probability distribution value is not greater than the probability threshold, it is determined that the target text sequence is not normalized.
[0083] After the server obtains the normalization classification result of the target text sequence, it can label the target text sequence according to the normalization classification result. The text labels include normalized and non-normalized. By analogy, all target text sequences uploaded by the terminal can be labeled with text labels to generate a classification list of target text sequences. When the terminal needs to query whether a certain target text sequence is normalized, the server can directly query the classification list to obtain the classification result of the target text sequence.
[0084] In addition, the trained multi-layer encoding model and classification task model (the first and second fully-connected layers) can be packaged, for example, packaged in a user-defined UDF function in Hive, and the packaged data is sent to each storage node in the preset cluster, so as to realize the distributed execution of the classification task and improve the text classification efficiency.
[0085] In the above technical solution, a multi-layer coding model is used to encode the target text sequence, and the encoding results of each encoding layer in the multi-layer coding model are obtained to get multiple enhanced semantic vectors of the target text sequence. Then, the multiple enhanced semantic vectors are fused to obtain a fused semantic vector. Finally, the fused semantic vector is classified to obtain the normative classification result of the target text sequence. That is, in the embodiments of the present invention, the model can be used to automatically classify the target text sequence to obtain the normative classification result of the target text sequence, which improves the processing efficiency compared with the manual operation method. In addition, since each encoding layer in the multi-layer coding model represents different semantics of the text sequence, by taking the encoding results of each encoding layer in the multi-layer coding model and fusing the multiple encoding results as the classification basis, the classification basis is made richer and the accuracy of the classification result is improved.
[0086] Figure 5 FIG. 0 is a structural diagram of an information processing device provided by an embodiment of the present invention, and the device is applicable to execute the information processing method provided by the embodiment of the present invention. As Figure 5 shown, the device may specifically include:
[0087] An encoding module 501, configured to encode a target text sequence by using a multi-layer coding model, and obtain the encoding results of each encoding layer in the multi-layer coding model to get multiple enhanced semantic vectors of the target text sequence;
[0088] A fusion module 502, configured to fuse the multiple enhanced semantic vectors to obtain a fused semantic vector;
[0089] A classification module 503, configured to classify the fused semantic vector to obtain the classification result of the target text sequence.
[0090] In one embodiment, the encoding module 501 encodes the target text sequence by using a multi-layer coding model, including:
[0091] Adding marker characters to the target text sequence;
[0092] Converting the target text sequence with the added marker characters into a target vector sequence;
[0093] Encoding the target vector sequence by using the multi-layer coding model.
[0094] In one embodiment, the encoding module 501 obtains the encoding results of each encoding layer in the multi-layer coding model to get multiple enhanced semantic vectors of the target text sequence, including:
[0095] Obtaining the encoding results of each encoding layer in the multi-layer coding model for the marker characters to get multiple enhanced semantic vectors of the target text sequence.
[0096] In one embodiment, the fusion module 502 fuses the multiple enhanced semantic vectors to obtain a fused semantic vector, including:
[0097] Using a fully connected layer to fuse the multiple enhanced semantic vectors to obtain the fused semantic vector.
[0098] In one embodiment, the fusion module 502 uses a fully connected layer to fuse the multiple enhanced semantic vectors to obtain the fused semantic vector, including:
[0099] Concatenating the multiple enhanced semantic vectors to obtain a concatenated semantic vector;
[0100] Using the fully connected layer to perform feature fusion and compression on the concatenated semantic vector to obtain the fused semantic vector.
[0101] In one embodiment, the classification module 503 classifies the fused semantic vector to obtain a normative classification result of the target text sequence, including:
[0102] Using a fully connected layer two to classify the fused semantic vector to obtain a normative classification result of the target text sequence.
[0103] In one embodiment, the classification module 503 uses a fully connected layer two to classify the fused semantic vector to obtain a normative classification result of the target text sequence, including:
[0104] Inputting the fused semantic vector into the fully connected layer two to generate a probability distribution value of the normative degree of the target text sequence;
[0105] Determining whether the target text sequence is normative according to the probability distribution value.
[0106] In one embodiment, one or more of the multi-layer encoding model, the fully connected layer one, and the fully connected layer two are trained using preset training data, and the preset training data is obtained according to the following method:
[0107] Obtaining a preset text sequence, and adding a text label to each preset text sequence, where the text label includes normative and non-normative, and using the preset text sequence with the added text label as the preset training data.
[0108] In one embodiment, the apparatus further includes:
[0109] A packaging module, configured to package the multi-layer encoding model, the fully connected layer one, and the fully connected layer two, and send the packaged data to each storage node in a preset cluster.
[0110] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the above-described functional module can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0111] The device according to the embodiment of the present invention encodes the target text sequence by using a multi-layer encoding model, obtains the encoding results of each encoding layer in the multi-layer encoding model, obtains multiple enhanced semantic vectors of the target text sequence, then fuses the multiple enhanced semantic vectors to obtain a fused semantic vector, and finally classifies the fused semantic vector to obtain a normative classification result of the target text sequence. That is, in the embodiment of the present invention, the model can be used to automatically classify the target text sequence to obtain a normative classification result of the target text sequence, which improves the processing efficiency compared with the manual operation method; in addition, since each encoding layer in the multi-layer encoding model represents different semantics of the text sequence, by taking the encoding results of each encoding layer in the multi-layer encoding model and fusing the multiple encoding results as the classification basis, the classification basis is made more abundant and the accuracy of the classification result is improved.
[0112] The embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the information processing method provided in any of the above embodiments.
[0113] The embodiment of the present invention also provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, it implements the information processing method provided in any of the above embodiments.
[0114] Figure 6 An exemplary system architecture 600 to which the information processing method or information processing device according to the embodiment of the present invention can be applied is shown.
[0115] As Figure 6 shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used to provide a medium for a communication link between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0116] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various client applications can be installed on terminal devices 601, 602, and 603, such as enterprise application clients, e-commerce application clients, web browser applications, search applications, instant messaging tools, and email clients, etc.
[0117] Terminal devices 601, 602, and 603 can be various electronic devices with a display screen and supporting various clients, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.
[0118] Server 605 can be a server that provides various services, such as a background management server that supports query requests for websites sent by users using terminal devices 601, 602, and 603. The background management server can process the received query requests and feedback the processing results to the terminal devices.
[0119] It should be noted that the information processing method provided by the embodiments of the present invention is generally executed by server 605. Correspondingly, the information processing device is generally arranged in server 605.
[0120] It should be understood, Figure 6 The numbers of terminal devices, networks, and servers in
[0121] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 7 Shown below is a schematic structural diagram of a computer system 700 of an electronic device suitable for implementing the embodiments of the present invention with reference to Figure 7 The shown electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0122] As Figure 7 shown, computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0123] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 710 as required so that a computer program read therefrom is installed into the storage section 708 as required.
[0124] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, the above functions defined in the system of the present invention are executed.
[0125] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0127] The modules and / or units involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes an encoding module, a fusion module, and a classification module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases.
[0128] As another aspect, the present invention further provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist independently without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: encoding a target text sequence using a multi-layer encoding model, and obtaining the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence; fusing the multiple enhanced semantic vectors to obtain a fused semantic vector; classifying the fused semantic vector to obtain a normative classification result of the target text sequence.
[0129] According to the technical solution of the embodiments of the present invention, a target text sequence can be encoded using a multi-layer encoding model, and the encoding results of each encoding layer in the multi-layer encoding model can be obtained to obtain multiple enhanced semantic vectors of the target text sequence. Then, the multiple enhanced semantic vectors can be fused to obtain a fused semantic vector. Finally, the fused semantic vector can be classified to obtain a normative classification result of the target text sequence. That is, in the embodiments of the present invention, the model can be used to automatically classify the target text sequence to obtain a normative classification result of the target text sequence, which improves the processing efficiency compared with the manual operation method. In addition, since each encoding layer in the multi-layer encoding model represents different semantics of the text sequence, by taking the encoding results of each encoding layer in the multi-layer encoding model and fusing the multiple encoding results as the classification basis, the classification basis is made more abundant, and the accuracy of the classification result is improved.
[0130] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An information processing method, characterized in that, Including: Encoding a target text sequence using a multi-layer encoding model, including: adding a marker character to the target text sequence, where the marker character is added at the starting position of the target text sequence; converting the target text sequence with the added marker character into a target vector sequence; encoding the target vector sequence using the multi-layer encoding model; wherein the multi-layer encoding model includes multiple encoding layers, and each encoding layer of the multi-layer encoding model represents different semantics for encoding the target text sequence; Obtaining the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence, including: obtaining the encoding results of each encoding layer in the multi-layer encoding model for the marker character to obtain multiple enhanced semantic vectors of the target text sequence; Wherein, each encoding layer of the multi-layer encoding model is composed of an attention layer, an intermediate layer, and a forward propagation layer; the attention layer is used to create a query sequence vector, a query index vector, and a value vector of the character itself for each character in the target text sequence through linear transformation, wherein the vectors corresponding to the first position in the query sequence vector, the query index vector, and the value vector of the character itself are all representation vectors of the marker character; the intermediate layer is used to perform feature extraction on the result output by the attention layer, then stack the target vector sequence and perform normalization processing; the forward propagation layer is used to add a residual connection; Fusing the multiple enhanced semantic vectors to obtain a fused semantic vector, including: using a first fully connected layer to fuse the multiple enhanced semantic vectors to obtain the fused semantic vector; Classifying the fused semantic vector to obtain a normative classification result of the target text sequence, including: using a second fully connected layer to classify the fused semantic vector to obtain a normative classification result of the target text sequence.
2. The information processing method according to claim 1, wherein The using the first fully connected layer to fuse the multiple enhanced semantic vectors to obtain the fused semantic vector includes: Concatenating the multiple enhanced semantic vectors to obtain a concatenated semantic vector; Using the first fully connected layer to perform feature fusion compression on the concatenated semantic vector to obtain the fused semantic vector.
3. The information processing method according to claim 1, characterized in that The using the second fully connected layer to classify the fused semantic vector to obtain a normative classification result of the target text sequence includes: Inputting the fused semantic vector into the second fully connected layer to generate a probability distribution value of the degree of normality of the target text sequence; Determining whether the target text sequence is normative according to the probability distribution value.
4. The information processing method according to any one of claims 2 to 3, characterized in that One or more of the multi-layer encoding model, the first fully connected layer, and the second fully connected layer are trained using preset training data, and the preset training data is obtained according to the following method: Obtaining a preset text sequence, and adding a text label to each preset text sequence, where the text label includes normative and non-normative, and using the preset text sequence with the added text label as the preset training data.
5. The information processing method according to claim 4, characterized in that, The method further includes: Packing the multi-layer encoding model, the first fully connected layer, and the second fully connected layer, and sending the packed data to each storage node in a preset cluster.
6. An information processing apparatus, characterized in that, Including: An encoding module, configured to encode a target text sequence by using a multi-layer encoding model, where the multi-layer encoding model includes multiple encoding layers, and the semantic representations encoded by each encoding layer of the multi-layer encoding model for the target text sequence are different; obtain the encoding results of each encoding layer in the multi-layer encoding model to obtain multiple enhanced semantic vectors of the target text sequence; A fusion module, configured to fuse the multiple enhanced semantic vectors to obtain a fused semantic vector; A classification module, configured to classify the fused semantic vector to obtain a normative classification result of the target text sequence; Wherein, the encoding module is specifically configured to: Add a marker character to the target text sequence, and the marker character is added at the starting position of the target text sequence; Convert the target text sequence with the added marker character into a target vector sequence; Encode the target vector sequence by using the multi-layer encoding model; Obtain the encoding results of each encoding layer in the multi-layer encoding model for the marker character to obtain multiple enhanced semantic vectors of the target text sequence; Wherein, each encoding layer of the multi-layer encoding model is composed of an attention layer, an intermediate layer, and a forward propagation layer; the attention layer is configured to create a query sequence vector, a query index vector, and a value vector of the character itself for each character in the target text sequence through a linear transformation, where the vectors corresponding to the first position in the query sequence vector, the query index vector, and the value vector of the character itself are all representation vectors of the marker character; the intermediate layer is configured to perform feature extraction on the result output by the attention layer, then stack the target vector sequence and perform a normalization process; the forward propagation layer is configured to add a residual connection; Wherein, the fusion module is specifically configured to: Fuse the multiple enhanced semantic vectors by using a first fully-connected layer to obtain the fused semantic vector; The classification module is specifically configured to: Classify the fused semantic vector by using a second fully-connected layer to obtain a normative classification result of the target text sequence.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the information processing method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the information processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for training business model and determining text classification categories
CN111737474A
Long text cascade classification method, system and device and storage medium
CN111930952A