A text processing method, apparatus, device, and storage medium

CN114266226BActive Publication Date: 2025-07-29ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111642879.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-29
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

[0004]以上操作过程中,在针对每一条文本序列进行编码时,只会依据当前文本序列包含的文本数据进行编码,其感受野(参考卷积处理中的感受野的概念)比较局限,编码效果并不理想,进而影响到待处理文本的分类效果

Benefits of technology

[0016] In the foregoing solution, first, the text to be processed is segmented into multiple text sequences, then each text sequence is encoded, and then the encoding results of the multiple text sequences are encoded to obtain an encoding result representing the semantics of the text to be processed, and finally, the text to be processed is classified based on the encoding result. Thus, the classification of long texts can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266226B_ABST
    Figure CN114266226B_ABST
Patent Text Reader

Abstract

The present application provides a text processing method, apparatus, device, and storage medium. The method may include: performing a segmentation operation on the text to be processed to obtain N text sequences; for each of the N text sequences, encoding the text sequence based on at least part of the text data in the text sequences adjacent to the text sequence before and after it to obtain the encoded text sequence; encoding the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, and determining the text type of the text to be processed according to the encoding result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and particularly to a text processing method, apparatus, device, and storage medium. Background Art

[0002] Major media platforms generate positive or negative public opinions related to companies. These public opinions have a great impact on companies. It is necessary to perceive and accurately distinguish the types of these public opinions.

[0003] Currently, public opinions may be embodied in text form. In related technologies, generally, the text to be processed is first split into several text sequences, and then each text sequence is encoded (encoding can be understood as feature extraction) to obtain the encoded text sequence. Then, each encoded text sequence is summarized to obtain the encoding result of the text to be processed, and the text to be processed is classified based on the encoding result of the text to be processed.

[0004] In the above operation process, when encoding each text sequence, only the text data included in the current text sequence is used for encoding, and its receptive field (referring to the concept of receptive field in convolutional processing) is relatively limited, and the encoding effect is not ideal, which in turn affects the classification effect of the text to be processed. Summary of the Invention

[0005] In view of this, this application discloses at least a text processing method. The method may include: performing a splitting operation on the text to be processed to obtain N text sequences; for each of the N text sequences, encoding the text sequence based on at least part of the text data in the text sequences adjacent to the text sequence before and after it to obtain the encoded text sequence; encoding the N encoded text sequences to obtain the encoding result corresponding to the text to be processed, and determining the text type of the text to be processed according to the encoding result.

[0006] In some embodiments, for each of the N text sequences, combining at least part of the text data in the text sequences adjacent to the text sequence before and after it to encode the text sequence to obtain the encoded text sequence, includes: performing multiple rounds of encoding operations on the N text sequences to obtain the N encoded text sequences; where each round of encoding operation is as follows: for each of the N text sequences obtained from the previous encoding operation, combining the text sequence with at least part of the text data in the text sequences adjacent to the text sequence before and after it to obtain a combined text sequence, encoding the combined text sequence to obtain the encoded combined text sequence, and deleting the encoded data corresponding to the at least part of the text data in the encoded combined text sequence to obtain the text sequence after the current encoding operation.

[0007] In some embodiments, the length of the text sequence is a first preset text length; for each of the N text sequences obtained from the previous encoding operation, combining the text sequence with at least partial text data in the text sequences adjacent to it before and after to obtain a combined text sequence, including: starting from the first text sequence of the N text sequences, sliding a preset window in the N text sequences with the first preset text length as the step size, and determining the segment contained within the preset window after each slide as the combined text sequence; the window size of the preset window is a second preset text length; the second preset text length is the sum of the first preset text length and the data length of the at least partial text data.

[0008] In some embodiments, the text sequence contains a preset character indicating the semantic information of the text sequence; encoding the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, including: summarizing the preset characters contained in each of the N text sequences to obtain a character sequence corresponding to the text to be processed; encoding the character sequence to obtain an encoding result corresponding to the text to be processed.

[0009] In some embodiments, for each of the N text sequences, encoding the text sequence based on at least partial text data in the text sequences adjacent to it before and after to obtain the encoded text sequence, including: based on a preset first encoding unit, for each of the N text sequences, combining at least partial text data in the text sequences adjacent to it before and after to encode the text sequence to obtain the encoded text sequence; encoding the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, including: based on a preset second encoding unit, encoding the N encoded text sequences to obtain an encoding result corresponding to the text to be processed.

[0010] In some embodiments, the first encoding unit and the second encoding unit include a BERT model, and the BERT model includes at least one Transformer layer; encoding the combined text sequence, including: using the Transformer layer corresponding to the current encoding operation included in the first encoding unit to encode the combined text sequence; encoding the character sequence, including: using at least one Transformer layer included in the second encoding unit to encode the character sequence.

[0011] In some embodiments, the BERT model is a model pre-trained through a text training sample set.

[0012] In some embodiments, the input length of the BERT model is a third preset text length; the third preset text length is greater than the second preset text length; the encoding of the combined text sequence using the Transformer layer corresponding to the current encoding operation included in the first encoding unit includes: performing a character completion operation on the combined text sequence to obtain a first input sequence with the third preset text length; inputting the first input sequence into the Transformer layer corresponding to the current encoding operation included in the first encoding unit for encoding; the encoding of the character sequence using at least one Transformer layer included in the second encoding unit includes: performing a character completion operation on the character sequence to obtain a second input sequence with the third preset text length; inputting the second input sequence into at least one Transformer layer included in the second encoding unit for encoding.

[0013] The present application also provides a text processing device, including: a segmentation module that performs a segmentation operation on the text to be processed to obtain N text sequences; a first encoding module that, for each of the N text sequences, encodes the text sequence based on at least part of the text data in the text sequences adjacent to the text sequence before and after it, to obtain the encoded text sequence; a second encoding and classification module that encodes the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, and determines the text type of the text to be processed according to the encoding result.

[0014] The present application also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the text processing method shown in any of the foregoing embodiments.

[0015] The present application also provides a computer-readable storage medium storing a computer program for causing a processor to execute the text processing method shown in any of the foregoing embodiments.

[0016] In the foregoing solution, first, the text to be processed is segmented into multiple text sequences, then each text sequence is encoded, and then the encoding results of the multiple text sequences are encoded to obtain an encoding result representing the semantics of the text to be processed, and finally, the text to be processed is classified based on the encoding result. Thus, the classification of long texts can be achieved.

[0017] Second, during the process of encoding each text sequence of the text to be processed, at least part of the text data in the adjacent text sequences before and after the text sequence can be combined to encode the text sequence, and the encoded text sequence can be obtained. In this way, it is equivalent to expanding the receptive field when encoding the text sequence. Compared with the related technology, a more accurate text sequence encoding effect can be obtained, and then a better text classification effect can be obtained.

[0018] Third, by performing multiple rounds of encoding operations on the N public opinion sequences, the encoded N public opinion sequences are obtained. Each time the input of the encoding combines at least part of the data of the adjacent public opinion sequences before and after, and the output of the encoding will delete the encoded data corresponding to the at least part of the data. Such an operation is similar to the convolution operation. As the number of encoding operations increases, the public opinion sequence can see more context information, solving the problem of information loss caused by public opinion segmentation. It is equivalent to further expanding the receptive field when encoding the public opinion sequence, obtaining a more accurate public opinion sequence encoding effect, and then a better public opinion classification effect can be obtained.

[0019] Fourth, the bert model can be used for encoding. The self-attention mechanism layer of the transformer included in it can be utilized. When encoding the target characters of the input public opinion, the semantic information between the target characters and other characters included in the input public opinion can be combined to obtain an accurate encoding result, and then an accurate public opinion analysis result can be obtained.

[0020] Fifth, the bert model pre-trained with a large number of training samples can be used for encoding, effectively migrating the general text knowledge to the current text task and improving the encoding effect.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in one or more embodiments of this application or the related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or the related technologies. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flowchart of a text processing method shown in an embodiment of this application;

[0024] Figure 2 It is a schematic diagram of a multi-round encoding operation shown in an embodiment of this application;

[0025] Figure 3 A schematic diagram of a multi-round encoding operation shown in an embodiment of the present application;

[0026] Figure 4 A schematic diagram of a public opinion analysis process shown in an embodiment of the present application;

[0027] Figure 5 A schematic structural diagram of a text processing device shown in an embodiment of the present application;

[0028] Figure 6 A schematic hardware structure diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners

[0029] Exemplary embodiments will be described in detail below, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0030] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. It should also be understood that the word "if" used herein can be interpreted as "when" or "while" or "in response to determining" depending on the context.

[0031] Based on this, the present application proposes a text processing method. In the process of encoding each text sequence of the text to be processed, at least part of the text data in the text sequences adjacent to the front and back of the text sequence can be combined to encode the text sequence, and the encoded text sequence is obtained. In this way, it is equivalent to expanding the receptive field when encoding the text sequence. Compared with the related art, a more accurate encoding effect of the text sequence can be obtained, and thus a better text classification effect can be obtained.

[0032] Please refer to Figure 1 , Figure 1 A schematic flowchart of a text processing method shown in an embodiment of the present application. Figure 1The text processing method shown can be applied to an electronic device. Among them, the electronic device can execute the text processing method by loading software logic corresponding to the text processing method. The type of the electronic device can be a laptop computer, a computer, a server, a mobile phone, a personal digital assistant (PDA), etc. The type of the electronic device is not particularly limited in this application. The electronic device can also be a client device or a server device, which is not particularly limited here.

[0033] As Figure 1 shown, the method may include S102 - S106. Unless otherwise specified, the execution order of these steps is not particularly limited in this application.

[0034] Among them, in S102, perform a segmentation operation on the text to be processed to obtain N text sequences.

[0035] The text to be processed is a text of a type that needs to be recognized. The text to be processed can have different classifications in different scenarios. For example, in the public opinion analysis scenario, the text to be processed can be public opinion text. Exemplarily, the public opinion text can include positive text and negative text. Of course, in some embodiments, the positive text and negative text can also be further classified. For another example, in the news classification scenario, the text to be processed can be news text. Exemplarily, the news text can include types such as sports, culture, and education.

[0036] According to a certain text length statistical method, the text length can be obtained. Exemplarily, the text length statistical method can include taking the number of characters included in the text as the text length. Among them, one Chinese character is regarded as one character, and one word (including English, French, etc.) is also regarded as one character. For example, the text length of "This comment is very nice" is 6.

[0037] In S102, the text to be processed can be segmented according to a first preset text length to obtain N text sequences.

[0038] The first preset text length is an empirical length and can be set according to requirements.

[0039] In some embodiments, in S102, the segmentation operation may be performed by using the method of window sliding. Exemplarily, a sliding window with both the window size and the moving step size being the first preset text length may be set; starting from the first character of the text to be processed, the sliding window is moved according to the moving step size, and after each movement, the characters within the sliding window are determined as a text sequence; when the sliding window finishes sliding, N text sequences of the first preset text length can be obtained. It can be understood that if the length of the Nth text sequence does not reach the first preset text length, a meaningless blank character (e.g., PAD character) can be used for the complement operation.

[0040] S104. For each of the N text sequences, combine at least part of the text data in the text sequences adjacent to the front and back of the text sequence to encode the text sequence, and obtain the encoded text sequence.

[0041] Based on the encoding mechanism, the encoded text sequence obtained after encoding can indicate the semantic information of the text sequence.

[0042] In this step, the data length of the at least part of the text data can be preset according to requirements.

[0043] Then, starting from the first text sequence of the N text sequences, each text sequence can be sequentially used as the current text sequence, and then the following steps are executed:

[0044] Obtain the first character matching the data length in the previous text sequence of the current text sequence, and the second character matching the data length in the next text sequence of the current text sequence; it should be noted that meaningless characters (e.g., PAD characters) can be used to fill the first character corresponding to the first text sequence and the second character corresponding to the last text sequence.

[0045] Combine the first character, the current text sequence, and the second character to obtain a combined text sequence;

[0046] Encode the combined text sequence to obtain an encoded combined text sequence, and in the encoded combined text sequence, delete the encoded data corresponding to the first character and the second character respectively to obtain the encoded text sequence.

[0047] The encoding involved in this step refers to the process of feature extraction using a neural network. In this step, the encoding can be performed through a natural language processing (NLP) model. In some embodiments, a transformer based on an autoregressive mechanism can be used for the encoding.

[0048] The transformer contains many self-attention mechanism layers. When encoding the target character of the input text, the semantic information between the target character and other characters included in the input text can be combined to obtain an accurate encoding result.

[0049] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a multi-round encoding operation shown in an embodiment of the present application.

[0050] Assume that the text to be processed is divided into 4 text sequences. As Figure 2 shown, for the 2nd text sequence, as Figure 2 the range indicated by the brackets, a part of the first characters at the front position in the 1st text sequence, the 2nd text sequence, and a part of the second characters at the rear position in the 3rd text sequence can be combined to obtain a combined text sequence. Then, the combined text sequence can be input into the transformer layer for encoding to obtain an encoded combined text sequence, and the encoded data corresponding to the first characters and the second characters are removed to obtain the encoded 2nd text sequence. In this way, during the encoding process, the context information of the text sequence can be combined to widen the encoding receptive field and improve the encoding effect.

[0051] S106. Encode the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, and determine the text type of the text to be processed according to the encoding result.

[0052] In this step, the N encoded text sequences can be summarized to obtain a summary result, and then the summary result is encoded to obtain an encoding result corresponding to the text to be processed. The encoding result combines the encoding information of the N text sequences, can indicate the semantic information of the text to be processed, and can be used for text classification of the text to be processed.

[0053] In some embodiments, the text sequence may include a preset character indicating the semantic information of the text sequence. The preset characters included in each of the N text sequences can be summarized to obtain a character sequence corresponding to the text to be processed. This can simplify the computation amount for obtaining the encoding result and improve the text processing efficiency.

[0054] In some embodiments, before encoding the text sequence, the preset character (e.g., CLS character) may be added to the text sequence. The preset character may serve as the first character of the text sequence. Then the text sequence is encoded to obtain the encoded text sequence. Based on the encoding mechanism, the preset character may indicate the semantic information of the text sequence. Then in S106, the preset characters included in each encoded text sequence may be aggregated to obtain a character sequence that can represent the semantics of the text to be processed, i.e., the encoding result.

[0055] Taking the transformer encoding as an example. The CLS character may be added to the text sequence. Then after encoding based on the transformer, the CLS character may indicate the semantic information of the text sequence. In S106, the CLS characters included in each encoded text sequence may be aggregated to obtain a CLS character sequence that can represent the semantics of the text to be processed, i.e., the encoding result.

[0056] After obtaining the encoding result, the text type of the text to be processed may be determined according to the encoding result.

[0057] The encoding result can well indicate the semantic information of the text to be processed. Inputting the encoding result into a classifier can obtain the text type of the text to be processed. The classifier may be obtained through supervised training with some text samples labeled with text types. The text samples may be the encoding results obtained by encoding using S102 - S106, and the text type is the true type of the text sample.

[0058] For example, taking the need to distinguish positive public opinion and negative public opinion as an example. The classifier is a binary classifier constructed based on the softmax function. When training the classifier, several public opinion samples may be obtained. The public opinion samples include the corresponding encoding results obtained by using S102 - S106 and the annotation information (indicating whether the public opinion sample is positive public opinion or negative public opinion). Then the binary classifier may be subjected to supervised training based on the public opinion samples (the supervised training process may refer to related technologies and will not be elaborated here). The trained binary classifier has the ability to distinguish positive public opinion and negative public opinion.

[0059] According to the solution described in S102 - S106, in the process of encoding each text sequence of the text to be processed, at least part of the text data in the text sequences adjacent to the front and back of the text sequence may be combined to encode the text sequence, so as to obtain the encoded text sequence. In this way, it is equivalent to expanding the receptive field when encoding the text sequence. Compared with the related technology, a more accurate text sequence encoding effect can be obtained, and thus a better text classification effect can be obtained.

[0060] In addition, in this method, the text to be processed is first segmented into multiple text sequences, then each text sequence is encoded, and then the encoded results of the multiple text sequences are encoded to obtain an encoded result representing the semantics of the text to be processed. Finally, the text to be processed is classified based on the encoded result. Thus, the classification of long texts can be achieved.

[0061] Taking the need to distinguish positive public opinion and negative public opinion as an example. According to S102 - S106, the public opinion to be processed can be split into N public opinion sequences. Then, for each public opinion sequence, the semantic information of its adjacent front and back public opinion sequences can be combined to broaden the receptive field and obtain the encoded public opinion sequence. After that, the encoded public opinion sequences can be aggregated to obtain the encoded result of the public opinion to be processed. Finally, the encoded result can be input into a pre-trained binary classifier to obtain the public opinion type of the text to be processed. It is not difficult to find that during the encoding of each public opinion sequence, the semantic information of its adjacent front and back public opinion sequences can be combined to broaden the receptive field, and a more accurate encoding effect of the public opinion sequence can be obtained. Furthermore, a better public opinion classification effect can be obtained.

[0062] In some embodiments, in order to further broaden the receptive field during encoding, convolution processing can be referred to, and multiple rounds of encoding operations can be repeatedly executed. As the number of encoding operations deepens, more context information of the text sequence can be seen, solving the problem of information loss caused by text segmentation. This is equivalent to further expanding the receptive field when encoding the text sequence, obtaining a more accurate encoding effect of the text sequence, and thus a better text classification effect can be obtained.

[0063] In S104, multiple rounds of encoding operations can be performed on the N text sequences to obtain the N encoded text sequences.

[0064] Among them, in the first round of encoding operation, for each text sequence among the N text sequences segmented in S102, the text sequence is combined with at least part of the text data in its adjacent front and back text sequences to obtain a combined text sequence, and the combined text sequence is encoded to obtain an encoded combined text sequence. In the encoded combined text sequence, the encoded data corresponding to the at least part of the text data is deleted to obtain the text sequence after the first round of encoding operation.

[0065] Starting from the second encoding operation, the input of each round of encoding operation is each text sequence among the N text sequences obtained in the previous encoding operation, and the remaining operations can refer to the first round of encoding operation, which will not be elaborated here.

[0066] By performing multiple rounds of encoding operations on the N text sequences, the encoded N text sequences are obtained. In each encoding, the input combines at least part of the data of the previous and subsequent text sequences, and the output of the encoding deletes the encoded data corresponding to the at least part of the data. Such an operation is similar to a convolution operation. As the number of encoding operations increases, the text sequences can see more context information, solving the problem of information loss caused by text segmentation. This is equivalent to further expanding the receptive field when encoding the text sequences, obtaining a more accurate encoding effect of the text sequences, and thus a better text classification effect.

[0067] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a multi-round encoding operation shown in an embodiment of this application. Exemplarily, Figure 3 schematically shows the process of the first two rounds of encoding operations. The subsequent encoding operations are similar to the shown encoding operations and are not shown in Figure 3 .

[0068] In some embodiments, a window sliding method can be used to obtain the combined text sequences.

[0069] Specifically, the preset window can be slid in the N text sequences with the first preset text length as the step size, and after each sliding, the segment contained in the preset window is determined as the combined text sequence; the window size of the preset window is the second preset text length; the second preset text length is the sum of the first preset text length and the data length of the at least part of the text data. The first preset text length is the length of the text sequence.

[0070] In some embodiments, the window center of the preset window can be placed at the center of the first text sequence of the N text sequences, and it is slid in the N text sequences with the first preset text length as the step size, and after each sliding, the segment contained in the preset window is determined as the combined text sequence.

[0071] As Figure 3 shown, in the first round of encoding operation, the preset window can start sliding from the 1st text sequence. After 4 slides, the Figure 3 shown 4 combined text sequences can be obtained. For example, for the 2nd text sequence, a combined text sequence composed of a part of the 1st text sequence, the 2nd text sequence, and a part of the 3rd text sequence can be obtained.

[0072] After obtaining the combined text sequences, the text combined text sequences can be respectively input into the transformer layer for encoding to obtain the corresponding encoded combined text sequences.

[0073] After that, data deletion operations can be performed. The encoded data corresponding to the at least part of the data is deleted from the combined text sequence after encoding to obtain the text sequence after the encoding operation. For example, for the combined text sequence corresponding to the 2nd text sequence, the encoded data corresponding to the 1st text sequence and the 3rd text sequence can be deleted from it to obtain the 2nd text sequence after encoding. It can be understood that since the 2nd text sequence incorporates context information during encoding, partial semantic information of the 1st text sequence and the 3rd text sequence will also be covered in the encoded 2nd text sequence.

[0074] Combining the encoded text sequences can obtain 4 text sequences after the first round of encoding.

[0075] Next, the 4 text sequences after the first round of encoding can be used as input, and the process of the first encoding operation can be repeated to obtain 4 text sequences after the second encoding operation. For example, for the 2nd text sequence, it will incorporate partial data of the 1st text sequence and the 3rd text sequence. And since the 3rd text sequence has been encoded and already contains partial semantics of the 4th text sequence, during the second encoding operation, the receptive field of the 2nd text sequence has expanded to the 4th text sequence, that is, the receptive field is larger. By analogy, as the number of encoding times increases, the receptive field of the text sequence will be larger, so that a more accurate encoding effect of the text sequence can be obtained, and then a better text classification effect can be obtained.

[0076] The following is an example description in combination with the public opinion analysis scenario.

[0077] In this scenario, for the relevant public opinions of the company, it can be identified whether they are positive public opinions or negative public opinions. In this scenario, each relevant public opinion can be used as the public opinion to be processed for classification. This public opinion analysis scenario can be implemented through a public opinion analysis system. The public opinion analysis system can include an encoding part and a classification part.

[0078] The encoding part can include a preset first encoding unit and a preset second encoding unit. The classification part can adopt a binary classifier constructed based on softmax that has been pre-trained. The binary classifier can be used to distinguish positive public opinions and negative public opinions.

[0079] The first encoding unit is used to encode the segmented public opinion sequence, and the second encoding unit is used to summarize and encode the encoded public opinion sequences output by the first encoding unit to obtain the encoding result of the public opinion to be processed. The output of the second encoding unit can be used as the input of the binary classifier to obtain the classification result for the public opinion to be processed.

[0080] The first encoding unit and the second encoding unit may include a BERT or ALBERT model. Here, taking the use of the BERT model for encoding as an example. The BERT model is a type of NLP model, which can be used for encoding (feature extraction) of text. The BERT model includes at least one Transformer layer. The Transformer layer is used to perform specific encoding tasks. The characters input to the Transformer and the characters output therefrom are in one-to-one correspondence.

[0081] The BERT model may be a model pre-trained through a text training sample set (the pre-training process may refer to related technologies and will not be elaborated here). Thus, the BERT model pre-trained using a large number of training samples can be used for encoding, effectively transferring and learning general semantic knowledge of the text to the current text task and improving the encoding effect.

[0082] The input length of the BERT model is a third preset text length; the third preset text length is greater than the second preset text length (the sum of the first preset text length and the length of at least part of the data).

[0083] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a public opinion analysis process shown in an embodiment of the present application. As Figure 4 shown, the method may include S401 - S404.

[0084] S401, according to the first preset text length, split the to-be-processed public opinion to obtain N public opinion sequences.

[0085] In this step, the split can be performed by using the window sliding method, which can refer to the relevant description of S102 and will not be elaborated here.

[0086] S402, use the first encoding unit to perform multiple rounds of first encoding on the N public opinion sequences to obtain the N public opinion sequences after multiple rounds of first encoding.

[0087] Figure 4 Only the process of performing the last first encoding on the Sj - th public opinion sequence by using the first encoding unit is schematically shown. As Figure 4 shown, before the last first encoding, the combined public opinion sequence corresponding to the Sj - th public opinion sequence can be obtained first. In this combined public opinion sequence, W0 to Wk - 1 are K first characters included in the Sj - 1 - th public opinion text, Wk to Wn are the Sj - th public opinion sequence, and Wn + 1 to Wn + k are K second characters included in the Sj + 1 - th public opinion text.

[0088] Before the first encoding, the combined public opinion sequence can be subjected to a character completion operation to obtain a first input sequence with a third preset text length. As Figure 4 shown, in the completion operation, a first CLS character for representing the semantics of Sj, a SEP character for indicating the boundary of the public opinion sequence, and a PAD character can be added.

[0089] Then, the first input sequence can be input into the transformer layer corresponding to the current encoding operation included in the first encoding unit, that is, the last transformer layer included in the first encoding unit, for encoding to obtain a first encoded combined public opinion sequence. It can be understood that for the combined public opinion sequence output after the last encoding, the encoded data corresponding to the first character and the second character can be selected to be removed to obtain the encoded public opinion sequence, or the output combined public opinion sequence can be selected to be retained. Regardless of whether the encoded data is removed, the first CLS character representing semantic information needs to be retained.

[0090] Multiple rounds of first encoding will be performed on N public opinion sequences to obtain N public opinion sequences after multiple rounds of first encoding. Among them, the first CLS characters representing semantics are retained in all N public opinion sequences.

[0091] S403, based on the second encoding unit, according to the N encoded public opinion sequences, obtain the encoding result corresponding to the to-be-processed public opinion.

[0092] In this step, the first CLS characters corresponding to the N public opinion sequences can be obtained and summarized to obtain a character sequence corresponding to the to-be-processed public opinion.

[0093] The character sequence can be subjected to a character completion operation to obtain a second input sequence with a third preset text length. In this completion operation, a second CLS character for indicating the semantic information of the to-be-processed public opinion can be added.

[0094] Then, the second input sequence can be input into at least one transformer layer included in the second encoding unit for second encoding to obtain the encoding result. The encoding result includes the second CLS character.

[0095] S404, according to the encoding result, determine the public opinion type of the to-be-processed public opinion.

[0096] In this step, the second CLS character included in the encoding result can be input into a binary classifier to obtain a first confidence level for predicting the to-be-processed public opinion as a positive public opinion and a second confidence level for predicting it as a negative public opinion, and the public opinion type corresponding to the higher confidence level among the two confidence levels can be selected as the public opinion type of the to-be-processed public opinion.

[0097] According to the foregoing solution, first, the text to be processed is segmented into multiple text sequences. Then, after each text sequence is encoded, the encoding results of the multiple text sequences are encoded to obtain an encoding result representing the semantics of the text to be processed. Finally, the text to be processed is classified based on the encoding result. Thus, the classification of long texts can be achieved.

[0098] Second, the bert model can be used for encoding. The self-attention mechanism layer of the transformer included therein can be utilized. When encoding the target characters of the input public opinion, the semantic information between the target characters and other characters included in the input public opinion can be combined to obtain an accurate encoding result, and then an accurate public opinion analysis result can be obtained.

[0099] Third, the bert model pre-trained with a large number of training samples can be used for encoding, effectively transferring the general semantic knowledge of the text to the current text task and improving the encoding effect.

[0100] Fourth, by performing multiple rounds of encoding operations on the N public opinion sequences, the N encoded public opinion sequences are obtained. Each time the input of the encoding combines at least part of the data of the adjacent public opinion sequences before and after, and the output of the encoding deletes the encoding data corresponding to the at least part of the data. Such an operation is similar to a convolution operation. As the number of encoding operations deepens, the public opinion sequence can see more context information, solving the problem of information loss caused by public opinion segmentation, which is equivalent to further expanding the receptive field when encoding the public opinion sequence and obtaining a more accurate encoding effect of the public opinion sequence, and then a better public opinion classification effect can be obtained.

[0101] Corresponding to any of the foregoing embodiments, the present application also proposes a text processing device 500.

[0102] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a text processing device shown in an embodiment of the present application. As Figure 5 shown, the device 500 may include:

[0103] A segmentation module 510 that performs a segmentation operation on the text to be processed to obtain N text sequences;

[0104] A first encoding module 520 that encodes each of the N text sequences based on at least part of the text data in the text sequences adjacent to the front and back of the text sequence to obtain the encoded text sequence;

[0105] The second encoding and classification module 530 encodes the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, and determines the text type of the text to be processed according to the encoding result.

[0106] In some embodiments, the first encoding module 520 is specifically configured to:

[0107] Perform multiple rounds of encoding operations on the N text sequences to obtain the N encoded text sequences; wherein, each round of encoding operation is as follows:

[0108] For each text sequence in the N text sequences obtained from the previous encoding operation, combine the text sequence with at least part of the text data in the text sequences adjacent to it before and after to obtain a combined text sequence, encode the combined text sequence to obtain an encoded combined text sequence, and in the encoded combined text sequence, delete the encoded data corresponding to the at least part of the text data to obtain the text sequence after this encoding operation.

[0109] In some embodiments, the first encoding module 520 is specifically configured to:

[0110] Starting from the first text sequence of the N text sequences, slide the preset window in the N text sequences with the first preset text length as the step size, and determine the segment contained in the preset window after each slide as the combined text sequence; the window size of the preset window is the second preset text length; the second preset text length is the sum of the first preset text length and the data length of the at least part of the text data.

[0111] In some embodiments, the second encoding and classification module 530 is specifically configured to:

[0112] Summarize the preset characters included in each text sequence in the N text sequences to obtain a character sequence corresponding to the text to be processed;

[0113] Encode the character sequence to obtain an encoding result corresponding to the text to be processed.

[0114] In some embodiments, the first encoding module 520 is specifically configured to:

[0115] Based on a preset first encoding unit, for each text sequence in the N text sequences, combine at least part of the text data in the text sequences adjacent to the text sequence before and after, and encode the text sequence to obtain the encoded text sequence;

[0116] The second encoding and classification module 530 is specifically configured to:

[0117] Based on a preset second encoding unit, encode the N encoded text sequences to obtain an encoding result corresponding to the text to be processed.

[0118] In some embodiments, the first encoding unit and the second encoding unit include a BERT model, and the BERT model includes at least one Transformer layer;

[0119] The first encoding module 520 is specifically configured to: encode the combined text sequence by using the Transformer layer corresponding to the current encoding operation included in the first encoding unit;

[0120] The second encoding and classification module 530 is specifically configured to: encode the character sequence by using at least one Transformer layer included in the second encoding unit.

[0121] In some embodiments, the BERT model is a model pre-trained by using a text training sample set.

[0122] In some embodiments, the input length of the BERT model is a third preset text length; the third preset text length is greater than the second preset text length;

[0123] The first encoding module 520 is specifically configured to:

[0124] Perform a character completion operation on the combined text sequence to obtain a first input sequence with a third preset text length;

[0125] Input the first input sequence into the Transformer layer corresponding to the current encoding operation included in the first encoding unit for encoding;

[0126] The second encoding and classification module 530 is specifically configured to:

[0127] Perform a character completion operation on the character sequence to obtain a second input sequence with a third preset text length;

[0128] Input the second input sequence into at least one Transformer layer included in the second encoding unit for encoding.

[0129] In the foregoing solution, first, the text to be processed is segmented into multiple text sequences, then each text sequence is encoded, and then the encoding results of the multiple text sequences are encoded to obtain an encoding result representing the semantics of the text to be processed. Finally, the text to be processed is classified based on the encoding result. Thus, the classification of long texts can be realized.

[0130] Second, during the process of encoding each text sequence of the text to be processed, at least part of the text data in the text sequences adjacent to the front and back of the text sequence can be combined to encode the text sequence, and the encoded text sequence can be obtained. In this way, it is equivalent to expanding the receptive field when encoding the text sequence. Compared with the related technology, a more accurate text sequence encoding effect can be obtained, and then a better text classification effect can be obtained.

[0131] Third, by performing multiple rounds of encoding operations on the N public opinion sequences, the encoded N public opinion sequences are obtained. Each time the input of the encoding combines at least part of the data of the front and back public opinion sequences, and the output of the encoding will delete the encoded data corresponding to the at least part of the data. Such an operation is similar to the convolution operation. As the number of encoding operations deepens, the public opinion sequence can see more context information, solving the problem of information loss caused by public opinion segmentation. It is equivalent to further expanding the receptive field when encoding the public opinion sequence, obtaining a more accurate public opinion sequence encoding effect, and then a better public opinion classification effect can be obtained.

[0132] Fourth, the bert model can be used for encoding. The self-attention mechanism layer of the transformer included in it can be utilized. When encoding the target character of the input public opinion, the semantic information between the target character and other characters included in the input public opinion can be combined to obtain an accurate encoding result, and then an accurate public opinion analysis result can be obtained.

[0133] Fifth, the bert model pre-trained with a large number of training samples can be used for encoding, effectively transferring the general text knowledge to the current text task and improving the encoding effect.

[0134] The embodiments of the text processing device shown in this application can be applied to an electronic device. Correspondingly, this application discloses an electronic device, which may include: a processor.

[0135] A memory for storing instructions executable by the processor.

[0136] Wherein, the processor is configured to call the executable instructions stored in the memory to implement the text processing method shown in any of the foregoing embodiments.

[0137] Please refer to Figure 6 , Figure 6 which is a schematic hardware structure diagram of an electronic device shown in an embodiment of this application.

[0138] As Figure 6As shown, the electronic device may include a processor for executing instructions, a network interface for network connection, a memory for storing operation data for the processor, and a non-volatile memory for storing corresponding instructions of the text processing device.

[0139] Among them, the embodiments of the device can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the electronic device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, in addition to Figure 6 the shown processor, memory, network interface, and non-volatile memory, the electronic device where the device is located in the embodiment usually may also include other hardware according to the actual functions of the electronic device, which will not be elaborated here.

[0140] It can be understood that, in order to improve the processing speed, the corresponding instructions of the text processing device may also be directly stored in the memory, which is not limited here.

[0141] This application proposes a computer-readable storage medium, and the storage medium stores a computer program, and the computer program can be used to enable the processor to execute the text processing method shown in any of the foregoing embodiments.

[0142] Those skilled in the art should understand that one or more embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (which may include but are not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0143] "And / or" in this application means at least one of the two. For example, "A and / or B" can include three scenarios: A, B, and "A and B".

[0144] Each embodiment in this application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the data processing device, since it is basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0145] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0146] Embodiments of the subject matter and the functional operations described in this application can be implemented in: digital electronic circuits, tangible computer software or firmware, computer hardware that may include the structures disclosed in this application and their structural equivalents, or a combination of one or more of them. Embodiments of the subject matter described in this application can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by a data processing apparatus or to control the operation of a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode and transmit information to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0147] The processes and logical flows described in this application can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as dedicated logic circuitry.

[0148] A computer suitable for executing a computer program can include, for example, a general and / or special purpose microprocessor, or any other type of CPU (processor). Generally, the CPU will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer can include a CPU for implementing or executing instructions and one or more memory devices for storing the instructions and data. Generally, the computer will also can include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.

[0149] Computer-readable media suitable for storing computer program instructions and data can include all forms of non-volatile memory, media, and memory devices, such as can include semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0150] Although this application contains many specific implementation details, these should not be construed as limiting the scope of any disclosure or the scope of what is claimed, but are mainly used to describe the features of specific embodiments of a particular disclosure. Certain features described in multiple embodiments in this application can also be implemented in combination in a single embodiment. On the other hand, various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may operate in certain combinations and even be initially claimed as such, one or more features from a claimed combination can in some cases be removed from that combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.

[0151] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system modules and components in the embodiments described should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0152] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processing depicted in the figures is not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0153] The foregoing are only preferred embodiments of one or more embodiments of the present application, and are not intended to limit one or more embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included within the scope of protection of one or more embodiments of the present application.

Claims

1. A text processing method, characterized in that, Including: Performing a segmentation operation on the text to be processed to obtain N text sequences; For each of the N text sequences, encoding the text sequence based on at least part of the text data in the text sequences adjacent to the text sequence before and after it, to obtain the encoded text sequence; Encoding the N encoded text sequences to obtain the encoding result corresponding to the text to be processed, and determining the text type of the text to be processed according to the encoding result; The step of, for each of the N text sequences, combining at least part of the text data in the text sequences adjacent to the text sequence before and after it, and encoding the text sequence to obtain the encoded text sequence, includes: Performing multiple rounds of encoding operations on the N text sequences to obtain the N encoded text sequences; where each round of encoding operation is as follows: For each of the N text sequences obtained from the previous encoding operation, combining the text sequence with at least part of the text data in the text sequences adjacent to it before and after to obtain a combined text sequence, encoding the combined text sequence to obtain the encoded combined text sequence, and deleting the encoded data corresponding to the at least part of the text data in the encoded combined text sequence to obtain the text sequence after the current encoding operation.

2. The method according to claim 1, characterized in that, The length of the text sequence is the first preset text length; the step of, for each of the N text sequences obtained from the previous encoding operation, combining the text sequence with at least part of the text data in the text sequences adjacent to it before and after to obtain a combined text sequence, includes: Starting from the first text sequence of the N text sequences, sliding a preset window in the N text sequences with the first preset text length as the step size, and determining the segment included in the preset window after each sliding as the combined text sequence; the window size of the preset window is the second preset text length; the second preset text length is the sum of the first preset text length and the data length of the at least part of the text data.

3. The method according to claim 2, wherein The text sequence includes a preset character indicating the semantic information of the text sequence; the step of encoding the N encoded text sequences to obtain the encoding result corresponding to the text to be processed, includes: Summarizing the preset characters included in each of the N text sequences to obtain a character sequence corresponding to the text to be processed; Encoding the character sequence to obtain the encoding result corresponding to the text to be processed.

4. The method according to claim 3, wherein The step of, for each of the N text sequences, encoding the text sequence based on at least part of the text data in the text sequences adjacent to the text sequence before and after it, to obtain the encoded text sequence, includes: Based on a preset first encoding unit, for each of the N text sequences, combining at least part of the text data in the text sequences adjacent to the text sequence before and after it, and encoding the text sequence to obtain the encoded text sequence; Encoding the N encoded text sequences to obtain an encoding result corresponding to the text to be processed includes: Encoding the N encoded text sequences based on a preset second encoding unit to obtain an encoding result corresponding to the text to be processed.

5. The method according to claim 4, wherein The first encoding unit and the second encoding unit include a BERT model, and the BERT model includes at least one Transformer layer; Encoding the combined text sequence includes: encoding the combined text sequence by using the Transformer layer corresponding to the current encoding operation included in the first encoding unit; Encoding the character sequence includes: encoding the character sequence by using at least one Transformer layer included in the second encoding unit.

6. The method according to claim 5, wherein The BERT model is a model pre-trained by using a text training sample set.

7. The method according to claim 6, characterized in that, The input length of the BERT model is a third preset text length; the third preset text length is greater than the second preset text length; Encoding the combined text sequence by using the Transformer layer corresponding to the current encoding operation included in the first encoding unit includes: Performing a character completion operation on the combined text sequence to obtain a first input sequence with the third preset text length; Inputting the first input sequence into the Transformer layer corresponding to the current encoding operation included in the first encoding unit for encoding; Encoding the character sequence by using at least one Transformer layer included in the second encoding unit includes: Performing a character completion operation on the character sequence to obtain a second input sequence with the third preset text length; Inputting the second input sequence into at least one Transformer layer included in the second encoding unit for encoding.

8. A text processing device, characterized in that, Includes: A splitting module that performs a splitting operation on the text to be processed to obtain N text sequences; A first encoding module that, for each of the N text sequences, encodes the text sequence based on at least partial text data in the text sequences adjacent to the front and back of the text sequence to obtain the encoded text sequence; A second encoding and classification module that encodes the N encoded text sequences to obtain an encoding result corresponding to the text to be processed, and determines the text type of the text to be processed according to the encoding result; The first encoding module specifically performs multiple rounds of encoding operations on the N text sequences to obtain the N encoded text sequences; wherein, each round of encoding operation is as follows: For each of the N text sequences obtained from the previous encoding operation, combining the text sequence with at least partial text data in the text sequences adjacent to the front and back of it to obtain a combined text sequence, encoding the combined text sequence to obtain an encoded combined text sequence, and deleting the encoded data corresponding to the at least partial text data from the encoded combined text sequence to obtain the text sequence after the current encoding operation.

9. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the text processing method according to any one of claims 1-7 by running the executable instructions.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for causing a processor to execute the text processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Question and answer text matching method based on multilayer semantic feature extraction structure

    CN111831789A

  • Transformer-based multi-feature Chinese and English sentiment classification method and Transformer-based multi-feature Chinese and English sentiment classification system

    CN111858932A