A text enhancement method, device, electronic device, and storage medium

By identifying and replacing similar text units of text units, the problem of low quality of text enhancement in the prior art is solved, and the balance of diversity and flexibility is achieved.

CN113822047BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110750553.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-02
Publication Date
2025-06-27
Estimated Expiration
2041-07-02

AI Technical Summary

Technical Problem

When existing text enhancement methods generate new text through back translation, it is easy to lead to low text quality and poor statements.

Method used

By obtaining the text units in the target text, identifying their number of characters and language types, selecting similar text units from the candidate text units according to the similarity and replacing them, and generating similar text matching the target text.

Benefits of technology

Taking into account the diversity of generated text and the clarity of statements, the text quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822047B_ABST
    Figure CN113822047B_ABST
Patent Text Reader

Abstract

The present application discloses a text enhancement method, apparatus, electronic device, and storage medium; at least one target text unit to be replaced can be selected from the text units of the target text; when the number of characters of the target text unit is greater than a preset number and its language type is the target language, a similar text unit of the target text unit is selected from the candidate text units based on the similarity between the target text unit and the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, a similar text unit of the target text unit is selected from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units; based on the similar text units of each target text unit, the target text units in the target text are replaced to generate a similar text. The present application is beneficial to taking into account both the diversity of the generated text and the smoothness of the sentences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a text enhancement method, apparatus, electronic device, and storage medium. Background Art

[0002] Natural language processing is an important direction in the fields of computer science and artificial intelligence. As an important application in natural language processing, text enhancement is widely used in expanding training samples. Text enhancement is a data augmentation technique, and its core idea is to increase the number of samples by slightly modifying a copy of the existing text data or creating new synthetic texts from the existing text data without changing the overall sentiment of the sentence.

[0003] In the current related technologies, generally back-translation is used for text enhancement, that is, based on machine translation technology, the original text is translated into another or multiple languages, and then translated back to obtain new samples. For example, the original Chinese text is translated into the corresponding English text, then the English text is translated into the corresponding Japanese text, and then the Japanese text is translated into Chinese text to obtain a new Chinese text with the same overall sentiment as the original Chinese text. However, this is likely to result in a low quality of the generated text and the problem of unsmooth sentences. Summary of the Invention

[0004] Embodiments of this application provide a text enhancement method, apparatus, electronic device, and storage medium, which are beneficial to taking into account the diversity of the generated text and the smoothness of the sentences.

[0005] Embodiments of this application provide a text enhancement method, including:

[0006] Obtain a target text, and select at least one target text unit to be replaced from the text units of the target text;

[0007] Identify the number of characters and the language type of the target text unit;

[0008] When the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units;

[0009] When the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units;

[0010] Based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text.

[0011] Correspondingly, an embodiment of the present application provides a text enhancement device, including:

[0012] An acquisition unit, configured to acquire a target text and select at least one target text unit to be replaced from the text units of the target text;

[0013] An identification unit, configured to identify the number of characters and the language type of the target text unit;

[0014] A first selection unit, configured to, when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units;

[0015] A second selection unit, configured to, when the number of characters of the target text unit is not greater than the preset number or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units;

[0016] A replacement unit, configured to replace the target text units in the target text based on the similar text units of each target text unit to generate a similar text that matches the target text.

[0017] Optionally, in some embodiments of the present application, the acquisition unit may include a frequency analysis subunit and a first selection subunit, as follows:

[0018] The frequency analysis subunit is configured to perform frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text;

[0019] The first selection subunit is configured to select at least one target text unit to be replaced from the text units of the target text based on the importance parameter.

[0020] Optionally, in some embodiments of the present application, the frequency analysis subunit may specifically be configured to: for each text unit in the target text, count the frequency of occurrence of the text unit in the target text to obtain the weight of the text unit in the target text; count the frequency of occurrence of the text unit in the sample text to obtain the reference weight of the text unit; and determine the importance parameter of the text unit according to the reference weight of the text unit and the weight in the target text.

[0021] Optionally, in some embodiments of the present application, the first selection unit may include a first addition subunit, a second selection subunit, and a third selection subunit, as follows:

[0022] The first addition subunit is configured to add a target text unit whose character number is greater than a preset number and whose language type is the target language to the first text sequence;

[0023] The second selection subunit is configured to select a target replacement text unit from the first text sequence according to a preset replacement ratio;

[0024] The third selection subunit is configured to select a similar text unit of the target replacement text unit from the candidate text units based on the similarity between the target replacement text unit and the candidate text units.

[0025] Optionally, in some embodiments of the present application, the second selection unit may include a second addition subunit, a fourth selection subunit, and a fifth selection subunit, as follows:

[0026] The second addition subunit is configured to add a target text unit whose character number is not greater than a preset number or whose language type is a non-target language to the second text sequence;

[0027] The fourth selection subunit is configured to select a target replacement text unit from the second text sequence based on a preset replacement ratio;

[0028] The fifth selection subunit is configured to select a similar text unit of the target replacement text unit from the candidate text units according to the similarity between the context information of the target replacement text unit and the candidate text units.

[0029] Optionally, in some embodiments of the present application, the first selection unit may include an extraction subunit, a calculation subunit, and a sixth selection subunit, as follows:

[0030] The extraction subunit is configured to perform feature extraction on the target text unit and the candidate text unit respectively to obtain the feature information of the target text unit and the feature information of the candidate text unit;

[0031] A calculation subunit, configured to calculate a similarity between the target text unit and the candidate text unit based on feature information of the target text unit and feature information of the candidate text unit;

[0032] A sixth selection subunit, configured to select a similar text unit of the target text unit from the candidate text units according to the similarity.

[0033] Optionally, in some embodiments of the present application, the second selection unit may include an acquisition subunit, a context extraction subunit, and a seventh selection subunit, as follows:

[0034] The acquisition subunit is configured to acquire feature information of a reference text unit in the target text, where the reference text unit is other text units in the target text except the target text unit;

[0035] The context extraction subunit is configured to perform feature extraction on the target text unit according to the feature information of the reference text unit to obtain context feature information corresponding to the context information of the target text unit;

[0036] The seventh selection subunit is configured to select a similar text unit of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text unit.

[0037] Optionally, in some embodiments of the present application, the context extraction subunit may specifically be configured to obtain weights corresponding to each reference text unit; and fuse the feature information of each reference text unit according to the weights to obtain context feature information corresponding to the context information of the target text unit.

[0038] Optionally, in some embodiments of the present application, the text enhancement device may further include a division unit, as follows:

[0039] The division unit is configured to divide the target text through a word segmentation model to obtain each text unit in the target text, where the word segmentation model is trained based on a sample text and a position tag sequence corresponding to the text unit in the sample text.

[0040] An electronic device provided by an embodiment of the present application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the text enhancement method provided by the embodiment of the present application.

[0041] In addition, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the text enhancement method provided by the embodiment of the present application are implemented.

[0042] An embodiment of the present application provides a text enhancement method, apparatus, electronic device, and storage medium. It can obtain a target text and select at least one target text unit to be replaced from the text units of the target text; identify the number of characters and language type of the target text unit; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, based on the similarity between the target text unit and candidate text units, select a similar text unit of the target text unit from the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, based on the context information of the target text unit and the similarity between the candidate text units, select a similar text unit of the target text unit from the candidate text units; based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text. The present application can adopt different methods to obtain similar text units of target text units based on the number of characters and language type of text units, and then replace the target text units, which is beneficial to taking into account the diversity of the generated text and the smoothness of the sentences. Description of the Drawings

[0043] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1a is a schematic diagram of the scenario of the text enhancement method provided by the embodiment of the present application;

[0045] Figure 1b is a flowchart of the text enhancement method provided by the embodiment of the present application;

[0046] Figure 1c is a model structure diagram of the text enhancement method provided by the embodiment of the present application;

[0047] Figure 1d is another model structure diagram of the text enhancement method provided by the embodiment of the present application;

[0048] Figure 1e is another model structure diagram of the text enhancement method provided by the embodiment of the present application;

[0049] Figure 1f It is an experimental illustration diagram of the text enhancement method provided by the embodiments of the present application;

[0050] Figure 1g It is another experimental illustration diagram of the text enhancement method provided by the embodiments of the present application;

[0051] Figure 2 It is another flowchart of the text enhancement method provided by the embodiments of the present application;

[0052] Figure 3 It is a schematic structural diagram of the text enhancement device provided by the embodiments of the present application;

[0053] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Specific Embodiments

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.

[0055] The embodiments of the present application provide a text enhancement method, device, electronic device, and storage medium. The text enhancement device can be specifically integrated in the electronic device, and the electronic device can be a device such as a terminal or a server.

[0056] It can be understood that the text enhancement method in this embodiment can be executed on the terminal, can also be executed on the server, or can be jointly executed by the terminal and the server. The above examples should not be construed as a limitation to the present application.

[0057] As Figure 1a shown, taking the terminal and the server jointly executing the text enhancement method as an example. The text enhancement system provided by the embodiments of the present application includes the terminal 10 and the server 11, etc.; the terminal 10 and the server 11 are connected through a network, for example, through a wired or wireless network connection, etc., where the text enhancement device can be integrated in the server.

[0058] Among them, the server 11 can be used to: obtain a target text, and select at least one target text unit to be replaced from the text units of the target text; identify the number of characters and language type of the target text unit; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units; replace the target text units in the target text based on the similar text units of each target text unit to generate a similar text that matches the target text; and send the similar text to the terminal 10. Among them, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers.

[0059] Among them, the terminal 10 can receive the similar text sent by the server 11 and effectively expand the training data for downstream tasks in natural language processing. Among them, the terminal 10 can include a mobile phone, a smart TV, a tablet computer, a laptop computer, a personal computer (PC), a wearable device, or an in-vehicle computer, etc. A client can also be set on the terminal 10, and the client can be an application client or a browser client, etc.

[0060] The steps of the above-mentioned server 11 for text enhancement can also be executed by the terminal 10.

[0061] The text enhancement method provided by the embodiments of the present application relates to natural language processing in the field of artificial intelligence. The embodiments of the present application are beneficial to taking into account the diversity of the generated text and the smoothness of the sentences.

[0062] Among them, artificial intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level and software-level technologies. Among them, artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0063] Among them, natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0064] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0065] This embodiment will be described from the perspective of a text enhancement device, which can be specifically integrated in an electronic device, and the electronic device can be a device such as a server or a terminal.

[0066] The text enhancement method of the embodiments of this application can be applied to various scenarios that require generating similar texts. For example, when training an NLP model, there are only a small number of training samples, and diverse training samples are needed. The text enhancement method provided by this embodiment can be used to provide diverse training samples for the NLP model training from a small amount of text, thereby alleviating the problem of lack of training samples and at the same time improving the generalization ability of the model.

[0067] As Figure 1b shown, the specific process of this text enhancement method can be as follows:

[0068] 101. Obtain a target text, and select at least one target text unit to be replaced from the text units of the target text.

[0069] Among them, the target text is the text to be data-augmented, which includes at least one text unit. The text unit can be a word or a character. Specifically, in this embodiment, it is necessary to generate a similar text that is semantically similar to the target text.

[0070] Among them, the target text can include at least one language type, and this embodiment does not limit the language type.

[0071] Among them, Data Augmentation is a data augmentation technique in data analysis. Specifically, it increases the number of samples by slightly modifying the copies of the existing data or creating new synthetic data from the existing data. In the process of model training in machine learning, this technique can play a role in regularization, helping to reduce overfitting and enhance the generalization performance of the model.

[0072] In this embodiment, some text units can be selected from the text units of the target text as the target text units to be replaced, and then similar text units similar to the target text units can be selected from the candidate text units. Based on the similar text units, the corresponding target text units in the target text are replaced to obtain the replaced target text. The replaced target text is also the text after data augmentation (specifically text augmentation) of the target text.

[0073] By performing text augmentation on the target text, new samples can be generated without changing the original sentiment of the target text, effectively expanding the training data for downstream tasks in NLP, providing diverse training samples, thus alleviating the problem of lack of training samples and at the same time enhancing the generalization ability of the model. The trained NLP model can be used in multiple downstream services such as video barrage classification, public account message recognition, and public account spam text interception.

[0074] Optionally, in this embodiment, before the step of "selecting at least one target text unit to be replaced from the text units of the target text", it may further include:

[0075] The target text is divided through a word segmentation model to obtain each text unit in the target text. The word segmentation model is trained based on the sample text and the position tag sequence corresponding to the text units in the sample text.

[0076] Among them, the word segmentation model can be a neural network, specifically a Bidirectional Encoder Representations from Transformers (Bert), or a Long Short-Term Memory (LSTM), a Bi-directional Long Short-Term Memory (BiLSTM), a Recurrent Neural Network (RNN), etc. The word segmentation model in this embodiment is not limited to the several types listed above.

[0077] Among them, the word segmentation model can be trained based on the sample text and the position tag sequence corresponding to the text units in the sample text. The position tag sequence of the text units can adopt the BIES paradigm, where BIES represents the position of a certain character in a word. B (begin) means that the character is the starting character in the word, I means that the character is the middle character in the word, E (end) means that the character is the ending character in the word, and S (single) can represent a single-character word.

[0078] In some embodiments, the target text includes Chinese text. Word segmentation of Chinese text is relatively special. Chinese word segmentation belongs to a sequence labeling task in NLP, which refers to assigning corresponding labels to the input sequence, such as a position tag sequence. Chinese word segmentation is the process of splitting a continuous sequence of Chinese characters according to certain specifications and recombining them into individual word sequences. Specifically, the above word segmentation model can be used to perform Chinese word segmentation on the Chinese text to obtain each text unit after the Chinese text is segmented. Each text unit can be denoted as a token.

[0079] In a specific scenario, the above word segmentation model (specifically, the Bert model) can be used to segment Chinese text (such as a Chinese single sentence), converting Chinese word segmentation into a sequence labeling problem for Chinese characters. To prevent unreasonable jump situations in the sequence labeling results output by the Bert model, a CRF model can be connected after the Bert model, as Figure 1c shown.

[0080] Figure 1c In, CLS can be regarded as a position tag sequence. The Bert model extracts features from each text unit in the single sentence according to the CLS flag, generating a set of feature vectors T1, T2…T N, and fine-tuned through a fully connected layer, which can be a CRF model. CRF, short for Conditional Random Fields, is a conditional random field. The CRF model can be regarded as a task-related layer of the Bert model.

[0081] During the training process of the word segmentation model, a position label sequence can be annotated for each sample text, and the sample text can be segmented through the word segmentation model. Based on the loss value between the word segmentation result and the position label sequence, the parameters in the word segmentation model (including the Bert model and the CRF model) are adjusted to obtain the trained word segmentation model.

[0082] Optionally, in this embodiment, the step of "selecting at least one target text unit to be replaced from the text units of the target text" may include:

[0083] Perform frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text;

[0084] Based on the importance parameter, select at least one target text unit to be replaced from the text units of the target text.

[0085] Optionally, in this embodiment, a text unit with an importance parameter greater than a preset value can be selected as the target text unit to be replaced, and the preset value can be set according to the actual situation; or based on the importance parameter, the text units of the target text can be sorted, such as sorting them from large to small, and the first n text units after sorting are used as the target text units to be replaced.

[0086] In some embodiments, before performing frequency analysis on each text unit in the target text, the text units in the target text can be filtered first. Specifically, the filtering process can include removing meaningless text units such as stop words, spaces, and punctuation marks from the segmented target text. It can be understood that the filtered characters can also be customarily added according to the application scenario, and this embodiment does not limit this. Among them, if the length of the target text after filtering is less than 2, that is, the number of text units included in the filtered target text is less than 2, the filtering process of the target text can be stopped and the original text can be returned.

[0087] Among them, the stop words can be set according to the actual situation, and it can be selected not to replace specified words and characters according to business needs. For example, some proper nouns can be specified as stop words and not replaced.

[0088] Optionally, in this embodiment, the step of "performing frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text" may include:

[0089] For each text unit in the target text, count the frequency of occurrence of the text unit in the target text to obtain the weight of the text unit in the target text;

[0090] Count the frequency of occurrence of the text unit in the sample text to obtain the reference weight of the text unit;

[0091] Determine the importance parameter of the text unit according to the reference weight of the text unit and its weight in the target text.

[0092] Among them, the frequency of occurrence of a text unit in the target text is specifically the word frequency of the text unit in the target text, which can be represented by TF. TF refers to the frequency of occurrence of a given word in the file, which is a normalization of the number of words to prevent it from biasing towards long files. Its calculation method is shown in formula (1):

[0093]

[0094] Among them, n i,j is the number of occurrences of the word in document d j , and the denominator is the sum of the number of occurrences of all words in document d j . k represents the number of different words in document d j . tf i,j represents the word frequency of the word in document d j . In this embodiment, document d j can specifically be the target text, and the word frequency tf i,j of the text unit in the target text can be directly used as the weight of the text unit in the target text.

[0095] Among them, the sample text can specifically be the text in the document library, and the reference weight of the text unit can specifically be represented by the inverse document frequency. The inverse document frequency of a certain text unit characterizes its frequency of occurrence in the corpus, and its calculation method is shown in formula (2):

[0096]

[0097] Among them, IDF is the inverse document frequency of word q i , N is the total number of documents in the library, n(q i ) is the number of documents containing word q i , and 0.5 is the smoothing coefficient.

[0098] Optionally, the step of "determining the importance parameter of the text unit according to the reference weight of the text unit and its weight in the target text" may include:

[0099] Fusing the reference weight of the text unit and its weight in the target text to obtain the importance parameter of the text unit.

[0100] Among them, there are many fusion methods. For example, the fusion method can be multiplication, etc., and this embodiment does not limit it. Specifically, the importance parameter of the text unit can be obtained by multiplying the term frequency TF and the inverse document frequency IDF of the text unit, and this importance parameter can be represented by TF-IDF.

[0101] The full name of TF-IDF is Term Frequency–Inverse Document Frequency, that is, term frequency-inverse document frequency. It is a commonly used weighting technique for information retrieval and text mining. TF-IDF is a statistical method used to evaluate the importance of a word for a document set or a single document in a corpus. The importance of a word increases in direct proportion to the number of times it appears in the document, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus.

[0102] In practical applications, since the term frequency is usually 1 in the short text scenario, the IDF can be directly used as the importance parameter of the text unit. In addition, it should be noted that the IDF can be calculated from the sample set or migrated from other scenarios according to business needs.

[0103] 102. Identify the number of characters and language type of the target text unit.

[0104] Among them, the target text unit can be one character, or two or more characters. In this embodiment, different methods can be used to obtain the similar text unit of the target text unit according to the number of characters of the target text unit.

[0105] For example, if the target text unit is the Chinese character "flow", which contains two characters, feature extraction can be performed on the target text unit to obtain the feature information of the target text unit. Based on the feature information of the target text unit and the feature information of the candidate text unit, the similar text unit of the target text unit is selected from the candidate text units.

[0106] Another example is that if the target text unit is the Chinese character "such as", which contains one character, the context information of the target text unit can be obtained, and the similar text unit of the target text unit is selected from the candidate text units according to the similarity between the context information and the candidate text units.

[0107] In addition, the target text may contain only one language type, or may contain two or more language types. For example, the target text may be short text data mixed with Chinese. Therefore, the target text units divided from the target text may belong to language type 1 or language type 2.

[0108] 103. When the number of characters in the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units.

[0109] Among them, the preset number and the target language can be set according to the actual situation, and this embodiment does not limit this. For example, the preset number can be 1 and the target language can be Chinese.

[0110] Optionally, in this embodiment, the step of "selecting a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units" may include:

[0111] Extract features from the target text unit and the candidate text units respectively to obtain the feature information of the target text unit and the feature information of the candidate text units;

[0112] Calculate the similarity between the target text unit and the candidate text units based on the feature information of the target text unit and the feature information of the candidate text units;

[0113] Select a similar text unit of the target text unit from the candidate text units according to the similarity.

[0114] Among them, features can be extracted from the target text unit and the candidate text units through a semantic model specifically. This feature extraction may include convolutional processing and pooling processing. Optionally, the semantic model may be a Word2vec (word to vector) model, a fastText (fast text) model, a Glove model, etc. Glove is a word embedding tool.

[0115] Among them, Word2vec is a type of neural network model, which mainly includes three-layer structures: an input layer, a projection layer, and an output layer. As Figure 1dAs shown, the two important models used in Word2vec are the CBOW model (Continuous Bag-of-Words Model) and the Skip-gram model (Continuous Skip-gram Model). Both models consist of three layers: an input layer, a projection layer, and an output layer. Specifically, the former predicts the semantics of the current word \(W\) t (the current word) given the context \(W\) t-2 , \(W\) t-1 , \(W\) t+1 , \(W\) t+2 (assuming a window size of 2), while the latter does the opposite. It predicts the context \(W\) t given the current word \(W\) t , \(W\) t-2 , \(W\) t-1 , \(W\) t+1 , \(W\) t+2 .

[0116] Among them, \(W\) t+1 represents the \((t + 1)\)-th window word, i.e., the first text unit after the central word; \(W\) t-1 represents the \((t - 1)\)-th window word, i.e., the first text unit before the central word, and so on. Here, the central word can specifically be the text unit to be semantically recognized.

[0117] In this embodiment, the feature information of the target text unit and the candidate text units can both be obtained by extracting through the Word2vec model. In a specific embodiment, before text augmentation of the target text, the feature vectors of the preset text units can be extracted through a trained Word2vec model and stored in a preset word vector table. The preset word vector table includes the mapping relationship between the preset text units and their corresponding feature vectors (specifically word vectors). When it is necessary to calculate the similarity between the target text unit and the candidate text units, the feature information of the target text unit and the feature information of each candidate text unit can be directly obtained from the preset word vector table for similarity calculation.

[0118] Among them, word vector (word embedding): Also known as word embedding, word representation, text representation, etc., it is a general term for language models and representation learning techniques in natural language processing (NLP). Conceptually, it refers to embedding a high-dimensional space with a dimension equal to the number of all words into a much lower-dimensional continuous vector space, and each word or phrase is mapped to a vector in the real number domain.

[0119] Among them, the step "calculating the similarity between the target text unit and the candidate text unit based on the feature information of the target text unit and the feature information of the candidate text unit" may include:

[0120] Calculating the vector distance between the feature vector of the target text unit and the feature vector of the candidate text unit;

[0121] Calculating the similarity between the target text unit and the candidate text unit according to the vector distance.

[0122] Among them, the vector distance can represent the similarity between the target text unit and the candidate text unit. The larger the vector distance, the smaller the similarity; on the contrary, the smaller the vector distance, the larger the similarity. This vector distance can specifically be a cosine distance, an Euclidean distance, etc., and this embodiment does not limit this.

[0123] In a specific embodiment, the feature vector of the target text unit and the feature vector of the candidate text unit can be multiplied, and the obtained product is the cosine similarity.

[0124] In this embodiment, after obtaining the similarity between the target text unit and each candidate text unit, the candidate text units with a similarity greater than the preset similarity threshold can be used as the similar text units of the target text unit, or based on the similarity, each candidate text unit can be sorted, such as sorting from large to small, to obtain the sorted candidate text units, and then the first n candidate text units in the sorted candidate text units can be used as the similar text units of the target text unit, where n can be set according to the actual situation.

[0125] Optionally, in this embodiment, the step "when the number of characters of the target text unit is greater than the preset number and the language type of the target text unit is the target language, selecting the similar text units of the target text unit from the candidate text units" may include:

[0126] Adding the target text unit with the number of characters greater than the preset number and the language type being the target language to the first text sequence;

[0127] Selecting the target replacement text unit from the first text sequence according to the preset replacement ratio;

[0128] Selecting the similar text units of the target replacement text unit from the candidate text units based on the similarity between the target replacement text unit and the candidate text units.

[0129] Among them, for the target text units that meet the conditions: the number of characters is greater than the preset number and the language type is the target language, they can be screened again. Specifically, a preset replacement ratio of target text units can be screened, and the screened target text units are used as target replacement text units, and the similarity between the target replacement text units and the candidate text units is calculated to select similar text units of the target replacement text units from the candidate text units.

[0130] Among them, the preset replacement ratio can be set according to the actual situation, and this embodiment does not limit it.

[0131] 104. When the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, similar text units of the target text unit are selected from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units.

[0132] Among them, if the target language is Chinese, the non-target language can be other language types except Chinese.

[0133] Optionally, in this embodiment, the step "when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, similar text units of the target text unit are selected from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units" may include:

[0134] Add target text units with the number of characters not greater than the preset number or the language type being a non-target language to the second text sequence;

[0135] Based on the preset replacement ratio, select target replacement text units from the second text sequence;

[0136] According to the similarity between the context information of the target replacement text unit and the candidate text units, select similar text units of the target replacement text unit from the candidate text units.

[0137] Among them, the preset replacement ratio can be set according to the actual situation, and this embodiment does not limit it.

[0138] Among them, for the target text units that meet the conditions: the number of characters is not greater than the preset number and the language type is a non-target language, they can be screened again. Specifically, a preset replacement ratio of target text units can be screened, and the screened target text units are used as target replacement text units, and the similarity between the target replacement text units and the candidate text units is calculated to select similar text units of the target replacement text units from the candidate text units.

[0139] Optionally, in some other embodiments, the step of "when the number of characters of the target text unit is not greater than a preset number, or the language type of the target text unit is a non-target language, selecting a similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units" may include:

[0140] Adding a target text unit with a non-target language type to a third text sequence;

[0141] Adding a target text unit with a character number not greater than a preset number and a target language type to a fourth text sequence;

[0142] Based on a preset replacement ratio, respectively selecting target replacement text units from the third text sequence and the fourth text sequence;

[0143] Selecting a similar text unit of the target replacement text unit from the candidate text units according to the similarity between the context information of the target replacement text unit and the candidate text units.

[0144] It can be understood that the third text sequence and the fourth text sequence can be regarded as subsequences of the second text sequence.

[0145] In a specific embodiment, the target text contains English and Chinese. After tokenizing and filtering it, the TF-IDF values corresponding to the text units in the target text can be obtained, and a mask sequence is generated according to the following rules:

[0146] If the text unit is a single Chinese character, set its mask to 1; if the text unit is a Chinese word (including two or more Chinese characters), set its mask to 2; if the text unit is an English word, set its mask to 3; otherwise, if the text unit is not in the vocabulary or other undefined situations, set its mask to 0.

[0147] Traverse the mask sequence and the TF-IDF sequence of the text units in the target text, and generate a sequence to be replaced according to the following rules:

[0148] (1) If the mask of the current text unit is 1 and the TF-IDF value is less than the TF-IDF-THRESHOLD global threshold, add the ID (Identity document, identity identification information) of the current text unit to sequence s1 (specifically, it can be the fourth text sequence in the above embodiment);

[0149] (2) If the mask of the current text unit is 2 and the TF-IDF value is greater than the global threshold TF-IDF-THRESHOLD, add the ID of the current text unit to sequence s2 (specifically, it can be the first text sequence in the above embodiments);

[0150] (3) If the mask of the current text unit is 3 and the TF-IDF value is greater than the global threshold TF-IDF-THRESHOLD, and the length after segmentation using the WordPiece algorithm is greater than 1, add the ID of the current text unit to sequence s3 (specifically, it can be the third text sequence in the above embodiments);

[0151] (4) Otherwise, skip this text unit.

[0152] Among them, TF-IDF-THRESHOLD is a global parameter, which can be specified by the user or replaced by the average value of the TF-IDF sequences of the input text units. This embodiment does not limit this.

[0153] Among them, the WordPiece algorithm is an algorithm that disassembles words into chunks. Specifically, it can separate the meaning and tense of the word itself, effectively reducing the number of the vocabulary and improving the training speed. Its main implementation method is through double-byte encoding.

[0154] Among them, the second text sequence in the above embodiments may include sequence s1 and sequence s3.

[0155] After obtaining sequences s1, s2, and s3, some text units can be selected from sequences s1, s2, and s3 as the target replacement text units to be replaced based on a preset replacement ratio. Specifically, for different sequences, the preset replacement ratio can be the same or different. This embodiment does not limit this.

[0156] In a specific embodiment, the number of target replacement text units for each sequence can be determined according to the lengths of sequences s1, s2, and s3 n can take 1 or 2 or 3, and its calculation method is shown in formulas (3)(4)(5)(6):

[0157]

[0158]

[0159] Among them, bool is a boolean operation. When the sequence length is 0, the boolean operation result is 0; when the sequence length is greater than 0, the boolean operation result is 1. len is a function to solve the sequence length, and the length of a string can be calculated through the len function.

[0160] Among them, SUB RATIOThe parameter represents the replacement ratio, which can be specified by the user and can be obtained from the SUB RATIO parameter to get the number of replaceable tokens SUB CNT , if SUB RATIO parameter is a floating-point number between [0, 1], then SUB CNT = len(token) * SUB RATIO ; if SUB RATIO parameter is an integer greater than 0, then SUB CNT = min(SUB RATIO , len(token)) - 1. To make the generated text more diverse, a random number can be taken between and SUB CNT for the calculated replacement amount, and its calculation method can be: For example, SUB RATIO can be set to 0.4 to enable random replacement.

[0161] Among them, is the number of target replacement text units in sequence s1; is the number of target replacement text units in sequence s2; is the number of target replacement text units in sequence s3.

[0162] In this embodiment, for the target replacement text units in sequence s2, text enhancement can be performed by replacement based on word vectors. For the target replacement text units in sequences s1 and s3, replacement based on context information can be used for text enhancement.

[0163] It should be noted that in this embodiment, for the text units in sequence s3, the selected similar text units of the text unit cannot contain the n-gram sub-words of the text unit, where gram is each byte segment.

[0164] Optionally, in this embodiment, the step of "selecting the similar text units of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units" may include:

[0165] Obtain the feature information of the reference text unit in the target text, where the reference text unit is other text units in the target text except the target text unit;

[0166] Extract the features of the target text unit according to the feature information of the reference text unit to obtain the context feature information corresponding to the context information of the target text unit;

[0167] Select a similar text unit of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text unit.

[0168] Among them, the context feature information corresponding to the context information of the target text can be obtained through a semantic model, and then a similar text unit of the target text unit is selected from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text unit. This semantic model can be an MLM model. Specifically, this semantic model can be BERT (Bidirectional Encoder Representations from Transformers), ALBERT (A Lite BERT), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), etc. This embodiment does not limit this.

[0169] Among them, Bert is an open-source sequential model based on the Transformer structure, and its model structure is as Figure 1e shown. Bert is composed of multiple layers of bidirectional Transformers connected together, generally 12 layers or 24 layers, etc. Figure 1e In it, E1, E2... EN represent inputs, T1, T2... TN represent outputs, J represents each Transformer node of Bert, and the connection lines represent the dependency relationships between words. The dependency relationships in Bert have both previous and subsequent contexts. Bert can be obtained through pre-training and fine-tuning. The training data of Bert is corpus information. Bert mainly includes two tasks during training. The first is to randomly remove words from the training corpus and replace them with masks, and let the model predict the removed words. The second task is that each training data is a pair of upper and lower sentences. Among them, the lower sentences of some training data are truly related to the upper sentences, while the lower sentences of some training data are unrelated to the upper sentences. The model needs to judge the relationship between the upper and lower sentences in the training data. The model is optimized based on the loss values of the two tasks. The training process of Bert can make full use of the context information, making the model have stronger expression ability. After completing the pre-training of the model, the model can also be fine-tuned for specific tasks. Fine-tuning is a commonly used transfer learning technique in the field of deep learning. Through fine-tuning, the model can better adapt to the language knowledge in specific scenarios.

[0170] Among them, since the self-attention mechanism in BERT considers context information during prediction, the text generation after word vector replacement using this method will be more grammatically smooth.

[0171] Among them, MLM (Masked Language Model) is a bidirectional language model. During training, this model randomly masks some text in the input, aiming to let the model predict the original vocabulary of the masked words based on its context. Different from the traditional left-to-right unidirectional language model pre-training, MLM can perform bidirectional deep pre-training by integrating the context on both sides.

[0172] Optionally, in this embodiment, the step of "extracting features from the target text unit according to the feature information of the reference text unit to obtain the context feature information corresponding to the context information of the target text unit" may include:

[0173] Obtain the weights corresponding to each reference text unit;

[0174] According to the weights, fuse the feature information of each reference text unit to obtain the context feature information corresponding to the context information of the target text unit.

[0175] Among them, there are various fusion methods, and this embodiment does not limit this. For example, it can be weighted fusion or feature splicing, etc.

[0176] Among them, the step of "selecting the similar text unit of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text units" may include:

[0177] Select the candidate text units with a similarity to the context feature information greater than the preset similarity threshold as the similar text units of the target text unit.

[0178] Optionally, in some embodiments, the candidate text units can also be sorted according to the similarity between the context feature information and the feature information of the candidate text units, such as sorting from large to small, to obtain the sorted candidate text units, and the first k text units of the sorted candidate text units are used as the similar text units of the target text unit.

[0179] Among them, before calculating the similarity between the context feature information and the feature information of the candidate text units, the context feature information and the feature information of the candidate text units can also be normalized using softmax, and softmax is a logistic regression algorithm.

[0180] Optionally, in this embodiment, the target text may include multiple target text units that meet the conditions: the number of characters is not greater than a preset number, or the language type is a non-target language. That is to say, there are multiple target text units in the target text that need to obtain similar text units based on context information. At this time, the similar text units of these target text units can be obtained in the order in which they appear in the target text.

[0181] For example, if there are three target text units in the target text that need to obtain similar text units based on context information, and based on their order of appearance in the target text, these three target text units are denoted as a, b, and c respectively. Similar text units of a can be obtained first. Specifically, the context feature information of the target text unit a can be obtained, and then the similar text unit of a can be selected from the candidate text units according to this context feature information. Then, the similar text units of the target text unit b are obtained. Specifically, the similar text unit of the target text unit a in the target text can be replaced according to the similar text unit of the target text unit a, and the updated target text is obtained. Then, the context information of the target text unit in the updated target text is obtained, and the similar text unit of b is selected from the candidate text units based on this context information. And so on, the similar text units of the target text unit c are obtained.

[0182] In a specific embodiment, if the preset number is 1 and the target language is Chinese, for a single Chinese character or a target text unit that does not belong to Chinese (such as an English word), the similar text unit can be selected from the candidate text units according to the similarity between its context information and the candidate text units; for a target text unit of a Chinese word (with 2 or more characters), the similar text unit can be selected from the candidate text units according to the feature information of the target text unit itself and the feature information of the candidate text units. In this way, different methods are adopted to obtain similar text units according to the number of characters and language types of the target text unit, which can make the generated similar text smoother and of higher quality. This is because the meaning of Chinese words is relatively rich. For the same character, the words formed by combining it with different characters may have different meanings. Therefore, it needs to be regarded as a whole. And because the Chinese sentence structure is relatively complex, if all target text units are directly used to generate similar text units in one way and then the corresponding target text units in the target text are replaced, it is easy to cause problems such as grammar errors and sentence structure contradictions in the generated similar text.

[0183] It should be noted that the similar text unit of the target text unit cannot be the target text unit itself.

[0184] 105. Replace the target text units in the target text with similar text units based on the similar text units of each target text unit, and generate similar text that matches the target text.

[0185] Among them, the similar text is the text obtained after text augmentation of the target text.

[0186] Optionally, in this embodiment, for each target text unit, there may be at least one similar text unit. If there are multiple similar text units for a certain target text unit, one target similar text unit can be selected from these multiple similar text units to replace the target text unit in the target text based on the target similar text unit. Among them, the selection method of the target similar text unit can be random, or the similar text unit with the highest similarity to the target text unit can be used as the target similar text unit, and this embodiment does not limit this.

[0187] In a specific scenario, during the development of some short text classification models, if the text enhancement scheme of this embodiment is adopted for text categories with poor classification and recognition effects, the effect can be improved by about 3% in the test set. The experimental index comparison is as Figure 1f and Figure 1g shown. Figure 1f is the classification effect index of the test set before text enhancement, Figure 1g is the classification effect index of the test set after text enhancement.

[0188] Among them, f1 is used as a comprehensive evaluation index, and the higher the f1 value, the better the classification effect.

[0189] It can be seen from the comparison that for the vulgar category with a relatively low recognition accuracy rate before text enhancement, the accuracy rate has been greatly improved, and the classification performance of other categories has also been improved to varying degrees.

[0190] The text enhancement method provided by this application can replace the words in the input sentence without changing the overall sentiment of the sentence, so as to generate new sample data. In the data preprocessing stage, the text data can be filtered to remove invalid spaces, delimiters, etc., and then Chinese word segmentation is performed, and the TF-IDF values of each text unit after word segmentation are calculated. Three sequences to be replaced (specifically, sequences s1, s2, and s3 in the above embodiments) are generated according to the TF-IDF and the type of text unit. For the text units in sequences s1 and s3, a Chinese pre-trained Masked Language Model can be used to predict and replace similar text units. For the text units in sequence s2, a Chinese pre-trained Word2vec model is used to search for and replace similar text units. This application creatively integrates traditional language models and deep bidirectional language models, and proposes a new text enhancement method and system, which can effectively expand high-quality training samples from a small number of samples for NLP tasks, thereby improving the generalization performance of downstream task models.

[0191] Specifically, the text enhancement method of this embodiment has strong adaptability to the Chinese scenario. Aiming at the vocabulary-level text enhancement problem in the Chinese scenario, Chinese word segmentation and TF-IDF are combined to select words to be replaced (i.e., similar text units), and traditional statistical language models and deep bidirectional language models are integrated for vocabulary replacement, taking into account the diversity of the generated text and the smoothness of the sentences. New samples are generated without changing the original sentiment of the text, which can effectively expand the training data for downstream tasks in NLP.

[0192] As can be seen from the above, this embodiment can obtain the target text, and select at least one target text unit to be replaced from the text units of the target text; identify the number of characters and language type of the target text unit; when the number of characters of the target text unit is greater than the preset number and the language type of the target text unit is the target language, based on the similarity between the target text unit and the candidate text units, select the similar text unit of the target text unit from the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, select the similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units; based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text. This application can adopt different methods to obtain the similar text units of the target text unit based on the number of characters and language type of the text unit, and then replace the target text unit, which is beneficial to taking into account the diversity of the generated text and the smoothness of the sentences.

[0193] According to the method described in the previous embodiments, the following will further elaborate in detail by taking the specific integration of the text enhancement device in the server as an example.

[0194] An embodiment of the present application provides a text enhancement method. As Figure 2 shown, the specific process of this text enhancement method can be as follows:

[0195] 201. The server obtains the target text.

[0196] Among them, the target text is the text to be data-enhanced, which includes at least one text unit. The text unit can be a word or a character. Specifically, in this embodiment, a similar text semantically similar to the target text needs to be generated.

[0197] Among them, the target text can include at least one language type, and this embodiment does not limit the language type.

[0198] 202. The server performs frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text.

[0199] Optionally, in this embodiment, the step of "performing frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text" may include:

[0200] For each text unit in the target text, count the frequency of the text unit appearing in the target text to obtain the weight of the text unit in the target text;

[0201] Count the frequency of the text unit appearing in the sample text to obtain the reference weight of the text unit;

[0202] Determine the importance parameter of the text unit according to the reference weight of the text unit and the weight in the target text.

[0203] Optionally, the step of "determining the importance parameter of the text unit according to the reference weight of the text unit and the weight in the target text" may include:

[0204] Fuse the reference weight of the text unit and the weight of the text unit in the target text to obtain the importance parameter of the text unit.

[0205] Among them, there are many fusion methods. For example, the fusion method can be multiplication, etc., and this embodiment does not limit this. Specifically, the importance parameter of the text unit can be obtained by multiplying the term frequency TF and the inverse document frequency IDF of the text unit, and this importance parameter can be represented by TF-IDF.

[0206] 203. The server selects at least one target text unit to be replaced from each text unit of the target text based on the importance parameter.

[0207] Optionally, in this embodiment, a text unit with an importance parameter greater than a preset value can be selected as the target text unit to be replaced, and the preset value can be set according to the actual situation; alternatively, based on the importance parameter, each text unit of the target text can be sorted, for example, sorted from large to small, to obtain the sorted text units, and the first n text units of the sorted text units can be used as the target text units to be replaced.

[0208] 204. The server identifies the number of characters and the language type of the target text unit.

[0209] Among them, the target text unit can be one character, or two or more characters. In this embodiment, different methods can be used to obtain the similar text unit of the target text unit according to the number of characters of the target text unit.

[0210] For example, if the target text unit is the Chinese character "flowing water", which contains two characters, feature extraction can be performed on the target text unit to obtain the feature information of the target text unit, and based on the feature information of the target text unit and the feature information of the candidate text unit, the similar text unit of the target text unit is selected from the candidate text units.

[0211] For another example, if the target text unit is the Chinese character "such as", which contains one character, the context information of the target text unit can be obtained, and based on the similarity between the context information and the candidate text unit, the similar text unit of the target text unit is selected from the candidate text units.

[0212] In addition, the target text may contain only one language type, or may contain two or more language types. For example, the target text can be short text data mixed with Chinese. Therefore, the target text units divided from the target text may belong to language type 1 or may belong to language type 2.

[0213] 205. When the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, the server selects the similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units.

[0214] Optionally, in this embodiment, the step of "selecting the similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units" may include:

[0215] Feature extraction is respectively performed on the target text unit and the candidate text units to obtain the feature information of the target text unit and the feature information of the candidate text units;

[0216] Based on the feature information of the target text unit and the feature information of the candidate text units, the similarity between the target text unit and the candidate text units is calculated;

[0217] According to the similarity, similar text units of the target text unit are selected from the candidate text units.

[0218] Among them, specifically, a semantic model can be used to perform feature extraction on the target text unit and the candidate text units, and this feature extraction can include convolutional processing and pooling processing. Optionally, the semantic model can be a Word2vec (word to vector) model, a fastText (fast text) model, a Glove model, etc. Glove is a word embedding tool.

[0219] 206. When the number of characters of the target text unit is not greater than a preset number, or the language type of the target text unit is a non-target language, the server selects similar text units of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units.

[0220] Optionally, in this embodiment, the step of "selecting similar text units of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units" may include:

[0221] Obtain the feature information of the reference text unit in the target text, where the reference text unit is other text units in the target text except the target text unit;

[0222] According to the feature information of the reference text unit, perform feature extraction on the target text unit to obtain the context feature information corresponding to the context information of the target text unit;

[0223] Select similar text units of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text units.

[0224] Among them, the context feature information corresponding to the context information of the target text can be obtained through a semantic model, and then, according to the similarity between the context feature information and the feature information of the candidate text units, the similar text units of the target text unit are selected from the candidate text units. The semantic model can be an MLM model. Specifically, the semantic model can be BERT (Bidirectional Encoder Representations from Transformers), ALBERT (A Lite BERT), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), etc., and this embodiment does not limit this.

[0225] Optionally, in this embodiment, the step of "extracting the features of the target text unit according to the feature information of the reference text unit to obtain the context feature information corresponding to the context information of the target text unit" may include:

[0226] Obtain the weights corresponding to each reference text unit;

[0227] According to the weights, fuse the feature information of each reference text unit to obtain the context feature information corresponding to the context information of the target text unit.

[0228] Among them, there are various fusion methods, and this embodiment does not limit this. For example, it can be weighted fusion or feature splicing, etc.

[0229] 207. The server replaces the target text units in the target text based on the similar text units of each target text unit to generate a similar text that matches the target text.

[0230] Among them, the similar text is the text obtained after text augmentation of the target text.

[0231] As can be seen from the above, in this embodiment, the server can obtain the target text; perform frequency analysis on each text unit in the target text to determine the importance parameters of each text unit in the target text; based on the importance parameters, select at least one target text unit to be replaced from each text unit of the target text; identify the number of characters and language type of the target text unit. When the number of characters of the target text unit is greater than the preset number and the language type of the target text unit is the target language, the server selects a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units. When the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, the server selects a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units. The server replaces the target text units in the target text based on the similar text units of each target text unit to generate a similar text that matches the target text. This application can adopt different methods to obtain the similar text units of the target text unit based on the number of characters and language type of the text unit, and then replace the target text unit, which is beneficial to taking into account the diversity of the generated text and the smoothness of the sentence.

[0232] To better implement the above method, an embodiment of the present application further provides a text enhancement device, as Figure 3 shown. The text enhancement device may include an acquisition unit 301, an identification unit 302, a first selection unit 303, a second selection unit 304, and a replacement unit 305, as follows:

[0233] (1) Acquisition unit 301;

[0234] The acquisition unit 301 is configured to acquire a target text and select at least one target text unit to be replaced from the text units of the target text.

[0235] Optionally, in some embodiments of the present application, the acquisition unit may include a frequency analysis subunit and a first selection subunit, as follows:

[0236] The frequency analysis subunit is configured to perform frequency analysis on each text unit in the target text to determine the importance parameters of each text unit in the target text;

[0237] The first selection subunit is configured to select at least one target text unit to be replaced from each text unit of the target text based on the importance parameters.

[0238] Optionally, in some embodiments of the present application, the frequency analysis subunit may specifically be configured to: for each text unit in the target text, count the frequency of occurrence of the text unit in the target text to obtain the weight of the text unit in the target text; count the frequency of occurrence of the text unit in the sample text to obtain the reference weight of the text unit; and determine the importance parameter of the text unit according to the reference weight of the text unit and the weight in the target text.

[0239] Optionally, in some embodiments of the present application, the text enhancement device may further include a partitioning unit, as follows:

[0240] The partitioning unit is configured to partition the target text through a word segmentation model to obtain each text unit in the target text, and the word segmentation model is trained based on the sample text and the position tag sequence corresponding to the text unit in the sample text.

[0241] (2) Identification unit 302;

[0242] The identification unit 302 is configured to identify the number of characters and the language type of the target text unit.

[0243] (3) First selection unit 303;

[0244] The first selection unit 303 is configured to, when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units.

[0245] Optionally, in some embodiments of the present application, the first selection unit may include a first addition subunit, a second selection subunit, and a third selection subunit, as follows:

[0246] The first addition subunit is configured to add a target text unit whose number of characters is greater than a preset number and whose language type is the target language to the first text sequence;

[0247] The second selection subunit is configured to select a target replacement text unit from the first text sequence according to a preset replacement ratio;

[0248] The third selection subunit is configured to select a similar text unit of the target replacement text unit from the candidate text units based on the similarity between the target replacement text unit and the candidate text units.

[0249] Optionally, in some embodiments of the present application, the first selection unit may include an extraction subunit, a calculation subunit, and a sixth selection subunit, as follows:

[0250] The extraction subunit is configured to perform feature extraction on the target text unit and the candidate text unit respectively, to obtain the feature information of the target text unit and the feature information of the candidate text unit;

[0251] The calculation subunit is configured to calculate the similarity between the target text unit and the candidate text unit based on the feature information of the target text unit and the feature information of the candidate text unit;

[0252] The sixth selection subunit is configured to select a similar text unit of the target text unit from the candidate text units according to the similarity.

[0253] (4) The second selection unit 304;

[0254] The second selection unit 304 is configured to, when the number of characters of the target text unit is not greater than a preset number, or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units.

[0255] Optionally, in some embodiments of the present application, the second selection unit may include a second addition subunit, a fourth selection subunit, and a fifth selection subunit, as follows:

[0256] The second addition subunit is configured to add a target text unit with the number of characters not greater than a preset number, or a language type of a non-target language, to the second text sequence;

[0257] The fourth selection subunit is configured to select a target replacement text unit from the second text sequence based on a preset replacement ratio;

[0258] The fifth selection subunit is configured to select a similar text unit of the target replacement text unit from the candidate text units according to the similarity between the context information of the target replacement text unit and the candidate text units.

[0259] Optionally, in some embodiments of the present application, the second selection unit may include an acquisition subunit, a context extraction subunit, and a seventh selection subunit, as follows:

[0260] The acquisition subunit is configured to acquire the feature information of a reference text unit in the target text, where the reference text unit is other text units in the target text except the target text unit;

[0261] A context extraction subunit, configured to extract features of the target text unit according to the feature information of the reference text unit, so as to obtain context feature information corresponding to the context information of the target text unit;

[0262] A seventh selection subunit, configured to select a similar text unit of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text units.

[0263] Optionally, in some embodiments of the present application, the context extraction subunit may specifically be configured to obtain weights corresponding to each reference text unit; and fuse the feature information of each reference text unit according to the weights, so as to obtain context feature information corresponding to the context information of the target text unit.

[0264] (5) Replacement unit 305;

[0265] The replacement unit 305 is configured to replace the target text units in the target text based on the similar text units of each target text unit, so as to generate a similar text that matches the target text.

[0266] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire a target text, and select at least one target text unit to be replaced from the text units of the target text; the recognition unit 302 can recognize the number of characters and the language type of the target text unit; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, the first selection unit 303 can select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, the second selection unit 304 can select a similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units; the replacement unit 305 can replace the target text units in the target text based on the similar text units of each target text unit, so as to generate a similar text that matches the target text. The present application can adopt different methods to obtain similar text units of target text units based on the number of characters and the language type of text units, and then replace the target text units, which is beneficial to taking into account the diversity of the generated text and the smoothness of sentences.

[0267] The embodiment of the present application further provides an electronic device, as Figure 4 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0268] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 of one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 4 the structure of the electronic device shown in

[0269] does not limit the electronic device. It may include more or fewer components than shown, or combine certain components, or have a different component arrangement. Among them:

[0270] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by invoking data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0271] The electronic device further includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0272] The electronic device may further include an input unit 404, which may be configured to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0273] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated herein. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0274] Obtain a target text, and select at least one target text unit to be replaced from the text units of the target text; identify the number of characters and the language type of the target text unit; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units; based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text.

[0275] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.

[0276] As described above, in this embodiment, the target text can be obtained, and at least one target text unit to be replaced can be selected from the text units of the target text; the number of characters and the language type of the target text unit can be identified; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, based on the similarity between the target text unit and the candidate text units, a similar text unit of the target text unit can be selected from the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, based on the context information of the target text unit and the similarity between the candidate text units, a similar text unit of the target text unit can be selected from the candidate text units; based on the similar text units of each target text unit, the target text units in the target text are replaced to generate a similar text that matches the target text. The present application can adopt different methods to obtain the similar text units of the target text units based on the number of characters and the language type of the text units, and then replace the target text units, which is beneficial to taking into account the diversity of the generated text and the smoothness of the sentences.

[0277] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0278] For this reason, an embodiment of the present application provides a storage medium in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps in any one of the text enhancement methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0279] Obtain the target text, and select at least one target text unit to be replaced from the text units of the target text; identify the number of characters and the language type of the target text unit; when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, based on the similarity between the target text unit and the candidate text units, select a similar text unit of the target text unit from the candidate text units; when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, based on the context information of the target text unit and the similarity between the candidate text units, select a similar text unit of the target text unit from the candidate text units; based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text.

[0280] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0281] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0282] Since the instructions stored in the storage medium can execute the steps in any of the text enhancement methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the text enhancement methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated here.

[0283] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the above text enhancement aspect.

[0284] The above has introduced in detail a text enhancement method, device, electronic device and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A text enhancement method, characterized in that, Including: Obtain the target text, and select at least one target text unit to be replaced from the text units of the target text; Identify the number of characters and the language type of the target text unit; When the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units; When the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units; Based on the similar text units of each target text unit, replace the target text units in the target text to generate a similar text that matches the target text.

2. The method according to claim 1, wherein The step of selecting at least one target text unit to be replaced from the text units of the target text includes: Perform frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text; Based on the importance parameter, select at least one target text unit to be replaced from the text units of the target text.

3. The method according to claim 2, wherein The step of performing frequency analysis on each text unit in the target text to determine the importance parameter of each text unit in the target text includes: For each text unit in the target text, count the frequency of occurrence of the text unit in the target text to obtain the weight of the text unit in the target text; Count the frequency of occurrence of the text unit in the sample text to obtain the reference weight of the text unit; Determine the importance parameter of the text unit according to the reference weight of the text unit and the weight in the target text.

4. The method according to claim 1, wherein The step of when the number of characters of the target text unit is greater than a preset number and the language type of the target text unit is the target language, selecting a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units includes: Add the target text units with the number of characters greater than the preset number and the language type being the target language to the first text sequence; Select target replacement text units from the first text sequence according to a preset replacement ratio; Based on the similarity between the target replacement text units and the candidate text units, select the similar text units of the target replacement text units from the candidate text units.

5. The method according to claim 1, characterized in that The step of when the number of characters of the target text unit is not greater than the preset number, or the language type of the target text unit is a non-target language, selecting a similar text unit of the target text unit from the candidate text units according to the context information of the target text unit and the similarity between the candidate text units includes: Add the target text units with the number of characters not greater than the preset number, or the language type being a non-target language to the second text sequence; Select a target replacement text unit from the second text sequence based on a preset replacement ratio; Select a similar text unit of the target replacement text unit from the candidate text units according to the similarity between the context information of the target replacement text unit and the candidate text units.

6. The method according to claim 1, characterized in that, The step of selecting a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units includes: Extract features from the target text unit and the candidate text units respectively to obtain the feature information of the target text unit and the feature information of the candidate text units; Calculate the similarity between the target text unit and the candidate text units based on the feature information of the target text unit and the feature information of the candidate text units; Select a similar text unit of the target text unit from the candidate text units according to the similarity.

7. The method according to claim 1, wherein The step of selecting a similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units includes: Obtain the feature information of a reference text unit in the target text, where the reference text unit is other text units in the target text except the target text unit; Obtain the weight corresponding to each reference text unit; Fuse the feature information of each reference text unit according to the weight to obtain the context feature information corresponding to the context information of the target text unit; Select a similar text unit of the target text unit from the candidate text units according to the similarity between the context feature information and the feature information of the candidate text units.

8. The method according to claim 1, wherein Before selecting at least one target text unit to be replaced from the text units of the target text, it further includes: Divide the target text through a word segmentation model to obtain each text unit in the target text, and the word segmentation model is trained based on a sample text and the position tag sequence corresponding to the text units in the sample text.

9. A text enhancement device, characterized in that, It includes: An acquisition unit, configured to acquire a target text and select at least one target text unit to be replaced from the text units of the target text; An identification unit, configured to identify the character number and language type of the target text unit; A first selection unit, configured to, when the character number of the target text unit is greater than a preset number and the language type of the target text unit is the target language, select a similar text unit of the target text unit from the candidate text units based on the similarity between the target text unit and the candidate text units; A second selection unit, configured to, when the character number of the target text unit is not greater than the preset number or the language type of the target text unit is a non-target language, select a similar text unit of the target text unit from the candidate text units according to the similarity between the context information of the target text unit and the candidate text units; A replacement unit, configured to replace a target text unit in the target text based on similar text units of each target text unit, so as to generate a similar text matching the target text.

10. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the operations in the text enhancement method according to any one of claims 1 to 8.

11. A storage medium, characterized in that, The storage medium stores multiple instructions, and the instructions are adapted to be loaded by a processor to execute the steps in the text enhancement method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Chinese error correction method and device, storage medium and electronic device

    CN110852087A

  • Text error correction method, system and device and readable storage medium

    CN112016310A