Text processing method, text classification method, device, equipment and storage medium

By extracting and splicing subtexts and preset characters from long texts, forming target splicing texts, and training the language model, solving the problems of high training costs and low classification accuracy in the existing technology, and achieving higher model performance and text classification accuracy.

CN114186060BActive Publication Date: 2025-06-17BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111449196.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-17
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

When training language models using long text, the prior art has high manual maintenance costs, the training language model performance is not ideal, and the accuracy of text classification is low.

Method used

By extracting the first subtext of the preset length from the to-process text and including the preset characters in the second subtext, splicing the preset characters and multiple characters in the first subtext, the target splicing text of the preset length is obtained, and the language model is trained.

Benefits of technology

The problem that the number of long text words does not meet the requirements of the language model is solved, the model's performance and recall accuracy are improved, and the accuracy of classifying texts is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114186060B_ABST
    Figure CN114186060B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a text processing method, a text classification method, an apparatus, a device, and a storage medium. The text classification method includes: obtaining a text to be processed; when the length of the text to be processed is greater than a preset length, extracting a first sub-text of the preset length from the text to be processed; when a second sub-text includes a preset character, splicing the preset character and multiple characters in the first sub-text to obtain a first target spliced text of the preset length; wherein, the second sub-text is the text in the text to be processed other than the first sub-text. The present disclosure not only solves the problem that the number of words in a long text does not meet the requirements of a language model, but also can intercept the first target spliced text representing the key characters of the core content of the text and the theme names to be monitored from the long text for model training, thereby improving the performance of the model and enabling the trained language model to have a higher accuracy rate when classifying texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a text processing method, a text classification method, an apparatus, a device, and a storage medium. Background Art

[0002] Due to the openness and dissemination characteristics of the Internet, it is necessary to monitor online public opinion and obtain an online public opinion analysis report. An online public opinion analysis platform generally obtains various comments, articles, news, etc. from the Internet, and then classifies the text such as the comments and articles. Since most of the texts on the Internet are long texts with a large number of words, and current machine learning algorithms are limited by machine memory and hardware configuration and cannot train all the content of the long text to obtain a classification model. Therefore, when inputting a long text into a language model for training and classification, it is often necessary to preprocess the long text to meet the requirements of the language model.

[0003] In related technologies, when training a language model with a long text, either the manual maintenance cost is relatively high, or the performance of the trained language model is not yet ideal, and the accuracy of text classification is relatively low. Therefore, it is necessary to improve the processing method of long texts and the text classification method so that they are applicable to some language models with better effects and improve the accuracy of text classification. Summary of the Invention

[0004] The present disclosure provides a text processing method, a text classification method, an apparatus, a device, and a storage medium to at least solve the problem that when training a language model with a long text in related technologies, either the manual maintenance cost is relatively high, or the performance of the trained language model is not yet ideal, and the accuracy of text classification is relatively low. The technical solution of the present disclosure is as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a text processing method is provided, including:

[0006] Obtaining a text to be processed;

[0007] In a case where the length of the text to be processed is greater than a preset length, extracting a first sub-text of the preset length from the text to be processed;

[0008] In a case where a second sub-text includes a preset character, splicing the preset character and multiple characters in the first sub-text to obtain a first target splicing text of the preset length;

[0009] wherein the second sub-text is the text in the text to be processed except the first sub-text.

[0010] In an exemplary embodiment, splicing the preset character and a plurality of characters in the first sub-text to obtain the first target splicing text with the preset length includes:

[0011] Extracting a preset number of characters from the first sub-text; the number of characters in the first sub-text except the preset number of characters is equal to the number of characters included in the preset character;

[0012] Splicing the preset number of characters and the preset character to obtain the first target splicing text.

[0013] In an exemplary embodiment, when the length of the text to be processed is greater than the preset length, extracting the first sub-text with the preset length from the text to be processed includes:

[0014] Taking the first character of the text to be processed as the starting position and extracting a first number of characters in the direction of the next character to obtain the first segment of text;

[0015] Taking the last character of the text to be processed as the ending position and extracting a second number of characters in the direction of the previous character to obtain the last segment of text;

[0016] Using the first segment of text and the last segment of text as the first sub-text; the difference between the first number and the second number is less than the preset number threshold.

[0017] In an exemplary embodiment, the preset character includes a third number of characters, and extracting a preset number of characters from the first sub-text includes:

[0018] Taking the first character of the first segment of text as the starting position and extracting a fourth number of characters in the direction of the next character;

[0019] Taking the last character of the last segment of text as the ending position and extracting a fifth number of characters in the direction of the previous character;

[0020] Using the fourth number of characters and the fifth number of characters as the preset number of characters;

[0021] Wherein, the fourth number is less than the first number, the fifth number is less than the second number, the sum of the first difference and the second difference is equal to the third number, the first difference represents the difference between the first number and the fourth number, and the second difference represents the difference between the second number and the fifth number.

[0022] In an exemplary embodiment, splicing the preset number of characters and the preset character to obtain the first target splicing text includes:

[0023] Concatenate the fourth quantity of characters, the preset character, and the fifth quantity of characters to obtain the first target concatenated text.

[0024] In an exemplary embodiment, the method further includes:

[0025] When the preset character is included in the second sub-text, determine the standard character corresponding to the preset character;

[0026] Concatenate the standard character and the plurality of characters to obtain the first target concatenated text.

[0027] In an exemplary embodiment, the determining the standard character corresponding to the preset character includes:

[0028] Perform word segmentation processing on the preset character to obtain a word segmentation result corresponding to the preset character;

[0029] Determine a target proper noun matching the word segmentation result from a preset word list; the preset word list stores a plurality of proper nouns through a double-array tree structure;

[0030] Based on preset mapping information, determine the standard character corresponding to the target proper noun; the preset mapping information represents the mapping relationship between proper nouns and standard characters.

[0031] In an exemplary embodiment, the double-array tree structure includes a root node and at least one leaf node. A proper noun is stored in the path between the root node and each leaf node. The number of the word segmentation results is multiple. The determining the target proper noun matching the word segmentation result from the preset word list includes:

[0032] Combine at least two adjacent word segmentation results to obtain a combined word segmentation result;

[0033] When the proper noun stored in the path between the root node and one of the leaf nodes matches the combined word segmentation result, use the proper noun stored in the path between the root node and one of the leaf nodes as the target proper noun matching the at least two adjacent word segmentation results; the one leaf node is a node among the at least one leaf node.

[0034] According to a second aspect of the embodiments of the present disclosure, a text classification method is provided, including:

[0035] Obtain a text to be classified;

[0036] When the length of the text to be classified is greater than a preset length, extract the first sub-text of the preset length from the text to be classified;

[0037] When the second sub - text includes a preset character, splice the preset character and multiple characters in the first sub - text to obtain a second target spliced text of the preset length; the second sub - text is the text in the text to be classified except the first sub - text.

[0038] Classify the second target spliced text through a preset language model; the preset language model is trained based on the first target spliced text in any of the above - mentioned embodiments.

[0039] According to the third aspect of the embodiments of the present disclosure, there is provided a text processing device, including:

[0040] A text - to - be - processed acquisition module, configured to acquire a text to be processed;

[0041] A first extraction module, configured to extract the first sub - text of the preset length from the text to be processed when the length of the text to be processed is greater than the preset length;

[0042] A first splicing module, configured to splice the preset character and multiple characters in the first sub - text to obtain a first target spliced text of the preset length when the second sub - text includes a preset character; wherein, the second sub - text is the text in the text to be processed except the first sub - text.

[0043] In an exemplary embodiment, the first splicing module includes:

[0044] A preset number of character extraction units, configured to extract a preset number of characters from the first sub - text; the number of characters in the first sub - text except the preset number of characters is equal to the number of characters included in the preset character.

[0045] A first target spliced text determination unit, configured to splice the preset number of characters and the preset character to obtain the first target spliced text.

[0046] In an exemplary embodiment, the first extraction module includes:

[0047] A first - paragraph text extraction unit, configured to extract a first number of characters in the direction of the next character starting from the first character of the text to be processed to obtain a first - paragraph text;

[0048] A last - paragraph text extraction unit, configured to extract a second number of characters in the direction of the previous character with the last character of the text to be processed as the termination position to obtain a last - paragraph text.

[0049] The first sub - text determination unit is configured to execute using the first - segment text and the last - segment text as the first sub - text; the difference between the first quantity and the second quantity is less than a preset quantity threshold.

[0050] In an exemplary embodiment, the preset characters include a third quantity of characters, and the preset - quantity character extraction unit includes:

[0051] A fourth - quantity character extraction sub - unit is configured to execute extracting a fourth quantity of characters in the direction of one character backward starting from the first character of the first - segment text.

[0052] A fifth - quantity character extraction sub - unit is configured to execute extracting a fifth quantity of characters in the direction of one character forward starting from the last character of the last - segment text.

[0053] A preset - quantity character determination sub - unit is configured to execute using the fourth quantity of characters and the fifth quantity of characters as the preset quantity of characters; wherein, the fourth quantity is less than the first quantity, the fifth quantity is less than the second quantity, the sum of the first difference and the second difference is equal to the third quantity, the first difference represents the difference between the first quantity and the fourth quantity, and the second difference represents the difference between the second quantity and the fifth quantity.

[0054] In an exemplary embodiment, the first target splicing - text determination unit is configured to execute splicing the fourth quantity of characters, the preset characters, and the fifth quantity of characters to obtain the first target splicing text.

[0055] In an exemplary embodiment, the device further includes:

[0056] A standard - character determination module is configured to execute determining the standard character corresponding to the preset character when the preset character is included in the second sub - text.

[0057] A second splicing module is configured to execute splicing the standard character and the plurality of characters to obtain the first target splicing text.

[0058] In an exemplary embodiment, the standard - character determination module includes:

[0059] A word - segmentation result determination unit is configured to execute performing word - segmentation processing on the preset character to obtain the word - segmentation result corresponding to the preset character.

[0060] A target proper noun determination unit, configured to determine a target proper noun that matches the word segmentation result from a preset vocabulary; the preset vocabulary stores multiple proper nouns through a double-array tree structure;

[0061] A standard character determination unit, configured to determine a standard character corresponding to the target proper noun based on preset mapping information; the preset mapping information represents the mapping relationship between proper nouns and standard characters.

[0062] In an exemplary embodiment, the double-array tree structure includes a root node and at least one leaf node, a proper noun is stored in the path between the root node and each leaf node, the number of word segmentation results is multiple, and the target proper noun determination unit includes:

[0063] A combined word segmentation result determination subunit, configured to combine at least two adjacent word segmentation results to obtain a combined word segmentation result;

[0064] A target proper noun determination subunit, configured to, when the proper noun stored in the path between the root node and one of the leaf nodes matches the combined word segmentation result, use the proper noun stored in the path between the root node and one of the leaf nodes as the target proper noun that matches the at least two adjacent word segmentation results; the one leaf node is a node among the at least one leaf node.

[0065] According to a fourth aspect of the embodiments of the present disclosure, a text classification device is provided, including:

[0066] A text to be classified acquisition module, configured to acquire a text to be classified;

[0067] A second extraction module, configured to, when the length of the text to be classified is greater than a preset length, extract a first sub-text of the preset length from the text to be classified;

[0068] A third splicing module, configured to, when the second sub-text includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain a second target spliced text of the preset length; the second sub-text is the text in the text to be classified other than the first sub-text;

[0069] A classification module, configured to classify the second target spliced text through a preset language model; the preset language model is trained based on the first target spliced text in any of the above embodiments.

[0070] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0071] A processor;

[0072] A memory for storing executable instructions of the processor;

[0073] Wherein, the processor is configured to execute the instructions to implement the text processing method described in any of the above embodiments or the text classification method described in any of the above embodiments.

[0074] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is caused to execute the text processing method described in any of the above embodiments or the text classification method described in any of the above embodiments.

[0075] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor implements the text processing method described in any of the above embodiments or the text classification method described in any of the above embodiments.

[0076] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0077] After obtaining the text to be processed, it is possible to determine whether the length of the text to be processed is greater than a preset length. If it is greater, a first sub-text of the preset length representing the core content of the text to be processed is extracted from the text to be processed, and it is determined whether a second sub-text other than the first sub-text in the text to be processed includes a preset character corresponding to the subject name to be monitored. If it includes, the preset character and multiple characters in the first sub-text are spliced to obtain a first target spliced text of the preset length. Training the language model with the first target spliced text of the preset length not only solves the problem that the number of words in the long text does not meet the requirements of the language model, but also can intercept the key characters representing the core content of the text and the first target spliced text of the subject name to be monitored from the long text for model training, thereby improving the performance and recall accuracy of the model, and making the trained language model have higher accuracy when classifying texts.

[0078] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0080] Figure 1The following is an application environment diagram of a text processing method shown according to an exemplary embodiment.

[0081] Figure 2 The following is a flowchart of a text processing method shown according to an exemplary embodiment.

[0082] Figure 3 The following is a flowchart of extracting a first sub - text of a preset length from a text to be processed according to an exemplary embodiment.

[0083] Figure 4 The following is a flowchart of concatenating a preset character and multiple characters in the first sub - text to obtain a first target concatenated text of a preset length according to an exemplary embodiment.

[0084] Figure 5 The following is a flowchart of extracting a preset number of characters from the first sub - text according to an exemplary embodiment.

[0085] Figure 6 The following is a flowchart of obtaining the first target concatenated text according to an exemplary embodiment.

[0086] Figure 7 The following is a flowchart of determining a standard character corresponding to the above - mentioned preset character according to an exemplary embodiment.

[0087] Figure 8 The following is a schematic diagram of a preset word list according to an exemplary embodiment.

[0088] Figure 9 The following is a flowchart of determining a target proper noun matching the word - segmentation result from a preset word list according to an exemplary embodiment.

[0089] Figure 10 The following is a flowchart of a text classification method according to an exemplary embodiment.

[0090] Figure 11 The following is a block diagram of a text processing device according to an exemplary embodiment.

[0091] Figure 12 The following is a block diagram of a text classification device according to an exemplary embodiment.

[0092] Figure 13 The following is a block diagram of an electronic device for text processing according to an exemplary embodiment. Detailed implementation manners

[0093] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0094] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0095] Since there are many long texts on the network, with a large number of characters, such as Weibo, many of which exceed 3000 characters. Due to the limitations of machine memory and hardware configuration in current machine learning algorithms, many algorithm models have limitations on the number of characters in the text, so they cannot learn all the content of long texts, resulting in the inability to learn some important information. For example, the ROBERTA (A Robustly Optimized BERT Pretraining Approach) model is a relatively good language understanding model, but it has a limitation on the number of characters in the text and can only support up to 512 characters at most. Therefore, it cannot give full play to the advantages of the ROBERTA model in training long texts.

[0096] Based on this, first, the embodiments of the present disclosure provide a text preprocessing method. By intercepting a relatively core first sub-text from the text to be processed, and when the second sub-text other than the first sub-text in the text to be processed includes a preset character, splicing the preset character and multiple texts in the first sub-text, and then training the language model, it can effectively solve the problem that the language model cannot train long texts due to the limitation of the number of characters, and can improve the training effect by intercepting the core content of the text according to the structural characteristics of the text and people's logical expression habits.

[0097] Please refer to Figure 1 , Figure 1 FIG. shows an application environment diagram of a text processing method according to an exemplary embodiment. The application environment may include a client 01 and a server 02. The client 01 can communicate with the server 02 in a wired or wireless manner, and the present disclosure does not make any limitations on this.

[0098] Among them, the client 01 can collect the text to be processed input by the user and send the text to be processed to the server 02. Optionally, the client 01 may include terminal devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices.

[0099] Optionally, the text to be processed input by the user may be collected according to user information. It should be noted that the user information involved in this disclosure (including but not limited to user device information, user personal information, etc.) is information that has been authorized by the user or fully authorized by all parties.

[0100] Among them, the server 02 can be used to obtain the text to be processed collected by the client 01, and when the length of the text to be processed is greater than a preset length, extract the first sub-text of the preset length from the text to be processed, and when the second sub-text other than the first sub-text in the text to be processed includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain a first target spliced text of the preset length. Optionally, the server 02 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0101] It should be noted that Figure 1 is only an example. In another exemplary embodiment, the text processing method provided by the embodiments of the present disclosure may also be applied to an application environment that only includes a client. Among them, the client may include terminal devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. After obtaining the text to be processed, the client may, when the length of the text to be processed is greater than a preset length, extract the first sub-text of the preset length from the text to be processed, and when the second sub-text other than the first sub-text in the text to be processed includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain the first target spliced text of the preset length described above.

[0102] Figure 2 is a flowchart of a text processing method shown according to an exemplary embodiment, asFigure 1 As shown, this method can be used for Figure 1 in an application environment including a client and a server, and includes the following steps.

[0103] In step S11, obtain the text to be processed.

[0104] The text to be processed in the embodiments of the present disclosure can be various texts for training a preset language model, such as various comments, articles, news, etc. on the network. The text to be processed can be texts in various languages such as Chinese text, English text, etc.

[0105] Exemplarily, the preset language model can include but is not limited to: ROBERTA model, BERT (Bidirectional Encoder Representation from Transformers) model, etc.

[0106] In some embodiments, after obtaining the text to be processed, subsequent steps can be directly executed.

[0107] In other embodiments, since the text to be processed usually contains some special characters such as numbers, letters, punctuation marks, emojis, space bars, etc., these special characters usually have no substantial meaning or have little impact on the meaning of the text to be processed, but instead increase the text length. Therefore, after obtaining the text to be processed, the special characters in the text to be processed can be deleted first, and after deleting these special characters, subsequent steps can be performed on the text to be processed to reduce the redundant information in the text to be processed, so that the trained language model has a higher accuracy rate when classifying the text.

[0108] In step S13, when the length of the above-mentioned text to be processed is greater than the preset length, a first sub-text of the preset length is extracted from the above-mentioned text to be processed.

[0109] In the embodiments of the present disclosure, a preset function for calculating the text length can be used to count the text length at the character level. And it is judged whether the text length is greater than the preset length.

[0110] Wherein, the preset length is determined based on the text length supported by the preset language model. Since the preset length is the reference length of the new text generated after truncating and splicing the text to be processed, therefore, the preset length can be less than or equal to the text length supported by the preset language model.

[0111] Taking the preset language model as the ROBERTA model as an example, the text length supported by the ROBERTA model is 512 character lengths. Then this preset length can be less than or equal to 512 character lengths. For example, it can be 512, or 256, or 64, etc. Specifically, it can be set according to the actual usage scenario, and a value with better training effect can be selected.

[0112] In an optional embodiment, in the above step S13, when the length of the above to-be-processed text is greater than the preset length, extracting the first sub-text of the above preset length from the above to-be-processed text may include:

[0113] When the length of the above to-be-processed text is greater than the preset length, using at least one specified character in the above to-be-processed text as a reference position, extracting the first sub-text of the above preset length from the above to-be-processed text.

[0114] In the embodiments of the present disclosure, if the to-be-processed text is longer than the preset length, the to-be-processed text can be truncated to generate a first sub-text that meets the requirements of the preset language model. Specifically, at least one specified character of the to-be-processed text can be used as a position reference, and the first sub-text is intercepted from the above to-be-processed text, and the length of the first sub-text is equal to the preset length. In order for the trained language model to achieve a better effect, the first sub-text can be the text representing the core idea of the to-be-processed text, that is, the relatively key content in the to-be-processed text, so that the language model can learn the key information of the text.

[0115] Among them, the specified character is used to determine the position of the core content of the to-be-processed text in the text. The specified character can be one character or multiple characters, and can be determined according to the distribution characteristics of the text structure and people's expression logic habits. For example, for Chinese, people's habitual expression method when writing articles is to adopt the article structures of "total - sub - total", "total - sub", or "sub - total". Therefore, for an article, its core idea and main arguments are very likely to be concentrated at the beginning or the end of the article. Thus, a paragraph of text at the beginning of the to-be-processed text, or a paragraph of text at the end of the text, or a paragraph of text is intercepted from both the beginning and the end can be input into the preset language model for training the model. Therefore, in some embodiments, the specified character can be one or more of the first character and the last character of the to-be-processed text.

[0116] In some scenarios, when people express their views in an article, they are accustomed to using some summarizing or generalizing words. For example, "Generally speaking", "To sum up", "In short", "Finally", etc. Therefore, the text following these words is also very likely to represent the core idea of the article. So, in a feasible embodiment, these text characters can also be used as a position reference to intercept a segment of text before or after this character from the text to be processed. Of course, there can be one or more specified characters, and the multiple intercepted characters can be a segment of text or multiple segments of text in the text to be processed.

[0117] Since in most cases, the core content of the article is concentrated at the beginning or end of the text to be processed, in a specific embodiment, the first character of the text to be processed can be used as the starting position, and the last character of the text to be processed can be used as the ending position to extract a first sub-text of a preset length from the text to be processed. Figure 3 is a flowchart showing an extraction of a first sub-text of a preset length from a text to be processed according to an exemplary embodiment. As Figure 3 shown, in the above step S13, in the case where the length of the above text to be processed is greater than the preset length, extracting the above first sub-text of the preset length from the above text to be processed may include:

[0118] In step S1301, taking the first character of the above text to be processed as the starting position, extracting a first number of characters in the direction of the next character to obtain a first segment of text.

[0119] In step S1303, taking the last character of the above text to be processed as the ending position, extracting a second number of characters in the direction of the previous character to obtain a last segment of text.

[0120] In step S1305, taking the above first segment of text and the above last segment of text as the above first sub-text; the difference between the above first number and the above second number is less than a preset number threshold.

[0121] Specifically, the first character of the text to be processed can be used as the starting position, and the first number of characters can be intercepted backward in the direction of the next character after the first character to obtain the first segment of text. Then, the last character of the text to be processed can be used as the ending position, and the second number of characters can be intercepted forward in the direction of the previous character of the last character to obtain the last segment of text. Then, the first segment of text and the last segment of text are used as the first sub-text, and the length of the first sub-text is equal to the preset length.

[0122] Optionally, the first quantity may be equal to the second quantity. For example, the number of characters in two segments of text intercepted from the beginning and the end of the text to be processed may be equal. In some embodiments, the first quantity and the second quantity may also be unequal, and the difference between the first quantity and the second quantity is less than a preset quantity threshold (which can be determined according to actual needs). For example, the number of characters intercepted from the beginning of the text to be processed is more, and the number of characters intercepted from the end of the text is less, or the number of characters intercepted from the beginning of the text is less, and the number of characters intercepted from the end of the text is more. For example, if the preset length is 128 characters, the first quantity and the second quantity may both be 64 characters, or one may be 60 characters and the other may be 68 characters.

[0123] In an operable embodiment, taking the ROBERTA model as an example of the preset language model, through a large number of experiments, it is found that when the preset length is 128 characters, starting from the first character of the text to be processed as the starting position, 64 consecutive characters are intercepted in the direction of the next character, and starting from the last character of the text to be processed as the termination position, 64 consecutive characters are intercepted in the direction of the previous character. Using the first sub-text intercepted in this way to train the ROBERTA model can greatly improve the performance of the model, and the accuracy of the trained model for text classification is also greatly improved.

[0124] In the embodiments of the present application, since in most cases, the core content of the article is concentrated at the beginning or the end of the text to be processed, thus in a specific embodiment, starting from the first character of the text to be processed as the starting position and the last character of the text to be processed as the termination position, a first sub-text of the preset length is extracted from the text to be processed, which can not only meet the requirements of the preset language model for the number of text characters, but also the first sub-text may contain the core content of the text to be processed, so that the preset language model trained through the first sub-text can learn the key information of the text to be processed, thereby improving the performance and recall accuracy of the model, and the accuracy of the trained model for text classification is also greatly improved.

[0125] In an alternative embodiment, if the length of the text to be processed is less than or equal to the preset length, the above truncation operation may not be performed on the text to be processed.

[0126] In another alternative embodiment, since the vector of the short text has fewer characters, many dimensions in its vector are filled with 0, that is, they have no actual meaning, resulting in an unsatisfactory final effect of the language model. If it is determined that the length of the text to be processed is less than the preset length, multiple characters can also be copied from the text to be processed so that the length of the text to be processed is equal to the preset length. For example, if the preset length is 128 characters and the text to be processed A has only 64 characters, the text to be processed A can be copied to obtain text A', and text A' and the text to be processed A are used as the first sub-text.

[0127] In step S15, when the second sub-text includes the preset character, the preset character and multiple characters in the first sub-text are spliced to obtain the first target spliced text of the preset length. Wherein, the second sub-text is the text in the text to be processed except the first sub-text.

[0128] Specifically, in the public opinion monitoring scenario, the preset character can be the name of the entity to be monitored in the public opinion monitoring scenario. Suppose it is necessary to monitor the positive and negative public opinions of a certain enterprise, then the entity name can be the enterprise name, the abbreviation of the enterprise name, the name of the main product corresponding to the enterprise, the abbreviation of the name of the main product corresponding to the enterprise, etc.

[0129] Optionally, the number of the preset characters can be one or multiple. For example, the preset character is the abbreviation of a certain enterprise, or the preset character is the abbreviation of a certain enterprise and the name of the product corresponding to the enterprise.

[0130] If the second sub-text includes the preset character, it indicates that the name of the entity to be monitored is not in the first sub-text that has been intercepted. In order to improve the performance of the model and make the trained language model have a higher accuracy rate when classifying the text, the preset character can be extracted from the second sub-text, and the preset character and multiple characters in the first preset text are spliced to obtain the first target spliced text.

[0131] Figure 4 It is a flowchart showing a method of splicing a preset character and multiple characters in a first sub-text to obtain a first target spliced text of a preset length. As Figure 4 shown, the splicing of the preset character and multiple characters in the first sub-text to obtain the first target spliced text of the preset length may include:

[0132] In step S1501, a preset number of characters are extracted from the first sub-text; the number of characters in the first sub-text except the preset number of characters is equal to the number of characters included in the preset character.

[0133] In step S1503, the above-mentioned preset number of characters and the above-mentioned preset character are concatenated to obtain the above-mentioned first target concatenated text.

[0134] Since the length of the finally concatenated first target concatenated text is the preset length, and the length of the first sub-text is also the preset length, if the preset character is directly concatenated with multiple characters in the first sub-text, the resulting first target concatenated text will be greater than the preset length. Based on this, a preset number of characters can be extracted from the first sub-text, and the characters in the first sub-text other than the above-mentioned preset number of characters are deleted (the number of characters in the first sub-text other than the above-mentioned preset number of characters is equal to the number of characters included in the above-mentioned preset character), and finally the preset number of characters and the above-mentioned preset character are concatenated to obtain the above-mentioned first target concatenated text.

[0135] In some embodiments, if the first sub-text is a text extracted from the text to be processed based on a specified character as a reference position, then some characters (the number of these characters is equal to the number of characters included in the above-mentioned preset character) can be deleted from the head, end or middle position of the first sub-text, so as to extract a preset number of characters.

[0136] In other embodiments, if the first sub-text is a text extracted from the text to be processed based on multiple specified characters as reference positions (for example, the first paragraph text in the above step S1301, the last paragraph text in the above step S1303), then some characters can be deleted from the first paragraph text, and some characters can be deleted from the last paragraph text (the sum of the number of characters deleted from the two paragraphs of text is equal to the number of characters included in the above-mentioned preset character), so as to extract a preset number of characters

[0137] Exemplarily, when concatenating the preset number of characters and the above-mentioned preset character, the concatenation can be performed in the order in which each segment of characters appears in the text. For example, the one that appears first is placed at the front, and the one that appears last is placed at the end. Of course, the concatenation can also be performed in the reverse order, or in a random order.

[0138] In the embodiments of the present disclosure, after deleting some characters from the first sub-text and then concatenating with the preset character, on the one hand, the resulting first target concatenated text can meet the requirements of the preset language model for the number of words in the text, and on the other hand, the first target concatenated text can contain both the core content of the text to be processed and the name of the subject to be monitored, so that the preset language model trained through the first target concatenated text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model for classifying the text.

[0139] Figure 5 is a flowchart showing the extraction of a preset number of characters from a first sub - text according to an exemplary embodiment. As Figure 5 shown, among the above - mentioned preset characters, there are a third number of characters. In the above step S1501, the extraction of a preset number of characters from the above - mentioned first sub - text may include:

[0140] In step S15011, starting from the first character of the above - mentioned first - paragraph text as the starting position, extract a fourth number of characters in the direction of the next character.

[0141] In step S15013, starting from the last character of the above - mentioned last - paragraph text as the ending position, extract a fifth number of characters in the direction of the previous character.

[0142] In step S15015, use the above - mentioned fourth number of characters and the above - mentioned fifth number of characters as the above - mentioned preset number of characters; wherein, the above - mentioned fourth number is less than the above - mentioned first number, the above - mentioned fifth number is less than the above - mentioned second number, the sum of the first difference and the second difference is equal to the above - mentioned third number, the above - mentioned first difference represents the difference between the above - mentioned first number and the above - mentioned fourth number, and the above - mentioned second difference represents the difference between the above - mentioned second number and the above - mentioned fifth number.

[0143] If the first sub - text includes the first - paragraph text and the last - paragraph text, then starting from the first character of the above - mentioned first - paragraph text as the starting position, intercept a fourth number of characters in the direction of the next character after the first character, and starting from the last character of the above - mentioned last - paragraph text as the ending position, intercept a fifth number of characters in the direction of the previous character before the last character, and use the fourth number of characters and the fifth number of characters as the above - mentioned preset number of characters. To ensure that the length of the finally spliced first target spliced text is the preset length, the sum of the first difference between the first number and the above - mentioned fourth number and the second difference between the second number and the above - mentioned fifth number is equal to the above - mentioned third number.

[0144] For example, if the first - paragraph text is 64 characters long and the last - paragraph text is 64 characters long, and the preset characters are 2 characters long, then the last character in the first - paragraph text can be deleted to obtain the fourth number of characters, and the first character in the last - paragraph text can be deleted to obtain the fifth number of characters.

[0145] Since in most cases, the core ideas of the text to be processed are concentrated at the beginning or the end of the text, deleting the last character of the first paragraph of the text and the first character of the last paragraph of the text can not only make the length of the first target spliced text obtained by the final splicing meet the requirements of the preset language model for the number of words in the text, but also avoid deleting the core keywords in the core ideas of the text to be processed, so that the preset language model trained by the first target spliced text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model in classifying the text.

[0146] In a feasible embodiment, splicing the above-mentioned preset number of characters and the above-mentioned preset character to obtain the above-mentioned first target spliced text may include:

[0147] Splice the above-mentioned fourth number of characters, the above-mentioned preset character, and the above-mentioned fifth number of characters to obtain the above-mentioned first target spliced text.

[0148] In the embodiments of the present disclosure, after obtaining the fourth number of characters and the fifth number of characters, the above-mentioned fourth number of characters, the above-mentioned preset character, and the above-mentioned fifth number of characters may be spliced to obtain the above-mentioned first target spliced text. On the one hand, the first target spliced text obtained by splicing can meet the requirements of the preset language model for the number of words in the text. On the other hand, the first target spliced text can contain both the core content of the text to be processed and the name of the subject to be monitored, so that the preset language model trained by the first target spliced text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model in classifying the text.

[0149] Exemplarily, when splicing the fourth number of characters, the fifth number of characters, and the preset character, the splicing can be performed in the order in which each paragraph of characters appears in the text. For example, the one that appears first is placed at the front, and the one that appears last is placed at the end. Of course, the splicing can also be performed in the reverse order, or in a random order.

[0150] In some embodiments, when the preset character is not included in the second sub - text, it indicates that the name of the subject to be monitored is already in the first sub - text. If the first sub - text is the text extracted from the text to be processed based on a specified character as the reference position, the first sub - text can be directly used as the first target concatenated text. If the first sub - text is the text extracted from the text to be processed based on multiple specified characters as the reference positions (for example, the first paragraph text in step S1301 and the last paragraph text in step S1303 above), the first paragraph text and the last paragraph text can be concatenated to obtain the first target concatenated text. The method of concatenating the first paragraph text and the last paragraph text to obtain the first target concatenated text can refer to the above steps S1501 - S1503 and will not be elaborated here.

[0151] Figure 6 is a flowchart of obtaining a first target concatenated text shown according to an exemplary embodiment. As Figure 6 shown, in an alternative embodiment, the above - mentioned method may further include:

[0152] In step S21, when the preset character is included in the second sub - text, determine the standard character corresponding to the preset character.

[0153] In step S23, concatenate the standard character and the multiple characters to obtain the first target concatenated text.

[0154] In some embodiments, the preset character existing in the second sub - text may not be the final standard character to be monitored. For example, if the subject to be monitored is an enterprise and the preset character is the product corresponding to the enterprise, when the preset character is included in the second sub - text, the standard character corresponding to the preset character can be determined first, and the standard character and the multiple characters can be concatenated to obtain the first target concatenated text.

[0155] In an alternative embodiment, the method of concatenating the standard character and the multiple characters may be as follows: Extract a preset number of characters from the multiple characters; the number of characters in the multiple characters except the preset number of characters is equal to the number of characters included in the standard character; concatenate the preset number of characters and the standard character to obtain the first target concatenated text.

[0156] In a specific embodiment, when the first sub - text includes the first - paragraph text and the above - mentioned last - paragraph text, the method of splicing the above - mentioned standard character and the above - mentioned multiple characters can be as follows: Assume that the preset character includes a sixth number of characters. Starting from the first character of the above - mentioned first - paragraph text as the starting position, extract a fourth number of characters in the direction of the next character; Taking the last character of the above - mentioned last - paragraph text as the ending position, extract a fifth number of characters in the direction of the previous character; Use the above - mentioned fourth number of characters and the above - mentioned fifth number of characters as the above - mentioned preset number of characters; The above - mentioned first difference represents the sum of the difference between the first number and the fourth number and the difference between the second number and the fifth number, which is equal to the sixth number. Splice the above - mentioned fourth number of characters, the above - mentioned standard character, and the above - mentioned fifth number of characters to obtain the above - mentioned first target spliced text.

[0157] It should be noted that multiple different types of preset characters may correspond to the same standard character. For example, when different types of preset characters are different products of the same enterprise, different products of the same enterprise can correspond to the same marked character (for example, the short name of the enterprise).

[0158] Since the standard character can more accurately reflect the subject to be monitored, by determining the standard character corresponding to the preset character and splicing the standard character with the preset character to obtain the first target spliced text, the first target spliced text can not only contain the core content of the text to be processed, but also accurately include the standard character corresponding to the subject to be monitored, so that the preset language model trained by the first target spliced text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model in classifying the text.

[0159] Figure 7 is a flowchart showing a method for determining the standard character corresponding to the above - mentioned preset character according to an exemplary embodiment. As Figure 7 shown, in an exemplary embodiment, in step S21, the method of determining the standard character corresponding to the above - mentioned preset character may include:

[0160] In step S2101, perform word - segmentation processing on the above - mentioned preset character to obtain the word - segmentation result corresponding to the above - mentioned preset character.

[0161] Exemplarily, a preset general word - segmentation model can be used to perform word - segmentation processing on the preset character to obtain the word - segmentation result corresponding to the preset character. For example, when the preset character is "cheesecake", using the general word - segmentation model to perform word - segmentation on "cheesecake", the obtained word - segmentation result is: cheese - cake. When the preset character is "tastes good", using the general word - segmentation model to perform word - segmentation on "tastes good", the obtained word - segmentation result is: taste - good.

[0162] Optionally, the preset general word segmentation model can be a hidden Markov model, a conditional random field model, etc.

[0163] In step S2103, determine a target proper noun that matches the above word segmentation result from a preset word list; the above preset word list stores multiple proper nouns through a double-array tree structure.

[0164] Figure 8 is a schematic diagram of a preset word list shown according to an exemplary embodiment. As Figure 8 shown, the preset word list stores multiple proper nouns through a double-array tree structure. Specifically, the double-array tree structure includes a root node and at least one leaf node, and a proper noun is stored in the path between the root node and each leaf node. For example, Figure 8 a proper noun "cheesecake" is stored in the path from the root node to the leaf node "cake" in, a proper noun "cheese open lid" is stored in the path from the root node to the leaf node "lid", and a proper noun "rich taste" is stored in the path from the root node to the leaf node "thick".

[0165] After obtaining the word segmentation result, it can be matched with the word segmentation result through the Figure 8 preset word list in to obtain the target proper noun corresponding to the word segmentation result.

[0166] In step S2105, based on the preset mapping information, determine the standard character corresponding to the above target proper noun; the above preset mapping information represents the mapping relationship between the proper noun and the standard character.

[0167] Exemplarily, a mapping relationship between multiple synonymous proper nouns and the same standard character can be established in advance. For example, the abbreviation of a certain enterprise can be used as the standard character, and the products corresponding to the enterprise include Product A, Product B, and Product C. Then, Product A, Product B, and Product C can be considered as synonymous proper nouns, and a mapping relationship between Product A, Product B, and Product C and the standard character can be established to obtain the preset mapping information.

[0168] After determining the target proper noun, the standard character corresponding to the target proper noun can be determined according to the preset mapping information established in advance. For example, if the target proper noun is Product A of a certain enterprise, then according to the preset mapping information, the standard character corresponding to Product A is determined to be the abbreviation of the enterprise.

[0169] In the embodiments of the present disclosure, by determining a target proper noun that matches the above-mentioned word segmentation result from a preset vocabulary, and based on preset mapping information representing the mapping relationship between proper nouns and standard characters, determining the standard characters corresponding to the above-mentioned target proper noun, the accuracy of determining the standard characters can be improved, ensuring that the first target concatenated text accurately reflects the subject to be detected, so that the preset language model trained through the first target concatenated text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model in classifying the text.

[0170] Figure 9 is a flowchart showing the determination of a target proper noun that matches the word segmentation result from a preset vocabulary according to an exemplary embodiment. As Figure 9 shown, in an exemplary embodiment, the above-mentioned double-array tree structure includes a root node and at least one leaf node, and a proper noun is stored in the path between the root node and each leaf node. If the number of the above-mentioned word segmentation results is multiple, then in the above-mentioned S2103, the determination of the target proper noun that matches the above-mentioned word segmentation result from the above-mentioned preset vocabulary may include:

[0171] In step S21031, at least two adjacent word segmentation results are combined to obtain a combined word segmentation result.

[0172] Specifically, when there are multiple word segmentation results, at least two adjacent word segmentation results can be combined to obtain a combined word segmentation result. For example, the preset character is "cheesecake", and the word segmentation result is: cheese - cake. Then cheese - cake can be combined to obtain a combined word segmentation result of "cheesecake". For example, the preset character is "tastes good", and the word segmentation result is: taste - good, then taste - good can be combined to obtain a combined word segmentation result of "tastes good".

[0173] In step S21033, when the proper noun stored in the path between the root node and one of the leaf nodes is the above-mentioned combined word segmentation result, the proper noun stored in the path between the root node and one of the leaf nodes is used as the target proper noun that matches the above-mentioned at least two adjacent word segmentation results; the one leaf node is a node among the above-mentioned at least one leaf node.

[0174] In the embodiments of the present disclosure, when using the preset vocabulary for matching, when the proper noun stored between the root node and a certain leaf node completely matches the combined word segmentation result, it is considered that the combined word segmentation result matches the preset vocabulary successfully, and the proper noun stored in the root node and a certain leaf node (the proper noun that completely matches successfully) is used as the target proper noun that matches the above-mentioned at least two adjacent word segmentation results.

[0175] Continue as Figure 8 shown, for the combined word segmentation result of "cheesecake", if it can exactly match the proper noun "cheesecake" stored in the path from the root node to "cake" in Figure 8 , then "cheesecake" is taken as the target proper noun of the word segmentation result "cheese - cake". For the combined word segmentation result of "tastes great", if it cannot exactly match the proper nouns stored in the path from the root node to any leaf node in Figure 8 , it is considered that the matching fails, and then the word segmentation result "taste - great" is filtered out.

[0176] In the embodiments of the present disclosure, the combined word segmentation result is matched with the proper nouns stored in the preset word list, and it is considered that the matching is successful only when the combined word segmentation result can be matched from the root node of the preset word list to any leaf node, thereby improving the accuracy of determining the target proper noun, ensuring that the first target concatenated text accurately reflects the subject to be detected and the core content of the text to be processed, so that the preset language model trained with the first target concatenated text can learn the key information of the text to be processed, thereby greatly improving the performance and recall accuracy of the model, and further improving the accuracy of the trained model in classifying the text.

[0177] In a feasible embodiment, after obtaining the first target concatenated text, the first target concatenated text can be input into the preset language model for further training to obtain a trained preset language model. The trained preset language model is no longer restricted by the number of characters, can adapt to long texts of any number of characters, and has a relatively high accuracy in text classification.

[0178] Figure 10 is a flowchart of a text classification method shown according to an exemplary embodiment. As Figure 10 shown, the above - mentioned text classification method may include:

[0179] In step S31, obtain the text to be classified.

[0180] In step S33, when the length of the text to be classified is greater than the preset length, extract the first sub - text of the preset length from the text to be classified.

[0181] In step S35, when the second sub - text includes the preset character, concatenate the preset character and multiple characters in the first sub - text to obtain the second target concatenated text of the preset length; the second sub - text is the text in the text to be classified except the first sub - text.

[0182] The above steps S31 - S35 are similar to the above steps S11 - S15 (simply modify the "text to be processed" in steps S11 - S15 to "text to be classified"), and will not be elaborated here.

[0183] In step S37, classify the above second target concatenated text through a preset language model; the above preset language model is trained based on the first target concatenated text in any of the above embodiments.

[0184] Specifically, the second target concatenated text can be input into the preset language model to obtain a classification result. Since the second target concatenated text can accurately reflect the core content of the subject to be detected and the text to be classified, inputting the second target concatenated text into the preset language model can learn the key information of the text to be classified, thereby greatly improving the performance of the model and the accuracy of text classification.

[0185] Figure 11 is a block diagram of a text processing device shown according to an exemplary embodiment. Refer to Figure 11 , the device may include a text - to - be - processed acquisition module 41, a first extraction module 43, and a first concatenation module 45.

[0186] The text - to - be - processed acquisition module 41 is configured to acquire the text to be processed.

[0187] The first extraction module 43 is configured to, when the length of the above text to be processed is greater than a preset length, extract a first sub - text of the above preset length from the above text to be processed.

[0188] The first concatenation module 45 is configured to, when the second sub - text includes preset characters, concatenate the preset characters and multiple characters in the above first sub - text to obtain the first target concatenated text of the above preset length; wherein, the above second sub - text is the text in the above text to be processed except the above first sub - text.

[0189] In an exemplary embodiment, the above first concatenation module 45 may include:

[0190] A preset number of character extraction units, configured to extract a preset number of characters from the above first sub - text; the number of characters in the above first sub - text except the above preset number of characters is equal to the number of characters included in the above preset characters.

[0191] A first target concatenated text determination unit, configured to concatenate the above preset number of characters and the above preset characters to obtain the above first target concatenated text.

[0192] In an exemplary embodiment, the above first extraction module 43 may include:

[0193] The first paragraph text extraction unit is configured to extract a first quantity of characters in the direction of the next character starting from the first character of the text to be processed above, to obtain the first paragraph text.

[0194] The last paragraph text extraction unit is configured to extract a second quantity of characters in the direction of the previous character with the last character of the text to be processed above as the termination position, to obtain the last paragraph text.

[0195] The first sub-text determination unit is configured to use the above first paragraph text and the above last paragraph text as the above first sub-text; the difference between the above first quantity and the above second quantity is less than a preset quantity threshold.

[0196] In an exemplary embodiment, among the above preset characters, there are a third quantity of characters, and the above preset quantity of character extraction unit may include:

[0197] The fourth quantity of character extraction sub-unit is configured to extract a fourth quantity of characters in the direction of the next character starting from the first character of the above first paragraph text.

[0198] The fifth quantity of character extraction sub-unit is configured to extract a fifth quantity of characters in the direction of the previous character with the last character of the above last paragraph text as the termination position.

[0199] The preset quantity of character determination sub-unit is configured to use the above fourth quantity of characters and the above fifth quantity of characters as the above preset quantity of characters; wherein, the above fourth quantity is less than the above first quantity, the above fifth quantity is less than the above second quantity, the sum of the first difference and the second difference is equal to the above third quantity, the first difference represents the difference between the above first quantity and the above fourth quantity, and the second difference represents the difference between the above second quantity and the above fifth quantity.

[0200] In an exemplary embodiment, the above first target splicing text determination unit is configured to splice the above fourth quantity of characters, the above preset characters, and the above fifth quantity of characters to obtain the above first target splicing text.

[0201] In an exemplary embodiment, the above device may further include:

[0202] The standard character determination module is configured to determine the standard character corresponding to the above preset character when the above preset character is included in the above second sub-text;

[0203] The second splicing module is configured to splice the above standard character and the above multiple characters to obtain the above first target splicing text.

[0204] In an exemplary embodiment, the above-mentioned standard character determination module may include:

[0205] A word segmentation result determination unit configured to perform word segmentation on the above-mentioned preset characters to obtain a word segmentation result corresponding to the above-mentioned preset characters.

[0206] A target proper noun determination unit configured to determine a target proper noun matching the above-mentioned word segmentation result from a preset word list; the above-mentioned preset word list stores multiple proper nouns through a double-array tree structure.

[0207] A standard character determination unit configured to determine a standard character corresponding to the above-mentioned target proper noun based on preset mapping information; the above-mentioned preset mapping information represents the mapping relationship between proper nouns and standard characters.

[0208] In an exemplary embodiment, the above-mentioned double-array tree structure includes a root node and at least one leaf node. A proper noun is stored in the path between the above-mentioned root node and each leaf node. The number of the above-mentioned word segmentation results is multiple. The above-mentioned target proper noun determination unit includes:

[0209] A combined word segmentation result determination subunit configured to perform combining at least two adjacent word segmentation results to obtain a combined word segmentation result.

[0210] A target proper noun determination subunit configured to, when the proper noun stored in the path between the above-mentioned root node and one of the leaf nodes matches the above-mentioned combined word segmentation result, use the proper noun stored in the path between the above-mentioned root node and one of the leaf nodes as the target proper noun matching the above-mentioned at least two adjacent word segmentation results; the above-mentioned one of the leaf nodes is a node among the above-mentioned at least one leaf node.

[0211] Figure 12 It is a block diagram of a text classification device shown according to an exemplary embodiment. Refer to Figure 12 , the device may include a text to be classified acquisition module 51, a second extraction module 53, a third splicing module 55, and a classification module 57.

[0212] The text to be classified acquisition module 51 is configured to acquire the text to be classified.

[0213] The second extraction module 53 is configured to, when the length of the above-mentioned text to be classified is greater than a preset length, extract a first sub-text of the above-mentioned preset length from the above-mentioned text to be classified.

[0214] The third splicing module 55 is configured to perform splicing of the preset character and multiple characters in the first sub-text when the second sub-text includes the preset character, so as to obtain the second target spliced text of the preset length; the second sub-text is the text in the text to be classified except the first sub-text.

[0215] The classification module 57 is configured to perform classification on the second target spliced text through a preset language model; the preset language model is trained based on the first target spliced text in any of the above embodiments.

[0216] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0217] In an exemplary embodiment, an electronic device is further provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored on the memory, the steps of the training method of any text processing model or the steps of any text processing method in the above embodiments are implemented.

[0218] The electronic device may be a terminal, a server, or a similar computing device. Taking the electronic device as a server as an example, Figure 13 FIG. 485 is a block diagram of an electronic device for text processing according to an exemplary embodiment. The electronic device 60 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 61 (the central processing unit 61 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 63 for storing data, and one or more storage media 62 for storing application programs 623 or data 622 (for example, one or more mass storage devices). Among them, the memory 63 and the storage media 62 may be transient storage or persistent storage. The program stored in the storage media 62 may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the central processing unit 61 may be set to communicate with the storage media 62 and execute a series of instruction operations in the storage media 62 on the electronic device 60. The electronic device 60 may further include one or more power supplies 66, one or more wired or wireless network interfaces 65, one or more input / output interfaces 64, and / or one or more operating systems 621, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0219] The input / output interface 64 can be used to receive or transmit data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the electronic device 60. In one example, the input / output interface 64 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In an exemplary embodiment, the input / output interface 64 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0220] Those of ordinary skill in the art can understand that Figure 13 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 60 may further include more or fewer components than Figure 13 shown, or have a different configuration from Figure 13 that shown.

[0221] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the computer-readable storage medium are executed by a processor of the electronic device, the electronic device can execute the steps of any text processing method or any text classification method in the above embodiments.

[0222] In an exemplary embodiment, a computer program product is further provided, including a computer program, which when executed by a processor implements the text processing method or the text classification method provided in any of the above embodiments.

[0223] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present disclosure can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0224] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0225] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A text processing method, characterized in that, Including: Obtain the text to be processed; When the length of the text to be processed is greater than a preset length, use at least one specified character in the text to be processed as a reference position, and extract the first sub-text of the preset length from the text to be processed; When the second sub-text includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain the first target spliced text of the preset length; The splicing the preset character and multiple characters in the first sub-text to obtain the first target spliced text of the preset length includes: extracting a preset number of characters from the first sub-text; the number of characters in the first sub-text other than the preset number of characters is equal to the number of characters included in the preset character; splicing the preset number of characters and the preset character to obtain the first target spliced text; Wherein, the second sub-text is the text in the text to be processed other than the first sub-text.

2. The text processing method according to claim 1, characterized in that, The step of, when the length of the text to be processed is greater than a preset length, using at least one specified character in the text to be processed as a reference position, and extracting the first sub-text of the preset length from the text to be processed includes: Taking the first character of the text to be processed as the starting position, and extracting the first number of characters in the direction of the next character to obtain the first segment of text; Taking the last character of the text to be processed as the termination position, and extracting the second number of characters in the direction of the previous character to obtain the last segment of text; Taking the first segment of text and the last segment of text as the first sub-text; the difference between the first number and the second number is less than a preset number threshold.

3. The text processing method according to claim 2, characterized in that, The preset character includes a third number of characters, and the extracting a preset number of characters from the first sub-text includes: Taking the first character of the first segment of text as the starting position, and extracting the fourth number of characters in the direction of the next character; Taking the last character of the last segment of text as the termination position, and extracting the fifth number of characters in the direction of the previous character; Taking the fourth number of characters and the fifth number of characters as the preset number of characters; Wherein, the fourth number is less than the first number, the fifth number is less than the second number, the sum of the first difference and the second difference is equal to the third number, the first difference represents the difference between the first number and the fourth number, and the second difference represents the difference between the second number and the fifth number.

4. The text processing method according to claim 3, characterized in that, The splicing the preset number of characters and the preset character to obtain the first target spliced text includes: Splicing the fourth number of characters, the preset character and the fifth number of characters to obtain the first target spliced text.

5. The text processing method according to any one of claims 1 to 4, characterized in that, The method further includes: When the second sub-text includes the preset character, determining the standard character corresponding to the preset character; Splicing the standard character and the multiple characters to obtain the first target spliced text.

6. The text processing method according to claim 5, characterized in that, The determining the standard character corresponding to the preset character includes: Perform word segmentation on the preset character to obtain the word segmentation result corresponding to the preset character; Determine a target proper noun matching the word segmentation result from a preset word list; the preset word list stores multiple proper nouns through a double-array tree structure; Based on preset mapping information, determine the standard character corresponding to the target proper noun; the preset mapping information represents the mapping relationship between the proper noun and the standard character.

7. The text processing method according to claim 6, characterized in that, The double-array tree structure includes a root node and at least one leaf node. The path between the root node and each leaf node stores a proper noun. The number of word segmentation results is multiple. The determining of the target proper noun matching the word segmentation result from the preset word list includes: Combine at least two adjacent word segmentation results to obtain a combined word segmentation result; When the proper noun stored in the path between the root node and one of the leaf nodes matches the combined word segmentation result, use the proper noun stored in the path between the root node and one of the leaf nodes as the target proper noun matching the at least two adjacent word segmentation results; the one leaf node is a node among the at least one leaf node.

8. A text classification method, characterized in that, Include: Obtain the text to be classified; When the length of the text to be classified is greater than a preset length, extract the first sub-text of the preset length from the text to be classified; When the second sub-text includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain the second target spliced text of the preset length; the second sub-text is the text in the text to be classified other than the first sub-text; Classify the second target spliced text through a preset language model; the preset language model is trained based on the first target spliced text in claim 1.

9. A text processing device, characterized in that, Include: A text acquisition module to be processed, configured to execute the acquisition of the text to be processed; A first extraction module, configured to execute when the length of the text to be processed is greater than a preset length, and extract the first sub-text of the preset length from the text to be processed with at least one specified character in the text to be processed as a reference position; A first splicing module, configured to execute when the second sub-text includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain the first target spliced text of the preset length; The first splicing module includes a preset number of character extraction units and a first target spliced text determination unit. The preset number of character extraction units is configured to execute extracting a preset number of characters from the first sub-text; the number of characters in the first sub-text other than the preset number of characters is equal to the number of characters included in the preset character; the first target spliced text determination unit is configured to execute splicing the preset number of characters and the preset character to obtain the first target spliced text; wherein, the second sub-text is the text in the text to be processed other than the first sub-text.

10. The text processing device according to claim 9, characterized in that, The first extraction module includes: The first paragraph text extraction unit is configured to extract a first number of characters in the direction of the next character starting from the first character of the text to be processed, to obtain the first paragraph text; The last paragraph text extraction unit is configured to extract a second number of characters in the direction of the previous character with the last character of the text to be processed as the termination position, to obtain the last paragraph text; The first sub-text determination unit is configured to use the first paragraph text and the last paragraph text as the first sub-text; the difference between the first number and the second number is less than a preset number threshold.

11. The text processing device according to claim 10, characterized in that, The preset characters include a third number of characters, and the preset number of character extraction unit includes: The fourth number of character extraction sub-unit is configured to extract a fourth number of characters in the direction of the next character starting from the first character of the first paragraph text; The fifth number of character extraction sub-unit is configured to extract a fifth number of characters in the direction of the previous character with the last character of the last paragraph text as the termination position; The preset number of character determination sub-unit is configured to use the fourth number of characters and the fifth number of characters as the preset number of characters; wherein, the fourth number is less than the first number, the fifth number is less than the second number, the sum of the first difference and the second difference is equal to the third number, the first difference represents the difference between the first number and the fourth number, and the second difference represents the difference between the second number and the fifth number.

12. The text processing device according to claim 11, characterized in that, The first target concatenated text determination unit is configured to concatenate the fourth number of characters, the preset characters, and the fifth number of characters to obtain the first target concatenated text.

13. The text processing device according to any one of claims 9 to 12, characterized in that, The device further includes: The standard character determination module is configured to determine the standard character corresponding to the preset character when the preset character is included in the second sub-text; The second concatenation module is configured to concatenate the standard character and the multiple characters to obtain the first target concatenated text.

14. The text processing device according to claim 13, characterized in that, The standard character determination module includes: The word segmentation result determination unit is configured to perform word segmentation processing on the preset character to obtain the word segmentation result corresponding to the preset character; The target proper noun determination unit is configured to determine the target proper noun matching the word segmentation result from a preset word list; the preset word list stores multiple proper nouns through a double-array tree structure; The standard character determination unit is configured to determine the standard character corresponding to the target proper noun based on preset mapping information; the preset mapping information represents the mapping relationship between proper nouns and standard characters.

15. The text processing device according to claim 14, characterized in that, The double-array tree structure includes a root node and at least one leaf node, and the path between the root node and each leaf node stores a proper noun. The number of the word segmentation results is multiple, and the target proper noun determination unit includes: The combined word segmentation result determination sub-unit is configured to combine at least two adjacent word segmentation results to obtain a combined word segmentation result; The target proper noun determination subunit is configured to, when the proper noun stored in the path between the root node and one of the leaf nodes matches the combined word segmentation result, use the proper noun stored in the path between the root node and one of the leaf nodes as the target proper noun that matches the at least two adjacent word segmentation results; where one of the leaf nodes is a node among the at least one leaf node.

16. A text classification device, characterized in that, Comprising: The text to be classified acquisition module is configured to acquire the text to be classified. The second extraction module is configured to, when the length of the text to be classified is greater than a preset length, extract the first sub-text of the preset length from the text to be classified. The third splicing module is configured to, when the second sub-text includes a preset character, splice the preset character and multiple characters in the first sub-text to obtain the second target spliced text of the preset length; the second sub-text is the text in the text to be classified other than the first sub-text. The classification module is configured to classify the second target spliced text through a preset language model; the preset language model is trained based on the first target spliced text in claim 1.

17. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the text processing method according to any one of claims 1 to 7 or the text classification method according to claim 8.

18. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, cause the electronic device to execute the text processing method according to any one of claims 1 to 7 or the text classification method according to claim 8.

19. A computer program product, comprising a computer program, characterized in that When the computer program is executed by the processor, it implements the text processing method according to any one of claims 1 to 7 or the text classification method according to claim 8.

Citation Information

Patent Citations

  • Answer generation method and device based on artificial intelligence, computer equipment and medium

    CN112417885A