Text processing method and apparatus, electronic device, and storage medium

By calculating the character noise probability in text fragments, the problem of low noise data recognition efficiency in the prior art is solved, and the rapid removal of noise characters in large-scale text data is achieved, and the efficiency and accuracy of training large models are improved.

WO2025140159A1PCT designated stage expired Publication Date: 2025-07-03VIVO MOBILE COMM CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141690
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the efficiency of identifying and removing noise data is inefficient, especially in large-scale text data. The manual annotation method is time-consuming and inefficient, and the accuracy of machine learning models is limited, resulting in the impact of the efficiency and accuracy of training large models.

Method used

By determining the character position and the central character position in the text fragment, calculating the noise probability, automatically filtering and processing noise characters, avoiding manual annotation, and improving the efficiency of identifying noise characters.

Benefits of technology

Quickly identifying and removing noise characters in large-scale text data improves the noise character recognition efficiency of electronic devices and improves the efficiency and accuracy of training large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141690_03072025_PF_FP_ABST
    Figure CN2024141690_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and discloses a a text processing method and apparatus, an electronic device, and a storage medium. The method comprises: on the basis of first text, acquiring at least one first text segment, each first text segment comprising at least two characters; on the basis of the position in a second text segment of a first character in a second text segment and the position in the second text segment of a central character in the second text segment, determining a first noise probability corresponding to the first character, the first noise probability being the probability that the first character is a noise character, the second text segment being one among the at least one first text segment, and the first character being a character in the first text segment; and on the basis of the first noise probability corresponding to each first character in each first text segment, processing the first text, to obtain second text.
Need to check novelty before this filing date? Find Prior Art

Description

Text processing method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311870122.1 and titled “Text processing method, device, electronic device and storage medium,” the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application belongs to the field of data processing technology, and specifically relates to a text processing method, device, electronic device and storage medium. Background Art

[0003] As electronic devices continue to develop, their functionality is also increasing. For example, electronic devices can retrieve text resources from web pages and determine the text required by the user from these resources. Typically, all text resources retrieved from web pages by electronic devices using crawler technology may contain noisy data. For example, web pages may contain many irregular characters, advertisements, navigation bars, statement templates, and other content that is meaningless to model training.

[0004] In related technologies, after obtaining text resources, electronic devices can identify unreasonable content in the data, namely noise data, through manual annotation, try to summarize some rules, and then use regular matching to filter out data that meets these rules.

[0005] However, when the amount of text resources reaches hundreds of millions, the process of identifying noise data in text resources through manual annotation is cumbersome and time-consuming. As a result, the efficiency of electronic devices in identifying noise data is low. Summary of the Invention

[0006] The purpose of the embodiments of the present application is to provide a text processing method, device, electronic device and storage medium that can improve the efficiency and accuracy of training large models.

[0007] In a first aspect, an embodiment of the present application provides a text processing method, which includes: based on a first text, obtaining at least one first text segment, each first text segment containing at least two characters; based on the position of the first character in the second text segment in the second text segment and the position of the center character in the second text segment in the second text segment, determining the first noise probability corresponding to the first character, the first noise probability being the probability that the first character is a noise character, the second text segment being one of at least one first text segment, and the first character being a character in the first text segment; processing the first text based on the first noise probability corresponding to each first character in each first text segment to obtain a second text.

[0008] In a second aspect, an embodiment of the present application provides a text processing device, which includes: an acquisition module, a determination module, and a processing module. The acquisition module is used to acquire at least one first text segment based on the first text, and each first text segment contains at least two characters. The determination module is used to determine the first noise probability corresponding to the first character based on the position of the first character in the second text segment acquired by the acquisition module in the second text segment and the position of the center character in the second text segment in the second text segment. The first noise probability is the probability that the first character is a noise character. The second text segment is one of the at least one first text segment, and the first character is a character in the first text segment. The processing module is used to process the first text based on the first noise probability corresponding to each first character in each first text segment determined by the determination module to obtain the second text.

[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0010] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0011] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0012] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.

[0013] In an embodiment of the present application, based on a first text, at least one first text segment is obtained, each first text segment contains at least two characters; based on the position of the first character in the second text segment in the second text segment and the position of the center character in the second text segment in the second text segment, the first noise probability corresponding to the first character is determined, the first noise probability being the probability that the first character is a noise character, the second text segment is one of at least one first text segment, and the first character is a character in the first text segment; based on the first noise probability corresponding to each first character in each first text segment, the first text is processed to obtain the second text. In this solution, the electronic device can filter out noise characters from the first text based on the first noise probability corresponding to the first character in the first text, and then process the noise characters without manual labeling, so that even if the amount of data in the first text is large, the electronic device can also quickly filter out noise characters from the first text, thereby improving the efficiency of the electronic device in identifying noise characters. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] FIG1 is a flowchart of a text processing method provided in an embodiment of the present application;

[0015] FIG2 is a second flowchart of a text processing method provided in an embodiment of the present application;

[0016] FIG3 is a third flowchart of a text processing method provided in an embodiment of the present application;

[0017] FIG4 is a fourth flowchart of a text processing method provided in an embodiment of the present application;

[0018] FIG5 is a fifth flowchart of a text processing method provided in an embodiment of the present application;

[0019] FIG6 is a schematic structural diagram of a text acquisition device provided in an embodiment of the present application;

[0020] FIG7 is a schematic diagram of a hardware structure of an electronic device provided in an embodiment of the present application;

[0021] FIG8 is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. Specific embodiments

[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0023] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0024] The terms "at least one" and "at least one of" in the specification and claims of this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".

[0025] The text processing method, device, electronic device and storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0026] The following is an explanation of the professional terms involved in this application.

[0027] Pre-trained large model: A model trained on large amounts of data using artificial intelligence technologies such as deep learning. The model has billions of parameters or more, and the amount of training data reaches TB or more.

[0028] Noise characters: Meaningless data included in the data used to train AI models. These interference data do not positively improve model training results and may even reduce them. For example, the noise characters could be "Add me to your favorites to see me first; only I know my taste, a solution to learning problems."

[0029] Polygram (n-gram): refers to n words that appear consecutively in a text. For example, after Chinese word segmentation of the sentence "The weather is very good today", we can get three words: "today", "weather", and "very good". If n is 1, then these three words constitute a unary grammar; if n is 2, then "the weather is very good today" and "the weather is very good" respectively constitute two bigrams, and so on.

[0030] With the continuous development of large-scale pre-trained model technology, the efficient training of high-quality pre-trained base models has become a research focus. High-quality pre-trained models are determined by massive amounts of high-quality corpus data. This corpus data comes from publicly available data collected from the internet, including Wikipedia, Q&A websites, personal blogs, forums, e-books, and scientific literature. Since data obtained from the internet is often interspersed with irregular characters, advertisements, navigation bars, statement templates, and other content that is meaningless to model training, we call this "noise data." Therefore, how to clean this noise data from massive text corpus data is crucial for improving model performance.

[0031] The industry is currently openly discussing some methods for cleaning data. The main methods are as follows:

[0032] Method 1: Identify unreasonable content in the data through manual labeling, try to summarize some patterns, and then use regular matching to filter out data that conforms to these patterns; Method 2: Use some open source machine learning models to screen noisy data.

[0033] However, for the above method 1, the amount of data required for the pre-training model is huge, and it is difficult to effectively identify all problematic data through manual labeling. In fact, the manually screened data accounts for less than one ten-thousandth of the total data volume.

[0034] Regarding the second method mentioned above, the accuracy of the machine learning model's ability to identify noise data is very limited. The data marked as noise data by the algorithm may be normal and useful data. Therefore, this result is uncontrollable and may cause a large amount of data to be wasted. Moreover, the proportion of noise data has not decreased, thus affecting the efficiency and accuracy of training large models.

[0035] In the text processing method, device, electronic device and storage medium provided in the embodiments of the present application, the electronic device can filter out noise characters from the first text based on the first noise probability corresponding to the first character in the first text, and then process the noise characters without manual labeling. Therefore, when the data volume of the first text is large, the electronic device can also quickly filter out noise characters from the first text, thereby improving the efficiency of the electronic device in recognizing noise characters.

[0036] The text processing method provided in the embodiment of the present application may be executed by a text acquisition device, which may be an electronic device or a functional module in an electronic device. The technical solution provided in the embodiment of the present application is described below using an electronic device as an example.

[0037] The present invention provides a text processing method, and Figure 1 shows a flowchart of the text processing method provided by the present invention. As shown in Figure 1, the text processing method provided by the present invention may include the following steps 201 to 203.

[0038] Step 201: The electronic device obtains at least one first text segment based on a first text.

[0039] In an embodiment of the present application, each of the at least one first text segment contains at least two characters.

[0040] Optionally, in an embodiment of the present application, the first text may be text obtained by the electronic device from a web page; or, the first text may be text obtained from text stored in the electronic device.

[0041] Illustratively, the first text may include but is not limited to at least one of the following: Chinese text, English text, digital text, etc.

[0042] For example, the first text may include at least one of an advertisement text, a page template text, a scientific literature text, and a navigation bar text.

[0043] Optionally, in an embodiment of the present application, the electronic device may perform segmentation processing on the first text to obtain at least one first text segment.

[0044] Optionally, in the embodiment of the present application, in combination with FIG1 , as shown in FIG2 , the above step 201 can be specifically implemented through the following steps 201a and 201b.

[0045] Step 201a: The electronic device divides the first text into segments to obtain at least one third text segment.

[0046] Optionally, in an embodiment of the present application, the electronic device may divide the first text using a first division method to obtain at least one third text segment.

[0047] Exemplarily, the first division method may be any one of the following: n-gram division, punctuation division, special character division, or preset sentence length division.

[0048] In one example, the electronic device can divide the first text into n-grams to obtain at least one third text segment. For example, the first text is first segmented, and each word is a token, so the result after segmentation is a token array. Starting from the first character of this array, every n consecutive tokens are taken out and combined into an n-gram segment, which is a segmented segment, that is, the third text segment mentioned above; assuming that the length of this token array is L, a total of L-n+1 n-gram segments can be combined. For example, if n is 2 and L is 10, then [0,1], [1,2], [2,3]…[8,9] can be combined, where the subscripts in the brackets are in the token array.

[0049] It should be noted that the value of n is greater than 1. The specific value needs to be verified based on the data size of the first text, and then the value with the best noise extraction effect is selected; or the above entire process is iterated multiple times, taking different values ​​each time, so as to cover different noise types.

[0050] In another example, the electronic device may divide the first text by using a period in the first text as a separator to obtain at least one third text segment.

[0051] In another example, the electronic device may divide the first text by using line breaks in the first text as separators to obtain at least one third text segment.

[0052] In another example, the electronic device may divide the first text according to a fixed character length to obtain at least one third text segment. For example, assuming that the character length determined by the electronic device is 15 characters, the electronic device may divide the first text according to the 15 characters to obtain at least one third text segment.

[0053] Step 201b: The electronic device deletes the start character and the end character in each third text segment to obtain at least one processed third text segment.

[0054] In the embodiment of the present application, the at least one first text segment is at least one processed third text segment.

[0055] It should be noted that, based on a large amount of data, it is found that noise characters generally appear at the beginning or end of a sentence. Therefore, after obtaining at least one third text segment, the electronic device can directly delete the starting character and ending character in each third text segment to obtain at least one processed third text segment, that is, the at least one first text segment mentioned above.

[0056] Optionally, in an embodiment of the present application, for a third text segment in at least one third text segment, the electronic device can position number each character in the third text segment, thereby obtaining a position number of each character; then, the electronic device can determine the starting character and the ending character in a third text segment through the position number of each character, and then delete the starting character and the ending character in a third text segment.

[0057] It should be noted that, for each third text segment in the at least one third text segment mentioned above, the electronic device can determine the starting character and the ending character in each third text segment in the at least one third text segment in the above-mentioned method, and then delete the starting character and the ending character in each third text segment in the at least one third text segment to obtain at least one first text segment.

[0058] In an embodiment of the present application, the electronic device can delete the starting character and the ending character in each third text segment in at least one third text segment based on a large amount of data, thereby first removing possible noise characters in the first text, and then reducing meaningless characters in at least one first text segment in the first text.

[0059] Step 202: The electronic device determines a first noise probability corresponding to the first character in the second text segment based on the position of the first character in the second text segment and the position of the central character in the second text segment.

[0060] In an embodiment of the present application, the above-mentioned first noise probability is the probability that the first character is a noise character, the second text segment is one of the at least one first text segment, and the first character is a character in the first text segment.

[0061] In other words, the present application can determine the noise probability corresponding to each first character in each text segment in at least one first text segment.

[0062] In the embodiment of the present application, the electronic device may determine the central character in the second text segment by using the position number of each first character in the second text segment.

[0063] Optionally, in an embodiment of the present application, the number of the above-mentioned central characters can be one or more.

[0064] Optionally, in an embodiment of the present application, the electronic device may calculate the first noise probability corresponding to each first character in at least one first text segment in the first text in sequence according to the above method to obtain the noise probabilities corresponding to all characters in the first text.

[0065] Optionally, in the embodiment of the present application, after obtaining the noise probabilities corresponding to all characters in the first text, the electronic device may store the noise probabilities corresponding to all characters in the first text.

[0066] Exemplarily, the electronic device may store all characters in the first text and the noise probabilities corresponding to all characters in the first text in a key-value pair storage manner. For example, taking a character and the noise probability corresponding to the character as an example, the electronic device may use a character in the first text as a key and the noise probability corresponding to the character as a value, and then store the character and the noise probability corresponding to the character in a key-value pair manner. In this way, the electronic device can directly obtain the noise probability corresponding to the character through the character.

[0067] Step 203: The electronic device processes the first text based on the first noise probability corresponding to each first character in each first text segment to obtain a second text.

[0068] Optionally, in an embodiment of the present application, the second text can be used for model training.

[0069] In an embodiment of the present application, when the noise probability corresponding to the first character in the second text segment is greater than a preset threshold, the electronic device can determine that the first character is a noise character, and then the electronic device can process the first character from the first text.

[0070] Optionally, in an embodiment of the present application, the first character may be one or more.

[0071] It should be noted that, for each first character in each first text segment in the first text, the electronic device can determine whether each first character in the first text is a noise character in the above manner, thereby processing the noise characters in the first text to obtain the second text.

[0072] Optionally, in an embodiment of the present application, after obtaining the second text, the electronic device may input the second text into the large model to train the large model, thereby obtaining a trained large model.

[0073] For example, the large model may include but is not limited to any of the following: a dialogue model, a translation model, and a text generation model, etc. The specific model may be determined according to actual usage and is not limited in the embodiments of this application.

[0074] Optionally, in the embodiment of the present application, in combination with FIG1 , as shown in FIG3 , the above step 203 may be specifically implemented through the following step 203a.

[0075] Step 203a: When the first noise probability corresponding to the first character is greater than a preset threshold, the electronic device deletes the first character from the first text.

[0076] Optionally, in an embodiment of the present application, when the first noise probability corresponding to the first character is less than or equal to a preset threshold, the electronic device may retain the first character.

[0077] Optionally, in an embodiment of the present application, after determining all noise characters in the first text, the electronic device may delete all noise characters at once to obtain the second text.

[0078] Optionally, in an embodiment of the present application, the electronic device obtains a fourth text after deleting noise characters in the first text, and then performs text normalization processing on the fourth text to obtain a fourth text after text normalization processing.

[0079] Illustratively, the text normalization processing performed on the fourth text may include deleting punctuation marks, control symbols, and other content without clear semantic meanings in the first text.

[0080] For example, the electronic device may perform text normalization processing on the fourth text by using a regular expression to obtain the fourth text after text normalization processing.

[0081] In this way, the electronic device can perform data cleaning on the characters in the fourth text by using regular expressions to reduce meaningless characters in the fourth text.

[0082] In the embodiment of the present application, the electronic device can delete the noise characters in the first text to obtain the second text, thereby ensuring that the second text contains less noise data.

[0083] In the text processing method provided in the embodiment of the present application, based on the first text, at least one first text segment is obtained, each first text segment contains at least two characters; based on the position of the first character in the second text segment in the second text segment and the position of the center character in the second text segment in the second text segment, the first noise probability corresponding to the first character is determined, the first noise probability being the probability that the first character is a noise character, the second text segment is one of the at least one first text segment, and the first character is a character in the first text segment; based on the first noise probability corresponding to each first character in each first text segment, the first text is processed to obtain the second text. In this solution, the electronic device can filter out noise characters from the first text based on the first noise probability corresponding to the first character in the first text, and then process the noise characters without manual labeling, so that even if the amount of data in the first text is large, the electronic device can also quickly filter out noise characters from the first text, thereby improving the efficiency of the electronic device in recognizing noise characters.

[0084] Optionally, in an embodiment of the present application, in combination with FIG1 , as shown in FIG4 , the above step 202 may be specifically implemented through the following steps 202a to 202c.

[0085] Step 202a: The electronic device numbers each character in the second text segment starting from the first character in the second text segment to obtain a position number for each character in the second text segment.

[0086] In the embodiment of the present application, the electronic device may start from the first character in the second text segment, number the position of each character in the second text segment, and obtain the position number of each character in the second text segment.

[0087] For example, assuming that the characters contained in the second text segment are "hello weather", the electronic device can start numbering from the character "you" and set the position number corresponding to the character "you" to 1; the electronic device can set the position number corresponding to the character "good" to 2; the electronic device can set the position number corresponding to the character "day" to 3; the electronic device can set the position number corresponding to the character "air" to 4; thereby, the position number of each character in the second text segment is 1, 2, 3, 4.

[0088] It should be noted that, for each first text segment in the at least one first text segment, the electronic device may number them starting from 1.

[0089] Step 202b: The electronic device determines the central character in the second text segment based on the position number of each character in the second text segment.

[0090] In the embodiment of the present application, after obtaining the position number of each character in the second text segment, the electronic device may obtain the central character in the second text segment through a center number function.

[0091] For example, the electronic device may obtain the central character in the second text segment by using the following formula 1, where formula 1 is specifically:

[0092] mid-value=medin(1...n) (1)

[0093] Wherein, mid-value is the center position number, medin() is the center number function, and 1…n is the position number of each character in the second text segment.

[0094] Furthermore, the electronic device may determine the character corresponding to the center position number as the center character.

[0095] Step 202c: The electronic device calculates a first probability value corresponding to the first character based on the absolute value of the difference between the position number of the first character and the position number of the center character.

[0096] In an embodiment of the present application, after obtaining the absolute value of the difference between the position number of the first character and the position number of the center character, the electronic device can obtain the first noise probability corresponding to the first character through a first function.

[0097] Optionally, in an embodiment of the present application, after obtaining the absolute value of the difference between the position number of the first character and the position number of the center character, hereinafter referred to as the first absolute value of the difference, the electronic device can take the negative of the first absolute value of the difference, and then input the negative of the first absolute value of the difference into the first function to obtain the first noise probability corresponding to the first character.

[0098] Exemplarily, the electronic device may obtain the negative number of the absolute value of the first difference by using the following formula 2, where formula 2 may specifically be:

[0099] difference=-|n-mid_value| (2)

[0100] Where difference is the negative of the absolute value of the first difference, n is the position number of the first character, and mid_value is the center position number.

[0101] Optionally, in an embodiment of the present application, the first function may include any one of the following: a sigmoid function, a softmax function, or a tanh function.

[0102] For example, the following uses the sigmoid function as an example to explain in detail how to obtain the first noise probability corresponding to the first character in this application. After obtaining the difference, the electronic device can input the difference into the following formula:

[0103] In formula 1, to obtain the first noise probability corresponding to the first character, formula 3 is specifically:

[0104] Among them, score is the first noise probability corresponding to the first character, difference is the negative of the absolute value of the first difference, is the Sigmod function.

[0105] It should be noted that the purpose of taking a negative number is to bring the distance value between the first character and the center character into the Sigmoid function to convert the distance into a probability between 0 and 1.

[0106] It can be understood that since the negative of the absolute value of the first difference is input into the Sigmund function, the smaller the distance between the first character in the second text segment and the center character, the greater the probability value of the first noise probability corresponding to the first character. That is, the greater the probability that the first character is a noise character.

[0107] It should be noted that, for each character in the first text, the electronic device can obtain the noise probability corresponding to each character in the above manner. To avoid repetition, it will not be described here.

[0108] In an embodiment of the present application, the electronic device obtains the first noise probability corresponding to the first character, and can thereby determine whether the first character is a noise character based on the first noise probability, without the need for manual detection, thereby improving the efficiency of the electronic device in determining noise characters.

[0109] Optionally, in an embodiment of the present application, in combination with FIG1 , as shown in FIG5 , after the above-mentioned step 203 , the text processing method provided in the embodiment of the present application further includes the following steps 301 to 303 .

[0110] Step 301: The electronic device aggregates at least one character based on the character content of each character in the second text to obtain at least one character set.

[0111] In the embodiment of the present application, each character set in the at least one character set includes at least one character.

[0112] In an embodiment of the present application, the electronic device may aggregate characters with the same character content to obtain at least one character set.

[0113] Exemplarily, the electronic device may aggregate characters with the same character content in the second text using an aggregation function to obtain at least one character set.

[0114] For example, the electronic device may aggregate characters with the same character content in the second text using a group by function to obtain at least one character set.

[0115] Step 302: The electronic device uses characters in a first character set in at least one character set as noise characters.

[0116] In an embodiment of the present application, the number of repetitions of characters in the first character set is greater than or equal to a preset threshold.

[0117] In an embodiment of the present application, the electronic device can count the number of repetitions of characters in the first character set through a counting function, thereby determining the characters in the first character set as noise characters when the number of repetitions of characters in the first character set is greater than or equal to a preset threshold.

[0118] Exemplarily, the electronic device may count the number of repetitions of characters in the first character set through a count function, and thereby determine a character in the first character set as a noise character if the number of repetitions of characters in the first character set is greater than or equal to a preset threshold.

[0119] It should be noted that, for each character set in the at least one character set, the electronic device can obtain the noise character in each character set through the above method, which will not be described again here to avoid repetition.

[0120] Optionally, in an embodiment of the present application, after the electronic device filters out the noise characters from the at least one character set, the electronic device may treat the noise characters filtered out from the at least one character set as a noise character set, and save the noise character set.

[0121] Step 303: The electronic device deletes the noise characters in the second text to obtain a third text.

[0122] Optionally, in an embodiment of the present application, the third text mentioned above can be used for model training.

[0123] In an embodiment of the present application, the electronic device can traverse the characters in the second text based on the noise character set, then determine the noise characters from the second text, and delete the noise characters in the second text to obtain the third text.

[0124] Exemplarily, the electronic device may start from the starting character in the second text, search for characters in the second text that are repeated in the noise character set in sequence, mark the characters in the second text that are repeated in the noise character set, and then delete the characters in the second text that are repeated in the noise character set after the traversal is completed.

[0125] Optionally, in an embodiment of the present application, the electronic device may start from the starting character in the second text, batch search for characters in the second text that are repeated in the noise character set, mark the characters in the second text that are repeated in the noise character set, and then delete the characters in the second text that are repeated in the noise character set after the traversal is completed.

[0126] Exemplarily, the electronic device can use a sliding window, starting from the starting character in the second text, to batch search for characters in the second text that are repeated in the noise character set, mark the characters in the second text that are repeated in the noise character set, and then delete the characters in the second text that are repeated in the noise character set after the traversal is completed.

[0127] In the embodiment of the present application, the electronic device can delete the noise characters in the second text to obtain the third text, and can completely remove the noise characters in the second text.

[0128] For example, the text processing method provided in the embodiment of the present application is explained in detail below through specific examples, which can be implemented through the following steps 20 to 27.

[0129] Step 20: The electronic device obtains the first text.

[0130] In the embodiment of the present application, the electronic device can obtain the first text from the Internet.

[0131] Exemplarily, the electronic device may obtain the first text from the Internet by using crawler technology.

[0132] Step 21: The electronic device divides the first text into text segments according to a segmentation strategy to obtain at least one text segment.

[0133] Optionally, in an embodiment of the present application, the above-mentioned division strategy may include any one of the following: n-gram grammar, punctuation division strategy, special symbol division strategy or line break division strategy.

[0134] Step 22: The electronic device deletes the first character from the first text segment based on the first probability value corresponding to the first character in the first text segment to obtain a second text.

[0135] It should be noted that the specific implementation process can be found in the above embodiments, and will not be described again here to avoid repetition.

[0136] Step 23: The electronic device performs text normalization processing on the second text to obtain a third text.

[0137] Step 24: The electronic device performs character aggregation based on repeated characters in the third text to obtain at least one character set.

[0138] Step 25: The electronic device counts the number of repeated characters in each character set in at least one character set, and then determines characters in the character set whose number of repeated characters is greater than a preset threshold as noise characters.

[0139] Step 26: The electronic device traverses the third text based on the noise character, and after traversing the third text, deletes the characters in the third text that are identical to the noise character to obtain a fourth text.

[0140] Step 27: The electronic device inputs the fourth text into the large model to train the large model.

[0141] It should be noted that the text processing method provided in the embodiments of the present application can be executed by a text acquisition device, an electronic device, or a functional module or entity in an electronic device. In the embodiments of the present application, the text acquisition device provided in the embodiments of the present application is described by taking the execution of the text processing method by the text acquisition device as an example.

[0142] FIG6 shows a possible structural diagram of a text processing device involved in an embodiment of the present application. As shown in FIG6 , the text processing device 70 may include: an acquisition module 71 , a determination module 72 , and a processing module 73 .

[0143] Among them, the acquisition module 71 is used to obtain at least one first text segment based on the first text, and each first text segment contains at least two characters. The determination module 72 is used to determine the first noise probability corresponding to the first character based on the position of the first character in the second text segment obtained by the acquisition module in the second text segment and the position of the center character in the second text segment in the second text segment. The first noise probability is the probability that the first character is a noise character, the second text segment is one of the at least one first text segment, and the first character is a character in the first text segment. The processing module 73 is used to process the first text based on the first noise probability corresponding to each first character in each first text segment determined by the determination module to obtain the second text.

[0144] In one possible implementation, the processing module 71 is further used to process the first text based on the noise probability corresponding to each character in each first text segment to obtain the second text, and then aggregate at least one character based on the character content of each character in the second text to obtain at least one character set, each character set containing at least one character; use characters in a first character set in the at least one character set as noise characters, and the number of repetitions of the characters in the first character set is greater than or equal to a preset threshold; delete the noise characters in the second text to obtain a third text.

[0145] In one possible implementation, the above-mentioned determination module 72 is specifically used to number each character in the second text segment starting from the first character in the second text segment to obtain the position number of each character in the second text segment; determine the central character in the second text segment based on the position number of each character in the second text segment; and calculate the first probability value corresponding to the first character based on the absolute value of the difference between the position number of the first character and the position number of the central character.

[0146] In a possible implementation, the processing module 73 is specifically configured to delete the first character from the first text when a first noise probability corresponding to the first character is greater than a preset threshold.

[0147] In one possible implementation, the acquisition module 71 is specifically configured to segment the first text to obtain at least one third text segment. The processing module 73 is further configured to delete the start and end characters in each third text segment to obtain at least one processed third text segment; wherein the at least one first text segment is the at least one processed third text segment.

[0148] An embodiment of the present application provides a text processing device, in which the electronic device can filter out noise characters from the first text based on a first noise probability corresponding to a first character in the first text, and then process the noise characters without manual labeling. Therefore, even if the data volume of the first text is large, the electronic device can also quickly filter out noise characters from the first text, thereby improving the efficiency of the electronic device in identifying noise characters.

[0149] The text processing device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0150] The text processing device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0151] The text processing device provided in the embodiment of the present application can implement each process implemented in the above embodiment. To avoid repetition, it will not be described here.

[0152] Optionally, as shown in Figure 7, an embodiment of the present application also provides an electronic device 90, including a processor 91 and a memory 92, and the memory 92 stores a program or instruction that can be run on the processor 91. When the program or instruction is executed by the processor 91, the various steps of the above-mentioned text processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0153] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0154] FIG8 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0155] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .

[0156] Those skilled in the art will appreciate that the electronic device 100 may further include a power source (such as a battery) for powering various components. The power source may be logically connected to the processor 110 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The electronic device structure shown in FIG8 does not limit the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components or arrange the components differently, which will not be described in detail here.

[0157] Among them, the processor 110 is used to obtain at least one first text segment based on the first text, each first text segment contains at least two characters; based on the position of the first character in the second text segment in the second text segment and the position of the central character in the second text segment in the second text segment, determine the first noise probability corresponding to the first character, the first noise probability is the probability that the first character is a noise character, the second text segment is one of the at least one first text segment, and the first character is a character in the first text segment; based on the first noise probability corresponding to each first character in each first text segment, the first text is processed to obtain the second text.

[0158] An embodiment of the present application provides an electronic device that can filter out noise characters from a first text based on a first noise probability corresponding to a first character in the first text, and then process the noise characters without manual labeling. Therefore, even when the data volume of the first text is large, the electronic device can quickly filter out noise characters from the first text, thereby improving the efficiency of the electronic device in identifying noise characters.

[0159] Optionally, in an embodiment of the present application, after processing the first text based on the noise probability corresponding to each character in each first text segment to obtain the second text, the above-mentioned processor 110 is further used to aggregate at least one character based on the character content of each character in the second text to obtain at least one character set, each character set containing at least one character; use characters in a first character set in the at least one character set as noise characters, and the number of repetitions of characters in the first character set is greater than or equal to a preset threshold; delete the noise characters in the second text to obtain a third text.

[0160] Optionally, in an embodiment of the present application, the above-mentioned processor 110 is specifically used to number each character in the second text segment starting from the first character in the second text segment to obtain a position number of each character in the second text segment; determine the central character in the second text segment based on the position number of each character in the second text segment; and calculate the first probability value corresponding to the first character based on the absolute value of the difference between the position number of the first character and the position number of the central character.

[0161] Optionally, in the embodiment of the present application, the processor 110 is specifically configured to delete the first character from the first text when a first noise probability corresponding to the first character is greater than a preset threshold.

[0162] Optionally, in an embodiment of the present application, the above-mentioned processor 110 is specifically used to divide the first text to obtain at least one third text segment; delete the starting character and the ending character in each third text segment to obtain at least one processed third text segment; wherein, at least one first text segment is at least one processed third text segment.

[0163] The electronic device provided in the embodiment of the present application can implement each process implemented in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0164] The beneficial effects of various implementations in this embodiment can be specifically referred to the beneficial effects of the corresponding implementations in the above method embodiment. To avoid repetition, they will not be described here.

[0165] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0166] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0167] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.

[0168] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0169] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0170] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0171] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0172] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned text processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0173] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0174] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0175] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A text processing method, the method comprising: Based on a first text, obtaining at least one first text segment, each of the first text segments containing at least two characters; Based on the position of a first character in a second text segment in the second text segment and the position of a central character in the second text segment in the second text segment, determining a first noise probability corresponding to the first character, the first noise probability being the probability that the first character is a noise character, the second text segment being one of the at least one first text segment, and the first character being a character in the first text segment; Based on the first noise probability corresponding to each first character in each of the first text segments, processing the first text to obtain a second text.

2. The method according to claim 1, wherein After processing the first text based on the noise probability corresponding to each character in each of the first text segments to obtain a second text, the method further comprises: Based on the character content of each character in the second text, performing an aggregation process on at least one character in the second text to obtain at least one character set, each of the character sets containing at least one character; Regarding the characters in a first character set among the at least one character set as noise characters, the number of repetitions of the characters in the first character set being greater than or equal to a preset threshold; Deleting the noise characters in the second text to obtain a third text.

3. The method according to claim 1, wherein The determining the first noise probability corresponding to the first character based on the position of the first character in the second text segment in the second text segment and the position of the central character in the second text segment in the second text segment comprises: Numbering each character in the second text segment starting from the first character in the second text segment to obtain the position numbers of each character in the second text segment; Based on the position numbers of each character in the second text segment, determining the central character in the second text segment; Calculating a first probability value corresponding to the first character based on the absolute value of the difference between the position number of the first character and the position number of the central character.

4. The method according to claim 1 or 3, wherein The processing the first text based on the first noise probability corresponding to each first character in each of the text segments comprises: In the case where the first noise probability corresponding to the first character is greater than a preset threshold, deleting the first character from the first text.

5. The method according to claim 1, wherein, The obtaining at least one first text segment based on the first text comprises: Dividing the first text to obtain at least one third text segment; Deleting the starting character and the ending character in each of the third text segments to obtain at least one processed third text segment; Wherein, the at least one first text segment is the at least one processed third text segment.

6. A text processing device, the device comprising: An obtaining module, a determining module, and a processing module; The obtaining module is configured to obtain at least one first text segment based on a first text, each of the first text segments containing at least two characters; The determining module is configured to determine a first noise probability corresponding to the first character based on the position of the first character in the second text segment obtained by the obtaining module in the second text segment and the position of the central character in the second text segment in the second text segment. The first noise probability is the probability that the first character is a noise character. The second text segment is one of the at least one first text segment, and the first character is a character in the first text segment. The processing module is configured to process the first text based on the first noise probability corresponding to each first character in each of the first text segments determined by the determining module to obtain a second text.

7. The apparatus according to claim 6, wherein The processing module is further configured to, after processing the first text based on the noise probability corresponding to each character in each of the first text segments to obtain a second text, perform an aggregation process on at least one character in the second text based on the character content of each character in the second text to obtain at least one character set, where each character set contains at least one character; use the characters in the first character set among the at least one character set as noise characters, and the repetition times of the characters in the first character set are greater than or equal to a preset threshold. Delete the noise characters in the second text to obtain a third text.

8. The apparatus according to claim 6, wherein, The determining module is specifically configured to number each character in the second text segment starting from the first character in the second text segment to obtain the position number of each character in the second text segment; determine the central character in the second text segment based on the position numbers of each character in the second text segment; calculate a first probability value corresponding to the first character based on the absolute value of the difference between the position number of the first character and the position number of the central character.

9. The device according to claim 6 or 8, wherein, The processing module is specifically configured to delete the first character from the first text when the first noise probability corresponding to the first character is greater than a preset threshold.

10. The device according to claim 6, wherein, The obtaining module is specifically configured to divide the first text to obtain at least one third text segment. The processing module is further configured to delete the starting character and the ending character in each of the third text segments to obtain at least one processed third text segment. Wherein, the at least one first text segment is the at least one processed third text segment.

11. An electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the text processing method according to any one of claims 1 to 5 are implemented.

12. [Corrected according to Rule 26 on 16.01.2025] A readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the text processing method according to any one of claims 1 to 5 are implemented.

13. A chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the steps of the text processing method according to any one of claims 1 to 5.

14. A computer program product, the program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the text processing method according to any one of claims 1 to 5.

15. An electronic device, characterized in that, The electronic device is configured to execute the steps of the text processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Identification method, device, server group and storage medium for noise words in text

    CN108304387A

  • Method, device and equipment for cleaning hidden characters in word document

    CN114239505A

  • Text recognition method and device and medium

    CN116110062A

  • Text extraction method and device

    CN116450812A

  • Energy big data cleaning method, device and equipment and storage medium

    CN116521654A