Corpus filtering method and device, equipment and medium
By preprocessing the filtered corpus and building a prefix tree, calculating the repetition rate of the corpus for filtering, the problem of low text filtration efficiency in the prior art is solved, and efficient text filtration is achieved.
Patent Information
- Application Number
- CN202311699754.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, text filtration efficiency is low, and a large number of rules and templates are required to be written manually, resulting in high labor costs and low filtration efficiency.
By preprocessing the filtered corpus, dividing it into multiple subtexts, and building a prefix tree, calculating the repetition rate of the corpus based on the number of nodes and the number of subtext characters in the prefix tree, and then filtering.
It improves the efficiency of text filtering, reduces the need for manual writing rules, and reduces labor costs.
Smart Images

Figure CN120144742A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a corpus filtering method, device, equipment and medium. Background Art
[0002] Data preprocessing is one of the important tasks for researchers in natural language processing. Its purpose is to process and screen the original text data to remove noise, errors and redundant content therein, so as to improve the data quality and accuracy of subsequent natural language processing tasks. For example, in natural language processing tasks such as text classification, question answering, dialogue, reading comprehension and entity extraction, we need to preprocess text data from different sources and in different forms to obtain clean high-quality corpus.
[0003] In the prior art, the corpus is generally filtered by means of rule matching based on regular expressions, etc. However, the defect of this method is that the labor cost is high, a large number of rules and templates need to be written manually, and the filtering efficiency is low. Therefore, when preprocessing text data from different sources and in different forms, how to improve the text filtering efficiency has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a corpus filtering method, device, equipment and medium to solve the problem of low text filtering efficiency.
[0005] In a first aspect, an embodiment of the present invention provides a corpus filtering method, which is characterized in that the corpus filtering method includes:
[0006] Preprocess the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero;
[0007] Obtain the string corresponding to each sub-text, and construct a prefix tree according to the strings of each sub-text. Wherein, a node in the prefix tree stores a character, and the node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents the any string, and the strings starting with the same character share the node corresponding to the same character;
[0008] Obtain the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters of all sub-texts, and filter the corpus to be filtered according to the repetition rate.
[0009] In a second aspect, an embodiment of the present invention provides a corpus filtering device, and the corpus filtering device includes:
[0010] A preprocessing module, configured to preprocess the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero;
[0011] A building block for obtaining a string corresponding to each sub - text, and constructing a prefix tree according to the strings of each sub - text. In the prefix tree, a node stores a character, and a node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents the any string, and strings starting with the same character share the node corresponding to the same character;
[0012] A filtering module for obtaining the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters of all sub - texts, and filtering the corpus to be filtered according to the repetition rate.
[0013] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the corpus filtering method as described in the first aspect.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer - readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the corpus filtering method as described in the first aspect.
[0015] The beneficial effects of the present invention compared with the prior art are as follows:
[0016] Pre - process the obtained corpus to be filtered to obtain N sub - texts, where N is an integer greater than zero. Obtain the string corresponding to each sub - text, and construct a prefix tree according to the strings of each sub - text. In the prefix tree, a node stores a character, and a node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents any string, and strings starting with the same character share the node corresponding to the same character. Obtain the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters of all sub - texts, and filter the corpus to be filtered according to the repetition rate. In this application, the text to be filtered is segmented into multiple sub - texts, the string corresponding to each sub - text is obtained, and a prefix tree is constructed according to the strings of each sub - text. For strings, by using the characteristic that the prefix tree can share common prefixes, each character in the corpus to be filtered can be quickly traversed, so that according to the calculated repetition rate of the characters in the corpus to be filtered, the corpus to be filtered can be quickly filtered, thereby improving the filtering efficiency. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of an application environment of a corpus filtering method provided by an embodiment of the present invention;
[0019] Figure 2 It is a schematic flowchart of a corpus filtering method provided by an embodiment of the present invention;
[0020] Figure 3 It is a schematic structural diagram of a corpus filtering device provided by an embodiment of the present invention;
[0021] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0023] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0024] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0025] It should also be understood that the term "and / or" used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] As used in the specification of the present invention and the appended claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0027] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0028] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0029] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results.
[0030] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0031] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0032] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.
[0033] A corpus filtering method provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server. Among them, the client includes but is not limited to computer devices such as palmtop computers, desktop computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, cloud computer devices, and personal digital assistants (PDAs). The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0034] See Figure 2 , which is a schematic flowchart of a corpus filtering method provided by an embodiment of the present invention. The above corpus filtering method can be applied to Figure 1 the server in Figure 2 as shown, and the corpus filtering method can include the following steps.
[0035] S201: Preprocess the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero.
[0036] In step S201, the obtained corpus to be filtered is preprocessed to obtain N sub-texts. Among them, the corpus to be filtered is the text in some training data. The preprocessing can perform segmentation processing on the corpus to be filtered. The sub-texts are short texts with shorter lengths after segmentation of the corpus to be filtered, and the corpus to be filtered can be segmented into multiple sub-texts.
[0037] In this embodiment, the corpus, that is, language materials, has multiple formats, such as voice format, text format, etc. The corpus to be filtered refers to the corpus in text format. In a possible implementation manner, the terminal obtaining the corpus to be filtered may include the terminal using web crawler technology to crawl the corpus to be filtered from the web page. Alternatively, after obtaining user authorization, the terminal records the user's usual voice and converts the voice into text to obtain multiple corpora to be filtered. Alternatively, the terminal receives multiple corpora to be filtered sent by the server or other terminals. Among them, the manner in which the server or other terminals obtain the multiple corpora to be filtered is the same as the manner in which this terminal obtains the multiple corpora to be filtered. Since the corpus to be filtered on the web page is published by the user on the web page, and the recorded voice is also generated by the user during language communication, the above-mentioned corpus to be filtered is language material that has actually appeared in the use of language. The multiple corpora to be filtered obtained can be used to train the language model. Considering that there may be a large number of duplicate corpora to be filtered in the multiple corpora to be filtered, these large numbers of duplicate corpora to be filtered will affect the training effect of the language model. Therefore, after obtaining the multiple corpora to be filtered, the terminal performs the following steps to filter the multiple corpora to be filtered to obtain the corpus to be filtered that meets the requirements of language model training.
[0038] In addition, when crawling the web page corpus to be filtered from the Internet, web page data can be crawled from forums, online encyclopedias, and the web. Preprocessing is performed on the obtained corpus to be filtered, that is, the corpus to be filtered. Among them, the preprocessing can segment the corpus to be filtered. When segmenting, the corpus to be filtered can be segmented into sub-texts within a preset length range. For example, the preset length range can be 64 characters - 512 characters. The preset length can be 64 characters, 128 characters, 256 characters, or 512 characters, or other character lengths within 64 - 512 characters can also be selected according to the actual application scenario. Alternatively, the corpus to be filtered can be segmented into equal-length texts, which is not limited in this embodiment.
[0039] It should be noted that when segmenting the corpus to be filtered into equal-length texts, the sliding window method can be used for segmentation. For example, if the corpus to be filtered is segmented step by step one by one according to lengths from 1 to 4, multiple sub-texts are obtained. For example, if the character length of the corpus to be filtered on a line is 10, when N = 1 and the window with a length of 1 is used for sliding segmentation once, 10 sub-texts are obtained; when N = 2 and the window with a length of 2 is used for sliding segmentation once, 9 sub-texts are obtained; when N = 3 and the window with a length of 3 is used for sliding segmentation once, 8 sub-texts are obtained; when N = 4 and the window with a length of 4 is used for sliding segmentation once, 7 sub-texts are obtained.
[0040] In this embodiment, preprocessing is performed on the corpus to be filtered to obtain N sub-texts, converting the long text into multiple short texts, which is convenient for constructing a prefix tree based on the short texts and calculating the corresponding repetition rate.
[0041] Optionally, before preprocessing the obtained corpus to be filtered to obtain N sub-texts, it further includes:
[0042] Removing texts that are different from the preset text types in the obtained original text to obtain the removed text, and determining the removed text as the corpus to be filtered.
[0043] In this embodiment, before preprocessing the obtained corpus to be filtered to obtain N sub-texts, it further includes removing texts that are different from the preset text types in the obtained original text to obtain the removed text, and determining the removed text as the corpus to be filtered. If the original text contains two or more text types, the different text types are removed, or the different text types are converted into the preset text type. For example, if the original text contains Chinese text type and English text type, and filtering the Chinese text type, the English text type needs to be removed, and the removed English text type is used as the corpus to be filtered, or the English text type is converted into the Chinese text type for filtering.
[0044] In this embodiment, removing texts that are different from the preset text types in the obtained original text to obtain the removed text, and determining the removed text as the corpus to be filtered, so that the characters in the nodes of the prefix tree constructed according to the sub-texts are of the same character type, avoiding the appearance of different types of characters, increasing the number of nodes in the prefix tree, and improving the calculation accuracy of the repetition rate.
[0045] Optionally, preprocessing the obtained corpus to be filtered to obtain N sub-texts includes:
[0046] Obtaining the splitting punctuation in the corpus to be filtered;
[0047] Splitting the filtered corpus at the splitting punctuation to obtain multiple split N sub-texts.
[0048] In this embodiment, obtaining the splitting punctuation in the corpus to be filtered and using the splitting punctuation as the basis for splitting. The splitting punctuation may include commas, periods, colons, semicolons, exclamation marks, question marks, dashes, ellipses, spaces, tab characters, and line breaks, etc. Splitting the filtered corpus at the splitting punctuation to obtain multiple split N sub-texts. For example, for "In the world of table tennis, he is the world champion and has won the singles world championship three times in a row; but he was once lost in the political whirlpool;", the split sub-texts are "In the world of table tennis", "He is the world champion", "Has won the singles world championship three times in a row", and "But he was once lost in the political whirlpool". The split sub-texts do not include the corresponding splitting punctuation, but only include the corresponding text information.
[0049] In this embodiment, the punctuation marks for segmentation are used as the segmentation markers, and the text is segmented at the punctuation marks for segmentation, so that the segmented sub-texts can represent complete semantic information.
[0050] Optionally, after preprocessing the obtained corpus to be filtered to obtain N sub-texts, the following steps are further included:
[0051] For any sub-text, obtain the string corresponding to each sub-text, and calculate the character length of the string;
[0052] If the character length is greater than the preset length threshold, filter the corpus to be filtered.
[0053] In this embodiment, for any sub-text, obtain the string corresponding to each sub-text, and calculate the character length of the string. If the character length is too long, it indicates that the quality of the corpus to be filtered is low, and the corpus to be filtered should be discarded. If the character length is greater than the preset length threshold, filter the corpus to be filtered. In this embodiment, the preset length threshold is set to 256. If the character length is greater than the preset length threshold, filter the corpus to be filtered.
[0054] In this implementation, for any sub-text, obtain the string corresponding to each sub-text, and calculate the character length of the string. If the character length is greater than the preset length threshold, filter the corpus to be filtered. Determine whether the quality of the corpus to be filtered is low based on the character length of the sub-text. If the quality of the corpus to be filtered is low, directly perform filtering processing without the need to calculate the repetition rate subsequently, improving the filtering efficiency.
[0055] S202: Obtain the string corresponding to each sub-text, and construct a prefix tree according to the strings of each sub-text. Among them, one node in the prefix tree stores one character, and the node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents any string, and the strings starting with the same character share the nodes corresponding to the same character.
[0056] In step S202, obtain the string corresponding to each sub-text, and construct a prefix tree according to the strings of each sub-text. Among them, the prefix tree is a tree-shaped data structure for storing strings. Each node of the tree stores one character in the string composed of the reading comprehension text. Multiple nodes on each path can be combined to form a complete string. Among them, the character refers to a glyph unit or symbol, including letters, numbers, and other symbols, as well as some functional symbols. A character is a general term for letters, numbers, and symbols in electronic computers or radio communications, and is the smallest data access unit in a data structure. A string is a sequence composed of multiple characters.
[0057] In this embodiment, a prefix tree is constructed for N sub-texts, where the prefix tree can be different types of prefix trees. For example, it can be a prefix tree or a compressed prefix tree. If it is a prefix tree, a node in the prefix tree stores a character, and a node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents any string, and the strings starting with the same character share the nodes corresponding to the same character. If it is a compressed prefix tree, the prefix tree is composed of multiple nodes connected by a parent-child node relationship, and each edge of the prefix tree represents a string containing at least one character, and the string corresponds to the state transition from the parent node to the child node of the edge.
[0058] In this embodiment, a prefix tree is constructed according to the strings of each sub-text. Each node in the prefix tree stores a character. During the storage process, the nodes in the same character can be merged, reducing the storage efficiency.
[0059] Optionally, obtain the string corresponding to each sub-text, and construct a prefix tree according to the strings of each sub-text, including:
[0060] Construct the root node of the prefix tree, and no character is stored in the root node;
[0061] Select any string as the first string, and construct a node tree connecting the root node according to the characters and their arrangement order in the first string. One node in the node tree stores one character;
[0062] Select any string from the remaining strings as the second string, and perform node matching on the existing node tree in turn according to the characters and their arrangement order in the second string;
[0063] If the first M characters in the second string match the corresponding nodes on the existing node tree, and the (M + 1)-th character does not match the node, then construct a new node tree. The nodes in the new node tree store the (M + 1)-th character to the last character in the second string in turn, and the first node of the new node tree is connected to the node that the M-th character matches on the existing node tree, where M is an integer greater than zero;
[0064] If the first M characters in the second string do not match the corresponding nodes on the existing node tree, then construct a new node tree connecting the root node according to the characters and their arrangement order in the second string. One node in the node tree stores one character;
[0065] Loop to select any string from the remaining strings as the second string, and perform node matching on the existing node tree in turn according to the characters and their arrangement order in the second string until all strings are traversed to obtain the prefix tree constructed by all node trees.
[0066] In this embodiment, the root node of the prefix tree is constructed. There is no character stored in the root node. When constructing the prefix tree, first construct the root node which does not store characters. Then construct the node tree of the root node behind the root node. One node in the node tree stores one character. According to the strings corresponding to each sub-text obtained, select one of the strings as the first string. According to the characters in the first string and their arrangement order, construct the node tree connecting the root node. One node in the node tree stores one character. Store the first character in the first string into the first node in the node tree, store the second character in the first string into the second node in the node tree, and so on until the last character in the first string. For example, if the string contains 10 characters, then the node tree contains 10 nodes which store the corresponding characters respectively.
[0067] Select any one of the remaining strings as the second string. According to the characters in the second string and their arrangement order, perform node matching on the existing node tree in sequence. If the first M characters in the second string match the corresponding nodes on the existing node tree, and the (M + 1)-th character does not match any node, then construct a new node tree. The nodes in the new node tree store the (M + 1)-th character to the last character in the second string in sequence, and the first node of the new node tree is connected to the node that the M-th character matches on the existing node tree, where M is an integer greater than zero.
[0068] For example, select the second string, compare the characters in the second string with those in the first string to determine whether they are equal to the first character in the first string. If they are equal, use the node where the first character in the first string is located as the node for the first character in the second string at the same time, that is, the first character in the second string shares a node with the character in the first string. After storing the first character in the second string, continue to compare the second character in the second string with the second character in the first string. If they are equal, use the node where the second character in the first string is located as the node for the second character in the second string at the same time. If they are not equal, construct a new node tree behind the node corresponding to the first character in the second string. The first node in the new node tree stores the second character in the first string, and the second node stores the third character in the second string. Store the remaining characters in the second string in sequence in the new node tree.
[0069] If the first M characters in the second string do not match the corresponding nodes in the existing node tree, a new node tree connecting to the root node is constructed according to the characters in the second string and their arrangement order. One character is stored in each node of the node tree. For example, compare the characters in the second string with those in the first string to determine whether they are equal to the first character in the first string. If not, continue to construct a new node tree behind the root node to store the first character in the second string. According to the order of the characters in the second string, store the corresponding characters into the nodes in the corresponding new node tree in sequence until all the characters in the second string are stored.
[0070] Loop to select any string from the remaining strings as the second string, and perform node matching on the existing node tree in sequence according to the characters in the second string and their arrangement order until all the strings are traversed to obtain the prefix tree constructed by all the node trees. For example, select the third string, compare the third string with the characters in the first string and the characters in the second string respectively to determine whether they are equal to the first character in the first string and the first character in the second string. If not, continue to construct a new node tree behind the root node, and the nodes in the node tree store the characters in the third string in sequence. If it is equal to the first character in the first string, use the node where the first character in the first string is located as the node for the first character in the third string at the same time. If it is equal to the first character in the second string, use the node where the first character in the second string is located as the node for the first character in the third string at the same time.
[0071] After storing the first character in the third string, if the first character in the third string is stored in the node where the first character in the first string is located, continue to compare the second character in the third string with the second character in the first string to determine whether they are equal. If it is equal to the second character in the first string, use the node where the second character in the first string is located as the node for the second character in the third string at the same time. If not, construct a new node tree, and the nodes in the new node tree store the second character in the third string and all the characters after the second character in sequence. Traverse each string to obtain the prefix trees corresponding to the strings in the N sub-texts.
[0072] It should be noted that after the prefix tree is constructed, it is also possible to check whether all sub-texts have been completely constructed or whether there are construction errors in the constructed prefix tree. When checking, each node tree in the prefix tree is queried in turn. First, check the nodes of the node tree connected to the root node. The number of nodes of the node tree connected to the root node should be equal to the number of different characters in the N sub-texts. For example, if the first characters of the strings corresponding to 10 sub-texts are a, a, a, b, b, c, c, d, d, d respectively, the number of unequal characters is 4. Therefore, the number of nodes of the node tree connected to the root node should be 4 nodes. If it is greater than 4, it is considered that there is an error in constructing the prefix tree and the same characters are not merged. If it is less than 4, it is considered that there is a string of sub-text that has not been constructed into a node tree. Then check the number of node trees in the prefix tree. Among them, the branch connecting from the root node to the last node in sequence is the node tree. Each node tree in the prefix tree represents a string, and the characters stored from the root node to the last node are in the same order as the characters in the string. The number of node trees should be equal to the number of different sub-texts. For example, if three of the 10 sub-texts have exactly the same characters, there should be 8 node trees in the prefix tree, that is, starting from the root node, there are 8 branches to the last node. If it is less than 8, it is considered that there is a string of sub-text that has not been constructed into a node tree. If it is greater than 8, it is considered that there are exactly the same node trees, and in this case, one of the node trees should be discarded.
[0073] In this embodiment, by constructing a prefix tree, the strings in the N sub-texts are divided into different node trees, and then the corresponding number of nodes is determined. When calculating the corresponding repetition rate, it is not necessary to calculate the repetition rates of different sub-texts multiple times and then integrate them to obtain the final repetition rate, thereby improving the efficiency of calculating the repetition rate.
[0074] S203: Obtain the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters in all sub-texts, and filter the corpus to be filtered according to the repetition rate.
[0075] In step S203, the repetition rate of the corpus to be filtered is obtained according to the number of nodes in the prefix tree and the number of characters in all sub-texts, and the corpus to be filtered is filtered according to the repetition rate. Among them, the number of characters in all sub-texts is the number of characters after removing the segmentation punctuation in the corpus to be filtered. Filtering the corpus to be filtered means determining the filtering method for the corpus to be filtered according to the repetition rate.
[0076] In this embodiment, according to the number of nodes in the prefix tree and the number of characters in all sub-texts, the repetition rate of the corpus to be filtered is obtained. When calculating the repetition rate, it can be calculated according to the ratio of the number of nodes in the prefix tree to the number of characters in all sub-texts, and 1 minus the ratio of the number of nodes in the prefix tree to the number of characters in all sub-texts is used as the repetition rate.
[0077] In another embodiment, when calculating the repetition rate, the corresponding repetition rate can also be calculated by calculating the number of node trees in the prefix tree. First, determine the number of node trees in the prefix tree, that is, the number of branches in the prefix tree. Determine the number of duplicate sub-texts according to the difference between the number of branches and the number of sub-texts. Among them, the duplicate sub-texts can be the same sub-text or different sub-texts. According to the number of nodes and the number of branches in the prefix tree, calculate the average number of nodes in each branch. The ratio of the product of the average number of nodes in each branch and the number of duplicate sub-texts to the number of characters in all sub-texts can be used as the repetition rate.
[0078] For example, if there are 10 sub-texts, and the number of node trees in the prefix tree is 8, that is, there are 8 branch trees, then there are two sub-texts whose string prefix characters are equal to the strings in the 8 branch trees. Calculate the average number of nodes in each of the 8 branch trees, multiply the average number of nodes by two, that is, the difference between the number of sub-texts and the number of node trees, and use the ratio of the corresponding product to the number of characters in all sub-texts as the repetition rate.
[0079] According to the repetition rate, filter the corpus to be filtered. In this embodiment, when the repetition rate is greater than the preset threshold, filter the corpus to be filtered, that is, delete the corpus to be filtered. When the repetition rate is less than the preset threshold, retain the corpus to be filtered.
[0080] In another embodiment, when filtering the corpus to be filtered according to the repetition rate, the corpus to be filtered can also be de-duplicated so that the repetition rate after de-duplication of the corpus to be filtered is less than the preset threshold, and the de-duplicated corpus to be filtered is retained.
[0081] Optionally, obtaining the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters in all sub-texts includes:
[0082] Calculate the proportional relationship between the number of nodes and the number of characters according to the number of nodes in the prefix tree and the number of characters in all sub-texts;
[0083] Calculate the repetition rate of the corpus to be filtered according to the proportional relationship.
[0084] In this embodiment, the repetition rate of the corpus to be filtered is calculated according to the number of nodes in the prefix tree and the number of characters in all sub-texts. The calculation formula is as follows:
[0085]
[0086] Among them, k is the repetition rate, q is the number of nodes in the prefix tree, and Q is the number of characters in all sub-texts. When calculating the repetition rate, if the prefix characters of the strings of the sub-texts are all different, the number of nodes in the prefix tree is equal to the number of characters in all sub-texts, and the corresponding repetition rate should be 0. If there are strings with the same prefix characters among the prefix characters of the strings of the sub-texts, the number of nodes in the prefix tree is less than the number of characters in all sub-texts, and the corresponding repetition rate is between 0 and 1. As there are more identical prefix characters among the prefix characters of the strings of the sub-texts, the number of nodes in the corresponding prefix tree is less, and the repetition rate value is larger.
[0087] In this embodiment, when calculating the repetition rate of the corpus to be filtered based on the number of nodes in the prefix tree and the number of characters in all sub-texts, it is not necessary to find the identical texts in each short text in the corpus to be filtered, nor to find the longest repeated text and the corresponding occurrence times in the corpus to be filtered. The repetition rate can be directly calculated based on the number of nodes in the prefix tree and the number of characters in all sub-texts, thereby improving the calculation efficiency of the repetition rate. And when calculating the repetition rate, all the characters in the sub-text are used, avoiding the influence of the punctuation marks for segmentation in the corpus to be filtered, thereby improving the calculation accuracy.
[0088] Optionally, filtering the corpus to be filtered according to the repetition rate includes:
[0089] If the number of samples in the training sample where the corpus to be filtered is located is greater than the preset number threshold, then determine whether the repetition rate is greater than the first preset threshold. If the repetition rate is greater than the first preset threshold, then filter the corpus to be filtered;
[0090] If the number of samples in the training sample where the corpus to be filtered is located is not greater than the preset number threshold, then determine whether the repetition rate is greater than the second preset threshold. If the repetition rate is greater than the second preset threshold, then locate the repeated texts in the corpus to be filtered. According to the location, filter the repeated texts to obtain the filtered corpus, and the repetition rate in the filtered corpus is between the first preset threshold and the second preset threshold.
[0091] In this embodiment, if the number of samples in the training sample where the corpus to be filtered is located is greater than a preset number threshold, then it is determined whether the repetition rate is greater than a first preset threshold. If the repetition rate is greater than the first preset threshold, the corpus to be filtered is filtered. The corpus to be filtered is the text data in the training sample. When processing the corpus to be filtered, it can be processed according to the number of samples in the training sample. If the number of samples in the training sample where the corpus to be filtered is located is greater than the preset number threshold, that is, the number of samples in the training sample is large, then it is determined whether the repetition rate is greater than the first preset threshold. If the repetition rate is greater than the first preset threshold, the corpus to be filtered is filtered. That is, when the number of samples in the training sample is large, it is directly determined whether to filter the corpus to be filtered according to the repetition rate. If the repetition rate is greater than the first preset threshold, all the corpus to be filtered is filtered out.
[0092] If the number of samples in the training sample where the corpus to be filtered is located is not greater than the preset number threshold, that is, when the training sample is small, it is not possible to directly filter out all the corpus to be filtered. Then it is determined whether the repetition rate is greater than a second preset threshold. If the repetition rate is greater than the second preset threshold, the repeated texts in the corpus to be filtered are located, and according to the location, the repeated texts are filtered to obtain the filtered corpus. The repetition rate in the filtered corpus is between the first preset threshold and the second preset threshold.
[0093] It should be noted that when locating the repeated texts in the corpus to be filtered and filtering the repeated texts according to the location, the strings in the N sub-texts can be arranged in a row in sequence, and the N sub-texts are sorted so that the same sub-texts are located in adjacent rows, and the repeated rows in each sub-text group are deleted. Before sorting, the same sub-texts may not be located in adjacent rows. After sorting, all the same sub-texts are located in adjacent rows. After de-duplicating the sorted sub-text groups, that is, after filtering the repeated texts, the repeated rows are deleted, and there are no repeated sub-texts in the sub-texts.
[0094] It should be noted that when sorting the sub-texts, the N sub-texts can be sorted by Linux sort (a sorting command) so that the same sub-texts are located in adjacent rows. The repeated texts in the N sub-texts can be filtered out by Linux uniq (a de-duplication command).
[0095] In this embodiment, when the repetition rate in the corpus to be filtered is high and the number of the training sample where the corpus to be filtered is located is small, the repeated texts in the corpus to be filtered are filtered, and the remaining texts in the corpus to be filtered are retained, so that the remaining texts can be used as training samples for model training to provide sufficient training samples for model training. When filtering, the sorting command and the de-duplication command are used, and the repeated sub-texts can be quickly filtered out, improving the filtering efficiency.
[0096] In another embodiment, when filtering duplicate texts, the simhash (Locality Sensitive Hashing) algorithm can be used. The main work of the simhash algorithm can be understood as dimensionality reduction of texts to generate a fingerprint. By comparing the Hamming distances of the fingerprints of different texts, the similarity between two texts can be determined. Among them, the Hamming distance refers to the number of different bits in the corresponding bits of two legal codes in information coding, which is called the code distance, and the number of different bits in the corresponding bits of two codewords is called the Hamming distance between the two codewords. In this embodiment, the fingerprint refers to the hash calculation of each sub-text to obtain a characteristic fingerprint value, and each sub-text corresponds to a string of characteristic fingerprint values. The characteristic fingerprint value is a string of codes, for example: 0111010000111100. The characteristic fingerprint values are binned, and the multiple segments of the to-be-query characteristic fingerprint values obtained after binning are respectively retrieved and filtered.
[0097] It should be noted that in this embodiment, the above binning can refer to segmenting the characteristic fingerprint values. After dividing them into several segments, each segment is filtered separately. Retrieval can be performed based on all the characteristic fingerprint values obtained after binning to determine whether there are identical parts in each segment of the to-be-query characteristic fingerprint values after segmentation of the same characteristic fingerprint value. Only when there are no identical parts retrieved in each segment of the to-be-query characteristic fingerprint values in the same characteristic fingerprint value can the corresponding sub-text be retained; if identical parts are retrieved in any segment of the to-be-query characteristic fingerprint values in the characteristic fingerprint value, it can be determined that the sub-text corresponding to the characteristic fingerprint value may have the same semantic content and needs to be filtered. Specifically, filtering can be performed based on the difference points of the characteristic fingerprint values.
[0098] In this embodiment, the characteristic fingerprint values in the sub-texts are extracted through the simhash algorithm, the characteristic fingerprint values are binned, and the multiple segments of the to-be-query characteristic fingerprint values obtained after binning are respectively retrieved and filtered, reducing the retrieval content and improving the filtering efficiency.
[0099] It should be noted that after filtering out the duplicate texts, the duplication rate of the filtered corpus is recalculated to make the duplication rate of the filtered corpus between the first preset threshold and the second preset threshold.
[0100] In another embodiment, after filtering out the duplicate texts, the duplication rate of the filtered corpus is recalculated, and the duplication rate of the filtered corpus can also be made less than the first preset threshold.
[0101] Preprocess the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero. Obtain the string corresponding to each sub-text. According to the strings of each sub-text, construct a prefix tree. In the prefix tree, a node stores a character. The node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents any string, and the strings starting with the same character share the nodes corresponding to the same character. According to the number of nodes in the prefix tree and the number of characters in all sub-texts, obtain the repetition rate of the corpus to be filtered. According to the repetition rate, perform filtering processing on the corpus to be filtered. In this application, the text to be filtered is segmented into multiple sub-texts, the string corresponding to each sub-text is obtained, and a prefix tree is constructed according to the strings of each sub-text. For the strings, using the characteristic that the prefix tree can share a common prefix, each character in the corpus to be filtered can be quickly traversed, so that according to the calculated repetition rate of the characters in the corpus to be filtered, the corpus to be filtered can be quickly filtered, thereby improving the filtering efficiency.
[0102] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a corpus filtering device provided by an embodiment of the present invention. In this embodiment, each unit included in the terminal is used to execute Figure 2 the corresponding steps in the embodiment. Specifically, please refer to Figure 2 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 3 , the corpus filtering device 30 includes: a preprocessing module 31, a construction module 32, and a filtering module 33.
[0103] The preprocessing module 31 is used to preprocess the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero.
[0104] The construction module 32 is used to obtain the string corresponding to each sub-text, and construct a prefix tree according to the strings of each sub-text. In the prefix tree, a node stores a character. The node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents any string, and the strings starting with the same character share the nodes corresponding to the same character.
[0105] The filtering module 33 is used to obtain the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters in all sub-texts, and perform filtering processing on the corpus to be filtered according to the repetition rate.
[0106] Optionally, the above-mentioned corpus filtering device 30 further includes:
[0107] a removal module, which is used to remove the texts different from the preset text type in the obtained original text to obtain the removed text, and determine the removed text as the text to be filtered.
[0108] Optionally, the above-mentioned corpus filtering device 30 further includes:
[0109] A calculation module, configured to obtain the string corresponding to each sub-text for any sub-text, and calculate the character length of the string.
[0110] A judgment module, configured to filter the corpus to be filtered if the character length is greater than a preset length threshold.
[0111] Optionally, the above-mentioned preprocessing module 31 includes:
[0112] An acquisition unit, configured to acquire the segmentation punctuation in the corpus to be filtered.
[0113] A segmentation unit, configured to segment the filtered corpus at the segmentation punctuation to obtain multiple segmented N sub-texts.
[0114] Optionally, the above-mentioned construction module 32 includes:
[0115] A root node construction unit, configured to construct the root node of the prefix tree, and no characters are stored in the root node.
[0116] A first selection unit, configured to select any string as the first string, and construct a node tree connecting the root node according to the characters and their arrangement order in the first string, and one character is stored in one node of the node tree.
[0117] A second selection unit, configured to select any string from the remaining strings as the second string, and sequentially perform node matching on the existing node tree according to the characters and their arrangement order in the second string.
[0118] A first new node tree construction unit, configured to construct a new node tree if the first M characters in the second string match the corresponding nodes on the existing node tree and the (M + 1)-th character does not match the node. The nodes in the new node tree sequentially store the (M + 1)-th character to the last character in the second string, and the first node of the new node tree is connected to the node where the M-th character matches on the existing node tree, and M is an integer greater than zero.
[0119] A second new node tree construction unit, configured to construct a new node tree connecting the root node according to the characters and their arrangement order in the second string if the first M characters in the second string do not match the corresponding nodes on the existing node tree, and one character is stored in one node of the node tree.
[0120] A loop unit, configured to loop through selecting any string from the remaining strings as the second string, and sequentially perform node matching on the existing node tree according to the characters and their arrangement order in the second string until all strings are traversed to obtain the prefix tree constructed by all node trees.
[0121] Optionally, the above filtering module 33 includes:
[0122] A first calculation unit, configured to calculate a proportional relationship between the number of nodes and the number of characters of all sub-texts according to the number of nodes in the prefix tree and the number of characters of all sub-texts.
[0123] A second calculation unit, configured to calculate the repetition rate of the corpus to be filtered according to the proportional relationship.
[0124] Optionally, the above filtering module 33 includes:
[0125] A first judgment unit, configured to, if the number of samples in the training sample where the corpus to be filtered is located is greater than a preset number threshold, judge whether the repetition rate is greater than a first preset threshold, and if the repetition rate is greater than the first preset threshold, filter the corpus to be filtered.
[0126] A second judgment unit, configured to, if the number of samples in the training sample where the corpus to be filtered is located is not greater than the preset number threshold, judge whether the repetition rate is greater than a second preset threshold, and if the repetition rate is greater than the second preset threshold, locate the repeated text in the corpus to be filtered, and filter the repeated text according to the location to obtain the filtered corpus, where the repetition rate in the filtered corpus is between the first preset threshold and the second preset threshold.
[0127] It should be noted that, for the information interaction, execution process, etc. between the above modules, units, and subunits, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0128] Figure 4 This is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in
[0129] ), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in any of the above-mentioned corpus filtering method embodiments are implemented. Figure 4 The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that
[0130] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0131] The memory includes a readable storage medium, internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. The memory may also be used to temporarily store the data that has been output or will be output.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0133] All or part of the processes in the above-mentioned method embodiments of the present invention can also be completed by a computer program product. When the computer program product runs on a computer device, the computer device can be made to execute the steps in the above-mentioned method embodiments when executed.
[0134] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0135] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0136] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.
[0137] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A corpus filtering method, characterized in that, the corpus filtering method includes: preprocessing the to-be-filtered corpus obtained to obtain N sub-texts, where N is an integer greater than zero; obtaining the string corresponding to each sub-text, and constructing a prefix tree according to the strings of each sub-text. Among them, one node in the prefix tree stores one character, and the node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents the any string, and the strings starting with the same character share the node corresponding to the same character; obtaining the repetition rate of the to-be-filtered corpus according to the number of nodes in the prefix tree and the number of characters of all sub-texts, and filtering the to-be-filtered corpus according to the repetition rate.
2. The corpus filtering method according to claim 1, characterized in that, the obtaining the string corresponding to each sub-text and constructing a prefix tree according to the strings of each sub-text includes: constructing the root node of the prefix tree, and no character is stored in the root node; selecting any string as the first string, and constructing a node tree connecting the root node according to the characters and their arrangement order in the first string, and one node in the node tree stores one character; selecting any string from the remaining strings as the second string, and sequentially performing node matching on the existing node tree according to the characters and their arrangement order in the second string; if the first M characters in the second string match the corresponding nodes on the existing node tree, and the (M + 1)-th character does not match the node, then construct a new node tree, and the nodes in the new node tree sequentially store the (M + 1)-th character to the last character in the second string, and the first node of the new node tree is connected to the node matched by the M-th character on the existing node tree, where M is an integer greater than zero; if the first M characters in the second string do not match the corresponding nodes on the existing node tree, then construct a new node tree connecting the root node according to the characters and their arrangement order in the second string, and one node in the node tree stores one character; circularly select any string from the remaining strings as the second string, and sequentially perform node matching on the existing node tree according to the characters and their arrangement order in the second string until all strings are traversed to obtain the prefix tree constructed by all node trees.
3. The corpus filtering method according to claim 1, characterized in that, the obtaining the repetition rate of the to-be-filtered corpus according to the number of nodes in the prefix tree and the number of characters of all sub-texts includes: calculating the proportional relationship between the number of nodes and the number of characters according to the number of nodes in the prefix tree and the number of characters of all sub-texts; calculating the repetition rate of the to-be-filtered corpus according to the proportional relationship.
4. The corpus filtering method according to claim 1, characterized in that, the filtering the to-be-filtered corpus according to the repetition rate includes: If the number of samples in the training sample where the corpus to be filtered is located is greater than a preset number threshold, determine whether the repetition rate is greater than a first preset threshold. If the repetition rate is greater than the first preset threshold, filter the corpus to be filtered. If the number of samples in the training sample where the corpus to be filtered is located is not greater than the preset number threshold, determine whether the repetition rate is greater than a second preset threshold. If the repetition rate is greater than the second preset threshold, locate the repeated texts in the corpus to be filtered, and filter the repeated texts according to the location to obtain the filtered corpus, where the repetition rate in the filtered corpus is between the first preset threshold and the second preset threshold.
5. The corpus filtering method according to claim 1, characterized in that, before preprocessing the obtained corpus to be filtered to obtain N sub-texts, it further includes: removing texts of a preset text type different from those in the obtained original text to obtain the removed text, and determining the removed text as the corpus to be filtered.
6. The corpus filtering method according to claim 1, characterized in that, preprocessing the obtained corpus to be filtered to obtain N sub-texts, including: obtaining the segmentation punctuation in the corpus to be filtered; performing segmentation on the filtered corpus at the segmentation punctuation to obtain multiple segmented N sub-texts.
7. The corpus filtering method according to claim 1, characterized in that, after preprocessing the obtained corpus to be filtered to obtain N sub-texts, it further includes: for any sub-text, obtaining the string corresponding to each sub-text and calculating the character length of the string; if the character length is greater than a preset length threshold, perform filtering processing on the corpus to be filtered.
8. A corpus filtering device, characterized in that, the corpus filtering device includes: a preprocessing module for preprocessing the obtained corpus to be filtered to obtain N sub-texts, where N is an integer greater than zero; a construction module for obtaining the string corresponding to each sub-text and constructing a prefix tree according to the strings of each sub-text, where a node in the prefix tree stores a character, and the node path formed by connecting the nodes corresponding to the characters in the order of the characters in any string represents the any string, and the strings starting with the same character share the node corresponding to the same character; a filtering module for obtaining the repetition rate of the corpus to be filtered according to the number of nodes in the prefix tree and the number of characters of all sub-texts, and performing filtering processing on the corpus to be filtered according to the repetition rate.
9. A computer device, characterized in that, the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the corpus filtering method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the corpus filtering method according to any one of claims 1 to 7.