Text processing method and apparatus
Patent Information
- Application Number
- CN202310028994.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-01-09
AI Technical Summary
[0004]然而,上述深度学习的方案速度较慢,过于耗费算力和内存,导致文本处理的效率低且准确性不高
[0021]本申请提供的文本处理方法,提取待处理文本中的目标文本段;基于目标文本段的字符顺序,对目标文本段进行分词,获得初始文本段和预设数量的初始分词,其中,初始文本段为目标文本段中除初始分词外剩余的文本段;将初始分词中的指定分词与初始文本段进行合并,获得更新后的目标文本段,并返回执行基于目标文本段的字符顺序,对目标文本段进行分词的步骤;在达到预设分词停止条件的情况下,获得待处理文本对应的分词集合。通过对目标文本段进行分词,获得初始文本段和预设数量的初始分词,将初始分词中的指定分词与初始文本段进行合并,对目标文本段进行更新,仅关注文本的局部语义,实现了高效、准确的文本处理。
Smart Images

Figure CN115994535B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a text processing method. This application also relates to a text processing apparatus, a computing device, and a computer-readable storage medium. Background Technology
[0002] With the development of Internet technology, in text processing tasks of Natural Language Processing (NLP), since the text content is usually large and the length is long, in order to facilitate users to obtain the effective information in the text, text segmentation can be performed before processing the text. Therefore, text segmentation has gradually become a research focus in natural language processing tasks.
[0003] In existing technologies, deep learning methods are typically used to transform the word segmentation problem into a sequence labeling problem, which labels the attributes of each character in the text to obtain the word segmentation results.
[0004] However, the aforementioned deep learning solutions are slow and consume too much computing power and memory, resulting in low efficiency and low accuracy in text processing. Summary of the Invention
[0005] In view of this, embodiments of this application provide a text processing method to address the technical deficiencies existing in the prior art. Embodiments of this application also provide a text processing apparatus, a computing device, and a computer-readable storage medium.
[0006] According to a first aspect of the embodiments of this application, a text processing method is provided, including:
[0007] Extract the target text segment from the text to be processed;
[0008] Based on the character order of the target text segment, the target text segment is segmented into words to obtain an initial text segment and a preset number of initial words. The initial text segment is the text segment remaining in the target text segment excluding the initial words.
[0009] The specified word segment in the initial word segmentation is merged with the initial text segment to obtain the updated target text segment, and the step of performing word segmentation on the target text segment based on the character order of the target text segment is returned.
[0010] When the preset segmentation stopping condition is met, the word set corresponding to the text to be processed is obtained.
[0011] According to a second aspect of the embodiments of this application, a text processing apparatus is provided, comprising:
[0012] The extraction module is configured to extract target text segments from the text to be processed;
[0013] The word segmentation module is configured to segment the target text segment based on the character order of the target text segment, and obtain an initial text segment and a preset number of initial words. The initial text segment is the text segment remaining in the target text segment excluding the initial words.
[0014] The merging module is configured to merge the specified words in the initial word segmentation with the initial text segment to obtain the updated target text segment, and return the steps of performing word segmentation on the target text segment based on the character order of the target text segment;
[0015] The acquisition module is configured to obtain the set of words corresponding to the text to be processed when the preset word segmentation stopping condition is met.
[0016] According to a third aspect of the embodiments of this application, a computing device is provided, comprising:
[0017] Memory and processor;
[0018] The memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions to implement the steps of the text processing method.
[0019] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the text processing method.
[0020] According to a fifth aspect of the present application, a chip is provided that stores a computer program, which, when executed by the chip, implements the steps of the text processing method.
[0021] The text processing method provided in this application extracts the target text segment from the text to be processed; based on the character order of the target text segment, it performs word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, wherein the initial text segment is the remaining text segment in the target text segment excluding the initial word segments; it merges the specified word segments from the initial word segments with the initial text segment to obtain an updated target text segment, and returns to the step of performing word segmentation on the target text segment based on the character order of the target text segment; and when a preset word segmentation stopping condition is met, it obtains the word set corresponding to the text to be processed. By performing word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, merging the specified word segments from the initial word segments with the initial text segment, and updating the target text segment, this method focuses only on the local semantics of the text, achieving efficient and accurate text processing. Attached Figure Description
[0022] Figure 1 This is a framework diagram of a text processing system provided in one embodiment of this application;
[0023] Figure 2 This is a flowchart of a text processing method provided in an embodiment of this application;
[0024] Figure 3 This is a flowchart illustrating a text processing method applied to the gaming field, provided in one embodiment of this application.
[0025] Figure 4 This is a schematic diagram of a text processing interface provided in an embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the structure of a text processing device provided in an embodiment of this application;
[0027] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this application. Detailed Implementation
[0028] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0029] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” used in one or more embodiments of this application refers to and includes any or all possible combinations of one or more associated listed items.
[0030] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this application, and similarly, second may also be referred to as first.
[0031] First, the terminology used in one or more embodiments of the present invention will be explained.
[0032] Term frequency (TF) refers to the number of times a given word appears in a document.
[0033] Optical Character Recognition (OCR) refers to the process of analyzing and recognizing textual data in image files to obtain text and layout information. In other words, it involves recognizing the text in an image and returning it as text.
[0034] Double-Array Trie: A double-array trie is a type of trie tree with low space complexity, used in word segmentation of languages with large character ranges (such as Chinese and Japanese). The principle of double array is that the trie tree, which originally required multiple arrays to represent, can be stored using only two data sets, which can greatly reduce space complexity.
[0035] Aho-Corasick automaton: The Aho-Corasick automaton is an extension of the trie algorithm and is a widely used algorithm for string manipulation.
[0036] Word cloud analysis: Word cloud analysis generates a visual word cloud by performing word frequency statistics on a text library. Compared to simple word frequency information, it is more suitable for use and display by non-professional data personnel.
[0037] With the rapid popularization of the Internet and smartphones, the amount of information that can be collected online has exploded. Traditional information processing and analysis methods are becoming increasingly inadequate. Therefore, it is necessary to introduce intelligent information processing and analysis methods based on data mining, machine learning, and even deep learning.
[0038] Taking Chinese text as an example, the first step in processing it using computer algorithms is usually word segmentation. The result of word segmentation is not only the basis for various subsequent algorithms, but it can also be directly processed into information such as word frequency for further analysis. The accuracy and effectiveness of word segmentation directly determine the accuracy and effectiveness of subsequent results.
[0039] It should be noted that the following three word segmentation schemes can be used to achieve text segmentation: The first is dictionary-based enumeration segmentation. Professional linguists can construct a large number of rules to assist in selecting the segmentation results. The second is to use machine learning methods, calculating the maximum probability of forming a word in the entire sentence to obtain the segmentation result. The third is to use deep learning / neural network methods, transforming the word segmentation problem into a sequence labeling problem (i.e., labeling each Chinese character in the sentence with its attributes such as the beginning / end / middle / single character of a word, etc.).
[0040] However, the aforementioned word segmentation schemes also have certain drawbacks. For example, the accuracy of machine learning methods in predicting word probability decreases in short sentences. As sentences become longer, the computation time of the algorithm also increases accordingly. Processing the entire corpus presents time pressure, and it is difficult to cope with the ever-increasing number of new words on the internet. Furthermore, deep learning / neural network methods are relatively slow and consume excessive computing power and memory.
[0041] To address the aforementioned issues, this application provides a text processing method that extracts a target text segment from the text to be processed; segments the target text segment into words based on its character order to obtain an initial text segment and a preset number of initial words, wherein the initial text segment is the remaining text segment in the target text segment excluding the initial words; merges a specified word from the initial words with the initial text segment to obtain an updated target text segment, and returns to the step of segmenting the target text segment into words based on its character order; and, upon reaching a preset segmentation stopping condition, obtains the word set corresponding to the text to be processed. By segmenting the target text segment to obtain an initial text segment and a preset number of initial words, merging the specified word from the initial words with the initial text segment, and updating the target text segment, this method focuses only on the local semantics of the text, achieving efficient and accurate text processing.
[0042] This application provides a text processing method. This application also relates to a text processing apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0043] See Figure 1 , Figure 1 This illustration shows a framework diagram of a text processing system according to an embodiment of the present application. The text processing system includes a server and a client.
[0044] The client is used to send the text to be processed to the server;
[0045] The server-side component extracts the target text segment from the text to be processed; it segments the target text segment into words based on the character order of the target text segment to obtain an initial text segment and a preset number of initial words, where the initial text segment is the remaining text segment in the target text segment excluding the initial words; it merges the specified words from the initial words with the initial text segment to obtain the updated target text segment, and returns the step of performing word segmentation on the target text segment based on the character order of the target text segment; when a preset word segmentation stopping condition is met, it obtains the word segmentation set corresponding to the text to be processed; and it sends the word segmentation set to the client.
[0046] It is worth noting that the text processing method provided in the embodiments of this application is generally executed by the server. However, in other embodiments of this application, the client may also have similar functions to the server, thereby executing the text processing method provided in the embodiments of this application. In other embodiments, the text processing method provided in the embodiments of this application may also be executed jointly by the client and the server.
[0047] The scheme of this application extracts the target text segment from the text to be processed; based on the character order of the target text segment, it performs word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, wherein the initial text segment is the remaining text segment in the target text segment excluding the initial word segments; the specified word segments in the initial word segments are merged with the initial text segment to obtain an updated target text segment, and the step of performing word segmentation on the target text segment based on the character order of the target text segment is returned; if a preset word segmentation stopping condition is met, the word set corresponding to the text to be processed is obtained. By performing word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, merging the specified word segments in the initial word segments with the initial text segment, and updating the target text segment, focusing only on the local semantics of the text, efficient and accurate text processing is achieved.
[0048] Figure 2 A flowchart of a text processing method according to an embodiment of this application is provided, which specifically includes the following steps:
[0049] Step 202: Extract the target text segment from the text to be processed.
[0050] In one or more embodiments of this application, a target text segment can be extracted from the text to be processed, and the target text segment can be further processed, thereby saving text processing time and improving text processing efficiency.
[0051] Specifically, the text to be processed is the object of text processing. The text to be processed can be text in different languages, such as English text, Chinese text, etc., and the selection is made according to the actual situation. This application embodiment does not impose any limitations on this. The core of this application embodiment lies in realizing text segmentation. For text in different languages, the segmentation process is basically the same. The segmentation process of Chinese text will be described in detail below.
[0052] It should be noted that there are multiple ways to obtain the text to be processed. In the first possible implementation of this application, the text to be processed can be received directly, or it can be obtained from a text library. In the second possible implementation of this application, an image to be processed can be obtained, and optical character recognition can be performed on the image to obtain the text to be processed. In the third possible implementation of this application, audio or video to be processed can be obtained, and speech-to-speech conversion can be performed on the audio or video to obtain the text to be processed.
[0053] In practical applications, there are multiple ways to extract the target text segment from the text to be processed, which can be selected specifically according to actual conditions. The embodiments of the present application do not make any limitation on this.
[0054] In a possible implementation of the present application, the target text segment in the text to be processed can be extracted by using a specific-domain thesaurus. That is, the step of extracting the target text segment from the text to be processed may include the following steps:
[0055] matching the text to be processed with the specific-domain thesaurus according to the character order of the text to be processed, and determining target word segmentation in the text to be processed, wherein the specific-domain thesaurus includes a plurality of specific-domain words;
[0056] taking the target word segmentation as a segmentation point, segmenting the text to be processed to obtain the target text segment.
[0057] Specifically, the character order of the text to be processed is the arrangement order of each character from front to back in the text to be processed. Assuming that the text to be processed is "Ni Hao Ya", the character order of the text to be processed is: the first character is "Ni", the second character is "Hao", and the third character is "Ya". The specific-domain thesaurus includes professional domain words that are focused on in specific projects. Taking the game field as an example, the specific-domain thesaurus includes proper nouns in games, titles for world construction, common idioms in player communities, and the like.
[0058] It should be noted that there are multiple ways to obtain the specific-domain thesaurus, which can be selected specifically according to actual conditions, and the embodiments of the present application do not make any limitation on this. In a possible implementation of the present application, the manually constructed specific-domain thesaurus can be directly obtained. In another possible implementation of the present application, dictionaries on the network can be collected, and these dictionaries can be filtered to obtain the specific-domain thesaurus. Meanwhile, the plurality of specific-domain words in the specific-domain thesaurus are stored in the format of a word list, so that it is not necessary to maintain complex information such as word frequency and part of speech. The words can take effect as long as they are added to the word list, which is convenient for relevant personnel who are not specialized in algorithms to operate.
[0059] Further, in order to improve the matching efficiency between the text to be processed and the specific-domain thesaurus, the data in the specific-domain thesaurus can be processed, and each specific-domain word can be processed into a data structure that is easy to quickly retrieve and then stored. When processing specific-domain words, a data structure of double-array Trie tree combined with AC automaton can be adopted. The specific processing method is selected according to actual conditions, and the embodiments of the present application do not make any limitation on this.
[0060] In practical applications, the text to be processed is matched with a domain-specific thesaurus based on the character order of the text to be processed. After determining the target word in the text to be processed, at least one character before the target word can be identified as the target text segment, and at least one character after the target word can be identified as the updated text to be processed. The process then returns to the previous steps of matching the text to be processed with the domain-specific thesaurus based on the character order of the text to determine the target word in the text to be processed. This process continues until all characters in the text to be processed have been matched, thus obtaining non-overlapping target word segments and target text segments.
[0061] For example, suppose the text to be processed is "Game A's New Year's benefits you absolutely can't miss! 520 lucky bags at 50% off". The domain-specific thesaurus includes the domain-specific terms "Game A" and "520 lucky bags". The text to be processed can be searched from beginning to end in the domain-specific thesaurus. Each time a domain-specific term that is as long as possible is found, the domain-specific term and the text portion before it are released. Then, the search starts from the next character after the domain-specific term until the text to be processed is completely searched. Then, the target words "Game A" and "520 lucky bags" in the text to be processed, and the target text segments "New Year's benefits you absolutely can't miss!" and "50% off" can be obtained.
[0062] The scheme applied in this application matches the text to be processed with a domain-specific thesaurus based on the character order of the text to be processed, thereby determining the target word segment in the text to be processed. The domain-specific thesaurus includes multiple domain-specific words. Using the target word segment as the dividing point, the text to be processed is segmented to obtain the target text segment. This ensures the accuracy of word formation for domain-specific words and further improves the accuracy of text processing.
[0063] In another possible implementation of this application, character recognition can be performed on the text to be processed to identify characters of a specified type in the text, and then these characters can be deleted to obtain the target text segment. Further, after matching the text to be processed with a domain-specific thesaurus based on the character order to determine the target word segment, character recognition can be performed on the text segments other than the target word segment to determine the target text segment. That is, the above-mentioned segmentation of the text to be processed using the target word segmentation point to obtain the target text segment can include the following steps:
[0064] Using the target word segmentation point as the segmentation point, the text to be processed is segmented to obtain candidate text segments;
[0065] Perform character recognition on candidate text segments to identify characters of a specified type within the candidate text segments;
[0066] Deleting characters of a specified type from a candidate text segment to obtain a target text segment, wherein the specified type comprises at least one of letters, numbers, and symbols.
[0067] It should be noted that the candidate text segment is a text segment remaining after removing target word segmentation from the text to be processed. Taking the target word segmentation as a segmentation point, segmenting the text to be processed to obtain the candidate text segment may specifically be deleting the target word segmentation from the text to be processed to obtain the candidate text segment. After obtaining the candidate text segment, there are multiple ways to perform character recognition on the candidate text segment and determine characters of the specified type in the candidate text segment, which are specifically selected according to actual conditions, and embodiments of the present application do not impose any limitation thereon.
[0068] In practical applications, the split continuous letters and numbers can be directly processed as a complete "word" without subsequent processing; each split symbol part is processed as a separate word.
[0069] In a possible implementation of the present application, the candidate text segment may be matched with a preset character library to determine characters of the specified type in the candidate text segment, wherein the preset character library includes a plurality of characters of types such as letters, numbers, and symbols.
[0070] In another possible implementation of the present application, the candidate text segment may be input into a character recognition model, and after processing by the character recognition model, characters of the specified type in the candidate text segment are obtained, wherein the character recognition model is trained based on a plurality of sample texts and execution type character labels carried by each sample text.
[0071] Further, after obtaining characters of the specified type in the candidate text segment, the characters of the specified type can be deleted from the candidate text segment to obtain the target text segment.
[0072] For example, assuming that the text to be processed is "Limited NPC Gift Pack!! 288 Zhang San, 388 Li Si", the specific domain vocabulary "Gift Pack" is included in the specific domain word library. According to the character order of the text to be processed, matching the text to be processed with the specific domain word library, determining that the target word segmentation in the text to be processed is "Gift Pack", taking the target word segmentation as the segmentation point, segmenting the text to be processed to obtain candidate text segments "Limited", "NPC", "了", "!! 288 Zhang San, 388 Li Si". Performing character recognition on the candidate text segments, determining that the letter in the candidate text segments is "NPC", the numbers are "288" and "388", and the symbols are "!", "!", and ",". Deleting the letters, numbers and symbols from the candidate text segments to obtain target text segments "Limited", "了", "Zhang San" and "Li Si".
[0073] The scheme of this application uses the target word segmentation as the segmentation point to segment the text to be processed, obtaining candidate text segments; character recognition is performed on the candidate text segments to determine characters of a specified type in the candidate text segments; the characters of the specified type are deleted from the candidate text segments to obtain the target text segment, wherein the specified type includes at least one of letters, numbers, and symbols. By processing characters of the specified type, a large number of unnecessary word segmentation processes are eliminated, improving the efficiency of text processing, especially in the context of the Internet where symbols are used extensively and Chinese and English are mixed.
[0074] Step 204: Based on the character order of the target text segment, segment the target text segment into words to obtain an initial text segment and a preset number of initial words. The initial text segment is the remaining text segment in the target text segment excluding the initial words.
[0075] In one or more embodiments of this application, after extracting the target text segment from the text to be processed, the target text segment can be further rooted based on the character order of the target text segment to obtain the initial text segment and a preset number of initial word segments.
[0076] Specifically, the specific value of the preset quantity is selected according to the actual situation, and this application embodiment does not impose any limitation on it. In this application embodiment, the preset quantity is preferably 3.
[0077] In practical applications, there are various ways to segment the target text segment based on the character order of the target text segment to obtain the initial text segment and a preset number of initial segments. The specific method to be selected depends on the actual situation, and this application embodiment does not impose any limitations on this.
[0078] In one possible implementation of this application, the target text segment can be segmented using a word segmentation tool based on the character order of the target text segment to obtain an initial text segment and a preset number of initial words. The word segmentation tool includes Jieba word segmentation tool, similarity word segmentation tool, etc., and the specific selection is made according to the actual situation. This application embodiment does not limit this in any way.
[0079] In another possible implementation of this application, a word feature library can be used to segment the target text segment to obtain an initial text segment and a preset number of initial words. That is, the above-mentioned segmentation of the target text segment based on the character order of the target text segment to obtain an initial text segment and a preset number of initial words may include the following steps:
[0080] Based on the character order of the target text segment and the word feature information of each word in the word feature library, the target text segment is segmented into words to obtain the initial text segment and a preset number of initial words.
[0081] Specifically, the word feature information refers to the attribute features of the word itself, which can be the word frequency information or the word weight information. The choice is made according to the actual situation, and this application embodiment does not impose any limitations on this. The word feature information of each word in the word feature library can be obtained through big data statistics, or the publicly available word frequency information of each word can be directly obtained as word feature information.
[0082] The solution of this application embodiment is to segment the target text segment based on the character order of the target text segment and the word feature information of each word in the word feature library, thereby obtaining the initial text segment and a preset number of initial words, which improves the accuracy of obtaining the initial words and the initial text segment.
[0083] In an optional embodiment of this application, before segmenting the target text segment based on the character order of the target text segment and the feature information of each word in the word feature library to obtain the initial text segment and a preset number of initial segments, the following steps may be included:
[0084] Obtain multiple sample words, where each sample word carries word feature information;
[0085] Multiple sample words are processed into linear arrays, and a word feature library is constructed based on the processed sample words.
[0086] In this application, there are multiple ways to obtain multiple sample words. One possible implementation is to manually input a large number of sample words to construct a word feature library. Another possible implementation is to read a large number of sample words from other data acquisition devices or databases to construct a word feature library.
[0087] It should be noted that after obtaining multiple sample words, a word feature library can be directly constructed based on these sample words. Furthermore, in order to speed up the word library retrieval, the sample words can be structurally processed by using a double-array Trie tree and an AC automaton to process the multiple sample words into a linear array, and then a word feature library can be constructed based on the processed sample words.
[0088] The scheme of this application embodiment obtains multiple sample words, wherein each sample word carries word feature information; the multiple sample words are processed into a linear array, and a word feature library is constructed based on the processed multiple sample words. This accelerates the word library retrieval speed and improves text processing efficiency.
[0089] In practical applications, there are various ways to segment the target text segment and obtain the initial text segment and a preset number of initial segments based on the character order of the target text segment and the word feature information of each word in the word feature library. The specific method is selected according to the actual situation, and this application embodiment does not limit it in any way.
[0090] In a possible implementation of the present application, the target text segment can be matched with the word feature library based on the character order of the target text segment, so as to determine a plurality of candidate words in the target text segment; the plurality of candidate words are sorted in descending order according to word feature information, the first preset number of candidate words are taken as initial words, and meanwhile, the preset number of initial words are deleted from the target text segment to obtain an initial text segment.
[0091] In another possible implementation of the present application, the word feature library can be used to traverse and retrieve the target text segment, so as to segment the consecutive preset number of initial words and the initial text segment with the highest word formation probability. That is, the above-mentioned segmenting the target text segment based on the character order of the target text segment and the word feature information of each word in the word feature library to obtain the initial text segment and the preset number of initial words may include the following steps:
[0092] matching the target text segment with the word feature library based on the character order of the target text segment, and determining a plurality of candidate words in the target text segment;
[0093] grouping the plurality of candidate words according to the preset number and the character order to obtain at least one candidate word group, wherein the candidate words in the candidate word group are consecutive;
[0094] calculating a word segmentation index of the at least one candidate word group according to the word feature information;
[0095] determining the preset number of initial words from the at least one candidate word group according to the word segmentation index;
[0096] deleting the preset number of initial words from the target text segment to obtain an initial text segment.
[0097] Specifically, a candidate word is a word that appears both in the word feature library and the target text segment. Assuming that the word feature library includes "you", "today", "really good", "hello", "really", "good", "look" and "nice", and the target text segment is "you look really nice today", although there are two words "you" and "good" in the target text segment, "hello" does not conform to the character order of the target text segment. Therefore, matching the target text segment with the word feature library based on the character order of the target text segment, the candidate words determined in the target text segment are "you", "today", "really good", "really", "good", "look" and "nice".
[0098] Further, when grouping the plurality of candidate words to obtain at least one candidate word group, in order to ensure the coherence of the character order, the initial words can be determined from the candidate words with earlier character order according to the preset number and the character order, that is, the preset number of candidate words with earlier character order are taken as one group.
[0099] For example, assuming the preset number is 3, the candidate word segmentation units "Ni (You)", "Jin Tian (today)", "Zhen Hao (really nice)", "Zhen (really)", "Hao (good / nice)", "Kan (look / watch)", and "Hao Kan (good-looking / nice)" are grouped according to the preset number 3 and the character order of the target text segment. It is required to ensure that the candidate word segmentation units in each group are continuous during grouping, so as to obtain candidate word segmentation group 1 ["Ni (You)", "Jin Tian (today)", "Zhen Hao (really nice)"] and candidate word segmentation group 2 ["Ni (You)", "Jin Tian (today)", "Zhen (really)"].
[0100] It should be noted that since each word in the word feature library carries word feature information, the word feature information of each word in each candidate word segmentation group can be multiplied to determine the word segmentation index of each candidate word segmentation group. After determining the word segmentation index of each candidate word segmentation group, the candidate word segmentation group with the largest word segmentation index is selected, and each candidate word segmentation unit in the candidate word segmentation group is used as the initial word segmentation of the preset number. For example, among the above two candidate word segmentation groups, the candidate word segmentation group 1 has the largest word segmentation index, so "Ni (You)", "Jin Tian (today)", and "Zhen Hao (really nice)" are used as the initial word segmentation of the preset number, and the initial word segmentation of the preset number is further deleted from the target text segment "Ni Jin Tian Zhen Hao Kan (You look really nice today)", so as to obtain the initial text segment "Kan (look)".
[0101] By applying the solution of the embodiments of the present application, based on the character order of the target text segment, the target text segment is matched with the word feature library to determine a plurality of candidate word segmentation units in the target text segment; the plurality of candidate word segmentation units are grouped according to the preset number and the character order, so as to obtain at least one candidate word segmentation group, wherein the candidate word segmentation units in the candidate word segmentation group are continuous; the word segmentation index of at least one candidate word segmentation group is calculated according to the word feature information; the initial word segmentation of the preset number is determined from at least one candidate word segmentation group according to the word segmentation index; and the initial word segmentation of the preset number is deleted from the target text segment to obtain the initial text segment, so that each word in the initial word segmentation of the preset number is more accurate, and the accuracy of text processing is further improved.
[0102] Step 206: combining the specified word segmentation in the initial word segmentation with the initial text segment to obtain an updated target text segment, and returning to execute the step of segmenting the target text segment based on the character order of the target text segment.
[0103] In one or more embodiments of the present application, after extracting a target text segment from a text to be processed; segmenting the target text segment based on the character order of the target text segment to obtain an initial text segment and the initial word segmentation of a preset number, wherein the initial text segment is the remaining text segment in the target text segment except the initial word segmentation, a specified word segmentation in the initial word segmentation is further combined with the initial text segment to obtain an updated target text segment, and the step of segmenting the target text segment based on the character order of the target text segment is executed again.
[0104] Specifically, the specified word segmentation is a word segmentation at a specified position among the initial word segmentations of a preset number, that is, the last word segmentation determined based on the character order of the target text segment. Assuming that the preset number of initial word segmentations are "你", "今天" and "真好", and the initial text segment is "看". The last word segmentation among the initial word segmentations is determined to be "真好", then according to the character order of the target text segment, the specified word segmentation "真好" and the initial text segment "看" are merged to obtain an updated target text segment "真好看".
[0105] Further, when merging the specified word segmentation with the initial text segment, the specified word segmentation may also be split, and the split specified word segmentation is merged with the initial text segment. For example, the specified word segmentation "真好" can be split into "真" and "好", and the split "好" is merged with the initial text segment "看" to obtain an updated target text segment "好看". Wherein, the splitting manner for the specified word segmentation is specifically selected according to actual situations, which is not limited in the embodiments of the present application.
[0106] It should be noted that, when determining the preset number of initial word segmentations from the target text segment, the preset number of initial word segmentations are usually the words with the highest word formation probability, so the association between words and the following text may be ignored. Taking the determination of the preset number of initial word segmentations as "你", "今天" and "真好" as an example, it is obvious that the semantics between "真好" and "看" is significantly different from the semantics of the target text segment "你今天真好看". Therefore, the specified word segmentation "真好" and the initial text segment "看" can be merged to obtain an updated target text segment "真好看", and then the step of segmenting the target text segment based on the character order of the target text segment is performed again, so that "真" and "好看" can be correctly segmented.
[0107] Step 208: when a preset word segmentation stopping condition is met, obtaining a word segmentation set corresponding to the text to be processed.
[0108] In one or more embodiments of the present application, after extracting a target text segment from the text to be processed; segmenting the target text segment based on the character order of the target text segment to obtain an initial text segment and a preset number of initial word segmentations, wherein the initial text segment is a remaining text segment in the target text segment except the initial word segmentations; merging the specified word segmentation in the initial word segmentations with the initial text segment to obtain an updated target text segment, and returning to perform the step of segmenting the target text segment based on the character order of the target text segment, further, when a preset word segmentation stopping condition is met, the word segmentation set corresponding to the text to be processed can be obtained.
[0109] Specifically, the word segmentation set refers to a set obtained by segmenting the text to be processed. The word segmentation set includes at least one of words and text segments, which is specifically selected according to actual situations, and the embodiments of the present application do not impose any limitation thereon.
[0110] For example, assuming the target text segment is "Secondly, the player's character armor does not exist", with a preset quantity of 3, the word segmentation process for the target text segment is as follows:
[0111] Step A: ["Secondly", "player", "of"] character armor does not exist.
[0112] Step B: [Secondly, the "player's" character armor does not exist.]
[0113] Step C: ["Secondly", "player", "character", "armor"] does not exist.
[0114] Step D: [Secondly, the armor of the player does not exist]
[0115] Step E: ["Secondly", "player", "character", "armor", "not", "exist"]
[0116] Therefore, the target text segment "Secondly, the player's character armor does not exist" is segmented into words, and the resulting word set is "secondly", "player", "of", "character", "armor", "not", "exist".
[0117] The scheme of this application extracts the target text segment from the text to be processed; based on the character order of the target text segment, it performs word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, wherein the initial text segment is the remaining text segment in the target text segment excluding the initial word segments; the specified word segments in the initial word segments are merged with the initial text segment to obtain an updated target text segment, and the step of performing word segmentation on the target text segment based on the character order of the target text segment is returned; if a preset word segmentation stopping condition is met, the word set corresponding to the text to be processed is obtained. By performing word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, merging the specified word segments in the initial word segments with the initial text segment, and updating the target text segment, focusing only on the local semantics of the text, efficient and accurate text processing is achieved.
[0118] In one optional embodiment of this application, after obtaining the target word segment, the specified type of characters, and the word segmentation set in the text to be processed, the target word segment, the specified type of characters, and the word segmentation set can be randomly combined to obtain the text processing result of the text to be processed. Further, to make the text processing result clearer, the target word segment, the specified type of characters, and the word segmentation set can be returned item by item in the character order of the text to be processed to obtain the text processing result. That is, after obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is reached, the following steps may also be included:
[0119] Sorting target word segments, specified type of characters and a word segmentation set according to the character order of the text to be processed to obtain a text processing result.
[0120] For example, assuming that the text to be processed is "限定NPC大礼包了!!288张三,388李四,漂亮极了", the text to be processed is matched with a domain-specific word bank according to the character order of the text to be processed, and the target word segment in the text to be processed is determined as "大礼包". With the target word segment as a splitting point, the text to be processed is split to obtain candidate text segments "限定NPC", "了" and "!!288张三,388李四". Character recognition is performed on the candidate text segments to determine that the letters in the candidate text segments are "NPC", the numbers are "288" and "388", and the symbols are "!", "!", "," and ",". The letters, numbers and symbols are removed from the candidate text segments to obtain target text segments "限定", "了", "张三", "李四" and "漂亮极了". Word segmentation is performed on the target text segments based on the character order of the target text segments to obtain initial text segments and a preset number of initial word segments, wherein the initial text segments are the remaining text segments in the target text segments excluding the initial word segments. Specified word segments in the initial word segments are merged with the initial text segments to obtain updated target text segments, and the step of performing word segmentation on the target text segments based on the character order of the target text segments is executed again. When a preset word segmentation stop condition is satisfied, a word segmentation set corresponding to the text to be processed is obtained, which includes "限定", "了", "张三", "李四", "漂亮" and "极了". The target word segments, the specified type of characters and the word segmentation set are sorted according to the character order of the text to be processed to obtain the text processing result: "限定", "NPC", "大礼包", "了", "!", "!", "288", "张三", ",", "388", "李四", ",", "漂亮", "极了".
[0121] With the solution of the embodiment of the present application, the target word segments, the specified type of characters and the word segmentation set are sorted according to the character order of the text to be processed to obtain the text processing result, which makes the text processing result clearer and improves user experience.
[0122] In practical applications, the preset word segmentation stop condition includes, but is not limited to, all characters in the target text segment have been segmented, the number of iterations reaches a preset number of iterations, and the number of initial word segments reaches a preset threshold, which is specifically selected according to actual conditions, and is not limited in any way in the embodiments of the present application.
[0123] In a first possible implementation of the present application, the preset word segmentation stop condition comprises that all characters in the target text segment have been segmented; obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stop condition is satisfied may comprise the following steps:
[0124] If all characters in the target text segment have been segmented, obtain the segmented word set corresponding to the text to be processed, where the segmented word set includes multiple words.
[0125] For example, assuming the target text segment is "You look really good today", the target text segment is segmented based on the character order of the target text segment to obtain an initial text segment and a preset number of initial segments. The specified segments in the initial segments are merged with the initial text segment to obtain an updated target text segment. The step of segmenting the target text segment based on the character order of the target text segment is then returned. If all characters in the target text segment have been segmented, the segmentation set corresponding to the text to be processed is "you", "today", "really", and "look good".
[0126] By applying the solution of this application embodiment, when all characters in the target text segment have been segmented, the word segmentation set corresponding to the text to be processed is obtained, thereby improving the efficiency and accuracy of text processing.
[0127] In the second possible implementation of this application, the preset word segmentation stopping condition includes a preset number of iterations; obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is reached may include the following steps:
[0128] Once the preset number of iterations is reached, a word segmentation set corresponding to the text to be processed is obtained, wherein the word segmentation set includes multiple words.
[0129] Specifically, the preset number of iterations is selected according to the actual situation, and this application embodiment does not impose any limitation on it.
[0130] For example, assuming the preset number of iterations is 2, taking the above example of the target text segment "Secondly, the player's character armor does not exist", the first iteration obtains the words "secondly" and "player", and the second iteration obtains the words "of" and "character". When the number of iterations reaches the preset number of iterations 2, the word segmentation set is "secondly", "player", "of", "character" and "armor does not exist" which has not yet been processed.
[0131] By applying the solution of this application embodiment, a word segmentation set corresponding to the text to be processed can be obtained after reaching a preset number of iterations, thereby improving the efficiency and accuracy of text processing.
[0132] In the third possible implementation of this application, the preset word segmentation stopping condition includes a preset threshold; obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is reached may include the following steps:
[0133] If the initial number of word segments reaches a preset threshold, obtain the preset threshold number of words;
[0134] Remove a preset threshold number of words from the text to be processed to obtain a segmented text segment;
[0135] Based on the segmented text segments and a preset threshold number of words, construct the segmented set corresponding to the text to be processed.
[0136] Specifically, the preset threshold is selected according to the actual situation, and this application embodiment does not impose any limitations on it.
[0137] For example, assuming the preset iteration count is 2, taking the above example of the target text segment "the second player's character armor does not exist", after obtaining the initial segmentation "the second" and "player", if it is determined that the number of initial segmentation words reaches the preset threshold of 2, then "the second" and "player" are deleted from "the second player's character armor does not exist", and the segmentation text segment "the character armor does not exist" is obtained, thereby obtaining the segmentation set: "the second", "player", "the character armor does not exist".
[0138] The solution of this application embodiment obtains a preset threshold number of words when the initial number of word segments reaches a preset threshold; the preset threshold number of words are then deleted from the text to be processed to obtain a word segmented text segment; based on the word segmented text segment and the preset threshold number of words, a word segmentation set corresponding to the text to be processed is constructed, thereby improving the efficiency and accuracy of text processing.
[0139] The text processing method provided in this application can be applied to different fields, such as e-commerce, gaming, etc. The specific application can be selected according to the actual situation, and this application does not limit it in any way.
[0140] Taking the gaming industry as an example, during the continuous operation of game products, game companies can obtain a large amount of player feedback, discussions, and derivative works on game content through relevant content output on social media platforms. This content is collectively referred to as game public opinion information. By organizing, processing, and analyzing this public opinion information, game product operators and developers can gain close access to players' genuine gaming emotions and needs, thereby making targeted improvements to the game product. Public opinion analysis is a crucial part of the subsequent iterative development of modern games.
[0141] Furthermore, internet corpora in the gaming field exhibit characteristics different from traditional long texts, primarily including: incomplete sentence elements such as subject, verb, and object; inconsistent sentence structure and grammar; a constant stream of new internet slang that is difficult to distinguish; and short individual sentences but a massive total volume, which severely tests word segmentation speed. This makes traditional word segmentation algorithms in the field of natural language processing ill-suited to the word segmentation requirements of internet corpora in the gaming field. Therefore, this application provides a method more suitable for handling the massive amounts of corpora in the gaming internet context.
[0142] The following is in conjunction with the appendix Figure 3 Taking the application of the text processing method provided in this application in the game field as an example, the text processing method will be further explained. Among other things, Figure 3 This application provides a flowchart illustrating a text processing method for the gaming industry, which includes the following steps:
[0143] Step 302: Match the text to be processed with the game domain lexicon according to the character order of the text to be processed to determine the target word in the text to be processed. The game domain lexicon includes multiple game domain words.
[0144] Step 304: Using the target word as the segmentation point, segment the text to be processed to obtain candidate text segments.
[0145] Step 306: Perform character recognition on the candidate text segment to determine the characters of a specified type in the candidate text segment.
[0146] Step 308: Delete characters of the specified type from the candidate text segment to obtain the target text segment, wherein the specified type includes at least one of letters, numbers, and symbols.
[0147] Step 310: Based on the character order of the target text segment and the word feature information of each word in the word feature library, segment the target text segment into words to obtain the initial text segment and a preset number of initial words.
[0148] Step 312: Merge the specified words in the initial word segmentation with the initial text segment to obtain the updated target text segment, and return to the step of performing word segmentation on the target text segment based on the character order of the target text segment and the word feature information of each word in the word feature library.
[0149] Step 314: If the preset segmentation stopping condition is met, obtain the segmentation set corresponding to the text to be processed.
[0150] Step 316: Based on the character order of the text to be processed, sort the target word segment, the characters of the specified type, and the word segmentation set to obtain the text processing result.
[0151] The text processing method provided by the embodiments of the present application is, first of all, very suitable for the word segmentation requirements of Internet corpora mainly in the forms of posts, bullet comments, comments, etc. It only focuses on the word formation probability of a preset number of locally consecutive words, which is more in line with the language habits of Internet users, and avoids complex word formation enumeration and probability calculation for the entire sentence of text. Secondly, according to the character order of the text to be processed, the text to be processed is matched with a domain-specific lexicon, which meets the demand of public opinion analysts in the game field for focusing on important words in the domain-specific lexicon. Through the retrieval and filtering of domain-specific words, the word formation accuracy of domain-specific words is guaranteed, so that the information of this part of words can be correctly presented in subsequent steps such as word frequency statistics and word cloud analysis. Meanwhile, the storage format of domain-specific words is a word list, which does not require maintaining complex information such as word frequency and part of speech, and can take effect as long as the words are added to the word list, which is convenient for non-professional relevant personnel to operate. In addition, character recognition is performed on the text to determine characters of a specified type, which eliminates a large number of unnecessary calculation processes, and the speed increase is particularly obvious in the Internet context where symbols are exaggeratedly used and Chinese and English are mixed. Finally, through the data structures of double-array Trie tree and AC automaton, the retrieval speed of the word feature bank and the domain-specific lexicon is greatly accelerated. Meanwhile, the word feature bank and the domain-specific lexicon can also be updated by means of new word discovery algorithm and manual update by business personnel, which ensures the accuracy in the text processing process.
[0152] Refer to Figure 4 , Figure 4 which shows a schematic diagram of a text processing interface provided by an embodiment of the present application. The text processing interface includes a text upload box, a "confirm" control, a "cancel" control and a text processing result display box. A user uploads a text to be processed in the text upload box, for example, "Limited NPC gift package now! ! 288 Zhang San, 388 Li Si", and clicks the "confirm" control, the server extracts the target text segment from the text to be processed; segments the target text segment based on the character order of the target text segment to obtain an initial text segment and a preset number of initial segmented words; merges a specified segmented word among the initial segmented words with the initial text segment to obtain an updated target text segment, and returns to perform the step of segmenting the target text segment based on the character order of the target text segment; when a preset word segmentation stop condition is met, a text processing result corresponding to the text to be processed is obtained: "Limited", "NPC", "gift package", "now", "!", "!", "288", "Zhang San", ",", "388", "Li Si", and the text processing result is displayed in the text processing result display box. Further, the text to be processed and the text processing result can also be displayed simultaneously in the text processing result display box.
[0153] It should be noted that the user can operate the control in any way, including clicking, double-clicking, touching, hovering the mouse, swiping, long-pressing, voice control, or shaking, depending on the actual situation. This application embodiment does not limit this in any way.
[0154] The solution implemented in this application displays the text processing results through a text processing interface, allowing users to intuitively see the results and improving the user experience.
[0155] Corresponding to the above method embodiments, this application also provides text processing apparatus embodiments. Figure 5 A schematic diagram of the structure of a text processing apparatus according to an embodiment of this application is shown. Figure 5 As shown, the device includes:
[0156] Extraction module 502 is configured to extract target text segments from the text to be processed;
[0157] The word segmentation module 504 is configured to segment the target text segment based on the character order of the target text segment to obtain an initial text segment and a preset number of initial words. The initial text segment is the text segment remaining in the target text segment excluding the initial words.
[0158] The merging module 506 is configured to merge the specified words in the initial word segmentation with the initial text segment to obtain the updated target text segment, and return the step of performing word segmentation on the target text segment based on the character order of the target text segment;
[0159] The module 508 is configured to obtain the set of words corresponding to the text to be processed when the preset word segmentation stopping condition is met.
[0160] Optionally, the extraction module 502 is further configured to match the text to be processed with a domain-specific thesaurus based on the character order of the text to be processed, and determine the target word in the text to be processed, wherein the domain-specific thesaurus includes multiple domain-specific words; and to segment the text to be processed using the target word as the segmentation point to obtain the target text segment.
[0161] Optionally, the extraction module 502 is further configured to segment the text to be processed using the target word segmentation as the segmentation point to obtain candidate text segments; to perform character recognition on the candidate text segments to determine characters of a specified type in the candidate text segments; and to delete characters of the specified type from the candidate text segments to obtain the target text segment, wherein the specified type includes at least one of letters, numbers, and symbols.
[0162] Optionally, the apparatus further includes a sorting module configured to sort the target word segment, characters of a specified type, and a set of word segments based on the character order of the text to be processed, thereby obtaining the text processing result.
[0163] Optionally, the word segmentation module 504 is further configured to segment the target text segment based on the character order of the target text segment and the word feature information of each word in the word feature library, so as to obtain an initial text segment and a preset number of initial words.
[0164] Optionally, the device further includes: a construction module configured to acquire multiple sample words, wherein the sample words carry word feature information; process the multiple sample words into a linear array, and construct a word feature library based on the processed multiple sample words.
[0165] Optionally, the word segmentation module 504 is further configured to match the target text segment with a word feature library based on the character order of the target text segment to determine multiple candidate words in the target text segment; group the multiple candidate words according to a preset number and character order to obtain at least one candidate word segmentation group, wherein the candidate words in the candidate word segmentation group are consecutive; calculate the word segmentation index of at least one candidate word segmentation group according to word feature information; determine a preset number of initial words from at least one candidate word segmentation group according to the word segmentation index; and delete the preset number of initial words from the target text segment to obtain an initial text segment.
[0166] Optionally, the preset word segmentation stopping condition includes that all characters in the target text segment have been segmented; the obtaining module 508 is further configured to obtain the word segmentation set corresponding to the text to be processed when all characters in the target text segment have been segmented, wherein the word segmentation set includes multiple words.
[0167] Optionally, the preset word segmentation stopping condition includes a preset number of iterations; the obtaining module 508 is further configured to obtain a word segmentation set corresponding to the text to be processed when the preset number of iterations is reached, wherein the word segmentation set includes multiple words.
[0168] Optionally, the preset word segmentation stopping condition includes a preset threshold; the obtaining module 508 is further configured to obtain a preset threshold number of words when the initial number of words reaches the preset threshold; delete the preset threshold number of words from the text to be processed to obtain a word segmented text segment; and construct a word segmentation set corresponding to the text to be processed based on the word segmented text segment and the preset threshold number of words.
[0169] The scheme of this application extracts the target text segment from the text to be processed; based on the character order of the target text segment, it performs word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, wherein the initial text segment is the remaining text segment in the target text segment excluding the initial word segments; the specified word segments in the initial word segments are merged with the initial text segment to obtain an updated target text segment, and the step of performing word segmentation on the target text segment based on the character order of the target text segment is returned; if a preset word segmentation stopping condition is met, the word set corresponding to the text to be processed is obtained. By performing word segmentation on the target text segment to obtain an initial text segment and a preset number of initial word segments, merging the specified word segments in the initial word segments with the initial text segment, and updating the target text segment, focusing only on the local semantics of the text, efficient and accurate text processing is achieved.
[0170] The above is an illustrative scheme of a text processing device according to this embodiment. It should be noted that the technical solution of this text processing device and the technical solution of the aforementioned text processing method belong to the same concept. Details not described in detail in the technical solution of the text processing device can be found in the description of the technical solution of the aforementioned text processing method. Furthermore, the components in the device embodiment should be understood as functional modules necessary to implement each step of the program flow or each step of the method; these functional modules are not actual functional divisions or separations. A device claim defined by such a set of functional modules should be understood as a functional module architecture that primarily implements the solution through the computer program described in the specification, and not as a physical device that primarily implements the solution through hardware.
[0171] Figure 6 This diagram illustrates a structural block diagram of a computing device according to an embodiment of this application. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0172] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0173] In one embodiment of this application, the aforementioned components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0174] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0175] The processor 620 is used to execute computer-executable instructions for the text processing method.
[0176] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described text processing method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described text processing method.
[0177] One embodiment of this application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, are used for a text processing method.
[0178] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-described text processing method belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the above-described text processing method.
[0179] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0180] An embodiment of this application also provides a chip that stores a computer program, which, when executed by the chip, implements the steps of the text processing method.
[0181] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0182] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0183] The preferred embodiments disclosed above are merely illustrative of this application. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this application. These embodiments are selected and specifically described in this application to better explain the principles and practical applications of this application, thereby enabling those skilled in the art to better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. A text processing method, characterized in that, include: Extract the target text segment from the text to be processed; Based on the character order of the target text segment, the target text segment is traversed and retrieved using a word feature library, and the target text segment is segmented into words to obtain an initial text segment and a preset number of initial words. The initial text segment is the remaining text segment in the target text segment excluding the initial words. Each word in the word feature library carries word feature information, which is the attribute feature of the word itself. The word feature library is constructed by performing structural processing on sample words, processing multiple sample words into a linear array, and constructing the word feature library based on the processed multiple sample words. The specified word segment in the initial word segmentation is merged with the initial text segment to obtain the updated target text segment, and the step of performing word segmentation on the target text segment based on the character order of the target text segment is returned. The specified word segmentation includes the last word segmentation determined based on the character order of the target text segment. When the preset segmentation stopping condition is met, the segmentation set corresponding to the text to be processed is obtained.
2. The method according to claim 1, characterized in that, The extraction of the target text segment from the text to be processed includes: Based on the character order of the text to be processed, the text to be processed is matched with a domain-specific lexicon to determine the target word segmentation in the text to be processed, wherein the domain-specific lexicon includes multiple domain-specific words; Using the target word as the segmentation point, the text to be processed is segmented to obtain the target text segment.
3. The method according to claim 2, characterized in that, The step of segmenting the text to be processed using the target word as the segmentation point to obtain the target text segment includes: Using the target word segmentation point as the segmentation point, the text to be processed is segmented to obtain candidate text segments; Character recognition is performed on the candidate text segment to determine characters of a specified type in the candidate text segment; The specified type of characters is deleted from the candidate text segment to obtain the target text segment, wherein the specified type includes at least one of letters, numbers, and symbols.
4. The method according to claim 3, characterized in that, After obtaining the word set corresponding to the text to be processed when the preset word segmentation stopping condition is met, the process further includes: Based on the character order of the text to be processed, the target word segment, the characters of the specified type, and the word segmentation set are sorted to obtain the text processing result.
5. The method according to claim 1, characterized in that, The step of segmenting the target text segment into words based on the character order of the target text segment to obtain an initial text segment and a preset number of initial words includes: Based on the character order of the target text segment and the word feature information of each word in the word feature library, the target text segment is segmented into words to obtain an initial text segment and a preset number of initial words.
6. The method according to claim 5, characterized in that, Before segmenting the target text segment based on the character order and feature information of each word in the word feature library to obtain the initial text segment and a preset number of initial words, the process further includes: Obtain multiple sample words, wherein the sample words carry word feature information; The multiple sample words are processed into a linear array, and a word feature library is constructed based on the processed sample words.
7. The method according to claim 5, characterized in that, The step of segmenting the target text segment into words based on the character order and word feature information of each word in the word feature library to obtain an initial text segment and a preset number of initial words includes: Based on the character order of the target text segment, the target text segment is matched with a word feature library to determine multiple candidate words in the target text segment; Based on the preset number and the character order, the multiple candidate word segments are grouped to obtain at least one candidate word segment group, wherein the candidate word segments in the candidate word segment group are consecutive; Based on the word feature information, calculate the word segmentation index of the at least one candidate word segmentation group; Based on the word segmentation index, determine the preset number of initial word segments from the at least one candidate word segmentation group; The initial text segment is obtained by deleting the preset number of initial words from the target text segment.
8. The method according to claim 1, characterized in that, The preset word segmentation stopping condition includes the fact that all characters in the target text segment have been segmented; the step of obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is met includes: If all characters in the target text segment have been segmented, a segmentation set corresponding to the text to be processed is obtained, wherein the segmentation set includes multiple words.
9. The method according to claim 1, characterized in that, The preset word segmentation stopping condition includes a preset number of iterations; obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is reached includes: Upon reaching a preset number of iterations, a word segmentation set corresponding to the text to be processed is obtained, wherein the word segmentation set includes multiple words.
10. The method according to claim 1, characterized in that, The preset word segmentation stopping condition includes a preset threshold; obtaining the word segmentation set corresponding to the text to be processed when the preset word segmentation stopping condition is met includes: If the number of initial word segments reaches the preset threshold, then the preset threshold number of words is obtained; The preset threshold number of words are deleted from the text to be processed to obtain a segmented text segment; Based on the segmented text segment and the preset threshold number of words, construct the segmented set corresponding to the text to be processed.
11. A text processing device, characterized in that, include: The extraction module is configured to extract target text segments from the text to be processed; The word segmentation module is configured to traverse and retrieve the target text segment based on the character order of the target text segment, and segment the target text segment into words to obtain an initial text segment and a preset number of initial words. The initial text segment is the remaining text segment in the target text segment excluding the initial words. Each word in the word feature library carries word feature information, which is the attribute feature of the word itself. The word feature library is constructed by performing structural processing on sample words, processing multiple sample words into a linear array, and constructing the word feature library based on the processed multiple sample words. The merging module is configured to merge a specified word from the initial word segmentation with the initial text segment to obtain an updated target text segment, and return to the step of performing word segmentation on the target text segment based on the character order of the target text segment, wherein the specified word segmentation includes the last word segmentation determined based on the character order of the target text segment; The acquisition module is configured to acquire the word set corresponding to the text to be processed when a preset word segmentation stopping condition is met.
12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the text processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium storing computer instructions, characterized in that, When executed by the processor, this instruction implements the steps of the text processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Word segmentation method and device, electronic equipment and storage medium
CN115409031A