Methods, apparatus, electronic devices and readable storage media for detecting mixed words
Patent Information
- Application Number
- CN202310828980.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-07-07
AI Technical Summary
[0003]本申请实施例的目的是提供一种混合词检测方法、装置和电子设备,能够解决相关技术中视频画面花屏的问题
Smart Images

Figure CN117033575B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, specifically relating to a mixed word detection method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the development and rise of internet social media, new hybrid words often emerge from the recombination of two existing words. Currently, the identification of such English hybrid words mainly relies on a continuously updated hybrid word text database. When a hybrid word is encountered, it can be queried and matched within this database. However, because matching depends on the content of this database, the accuracy of hybrid word detection and matching is low when encountering hybrid words that are not found in the database. Summary of the Invention
[0003] The purpose of this application is to provide a mixed word detection method, apparatus, and electronic device that can solve the problem of video screen distortion in related technologies.
[0004] In a first aspect, embodiments of this application provide a method for detecting mixed words, the method comprising:
[0005] Obtain the word to be detected, the prefix word list, and the suffix word list;
[0006] The similarity of the word to be detected with the prefix words in the prefix word list and the suffix words in the suffix word list is calculated respectively;
[0007] Based on the similarity calculation results, it is determined whether the word to be detected is a mixed word.
[0008] Optionally, obtaining the prefix and suffix word lists includes:
[0009] Obtain a hybrid lexicon, which includes K hybrid words, where K is an integer greater than 1;
[0010] Based on the trie data structure, the K mixed words are split into K first prefix words and K first suffix words;
[0011] The prefix word table is constructed based on the K first prefix words;
[0012] The suffix word table is constructed based on the K first suffix words.
[0013] Optionally, after calculating the similarity between the word to be detected and the prefix words in the prefix word list and the suffix words in the suffix word list, the method further includes:
[0014] Obtain M target prefix words similar to the word to be detected from the prefix word list, and obtain N target suffix words similar to the word to be detected from the suffix word list, where M and N are both integers greater than or equal to 1.
[0015] Optionally, obtaining M target prefix words similar to the word to be detected in the prefix word table, and obtaining N target suffix words similar to the word to be detected in the suffix word table, includes:
[0016] According to the similarity matching algorithm, the word to be detected is compared with the prefix words in the prefix word table to obtain the first similarity result corresponding to each prefix word in the prefix word table;
[0017] According to the similarity matching algorithm, the word to be detected is compared with the suffixes in the suffix word list to obtain the second similarity result corresponding to each suffix in the suffix word list;
[0018] Based on the first similarity result corresponding to each prefix word in the prefix word table, a first prefix word sequence is determined, the first prefix word sequence including the M target prefix words arranged in descending order of the first similarity result;
[0019] Based on the second similarity result corresponding to each suffix in the suffix word list, a second suffix word sequence is determined. The second suffix word sequence includes the N target suffix words arranged in descending order of the second similarity result.
[0020] Optionally, determining whether the word to be detected is a mixed word based on the similarity calculation result includes:
[0021] Based on the similarity calculation results, the M target prefix words and the N target suffix words are obtained;
[0022] The M target prefixes and N target suffixes are concatenated to obtain S target mixed words, where S is the product of M and N;
[0023] Obtain the average edit distance between the S target mixed words and the word to be detected;
[0024] The average edit distance is used to determine whether the word to be detected is a mixed word.
[0025] Optionally, obtaining the average edit distance between the S target mixed words and the word to be detected includes:
[0026] Obtain the number of characters in the word to be detected;
[0027] Obtain the S first edit distances between the S target mixed words and the word to be detected;
[0028] Based on each of the first edit distances and the number of characters, obtain S second edit distances;
[0029] The average of the S second edit distances is taken as the average edit distance.
[0030] Optionally, determining whether a word is a hybrid word based on the average edit distance includes:
[0031] Determine whether the average edit distance is less than or equal to a preset threshold;
[0032] If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word.
[0033] Secondly, embodiments of this application provide a mixed word detection device, including:
[0034] The acquisition module is used to acquire the word to be detected, the prefix word list, and the suffix word list;
[0035] The processing module is used to calculate the similarity between the word to be detected and the prefix words in the prefix word list and the suffix words in the suffix word list, respectively.
[0036] The judgment module is used to determine whether the word to be detected is a mixed word based on the similarity calculation result.
[0037] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the mixed word detection method as described in the first aspect.
[0038] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the mixed word detection method as described in the first aspect.
[0039] In this embodiment, based on pre-acquired prefix and suffix word lists, the word to be detected is compared with the prefixes in the prefix list and the suffixes in the suffix list for similarity calculation. Based on the similarity calculation results, multiple prefixes and suffixes with high similarity to the word to be detected are identified. Then, by concatenating the prefixes and suffixes, it is determined whether there are any words similar to the word to be detected, thereby determining whether the word to be detected is a mixed word. This eliminates the need to store a large number of mixed words in a text library, reducing storage space requirements, and effectively addresses the identification problems of the word to be detected, improving the accuracy of word recognition and matching. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating a mixed word detection method provided in an embodiment of this application;
[0042] Figure 2 yes Figure 1 A flowchart illustrating the process of obtaining prefix and suffix word lists;
[0043] Figure 3 yes Figure 1 A flowchart illustrating the process of determining whether the word to be detected is a mixed word;
[0044] Figure 4 It is a cumulative distribution map of the S first edit distances between the S target mixed words and the words to be detected;
[0045] Figure 5 It is a cumulative distribution map of the S second edit distances between the S target mixed words and the words to be detected;
[0046] Figure 6 This is a schematic diagram of the structure of a mixed word detection device provided in an embodiment of this application;
[0047] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0049] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0050] The following description, in conjunction with the accompanying drawings, details the mixed word detection method, apparatus, electronic device, and readable storage medium provided in this application through specific embodiments and application scenarios.
[0051] Please refer to Figure 1 , Figure 1 This is a flowchart of a mixed word detection method provided in an embodiment of this application.
[0052] like Figure 1 As shown, the mixed word detection method includes the following steps:
[0053] Step 101: Obtain the word to be detected, the prefix word list, and the suffix word list.
[0054] It is worth mentioning that the word to be detected in this application can be a mixed English word, such as the English word "Brunch," which is composed of the words "Breakfast" and "Lunch." The word to be detected can also be a mixed Chinese word, or a mixed Chinese and English word; this application does not impose specific restrictions on this.
[0055] Specifically, obtaining the prefix and suffix word lists mentioned above can be achieved by splitting the mixed words in an existing hybrid word list, classifying the prefix and suffix words, and assigning them to the respective prefix and suffix word lists. Alternatively, it can be based on existing word roots, dividing the word roots into prefix word roots to form a prefix word list and suffix word roots to form a suffix word list. This division into prefix and suffix word lists reduces the storage of duplicate content and improves storage efficiency.
[0056] Step 102: Calculate the similarity between the word to be detected and the prefix words in the prefix word database and the suffix words in the suffix word database.
[0057] Understandably, the similarity between the word to be detected and prefixes in the prefix word database is calculated, and the similarity between the word to be detected and prefixes in the suffix word database is also calculated. The similarity calculation identifies prefixes and suffixes in the prefix word database that have high similarity to the word to be detected. Specifically, based on the similarity results between each prefix and the word to be detected in the prefix word database, the prefixes can be ranked, and multiple prefixes with high similarity can be identified based on this ranking. Similarly, multiple suffixes with high similarity can also be identified. This application does not impose any restrictions on the selected prefixes and suffixes, their number, or their order.
[0058] Furthermore, the aforementioned similarity calculation can specifically employ a similarity algorithm, such as a string similarity comparison algorithm (e.g., Jaro–Winkler similarity), or other algorithms capable of determining the similarity between the word to be detected and another word.
[0059] Step 103: Based on the similarity calculation results, determine whether the word to be detected is a mixed word.
[0060] In this embodiment, based on the aforementioned similarity calculation results, it is determined whether the word to be detected is a mixed word. The similarity calculation results can be multiple prefixes and suffixes similar to the word to be detected, as determined by the similarity algorithm. Multiple complete mixed words can be generated by concatenating these prefixes and suffixes. Then, the edit distances of these generated mixed words and the word to be detected are determined separately. All obtained edit distances are normalized to determine the average edit distance, which can be used as an indicator to determine whether the word to be detected is a mixed word. By comparing the average edit distance with a preset threshold, it is determined whether the word to be detected is a mixed word. In this way, normalization reduces the impact of string lengths of different concatenated words on the word to be detected and effectively filters out non-mixed words, improving the success rate of mixed word detection.
[0061] Optionally, such as Figure 2 As shown, obtaining the prefix word list and suffix word list includes:
[0062] Step 201: Obtain a hybrid vocabulary, which includes K hybrid words, where K is an integer greater than 1;
[0063] Step 202: Based on the trie data structure, split the K mixed words into K first prefix words and K first suffix words;
[0064] Step 203: Construct the prefix word table based on the K first prefix words;
[0065] Step 204: Construct the suffix word database based on the K first suffix words.
[0066] For example, the prefix and suffix word lists can be obtained from an existing hybrid word list. By splitting the hybrid words in the hybrid word list, the prefixes and suffixes of the existing hybrid words are obtained, and these prefixes are combined into a prefix word list, and these suffixes are combined into a suffix word list. Specifically, the splitting of existing hybrid words can use a dictionary Trie tree data structure, or other methods can be used to split the hybrid words; this application does not impose any limitations on this. By determining the aforementioned prefix and suffix word lists, the duplication of prefixes and suffixes can be reduced. By summarizing multiple distinct prefix or suffix word lists, the duplication of prefix or suffix content can be reduced, storage space can be saved, and storage efficiency can be improved.
[0067] Optionally, after calculating the similarity between the word to be detected and the prefix words in the prefix word list and the suffix words in the suffix word list, the method further includes:
[0068] Obtain M target prefix words similar to the word to be detected from the prefix word list, and obtain N target suffix words similar to the word to be detected from the suffix word list, where M and N are both integers greater than or equal to 1.
[0069] In practical implementation, based on the similarity calculation results, prefixes similar to the word to be detected in the prefix word list are identified as target prefixes, and suffixes similar to the word to be detected in the suffix word list are identified as target suffixes. This application does not impose specific restrictions on the number of target prefixes and target suffixes; they can be the same, or they can be determined based on the characteristics of the prefixes and suffixes of the word to be detected. For example, if the prefix of the word to be detected is short and concise, but the suffix is more complex, the number of target prefixes can be set to be less than the number of target suffixes, or the number of target prefixes and target suffixes can be set to be the same. By identifying target prefixes and target suffixes that are similar to the prefix and suffix of the word to be detected, respectively, the identification and detection process of the word to be detected is simplified, storage space is saved, and detection efficiency is improved.
[0070] Optionally, obtaining M target prefix words similar to the word to be detected in the prefix word table, and obtaining N target suffix words similar to the word to be detected in the suffix word table, includes:
[0071] According to the similarity matching algorithm, the word to be detected is compared with the prefix words in the prefix word table to obtain the first similarity result corresponding to each prefix word in the prefix word table;
[0072] According to the similarity matching algorithm, the word to be detected is compared with the suffixes in the suffix word list to obtain the second similarity result corresponding to each suffix in the suffix word list;
[0073] Based on the first similarity result corresponding to each prefix word in the prefix word table, a first prefix word sequence is determined, the first prefix word sequence including the M target prefix words arranged in descending order of the first similarity result;
[0074] Based on the second similarity result corresponding to each suffix in the suffix word list, a second suffix word sequence is determined. The second suffix word sequence includes the N target suffix words arranged in descending order of the second similarity result.
[0075] Understandably, a string similarity matching algorithm (e.g., Jaro–Winkler similarity) can be used to determine target prefix words in a prefix word list that have high similarity to the word to be detected, and target suffix words in a suffix word list that have high similarity to the word to be detected. Prefix words are sorted from highest to lowest based on their first similarity score, and suffix words are sorted from lowest to highest based on their second similarity score. Target prefix words can be the M most similar words identified in the sorted sequence, and target suffix words can be the N most similar words identified in the sorted sequence. Target prefix words can also be several consecutive prefix words identified in the sorted sequence, and target suffix words can also be several consecutive suffix words identified in the sorted sequence; this application does not impose any restrictions here. In this way, by determining the target prefix words and target suffix words with the highest similarity to the word to be detected, the detection efficiency and accuracy can be improved.
[0076] Optionally, such as Figure 3 As shown, determining whether the word to be detected is a mixed word based on the similarity calculation result includes:
[0077] Step 301: Based on the similarity calculation results, obtain the M target prefix words and the N target suffix words;
[0078] Step 302: Concatenate the M target prefixes and the N target suffixes to obtain S target mixed words, where S is the product of M and N;
[0079] Step 303: Obtain the average edit distance between the S target mixed words and the word to be detected;
[0080] Step 304: Determine whether the word to be detected is a mixed word based on the average edit distance.
[0081] In this embodiment, after determining the target prefix and target suffix, each target prefix and target suffix can be concatenated, specifically using a Cartesian product to obtain M and N target mixed words. The target mixed words can be compared with the word to be detected to determine the edit distance between the word to be detected and each target mixed word. Then, the average edit distance is determined based on the edit distance, which is used to determine whether the word to be detected is a mixed word. By comparing the word to be detected with the target mixed words formed by concatenating its closest target prefix and target suffix, the edit distance is determined. Based on the edit distance, the similarity between the word to be detected and the target mixed words is determined to judge whether it is a mixed word. This effectively improves the detection accuracy, reduces the dependence on existing word samples in the mixed word library, and reduces the storage space requirements.
[0082] Optionally, obtaining the average edit distance between the S target mixed words and the word to be detected includes:
[0083] Obtain the number of characters in the word to be detected;
[0084] Obtain the S first edit distances between the S target mixed words and the word to be detected;
[0085] Based on each of the first edit distances and the number of characters, obtain S second edit distances;
[0086] The average of the S second edit distances is taken as the average edit distance.
[0087] Specifically, the average edit distance between the target hybrid word and the word to be detected can be determined through normalization. Normalization can be achieved by dividing the first edit distance between each target hybrid word and the word to be detected by the number of characters in the corresponding target hybrid word, resulting in a second edit distance. The second edit distance removes the influence of different character lengths on the edit distance, enabling more accurate identification of the word to be detected. Subsequently, the average of each second edit distance is determined, and this average is used as the average edit distance. For example, refer to... Figure 4 and Figure 5 We can introduce the cumulative distribution function (CDF) as the ordinate and the graph edit distance (GED) as the abscissa to calculate the first and second edit distances. Figure 4It is a graph of the cumulative distribution function of the S first edit distances between the S target mixed words and the word to be detected. Figure 5 This is a cumulative distribution map of the S second edit distances between the S target mixed words and the word to be detected. Figure 5 The data in the middle removes the influence of different string lengths of the target mixed words, relative to Figure 4 The data is more accurate, and the average edit distance of the identified words is more precise.
[0088] Optionally, determining whether a word is a hybrid word based on the average edit distance includes:
[0089] Determine whether the average edit distance is less than or equal to a preset threshold;
[0090] If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word.
[0091] In one embodiment of this application, determining whether a word to be detected is a hybrid word based on the average edit distance can be achieved by comparing the average edit distance with a preset threshold. If the average edit distance is less than or equal to the preset threshold, it indicates a high probability that the word to be detected is a hybrid word, and the word to be detected can be identified as a hybrid word. The preset threshold can be obtained by training a hybrid word detection model using hybrid words as training samples. The cumulative distribution relationship formed by the combination of all trained hybrid words and their corresponding source words is determined through the training results of the hybrid word detection model. For example... Figure 5 The preset threshold can be set to 0.6. When the preset threshold is 0.6, the accuracy of word recognition can reach more than 93%, and it can effectively filter out non-mixed words in the target mixed words.
[0092] In another specific embodiment of this application, the mixed word detection method can be applied to text detection to detect whether the text content contains mixed words. Specifically, the text content to be identified can be split into individual words. Then, the similarity between each word in the text and the prefixes in the prefix word list is calculated, and the similarity between each word and the suffixes in the suffix word list is also calculated. Based on the similarity calculation results, the prefixes and suffixes corresponding to each word are determined. These prefixes and suffixes are then combined to obtain multiple mixed words. The edit distance algorithm is used to calculate the edit distance between each word and its corresponding multiple mixed words in the detected text. The edit distance is then combined to determine the possible mixed words in the detected text. This effectively improves the efficiency and recall rate of finding mixed words in text, effectively saves storage space, adapts to updates of mixed words, and improves the success rate of mixed word detection and recognition.
[0093] Please refer to Figure 6This application provides a mixed word detection device 400, including:
[0094] The acquisition module 401 is used to acquire the word to be detected, the prefix word list, and the suffix word list;
[0095] Processing module 402 is used to calculate the similarity between the word to be detected and the prefix words in the prefix word list and the suffix words in the suffix word list, respectively.
[0096] The judgment module 403 is used to determine whether the word to be detected is a mixed word based on the similarity calculation result.
[0097] Optionally, the acquisition module 401 includes:
[0098] A hybrid lexicon module is used to obtain a hybrid lexicon, which includes K hybrid words, where K is an integer greater than 1;
[0099] The splitting module is used to split the K mixed words into K first prefix words and K first suffix words based on a trie data structure;
[0100] The first construction module is used to construct the prefix word table library based on the K first prefix words;
[0101] The second construction module is used to construct the suffix word table library based on the K first suffix words.
[0102] Optionally, the processing module 402 is further configured to:
[0103] Obtain M target prefix words similar to the word to be detected from the prefix word list, and obtain N target suffix words similar to the word to be detected from the suffix word list, where M and N are both integers greater than or equal to 1.
[0104] Optionally, the processing module 402 is further configured to:
[0105] According to the similarity matching algorithm, the word to be detected is compared with the prefix words in the prefix word table to obtain the first similarity result corresponding to each prefix word in the prefix word table;
[0106] According to the similarity matching algorithm, the word to be detected is compared with the suffixes in the suffix word list to obtain the second similarity result corresponding to each suffix in the suffix word list;
[0107] Based on the first similarity result corresponding to each prefix word in the prefix word table, a first prefix word sequence is determined, the first prefix word sequence including the M target prefix words arranged in descending order of the first similarity result;
[0108] Based on the second similarity result corresponding to each suffix in the suffix word list, a second suffix word sequence is determined. The second suffix word sequence includes the N target suffix words arranged in descending order of the second similarity result.
[0109] Optionally, the determination module 403 includes:
[0110] The similarity result module is used to obtain the M target prefix words and the N target suffix words based on the similarity calculation results;
[0111] The mixed word concatenation module is used to concatenate the M target prefix words and the N target suffix words to obtain S target mixed words, where S is the product of M and N;
[0112] The average edit distance module is used to obtain the average edit distance between the S target mixed words and the word to be detected;
[0113] The mixed word determination module is used to determine whether the word to be detected is a mixed word based on the average edit distance.
[0114] Optionally, the average edit distance module is used for:
[0115] Obtain the number of characters in the word to be detected;
[0116] Obtain the S first edit distances between the S target mixed words and the word to be detected;
[0117] Based on each of the first edit distances and the number of characters, obtain S second edit distances;
[0118] The average of the S second edit distances is taken as the average edit distance.
[0119] Optionally, the mixed word determination module is used for:
[0120] Determine whether the average edit distance is less than or equal to a preset threshold;
[0121] If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word.
[0122] The mixed word detection device 400 provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments described above can achieve the same beneficial effects, and will not be repeated here to avoid repetition.
[0123] Please see Figure 7 , Figure 7 This is a structural diagram of an electronic device provided in an embodiment of this application, such as... Figure 7 As shown, the electronic device includes: a processor 501, a memory 502, and a program or instructions stored in the memory 502 and executable on the processor 501. The processor 501 is used to read the program or instructions in the memory 502. The electronic device also includes a bus interface and a transceiver 503.
[0124] Among them, Figure 7 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 501) and memory (memory 502). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 503 can be multiple elements, including transmitters and transceivers, providing a unit for communicating with various other devices over a transmission medium. Processor 501 is responsible for managing the bus architecture and general processing, and memory 502 can store data used by processor 501 during operation.
[0125] The processor 501 is used to read programs or instructions from the memory 502 and execute the following steps:
[0126] Obtain the word to be detected, the prefix word list, and the suffix word list;
[0127] The similarity of the word to be detected with the prefix words in the prefix word list and the suffix words in the suffix word list is calculated respectively;
[0128] Based on the similarity calculation results, it is determined whether the word to be detected is a mixed word.
[0129] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0130] Obtain a hybrid lexicon, which includes K hybrid words, where K is an integer greater than 1;
[0131] Based on the trie data structure, the K mixed words are split into K first prefix words and K first suffix words;
[0132] The prefix word table is constructed based on the K first prefix words;
[0133] The suffix word table is constructed based on the K first suffix words.
[0134] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0135] Obtain M target prefix words similar to the word to be detected from the prefix word list, and obtain N target suffix words similar to the word to be detected from the suffix word list, where M and N are both integers greater than or equal to 1.
[0136] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0137] According to the similarity matching algorithm, the word to be detected is compared with the prefix words in the prefix word table to obtain the first similarity result corresponding to each prefix word in the prefix word table;
[0138] According to the similarity matching algorithm, the word to be detected is compared with the suffixes in the suffix word list to obtain the second similarity result corresponding to each suffix in the suffix word list;
[0139] Based on the first similarity result corresponding to each prefix word in the prefix word table, a first prefix word sequence is determined, the first prefix word sequence including the M target prefix words arranged in descending order of the first similarity result;
[0140] Based on the second similarity result corresponding to each suffix in the suffix word list, a second suffix word sequence is determined. The second suffix word sequence includes the N target suffix words arranged in descending order of the second similarity result.
[0141] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0142] Based on the similarity calculation results, the M target prefix words and the N target suffix words are obtained;
[0143] The M target prefixes and N target suffixes are concatenated to obtain S target mixed words, where S is the product of M and N;
[0144] Obtain the average edit distance between the S target mixed words and the word to be detected;
[0145] The average edit distance is used to determine whether the word to be detected is a mixed word.
[0146] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0147] Obtain the number of characters in the word to be detected;
[0148] Obtain the S first edit distances between the S target mixed words and the word to be detected;
[0149] Based on each of the first edit distances and the number of characters, obtain S second edit distances;
[0150] The average of the S second edit distances is taken as the average edit distance.
[0151] Optionally, processor 501 is configured to read programs or instructions from memory 502 and perform the following steps:
[0152] Determine whether the average edit distance is less than or equal to a preset threshold;
[0153] If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word.
[0154] The electronic device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments described above can achieve the same beneficial effects, and will not be repeated here to avoid repetition.
[0155] This application embodiment also provides a readable storage medium storing a program or instructions that, when executed by a processor, implement the above-described functionality. Figure 1 The various processes of the hybrid word detection method embodiment described herein can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0156] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0157] This application embodiment also provides a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the above. Figure 1 The various processes of the hybrid word detection method embodiment described herein can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0158] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0159] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0161] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for detecting mixed words, characterized in that, The method includes: Obtain the word to be detected, the prefix word list, and the suffix word list; wherein, the word to be detected is a mixed English word to be detected; The word to be detected is compared with the prefix words in the prefix word database and the suffix words in the suffix word database using a similarity algorithm. Based on the similarity calculation results, it is determined whether the word to be detected is a mixed word; The acquisition of the prefix and suffix word lists includes: Obtain a hybrid lexicon, which includes K hybrid words, where K is an integer greater than 1; Based on the trie data structure, the K mixed words are split into K first prefix words and K first suffix words; The prefix word table is constructed based on the K first prefix words; The suffix word library is constructed based on the K first suffix words; After calculating the similarity between the word to be detected and the prefixes in the prefix word list and the suffixes in the suffix word list using a similarity algorithm, the method further includes: Obtain M target prefix words similar to the word to be detected from the prefix word list, and obtain N target suffix words similar to the word to be detected from the suffix word list, where M and N are both integers greater than or equal to 1; The step of determining whether the word to be detected is a mixed word based on the similarity calculation result includes: Based on the similarity calculation results, the M target prefix words and the N target suffix words are obtained; The M target prefixes and N target suffixes are concatenated to obtain S target mixed words, where S is the product of M and N; Obtain the average edit distance between the S target mixed words and the word to be detected; Based on the average edit distance, it is determined whether the word to be detected is a mixed word; The step of determining whether the word to be detected is a mixed word based on the average edit distance includes: Determine whether the average edit distance is less than or equal to a preset threshold; If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word; The preset threshold is obtained by training a mixed word detection model using mixed words as training samples.
2. The method according to claim 1, characterized in that, The step of obtaining M target prefix words similar to the word to be detected in the prefix word table and N target suffix words similar to the word to be detected in the suffix word table includes: According to the similarity matching algorithm, the word to be detected is compared with the prefix words in the prefix word table to obtain the first similarity result corresponding to each prefix word in the prefix word table; According to the similarity matching algorithm, the word to be detected is compared with the suffixes in the suffix word list to obtain the second similarity result corresponding to each suffix in the suffix word list; Based on the first similarity result corresponding to each prefix word in the prefix word table, a first prefix word sequence is determined, the first prefix word sequence including the M target prefix words arranged in descending order of the first similarity result; Based on the second similarity result corresponding to each suffix in the suffix word list, a second suffix word sequence is determined. The second suffix word sequence includes the N target suffix words arranged in descending order of the second similarity result.
3. The method according to claim 1, characterized in that, The step of obtaining the average edit distance between the S target mixed words and the word to be detected includes: Obtain the number of characters in the word to be detected; Obtain the S first edit distances between the S target mixed words and the word to be detected; Based on each of the first edit distances and the number of characters, obtain S second edit distances; The average of the S second edit distances is taken as the average edit distance.
4. A mixed word detection device, characterized in that, include: The acquisition module is used to acquire the word to be detected, the prefix word list, and the suffix word list; wherein, the word to be detected is a mixed English word to be detected; The processing module is used to calculate the similarity between the word to be detected and the prefix words in the prefix word list and the suffix words in the suffix word list using a similarity algorithm. The judgment module is used to determine whether the word to be detected is a mixed word based on the result of the similarity calculation; The acquisition module includes: A hybrid lexicon module is used to obtain a hybrid lexicon, which includes K hybrid words, where K is an integer greater than 1; The splitting module is used to split the K mixed words into K first prefix words and K first suffix words based on a trie data structure; The first construction module is used to construct the prefix word table library based on the K first prefix words; The second construction module is used to construct the suffix word library based on the K first suffix words; The processing module is further configured to: obtain M target prefix words similar to the word to be detected in the prefix word list, and obtain N target suffix words similar to the word to be detected in the suffix word list, wherein M and N are both integers greater than or equal to 1; The judgment module includes: The similarity result module obtains the M target prefix words and the N target suffix words based on the similarity calculation results; The mixed word concatenation module is used to concatenate the M target prefix words and the N target suffix words to obtain S target mixed words, where S is the product of M and N; The average edit distance module is used to obtain the average edit distance between the S target mixed words and the word to be detected; The mixed word determination module is used to determine whether the word to be detected is a mixed word based on the average edit distance; The mixed word detection module is used for: Determine whether the average edit distance is less than or equal to a preset threshold; If the average edit distance is less than or equal to the preset threshold, the word to be detected is determined to be a mixed word; The preset threshold is obtained by training a mixed word detection model using mixed words as training samples.
5. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the mixed word detection method as described in any one of claims 1-3.
6. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the mixed word detection method as described in any one of claims 1-3.
Citation Information
Patent Citations
Word table storage management method and device, electronic equipment and storage medium
CN109739948A
Sensitive word detection method and device, electronic equipment and storage medium
CN112364637A
Word recognition system
JP1985000583A