Vocabulary detection method, apparatus, and related device
Patent Information
- Application Number
- CN202410652424.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-24
AI Technical Summary
[0004]本公开的目的在于提供一种词汇检测方法、装置及相关设备,用于解决相关技术对敏感词的识别准确率较低的技术问题
[0010]在本申请中,利用大语言模型确定待检测词汇与敏感词库中某一敏感词的字词较为相似的情况下,使用大语言模型,结合待检测词汇的上下文信息,对待检测词汇的语义作进一步分析,并在所分析语义与敏感词库中某一敏感词的语义也相似的情况下,将待检测词汇最终确定为待过滤词汇,以通过词汇相似判别和语义相似判别来综合确定待检测词汇是否为敏感词,进而提升对敏感词的识别准确率。
Smart Images

Figure CN118484713B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of artificial intelligence, specifically to a word detection method, apparatus, and related equipment. Background Technology
[0002] Sensitive words generally refer to words that have violent tendencies, unhealthy connotations, or uncivilized meanings.
[0003] In related technologies, sensitive words are set to ensure the legality and compliance of displayed content. However, these technologies are not accurate enough in identifying sensitive words, resulting in poor filtering effects. Summary of the Invention
[0004] The purpose of this disclosure is to provide a vocabulary detection method, apparatus, and related equipment to solve the technical problem of low accuracy in identifying sensitive words in related technologies.
[0005] In a first aspect, embodiments of this application provide a vocabulary detection method, the method comprising: Based on a pre-set large language model, the first sensitive word with the highest word similarity to the word to be detected is determined from the pre-set sensitive word library; When the word similarity between the first sensitive word and the word to be detected is greater than or equal to the first threshold, based on the text block to which the word to be detected belongs, the large language model is used to perform semantic analysis on the word to be detected to obtain the semantic data of the word to be detected. The text block includes the word to be detected and the context information of the word to be detected. Based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data is determined in the sensitive word library; If the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0006] Secondly, embodiments of this application also provide a vocabulary detection device, the device comprising: The first determination module is used to determine the first sensitive word with the highest word similarity to the word to be detected in the preset sensitive word library based on the pre-set large language model. The semantic analysis module is used to perform semantic analysis on the word to be detected based on the text block to which the word to be detected belongs, using the large language model, when the word similarity between the first sensitive word and the word to be detected is greater than or equal to a first threshold, to obtain the semantic data of the word to be detected. The text block includes the word to be detected and the context information of the word to be detected. The second determining module is used to determine, based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library; The third determining module is used to determine the word to be detected as a word to be filtered when the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to a second threshold.
[0007] Thirdly, this application provides an electronic device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.
[0008] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] Fifthly, this application provides a computer program product including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] In this application, when a large language model is used to determine that the word to be detected is similar to a certain sensitive word in the sensitive word library, the large language model is used in conjunction with the contextual information of the word to be detected to further analyze the semantics of the word to be detected. If the analyzed semantics are also similar to the semantics of a certain sensitive word in the sensitive word library, the word to be detected is finally determined as a word to be filtered. In order to comprehensively determine whether the word to be detected is a sensitive word through word similarity discrimination and semantic similarity discrimination, the accuracy of sensitive word identification is improved. Attached Figure Description
[0011] Figure 1 This is a schematic flowchart of a vocabulary detection method provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of a vocabulary detection device provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0012] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0013] This disclosure provides a vocabulary detection method, such as... Figure 1 As shown, the vocabulary detection method includes: Step 101: Based on the pre-set large language model, determine the first sensitive word with the highest word similarity to the word to be detected in the pre-set sensitive word library.
[0014] The Large Language Model (LLM) referred to in this application is an artificial intelligence model designed to understand and generate human language. Trained on large amounts of text data, it can perform a wide range of tasks, such as text summarization, translation, sentiment analysis, and more.
[0015] For example, the sensitive word library may include words with violent tendencies, unhealthy connotations, or uncivilized language, as well as words that contain false claims of efficacy, are suspected of superstition, or violate authority. It may also include sensitive words that the user defines according to their actual needs and preferences.
[0016] In one example, the above-mentioned lexical similarity can be understood as the proportion of the number of characters of the target word included in the reference word to the total number of characters of the target word. For example, if the target word is set as abc, the first reference word is abcd, the second reference word is abd, and the third reference word is ab, then the lexical similarity between the target word and the first reference word is 100%, the lexical similarity between the target word and the second reference word is 66.7%, and the lexical similarity between the target word and the third reference word is 66.7%. The word to be detected is the target word, and the sensitive words in the sensitive word library are the reference words.
[0017] In another example, the above-mentioned lexical similarity can also be understood as the ratio of the intersection and union of characters between two words to be compared. For example, if we still set the target word as abc, the first reference word as abcd, the second reference word as abd, and the third reference word as ab, then the lexical similarity between the target word and the first reference word is 75%, the lexical similarity between the target word and the second reference word is 50%, and the lexical similarity between the target word and the third reference word is 66.7%.
[0018] In the application, a trie can be constructed based on multiple sensitive words included in the sensitive word library. The trie is a multi-branch tree structure, where each node represents a character, and each path in the trie represents a sensitive word. The word ending tag is used to indicate the complete sensitive word corresponding to the path. Since the multi-branch tree structure supports prefix matching skipping, when the large language model matches the word to be detected in the trie, if a non-matching character is encountered, it can jump to the next possible matching position based on the prefix matching information to reduce unnecessary comparison operations and improve character matching efficiency. Subsequently, characters with a matching degree exceeding a preset degree are used as candidate sensitive words, and the first sensitive word is determined from the candidate sensitive words to reduce the number of sensitive words whose word similarity needs to be calculated and improve the determination efficiency of the first sensitive word. Here, the character matching degree can be understood as the proportion of the number of characters in the word to be detected included in the sensitive word to the total number of characters in the word to be detected.
[0019] For example, the word to be detected can be any word in the user-input text, any word obtained after performing text recognition on the user-input video or image, or any word obtained after performing speech-to-text conversion on the user-input audio.
[0020] Step 102: When the lexical similarity between the first sensitive word and the word to be detected is greater than or equal to the first threshold, based on the text block to which the word to be detected belongs, the large language model is used to perform semantic analysis on the word to be detected to obtain the semantic data of the word to be detected.
[0021] The text block includes the word to be detected and its contextual information.
[0022] Step 103: Based on the large language model, determine the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library.
[0023] In addition to storing multiple sensitive words, the aforementioned sensitive word database also stores the sensitive semantics corresponding to each of the multiple sensitive words. By calculating the semantic similarity between the semantic data and the multiple sensitive semantics, the second sensitive word can be determined.
[0024] Step 104: If the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0025] In this application, when a large language model is used to determine that the word to be detected is similar to a certain sensitive word in the sensitive word library, the large language model is used in conjunction with the contextual information of the word to be detected to further analyze the semantics of the word to be detected. If the analyzed semantics are also similar to the semantics of a certain sensitive word in the sensitive word library, the word to be detected is finally determined as a word to be filtered. In order to comprehensively determine whether the word to be detected is a sensitive word through word similarity discrimination and semantic similarity discrimination, the accuracy of sensitive word identification is improved.
[0026] For example, after determining that the word to be detected is a word to be filtered, the word to be detected can be blocked or deleted, or the text block to which the word to be detected belongs can be blocked or deleted. This application does not limit the specific filtering mechanism for the determined word to be filtered.
[0027] In some implementations, users are allowed to provide feedback on misjudgments of sensitive word detection results.
[0028] For example, after identifying the words to be detected in the user-uploaded text as sensitive words, it can support the user's feedback on the false positives of the sensitive word identification results (referring to non-sensitive words being identified as sensitive words). The false positive feedback is verified by human. If the false positive is true, the false positive feedback and the corresponding words to be detected are marked as non-sensitive words, and the sensitive word filtering mechanism of the large language model is optimized based on the identified non-sensitive words. Furthermore, after identifying the words to be detected in the text uploaded by the first user as sensitive words, it can support feedback from the second user on the misjudgment results of the sensitive word identification results (referring to the sensitive word being identified as a non-sensitive word). The feedback on the misjudgment is verified by a human. If the misjudgment is true, the feedback on the misjudgment and the corresponding words to be detected are marked as sensitive words, and the sensitive word filtering mechanism of the large language model is optimized based on the identified sensitive words.
[0029] In one embodiment, the sensitive word database includes N sensitive word subsets, and the N sensitive word subsets correspond one-to-one with N languages, where N is an integer greater than 1; The step of determining the first sensitive word with the highest lexical similarity to the word to be detected from a pre-set sensitive word library based on a pre-set large language model includes: Based on a pre-set large language model, the first sensitive word with the highest word similarity to the word to be detected is determined in the target sensitive word subset. The target sensitive word subset is the subset of sensitive words in the same language as the word to be detected in the N sensitive word subsets. The step of determining the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library based on the large language model includes: Based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data is determined in the target sensitive word subset.
[0030] In this embodiment, a subset of sensitive words corresponding to multiple languages is constructed to support sensitive word filtering for multilingual texts, thus avoiding the problem that large language models cannot identify sensitive words in some languages in a multilingual environment.
[0031] In applications, due to grammatical, lexical, and cultural differences between different languages, large language models can set different word similarity and semantic similarity calculation mechanisms for different sensitive word subsets. For example, a large language model can set a dedicated semantic feature extraction scheme for each sensitive word subset, and the semantic similarity can be understood as the feature similarity between the semantic features extracted from different words.
[0032] In one embodiment, after determining that the word to be detected is a word to be filtered, the method further includes: The words to be filtered are translated to obtain N-1 translated words. The N-1 translated words correspond one-to-one with the N-1 languages other than the language corresponding to the words to be filtered. The words to be filtered are added to the target sensitive word subset, and the N-1 translated words are added to the N-1 sensitive word subsets corresponding to the N-1 languages respectively.
[0033] In applications, when N languages may include less commonly spoken languages, collecting sensitive words for these languages is more difficult. Therefore, the number of sensitive words in the subset corresponding to these less commonly spoken languages is usually small. In this case, transfer learning or cross-language training methods can be used to configure an initial language processing module for the less commonly spoken languages. Subsequently, the sensitive word content of the subset corresponding to these less commonly spoken languages is enriched by translating the words to be filtered determined in the application stage into the corresponding words in the less commonly spoken languages. This optimizes the initial language processing module configured for these less commonly spoken languages and improves the accuracy of filtering sensitive words in these languages.
[0034] In one embodiment, the method further includes: If the lexical similarity between the first sensitive word and the word to be detected is less than the first threshold, according to the large language model, in the sensitive word library, the third sensitive word with the highest attribute similarity between its character attributes and the word to be detected is determined, wherein the character attributes are used to represent the glyphs and / or pronunciations of the characters included in the word; If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the word to be detected is determined to be a word to be filtered.
[0035] In this embodiment, when the characters of the word to be detected are significantly different from those of each sensitive word in the sensitive word library, the differences between the shape and / or sound of the word to be detected and the shape and / or sound of each sensitive word in the sensitive word library are compared. When the differences are small, the word to be detected is identified as the word to be filtered, so as to prevent illegal users from bypassing the sensitive word filtering mechanism by means of homophones or pictographs. This can further improve the accuracy of sensitive word filtering.
[0036] For example, filtering sensitive words based on similar character shapes could be as follows: if "illiterate" is set as a sensitive word, "zhangyu" which has a similar character shape to "illiterate" would also be identified as a sensitive word.
[0037] The case of filtering sensitive words based on similar pronunciation is as follows: if "illiterate" is set as a sensitive word, "mosquito" which has a similar pronunciation to "illiterate" will also be identified as a sensitive word.
[0038] In one embodiment, determining the word to be detected as a word to be filtered when the attribute similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to a third threshold includes: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the semantic analysis of the word to be detected is performed using the large language model based on the text block to which the word to be detected belongs, and the semantic data of the word to be detected is obtained. If the semantic similarity between the semantic data and the lexical semantics of the third sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0039] In this embodiment, when a sensitive word with similar pronunciation and / or shape is found to be present in the word to be detected, the semantics of the word to be detected are further analyzed to perform semantic similarity detection in the sensitive word database. Only when the detection result indicates that there is a sensitive word with similar semantics will the word to be detected be identified as a word to be filtered. Otherwise, if the semantic similarity between the semantic data and the semantics of the third sensitive word is less than the second threshold, the word to be detected will not be identified as a word to be filtered, so as to avoid the situation where non-sensitive words are mistakenly detected as sensitive words.
[0040] In one embodiment, determining the word to be detected as a word to be filtered when the attribute similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to a third threshold includes: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the large language model is used to perform semantic analysis on the background text of the word to be detected to obtain the first text semantics. The word to be detected in the text block is replaced with the third sensitive word to obtain the replaced text block; The semantics of the replaced text block are obtained by performing semantic analysis on the large language model to obtain the second text semantics. If the semantic similarity between the first text semantic and the second text semantic is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0041] In this embodiment, when a third sensitive word with a similar pronunciation and / or shape is found to be a word to be detected, the semantics of the third sensitive word in the context of the word to be detected are analyzed to determine whether the word to be detected is identified as a word to be filtered. The above measures can avoid full semantic similarity comparison of multiple sensitive words within the sensitive word, thus improving the detection efficiency of the word to be detected while reducing the false detection probability of the word to be detected.
[0042] See Figure 2 , Figure 2 This is a vocabulary detection device provided in an embodiment of the present disclosure, such as... Figure 2 As shown, the vocabulary detection device 200 includes: The first determining module 201 is used to determine the first sensitive word with the highest word similarity to the word to be detected in a preset sensitive word library based on a pre-set large language model. The semantic analysis module 202 is used to perform semantic analysis on the word to be detected based on the text block to which the word to be detected belongs, using the large language model, when the word similarity between the first sensitive word and the word to be detected is greater than or equal to a first threshold, to obtain the semantic data of the word to be detected. The text block includes the word to be detected and the context information of the word to be detected. The second determining module 203 is used to determine, based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library; The third determining module 204 is used to determine the word to be detected as a word to be filtered when the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to a second threshold.
[0043] In one embodiment, the sensitive word database includes N sensitive word subsets, and the N sensitive word subsets correspond one-to-one with N languages, where N is an integer greater than 1; The first determining module 201 is specifically used for: Based on a pre-set large language model, the first sensitive word with the highest word similarity to the word to be detected is determined in the target sensitive word subset. The target sensitive word subset is the subset of sensitive words in the same language as the word to be detected in the N sensitive word subsets. The second determining module 203 is specifically used for: Based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data is determined in the target sensitive word subset.
[0044] In one embodiment, the device 200 further includes: The translation module is used to translate the words to be filtered to obtain N-1 translated words. The N-1 translated words correspond one-to-one with the N-1 languages other than the language corresponding to the words to be filtered. The vocabulary supplementation module is used to add the words to be filtered to the target sensitive word subset, and to add the N-1 translated words to the N-1 sensitive word subsets corresponding to the N-1 languages respectively.
[0045] In one embodiment, the device 200 further includes: The fourth determining module is used to determine, in the sensitive word library, the third sensitive word with the highest attribute similarity to the character attributes of the word to be detected, according to the large language model, when the lexical similarity between the first sensitive word and the word to be detected is less than the first threshold. The character attributes are used to represent the glyphs and / or pronunciations of the characters included in the word. The fifth determining module is used to determine the word to be detected as a word to be filtered when the attribute similarity between the character attribute of the third sensitive word and the character attribute of the word to be detected is greater than or equal to the third threshold.
[0046] In one embodiment, the fifth determining module is specifically used for: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the semantic analysis of the word to be detected is performed using the large language model based on the text block to which the word to be detected belongs, and the semantic data of the word to be detected is obtained. If the semantic similarity between the semantic data and the lexical semantics of the third sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0047] In one embodiment, the fifth determining module is specifically used for: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the large language model is used to perform semantic analysis on the background text of the word to be detected to obtain the first text semantics. The word to be detected in the text block is replaced with the third sensitive word to obtain the replaced text block; The semantics of the replaced text block are obtained by performing semantic analysis on the large language model to obtain the second text semantics. If the semantic similarity between the first text semantic and the second text semantic is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
[0048] The vocabulary detection device 200 provided in this embodiment can implement the various processes in the above-described vocabulary detection method embodiments, and will not be described again here to avoid repetition.
[0049] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0050] Figure 3 A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0051] like Figure 3 As shown, device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 302 or a computer program loaded from storage unit 308 into random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. Input / output (I / O) interface 305 is also connected to bus 304.
[0052] Multiple components in device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of monitors, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0053] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the vocabulary detection method. For example, in some embodiments, the vocabulary detection method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the vocabulary detection method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform a vocabulary detection method by any other suitable means (e.g., by means of firmware).
[0054] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0055] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0056] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0057] As used herein, the term "machine-readable medium" refers to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0058] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0059] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0060] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0061] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0062] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0063] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A vocabulary detection method, characterized in that, The method includes: Based on a pre-set large language model, the first sensitive word with the highest word similarity to the word to be detected is determined from the pre-set sensitive word library; When the word similarity between the first sensitive word and the word to be detected is greater than or equal to the first threshold, based on the text block to which the word to be detected belongs, the large language model is used to perform semantic analysis on the word to be detected to obtain the semantic data of the word to be detected. The text block includes the word to be detected and the context information of the word to be detected. Based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data is determined in the sensitive word library; If the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered. The method further includes: If the lexical similarity between the first sensitive word and the word to be detected is less than the first threshold, according to the large language model, in the sensitive word library, the third sensitive word with the highest attribute similarity between its character attributes and the word to be detected is determined, wherein the character attributes are used to represent the glyphs and / or pronunciations of the characters included in the word; If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the word to be detected is determined to be a word to be filtered. Specifically, when the attribute similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to a third threshold, the word to be detected is determined as a word to be filtered, including: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the large language model is used to perform semantic analysis on the background text of the word to be detected to obtain the first text semantics. The word to be detected in the text block is replaced with the third sensitive word to obtain the replaced text block; The semantics of the replaced text block are obtained by performing semantic analysis on the large language model to obtain the second text semantics. If the semantic similarity between the first text semantic and the second text semantic is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
2. The method according to claim 1, characterized in that, The sensitive word database includes N sensitive word subsets, and each of the N sensitive word subsets corresponds one-to-one with N languages, where N is an integer greater than 1; The step of determining the first sensitive word with the highest lexical similarity to the word to be detected from a pre-set sensitive word library based on a pre-set large language model includes: Based on a pre-set large language model, the first sensitive word with the highest word similarity to the word to be detected is determined in the target sensitive word subset. The target sensitive word subset is the subset of sensitive words in the same language as the word to be detected in the N sensitive word subsets. The step of determining the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library based on the large language model includes: Based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data is determined in the target sensitive word subset.
3. The method according to claim 2, characterized in that, After determining that the word to be detected is the word to be filtered, the method further includes: The words to be filtered are translated to obtain N-1 translated words. The N-1 translated words correspond one-to-one with the N-1 languages other than the language corresponding to the words to be filtered. The words to be filtered are added to the target sensitive word subset, and the N-1 translated words are added to the N-1 sensitive word subsets corresponding to the N-1 languages respectively.
4. The method according to claim 1, characterized in that, If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to a third threshold, the word to be detected is determined to be a word to be filtered, including: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the semantic analysis of the word to be detected is performed using the large language model based on the text block to which the word to be detected belongs, and the semantic data of the word to be detected is obtained. If the semantic similarity between the semantic data and the lexical semantics of the third sensitive word is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
5. A vocabulary detection device, characterized in that, The device includes: The first determination module is used to determine the first sensitive word with the highest word similarity to the word to be detected in the preset sensitive word library based on the pre-set large language model. The semantic analysis module is used to perform semantic analysis on the word to be detected based on the text block to which the word to be detected belongs, using the large language model, when the word similarity between the first sensitive word and the word to be detected is greater than or equal to a first threshold, to obtain the semantic data of the word to be detected. The text block includes the word to be detected and the context information of the word to be detected. The second determining module is used to determine, based on the large language model, the second sensitive word with the highest semantic similarity to the semantic data in the sensitive word library; The third determining module is used to determine the word to be detected as the word to be filtered when the semantic similarity between the semantic data and the lexical semantics of the second sensitive word is greater than or equal to a second threshold. The device further includes: The fourth determining module is used to determine, in the sensitive word library, the third sensitive word with the highest attribute similarity to the character attributes of the word to be detected, according to the large language model, when the lexical similarity between the first sensitive word and the word to be detected is less than the first threshold. The character attributes are used to represent the glyphs and / or pronunciations of the characters included in the word. The fifth determining module is used to determine the word to be detected as a word to be filtered when the attribute similarity between the character attribute of the third sensitive word and the character attribute of the word to be detected is greater than or equal to the third threshold. The fifth determining module is specifically used for: If the similarity between the character attributes of the third sensitive word and the character attributes of the word to be detected is greater than or equal to the third threshold, the large language model is used to perform semantic analysis on the background text of the word to be detected to obtain the first text semantics. The word to be detected in the text block is replaced with the third sensitive word to obtain the replaced text block; The semantics of the replaced text block are obtained by performing semantic analysis on the large language model to obtain the second text semantics. If the semantic similarity between the first text semantic and the second text semantic is greater than or equal to the second threshold, the word to be detected is determined to be a word to be filtered.
6. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Sensitive word detection method and device, electronic equipment and storage medium
CN114417881A
Text processing method and device, medium and electronic equipment
CN115146589A