Keyword extraction method and device, electronic equipment and computer readable storage medium

Through the data enhancement and keyword combination technology of social network platforms, the problems of low efficiency and strong subjectivity of manual analysis are solved, and comprehensive, objective extraction and efficient analysis of key information are achieved.

CN120278150APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410021720.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有技术中,人工分析关键词效率低、主观性强,难以全面覆盖社交网络平台上海量数据中的隐藏关键词,导致关键信息丢失。

Method used

By obtaining the object detection text, data enhancement processing is performed, the initial keyword set is extracted, and the frequent item keyword set is combined based on the co-occurrence relationship, the associated keyword set is obtained and the security level is determined, and the target keyword is finally filtered out.

Benefits of technology

It realizes comprehensive and objective extraction of key information on social network platforms, improves analysis efficiency, reduces subjectivity, can automatically filter irrelevant content, and provides consistent analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278150A_ABST
    Figure CN120278150A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a keyword extraction method and device, electronic equipment and a computer readable storage medium. According to the embodiment of the invention, after a target detection text of a target detection object is obtained and at least one original detection statement is extracted from the target detection text, data enhancement is performed on the original detection statement, and then at least one keyword is extracted from a target detection statement set to obtain an initial keyword set; initial keywords in the initial keyword set are combined according to a co-occurrence relation to obtain at least one frequent item keyword set, then, an associated keyword set having a preset similarity with the frequent item keyword set is obtained, an associated object and the security level of the associated object are determined according to the associated keyword set, and then, based on the security level, the security level of the associated object is determined; screening out at least one target keyword from the frequent item keyword set; according to the scheme, the keywords can be extracted comprehensively, efficiently and objectively; the embodiment of the invention can be applied to scenes such as natural language processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a keyword extraction method, device, electronic device and computer-readable storage medium. Background Art

[0002] Keyword extraction plays an important role in many application scenarios. For example, on social networking platforms, users may be involved in inappropriate behavior, low-quality information dissemination and other issues during the interaction process. Keyword extraction can identify information related to these issues to maintain the safety and health of the social environment. However, due to the large number of users on social networking platforms, a huge amount of data will be generated accordingly. In addition, since the interactive content between users changes dynamically over time, new threats and problems may arise at any time, causing the confrontation between network platforms and improper industries to escalate. This requires keyword extraction to have the ability to respond in real time or even in advance, so as to mine key information with strong confusion or concealment in massive data, so as to quickly respond to new threats and maliciousness.

[0003] Currently, the analysis and extraction of keywords from massive data mainly rely on manual labor, which is inefficient and highly subjective. In addition, due to the limitations of the staff's own cognition and subjective judgment, manual keyword analysis is likely to ignore some hidden keywords, and it is difficult to cover all possible keywords, resulting in the loss of key information, which has obvious limitations. Summary of the invention

[0004] The embodiments of the present application provide a keyword extraction method, device, electronic device and computer-readable storage medium, which can extract keywords comprehensively, efficiently and objectively.

[0005] The present application embodiment provides a keyword extraction method, including:

[0006] Acquire a target detection text of a target detection object, and extract at least one original detection sentence from the target detection text;

[0007] Performing data enhancement on the original detection sentence to obtain a data enhanced sentence;

[0008] Extracting at least one keyword from a target detection sentence set to obtain an initial keyword set, wherein the target detection sentence set includes the original detection sentence and the data enhancement sentence;

[0009] Combining the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent keyword set;

[0010] Obtain an associated keyword set that has a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set;

[0011] Based on the security level, screen out at least one target keyword from the frequent item keyword set.

[0012] Correspondingly, an embodiment of the present application provides a keyword extraction device, including:

[0013] A data acquisition unit, configured to acquire a target detection text of a target detection object, and extract at least one original detection statement from the target detection text;

[0014] A data enhancement unit, configured to perform data enhancement on the original detection statement to obtain a data-enhanced statement;

[0015] A keyword extraction unit, configured to extract at least one keyword from a target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data-enhanced statement;

[0016] A keyword combination unit, configured to combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set;

[0017] An association relationship determination unit, configured to obtain an associated keyword set that has a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set;

[0018] A keyword determination unit, configured to screen out at least one target keyword from the frequent item keyword set based on the security level.

[0019] In some embodiments, the data acquisition unit may specifically be configured to extract at least one target character from the target detection text, where the target character includes at least one of a character or a number; combine the target characters to obtain at least one original detection statement.

[0020] In some embodiments, the data enhancement unit may specifically be configured to perform word segmentation processing on the original detection statement to obtain at least one original word segmentation segment; perform multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-enhanced statement corresponding to the original detection statement.

[0021] In some embodiments, the data augmentation unit may be specifically configured to extract features from the original word segmentation fragments to obtain at least one fragment feature; determine at least one replacement content corresponding to the original word segmentation fragments based on the fragment features; and replace the original word segmentation fragments in the original detection statement with the replacement content to obtain a data augmentation statement corresponding to the original detection statement.

[0022] In some embodiments, the replacement content includes at least one of semantic replacement content, homophone replacement content, or homograph replacement content. The data augmentation unit may be specifically configured to, if the fragment features include semantic features, replace the original word segmentation fragments based on the semantic features to obtain semantic replacement content, where the semantic replacement content includes the content after equivalent replacement or synonymous replacement of the original word segmentation fragments; if the fragment features include syllable features, perform homophone replacement on the original word segmentation fragments based on the syllable features to obtain homophone replacement content; if the fragment features include appearance features, perform homograph replacement on the original word segmentation fragments based on the appearance features to obtain homograph replacement content.

[0023] In some embodiments, the data augmentation unit may be specifically configured to find word segmentation fragments containing text content from the original word segmentation fragments in the original detection statement; find adjacent word segmentation fragments having an adjacent relationship with the word segmentation fragments containing text content from the original word segmentation fragments, where the adjacent word segmentation fragments include text content; exchange the positions of the word segmentation fragments containing text content and the adjacent word segmentation fragments to obtain a reversed statement corresponding to the original detection statement, and use the reversed statement as the data augmentation statement corresponding to the original detection statement.

[0024] In some embodiments, the keyword extraction unit may be specifically configured to obtain a historical detection statement set of a historical detection object; perform word segmentation processing on the target detection statements in the target detection statement set to obtain at least one target word segmentation fragment; determine the target association degree of each target word segmentation fragment with the target detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set and the target detection statement set; and filter out at least one word segmentation fragment from the target word segmentation fragments based on the target association degree to obtain an initial keyword set.

[0025] In some embodiments, the keyword extraction unit may specifically be configured to calculate a first correlation degree between each target word segmentation fragment and the historical detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set; calculate a second correlation degree between each target word segmentation fragment and the target detection statement set based on the recurrence frequency of each target word segmentation fragment in the target detection statement set; and fuse the first correlation degree and the second correlation degree to obtain a target correlation degree between each target word segmentation fragment and the target detection statement set.

[0026] In some embodiments, the keyword extraction unit may specifically be configured to obtain the total number of statements in the historical detection statement set; obtain the number of target statements in the historical detection statement set that contain the current target word segmentation fragment; and calculate the ratio of the number of target statements to the total number of statements to obtain the first correlation degree between the current target word segmentation fragment and the historical detection statement set.

[0027] In some embodiments, the keyword extraction unit may specifically be configured to obtain the total number of fragments of the target word segmentation fragments contained in the target detection statement set; determine the number of target fragments of the current target word segmentation fragment contained in the target detection statement set; and determine the second correlation degree between the current target word segmentation fragment and the target detection statement set based on the ratio of the number of target fragments to the total number of fragments.

[0028] In some embodiments, the keyword extraction unit may specifically be configured to screen out at least one word segmentation fragment with a target correlation degree greater than a preset correlation degree in the target word segmentation fragments as potential keywords; obtain a stop word database corresponding to the target detection object, and clean the potential keywords based on the stop word database to obtain an initial keyword set.

[0029] In some embodiments, the keyword combination unit may specifically be configured to identify the word positions of each initial keyword in the initial keyword set in the target detection statement, and divide the initial keywords based on the word positions to obtain at least one initial keyword subset; construct a frequent pattern tree according to the initial keyword subset, where the frequent pattern tree represents the co-occurrence relationship between the initial keywords; and screen out keywords that meet a preset co-occurrence frequency from the initial keyword set to obtain at least one frequent item keyword set.

[0030] In some embodiments, the association relationship determination unit may specifically be configured to obtain a historical keyword set of a historical detection object, where the historical keyword set includes at least one historical keyword subset; screen out a historical keyword subset from the historical keyword set that has a preset similarity with the frequent item keyword set, and use the screened-out historical keyword subset as the associated keyword set; screen out the historical detection object corresponding to the associated keyword set from the historical detection object to obtain an associated object; and determine the security level of the associated object based on the object identifier of the associated object.

[0031] In some embodiments, the keyword determination unit may specifically be configured to compare the security level with a preset security level, and based on the comparison result, determine a security associated object whose security level is greater than the preset security level and the number of security objects of the security associated object from the associated objects; determine the total number of objects of the associated objects, and based on the ratio between the number of security objects and the total number of objects, determine a target keyword from the frequent item keyword set.

[0032] In some embodiments, the keyword determination unit may specifically be configured to, if the ratio between the number of security objects and the total number of objects is greater than or equal to a preset threshold, use the keywords included in the frequent item keyword set as target keywords; or if the ratio between the number of security objects and the total number of objects is less than the preset threshold, use the keywords included in the frequent item keyword set as invalid keywords.

[0033] In some embodiments, the keyword extraction device may further include a risk control unit, configured to determine the risk level of the target detection object based on the target keyword; and restrict the interaction permissions of the target detection object according to the risk level.

[0034] In addition, an embodiment of the present application further provides an electronic device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to execute the keyword extraction method provided by the embodiment of the present application.

[0035] In addition, an embodiment of the present application further provides a computer program product, including a computer program or instruction, where when the computer program or instruction is executed by a processor, the steps in the keyword extraction method provided by the embodiment of the present application are implemented.

[0036] In addition, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the keyword extraction methods provided by the embodiment of the present application.

[0037] In the embodiment of the present application, after obtaining the target detection text of the target detection object and extracting at least one original detection sentence from the target detection text, the original detection sentence is data enhanced to obtain a data enhanced sentence; then, at least one keyword is extracted from the target detection sentence set to obtain an initial keyword set, and the target detection sentence set includes the original detection sentence and the data enhanced sentence; then, the initial keywords in the initial keyword set are combined according to the co-occurrence relationship to obtain at least one frequent keyword set; then, an associated keyword set with a preset similarity to the frequent keyword set is obtained, and the associated object and the security level of the associated object are determined according to the associated keyword set; then, based on the security level, at least one target keyword is screened out from the frequent keyword set. Since the scheme can perform data enhancement on the existing original detection sentence to increase the diversity of data samples, hidden keywords and keyword information that may be used by improper industries in the future can be inferred based on the diversity samples; then, the co-occurrence relationship between the extracted keywords is identified, and they are screened and combined according to the frequent pattern between the keywords; and the combined keywords are further filtered in combination with the account quality level of the associated object to obtain keywords that can comprehensively and objectively reflect the target detection object. Compared with manual keyword analysis, this solution can dig out key information hidden in text content and automatically filter out irrelevant content, thereby improving the comprehensiveness and objectivity of the analysis results; this solution can handle large-scale text content and improve analysis efficiency; and this solution can provide consistent analysis results, thereby avoiding the subjective problems in manual analysis, and has the advantages of automatic filtering, wide coverage and low time consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0039] Figure 1 It is a schematic diagram of an application scenario of the keyword extraction method provided in an embodiment of the present application;

[0040] Figure 2 It is a flowchart of the keyword extraction method provided in the embodiment of the present application;

[0041] Figure 3 It is a schematic diagram of data enhancement in the keyword extraction method provided in an embodiment of the present application;

[0042] Figure 4 It is a schematic diagram of word segmentation processing in the keyword extraction method provided in the embodiment of the present application;

[0043] Figure 5 It is a schematic diagram in the keyword extraction method provided by the embodiments of the present application;

[0044] Figure 6 It is another schematic flowchart of the keyword extraction method provided by the embodiments of the present application;

[0045] Figure 7 It is a schematic structural diagram of the keyword extraction device provided by the embodiments of the present application;

[0046] Figure 8 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0047] Next, the technical solutions will be clearly and completely described in conjunction with the accompanying drawings therein. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0048] The embodiments of the present application provide a keyword extraction method, a device, and a computer-readable storage medium. Among them, the keyword extraction device can be integrated in an electronic device, and the electronic device can be a server or a user terminal and other devices.

[0049] Among them, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud preset databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0050] Figure 1 It shows a schematic diagram of the application scenario of the keyword extraction method provided by the embodiments of the present application. As Figure 1As shown, take the case where the keyword extraction device is integrated in an electronic device, and the electronic device is a server. On a social networking platform, user terminals where users are located (including at least client A and client B) can respectively establish communication connections with the server, and communication connections can be established between each user terminal. For example, client A and client B can respectively establish communication connections with the server, and a communication connection can also be established between client A and client B. When a communication connection is established between client A and client B, during the process of information interaction between the two, if the user of client A believes that the content sent by client B has problems such as illegal acts, harassment, and malicious information dissemination, the user can report client B to the server and provide corresponding clues. Specifically, client A can send its interaction content with client B (such as communication content) to the server. At this time, client B is the target detection object to be reported, and the interaction content sent by client A regarding client A and client B is the source of the target detection text. If the interaction content between client A and client B is text content, the interaction content can be directly used as the target detection text; if the interaction content between client A and client B is other types of content other than text form (such as audio content, video content), the interaction content can be converted into text content, and the text content obtained after conversion is used as the target detection text.

[0051] After the server obtains the target detection text of the target detection object and extracts at least one original detection statement from the target detection text, it performs data enhancement on the original detection statement to obtain a data-enhanced statement; then, it extracts at least one keyword from the target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data-enhanced statement; then, it combines the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; then, it obtains an associated keyword set having a preset similarity with the frequent item keyword set, and determines the associated object and the security level of the associated object according to the associated keyword set; then, based on the security level, it filters out at least one target keyword from the frequent item keyword set.

[0052] It should be noted that Figure 1 the number of clients can be multiple, and multiple clients can send the target detection text for the same target detection object to the server at the same time, or can send the target detection text for different target detection objects to the server at different times.

[0053] The keyword extraction method provided by the embodiments of this application relates to the natural language processing (NLP) direction in artificial intelligence (AI). The embodiments of this application can obtain the target detection text of the target detection object, extract at least one original detection statement from the target detection text, and then perform data augmentation on the original detection statement to obtain a data-augmented statement; then, extract at least one keyword from the target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data-augmented statement; then, combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; then, obtain an associated keyword set that has a preset similarity with the frequent item keyword set, and determine the associated object and the security level of the associated object according to the associated keyword set; then, based on the security level, screen out at least one target keyword from the frequent item keyword set

[0054] Artificial intelligence is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0055] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0056] Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0057] Among them, it can be understood that in the specific implementation of this application, relevant data such as the object identifier and object image of the object are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0058] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0059] This embodiment will be described from the perspective of the keyword extraction device. The keyword extraction device can be specifically integrated in an electronic device, which can be a server or a terminal device, etc.; among them, the terminal can include devices such as a tablet computer, a laptop computer, a personal computer (PC, Personal Computer), a wearable device, a virtual reality device, or other intelligent devices that can generate mirror files.

[0060] A keyword extraction method includes: obtaining the target detection text of the target detection object, and extracting at least one original detection statement from the target detection text; performing data augmentation on the original detection statement to obtain a data-augmented statement; extracting at least one keyword from the target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data-augmented statement; combining the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; obtaining an associated keyword set having a preset similarity with the frequent item keyword set, and determining the associated object and the security level of the associated object according to the associated keyword set; screening at least one target keyword from the frequent item keyword set based on the security level.

[0061] Figure 2 Shows a schematic flowchart of the keyword extraction method provided by the embodiment of the present application. As Figure 2 shown, the specific process of this keyword extraction method is as follows:

[0062] 101. Obtain the target detection text of the target detection object, and extract at least one original detection statement from the target detection text.

[0063] In the process of a user interacting with other users on a social network platform, the user may receive some content involving illegal acts, harassment, malicious information, etc. At this time, the user can report to the social network platform the user who sent the above content. When a user reports another user to the social network platform, the user can provide corresponding reporting clues or complaint evidence for the social network platform to evaluate the reported user.

[0064] Specifically, when a user reports another user to the social network platform, the user can send the account number of the reported user to the server of the social network platform (hereinafter referred to as the server) as a reporting clue, or send the communication content between the user and the reported user to the server as complaint evidence, or only send the content sent by the reported user to the user to the server as complaint evidence. The communication content between the user and the reported user may include at least one of video, audio, or text sent by the reported user to the user. Similarly, the form of the content sent by the reported user to the user may also include at least one of video, audio, or text. To facilitate the extraction of key information in the complaint evidence provided by the user, the server can convert the video content and audio content into text content, or the server can also directly extract text content from the complaint evidence and finally extract key information from the content in text form.

[0065] In the embodiments of the present application, the target detection object can be understood as the reported user. The target detection text can be understood as the communication content finally presented in text form, or the target detection text can be understood as the content sent by the reported user to the user finally presented in text form. After the server obtains the target detection text of the target detection object, the server can preprocess the target detection text, screen out irrelevant characters in the target detection text, and combine the remaining relevant characters into statements to obtain at least one original detection statement. Among them, the relevant characters can be letters and numbers, and the irrelevant characters can be other characters other than letters and numbers, such as underscores, mathematical operators, punctuation marks, etc.

[0066] Among them, there are various ways to extract at least one original detection statement from the target detection text. For example, it may include: extracting at least one target character from the target detection text, where the target character includes at least one of text or numbers; combining the target characters to obtain at least one original detection statement. The target character can be understood as the relevant characters filtered out from the target detection text. The type of the target character can be at least one of text or numbers, that is, the target characters can all be text, or all be numbers, or can include both text and numbers. The text in the target characters can include Chinese characters, English letters, and also the characters corresponding to other languages. For example, the text in the target characters can all be Chinese characters, or part of them can be Chinese characters and part can be English characters. Of course, it can also be a combination of characters corresponding to other languages, which will not be listed one by one here. The numbers in the target characters can include Arabic numerals, Roman numerals, and also Chinese numerals, etc. For example, the numbers in the target characters can all be Arabic numerals, or part of them can be Arabic numerals and part can be Chinese numerals, or it can be a combination of Arabic numerals, Roman numerals, or Chinese numerals, which will not be listed one by one here.

[0067] In the embodiments of the present application, the target characters can be screened out from the target detection text by using the regular matching method, and then the screened target characters are arranged in the order in which they appear in the target detection text and combined together to obtain at least one original detection statement. Similarly, the irrelevant characters (such as underlines, mathematical operators, punctuation marks, etc.) in the target detection text can also be identified by using the regular matching method. After filtering out the irrelevant characters, the remaining characters are the relevant characters containing text and / or numbers. The remaining characters are arranged in the order in which they appear in the target detection text and combined together to obtain at least one original detection statement.

[0068] The original detection statement can be understood as the statement obtained by filtering out the irrelevant characters from each statement contained in the target detection text. At this time, in the process of combining the target characters, according to the sentence-breaking method of each statement in the target detection text, the target characters can be directly sorted in order and combined to obtain the corresponding original detection statement.

[0069] In the specific application process, the length of each sentence in the target detection text is different, and the number of irrelevant characters contained in each sentence is also different. Accordingly, after filtering the irrelevant characters in the target detection text, the length of the sentence obtained by sorting and combining the remaining relevant characters in order is also different. Considering that the length and number of the original detection text will affect the efficiency of further data processing, at this time, in the process of extracting the original detection sentence from the target detection text, the length of each original detection sentence and the number of original detection sentences can be controlled within an appropriate range. For example, for a long sentence in the target detection text, the irrelevant characters in the long sentence can be directly filtered, and the remaining target characters can be sorted according to their original positions in the sentence, and recombined together to obtain the corresponding original detection sentence. For short sentences mixed in long sentences, if there are a large number of short sentences mixed between two long sentences, these short sentences can be combined into a long sentence, and then according to the processing method for the long sentence, a corresponding original detection sentence is obtained. If the number of short sentences sandwiched between two long sentences is small, these short sentences can be directly combined with one of the adjacent long sentences to obtain a combined long sentence, and then a corresponding original detection sentence is obtained according to the processing method for the long sentence.

[0070] It should be understood that if the proportion of irrelevant characters in a long sentence is relatively large, it can be combined with adjacent sentences to ensure that the original detection sentence obtained has a uniform length.

[0071] The original detection sentence can also be understood as the sentence obtained after filtering out irrelevant characters from each paragraph contained in the target detection text. At this time, in the process of combining the target characters, each paragraph in the target detection text can be regarded as a sentence, and the target characters are sorted in order according to the segmentation method of each paragraph in the target detection text, and combined together to obtain the corresponding original detection sentence.

[0072] In the specific application process, the text length of the target detection text, the number of paragraphs in the target detection text, the length of each sentence in the target detection text, and the number of sentences contained in the target detection text can be comprehensively considered, and an appropriate method can be selected to combine them to obtain a moderate number of original detection sentences with moderate length of each sentence.

[0073] 102. Perform data enhancement on the original detection sentence to obtain a data enhanced sentence.

[0074] Data Augmentation is a method of expanding the training data set by using a small amount of data to generate more similar generated data through prior knowledge. In the embodiments of the present application, after extracting at least one original detection statement from the target detection text, data augmentation can be performed on each of the extracted original detection statements to obtain at least one data-augmented statement. Specifically, the same data augmentation method can be used to process each of the extracted original detection statements, or different data augmentation methods can be used to process each of the extracted original detection statements. Data augmentation can be performed multiple times on each original detection statement, or data augmentation can be performed once on each original detection statement.

[0075] Figure 3 FIG. shows a schematic diagram of data augmentation in the keyword extraction method provided by the embodiments of the present application. As Figure 3 shown, there are various ways to perform data augmentation on the original detection statement to obtain a data-augmented statement. For example, it may include: performing word segmentation on the original detection statement to obtain at least one original word segmentation segment; performing multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-augmented statement corresponding to the original detection statement.

[0076] Performing word segmentation on the original detection statement means recombining the continuous text and numbers in the original detection statement into a word sequence according to certain specifications. As mentioned above, the target text may include Chinese characters, English letters, and characters corresponding to other languages. For different text types, different word segmentation methods can be used for word segmentation due to language structure differences. For the convenience of explanation, the embodiments of the present application will take the target text as Chinese characters and the target numbers as Arabic numerals as an example for illustration.

[0077] In the embodiment of the present application, a common Chinese word segmentation tool can be used to perform word segmentation on the original detection sentence. For example, the Jieba word segmentation tool (also known as the Jieba word segmentation tool) can be used to perform word segmentation on the original detection sentence, SnowNLP (a Chinese word segmentation tool) can be used to perform word segmentation on the original detection sentence, THULAC (a Chinese word segmentation tool) can be used to perform word segmentation on the original detection sentence, and NLPIR-ICTCLAS (a Chinese word segmentation system) can also be used to perform word segmentation on the original detection sentence. Of course, other types of Chinese word segmentation tools can also be used to perform word segmentation on the original detection sentence, which are not listed one by one here. Among them, taking the Jieba word segmentation tool as an example, it can provide three word segmentation modes, namely, precise mode, full mode and search engine mode. In the above three modes, the precise mode can cut the sentence most accurately; the full mode can quickly scan out all the words that can be formed into words in the sentence; the search engine mode can segment the long words again on the basis of the precise mode. In an embodiment of the present application, any one of the above three modes can be used to perform word segmentation processing on the original detection sentence, or the above three modes can be combined to perform word segmentation processing on the original detection sentence.

[0078] After the original detection sentence is processed by word segmentation, at least one original word segmentation fragment can be obtained. The type of the original word segmentation fragment can include at least one of a phrase, a word, a single number, and a single character. For example, if the original detection sentence is "The weather is good today", three original word segmentation fragments can be obtained after word segmentation processing, namely "today", "weather", and "good". When the original detection sentence contains numbers, the numbers can be counted separately into an original word segmentation fragment. For example, if the original detection sentence is "I spent 100 yuan on a meal today", the following original word segmentation fragments can be obtained after word segmentation processing, namely "I", "today", "meal", "spend", "100", and "yuan". When the numbers in the original detection sentence are distributed at intervals, the numbers distributed at intervals can be counted into an original word segmentation fragment respectively. For example, if the original detection sentence is "Xiao Ming is 10 years old and Xiao Hua is 8 years old", the following original word segmentation fragments can be obtained after word segmentation processing, namely "Xiao Ming", "10", "years old", "Xiao Hua", "8", and "years old".

[0079] It should be understood that the word segmentation processing results in the above examples are only for explaining the technical solutions of the embodiments of the present application. In actual applications, different word segmentation processing results may be obtained for the original detection sentences in the above examples. Therefore, the above examples do not have a limiting effect on the technical solutions of the present application.

[0080] After obtaining the original tokenized segments corresponding to the original detection statement, the original tokenized segments can be transformed from multiple dimensions to obtain at least one data augmentation statement corresponding to the original detection statement. Specifically, a certain original tokenized segment in the original detection statement can be transformed from multiple dimensions, and each dimension transformation can obtain a data augmentation statement; alternatively, a certain original tokenized segment in the original detection statement can be transformed from a certain dimension to obtain a data augmentation statement; or a certain original tokenized segment in the original detection statement can be transformed from a certain dimension for each original tokenized segment; in addition, each original tokenized segment in the original detection statement can be transformed from multiple dimensions.

[0081] In the embodiments of the present application, performing multi-dimensional transformation on the original tokenized segments in the original detection statement can be understood as transforming the original tokenized segments in the original detection statement by using different transformation methods. There are various ways to perform multi-dimensional transformation on the original tokenized segments in the original detection statement to obtain the data augmentation statement corresponding to the original detection statement. For example, it can include: extracting features from the original tokenized segments to obtain at least one segment feature; based on the segment feature, determining at least one replacement content corresponding to the original tokenized segment; and replacing the original tokenized segment in the original detection statement with the replacement content to obtain the data augmentation statement corresponding to the original detection statement.

[0082] The segment feature can include the semantic feature of the original tokenized segment, and the semantic feature can represent the language meaning of the original tokenized segment; the segment feature can also include the syllable feature of the original tokenized segment, and the syllable feature can represent the pronunciation of the original tokenized segment; the segment feature can also include the appearance feature of the original tokenized segment, and the appearance feature can represent the writing form of the original tokenized segment. In the embodiments of the present application, for different types of segment features extracted, after transforming the original tokenized segment, the corresponding replacement content can be obtained. For example, after transforming the original tokenized segment based on the semantic feature, the corresponding semantic replacement content can be obtained; after transforming the original tokenized segment based on the syllable feature, the corresponding homophone replacement content can be obtained; after transforming the original tokenized segment based on the appearance feature, the corresponding homograph replacement content can be obtained. Therefore, the replacement content can include at least one of semantic replacement content, homophone replacement content, or homograph replacement content.

[0083] In the embodiments of the present application, for different types of segment features extracted, different transformation methods (i.e., from different dimensions) can be used to transform the original tokenized segment to obtain the corresponding replacement content. Based on the segment feature, determining at least one replacement content corresponding to the original tokenized segment can include the following situations:

[0084] (1) If the fragment feature includes a semantic feature, the original word segmentation fragment is replaced based on the semantic feature to obtain semantic replacement content, and the semantic replacement content includes the content after the original word segmentation fragment is equivalently replaced or synonymously replaced. As mentioned above, the semantic feature can characterize the linguistic meaning of the original word segmentation fragment. If the fragment feature of the original word segmentation fragment includes a semantic feature, the original word segmentation fragment can be synonymously replaced, or the original word segmentation fragment can be equivalently replaced. For example, if the original word segmentation fragment is "不要", the original word segmentation fragment can be synonymously replaced based on its semantic feature, and the obtained semantic replacement content can include "勿", "不", "請勿", etc. If the original word segmentation fragment is "1", the original word segmentation fragment can be equivalently replaced based on its semantic feature, and the obtained semantic replacement content can include "一", "壹", "①", etc.

[0085] It should be understood that after the original word segmentation fragments are transformed based on semantic features, the resulting semantic replacement content may change from the original text to numbers, may change from the original numbers to text, may change from the original single characters or words to words or phrases composed of two or more characters or words, and may change from the original words or source domains composed of two or more characters or words to single characters or words.

[0086] (2) If the segment features include syllable features, homophonic replacement is performed on the original word segment segment based on the syllable features to obtain homophonic replacement content. As mentioned above, the syllable features can characterize the pronunciation of the original word segment segment. If the segment features of the original word segment segment include syllable features, homophonic replacement can be performed on the original word segment segment to obtain corresponding homophonic replacement content. For example, if the original word segment segment is "的", homophonic replacement is performed on the original word segment segment based on its pronunciation features, and the obtained homophonic replacement content may include "地", "得", "德", "嘚", "德", etc.

[0087] (3) If the segment feature includes an appearance feature, the original segment segment is replaced with an isomorphic replacement based on the appearance feature to obtain isomorphic replacement content. As mentioned above, the appearance feature can characterize the written form of the original segment segment. If the segment feature of the original segment segment includes an appearance feature, the original segment segment can be replaced with an isomorphic replacement to obtain corresponding isomorphic replacement content. For example, if the original segment segment is "人", based on its appearance feature after writing, the original segment segment is replaced with an isomorphic replacement, and the obtained isomorphic replacement content may include "入". If the original segment segment is "已", based on its appearance feature after writing, the original segment segment is replaced with an isomorphic replacement, and the obtained isomorphic replacement content may include "己", "巳", etc.

[0088] In addition to the multi-dimensional transformation methods adopted above (such as equivalent replacement, synonym replacement, homophone replacement, homograph replacement, etc.), which perform multi-dimensional transformation on the original word segmentation fragments in the original detection statement to obtain the data augmentation statement corresponding to the original detection statement, it can also include: finding the word segmentation fragments containing text content from the original word segmentation fragments in the original detection statement; finding the adjacent word segmentation fragments having an adjacent relationship with the word segmentation fragments from the original word segmentation fragments, where the adjacent word segmentation fragments include text content; swapping the positions of the word segmentation fragments and the adjacent word segmentation fragments to obtain the inverted statement corresponding to the original detection statement, and using the inverted statement as the data augmentation statement corresponding to the original detection statement. The word segmentation fragment can be understood as the original word segmentation fragment that only contains text and does not contain numbers. The word segmentation fragment can be a single character, or a word or phrase composed of two or more characters. The adjacent word segmentation fragment can also be understood as the original word segmentation fragment that only contains text and does not contain numbers. The adjacent word segmentation fragment can be a single character, or a word or phrase composed of two or more characters. The adjacent word segmentation fragment is adjacent to the word segmentation fragment, and the adjacent word segmentation fragment can be located before or after the word segmentation fragment.

[0089] For example, taking the original detection statement "I spent 100 yuan on eating today" and the obtained original word segmentation fragments (respectively "I", "today", "eating", "spent", "100", "yuan") as an example for illustration. Among them, since the original word segmentation fragments "I", "today", "eating", "spent", and "yuan" do not contain numbers, they can all be used as word segmentation fragments. For the case where the word segmentation fragment is "today", the adjacent word segmentation fragments having an adjacent relationship with this word segmentation fragment include "I" and "eating", where "I" is located before "today" and "eating" is located after "today". At this time, the order of "I" and "today" can be swapped to obtain the inverted statement "Today I spent 100 yuan on eating"; or the order of "eating" and "I" can be swapped to obtain the inverted statement "I ate today and spent 100 yuan". For the case where the word segmentation fragment is "spent", since "100" behind "spent" includes numbers, the adjacent word segmentation fragment having an adjacent relationship with this word segmentation fragment ("spent") is "eating". At this time, the order of "eating" and "spent" can be swapped to obtain the inverted statement "I spent eating today for 100 yuan". The inverted statement obtained by swapping the positions of the word segmentation fragment and the adjacent word segmentation fragment is the data augmentation statement corresponding to the original detection statement.

[0090] As mentioned above, in actual applications, different word segmentation processing results may be obtained for the original detection sentences in the above examples. For example, if the original detection sentence is "I spent 100 yuan on meals today", the original word segmentation fragments obtained are "I", "today", "eat", "meal", "spend", "fee", "100", "yuan". Among them, since the original word segmentation fragments of "I", "today", "eat", "meal", "spend", "fee", "yuan" do not contain numbers, they can all be used as text segmentation fragments. For the case where the text segmentation fragment is "meal", the adjacent segmentation fragments corresponding to the text segmentation fragment include "eat" and "spend". At this time, the order of "eat" and "meal" can be reversed, and the reversed sentence obtained is "I spent 100 yuan on meals today"; the order of "meal" and "spend" can also be reversed, and the reversed sentence obtained is "I spent 100 yuan on meals today".

[0091] It should be understood that when the original detection sentence contains multiple text segmentation fragments, the corresponding adjacent word segmentation fragments can be determined according to the adjacent relationship between the multiple text segmentation fragments, and the order of one or more of the text segmentation fragments and the corresponding adjacent word segmentation fragments can be randomly reversed to obtain at least one reversed sentence.

[0092] In the embodiment of the present application, the characters contained in the original detection sentence can be directly distinguished according to individual characters, and the adjacent relationship can be determined according to the arrangement order, and then the order of the individual characters with the adjacent relationship can be randomly reversed to obtain at least one reversed sentence corresponding to the original detection sentence. For example, taking the original detection sentence "The weather is good today" as an example, after distinguishing according to the individual characters therein and randomly reversing the individual characters therein according to the adjacent relationship, the reversed sentence obtained can be "Is the weather good today?", "The weather is good today?", "Is the weather bad today?", etc.

[0093] It should be understood that the above-mentioned various methods of performing multi-dimensional transformation on the original word segmentation fragments in the original detection sentence to obtain the data enhancement sentence corresponding to the original detection sentence can be used simultaneously, can be used in combination in any combination, or can be used separately.

[0094] In an embodiment of the present application, the original word segmentation fragments in the original detection sentence are transformed through the above-mentioned multi-dimensional transformation (such as equivalent replacement, synonymous replacement, homophonic replacement, homomorphic replacement or order reversal), which can not only generate more text sentences, but also mine the keyword information hidden in the original detection sentence and the keyword information that may be used by the improper industry in the future, thereby improving the generalization ability of keyword extraction.

[0095] 103. Extract at least one keyword from the target detection statement set to obtain an initial keyword set. The target detection statement set includes original detection statements and data augmentation statements.

[0096] After data augmentation is performed on the original detection statements to generate multiple data augmentation statements, the number of target detection statements available for key extraction increases significantly. The target detection statements can be original detection statements or data augmentation statements. Since the target detection statements belong to the target detection statement set, the target detection statement set includes original detection statements and data augmentation statements. Among them, when extracting keywords from the target detection statement set, keywords can be extracted from at least one target detection statement in the target detection statement set, and the extracted keywords form an initial keyword set corresponding to the target detection object.

[0097] There are various ways to extract keywords from at least one target detection statement in the target detection statement set. For example, it can be to extract at least one keyword from the original detection statements in the target detection statement set, or to extract at least one keyword from the data augmentation statements in the target detection statement set, or to extract at least one keyword from both the original detection statements and the data augmentation statements in the target detection statement set simultaneously.

[0098] Among them, extracting at least one keyword from the target detection statement set to obtain an initial keyword set may include: obtaining the historical detection statement set of the historical detection object; performing word segmentation on the target detection statements in the target detection statement set to obtain at least one target word segmentation fragment; determining the target relevance of each target word segmentation fragment to the target detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set and the target detection statement set; and screening out at least one word segmentation fragment from the target word segmentation fragments based on the target relevance to obtain an initial keyword set.

[0099] The historical detection object can be understood as a user who has been reported before the keywords included in the target detection sentence set are extracted. The historical detection sentence set includes multiple historical detection sentence subsets, and each historical detection object corresponds to a historical detection sentence subset. For each historical detection object, historical detection text can be extracted from the reported content for this historical detection object, and at least one historical detection sentence can be extracted from the historical detection text, so that the historical detection sentence subset of this historical detection object can be obtained. If data augmentation is performed on the historical detection sentences, correspondingly, the historical detection sentence subset can also include the data augmentation sentences obtained after data augmentation. In the embodiments of the present application, in the historical detection sentence subset of each historical detection object, it can include only the historical detection sentences extracted from the historical detection text, or only the data augmentation sentences obtained by performing data augmentation based on the historical detection sentences, or it can include both historical detection sentences and data augmentation sentences.

[0100] Since in the stage of performing data augmentation on the original detection sentences, the original tokenized segments can be transformed based on at least one of the semantic features, syllable features, and appearance features of the original tokenized segments, and the order of the word tokenized segments in the original tokenized segments can also be reversed based on the adjacency relationship, the content in the data augmentation sentences is different from that in the original detection sentences. At this time, the target detection sentences in the target detection sentence set can be tokenized to obtain at least one target tokenized segment. As mentioned above, the original detection sentences can be tokenized in the stage of performing data augmentation on the original detection sentences. Therefore, in the process of tokenizing the target detection sentences, only the data augmentation sentences among them can be tokenized, and the original detection sentences are not tokenized repeatedly. Of course, in the process of tokenizing the target detection sentences, the original detection sentences can also be tokenized, and the data augmentation sentences can also be tokenized.

[0101] In the embodiments of the present application, the target detection sentences can be tokenized using the same tokenization process as the original detection sentences. For example, common Chinese tokenization tools (such as Jieba tokenization tool, SnowNLP, THULAC, NLPIR-ICTCLAS Chinese tokenization system, etc.) can be used to tokenize the original detection sentences. Taking the Jieba tokenization tool as an example, a directed acyclic graph (i.e., DAG graph) can be constructed based on the local word library in a forward traversal manner, so as to dynamically plan to find the maximum probability path and find the maximum segmentation combination based on word frequency, thereby dividing the target detection sentence into individual target tokenized segments.

[0102] Figure 4 Shows the tokenization process schematic diagram in the keyword extraction method provided by the embodiments of the present application, such as Figure 4As shown, taking "designated brand cars" as an example, first, a word lookup tree (Trie tree) model is established based on the local prefix dictionary, and a directed acyclic graph is generated for all possible word combinations of Chinese characters in the sentence through efficient word graph scanning. Second, the maximum probability path is found through dynamic programming to find the maximum segmentation combination based on word frequency. Specifically, taking the word "cars" in Figure 4 as an example, the corresponding probability path can be expressed as Calculating each probability path from back to front, if it is found that the probability of combining "qi" and "che" is greater than the probability of separating them, they are combined into "cars". And so on, finally, the maximum probability path is obtained, and thus the maximum probability segmentation combination corresponding to the sentence is obtained.

[0103] If the content (such as characters, words or phrases, etc.) in the target detection sentence is not included in the local prefix dictionary, the content in the target detection sentence can be segmented based on the hidden Markov model (also known as the HMM model).

[0104] After segmenting the target detection sentences in the target detection sentence set to obtain at least one target segmentation segment, the recurrence frequency of each target segmentation segment can be counted. The recurrence frequency of the target segmentation segment can include its recurrence frequency in the historical detection sentence set and can also include its recurrence frequency in the target detection sentence set. Among them, the recurrence frequency of the target segmentation segment in the historical detection sentence set can be understood as the number of times the target segmentation segment appears in the historical detection sentence set, and the recurrence frequency of the target segmentation segment in the target detection sentence set can be understood as the number of times the target segmentation segment appears in the target detection sentence set. The more times it appears, the greater the corresponding recurrence frequency.

[0105] After determining the recurrence frequency of each target segmentation segment, its target relevance to the target detection sentence set can be determined based on the recurrence frequency. Among them, determining the target relevance of each target segmentation segment to the target detection sentence set based on the recurrence frequency of each target segmentation segment in the historical detection sentence set and the target detection sentence set can include: calculating the first relevance of each target segmentation segment to the historical detection sentence set based on the recurrence frequency of each target segmentation segment in the historical detection sentence set; calculating the second relevance of each target segmentation segment to the target detection sentence set based on the recurrence frequency of each target segmentation segment in the target detection sentence set; and fusing the first relevance and the second relevance to obtain the target relevance of each target segmentation segment to the target detection sentence set.

[0106] For a certain target word segmentation fragment, its first association degree with the historical detection statement set can be understood as: the inverse document frequency (abbreviated as IDF) of the target word segmentation fragment; its second association degree with the target detection statement set can be understood as: the term frequency (abbreviated as TF) of the target word segmentation fragment. The target association degree obtained by fusing the first association degree and the second association degree can be understood as the multiplication of the inverse document frequency and the term frequency corresponding to the target word segmentation fragment, that is, calculating the TF-IDF value.

[0107] The main idea of TF-IDF is: If a certain word or phrase appears frequently (TF) in an article and rarely appears in other articles, it is considered that this word or phrase has good category discrimination ability and is suitable for classification. Correspondingly, in the embodiments of the present application, by using the above method to calculate the target association degree of each target word segmentation fragment, it is possible to well distinguish which target word segmentation fragments in the target detection statement set are suitable as keywords, so that the initial keywords in the target detection statement set can be determined according to the magnitude of the target association degree.

[0108] Among them, determining the first association degree of each target word segmentation fragment with the historical detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set may include: obtaining the total number of statements in the historical detection statement set; obtaining the number of target statements in the historical detection statement set that contain the current target word segmentation fragment; calculating the ratio of the number of target statements to the total number of statements to obtain the first association degree of the current target word segmentation fragment with the historical detection statement set.

[0109] Among them, there are various ways to obtain the number of target statements in the historical detection statement set that contain the current target word segmentation fragment. For example, each subset of historical detection statements corresponding to each historical detection object can be regarded as a long statement, then there are multiple long statements in the historical detection statement set, and at this time, the total number of statements in the historical detection statement set is the same as the number of historical detection objects. For example, it is also possible to directly add up the historical detection statements included in each subset of historical detection statements, and the added value is based on the total number of statements in the historical detection statement set.

[0110] In the embodiments of the present application, the calculation formula of the first association degree can be expressed as:

[0111]

[0112] Among them, the target word segmentation fragment is the current target word segmentation fragment, and the larger the IDF value, the lower the frequency of the target word segmentation fragment appearing in the historical detection statement set.

[0113] Among them, based on the second association between each target word segmentation fragment and the target detection sentence set, it can include: obtaining the total number of target word segmentation fragments contained in the target detection sentence set; determining the target fragment number of the current target word segmentation fragment contained in the target detection sentence set; based on the ratio of the target fragment number to the total number of fragments, determining the second association between the current target word segmentation fragment and the target detection sentence set.

[0114] In the embodiment of the present application, the calculation formula of the second correlation degree can be expressed as:

[0115]

[0116] Among them, the larger the TF value, the higher the frequency of occurrence of the target word segment in the target detection sentence set.

[0117] Among them, based on the target relevance, at least one word segmentation fragment is screened out from the target word segmentation fragments to obtain an initial keyword set, which may include: screening out at least one word segmentation fragment whose target relevance is greater than a preset relevance from the target word segmentation fragments as a potential keyword; obtaining a stop word database corresponding to the target detection object, and based on the stop word database, cleaning the potential keywords to obtain an initial keyword set.

[0118] After calculating the target relevance of each target word segment, the initial keyword can be selected according to the size of each target relevance. Specifically, each target relevance can be compared with the preset relevance, and the target word segment with a target relevance greater than the preset relevance can be screened out. At this time, the screened target word segment can be used as a potential keyword. It is also possible to arrange each target relevance in order from high to low relevance, and take the target word segment corresponding to at least one target relevance ranked at the top as a potential keyword. At this time, the target relevance corresponding to the potential keyword is greater than or equal to the preset relevance.

[0119] After determining the potential keywords, the potential keywords may be compared with the contents recorded in the stop word database, and the potential keywords that do not appear in the stop word database may be the initial keywords in the initial keyword set.

[0120] It should be understood that in the embodiments of the present application, a stop word database can also be directly used to clean the target word segmentation fragments obtained after word segmentation processing, and the target word segmentation fragments that do not appear in the stop word database are screened out, and then based on the recurrence frequency of each screened target word segmentation fragment in the historical detection sentence set and the target detection sentence set, the target correlation of each target word segmentation fragment with the target detection sentence set is determined, and the target word segmentation fragment with a target correlation greater than a preset correlation is used as the initial keyword.

[0121] In the embodiments of the present application, at least one keyword is extracted from a target detection statement set including an original detection statement and a data augmentation statement to obtain an initial keyword set. Since the target detection statement set also includes a data augmentation statement, and as described above, the data augmentation statement may include keyword information hidden in the original detection statement and keyword information that may be used by improper industries in the future. Therefore, keywords hidden in the original detection statement and keywords that may be used by improper industries in the future can be extracted from the target detection statement set, thereby improving the generalization ability of the extracted keywords.

[0122] 104. The initial keywords in the initial keyword set are combined according to the co-occurrence relationship to obtain at least one frequent item keyword set.

[0123] In the embodiments of the present application, the co-occurrence relationship can be understood as the probability of co-occurrence between the initial keywords in the initial keyword set, which can reflect the association relationship between the initial keywords. The greater the probability of co-occurrence between some initial keywords, the stronger the co-occurrence relationship of these initial keywords. Although improper industries have made special treatments on certain words or expressions, in the content sent by improper industries, there are obvious co-occurrence relationships among some words and expressions. Even though the specific text content sent is different, they usually revolve around specific keywords or phrases. For example, in the fraud of impersonating a teacher, certain keywords in the content sent by improper industries always appear simultaneously, such as "parent", "material", and "tuition fee", etc. Based on this co-occurrence relationship, potential fraudulent texts can be identified.

[0124] The frequent item keyword set can be understood as a set of initial keywords that often co-occur in the initial keyword set. Different frequent item keyword sets can be generated according to the different co-occurrence relationships between the initial keywords. Based on this, by mining the co-occurrence relationship between the initial keywords, keyword combinations (frequent item sets) that frequently co-occur in the initial keywords can be identified and extracted, thereby obtaining at least one frequent item keyword set.

[0125] Among them, there are various ways to combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set. For example, it may include: identifying the word positions of each initial keyword in the initial keyword set in the target detection statement, and based on the word positions, dividing the initial keywords to obtain at least one initial keyword subset; constructing a frequent pattern tree according to the initial keyword subset, and the frequent pattern tree represents the co-occurrence relationship between the initial keywords; screening out keywords that meet the preset co-occurrence frequency from the initial keyword set based on the co-occurrence relationship to obtain at least one frequent item keyword set.

[0126] The frequent pattern tree (also known as FP-tree) is a data structure for efficiently mining frequent patterns. The FP-tree is an extension based on the word search tree (also known as Trie tree) and is used to store and represent frequent patterns in data. In the embodiments of the present application, the frequent pattern growth algorithm (also known as FP-growth algorithm) can be used to traverse all the initial keywords in the initial keyword set to construct a frequent pattern tree, and the frequent item keyword set can be mined from the frequent pattern tree. Specifically, the word position of each initial keyword in the initial keyword set can be determined first. The word position can be understood as the arrangement order of the initial keywords in the initial keyword set. After determining the word position of each initial keyword, the initial keywords can be divided into multiple initial keyword subsets according to the arrangement order of the initial keywords. Among them, each initial keyword subset contains multiple initial keywords, and the initial keywords in each initial keyword subset are arranged in their original relative order in the initial keyword set. Since each initial keyword subset can contain multiple initial keywords, the multiple initial keywords in each initial keyword subset can be used as a traversal statement after being arranged in order, and multiple initial keyword subsets can be regarded as multiple traversal statements. During the traversal of each traversal statement, a frequent pattern tree can be constructed according to the arrangement order of the initial keywords included therein.

[0127] There are various ways to divide the initial keywords into multiple initial keyword subsets according to the arrangement order of the initial keywords. For example, each target detection statement to which the initial keyword belongs can be used as a division criterion, and the initial keywords belonging to the same target detection statement can be arranged in order to obtain a corresponding initial keyword subset. Of course, two or more target detection statements with adjacent relationships can also be merged into one statement, and the merged statement can be used as a division criterion, and the initial keywords belonging to the statement can be arranged in order to obtain a corresponding initial keyword subset.

[0128] By constructing a frequent pattern tree, the co-occurrence relationship between the initial keywords can be discovered. Thus, based on the co-occurrence relationship, the keywords that meet the preset co-occurrence frequency can be screened out from the initial keyword set to obtain at least one frequent item keyword set. For example, if the initial keyword set contains initial keywords such as "parent", "diaper", "information", "beer", "tuition fee", "orange juice" and "supermarket". By constructing a frequent pattern tree, it is found that the co-occurrence frequencies of "parent", "information" and "tuition fee" are relatively high, and the co-occurrence frequencies of "diaper", "beer" and "supermarket" are relatively high, while the co-occurrence frequency of "orange juice" with other initial keywords is relatively low. At this time, two frequent item keyword sets can be formed. One of the frequent item keyword sets contains the following keywords: "parent", "information" and "tuition fee"; the other frequent item keyword set contains the following keywords: "diaper", "beer" and "supermarket".

[0129] In the embodiments of the present application, by identifying the co-occurrence relationship between the extracted keywords and screening and combining them according to the frequent patterns between the keywords, keywords with a relatively high degree of association in the target detection text can be screened out and grouped into a corresponding frequent item set (i.e., a frequent item keyword set). Furthermore, based on the obtained at least one frequent item set and the keywords included in each frequent item set, potential fraudulent texts in the target detection text can be identified.

[0130] 105. Obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set.

[0131] As described above, after combining the initial keywords in the initial keyword set according to the co-occurrence relationship, at least one frequent item keyword set can be obtained. On this basis, each frequent item keyword set can be compared with the historical keyword set of the historical detection object, and an associated keyword set having a preset similarity with the frequent item keyword set can be screened out from the historical keyword set. Therefore, the associated keyword set corresponds to the historical detection object, and the frequent item keyword set corresponds to the target detection object. A frequent item keyword set may be associated with multiple associated keyword sets or may not have a corresponding associated keyword set.

[0132] In the embodiments of the present application, the similarity between the initial keywords included in the frequent item keyword set and the keywords in the associated keyword set is greater than or equal to the preset similarity. Among them, the similarity between the initial keywords included in the frequent item keyword set and the keywords in the associated keyword set can be understood as: the percentage of the number of initial keywords in the frequent item keyword set that appear in the associated keyword set in the total number of initial keywords in the frequent item keyword set. The similarity between the initial keywords included in the frequent item keyword set and the keywords in the associated keyword set can also be understood as: the overlap degree between the keywords included in the frequent item keyword set and the associated keyword set respectively. If the overlap degree between the keywords included in the frequent item keyword set and the associated keyword set respectively is higher, that is, the proportion of the keywords that appear in both the frequent item keyword set and the associated keyword set is higher, then the similarity between the initial keywords included in the frequent item keyword set and the keywords in the associated keyword set is higher.

[0133] In the embodiments of the present application, the specific value of the preset similarity can be set or adjusted according to the actual situation. For example, the preset similarity can be any percentage between 0-100%, where the larger the percentage value, the higher the similarity. If the preset similarity is 100%, it is required that the initial keywords included in the frequent item keyword set are exactly the same as the keywords in the associated keyword set. Specifically, for example, the preset similarity can be 50%, 60%, 70%...

[0134] Since the associated keyword set corresponds to the historical detection object, after determining the associated keyword set having a preset similarity with the frequent item keyword set, the corresponding historical detection object (i.e., the associated object) and the security level of the historical detection object (i.e., the security level of the associated object) can be determined according to the associated keyword set.

[0135] Among them, obtaining the associated keyword set having a preset similarity with the frequent item keyword set and determining the associated object according to the associated keyword set may include: obtaining the historical keyword set of the historical detection object, where the historical keyword set includes at least one historical keyword subset; screening out the historical keyword subset having a preset similarity with the frequent item keyword set from the historical keyword set, and using the screened historical keyword subset as the associated keyword set; screening out the historical detection object corresponding to the associated keyword set from the historical detection objects to obtain the associated object; determining the security level of the associated object based on the object identifier of the associated object.

[0136] As mentioned above, the historical detection object can be understood as a user who has been reported before the keywords included in the target detection statement set are extracted. For each historical detection object, the server has performed keyword extraction on the corresponding historical detection text of it, and may obtain at least one keyword. The historical keyword subset can be understood as a set composed of at least one keyword extracted from the historical detection text corresponding to each historical detection object, and one historical detection object corresponds to one historical keyword subset. Correspondingly, the historical keyword set includes multiple historical keyword subsets corresponding to multiple historical detection objects.

[0137] In the embodiment of the present application, each frequent item keyword subset of the target detection object can be compared with each historical keyword subset, and according to the comparison result, the historical keyword subset having a similarity greater than the preset similarity with the frequent item keyword set is determined, and the historical keyword subset satisfying the above conditions is used as the associated keyword set corresponding to the frequent item keyword set. At the same time, the historical detection object corresponding to the historical keyword subset satisfying the above conditions can be used as the associated object. Since the server has performed keyword extraction on the historical detection text corresponding to the historical detection object, that is, the server has evaluated the security level of the historical detection object, the server can add an object identifier to the evaluated historical detection object, and the security level of the historical detection object can be determined through the object identifier. Therefore, after determining the associated object, the security level of the associated object can be determined based on the object identifier of the associated object.

[0138] Figure 5 Shows a schematic diagram in the keyword extraction method provided by the embodiment of the present application. As Figure 5As shown, if the initial keywords in the initial keyword set are combined according to the co-occurrence relationship, two frequent item keyword sets are obtained, namely frequent item keyword set A and frequent item keyword set B. Among them, the 5 initial keywords included in frequent item keyword set A are A1, A2, A3, A4, and A5 respectively; the 6 initial keywords included in frequent item keyword set B are B1, B2, B3, B4, B5, and B6 respectively.

[0139] Compare frequent item keyword set A and frequent item keyword set B with the 6 historical keyword subsets included in the historical keyword set respectively. It is determined that the historical keyword subsets having a preset similarity with frequent item keyword set A include historical keyword subset -1, historical keyword subset -2, and historical keyword subset -3. Correspondingly, the associated keyword set corresponding to frequent item keyword set A includes historical keyword subset -1, historical keyword subset -2, and historical keyword subset -3. It is determined that the historical keyword subsets having a preset similarity with frequent item keyword set B include historical keyword subset -3, historical keyword subset -4, and historical keyword subset -5. Correspondingly, the associated keyword set corresponding to frequent item keyword set B includes historical keyword subset -3, historical keyword subset -4, and historical keyword subset -5. Since the keywords included in historical keyword subset -6 have a low association degree with the initial keywords included in frequent item keyword set A and frequent item keyword set B, historical keyword subset -6 is not the associated keyword set corresponding to the target detection object.

[0140] After respectively determining the key keyword sets corresponding to frequent item keyword set A and frequent item keyword set B, the associated object and the security level of the associated object can be determined according to the associated keyword set. Among them, historical keyword subset -1 corresponds to historical detection object -1, and historical detection object -1 has a high security level; historical keyword subset -2 corresponds to historical detection object -2, and historical detection object -2 has a high security level; historical keyword subset -3 corresponds to historical detection object -3, and historical detection object -3 has a low security level; historical keyword subset -4 corresponds to historical detection object -4, and historical detection object -4 has a low security level; historical keyword subset -5 corresponds to historical detection object -5, and historical detection object -5 has a high security level.

[0141] It should be understood that Figure 5 The content shown is only for illustrative purposes of the solution of the embodiments of the present application, and it does not have a restrictive effect on the actual association situation of the embodiments of the present application.

[0142] 106. Based on the security level, screen at least one target keyword from the frequent item keyword set.

[0143] For each frequent item keyword set, after determining its corresponding associated keyword set, the associated object corresponding to the associated keyword set, and the security level of the associated object, it is possible to determine whether the keywords included in the frequent item keyword set can ultimately be used as target keywords based on the security level of the associated object.

[0144] Specifically, based on the security level, screening at least one target keyword from the frequent item keyword set may include: comparing the security level with a preset security level, and determining, based on the comparison result, the secure associated objects with a security level higher than the preset security level and the number of secure objects of the secure associated objects from the associated objects; determining the total number of associated objects, and determining the target keyword from the frequent item keyword set based on the ratio between the number of secure objects and the total number of objects.

[0145] In the embodiments of the present application, there can be various ways to classify the security level. For example, as Figure 5 shown, the security level can be divided into three levels: high, medium, and low. Of course, other ways can also be used to divide the security level. For example, the security level can be divided from high to low into level 1, level 2, level 3, level 4, etc., which will not be enumerated one by one here. For the sake of consistency, the classification method of the preset security level can refer to the security level. Taking the preset security level as medium as an example for illustration, for the frequent item keyword set A, since the associated keyword set corresponding to the frequent item keyword set A includes historical keyword subset -1, historical keyword subset -2, and historical keyword subset -3, therefore, the total number of associated objects corresponding to the frequent item keyword set A is three, among which, there are two secure associated objects with a security level higher than the preset security level, namely: historical detection object -1 corresponding to historical keyword subset -1, historical detection object -2 corresponding to historical keyword subset -1. At this time, it is possible to determine whether the keywords included in the frequent item keyword set A can be used as target keywords through the ratio between the number of secure objects and the total number of objects.

[0146] Similarly, for the frequent item keyword set B, since the associated keyword set corresponding to the frequent item keyword set B includes historical keyword subset -3, historical keyword subset -4, and historical keyword subset -5, therefore, the total number of associated objects corresponding to the frequent item keyword set B is three, among which, there is one secure associated object with a security level higher than the preset security level, which is: historical detection object -5 corresponding to historical keyword subset -5. At this time, it is possible to determine whether the keywords included in the frequent item keyword set B can be used as target keywords through the ratio between the number of secure objects and the total number of objects.

[0147] Among them, determining target keywords from the frequent item keyword set based on the ratio between the number of security objects and the total number of objects may include: if the ratio between the number of security objects and the total number of objects is greater than or equal to a preset threshold, taking the keywords included in the frequent item keyword set as target keywords; or, if the ratio between the number of security objects and the total number of objects is less than the preset threshold, taking the keywords included in the frequent item keyword set as invalid keywords. In the embodiments of the present application, the ratio between the number of security objects and the total number of objects can be compared with the preset threshold, and based on the comparison result, it can be determined whether the keywords included in the frequent item keyword set can be used as target keywords. In the embodiments of the present application, the specific value of the preset threshold can be set or adjusted according to the actual situation. For example, the preset threshold can be set to any percentage between 0-100%. Specifically, for example, the preset similarity can be 50%, 60%, 70%... For a certain frequent item keyword set, if the ratio between the number of security objects corresponding to the frequent item keyword set and the total number of associated objects corresponding to the frequent item keyword set is not less than the preset threshold, all the keywords included in the frequent item keyword set can be used as target keywords. On the contrary, if the ratio between the number of security objects corresponding to the frequent item keyword set and the total number of associated objects corresponding to the frequent item keyword set is less than the preset threshold, the keywords included in the frequent item keyword set cannot be used as target keywords. At this time, the keywords included in the frequent item keyword set can be used as invalid keywords, and then, the invalid keywords can be blocked or deleted.

[0148] Taking the preset threshold of 50% as an example for illustration, as Figure 5 shown, for the frequent item keyword set A, as described above, since the total number of associated objects corresponding to the frequent item keyword set A is three, and the number of security associated objects corresponding to the frequent item keyword set A is two, the ratio between the number of security associated objects and the total number of objects is 2 / 3, approximately 66.7%. This ratio is greater than the preset threshold. At this time, all the keywords in the frequent item keyword set A can be used as target keywords.

[0149] For the frequent item keyword set B, as described above, since the total number of associated objects corresponding to the frequent item keyword set B is three, and the number of security associated objects corresponding to the frequent item keyword set B is one, the ratio between the security associated object and the total number of objects is 1 / 3, approximately 33.3%. This ratio is less than the preset threshold. At this time, all the keywords in the frequent item keyword set B can be used as invalid keywords and cannot be included in the scope of the final target keywords.

[0150] In the embodiments of the present application, the keywords corresponding to the frequent item keyword set are further filtered in combination with the account quality level of the associated object, and the keywords in the frequent item keyword set with a relatively high proportion of high-quality level users in the associated object are retained, while the keywords in the frequent item keyword set with a relatively high proportion of low-quality level users in the associated object are filtered out. The obtained target keywords can more comprehensively and objectively reflect the relevant information of the target detection object, thereby improving the accuracy of keyword extraction.

[0151] Among them, after screening at least one target keyword from the frequent item keyword set based on the security level, it may further include: determining the risk level of the target detection object based on the target keyword; restricting the interaction permissions of the target detection object according to the risk level. Since the target keyword can reflect the relevant information of the target detection object, after determining the target keyword, the risk level of the target detection object can be evaluated based on the target keyword. The risk level of the target detection object is opposite to the security level, and the higher the risk level, the lower the security level. For example, if the target keyword contains keywords related to fraud, the risk level of the target detection object can be rated as a high risk level. It should be understood that there are various ways to determine the risk level of the target detection object using the target keyword in practical applications, and they are not listed one by one here.

[0152] After determining the risk level of the target detection object, the interaction permissions of the target detection object can be restricted according to the risk level. For example, if the risk level of the target detection object is a high risk level, the account of the target detection object can be directly frozen, and it is prohibited from sending information to other users (including the reporting user). If the risk level of the target detection object is a medium risk level, a warning message can be sent to the target detection object, and it is prohibited from sending information to the reporting user. If the risk level of the target detection object is a low risk level, a reminder message can be sent to the target detection object.

[0153] It should be understood that the above-mentioned ways of restricting the interaction permissions of the target detection object are only examples of the ways of restricting the interaction permissions that can actually be adopted, and do not limit the ways of restricting the interaction permissions that can actually be adopted.

[0154] As can be seen from the above, in the embodiment of the present application, after obtaining the target detection text of the target detection object and extracting at least one original detection statement from the target detection text, data augmentation is performed on the original detection statement to obtain a data-augmented statement; then, at least one keyword is extracted from the target detection statement set to obtain an initial keyword set, and the target detection statement set includes the original detection statement and the data-augmented statement; then, the initial keywords in the initial keyword set are combined according to the co-occurrence relationship to obtain at least one frequent item keyword set; then, an associated keyword set with a preset similarity to the frequent item keyword set is obtained, and the associated object and the security level of the associated object are determined according to the associated keyword set; then, based on the security level, at least one target keyword is screened out from the frequent item keyword set. Since this solution can perform data augmentation on the existing original detection statements to increase the diversity of data samples, hidden keywords and keyword information that may be used by improper industries in the future can be inferred based on the diverse samples; then, the co-occurrence relationship between the extracted keywords is identified, and screening and combination are performed according to the frequent patterns between the keywords; further filtering is performed on the combined keywords in combination with the account quality level of the associated object to obtain keywords that can comprehensively and objectively reflect the target detection object. Compared with manual keyword analysis, this solution can mine the key information hidden in the text content, automatically filter out irrelevant content, thereby improving the comprehensiveness and objectivity of the analysis results; this solution can process large-scale text content and improve the analysis efficiency; moreover, this solution can provide consistent analysis results, thereby avoiding the subjectivity problem in manual analysis, and has the advantages of automatic filtering, wide coverage, and low time consumption.

[0155] According to the method described in the above embodiment, the following will give a further detailed description by way of examples.

[0156] In this embodiment, it will be described by taking the keyword extraction device specifically integrated in an electronic device, the electronic device being a server, and the server being a server as an example.

[0157] Figure 6 Another flowchart of the keyword extraction method provided by the embodiment of the present application is shown. As Figure 6 shown, a keyword extraction method has the following specific process:

[0158] 201. The server obtains the target detection text of the target detection object and extracts at least one original detection statement from the target detection text.

[0159] For example, the server can extract at least one target character from the target detection text, where the target character can include at least one of text or numbers, and combine the target characters to obtain at least one original detection statement.

[0160] 202. The server performs word segmentation on the original detection statement to obtain at least one original word segmentation fragment.

[0161] For example, the server can use common Chinese word segmentation tools (such as Jieba word segmentation tool, SnowNLP, THULAC, NLPIR-ICTCLAS Chinese word segmentation system, etc.) to perform word segmentation on the original detection statement, thereby obtaining at least one original word segmentation fragment. Among them, the types of original word segmentation fragments can include at least one of phrases, words, single digits, and single characters.

[0162] 203. The server performs multi-dimensional transformation on the original word segmentation fragment in the original detection statement to obtain a data augmentation statement corresponding to the original detection statement.

[0163] For example, the server can extract features from the original word segmentation fragment to obtain at least one fragment feature; based on the fragment feature, determine at least one replacement content corresponding to the original word segmentation fragment; replace the original word segmentation fragment in the original detection statement with the replacement content to obtain a data augmentation statement corresponding to the original detection statement.

[0164] For example, the server can also find a literal word segmentation fragment containing literal content from the original word segmentation fragments in the original detection statement; find an adjacent word segmentation fragment having an adjacent relationship with the literal word segmentation fragment from the original word segmentation fragments, and the adjacent word segmentation fragment includes literal content; exchange the positions of the literal word segmentation fragment and the adjacent word segmentation fragment to obtain a reversed statement corresponding to the original detection statement, and use the reversed statement as a data augmentation statement corresponding to the original detection statement.

[0165] 204. The server extracts at least one keyword from the target detection statement set to obtain an initial keyword set, and the target detection statement set includes the original detection statement and the data augmentation statement.

[0166] For example, the server can obtain a historical detection statement set of a historical detection object; perform word segmentation on the target detection statements in the target detection statement set to obtain at least one target word segmentation fragment; based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set and the target detection statement set, determine the target correlation degree of each target word segmentation fragment with the target detection statement set; based on the target correlation degree, screen out at least one word segmentation fragment from the target word segmentation fragments to obtain an initial keyword set.

[0167] 205. The server combines the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set.

[0168] For example, the server can identify the word positions of each initial keyword in the initial keyword set in the target detection statement, and based on the word positions, partition the initial keywords to obtain at least one initial keyword subset; construct a frequent pattern tree according to the initial keyword subset, where the frequent pattern tree represents the co-occurrence relationship between the initial keywords; screen out the keywords that meet the preset co-occurrence frequency from the initial keyword set to obtain at least one frequent item keyword set.

[0169] 206. The server obtains an associated keyword set that has a preset similarity with the frequent item keyword set, and determines the associated object and the security level of the associated object according to the associated keyword set.

[0170] For example, the server can obtain the historical keyword set of the historical detection object, where the historical keyword set includes at least one historical keyword subset; screen out the historical keyword subset that has a preset similarity with the frequent item keyword set from the historical keyword set, and use the screened historical keyword subset as the associated keyword set; screen out the historical detection object corresponding to the associated keyword set in the historical detection objects to obtain the associated object; determine the security level of the associated object based on the object identifier of the associated object.

[0171] 207. The server screens out at least one target keyword from the frequent item keyword set based on the security level.

[0172] For example, the server can compare the security level with the preset security level, and based on the comparison result, determine the secure associated objects with a security level higher than the preset security level and the number of secure objects of the secure associated objects among the associated objects; determine the total number of associated objects, and based on the ratio between the number of secure objects and the total number of objects, determine the target keyword from the frequent item keyword set.

[0173] For example, if the ratio between the number of secure objects and the total number of objects is greater than or equal to the preset threshold, the server can use the keywords included in the frequent item keyword set as the target keywords.

[0174] For example, if the ratio between the number of secure objects and the total number of objects is less than the preset threshold, the server can use the keywords included in the frequent item keyword set as invalid keywords.

[0175] 208. The server determines the risk level of the target detection object based on the target keyword, and restricts the interaction permissions of the target detection object according to the risk level.

[0176] For example, if the risk level of the target detection object is a high-risk level, the server can directly freeze the account of the target detection object and prohibit it from sending information to other users (including the reporting user).

[0177] For example, if the risk level of the target detection object is a medium risk level, the server can send a warning message to the target detection object and prohibit it from sending information to the reporting user.

[0178] For example, if the risk level of the target detection object is a low risk level, the server may send a reminder message to the target detection object.

[0179] As can be seen from the above, after obtaining the target detection text of the target detection object and extracting at least one original detection sentence from the target detection text, the embodiment of the present application performs data enhancement on the original detection sentence to obtain a data enhancement sentence; then, extract at least one keyword from the target detection sentence set to obtain an initial keyword set, and the target detection sentence set includes the original detection sentence and the data enhancement sentence; then, the initial keywords in the initial keyword set are combined according to the co-occurrence relationship to obtain at least one frequent keyword set; then, an associated keyword set with a preset similarity to the frequent keyword set is obtained, and the associated object and the security level of the associated object are determined according to the associated keyword set; then, based on the security level, at least one target keyword is screened out from the frequent keyword set. Since the scheme can perform data enhancement on the existing original detection sentence to increase the diversity of data samples, hidden keywords and keyword information that may be used by improper industries in the future can be inferred based on the diversity samples; then, the co-occurrence relationship between the extracted keywords is identified, and they are screened and combined according to the frequent pattern between the keywords; and the combined keywords are further filtered in combination with the account quality level of the associated object to obtain keywords that can comprehensively and objectively reflect the target detection object. Compared with manual keyword analysis, this solution can dig out key information hidden in text content and automatically filter out irrelevant content, thereby improving the comprehensiveness and objectivity of the analysis results; this solution can handle large-scale text content and improve analysis efficiency; and this solution can provide consistent analysis results, thereby avoiding the subjective problems in manual analysis, and has the advantages of automatic filtering, wide coverage and low time consumption.

[0180] In order to better implement the above method, an embodiment of the present application also provides a keyword extraction device, which can be integrated into a network device, such as a server or a terminal, and the terminal may include a tablet computer, a laptop computer and / or a personal computer.

[0181] Figure 7 FIG. 1 shows a schematic diagram of the structure of a keyword extraction device provided in an embodiment of the present application. Figure 7 As shown, the keyword extraction device may include a data acquisition unit 301, a data enhancement unit 302, a keyword extraction unit 303, a keyword combination unit 304, an association determination unit 305 and a keyword determination unit 306, as follows:

[0182] (1) Data acquisition unit 301;

[0183] The data acquisition unit 301 is configured to acquire the target detection text of the target detection object, and extract at least one original detection statement from the target detection text.

[0184] For example, the data acquisition unit 301 can specifically be configured to extract at least one target character from the target detection text, where the target character includes at least one of letters or numbers; combine the target characters to obtain at least one original detection statement.

[0185] (2) Data enhancement unit 302;

[0186] The data enhancement unit 302 is configured to perform data enhancement on the original detection statement to obtain a data-enhanced statement.

[0187] For example, the data enhancement unit 302 can specifically be configured to perform word segmentation on the original detection statement to obtain at least one original word segmentation segment; perform multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-enhanced statement corresponding to the original detection statement.

[0188] (3) Keyword extraction unit 303;

[0189] The keyword extraction unit 303 is configured to extract at least one keyword from the target detection statement set to obtain an initial keyword set.

[0190] For example, the keyword extraction unit 303 can specifically be configured to obtain the historical detection statement set of the historical detection object; perform word segmentation on the target detection statements in the target detection statement set to obtain at least one target word segmentation segment; determine the target relevance of each target word segmentation segment to the target detection statement set based on the recurrence frequency of each target word segmentation segment in the historical detection statement set and the target detection statement set; and screen out at least one word segmentation segment from the target word segmentation segments based on the target relevance to obtain an initial keyword set.

[0191] (4) Keyword combination unit 304;

[0192] The keyword combination unit 304 is configured to combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set.

[0193] For example, the keyword combination unit 304 can be specifically used to identify the word positions of each initial keyword in the initial keyword set in the target detection statement, and based on the word positions, divide the initial keywords to obtain at least one initial keyword subset; construct a frequent pattern tree according to the initial keyword subset, and the frequent pattern tree represents the co-occurrence relationship between the initial keywords; screen out the keywords that meet the preset co-occurrence frequency from the initial keyword set to obtain at least one frequent item keyword set.

[0194] (5) Association relationship determination unit 305;

[0195] The association relationship determination unit 305 is used to obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set.

[0196] For example, the association relationship determination unit 305 can be specifically used to obtain the historical keyword set of the historical detection object; screen out the historical keyword subset having a preset similarity with the frequent item keyword set from the historical keyword set, and use the screened historical keyword subset as the associated keyword set; screen out the historical detection object corresponding to the associated keyword set in the historical detection object to obtain the associated object.

[0197] (6) Keyword determination unit 306;

[0198] The keyword determination unit 306 is used to screen out at least one target keyword from the frequent item keyword set based on the security level.

[0199] For example, the keyword determination unit 306 can be specifically used to compare the security level with a preset security level, determine the secure associated object and the number of secure objects of the secure associated object with a security level greater than the preset security level from the associated object based on the comparison result; determine the total number of associated objects, and determine the target keyword from the frequent item keyword set based on the ratio between the number of secure objects and the total number of objects.

[0200] In specific implementation, the above units can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above units, reference can be made to the method embodiments described above, which will not be elaborated here.

[0201] As can be seen from the above, in the embodiment of the present application, after the data acquisition unit 301 acquires the target detection text of the target detection object and extracts at least one original detection sentence from the target detection text, the data enhancement unit 302 performs data enhancement on the original detection sentence to obtain a data enhanced sentence; thereafter, the keyword extraction unit 303 extracts at least one keyword from the target detection sentence set to obtain an initial keyword set, and the target detection sentence set includes the original detection sentence and the data enhanced sentence; then, the keyword combination unit 304 combines the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent keyword set; then, the association relationship determination unit 305 acquires an associated keyword set having a preset similarity with the frequent keyword set, and determines the associated object and the security level of the associated object according to the associated keyword set; then, the keyword determination unit 306 filters out at least one target keyword from the frequent keyword set based on the security level. Since the scheme can perform data enhancement on the existing original detection sentences to increase the diversity of data samples, it can infer the hidden keywords and the keyword information that may be used by the improper industry in the future based on the diversity samples; then, the co-occurrence relationship between the extracted keywords is identified, and the keywords are screened and combined according to the frequent patterns between the keywords; the combined keywords are further filtered in combination with the account quality level of the associated object to obtain keywords that can comprehensively and objectively reflect the target detection object. Compared with manual keyword analysis, this scheme can mine the key information hidden in the text content and automatically filter out irrelevant content, thereby improving the comprehensiveness and objectivity of the analysis results; this scheme can handle large-scale text content and improve analysis efficiency; and this scheme can provide consistent analysis results, thereby avoiding the subjective problems in manual analysis, and has the advantages of automatic filtering, wide coverage and low time consumption.

[0202] The present application also provides an electronic device, such as Figure 8 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:

[0203] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will appreciate that Figure 8 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0204] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by calling the data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0205] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0206] The electronic device further includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0207] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0208] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:

[0209] Obtain the object detection text of the target detection object, and extract at least one original detection statement from the object detection text; perform data augmentation on the original detection statement to obtain a data-augmented statement; extract at least one keyword from the object detection statement set to obtain an initial keyword set, where the object detection statement set includes the original detection statement and the data-augmented statement; combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set; based on the security level, screen out at least one target keyword from the frequent item keyword set.

[0210] For example, an electronic device can obtain the object detection text of the target detection object, and extract at least one original detection statement from the object detection text; perform word segmentation on the original detection statement to obtain at least one original word segmentation segment; perform multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain a data-augmented statement corresponding to the original detection statement; extract at least one keyword from the object detection statement set to obtain an initial keyword set, where the object detection statement set includes the original detection statement and the data-augmented statement; combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set; based on the security level, screen out at least one target keyword from the frequent item keyword set; determine the risk level of the target detection object based on the target keyword, and limit the interaction permission of the target detection object according to the risk level, and so on.

[0211] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0212] As can be seen from the above, after obtaining the target detection text of the target detection object and extracting at least one original detection sentence from the target detection text, the embodiment of the present application performs data enhancement on the original detection sentence to obtain a data enhancement sentence; then, extract at least one keyword from the target detection sentence set to obtain an initial keyword set, and the target detection sentence set includes the original detection sentence and the data enhancement sentence; then, the initial keywords in the initial keyword set are combined according to the co-occurrence relationship to obtain at least one frequent keyword set; then, an associated keyword set with a preset similarity to the frequent keyword set is obtained, and the associated object and the security level of the associated object are determined according to the associated keyword set; then, based on the security level, at least one target keyword is screened out from the frequent keyword set. Since the scheme can perform data enhancement on the existing original detection sentence to increase the diversity of data samples, hidden keywords and keyword information that may be used by improper industries in the future can be inferred based on the diversity samples; then, the co-occurrence relationship between the extracted keywords is identified, and they are screened and combined according to the frequent pattern between the keywords; and the combined keywords are further filtered in combination with the account quality level of the associated object to obtain keywords that can comprehensively and objectively reflect the target detection object. Compared with manual keyword analysis, this solution can dig out key information hidden in text content and automatically filter out irrelevant content, thereby improving the comprehensiveness and objectivity of the analysis results; this solution can handle large-scale text content and improve analysis efficiency; and this solution can provide consistent analysis results, thereby avoiding the subjective problems in manual analysis, and has the advantages of automatic filtering, wide coverage and low time consumption.

[0213] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0214] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any keyword extraction method provided in the embodiment of the present application. For example, the instructions can execute the following steps:

[0215] Obtain the object detection text of the target detection object, and extract at least one original detection statement from the object detection text; perform data augmentation on the original detection statement to obtain a data-augmented statement; extract at least one keyword from the object detection statement set to obtain an initial keyword set, where the object detection statement set includes the original detection statement and the data-augmented statement; combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine the associated object and the security level of the associated object according to the associated keyword set; based on the security level, screen out at least one target keyword from the frequent item keyword set.

[0216] For example, obtain the object detection text of the target detection object, and extract at least one original detection statement from the object detection text; perform word segmentation on the original detection statement to obtain at least one original word segmentation segment; perform multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-augmented statement corresponding to the original detection statement; extract at least one keyword from the object detection statement set to obtain an initial keyword set, where the object detection statement set includes the original detection statement and the data-augmented statement; combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine the associated object and the security level of the associated object according to the associated keyword set; based on the security level, screen out at least one target keyword from the frequent item keyword set; determine the risk level of the target detection object based on the target keyword, and limit the interaction permission of the target detection object according to the risk level, and so on.

[0217] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0218] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0219] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the keyword extraction methods provided in the embodiments of the present application, the beneficial effects achievable by any of the keyword extraction methods provided in the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments, which will not be elaborated here.

[0220] Wherein, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various alternative implementations of the above data access aspect.

[0221] The above has introduced in detail a keyword extraction method, apparatus, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A keyword extraction method, characterized in that, Including: Obtain the target detection text of the target detection object, and extract at least one original detection statement from the target detection text; Perform data augmentation on the original detection statement to obtain a data-augmented statement; Extract at least one keyword from the target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data-augmented statement; Combine the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; Obtain an associated keyword set having a preset similarity with the frequent item keyword set, and determine an associated object and the security level of the associated object according to the associated keyword set; Based on the security level, screen out at least one target keyword from the frequent item keyword set.

2. The keyword extraction method according to claim 1, wherein, The extracting at least one original detection statement from the target detection text includes: Extract at least one target character from the target detection text, where the target character includes at least one of text or numbers; Combine the target characters to obtain at least one original detection statement.

3. The keyword extraction method according to claim 1, wherein The performing data augmentation on the original detection statement to obtain a data-augmented statement includes: Perform word segmentation on the original detection statement to obtain at least one original word segmentation segment; Perform multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-augmented statement corresponding to the original detection statement.

4. The keyword extraction method according to claim 3, wherein The performing multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-augmented statement corresponding to the original detection statement includes: Extract features of the original word segmentation segment to obtain at least one segment feature; Based on the segment feature, determine at least one replacement content corresponding to the original word segmentation segment; Replace the original word segmentation segment in the original detection statement with the replacement content to obtain the data-augmented statement corresponding to the original detection statement.

5. The keyword extraction method according to claim 4, wherein The replacement content includes at least one of semantic replacement content, homophone replacement content or homograph replacement content. The determining at least one replacement content corresponding to the original word segmentation segment based on the segment feature includes: If the segment feature includes a semantic feature, replace the original word segmentation segment based on the semantic feature to obtain semantic replacement content, where the semantic replacement content includes the content after equivalent replacement or synonymous replacement of the original word segmentation segment; If the segment feature includes a syllable feature, perform homophone replacement on the original word segmentation segment based on the syllable feature to obtain homophone replacement content; If the segment feature includes an appearance feature, perform homograph replacement on the original word segmentation segment based on the appearance feature to obtain homograph replacement content.

6. The keyword extraction method according to claim 3, characterized in that The performing multi-dimensional transformation on the original word segmentation segment in the original detection statement to obtain the data-augmented statement corresponding to the original detection statement includes: Search for a word segmentation segment containing text content from the original word segmentation segments in the original detection statement; Search for an adjacent word segmentation segment having an adjacent relationship with the word segmentation segment containing text content from the original word segmentation segments, where the adjacent word segmentation segment includes text content; Swap the positions of the word segmentation fragments and the adjacent word segmentation fragments to obtain an inverted statement corresponding to the original detection statement, and use the inverted statement as the data augmentation statement corresponding to the original detection statement.

7. The keyword extraction method according to claim 1, wherein Extract at least one keyword from the target detection statement set to obtain an initial keyword set, including: Obtain the historical detection statement set of the historical detection object; Perform word segmentation on the target detection statements in the target detection statement set to obtain at least one target word segmentation fragment; Based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set and the target detection statement set, determine the target relevance of each target word segmentation fragment to the target detection statement set; Based on the target relevance, screen out at least one word segmentation fragment from the target word segmentation fragments to obtain an initial keyword set.

8. The keyword extraction method according to claim 7, wherein The step of determining the target relevance of each target word segmentation fragment to the target detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set and the target detection statement set includes: Based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set, calculate the first relevance of each target word segmentation fragment to the historical detection statement set; Based on the recurrence frequency of each target word segmentation fragment in the target detection statement set, calculate the second relevance of each target word segmentation fragment to the target detection statement set; Fuse the first relevance and the second relevance to obtain the target relevance of each target word segmentation fragment to the target detection statement set.

9. The keyword extraction method according to claim 8, characterized in that, The step of determining the first relevance of each target word segmentation fragment to the historical detection statement set based on the recurrence frequency of each target word segmentation fragment in the historical detection statement set includes: Obtain the total number of statements in the historical detection statement set; Obtain the number of target statements in the historical detection statement set that contain the current target word segmentation fragment; Calculate the ratio of the number of target statements to the total number of statements to obtain the first relevance of the current target word segmentation fragment to the historical detection statement set.

10. The keyword extraction method according to claim 8, wherein The step of determining the second relevance of each target word segmentation fragment to the target detection statement set based on... Obtain the total number of word segmentation fragments contained in the target detection statement set; Determine the number of target fragments of the current target word segmentation fragment contained in the target detection statement set; Based on the ratio of the number of target fragments to the total number of fragments, determine the second relevance of the current target word segmentation fragment to the target detection statement set.

11. The keyword extraction method according to claim 7, wherein The step of screening out at least one word segmentation fragment from the target word segmentation fragments based on the target relevance to obtain an initial keyword set includes: Screen out at least one word segmentation fragment with a target relevance greater than a preset relevance in the target word segmentation fragments as potential keywords; Obtain the stop word database corresponding to the target detection object, and based on the stop word database, clean the potential keywords to obtain an initial keyword set.

12. The keyword extraction method according to claim 1, wherein The step of combining the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set includes: Identify the word positions of each initial keyword in the initial keyword set in the target detection statement, and based on the word positions, partition the initial keywords to obtain at least one initial keyword subset; Construct a frequent pattern tree according to the initial keyword subset, where the frequent pattern tree represents the co-occurrence relationship between the initial keywords; Based on the co-occurrence relationship, screen out keywords that meet the preset co-occurrence frequency from the initial keyword set to obtain at least one frequent item keyword set.

13. The keyword extraction method according to claim 1, wherein The obtaining of the associated keyword set having a preset similarity with the frequent item keyword set and determining the associated object and the security level of the associated object according to the associated keyword set includes: Obtain the historical keyword set of historical detection objects, where the historical keyword set includes at least one historical keyword subset; Screen out the historical keyword subset having a preset similarity with the frequent item keyword set from the historical keyword set, and use the screened historical keyword subset as the associated keyword set; Screen out the historical detection objects corresponding to the associated keyword set from the historical detection objects to obtain the associated object; Determine the security level of the associated object based on the object identifier of the associated object.

14. The keyword extraction method according to claim 1, wherein The screening out of at least one target keyword from the frequent item keyword set based on the security level includes: Compare the security level with a preset security level, and based on the comparison result, determine the secure associated object with a security level higher than the preset security level and the number of secure objects of the secure associated object from the associated objects; Determine the total number of objects of the associated object, and based on the ratio between the number of secure objects and the total number of objects, determine the target keyword from the frequent item keyword set.

15. The keyword extraction method according to claim 14, wherein The determining of the target keyword from the frequent item keyword set based on the ratio between the number of secure objects and the total number of objects includes: If the ratio between the number of secure objects and the total number of objects is greater than or equal to a preset threshold, use the keywords included in the frequent item keyword set as the target keyword; or If the ratio between the number of secure objects and the total number of objects is less than the preset threshold, use the keywords included in the frequent item keyword set as invalid keywords.

16. The keyword extraction method according to claim 1, wherein After screening out at least one target keyword from the frequent item keyword set based on the security level, it further includes: Determine the risk level of the target detection object based on the target keyword; Restrict the interaction permissions of the target detection object according to the risk level.

17. A keyword extraction device, characterized in that It includes: A data acquisition unit that acquires the target detection text of the target detection object and extracts at least one original detection statement from the target detection text; A data enhancement unit that enhances the original detection statement to obtain a data enhanced statement; A keyword extraction unit that extracts at least one keyword from the target detection statement set to obtain an initial keyword set, where the target detection statement set includes the original detection statement and the data enhanced statement; The keyword combination unit combines the initial keywords in the initial keyword set according to the co-occurrence relationship to obtain at least one frequent item keyword set; The association relationship determination unit obtains an associated keyword set having a preset similarity with the frequent item keyword set, and determines an associated object and the security level of the associated object according to the associated keyword set; The keyword determination unit filters out at least one target keyword from the frequent item keyword set based on the security level.

18. An electronic device, characterized in that, It includes a processor and a memory. The memory stores an application program, and the processor is used to run the application program in the memory to execute the steps in the keyword extraction method according to any one of claims 1 to 16.

19. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps in the keyword extraction method according to any one of claims 1 to 16 are implemented.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the steps in the keyword extraction method according to any one of claims 1 to 16.