Text keyword processing method and device, electronic equipment and storage medium

Through TF-IDF technology and negative feedback feature matching, keywords are automatically extracted and combined, which solves the problems of lag and low efficiency of text keyword extraction in the prior art, and achieves efficient and low-cost keyword processing.

CN120296143APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410033404.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the text keyword automation extraction method has problems of lag and low efficiency, and the manual review cost is high, making it difficult to adapt to the expansion and complexity of business scenarios.

Method used

The candidate keywords in the sample text are extracted through TF-IDF technology, combined into candidate combination keywords, and compared with multiple sample texts, the target combination keywords are determined based on the negative feedback characteristics, and added to the keyword lexicon for matching of the text to be detected.

Benefits of technology

It improves the efficiency and coverage of keyword extraction, reduces labor costs, enhances the timeliness and accuracy of keyword processing, and avoids the limitations of business experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296143A_ABST
    Figure CN120296143A_ABST
Patent Text Reader

Abstract

The invention provides a text keyword processing method and device, electronic equipment, a computer program product and a computer readable storage medium. The method comprises the steps of obtaining a text set; extracting a plurality of candidate keywords from the plurality of sample texts; combining a plurality of candidate keywords extracted from the same sample text to obtain a plurality of candidate combined keywords; each candidate combined keyword is compared with the multiple sample texts, negative feedback features of each candidate combined keyword are obtained, and the negative feedback features are used for representing the matching degree of the candidate combined keywords and the multiple sample texts; based on the negative feedback feature of each candidate combined keyword, determining a target combined keyword from the plurality of candidate combined keywords; and adding the target combined keyword into a keyword library. According to the method and the device, the keyword extraction efficiency and the coverage rate can be improved, and meanwhile, the keyword processing timeliness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for processing keywords of text. Background Art

[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. The processing of keywords in text involves important technologies such as text processing and semantic understanding in natural language processing technology. In the related art, the method of extracting keyword combinations by text tokenization and then sending them for manual review has two problems. First, due to the limitation of business experience, the extracted keywords are lagging. Second, with the continuous expansion and complexity of business scenarios, the efficiency of extracting keywords is low and the cost is high. Summary of the Invention

[0003] Embodiments of the present application provide a method, apparatus, electronic device, computer program product, and computer-readable storage medium for processing keywords of text, which can improve the efficiency and coverage rate of keyword extraction, and at the same time enhance the timeliness of keyword processing.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] Embodiments of the present application provide a method for processing keywords of text, the method comprising:

[0006] Obtain a text set, where the text set includes a plurality of sample texts, and the plurality of sample texts have a negative feedback intention;

[0007] Extract a plurality of candidate keywords from the plurality of sample texts;

[0008] Combine the plurality of candidate keywords extracted from the same sample text to obtain a plurality of candidate combined keywords;

[0009] Compare each candidate combined keyword with the plurality of sample texts respectively to obtain a negative feedback feature of each candidate combined keyword, where the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the plurality of sample texts;

[0010] Determine a target combined keyword from the plurality of candidate combined keywords based on the negative feedback feature of each candidate combined keyword;

[0011] Add the target combined keyword to a keyword library, where the keyword library is used to compare with a text to be detected to determine whether the text to be detected has the negative feedback intention.

[0012] An embodiment of the present application provides a keyword processing device for text, and the device includes:

[0013] An acquisition module, configured to acquire a text set, where the text set includes multiple sample texts, and the multiple sample texts have a negative feedback intention;

[0014] A processing module, configured to extract multiple candidate keywords from the multiple sample texts; combine the multiple candidate keywords extracted from the same sample text to obtain multiple candidate combined keywords;

[0015] A comparison module, configured to compare each candidate combined keyword with the multiple sample texts respectively to obtain a negative feedback feature of each candidate combined keyword, where the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the multiple sample texts;

[0016] A determination module, configured to determine a target combined keyword from the multiple candidate combined keywords based on the negative feedback feature of each candidate combined keyword; add the target combined keyword to a keyword library, where the keyword library is used to compare with a text to be detected to determine whether the text to be detected has the negative feedback intention.

[0017] An embodiment of the present application provides an electronic device, and the electronic device includes:

[0018] A memory, configured to store computer-executable instructions;

[0019] A processor, configured to implement the keyword processing method for text provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are used to implement the keyword processing method for text provided by the embodiment of the present application when being executed by a processor.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where the computer program or computer-executable instructions implement the keyword processing method for text provided by the embodiment of the present application when being executed by a processor.

[0022] The embodiment of the present application has the following beneficial effects:

[0023] By extracting the keywords in the sample texts with negative feedback intention, the embodiments of the present application effectively reduce the labor cost and improve the efficiency of keyword extraction compared with the method of manually reviewing keywords in the related art; combining the extracted keywords to obtain multiple groups of candidate combined keywords realizes the automatic extraction and combination of keywords and improves the coverage rate of keyword extraction; matching the candidate combined keywords with multiple sample texts to obtain the negative feedback features of the candidate combined keywords, determining the target combined keywords based on the negative feedback features of the candidate keyword combinations, and supplementing the target combined keywords as new keywords to the keyword library, avoiding the lag in keyword extraction caused by the limitations of business experience in the related art, enhancing the timeliness of keyword processing and the accuracy of using keywords for text detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 FIG. 6 is a schematic structural diagram of a keyword processing system 100 for texts provided by an embodiment of the present application;

[0025] Figure 2 FIG. 7 is a schematic structural diagram of a server 200 provided by an embodiment of the present application;

[0026] Figure 3A FIG. 8 is a schematic flowchart of a keyword processing method for texts provided by an embodiment of the present application;

[0027] Figure 3B FIG. 9 is a schematic flowchart of extracting candidate keywords provided by an embodiment of the present application;

[0028] Figure 3C FIG. 10 is a schematic flowchart of combining candidate keywords provided by an embodiment of the present application;

[0029] Figure 3D FIG. 11 is a schematic flowchart of determining the first negative feedback feature provided by an embodiment of the present application;

[0030] Figure 3E FIG. 12 is a schematic flowchart of determining the second negative feedback feature provided by an embodiment of the present application;

[0031] Figure 3F FIG. 13 is a schematic flowchart of determining the target combined keywords provided by an embodiment of the present application;

[0032] Figure 3G FIG. 14 is a schematic flowchart of the training process of a classification model provided by an embodiment of the present application;

[0033] Figure 3H FIG. 15 is a schematic flowchart of filtering candidate combined keywords provided by an embodiment of the present application;

[0034] Figure 3I FIG. 16 is a schematic flowchart of the online process of the target combined keywords provided by an embodiment of the present application;

[0035] Figure 3J is a schematic flow chart for detecting a text to be detected provided by an embodiment of the present application;

[0036] Figure 3K is a schematic flow chart for manual submission for review provided by an embodiment of the present application;

[0037] Figure 4 is a schematic diagram of the principle of a classification model provided by an embodiment of the present application;

[0038] Figure 5 is a schematic diagram of a filtering threshold interval and an effective threshold interval provided by an embodiment of the present application;

[0039] Figure 6 is a schematic diagram of the principle of gray-scale online and official online provided by an embodiment of the present application;

[0040] Figure 7 is a schematic flow chart of a method for processing keywords of negative feedback text provided by an embodiment of the present application. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0042] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0043] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0044] In the embodiments of the present application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.

[0045] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0046] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0047] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0048] 1) Suppression operation: An operation to suppress the publishing behavior of an abnormal account. For example, logging off the account and prohibiting login.

[0049] 2) Normal account: An account that publishes positive emotions (happy) or neutral emotions (bored, apathetic, and calm) and does not publish negative emotions (sad, angry). For example, an account with a long active time and regular social activities (such as interacting with friends via messages, posting and commenting on Moments, reading official account articles, etc.), and does not publish text with negative emotions.

[0050] 3) Abnormal account: An account that publishes text carrying negative emotions. For example, publishing text with negative emotions in an instant messaging client.

[0051] 4) Negative feedback text: Text with negative emotions. For example, the text with negative emotions published by an abnormal account selected by a user when reporting an abnormal account in an instant messaging client.

[0052] 5) Term Frequency: The frequency of a certain word appearing in an article.

[0053] 6) Term Frequency–Inverse Document Frequency (TF-IDF) is a statistical method used to calculate the importance of words in a document, evaluating the importance of a word for a document set or a single document in a corpus. The importance of a word increases proportionally with the number of times it appears in a document, but decreases inversely with its frequency in the corpus. The main idea of TF-IDF is that if a certain word has a high term frequency (TF) in an article and rarely appears in other articles, then this word or phrase is considered to have good category discrimination ability and is suitable for classification.

[0054] 7) Gray-scale launch: In contrast to the official launch, the gray-scale launch is a test launch phase before the official launch of a software product or function. The official launch is the phase when the software product or function is officially released for users to use. In the gray-scale launch, according to the priority of product requirements, the core requirements are selected, and it is launched quickly while meeting the basic requirements of users. Software products or functions are tested and tried out through mechanisms such as traffic restriction and white list to collect users' opinions, thereby obtaining potential user requirements and forming more targeted optimization plans for the follow-up.

[0055] 8) Negative feedback intention: That is, the intention to express opposition. For example, when instant messaging client users conduct network interactions, they express opposition to the text with negative emotions posted by abnormal accounts.

[0056] 9) Sample text: Text with negative feedback intention. For example, sample texts expressing opposition to various types of improper behaviors. Various types of deceptive behaviors can include malicious brushing of orders, behaviors that undermine public safety (such as running red lights), and uncivilized behaviors (such as uncivil remarks).

[0057] 10) Valid sample text: Text with negative feedback intention confirmed manually.

[0058] The current method for automatically extracting text keywords in related technologies mainly uses text segmentation methods (such as Chinese Jieba segmentation / TF-IDF and other algorithms) to extract keyword combinations and then send them for manual review. The reviewers first determine whether the keywords meet the specific scenario requirements based on business characteristics, and secondly, judge the accuracy of the keywords according to existing data and experience, so as to classify the new word combinations into different levels. Reviewers often need to undergo a large amount of training to obtain sufficient business experience, otherwise they cannot judge whether the keyword combinations meet the existing technical solutions and have sufficient accuracy.

[0059] The advantage of this solution is that the reviewers can ensure that the new words meet the business requirements and have a high enough accuracy rate, so as to continuously and stably empower downstream business scenarios. However, this solution also has disadvantages. First, due to the limitations of business experience, reviewers can only retain the keyword combinations they are familiar with and cannot truly discover new words, and the results are often lagging. Second, as business scenarios continue to expand and become more complex, the volume of automatically extracted keywords has increased rapidly, requiring more reviewers, resulting in high labor costs and gradually decreasing benefits to the business.

[0060] Based on the above analysis, the applicant found that in the related art, the method of extracting keyword combinations by text segmentation and then sending them for manual review has the problems of lagging extracted keywords, low efficiency, and high cost. To address the above problems, the embodiments of the present application provide a method for processing keywords of text, which can improve the efficiency and coverage of keyword extraction, and enhance the timeliness of keyword processing.

[0061] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for processing keywords of text, which can improve the efficiency and coverage of keyword extraction, and enhance the timeliness of keyword processing. The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smart phones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can be implemented as a server.

[0062] The embodiments of the present application can be implemented through artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0063] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0064] The keyword processing method of the text provided in the embodiment of the present application is implemented by natural language processing technology, which is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing involves natural language, that is, the language used by people in daily life, which is closely related to linguistic research; it also involves computer science and mathematics, an important technology for model training in the field of artificial intelligence, and pre-training models, which are developed from large language models (Large Language Model) in the field of NLP. After fine-tuning, large language models can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies.

[0065] See also Figure 1 , Figure 1 1 is a schematic diagram of the architecture of a text keyword processing system 100 provided in an embodiment of the present application, for implementing a text keyword processing application, for example, Figure 1 The keyword processing system 100 of the text involves a server 200, a network 300 and a terminal 400. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0066] In some embodiments, terminal 400 is used to run an instant messaging client, obtain sample text with negative feedback intention from a local environment; identify the sample text with negative feedback intention, extract and combine keywords of the sample text; and build a keyword thesaurus based on the negative feedback features of the combined keywords; and detect the text to be detected using the keywords of the keyword thesaurus.

[0067] In some embodiments, the terminal 400 is used to run an instant messaging client. The instant messaging client submits a sample text with a negative feedback intention to the server 200 through the network 300 and displays it on the graphical interface 410 (exemplarily shows the graphical interface 410-1); the server 200 is used to identify the sample text with a negative feedback intention, extract and combine the keywords of the sample text, and construct a keyword library according to the negative feedback characteristics of the combined keywords. The instant messaging client submits a text to be detected for the account feedback of the instant messaging client to the server 200; the server 200 is used to receive the text to be detected, detect the text to be detected, and process the account according to the detection result.

[0068] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through a wired or wireless communication method, which is not limited in the embodiments of the present application.

[0069] See Figure 2 , Figure 2 is a schematic structural diagram of the server 200 provided by the embodiments of the present application. Figure 2 The server 200 shown includes at least one processor 210, a memory 250, and at least one network interface 220. Each component in the terminal 400 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 2 all kinds of buses are labeled as the bus system 240.

[0070] The processor 210 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0071] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 250 optionally includes one or more storage devices that are physically remote from the processor 210.

[0072] The memory 250 includes volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0073] In some embodiments, the memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0074] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0075] The network communication module 252 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.;

[0076] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is a keyword processing device 253 for the text stored in the memory 250, which can be software in the form of programs and plugins, etc., including the following software modules: an acquisition module 2531, a processing module 2532, a comparison module 2533, and a determination module 2534. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.

[0077] In some embodiments, a terminal or a server may implement the keyword processing method for text provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions may be microprogram-level commands, machine instructions, or software instructions. The computer program may be a native program or a software module in an operating system; it may be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP; or it may be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded into a browser environment to run. In short, the above computer-executable instructions may be instructions in any form, and the above computer programs may be application programs, modules, or plug-ins in any form.

[0078] The exemplary applications and implementations of the electronic device provided in the embodiments of the present application will be combined to illustrate the keyword processing method for text provided in the embodiments of the present application.

[0079] Next, the keyword processing method for text provided in the embodiments of the present application will be described. As mentioned above, the electronic device implementing the keyword processing method for text in the embodiments of the present application may be a terminal or a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.

[0080] See Figure 3A , Figure 3A is a schematic flowchart of the keyword processing method for text provided in the embodiments of the present application. Taking the server as the main body, it will be described in combination with Figure 3A the steps 101 to 106 shown.

[0081] In step 101, a text set is obtained, where the text set includes multiple sample texts, and the multiple sample texts have a negative feedback intention.

[0082] In some embodiments, the obtained text set may be a set composed of sample texts submitted by a user received by the terminal through an instant messaging client, where the sample texts have a negative feedback intention.

[0083] For example, during the network interaction of a user of the instant messaging client, when the content published by user A has a negative impact on other users, other users can report by selecting some or all of the content. The instant messaging client can receive in real time the text information corresponding to the content submitted by other users as the sample text corresponding to user A, and the sample text is marked as having a negative feedback intention by at least one user.

[0084] In step 102, multiple candidate keywords are extracted from the multiple sample texts.

[0085] In some embodiments, candidate keywords are extracted from the sample texts of the text collection by using the TF-IDF technique. Refer to Figure 3B , Figure 3B which is a schematic flowchart of extracting candidate keywords provided by an embodiment of the present application. Figure 3A Step 102 of Figure 3B can be implemented by steps 1021 to 1027 of

[0086] and the following is a specific description.

[0087] In step 1021, the keywords of each sample text and the total number of keywords of each sample text are obtained.

[0088] For each keyword obtained from each sample text in step 1021, the following steps 1022 to 1027 are executed.

[0089] In step 1022, the number of times the keyword appears in the sample text is obtained.

[0090] In some embodiments, the number of times the keyword appears in the sample text can be counted by a counting class. For example, first split the sample text into a keyword list, then count the number of times each keyword appears in the list, that is, the word frequency, and finally output the number of times the keyword appears.

[0091] In step 1023, the ratio of the number of times the keyword appears in the sample text to the total number of sample texts is determined as the first parameter value.

[0092] In some embodiments, the ratio of the number of times the keyword appears in the sample text to the total number of sample texts is the word frequency.

[0093] For example, the keywords of sample text 1 include keyword 1, keyword 2, and keyword 3. Among them, keyword 1 appears 10 times in sample text 1, and the total number of times of keyword 1, keyword 2, and keyword 3 in sample text 1 is 50, then the word frequency of keyword 1 is 0.2.

[0094] In step 1024, the total number of sample texts in the text collection and the number of sample texts containing the keyword are obtained.

[0095] In some embodiments, the total number of sample texts in the text collection and the number of sample texts containing the keyword can be counted by a fuzzy statistical function. For example, the total number of sample texts in the text collection and the number of sample texts containing the keyword can be counted by the countif function.

[0096] In step 1025, determine the ratio of the total number of sample texts in the text set to the number of sample texts containing the keyword, and use it as the second parameter value.

[0097] For example, if the total number of sample texts in the text set is 100 and the number of sample texts containing keyword 1 in the above example is 25, then the second parameter value corresponding to keyword 1 is 4.

[0098] In step 1026, fuse the first parameter value and the second parameter value into the third parameter value of the keyword.

[0099] In some embodiments, fusing the first parameter value and the second parameter value to obtain the third parameter value of the keyword can be achieved by performing the following processing: First, calculate the logarithm of the second parameter value to obtain the inverse document frequency of the keyword; then, use the product of the first parameter value and the inverse document frequency of the keyword as the third parameter value of the keyword.

[0100] For example, the second parameter value corresponding to keyword 1 in the above example is 4. Then, the logarithm of the second parameter value, that is, the inverse document frequency is 2. The third parameter value corresponding to keyword 1 is the product of the word frequency corresponding to keyword 1 and the inverse document frequency, which is 0.4.

[0101] In step 1027, determine candidate keywords based on the third parameter value of each keyword.

[0102] In some embodiments, based on the third parameter value of each keyword, use at least one target keyword starting from the head in the descending order result as the candidate keyword.

[0103] For example, the specific number of at least one can be preset to 20% of the total number of keywords. For example, if the total number of target keywords is 50, then use the first 20% of the target keywords, that is, the first 10 target keywords, as the candidate keywords.

[0104] Continue to refer to Figure 3A , and continue to describe step 102 above.

[0105] In step 103, combine multiple candidate keywords extracted from the same sample text to obtain multiple candidate combined keywords.

[0106] In some embodiments, refer to Figure 3C , Figure 3C is a schematic flowchart of the process of combining candidate keywords provided by an embodiment of the present application. Figure 3A Step 103 of Figure 3C can be implemented by steps 1031 to 1032 of

[0107] In step 1031, N candidate keywords are extracted from each sample text, where N is an integer constant greater than or equal to 2.

[0108] In some embodiments, let n be an integer variable and 2 ≤ n ≤ N. Traverse n to perform the processing of the following step 1032.

[0109] In step 1032, any n candidate keywords are selected from the N candidate keywords, and the n candidate keywords are concatenated to obtain a candidate combined keyword.

[0110] In some embodiments, the N candidate keywords are candidate keywords among the L candidate keywords extracted from each sample text, where L is greater than N and N is a preset value.

[0111] In some embodiments, only when L is greater than N can various situations of keyword combinations be fully explored, realizing various possibilities of combining candidate keywords. For example, when L is 7 and N is 4 or 5, the combination methods of the candidate combined keywords obtained by combining the candidate keywords are more diverse.

[0112] Exemplarily, when N is 5, the 5 candidate keywords are candidate keyword 1, candidate keyword 2, candidate keyword 3, candidate keyword 4, and candidate keyword 5 respectively; when 3 candidate keywords are selected from the 5 candidate keywords, the obtained candidate combined keyword can be {candidate keyword 1, candidate keyword 2, candidate keyword 3}.

[0113] In the embodiments of the present application, after filtering the extracted keywords for common words, different keywords are automatically combined to obtain multiple groups of candidate combined keywords, realizing automatic extraction and combination of keywords, effectively reducing the labor cost, and improving the extraction efficiency and coverage rate of keywords.

[0114] Continue to refer to Figure 3A , and continue to describe step 103 above.

[0115] In step 104, each candidate combined keyword is respectively compared with multiple sample texts to obtain a negative feedback feature of each candidate combined keyword, where the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the multiple sample texts.

[0116] In some embodiments, the negative feedback feature includes the number of sample texts hit by the candidate combined keyword and the number of valid sample texts with negative feedback intention hit by the candidate combined keyword.

[0117] In some embodiments, refer to Figure 3D , Figure 3D is a schematic flowchart of the process for determining the first negative feedback feature provided by the embodiments of the present application.Figure 3A Step 104 can be implemented by performing the processing from step 1041A to step 1042A for each candidate combined keyword. The following is a specific description.

[0118] In step 1041A, the candidate combined keywords are respectively compared with multiple sample texts to obtain a first comparison result.

[0119] In some embodiments, the candidate combined keywords are respectively matched with the keywords in multiple sample texts, and the obtained matching result is the first comparison result.

[0120] For example, if the candidate combined keyword A, candidate combined keyword B, and candidate combined keyword C are respectively matched with 100 sample texts, for sample text 1 among the 100 sample texts, if the matching result is that sample text 1 contains candidate combined keyword A and does not contain candidate combined keyword B and candidate combined keyword C, it indicates that candidate combined keyword A hits sample text 1, and candidate combined keyword B and candidate combined keyword C do not hit sample text 1.

[0121] In step 1042A, in response to the first comparison result indicating that the candidate combined keyword hits at least one sample text, the number of sample texts hit by the candidate combined keyword is used as the first negative feedback feature of the candidate combined keyword.

[0122] In some embodiments, the candidate combined keyword hits at least one sample text, that is, at least one sample text contains the candidate combined keyword.

[0123] For example, if the candidate combined keyword A, candidate combined keyword B, and candidate combined keyword C mentioned above are respectively compared with 100 sample texts, among the 100 sample texts, 40 sample texts contain candidate combined keyword A, 50 sample texts contain candidate combined keyword B, and 30 sample texts contain candidate combined keyword C; that is, candidate combined keyword A hits 40 sample texts, candidate combined keyword B hits 50 sample texts, and candidate combined keyword B hits 30 sample texts; therefore, the first negative feedback feature of candidate combined keyword A is 40, the first negative feedback feature of candidate combined keyword B is 50, and the first negative feedback feature of candidate combined keyword C is 30.

[0124] In some embodiments, the sample text carries an attribute label, where the attribute label indicates whether the sample text is a valid sample text with a negative feedback intention. For example, label 1 is used to indicate that the sample text is a valid sample text with a negative feedback intention, and label 0 is used to indicate that the sample text is a valid sample text without a negative feedback intention.

[0125] In some embodiments, the valid sample text includes sample texts of multiple deception types, which are used to characterize different types of negative feedback intentions, such as the deception type of placing multiple orders within a short period of time.

[0126] In some embodiments, the valid sample text is manually calibrated text with negative feedback intention, such as sample texts expressing objections to various types of improper behaviors. The multiple deception behaviors may include malicious brushing of orders, behaviors that undermine public safety (such as running a red light), and uncivilized behaviors (such as uncivil remarks).

[0127] See Figure 3E , Figure 3E which is a schematic flowchart of determining the second negative feedback feature provided by an embodiment of the present application. Figure 3A Step 104 of Figure 3E can be implemented by the processing from step 1041B to step 1042B of

[0128] In step 1041B, the candidate combined keywords are respectively compared with the valid sample text to obtain a second comparison result.

[0129] In some embodiments, the candidate combined keywords are respectively matched with the keywords in the valid sample text, and the obtained matching result is the second comparison result.

[0130] For example, if there are 60 valid sample texts labeled with label 1 among the 100 sample texts above, the candidate combined keyword A, candidate combined keyword B, and candidate combined keyword C are respectively matched with the 60 valid sample texts. For the sample text 2 among the 60 valid sample texts, if the matching result is that the sample text 2 contains the candidate combined keyword A and does not contain the candidate combined keyword B and candidate combined keyword C, it indicates that the candidate combined keyword A hits the sample text 2, and both the candidate combined keyword B and candidate combined keyword C do not hit the sample text 2.

[0131] In step 1042B, in response to the second comparison result indicating that the candidate combined keyword hits the valid sample text, the number of valid sample texts hit by the candidate combined keyword is used as the second negative feedback feature of the candidate combined keyword.

[0132] Exemplarily, if there are 60 valid sample texts with label 1 among the 100 sample texts above, then the candidate combined keyword A, candidate combined keyword B, and candidate combined keyword C above are compared with the 60 valid sample texts respectively. Among them, 30 sample texts contain candidate combined keyword A, 40 sample texts contain candidate combined keyword B, and 20 sample texts contain candidate combined keyword C; that is to say, candidate combined keyword A hits 30 sample texts, candidate combined keyword B hits 20 sample texts, and candidate combined keyword B hits 40 sample texts; therefore, the second negative feedback feature of candidate combined keyword A is 30, the second negative feedback feature of candidate combined keyword B is 40, and the second negative feedback feature of candidate combined keyword C is 20.

[0133] Continue to refer to Figure 3A , and continue to explain according to step 104 above.

[0134] In step 105, based on the negative feedback features of each candidate combined keyword, the target combined keyword is determined from multiple candidate combined keywords.

[0135] In some embodiments, refer to Figure 3F , Figure 3F is a schematic flowchart of the process for determining the target combined keyword provided by the embodiments of the present application. Figure 3A Step 105 of Figure 3F can be implemented by steps 1051 to 1052 of

[0136] In step 1051, according to the negative feedback features of each candidate combined keyword, the candidate combined keywords are sorted in descending order.

[0137] In some embodiments, according to the negative feedback features of each candidate combined keyword, sorting the candidate combined keywords in descending order may be to weight the first negative feedback feature and the second negative feedback feature of each candidate combined keyword, and sort the candidate combined keywords in descending order according to the weighted average of the first negative feedback feature and the second negative feedback feature.

[0138] Exemplarily, for the candidate combined keyword A in the above example, the first negative feedback feature is 40, and the second negative feedback feature is 30. The average of the two can be taken, that is, the negative feedback feature of the candidate combined keyword A is 35; for the candidate combined keyword B in the above example, the first negative feedback feature is 50, and the second negative feedback feature is 40. Taking the average of the two, that is, the negative feedback feature of the candidate combined keyword B is 45; for the candidate combined keyword C in the above example, the first negative feedback feature is 30, and the second negative feedback feature is 20. Taking the average of the two, that is, the negative feedback feature of the candidate combined keyword C is 25. Then, arranging the candidate combined keyword A, the candidate combined keyword B, and the candidate combined keyword C in descending order according to the negative feedback feature is: the candidate combined keyword B, the candidate combined keyword A, the candidate combined keyword C.

[0139] In step 1052, at least one candidate combined keyword starting from the head in the descending order result is used as the target combined keyword.

[0140] In some embodiments, the candidate combined keyword with a higher ranking in the descending order result, that is, the candidate combined keyword with a larger negative feedback feature, is used as the target combined keyword.

[0141] Exemplarily, the candidate combined keyword B and the candidate combined keyword A with a higher ranking above can be used as the target combined keywords.

[0142] Continue to refer to Figure 3A and continue to describe step 105 above.

[0143] In step 106, the target combined keyword is added to the keyword library, where the keyword library is used to compare with the text to be detected to determine whether the text to be detected has a negative feedback intention.

[0144] In the embodiment of the present application, by matching the candidate combined keyword with the sample text, the negative feedback feature of the candidate combined keyword is obtained. Based on the negative feedback feature of the candidate combined keyword, after setting a threshold, the candidate combined keyword with a higher negative feedback feature is left and used as a new keyword to be supplemented to the database, which improves the efficiency and coverage rate of keyword extraction while enhancing the timeliness of keyword processing.

[0145] In some embodiments, before executing Figure 3A step 102, at least one of the following sample texts is filtered out from the sample texts in the text set: the sample text that does not contain negative feedback evidence, and the sample text with a length less than the length threshold.

[0146] Exemplarily, when the sample text in the text set contains insufficient negative feedback evidence, or the length of the sample text is less than the length threshold, the corresponding sample text can be filtered. For example, when the length of the sample text is less than 10, it is defaulted that the sample text does not contain valuable information, and the corresponding sample text can be filtered out.

[0147] In some embodiments, the sample text that does not contain negative feedback evidence is determined by calling a pre-trained classification model. Refer to Figure 3G , Figure 3G which is a schematic diagram of the training process of the classification model provided by the embodiments of the present application. The training of the classification model can be implemented through Figure 3G steps 201 to 204, which are specifically described below.

[0148] In step 201, training sample texts are obtained.

[0149] In some embodiments, the training sample texts include sample texts containing negative feedback evidence and sample texts not containing negative feedback evidence, and are obtained by receiving the sample text data submitted by the user through an instant messaging client and performing data preprocessing.

[0150] In step 202, the training sample texts are labeled with classification labels to form a training text set, where the types of classification labels include having negative feedback evidence and not having negative feedback evidence.

[0151] In some embodiments, before classifying the sample text through the classification model, the parameters of the classification model need to be initialized. For example, the parameters of the classification model can be randomly assigned values.

[0152] In step 203, the training sample texts in the training text set are classified through the classification model to obtain the predicted labels of the training sample texts in the training text set.

[0153] In some embodiments, for whether the training sample contains negative feedback evidence, two predicted labels of having negative feedback evidence or not having negative feedback evidence are obtained.

[0154] In step 204, the loss between the classification label and the predicted label is determined, and the parameters of the classification model are updated according to the loss.

[0155] Exemplarily, refer to Figure 4 , Figure 4 which is a schematic diagram of the principle of the classification model provided by the embodiments of the present application. The classification model includes a semantic understanding model and a classifier. As Figure 4As shown, first, the obtained training sample text is input into the semantic understanding model, and the output result is the semantic vector corresponding to the sample text; second, the semantic vector is input into the classifier to obtain the classification result, which is used as the predicted label corresponding to each training sample text; then, the loss between the predicted label and the classification label is calculated through the loss function, and the loss is backpropagated to obtain the gradient of each parameter. According to the gradient of each parameter, the optimization function is used to update the parameters.

[0156] Exemplarily, the loss function can use softmax, and the optimization functions include: gradient descent, momentum optimization algorithm, adaptive learning rate algorithm, etc. When updating the parameters, only the parameters in the semantic understanding model or the parameters in the classifier can be updated, or both can be updated.

[0157] Finally, repeat the above steps until the set number of iterations is reached or the loss function converges, obtaining a classification model to realize the classification of whether the sample text contains negative feedback evidence.

[0158] In some embodiments, after performing Figure 3A step 106, a filtering operation can be performed on the candidate combined keywords. Refer to Figure 3H , Figure 3H which is a schematic flowchart of filtering candidate combined keywords provided by an embodiment of the present application and is implemented through Figure 3H steps 301 to 303, which will be specifically described below.

[0159] In step 301, the filtering threshold interval of the candidate combined keywords is obtained.

[0160] In some embodiments, the filtering threshold interval is used to perform a filtering operation on the candidate combined keywords based on the target condition, and the target condition is the filtering threshold interval.

[0161] Exemplarily, the filtering threshold interval can be set to 50% to 80%. That is to say, when the negative feedback feature ratio corresponding to the negative feedback feature of the candidate combined keyword is between 50% and 80%, the corresponding candidate combined keyword is filtered.

[0162] The processing of steps 302 and 303 is performed for each candidate combined keyword.

[0163] In step 302, the ratio of the negative feedback feature of the candidate combined keyword to the total number of sample texts in the text set is used as the negative feedback feature ratio corresponding to the negative feedback feature.

[0164] Exemplarily, in the above example, if the first negative feedback feature of candidate combined keyword A is 40 and the total number of sample texts is 100, then the negative feedback feature ratio corresponding to the first negative feedback feature of candidate combined keyword A is 0.4.

[0165] In some embodiments, the negative feedback feature includes at least one of a first negative feedback feature and a second negative feedback feature. In one case, the negative feedback feature includes the first negative feedback feature or the second negative feedback feature; in another case, the negative feedback feature includes both the first negative feedback feature and the second negative feedback feature.

[0166] In step 303, in response to the negative feedback feature ratio being within the filtering threshold range, the candidate combined keywords are filtered out from the keyword library.

[0167] In some embodiments, for the two different cases of the negative feedback feature in the above embodiments, different filtering threshold ranges can be set.

[0168] For example, when the negative feedback feature includes one of the first negative feedback feature and the second negative feedback feature, only the filtering threshold range needs to be set for the first negative feedback feature or the second negative feedback feature of the candidate combined keywords. For example, when the negative feedback feature only includes the first negative feedback feature, the filtering threshold range is set to 50% to 80%. That is to say, when the negative feedback feature ratio corresponding to the first negative feedback feature of the candidate combined keywords is between 50% and 80%, the corresponding candidate combined keywords are filtered.

[0169] For example, when the negative feedback features of the candidate combined keywords include both the first negative feedback feature and the second negative feedback feature, the filtering threshold ranges need to be set separately for the first negative feedback feature and the second negative feedback feature of the selected combined keywords. When the candidate combined keywords simultaneously meet the filtering threshold ranges corresponding to the first negative feedback feature and the second negative feedback feature respectively, the candidate combined keywords are filtered. As a special case, the filtering threshold range corresponding to the first negative feedback feature is set to 50% to 80%, and the filtering threshold range corresponding to the second negative feedback feature is set to 0%. That is to say, as long as the negative feedback feature ratio corresponding to the first negative feedback feature of the candidate combined keywords is between 50% and 80%, regardless of the value of the negative feedback feature ratio corresponding to the second negative feedback feature, the corresponding candidate combined keywords will be filtered.

[0170] In some embodiments, after performing Figure 3A step 106, the filtering operation can be performed on the candidate combined keywords. Refer to Figure 3I , Figure 3I which is a schematic flowchart of the online process of the target combined keywords provided by the embodiments of the present application, and is implemented through Figure 3I steps 401 to 406, which will be specifically described below.

[0171] In step 401, the effective threshold range of the candidate combined keywords is obtained, where the threshold in the filtering threshold range is greater than the threshold in the effective threshold range.

[0172] In some embodiments, an effective threshold range is used to determine target combined keywords.

[0173] For example, if the filtering threshold range is set to 50% to 80% and the effective threshold range is set to 20% to 50%, then for any threshold A in the filtering threshold range, it is always greater than threshold B in the effective threshold range.

[0174] In step 402, in response to the negative feedback feature ratio being within the effective threshold range, a gray-scale online operation is performed on the keyword library.

[0175] For example, referring to Figure 5 , Figure 5 is a schematic diagram of the filtering threshold range and the effective threshold range provided by an embodiment of the present application. Among them, the threshold percentage represents the ratio of the negative feedback features of the candidate combined keywords to the total number of sample texts in the text set, that is, the negative feedback feature ratio corresponding to the negative feedback features. As Figure 5 shown, the threshold percentage with the negative feedback feature ratio between 20% and 50% is used as the effective threshold range, and the threshold percentage with the negative feedback feature ratio between 50% and 80% is used as the filtering threshold range.

[0176] In some embodiments, after performing the gray-scale online operation on the keyword library, the following processing can be performed: detecting the text to be detected based on the keyword library, and performing an inhibition process on the publishing accounts of the text to be detected with negative feedback intent.

[0177] In step 403, based on the detection results of the gray-scale online, the detection accuracy indicators of multiple target combined keywords in the keyword library are determined.

[0178] In some embodiments, the detection results of the gray-scale online include the number of objections raised by the publishing accounts against the inhibition process. Based on the detection results of the gray-scale online, the detection accuracy indicators of multiple target combined keywords in the keyword library can be achieved by performing the following processing: obtaining the number corresponding to each target combined keyword, and determining the detection accuracy indicator of the target combined keyword based on the number, where the number is negatively correlated with the detection accuracy indicator.

[0179] For example, if the candidate combined keyword A in the above example is used as a target combined keyword 1 in the keyword library and is used to detect 100 texts to be detected, when the number of objections raised by the publishing accounts of the texts to be detected against the inhibition process is 10, that is, among the publishing accounts of 100 texts to be detected, 90 publishing account users consider the target combined keyword 1 to be effective, then the detection accuracy indicator of the target combined keyword 1 is 0.9. As a special case, if the number of objections raised by the publishing accounts against the inhibition process is 0, then the detection accuracy indicator is 1.

[0180] In some embodiments, the negative correlation between the number of objections raised by the publishing account against the suppression process and the detection accuracy index means that the reciprocal of the number of objections raised by the publishing account against the suppression process is positively correlated with the detection accuracy index. The negative correlation between the number and the detection accuracy index is manifested as the more objections the publishing account raises against the suppression process, the smaller the value representing the detection accuracy index. That is to say, the result of detecting the text to be detected using the corresponding target combination keywords is less accurate.

[0181] Exemplarily, the positive correlation between the reciprocal of the number of objections raised by the publishing account against the suppression process and the detection accuracy index can be a linear positive correlation. For example, if parameter x represents the reciprocal of the number of objections raised by the publishing account against the suppression process and parameter y represents the detection accuracy index, when y = ax + b, where a > 0; then, the correlation between parameter x and parameter y is a linear positive correlation. That is to say, the correlation between the detection accuracy index and the reciprocal of the number of objections raised by the publishing account against the suppression process is a linear positive correlation; the positive correlation can also be a non-linear positive correlation. For example, y = aln x + b, where a > 0; then, the correlation between parameter x and y is a non-linear positive correlation. That is to say, the correlation between the detection accuracy index and the reciprocal of the number of objections raised by the publishing account against the suppression process is a non-linear positive correlation.

[0182] In some embodiments, referring to Figure 3J , Figure 3J is a schematic flowchart of the process for detecting the text to be detected provided by an embodiment of the present application. The process of detecting the text to be detected based on the keyword library is implemented through Figure 3J steps 501 to 503 below. The following is a specific description.

[0183] In step 501, multiple keywords to be detected are extracted from the text to be detected.

[0184] In some embodiments, extracting multiple keywords to be detected from the text to be detected can be achieved by performing the following processing: First, obtain the keywords of each text to be detected and the total number of keywords of each text to be detected; at the same time, for each keyword in each text to be detected, obtain the number of times the keyword appears in the text to be detected; Second, determine the ratio of the number of times the keyword appears in the text to be detected to the total number of the text to be detected as the first parameter value; Third, obtain the total number of texts to be detected in the text set and the number of texts to be detected containing the keyword, and determine the ratio of the total number of texts to be detected in the text set to the number of texts to be detected containing the keyword as the second parameter value; Then, fuse the first parameter value and the second parameter value into the third parameter value of the keyword; Finally, based on the third parameter value of each keyword, determine the keyword to be detected.

[0185] In step 502, the keyword to be detected is compared with each keyword in the keyword library to obtain a third comparison result.

[0186] In step 503, in response to the third comparison result indicating that the number of times multiple keywords to be detected hit the keyword library exceeds the threshold number of times, it is determined that the text to be detected has a negative feedback intention.

[0187] In some embodiments, the threshold number of times is used to characterize the matching degree between the keywords of the text to be detected and the keywords in the keyword library, and can be a preset value.

[0188] Continue to refer to Figure 3I and continue to describe based on step 403 above.

[0189] In step 404, in response to the detection accuracy index of the target combined keyword being less than or equal to the detection accuracy index threshold, the target combined keyword is deleted from the keyword library.

[0190] In some embodiments, a preset detection accuracy index threshold is obtained, and in response to the detection accuracy index of the target combined keyword for gray-scale online launch being less than or equal to the detection accuracy index threshold, the target combined keyword is deleted from the keyword library.

[0191] For example, if the detection accuracy index threshold is 0.6, the detection accuracy index of target combined keyword 1 is 0.9, the detection accuracy index of target combined keyword 2 is 0.7, and the detection accuracy index of target combined keyword 3 is 0.5; then the detection accuracy index of target combined keyword 3 is less than 0.6, so target combined keyword 3 is to be deleted from the keyword library.

[0192] In step 405, in response to the detection accuracy index of the target combined keyword being greater than the detection accuracy index threshold, the target combined keyword is determined as a valid combined keyword.

[0193] In some embodiments, a preset detection accuracy index threshold is obtained, and in response to the detection accuracy index of the target combined keyword for gray-scale online launch being greater than the detection accuracy index threshold, the target combined keyword is determined as a valid combined keyword.

[0194] For example, in the above example, the detection accuracy index of target combined keyword 1 is 0.9, the detection accuracy index of target combined keyword 2 is 0.7, and the detection accuracy indexes of target combined keyword 1 and target combined keyword 2 are both greater than the detection accuracy index threshold, then target combined keyword 1 and target combined keyword 2 are determined as valid combined keywords.

[0195] In step 406, an operation of formal online launch is performed on the valid combined keywords in the keyword library.

[0196] Exemplarily, refer to Figure 6 , Figure 6 which is the schematic diagram of gray-scale online and official online of the embodiments of the present application; according to Figure 6 the time sequence shown, after the keyword library is gray-scale online, first, calculate the detection accuracy of the target combined keywords in the keyword library undergoing gray-scale online. After the accuracy detection is completed, based on the detection accuracy results, divide the target combined keywords in the keyword library undergoing gray-scale online into two parts: valid combined keywords and invalid combined keywords; then, perform the official online operation on the valid combined keywords in the keyword library until the online is completed.

[0197] In some embodiments, when the number of keywords in the keyword library is less than the number threshold, manual review can be performed. Refer to Figure 3K , Figure 3K which is the schematic diagram of the manual submission review process provided by the embodiments of the present application, and is implemented through Figure 3K steps 601 to 604 below, and the specific description is as follows.

[0198] In step 601, send multiple texts to be detected to the manual review seat.

[0199] In some embodiments, the terminal sends multiple texts to be detected submitted by the instant messaging client to the manual review seat.

[0200] In step 602, obtain the review result of the manual review seat, where the review result includes whether the text to be detected has a negative feedback intention, and the combined keyword to be detected in the text to be detected when the text to be detected has a negative feedback intention.

[0201] In some embodiments, when the text to be detected has a negative feedback intention, first extract the corresponding keyword to be detected from the text to be detected, and then combine the keywords to be detected to obtain the combined keyword to be detected.

[0202] In step 603, determine the negative feedback feature ratio corresponding to the negative feedback feature of the combined keyword to be detected.

[0203] In some embodiments, use the ratio of the negative feedback feature of each combined keyword to be detected to the total number of sample texts in the text set as the negative feedback feature ratio corresponding to the negative feedback feature.

[0204] In some embodiments, the negative feedback feature includes at least one of the first negative feedback feature and the second negative feedback feature. In one case, the negative feedback feature includes the first negative feedback feature or the second negative feedback feature; in another case, the negative feedback feature includes both the first negative feedback feature and the second negative feedback feature.

[0205] In step 604, in response to the negative feedback feature ratio corresponding to the negative feedback feature of the combination keyword to be detected being within the effective threshold range, the combination keyword to be detected will be added to the keyword library.

[0206] In the embodiment of the present application, by extracting keywords from the sample text with negative feedback intention and combining the extracted keywords, multiple groups of candidate combination keywords are obtained, realizing automatic extraction and combination of keywords, and improving the coverage rate of keyword extraction; matching the candidate combination keywords with multiple sample texts to obtain the negative feedback features of the candidate combination keywords, determining the target combination keywords based on the negative feedback features of the candidate keyword combinations, and supplementing the target combination keywords as new keywords to the keyword library, avoiding the lag of keyword extraction caused by the limitation of business experience in the related technology, enhancing the timeliness of keyword processing and the accuracy of keywords used for text detection.

[0207] Next, an exemplary application of the keyword processing method for the text proposed in the embodiment of the present application in an actual application scenario will be described.

[0208] There are billions of active users on the social network platform every day. The huge social volume and the complexity of the social network make some instant messaging APPs receive a large number of sample texts with negative feedback intention submitted by user accounts every day. In order to improve the user experience and reasonably feedback user needs, the instant messaging client needs to promptly identify the sample text, so as to accurately and efficiently process the accounts that publish the sample texts with negative feedback intention according to the detection results of the sample texts submitted by users, and promptly contain the negative impact caused by abnormal users to other normal users in the social network. Although the related technology that combines intelligence with manual review has a relatively high accuracy rate, it has a slow effect, low efficiency, and high cost.

[0209] The keyword processing method for text proposed in the embodiments of the present application. First, the negative feedback text (sample text) after data cleaning is subjected to Chinese word segmentation using the TF-IDF technology to extract the keywords in the negative feedback text. Secondly, after filtering the common words from the keywords obtained by the word segmentation result, different keywords (candidate keywords) are automatically combined. The length of the keyword combination is the number of keywords therein. For example, if there are 3 keywords in the combination, in the form of "word 1|word 2|word 3", the length is three, and multiple groups of new keyword combinations (candidate combined keywords) are obtained, which are marked as (keyword combination, length). Then, the keyword combination is matched with the historical negative feedback text to obtain two negative feedback features of the keyword combination. One is the number of negative feedback texts that the keyword combination can match, and the other is the number of effective negative feedback texts hit by the keyword combination. Finally, based on the negative feedback features of the keyword combination, after setting a threshold, the keyword combinations with higher negative feedback features are left and supplemented into the database as new keywords, solving the problems of low efficiency, lag, and low coverage rate in the existing solutions.

[0210] See Figure 7 , Figure 7 is a schematic flowchart of the keyword processing method for negative feedback text provided by the embodiments of the present application. Next, it will be specifically described in conjunction with Figure 7 the steps 701 to 707 shown.

[0211] In step 701, obtain negative feedback text.

[0212] Obtain the negative feedback text data submitted by the terminal. Among them, the negative feedback text data includes real-time negative feedback text data and historical negative feedback text data. The content of the negative feedback text data includes the text information, picture information, and user description information selected by the user when submitting. The historical negative feedback text data also includes whether at least one of the negative feedback texts therein is marked as an effective negative feedback text.

[0213] In step 702, filter the negative feedback text.

[0214] When filtering the obtained negative feedback text, on the one hand, it is necessary to filter the negative feedback text with insufficient evidence or very little description. On the other hand, the negative feedback text contains the information flow of the negative feedback text submitter and the information flow of the negative feedback text publisher. Since the information flow of the negative feedback text submitter is not the focus of identifying the negative feedback intention, only the part of the information flow of the negative feedback text publisher in the negative feedback text needs to be extracted. Therefore, the information flow of the negative feedback text submitter should be filtered out.

[0215] In step 703, extract keywords.

[0216] Extract keywords from the filtered negative feedback text using the TF-IDF algorithm. Specifically, for sentence A, perform word segmentation on sentence A using TF-IDF, and the word segmentation result is the set S = {k1, k2, … k n ,}, and usually, during the word segmentation process, filtering operations for common words will be performed.

[0217] For example, for the negative feedback text "The environment here is too bad. Let's go somewhere else to play", the word segmentation result may be S = {"here", "environment", "bad", "we", "change place", "play"}, and here, "we", as a common personal pronoun, has little information content and insufficient evidence, so it will be filtered.

[0218] In step 704, combine keywords.

[0219] For sentence A, the result of TF-IDF word segmentation is the set S = {k1, k2, … k n}, where n is the number of keywords extracted. After filtering out the keywords with insufficient evidence from the keywords in set S, the keyword combination length is L. For example, in the above example, S = {"here", "environment", "bad", "we", "change place", "play"}, then the subscript n in set S = 6; after filtering out "we", L = 5.

[0220] Randomly select N keywords from the keyword combinations with a keyword combination length of L for combination. N is an integer constant greater than or equal to 2, and the combination result is represented as R = {C1, C2, … C m}. Among them, each element in R represents a keyword combination containing the selected N keywords, and C m represents the combination containing all the keywords in the keyword combination with a combination length of L, and C m = {k i1 , k i2 , … k iL}. For example, in the above example, after filtering out "me", the remaining keywords in set S are "work", "first time", "tuition fee", "exchange with you"; randomly select N keywords from the keyword combinations with a keyword combination length of L for combination, and the elements in the combination result R can be C1 = {"here", "environment"}, C2 = {"here", "bad"}, C3 = {"here", "change place"}, C m = {"here", "environment", "bad", "we", "change place", "play"}.

[0221] In step 705, calculate the negative feedback feature.

[0222] For the extracted keyword combinations, calculate two negative feedback features: one is the number of negative feedback texts hit by the keyword combination; the other is the number of negative feedback texts that are confirmed as valid negative feedback texts hit by the keyword combination. Use these two negative feedback features to measure the accuracy rate of the keyword combination in detecting negative feedback texts, which is expressed as:

[0223]

[0224] Among them, represents the accuracy rate of the keyword combination in detecting negative feedback texts, ExposeNum represents the number of negative feedback texts hit by the keyword combination, and ValidExposeNum represents the number of negative feedback texts that are confirmed as valid negative feedback texts hit by the keyword combination.

[0225] For example, take 100 negative feedback texts in the historical negative feedback text library for matching, and 50 are hit, among which 10 are confirmed as valid complaints. Generally speaking, the more negative feedback texts hit by the keyword combination, the greater the potential negative feedback degree of the keyword combination, and the more the number of negative feedback texts that are confirmed as valid negative feedback texts hit by the keyword combination, the higher the accuracy rate of the keyword combination in identifying negative feedback texts may be.

[0226] In step 706, a gray-scale online launch is carried out.

[0227] The higher the negative feedback feature of a keyword, the higher the effective rate of the negative feedback texts it hits.

[0228] Before the official online launch, first of all, filter out the keyword combinations with very high hit times to reduce the misjudgment of automated strategies. Usually, the preset threshold for filtering invalid keyword combinations is 50%-80%.

[0229] Exemplarily, obtain a set of negative feedback texts with negative feedback intentions marked with valid negative feedback types, where the valid negative feedback types include type A, type B, and type C; if the negative feedback texts of type A account for 20%-30% of the number of negative feedback texts in the negative feedback text set, and the proportion of the keyword combination used to detect type A that hits negative feedback texts is 70%-80%, it means that the combination keyword cannot effectively identify the negative feedback texts of type A, indicating that the keyword combination is very common and a filtering operation is performed on it.

[0230] Then, a preset effective threshold is set to retain keyword combinations with relatively high negative feedback features, and a negative feedback text recognition strategy is formulated to perform a gray-scale online operation on the keyword combinations. Usually, the preset effective threshold is 20%-50%. Based on the keyword combinations in the gray-scale online operation, the negative feedback text to be detected is detected. If the negative feedback text is confirmed to be effective, the user account that publishes the relevant negative feedback text will be suppressed according to the severity. Then, according to the feedback of the user on the suppression operation, the effectiveness of the keyword combination is judged to determine the effective combined keywords. Finally, the effective combined keywords in the keyword library are officially launched.

[0231] In step 707, it is submitted for manual review.

[0232] In the initial stage of constructing the keyword library, since there are fewer keywords in the library, the negative feedback text to be detected can be sent to the manual review seat; based on the combination keywords to be detected extracted from the negative feedback text to be detected with negative feedback intention and the negative feedback features of the combination keywords to be detected, the review result of the combination keywords to be detected by the manual review seat is obtained, so as to determine the target combination keywords and add them to the keyword library as a supplement to the automated strategy.

[0233] In the embodiment of the present application, first, the TF-IDF technology is used to extract keywords in the negative feedback text. After filtering the common words from the keywords obtained by the word segmentation result, different keywords are automatically combined to obtain multiple groups of new keyword combinations (candidate combined keywords); then, the keyword combinations are matched with the historical negative feedback text to obtain two negative feedback features of the keyword combinations. One is the number of negative feedback texts that the keyword combination can match, and the other is the number of effective negative feedback texts hit by the keyword combination; finally, based on the negative feedback features of the keyword combination, after setting a threshold, the keyword combinations with relatively high negative feedback features are retained and used as new keywords to supplement the database, which improves the efficiency and coverage of keyword extraction and enhances the timeliness of keyword processing.

[0234] Next, the implementation of the keyword processing device 253 for the text provided in the embodiment of the present application as an exemplary structure of a software module will be further described. In some embodiments, as Figure 2 shown, the software module in the keyword processing device 253 for the text stored in the memory 250 may include:

[0235] An acquisition module 2531, configured to acquire a text set, where the text set includes multiple sample texts, and the multiple sample texts have negative feedback intention.

[0236] A processing module 2532, configured to extract multiple candidate keywords from the multiple sample texts; combine the multiple candidate keywords extracted from the same sample text to obtain multiple candidate combined keywords.

[0237] A comparison module 2533 is configured to compare each candidate combined keyword with a plurality of sample texts respectively to obtain a negative feedback feature of each candidate combined keyword, where the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the plurality of sample texts.

[0238] A determination module 2534 is configured to determine a target combined keyword from a plurality of candidate combined keywords based on the negative feedback feature of each candidate combined keyword; add the target combined keyword to a keyword library, where the keyword library is used to compare with a text to be detected to determine whether the text to be detected has a negative feedback intention.

[0239] In some embodiments, the processing module 2532 is further configured to obtain the keywords of each sample text and the total number of keywords of each sample text; for each keyword in each sample text, perform the following processing: obtain the number of times the keyword appears in the sample text; determine the ratio of the number of times the keyword appears in the sample text to the total number of the sample text as a first parameter value; obtain the total number of sample texts in the text set and the number of sample texts containing the keyword; determine the ratio of the total number of sample texts in the text set to the number of sample texts containing the keyword as a second parameter value; fuse the first parameter value and the second parameter value into a third parameter value of the keyword; determine candidate keywords based on the third parameter value of each keyword.

[0240] In some embodiments, the processing module 2532 is further configured to filter out at least one of the following sample texts from the sample texts in the text set: sample texts without negative feedback evidence, sample texts with a length less than a length threshold.

[0241] In some embodiments, the sample text without negative feedback evidence is determined by calling a pre-trained classification model. The processing module 2532 is further configured to obtain training sample texts; label the training sample texts with classification labels to form a training text set, where the classification labels include having negative feedback evidence and not having negative feedback evidence; classify the training sample texts in the training text set through the classification model to obtain predicted labels of the training sample texts in the training text set; determine the loss between the classification label and the predicted label, and update the parameters of the classification model according to the loss.

[0242] In some embodiments, the processing module 2532 is further configured to extract N candidate keywords from each sample text, where N is an integer constant greater than or equal to 2; let n be an integer variable and 2≤n≤N, traverse n to perform the following processing: select any n candidate keywords from the N candidate keywords, and splice the n candidate keywords to obtain a candidate combined keyword.

[0243] In some embodiments, the N candidate keywords are candidate keywords among the L candidate keywords extracted from each sample text, where L is greater than N, and N is a preset value.

[0244] In some embodiments, the comparison module 2533 is further configured to perform the following processing for each candidate combined keyword: comparing the candidate combined keyword with a plurality of sample texts respectively to obtain a first comparison result; in response to the first comparison result indicating that the candidate combined keyword hits at least one sample text, taking the number of sample texts hit by the candidate combined keyword as the first negative feedback feature of the candidate combined keyword.

[0245] In some embodiments, the sample texts are marked with attribute tags, where the attribute tags indicate whether the sample texts are valid sample texts with negative feedback intention. The comparison module 2533 is further configured to perform the following processing for each candidate combined keyword: comparing the candidate combined keyword with the valid sample texts respectively to obtain a second comparison result; in response to the second comparison result indicating that the candidate combined keyword hits the valid sample texts, taking the number of valid sample texts hit by the candidate combined keyword as the second negative feedback feature of the candidate combined keyword.

[0246] In some embodiments, the processing module 2532 is further configured to obtain the filtering threshold interval of the candidate combined keyword; perform the following processing for each candidate combined keyword: taking the ratio of the negative feedback feature of the candidate combined keyword to the total number of sample texts in the text set as the negative feedback feature ratio corresponding to the negative feedback feature; in response to the negative feedback feature ratio being within the filtering threshold interval, filtering out the candidate combined keyword from the keyword library.

[0247] In some embodiments, the processing module 2532 is further configured to obtain the effective threshold interval of the candidate combined keyword, where the threshold in the filtering threshold interval is greater than the threshold in the effective threshold interval; in response to the negative feedback feature ratio being within the effective threshold interval, perform a gray-scale online operation on the keyword library; based on the detection result of the gray-scale online operation, determine the detection accuracy index of multiple target combined keywords in the keyword library; in response to the detection accuracy index of the target combined keyword being less than or equal to the detection accuracy index threshold, delete the target combined keyword from the keyword library; in response to the detection accuracy index of the target combined keyword being greater than the detection accuracy index threshold, determine the target combined keyword as a valid combined keyword; perform an official online operation on the valid combined keywords in the keyword library.

[0248] In some embodiments, the processing module 2532 is further configured to detect the text to be detected based on a keyword library, and perform suppression processing on the publishing account of the text to be detected with a negative feedback intention; the detection result of the gray-scale online includes the number of times the publishing account raises an objection to the suppression processing; obtain the number corresponding to each target combined keyword, and determine the detection accuracy index of the target combined keyword based on the number, where the number is negatively correlated with the detection accuracy index.

[0249] In some embodiments, the processing module 2532 is further configured to, when the number of keywords in the keyword library is less than a quantity threshold, send multiple texts to be detected to a manual review seat; obtain the review result of the manual review seat, where the review result includes whether the text to be detected has a negative feedback intention, and the combined keyword to be detected in the text to be detected when the text to be detected has a negative feedback intention; determine the negative feedback feature ratio corresponding to the negative feedback feature of the combined keyword to be detected; in response to the negative feedback feature ratio corresponding to the negative feedback feature of the combined keyword to be detected being within an effective threshold range, add the combined keyword to be detected to the keyword library.

[0250] In some embodiments, the determination module 2534 is further configured to sort the candidate combined keywords in descending order according to the negative feedback feature of each candidate combined keyword; use at least one candidate combined keyword starting from the head in the sorted result in descending order as the target combined keyword.

[0251] In some embodiments, the comparison module 2533 is further configured to extract multiple keywords to be detected from the text to be detected; compare the keywords to be detected with each keyword in the keyword library to obtain a third comparison result; in response to the third comparison result indicating that the number of times the multiple keywords to be detected match the keyword library exceeds a number threshold, determine that the text to be detected has a negative feedback intention.

[0252] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the keyword processing method of the text in the above embodiment of the present application.

[0253] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where computer-executable instructions or a computer program are stored, and when the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the keyword processing method of the text provided in the embodiment of the present application, for example, Figure 3A the keyword processing method of the text shown.

[0254] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0255] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0256] As an example, the computer-executable instructions may or may not correspond to files in a file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (such as files that store one or more modules, subroutines, or code portions).

[0257] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0258] In summary, through the embodiments of the present application, the TF-IDF technology is used to extract keywords from the sample text, and after filtering the extracted keywords for common words, different keywords are automatically combined to obtain multiple groups of candidate combined keywords, realizing automatic extraction and combination of keywords, effectively reducing the labor cost, and improving the efficiency and coverage rate of keyword extraction; the candidate combined keywords are matched with the historical negative feedback text to obtain the negative feedback characteristics of the candidate combined keywords. Based on the negative feedback characteristics of the candidate combined keywords, after setting a threshold, the candidate combined keywords with higher negative feedback characteristics are left and used as new keywords to supplement the database, enhancing the timeliness of keyword processing while improving the efficiency and coverage rate of keyword extraction.

[0259] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for processing keywords of a text, characterized in that The method includes: Obtaining a text set, where the text set includes multiple sample texts, and the multiple sample texts have a negative feedback intention; Extracting multiple candidate keywords from the multiple sample texts; Combining the multiple candidate keywords extracted from the same sample text to obtain multiple candidate combined keywords; Comparing each candidate combined keyword with the multiple sample texts respectively to obtain the negative feedback feature of each candidate combined keyword, where the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the multiple sample texts; Determining a target combined keyword from the multiple candidate combined keywords based on the negative feedback feature of each candidate combined keyword; Adding the target combined keyword to a keyword library, where the keyword library is used to compare with a text to be detected to determine whether the text to be detected has the negative feedback intention.

2. The method according to claim 1, wherein The extracting multiple candidate keywords from the multiple sample texts includes: Obtaining the keywords of each sample text and the total number of keywords of each sample text; Performing the following processing for each keyword in each sample text: Obtaining the number of times the keyword appears in the sample text; Determining the ratio of the number of times the keyword appears in the sample text to the total number of the sample text as a first parameter value; Obtaining the total number of the sample texts in the text set and the number of sample texts containing the keyword; Determining the ratio of the total number of the sample texts in the text set to the number of sample texts containing the keyword as a second parameter value; Fusing the first parameter value and the second parameter value into a third parameter value of the keyword; Determining the candidate keywords based on the third parameter value of each keyword.

3. The method according to claim 2, wherein Before extracting multiple candidate keywords from the multiple sample texts, the method further includes: Filtering out at least one of the following sample texts from the sample texts in the text set: the sample text that does not contain negative feedback evidence, the sample text with a length less than a length threshold.

4. The method according to claim 3, wherein The sample text that does not contain negative feedback evidence is determined by calling a pre-trained classification model; The method further includes: Training the classification model in the following manner: Obtaining training sample texts; Labeling the training sample texts with classification labels to form a training text set, where the types of the classification labels include having the negative feedback evidence and not having the negative feedback evidence; Classifying the training sample texts in the training text set through the classification model to obtain the predicted labels of the training sample texts in the training text set; Determining the loss between the classification label and the predicted label, and updating the parameters of the classification model according to the loss.

5. The method according to any one of claims 1 to 4, characterized in that, The combining the multiple candidate keywords extracted from the same sample text to obtain multiple candidate combined keywords includes: Extract N candidate keywords from each of the sample texts, where N is an integer constant greater than or equal to 2; Let n be an integer variable, and 2 ≤ n ≤ N. Traverse n to perform the following processing: Select any n of the N candidate keywords, and concatenate the n candidate keywords to obtain a candidate combined keyword.

6. The method according to claim 5, wherein the N candidate keywords are candidate keywords among L candidate keywords extracted from each of the sample texts, where L is greater than N and N is a preset value.

7. The method according to claim 1, characterized in that The step of comparing each candidate combined keyword with the multiple sample texts respectively to obtain the negative feedback feature of each candidate combined keyword includes: Perform the following processing for each candidate combined keyword: Compare the candidate combined keyword with the multiple sample texts respectively to obtain a first comparison result; In response to the first comparison result indicating that the candidate combined keyword hits at least one of the sample texts, use the number of the sample texts hit by the candidate combined keyword as the first negative feedback feature of the candidate combined keyword.

8. The method according to claim 1 or 7, wherein the sample text carries an attribute label, where the attribute label indicates whether the sample text is a valid sample text with a negative feedback intention; The step of comparing each candidate combined keyword with the multiple sample texts respectively to obtain the negative feedback feature of each candidate combined keyword includes: Perform the following processing for each candidate combined keyword: Compare the candidate combined keyword with the valid sample texts respectively to obtain a second comparison result; In response to the second comparison result indicating that the candidate combined keyword hits the valid sample text, use the number of the valid sample texts hit by the candidate combined keyword as the second negative feedback feature of the candidate combined keyword.

9. The method according to any one of claims 1 to 4, characterized in that After adding the target combined keyword to the keyword library, the method further includes: Obtain the filtering threshold interval of the candidate combined keyword; Perform the following processing for each candidate combined keyword: Use the ratio of the negative feedback feature of the candidate combined keyword to the total number of the sample texts in the text set as the negative feedback feature ratio corresponding to the negative feedback feature; In response to the negative feedback feature ratio being within the filtering threshold interval, filter out the candidate combined keyword from the keyword library.

10. The method according to claim 9, wherein After adding the target combined keyword to the keyword library, the method further includes: Obtain the effective threshold interval of the candidate combined keyword, where the threshold in the filtering threshold interval is greater than the threshold in the effective threshold interval; In response to the negative feedback feature ratio being within the effective threshold interval, perform a gray-scale online operation on the keyword library; Based on the detection result of the gray-scale online operation, determine the detection accuracy index of the multiple target combined keywords in the keyword library; In response to the detection accuracy index of the target combined keyword being less than or equal to the detection accuracy index threshold, delete the target combined keyword from the keyword library; In response to the detection accuracy index of the target combined keyword being greater than the detection accuracy index threshold, determine the target combined keyword as a valid combined keyword; Perform an operation of officially going online on the valid combined keywords in the keyword library.

11. The method according to claim 10, wherein After the operation of gray-scale going online on the keyword library, the method further includes: Detect the text to be detected based on the keyword library, and perform suppression processing on the publishing account of the text to be detected with negative feedback intention; The detection result of the gray-scale going online includes the number of times the publishing account raises an objection to the suppression processing. Based on the detection result of the gray-scale going online, determining the detection accuracy index of multiple target combined keywords in the keyword library includes: Obtain the number of times corresponding to each target combined keyword, and determine the detection accuracy index of the target combined keyword based on the number of times, wherein the number of times is negatively correlated with the detection accuracy index.

12. The method according to claim 11, wherein The detecting the text to be detected based on the keyword library includes: Extract multiple keywords to be detected from the text to be detected; Compare the keywords to be detected with each keyword in the keyword library to obtain a third comparison result; In response to the third comparison result indicating that the number of times the multiple keywords to be detected hit the keyword library exceeds the number threshold, determine that the text to be detected has the negative feedback intention.

13. The method according to claim 10, characterized in that, When the number of keywords in the keyword library is less than the number threshold, the method further includes: Send multiple texts to be detected to the manual review seat; Obtain the review result of the manual review seat, wherein the review result includes whether the text to be detected has a negative feedback intention, and the combined keyword to be detected in the text to be detected when the text to be detected has the negative feedback intention; Determine the negative feedback feature ratio corresponding to the negative feedback feature of the combined keyword to be detected; In response to the negative feedback feature ratio corresponding to the negative feedback feature of the combined keyword to be detected being within the valid threshold range, add the combined keyword to be detected from the keyword library.

14. The method according to any one of claims 1 to 4, characterized in that, The determining the target combined keyword from multiple candidate combined keywords based on the negative feedback feature of each candidate combined keyword includes: Arrange the candidate combined keywords in descending order according to the negative feedback feature of each candidate combined keyword; Use at least one candidate combined keyword starting from the head in the result of the descending order arrangement as the target combined keyword.

15. A keyword processing device for text, characterized in that, The apparatus includes: An acquisition module, configured to acquire a text set, wherein the text set includes multiple sample texts, and the multiple sample texts have negative feedback intention; A processing module extracts a plurality of candidate keywords from the plurality of sample texts; combines the plurality of candidate keywords extracted from the same sample text to obtain a plurality of candidate combined keywords; A comparison module is configured to compare each of the candidate combined keywords with the plurality of sample texts respectively to obtain a negative feedback feature of each of the candidate combined keywords, wherein the negative feedback feature is used to characterize the matching degree between the candidate combined keyword and the plurality of sample texts; A determination module is configured to determine a target combined keyword from the plurality of candidate combined keywords based on the negative feedback feature of each of the candidate combined keywords; add the target combined keyword to a keyword library, wherein the keyword library is used to compare with a text to be detected to determine whether the text to be detected has the negative feedback intention.

16. An electronic device, characterized in that, The electronic device includes: A memory for storing computer-executable instructions; A processor, when executing the computer-executable instructions stored in the memory, implements the keyword processing method of the text according to any one of claims 1 to 14.

17. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by the processor, the keyword processing method of the text according to any one of claims 1 to 14 is implemented.

18. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by the processor, the keyword processing method of the text according to any one of claims 1 to 14 is implemented.