Sensitive topic classification detection method, storage medium and electronic device
Through small sample learning and Fine-Tuning methods, the topic pre-training model and Softmax classifier are used to solve the problem of unclear classification of sensitive topics, and more precise sensitivity topic detection and generalization performance improvements are achieved.
Patent Information
- Application Number
- CN202311873222.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the classification of sensitive topics is not clear, resulting in the reasonable inquiry request being accidentally injured or unable to output, and the generalization performance of sensitive topic detection is poor, mainly due to incomplete enumeration of sensitive words and insufficient training corpus.
The small sample learning method is used to extract feature vectors through topic pre-training models, combine Softmax classifier and similarity matching, and use a small amount of labeled data for model training and Fine-Tuning to improve the accuracy of sensitive topic detection.
It realizes more accurate classification of sensitive topics, avoids misleading reasonable requests, reduces data preparation costs and resource consumption, and improves the generalization performance of detection.
Smart Images

Figure CN120277381A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of language recognition and classification, and in particular to a sensitive topic classification detection method, storage medium and electronic device. Background Art
[0003] The commonly used intervention methods for sensitive topics include: detecting the content of the conversation through sensitive word detection; and determining whether the topic is a sensitive topic and what type of sensitive topic it is through detection models.
[0004] There are several problems with using sensitive word detection: intervening in queries containing sensitive words may cause originally reasonable queries to be accidentally affected; enumerating sensitive words may not be able to enumerate all sensitive words related to sensitive topics.
[0005] When using models to train sensitive topic detection, a large amount of sensitive topic training corpus is required, and these sensitive topic corpora are often very difficult to obtain. Therefore, model-based sensitive topic detection often has poor generalization performance. Summary of the invention
[0006] The present application provides a sensitive topic classification detection method, storage medium and electronic device, which are used to solve the defects in the prior art that the classification of sensitive topics is unclear, resulting in unreasonable query request output and reasonable query request failure to output, and realizes more accurate classification and filtering of sensitive topics in questions.
[0007] This application provides a sensitive topic classification detection method, including:
[0008] Obtaining a target query input by a user, and extracting a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained based on a training set;
[0009] Acquire the configured sensitive topic samples of each category, and extract the sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category;
[0010] Performing similarity matching between the target feature vector and the mean vector of each category, and pre-determining the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic;
[0011] The highest similarity is compared with a preset threshold, and whether the target inquiry is a target sensitive topic is determined based on the comparison result.
[0012] A sensitive topic classification and detection method provided by the present application trains a topic pre-training model, which specifically includes: randomly selecting a single text from a labeled dataset as an anchor; randomly selecting a text from the topic category to which the anchor belongs as a positive sample; excluding the category to which the anchor belongs, and randomly extracting a single text from the remaining categories as a negative sample; inputting the anchor, the positive sample, and the negative sample into a neural network to respectively obtain three feature vectors; calculating a first distance between the feature vectors of the positive sample and the anchor in the feature space, and calculating a second distance between the feature vectors of the negative sample and the anchor in the feature space; determining a loss function according to the first distance and the second distance; determining the gradient of the neural network parameters according to the loss function, and updating the parameters through a gradient descent algorithm until the loss function is 0.
[0013] A sensitive topic classification and detection method provided by the present application, the calculating the first distance between the feature vectors of the positive sample and the anchor in the feature space, and calculating the second distance between the feature vectors of the negative sample and the anchor in the feature space, specifically includes: obtaining a first difference between the feature vectors of the positive sample and the anchor; calculating the square of the two-norm of the first difference as the first distance between the positive sample and the anchor; obtaining a second difference between the feature vectors of the negative sample and the anchor; calculating the square of the two-norm of the second difference as the second distance between the negative sample and the anchor.
[0014] A sensitive topic classification and detection method provided by the present application, after the training of the topic pre-training model is completed, the method further includes: learning the parameters of the Softmax classifier based on the configured sensitive topic samples to improve the prediction accuracy; the classifier is used for topic classification of a target query.
[0015] A sensitive topic classification and detection method provided by the present application, the learning the parameters of the Softmax classifier based on the configured sensitive topic samples specifically includes: inputting the sample feature vectors of the sensitive topic samples of each category into the Softmax classifier, and the classifier outputs a probability distribution, and each element in the probability distribution corresponds to a probability value of a category prediction; measuring the difference between the label of the sensitive topic sample and the probability value through a cross-entropy loss function, where each sensitive topic sample corresponds to a cross-entropy; adding all the cross-entropies in the sensitive topic samples to obtain a target loss function, and minimizing the target loss function to learn the parameters in the classifier.
[0016] A sensitive topic classification and detection method provided by the present application trains a topic pre-training model, specifically including: inputting a first text into the topic pre-training model to obtain a first feature vector; inputting a second text into the topic pre-training model to obtain a second feature vector; obtaining an absolute value vector of the difference between the first feature vector and the second feature vector; converting the absolute value vector into a scalar, and obtaining a similarity prediction between the first text and the second text according to the scalar; determining a loss function according to the similarity prediction and a text label, where the text label is the true similarity between the first text and the second text; calculating a gradient through backpropagation according to the loss function, and updating parameters in the topic pre-training model through a gradient descent algorithm until the loss function is equal to 0.
[0017] The present application also provides a sensitive topic classification and detection device, including:
[0018] A target feature vector extraction module, configured to obtain a target query input by a user, and extract a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained according to a training set, and the training set is text data;
[0019] A sample feature vector extraction module, configured to obtain sensitive topic samples of each category configured, and extract sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category;
[0020] A mean vector obtaining module, configured to average the sample feature vectors of each category to obtain a mean vector of the sensitive topic samples of each category;
[0021] A similarity matching module, configured to perform similarity matching between the target feature vector and the mean vectors of each category, and pre-determine the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic;
[0022] A sensitive topic judgment module, configured to compare the highest similarity with a preset threshold, and determine whether the target query is a target sensitive topic according to the comparison result.
[0023] A sensitive topic classification and detection device provided by the present application, a model training module, is used to train a topic pre-training model, specifically including: randomly selecting a single text from a labeled dataset as an anchor; randomly selecting a text from the topic category to which the anchor belongs as a positive sample; excluding the category to which the anchor belongs, and randomly extracting a single text from the remaining categories as a negative sample; inputting the anchor, the positive sample, and the negative sample into a neural network to respectively obtain three feature vectors; calculating a first distance between the feature vectors of the positive sample and the anchor in the feature space, and calculating a second distance between the feature vectors of the negative sample and the anchor in the feature space; determining a loss function according to the first distance and the second distance; determining a gradient with respect to the neural network parameters according to the loss function, and updating the parameters through a gradient descent algorithm until the loss function is 0.
[0024] The present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the computer program to implement the sensitive topic classification and detection method as described in any one of the above.
[0025] The present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein the program runs to implement the sensitive topic classification and detection method as described in any one of the above.
[0026] The present application also provides a computer program product, including a computer program, and the computer program implements the sensitive topic classification and detection method as described in any one of the above when executed by a processor.
[0027] The sensitive topic classification and detection method, storage medium, and electronic device provided by the present application detect sensitive topics through a small-sample learning method, thereby avoiding a large amount of labeled data, avoiding the high cost of data preparation in certain specific applications, and saving resources such as manpower, time, and computing; through the small-sample learning method, it avoids the misjudgment of originally reasonable requests caused by sensitive word detection; it avoids the lack of generalization of sensitive topic detection due to the absence of sensitive words and labeled samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0030] Figure 1 Schematic diagram of the hardware environment for a sensitive topic classification and detection method provided by an embodiment of the present application;
[0031] Figure 2 Schematic diagram of the sensitive topic configuration provided according to an embodiment of the present application;
[0032] Figure 3 Schematic flowchart of a sensitive topic classification and detection method provided by the present application;
[0033] Figure 4 Schematic diagram of the sensitive topic classification and detection framework provided by the present application;
[0034] Figure 5 Schematic diagram of the structure of a sensitive topic classification and detection device provided by the present application;
[0035] Figure 6 Schematic diagram of the structure of an electronic device provided by the present application. Detailed implementation manners
[0036] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0038] According to one aspect of the embodiments of the present application, a sensitive topic classification and detection method is provided. The sensitive topic classification and detection method is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, intelligent home appliance ecosystem, Intelligence House ecosystem, etc. Optionally, in this embodiment, the above-mentioned sensitive topic classification and detection method can be applied to, for example, Figure 1 the hardware environment composed of the terminal device 102 and the server 104 as shown. As Figure 1 shown, the server 104 is connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server, and is used to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server, and are used to provide data operation services for the server 104.
[0039] The above network can include, but is not limited to, at least one of the following: wired network, wireless network. The above wired network can include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart clothes hanger, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.
[0040] Some concepts related to the present application will be explained below.
[0041] Few-shot learning aims to build an accurate machine learning model with less training data. Since the dimension of the input data is a factor determining the resource cost (such as time cost, computing cost, etc.), people can reduce the data analysis / machine learning cost by using few-shot learning.
[0042] Cross Entropy is an important concept in Shannon's information theory, mainly used to measure the differential information between two probability distributions. The significance of cross entropy is the difficulty of text recognition using this model, or from the perspective of compression, how many bits are used to encode each word on average. The significance of complexity is the average number of branches used by this model to represent this text, and its reciprocal can be regarded as the average probability of each word. Smoothing means assigning a probability value to unobserved n-grams to ensure that the word sequence can always obtain a probability value through the language model. In feature engineering, it can be used to measure the similarity between two random variables. In the language model (NLP), since the true distribution p is unknown and the model in the language model is obtained through the training set, cross entropy is to measure the accuracy of this model on the test set.
[0043] The principle of Fine-Tuning is to utilize the known network structure and known network parameters, modify the output layer to our own layer, and fine-tune the parameters of several layers before the last layer. In this way, the powerful generalization ability of the deep neural network is effectively utilized, and the design of complex models and time-consuming training are avoided. Therefore, Fine Tuning is a more appropriate choice when the amount of data is insufficient.
[0044] When intervening in sensitive topics, it is often through two methods: sensitive word detection and judging sensitive topics through a model. Intervening in queries containing sensitive words through sensitive word detection may cause reasonable queries to be misjudged, and the enumeration of sensitive words may not cover all related topics; when detecting through a model, a large amount of labeled data is often required for training. However, the corpus of sensitive topics is very rare. Therefore, in actual use, the generalization effect of the model is often not very good.
[0045] The existing technologies for intervening in sensitive data have the following defects:
[0046] 1. Intervening in queries containing sensitive words may cause reasonable queries to be misjudged;
[0047] 2. The enumeration of sensitive words may not be able to enumerate all sensitive words related to sensitive topics completely;
[0048] 3. The labeled data for sensitive topics is scarce, so the generalization performance of sensitive topic detection is often not good.
[0049] This patent overcomes, on the one hand, the queries that misjudge reasonable queries through sensitive word detection by means of few-shot learning; on the other hand, it overcomes the queries in which the generalization effect of the model is not good due to insufficient training corpus in ordinary models.
[0050] The following is a specific explanation.
[0051] Before using the model, it needs to be trained. Two training methods are provided below.
[0052] The first method: Train the topic pre-training model according to a sensitive topic classification and detection method provided by this application. Specifically, it includes: randomly selecting a single text from the labeled dataset as an anchor; randomly selecting a text from the topic category to which the anchor belongs as a positive sample; excluding the category to which the anchor belongs, and randomly extracting a single text from the remaining categories as a negative sample; inputting the anchor, positive sample, and negative sample into the neural network to obtain three feature vectors respectively; calculating the first distance between the feature vectors of the positive sample and the anchor in the feature space, and calculating the second distance between the feature vectors of the negative sample and the anchor in the feature space; determining the loss function according to the first distance and the second distance; determining the gradient with respect to the neural network parameters according to the loss function, and updating the parameters through the gradient descent algorithm until the loss function is 0.
[0053] According to a sensitive topic classification and detection method provided by this application, calculate the first distance between the feature vectors of the positive sample and the anchor in the feature space, and calculate the second distance between the feature vectors of the negative sample and the anchor in the feature space. Specifically, it includes: obtaining the first difference between the feature vectors of the positive sample and the anchor; calculating the square of the two-norm of the first difference as the first distance between the positive sample and the anchor; obtaining the second difference between the feature vectors of the negative sample and the anchor; calculating the square of the two-norm of the second difference as the second distance between the negative sample and the anchor.
[0054] Among them, the first distance should be very small because the positive sample and the anchor belong to the same category, and the smaller the better; the second distance should be very large because the negative sample and the anchor belong to different categories, and the larger the better.
[0055] The second method: Train the topic pre-training model according to a sensitive topic classification and detection method provided by this application. Specifically, it includes: inputting the first text into the topic pre-training model to obtain a first feature vector; inputting the second text into the topic pre-training model to obtain a second feature vector; obtaining the absolute value vector of the difference between the first feature vector and the second feature vector; converting the absolute value vector into a scalar, and obtaining the similarity prediction between the first text and the second text according to the scalar; determining the loss function according to the similarity prediction and the text label, where the text label is the true similarity between the first text and the second text; calculating the gradient through backpropagation according to the loss function, and updating the parameters in the topic pre-training model through the gradient descent algorithm until the loss function is equal to 0.
[0056] Specifically, in the case of few-shot learning with a small amount of labeled data, the knowledge already learned by the model is utilized to perform task classification. Therefore, providing a large amount of unlabeled topic data is also important for few-shot learning of sensitive topic detection. The small amount of labeled data set is text data marked with sensitive topic categories, which can be used to adjust the parameters of the classifier. The unlabeled data is text data without labels, which can be used to adjust the parameters of the model embedding function.
[0057] Configure the sensitive topics in the conversation, including: configuration of sensitive topic types, configuration of sensitive topic examples, and configuration of expected responses to sensitive topics. The configuration process can present what the sensitive topics are, as well as the corresponding examples and expected responses. For example, Figure 1 as shown, it includes the configuration of sensitive topics such as games and commodities. For example, in games, there are often abusive remarks such as personal attacks and scapegoating, and the above remarks can be used as sensitive topics; in the product introduction, the sale of pirated, burned, and illegal commodities can also be used as sensitive topics.
[0058] The unlabeled data can be classified according to the configuration in the sensitive topic configuration module, and each category can learn the corresponding knowledge. For example, for the sensitive topic of commodities, concepts, introductions, etc. related to commodities can be crawled as unlabeled data for pre-training the model, and the model will learn more knowledge related to commodities.
[0059] Extract feature vectors from the text through the convolutional neural network in the topic pre-training model, and map the text into the feature space for easy similarity comparison.
[0060] According to a sensitive topic classification and detection method provided by the present application, after the topic pre-training model is trained, the parameters of the Softmax classifier are learned based on the configured sensitive topic samples to improve the prediction accuracy; the classifier is used to perform topic classification on the target query.
[0061] According to a sensitive topic classification and detection method provided by the present application, learning the parameters of the Softmax classifier based on the configured sensitive topic samples specifically includes: inputting the sample feature vectors of the sensitive topic samples of each category into the Softmax classifier, and the classifier outputs a probability distribution, and each element in the probability distribution corresponds to a probability value of a category prediction; the difference between the label of the sensitive topic sample and the probability value is measured by the cross-entropy loss function, where each sensitive topic sample corresponds to a cross-entropy; adding up all the cross-entropies in the sensitive topic samples to obtain the target loss function, and minimizing the target loss function to learn the parameters in the classifier.
[0062] Specifically, the loss function is used to measure the difference between the label and the predicted probability. After the model is pre-trained, Fine-Tuning can be used to further improve the prediction accuracy. Initialize the weight parameters in the Softmax classifier as a matrix formed by the mean vectors of each category, and then input the feature vector into the Softmax classifier;
[0063] Multiply the matrix composed of each mean vector by the target feature vector, perform a Softmax transformation on the multiplication result to obtain the p vector, and the p vector is the probability distribution; the category corresponding to the largest element value in the p vector is the topic category of the target question.
[0064] Since the amount of data for sensitive topic matching is small, entropy regularization can be used to prevent overfitting in the Softmax classifier.
[0065] Knowledge related to topics is very important for the inspection of sensitive topics. Therefore, we pre-train different sensitive topics separately to let the model learn the knowledge related to sensitive topics. To avoid the influence between topics, a model can be pre-trained for each different sensitive topic. The method is to perform fine-tuning on the basis of the Bert series or GPT series model with unsupervised data related to the topic to make the classification more accurate.
[0066] Figure 3 A sensitive topic classification and detection method provided by an embodiment of the present application may include the following steps:
[0067] S310: Obtain the target question input by the user, and extract the target feature vector of the target question through the topic pre-training model; the topic pre-training model is trained according to the training set, and the training set is text data;
[0068] S320: Obtain the sensitive topic samples of each configured category, and extract the sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category;
[0069] S330: Average the sample feature vectors of each category to obtain the mean vector of the sensitive topic samples of each category;
[0070] S340: Perform similarity matching between the target feature vector and the mean vectors of each category, and pre-determine the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic;
[0071] S350: Compare the highest similarity with a preset threshold, and determine whether the target question is a target sensitive topic according to the comparison result.
[0072] Specifically, as Figure 4The figure shows the model architecture diagram for matching user queries. The topic pre-training model is trained for similarity using a large amount of unlabeled data. After the training is completed, the obtained similarity function is used for prediction. The query is compared with each configured sensitive topic one by one for similarity, and the sensitive topic with the highest similarity is used as the category of the query. Then, the highest similarity is compared with the preset similarity threshold. If the highest similarity is greater than the preset similarity threshold, it indicates that the query belongs to a sensitive topic; if the highest similarity is less than the preset similarity threshold, it indicates that the query belongs to a non-sensitive topic.
[0073] Match the user query with the sensitive topic samples. The matching process is as follows:
[0074] a. First, input the user's query into the topic pre-training model respectively to obtain the relevant feature vector representations;
[0075] b. Input the configured topic samples into their respective topic pre-training models respectively to obtain the corresponding sample representations;
[0076] c. Average the feature vector representations of the topic samples obtained in b to find the center point of the sample representations, that is, the mean vector;
[0077] d. Calculate the distances between the query vector representations obtained in a and the sensitive topic center point representations obtained in c respectively;
[0078] e. Obtain the topic with the closest distance as the target sensitive topic;
[0079] f. Compare the shortest distance with the preset distance threshold. If the shortest distance is greater than the preset distance threshold, the query is a non-sensitive topic. If the shortest distance is less than the preset distance threshold, determine that the query is the target sensitive topic and output it.
[0080] If the topic with the closest distance is a non-sensitive topic, output that the query is a non-sensitive topic. If the topic with the closest distance is a sensitive topic, only output which type of sensitive topic it belongs to and do not make a reply.
[0081] The following describes the sensitive topic classification detection device provided by this application. As Figure 5 shown, the sensitive topic classification detection device described below can be mutually referred to with the sensitive topic classification detection method described above.
[0082] A sensitive topic classification detection device includes:
[0083] A target feature vector extraction module 510 is configured to obtain a target query input by a user, and extract a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained according to a training set, and the training set is text data.
[0084] A sample feature vector extraction module 520 is configured to obtain sensitive topic samples of each configured category, and extract sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category.
[0085] A mean vector obtaining module 530 is configured to average the sample feature vectors of each category to obtain a mean vector of the sensitive topic samples of each category.
[0086] A similarity matching module 540 is configured to perform similarity matching between the target feature vector and the mean vectors of each category, and pre-determine the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic.
[0087] A sensitive topic determination module 550 is configured to compare the highest similarity with a preset threshold, and determine whether the target query is a target sensitive topic according to the comparison result.
[0088] According to a sensitive topic classification and detection device provided by the present application, a model training module 508 is configured to train the topic pre-training model, specifically including: randomly selecting a single text from a labeled data set as an anchor point; randomly selecting a text from the topic category to which the anchor point belongs as a positive sample; excluding the category to which the anchor point belongs, and randomly extracting a single text from the remaining categories as a negative sample; inputting the anchor point, the positive sample, and the negative sample into a neural network to respectively obtain three feature vectors; calculating a first distance between the feature vectors of the positive sample and the anchor point in the feature space, and calculating a second distance between the feature vectors of the negative sample and the anchor point in the feature space; determining a loss function according to the first distance and the second distance; determining a gradient of the neural network parameters according to the loss function, and updating the parameters through a gradient descent algorithm until the loss function is 0.
[0089] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown in Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 660. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 660. The processor 610 may call the logical instructions in the memory 630 to execute a sensitive topic classification detection method, which includes: obtaining a target query input by a user, and extracting a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained according to a training set; obtaining sensitive topic samples of each configured category, and extracting sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category; performing similarity matching between the target feature vector and the mean vectors of each category, and pre-determining the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic; comparing the highest similarity with a preset threshold, and determining whether the target query is a target sensitive topic according to the comparison result.
[0090] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0091] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program, which can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the sensitive topic classification and detection method provided by each of the above methods. The method includes: obtaining a target query input by a user, and extracting a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained according to a training set; obtaining sensitive topic samples of each configured category, and extracting sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category; performing similarity matching between the target feature vector and the mean vectors of each category, and preliminarily determining the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic; comparing the highest similarity with a preset threshold, and determining whether the target query is a target sensitive topic according to the comparison result.
[0092] On another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it executes the sensitive topic classification and detection method provided by each of the above methods. The method includes: obtaining a target query input by a user, and extracting a target feature vector of the target query through a topic pre-training model; the topic pre-training model is trained according to a training set; obtaining sensitive topic samples of each configured category, and extracting sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category; performing similarity matching between the target feature vector and the mean vectors of each category, and preliminarily determining the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic; comparing the highest similarity with a preset threshold, and determining whether the target query is a target sensitive topic according to the comparison result.
[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A sensitive topic classification and detection method, characterized in that, Including: Obtain the target query input by the user, and extract the target feature vector of the target query through the topic pre-training model; The topic pre-training model is trained according to the training set; Obtain the sensitive topic samples of each configured category, and extract the sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category; Average the sample feature vectors of each category to obtain the mean vector of the sensitive topic samples of each category; Perform similarity matching between the target feature vector and the mean vectors of each category, and pre-determine the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic; Compare the highest similarity with a preset threshold, and determine whether the target query is a target sensitive topic according to the comparison result.
2. The sensitive topic classification and detection method according to claim 1, wherein Train the topic pre-training model, specifically including: Randomly select a single text from the labeled dataset as the anchor point; Randomly select a text from the topic category to which the anchor point belongs as the positive sample; Exclude the category to which the anchor point belongs, and randomly extract a single text from the remaining categories as the negative sample; Input the anchor point, the positive sample, and the negative sample into the neural network, and obtain three feature vectors respectively; Calculate the first distance between the feature vectors of the positive sample and the anchor point in the feature space, and calculate the second distance between the feature vectors of the negative sample and the anchor point in the feature space; Determine the loss function according to the first distance and the second distance; Determine the gradient of the neural network parameters according to the loss function, and update the parameters through the gradient descent algorithm until the loss function is 0.
3. The sensitive topic classification and detection method according to claim 2, characterized in that, The calculating the first distance between the feature vectors of the positive sample and the anchor point in the feature space, and calculating the second distance between the feature vectors of the negative sample and the anchor point in the feature space specifically includes: Obtain the first difference between the feature vector of the positive sample and the feature vector of the anchor point; Calculate the square of the two-norm of the first difference as the first distance between the positive sample and the anchor point; Obtain the second difference between the feature vector of the negative sample and the feature vector of the anchor point; Calculate the square of the two-norm of the second difference as the second distance between the negative sample and the anchor point.
4. The sensitive topic classification and detection method according to claim 1, characterized in that After the training of the topic pre-training model is completed, the method further includes: Learn the parameters of the Softmax classifier based on the configured sensitive topic samples to improve the prediction accuracy; The classifier is used to perform topic classification on the target query.
5. The sensitive topic classification and detection method according to claim 4, characterized in that The learning the parameters of the Softmax classifier based on the configured sensitive topic samples specifically includes: Input the sample feature vectors of the sensitive topic samples of each category into the Softmax classifier, and the classifier outputs a probability distribution, and each element in the probability distribution corresponds to a probability value of a category prediction; Measure the difference between the label of the sensitive topic sample and the probability value through the cross-entropy loss function, where each sensitive topic sample corresponds to a cross-entropy; Add up all the cross-entropies in the sensitive topic samples to obtain the target loss function, and minimize the target loss function to learn the parameters in the classifier.
6. The sensitive topic classification and detection method according to claim 1, characterized in that Train the topic pre-training model, specifically including: Input the first text into the topic pre-training model to obtain the first feature vector; Input the second text into the topic pre-training model to obtain a second feature vector; Obtain an absolute value vector of the difference between the first feature vector and the second feature vector; Convert the absolute value vector into a scalar, and obtain a similarity prediction of the first text and the second text according to the scalar; Determine a loss function according to the similarity prediction and a text label, where the text label is the true similarity between the first text and the second text; Calculate a gradient through backpropagation according to the loss function, and update parameters in the topic pre-training model through a gradient descent algorithm until the loss function is equal to 0.
7. A sensitive topic classification and detection device, characterized in that, including: A target feature vector extraction module, configured to obtain a target query input by a user, and extract a target feature vector of the target query through a topic pre-training model; The topic pre-training model is trained according to a training set; A sample feature vector extraction module, configured to obtain sensitive topic samples of each category configured, and extract sample feature vectors of the sensitive topic samples of each category through the topic pre-training model corresponding to each category; A mean vector obtaining module, configured to average the sample feature vectors of each category to obtain a mean vector of the sensitive topic samples of each category; A similarity matching module, configured to perform similarity matching between the target feature vector and the mean vectors of each category, and pre-determine the sensitive topic corresponding to the category with the highest similarity as the target sensitive topic; A sensitive topic determination module, configured to compare the highest similarity with a preset threshold, and determine whether the target query is a target sensitive topic according to the comparison result.
8. The sensitive topic classification and detection device according to claim 7, wherein A model training module, configured to train the topic pre-training model, specifically including: Randomly select a single text from a labeled data set as an anchor point; randomly select a text from the topic category to which the anchor point belongs as a positive sample; exclude the category to which the anchor point belongs, and randomly extract a single text from the remaining categories as a negative sample; input the anchor point, the positive sample, and the negative sample into a neural network, and respectively obtain three feature vectors; calculate a first distance between the feature vectors of the positive sample and the anchor point in the feature space, and calculate a second distance between the feature vectors of the negative sample and the anchor point in the feature space; determine a loss function according to the first distance and the second distance; determine a gradient of the neural network parameters according to the loss function, and update the parameters through a gradient descent algorithm until the loss function is 0.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 7 when running.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.