Text auditing method and device, storage medium and electronic equipment

By using the recognition model trained by the original sample and the target sample, keyword replacement and review of the published barrage and comments, the problem of insufficient audit accuracy in the existing technology is solved and higher audit accuracy is achieved.

CN119988619APending Publication Date: 2025-05-13BEIJING IQIYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411892029.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has insufficient accuracy when reviewing barrage and comments, making it difficult to effectively identify and block unqualified content.

Method used

The published target text is reviewed by using a recognition model trained from the original sample and the target sample. The target sample is obtained by keyword substitution of the original sample, and the replacement position and probability are determined according to the specific category.

Benefits of technology

The accuracy of auditing of target text is improved, and the trained recognition model can more accurately identify inappropriate barrage and comments, thereby improving the overall effect of the audit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988619A_ABST
    Figure CN119988619A_ABST
Patent Text Reader

Abstract

The invention relates to a text auditing method and device, a storage medium and electronic equipment. The method comprises an input module used for inputting a target text into a target recognition model after obtaining the target text to be published, the target recognition model being a recognition model obtained by training an original sample and a target sample, and the target recognition model being a recognition model obtained by training the original sample and the target sample; the target sample is a sample obtained by performing keyword replacement on the original sample; the acquisition module is used for acquiring the category of the target text output by the target recognition model; and the publishing module is used for publishing the target text under the condition that the category is the target category. According to the method and the device, the technical problem of inaccurate review of unqualified bullet screens and comments is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bullet screen and comment review, and in particular to a text review method, device, storage medium and electronic device. Background Art

[0002] With the development of the Internet, users can communicate with multimedia programs and other users by sending bullet screens and comments while watching multimedia programs through the client. However, the bullet screens and comments sent by users need to be reviewed.

[0003] How to accurately review unqualified barrages and comments and block them is an urgent problem that needs to be solved. Summary of the invention

[0004] The present application provides a text review method, device, storage medium and electronic device to solve the technical problem of inaccurate review of unqualified barrages and comments.

[0005] In a first aspect, the present application provides a text review method, comprising: after obtaining a target text to be published, inputting the target text into a target recognition model, wherein the target recognition model is a recognition model trained using original samples and target samples, and the target samples are samples obtained by replacing keywords in the original samples; obtaining the category of the target text output by the target recognition model; and publishing the target text when the category is the target category.

[0006] In the second aspect, the present application provides a text review device, comprising: an input module, used to input the target text to be published into a target recognition model after acquiring the target text to be published, wherein the target recognition model is a recognition model trained using original samples and target samples, and the target samples are samples obtained by replacing keywords in the original samples; an acquisition module, used to acquire the category of the target text output by the target recognition model; and a publishing module, used to publish the target text when the category is the target category.

[0007] As an optional example, the above-mentioned device also includes: a first training module, which is used to obtain the above-mentioned original sample and the first category to which the above-mentioned original sample belongs before obtaining the target text to be published; determine the position of the to-be-replaced keyword in the above-mentioned original sample and the replacement probability of the above-mentioned keyword to be replaced according to the above-mentioned first category; perform a replacement operation on the above-mentioned original sample according to the position of the above-mentioned keyword to be replaced and the replacement probability of the above-mentioned keyword to be replaced to obtain the above-mentioned target sample.

[0008] As an optional example, the first training module includes: a first replacement unit, used to determine the number of times a replaceable word appears in each position in the original sample of the first category; determine the position with the highest number of times as the position of the keyword to be replaced; and determine the ratio of the number of times to the total number of times as the replacement probability of the keyword to be replaced.

[0009] As an optional example, the first training module includes: a second replacement unit, used to determine candidate words for the keywords to be replaced according to the first category; and to replace the keywords to be replaced with the candidate words according to the replacement probability to obtain the target sample.

[0010] As an optional example, the first training module includes: a correction unit, used to obtain the original category obtained by initially dividing the original sample; and correct the original category to obtain the first category of the original sample.

[0011] As an optional example, the above-mentioned device also includes: a second training module, which is used to obtain the above-mentioned original sample and the sample features of the above-mentioned original sample before obtaining the target text to be published; linearly change the above-mentioned sample features to obtain the query vector, key vector and value vector of the above-mentioned sample features; determine the feature with the highest weight among the above-mentioned sample features according to the above-mentioned query vector, the above-mentioned key vector and the above-mentioned value vector; determine the text corresponding to the feature with the highest weight as the keyword to be replaced; use the candidate word to replace the above-mentioned keyword to be replaced to obtain the target sample.

[0012] As an optional example, the second training module is further used to determine the feature with the highest weight among the sample features by using the following formula (1):

[0013]

[0014] Among them, Attention(Q,K,V) is the weight, Q is the query vector, K is the key vector, V is the value vector, and d k is the vector dimension of Q above.

[0015] In a third aspect, the present application provides an electronic device comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the memory stores a computer program, and the processor is configured to implement any one of the above-mentioned text review methods when executing the computer program.

[0016] In a fourth aspect, the present application also provides a computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute any of the above-mentioned text review methods of the present application.

[0017] The above-mentioned technical scheme provided by the embodiment of the present application has the following advantages over the prior art: the scheme provided by the embodiment of the present application, after obtaining the target text to be published, inputs the above-mentioned target text into the target recognition model, wherein the above-mentioned target recognition model is a recognition model obtained by training using the original sample and the target sample, and the above-mentioned target sample is a sample obtained by replacing the keywords of the above-mentioned original sample; the category of the above-mentioned target text output by the above-mentioned target recognition model is obtained; when the above-mentioned category is the target category, the above-mentioned target text is published, so that the target sample replaced by the original sample can be used together with the original sample to train the recognition model, and the recognition accuracy of the trained recognition model is higher. The target text to be published is recognized by the trained recognition model, thereby improving the review accuracy of the target text. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0021] Figure 1 A flowchart of a text review method provided in an embodiment of the present application;

[0022] Figure 2 A flowchart of another text review method provided in an embodiment of the present application;

[0023] Figure 3 A flowchart of another text review method provided in an embodiment of the present application;

[0024] Figure 4 A flowchart of another text review method provided in an embodiment of the present application;

[0025] Figure 5 A schematic diagram of the structure of a text review device provided in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0028] The disclosure below provides many different embodiments or examples to implement different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0029] In order to solve the technical problem of inaccurate review of unqualified barrages and comments in the prior art, the present application provides a text review method that can achieve the effect of improving the review accuracy of the target text.

[0030] Figure 1 A flowchart of a text review method provided in an embodiment of the present application. Figure 1 As shown, the above text review method includes:

[0031] S102, after obtaining the target text to be published, inputting the target text into a target recognition model, wherein the target recognition model is a recognition model trained using the original sample and the target sample, and the target sample is a sample obtained by replacing the keywords of the original sample;

[0032] S104, obtaining the category of the target text output by the target recognition model;

[0033] S106: When the above category is a target category, publish the above target text.

[0034] This embodiment can be applied in the process of automatically reviewing bullet screens and comments. During the playback of multimedia programs, users may post inappropriate bullet screens, comments, and other content, which are not suitable for other users to watch. Therefore, this part of the content should be blocked and not allowed to be published. How to accurately review this part of the bullet screen and comments and identify the inappropriate bullet screen and comments is a problem that needs to be solved.

[0035] This embodiment proposes the above-mentioned text review method, which uses the original text and the target text obtained by replacing the keywords of the original text to train the recognition model to identify whether the target text is appropriate and whether it is allowed to be published.

[0036] The target text mentioned above is the text to be reviewed, such as bullet comments, comments, etc. Since the contents of bullet comments, comments, etc. are real-time, this embodiment can display the bullet comments and comments posted by the user locally on the client after the user posts the bullet comments and comments through the client, and after the background obtains the bullet comments and comments, it will use the text review method of this embodiment to review the bullet comments and comments, so that after the review is passed, it can be published to other users. If the review fails, it cannot be published.

[0037] The above-mentioned original samples can be samples used to train the recognition model. The samples can include bullet comments, comments and corresponding types. The types can be one or more of the above-mentioned appropriate types or inappropriate types. A bullet comment or comment cannot include both appropriate types and inappropriate types. Keywords can be determined in the original samples. After the keywords are determined, the keywords are replaced to obtain the target samples.

[0038] The above-mentioned target recognition model is a model for training to identify the category of the target text. The categories in the present embodiment can be divided into multiple categories, for example, to identify whether the target text is an inappropriate text, inappropriate texts are also divided into many categories, such as abusive type, spoiler type, etc., and appropriate texts can also be divided into many categories, such as encouraging type, friendly type, etc. By classifying the text, multiple categories are obtained, and some of the categories are determined to be inappropriate categories. If the target text is identified as an inappropriate category, it cannot be published. And other categories are appropriate categories. If the target text is identified as an appropriate category, it is allowed to be published.

[0039] The solution provided by the embodiment of the present application, after obtaining the target text to be published, inputs the above target text into the target recognition model, wherein the above target recognition model is a recognition model obtained by training using the original sample and the target sample, and the above target sample is a sample obtained by replacing the keywords of the above original sample; the category of the above target text output by the above target recognition model is obtained; when the above category is the target category, the above target text is published, so that the target sample replaced by the original sample can be used together with the original sample to train the recognition model, and the recognition accuracy of the trained recognition model is higher. The target text to be published is recognized by the trained recognition model, thereby improving the review accuracy of the target text.

[0040] As an alternative example, Figure 2 As shown, before obtaining the target text to be published, the method further includes:

[0041] S202, obtaining the original sample and the first category to which the original sample belongs;

[0042] S204, determining the position of the keyword to be replaced in the original sample and the replacement probability of the keyword to be replaced according to the first category;

[0043] S206 , performing a replacement operation on the original sample according to the position of the keyword to be replaced and the replacement probability of the keyword to be replaced to obtain the target sample.

[0044] The first category is the classification category of the original sample. There can be one or more classification categories. According to the different classification categories, the position of the keyword to be replaced in the original sample and the replacement probability of being replaced are also different. The position of the keyword to be replaced is the position of the keyword to be replaced. In this embodiment, the keyword to be replaced is determined according to the position. The position is different for different categories of different original samples, and the result is that the determined keywords to be replaced are different. For example, for an original sample of a certain category, according to the category, the first two words of the original sample are determined to be replaced. For another category, the middle two words of the original sample are replaced, and so on. The above replacement probability is the possibility of the keyword to be replaced being replaced. A replacement probability of 50% means that the keyword to be replaced may have a 50% chance of being replaced, while a replacement probability of 100% means that the keyword to be replaced must be replaced, and a replacement probability of 0% means that the keyword to be replaced cannot be replaced.

[0045] As an optional example, the above-mentioned determining the position of the keyword to be replaced in the above-mentioned original sample and the replacement probability of the above-mentioned keyword to be replaced according to the above-mentioned first category includes: determining the number of times the replaceable word appears in each position in the above-mentioned original sample of the above-mentioned first category; determining the position with the highest number of times as the position of the above-mentioned keyword to be replaced; and determining the ratio of the above number of times to the total number of times as the replacement probability of the above-mentioned keyword to be replaced.

[0046] In this embodiment, when determining the position of the keyword to be replaced, the number of times the replaceable word appears in each position in the original sample can be determined according to the category of the original sample. For example, for an original sample, the original sample may include multiple replaceable words, each replaceable word is located at a different position, and the number of times the replaceable words in different positions appear in each position is 1. For different original samples, the number of times each replaceable word appears in each position is counted. Since there are many original samples, the number of times it appears in each position is not 1, and the counted number of times is large or small. Among the counted numbers, the position in the original sample corresponding to the highest number is the position of the keyword to be replaced, and the ratio of the counted number of times at this position to the total number of times is the replacement probability of the keyword to be replaced.

[0047] For example, take two original samples in a category as an example. The text content of the original samples is "123456" and "abcdef". Among them, "123456" includes two replaceable words, "12" and "56", and "abcdef" includes one replaceable word, "ab". It can be found that "12" is located in the first two words of the first original sample, so the first two words are recorded as a position with a frequency of 1, and "56" is the fifth and sixth words of the original sample, so the fifth and sixth words are recorded as a position with a frequency of 1, and "ab" is located in the first two words of the second original sample, so the first two words are recorded as a position with a frequency increase of 1, becoming 2, and there is no replaceable word in the position of the third and fourth words, so the record number is 0. Then, for this type of original sample, the first two words have the largest number of replacements, so the positions of the first two words are used as the positions of the keywords to be replaced, and the first two words are used as keywords to be replaced, and the replacement probability is 66.6%.

[0048] As an optional example, the above-mentioned replacement operation is performed on the above-mentioned original sample according to the position of the above-mentioned keyword to be replaced and the replacement probability of the above-mentioned keyword to be replaced to obtain the above-mentioned target sample, including: determining the candidate words of the above-mentioned keyword to be replaced according to the above-mentioned first category; according to the above-mentioned replacement probability, using the above-mentioned candidate words to replace the above-mentioned keyword to be replaced to obtain the above-mentioned target sample.

[0049] In this embodiment, for each category of original samples, a candidate word set can be prepared, and the candidate words in the candidate word set can be different according to different categories. For each category of original samples, the candidate words can be used to replace the keywords to be replaced. When replacing, the replacement probability of the keywords to be replaced in the original samples must be followed. After the replacement, the target sample is obtained.

[0050] In this embodiment, for each original sample, each candidate word in the candidate word set can be used to replace it to obtain a target sample. If there are 3 original samples and 3 candidate words, then when the replacement probability is 1, 9 target samples will be obtained. Adding the 3 original samples, there will be a total of 12 samples used to train the recognition model. And the combination of the candidate words and the positions of the candidate words in the 12 samples is different. If the replacement probability is not 1, it means that it may be replaced by the candidate word or may not be replaced by the candidate word. If the replacement probability is 1 / 3, then the 3 original samples and the 3 candidate words will also obtain 9 target samples, but 6 of the 9 target samples are the same as the original samples and have not been replaced by the candidate words.

[0051] As an optional example, obtaining the first category to which the original sample belongs includes: obtaining the original category obtained by initially dividing the original sample; and correcting the original category to obtain the first category of the original sample.

[0052] In this embodiment, the original samples can be preliminarily divided by the classification model, which can be a neural network model for classification. The classification model classifies the batch of original samples. The purpose of classification is to classify similar original samples into one category, rather than outputting the classification of the original samples as in the above-mentioned recognition model. Then, for the batch of original samples, they can be divided into multiple original sample sets. The original samples in the original sample sets have high similarity, while the original samples in different original sample sets have low similarity. For the original sample sets obtained after classification, the engineering personnel participate in manual classification, and correct the categories of the original samples in the original sample sets, so as to obtain accurately classified original samples.

[0053] For example, 10,000 original samples are classified into 8 categories by the classification model. Each category includes at least one original sample. The engineering staff assigns categories to the original samples of the 8 categories and corrects the original samples in each category. If there are original samples that are incorrectly classified, new categories are opened or moved to other categories among the 8 categories.

[0054] After correction, 8 or more categories of original samples are obtained, and the category of each original sample is accurate.

[0055] As an alternative example, Figure 3 As shown, before obtaining the target text to be published, the method further includes:

[0056] S302, obtaining the original sample and the sample features of the original sample;

[0057] S304, performing linear transformation on the sample features to obtain a query vector, a key vector, and a value vector of the sample features;

[0058] S306, determining a feature with the highest weight among the sample features according to the query vector, the key vector, and the value vector;

[0059] S308, determining the text corresponding to the feature with the highest weight as the keyword to be replaced;

[0060] S310, using candidate words to replace the above keywords to be replaced to obtain target samples.

[0061] In this embodiment, the sample features of the original samples mentioned above may be features extracted by a feature extraction model, and the features may have multiple dimensions, for example, may be sample features of 768 dimensions.

[0062] The above linear transformation of sample features can be used to map 768-dimensional sample features into a high-dimensional space, such as mapping 768-dimensional sample features into 2000-dimensional sample features, where the 2000-dimensional sample features include query vectors, key vectors, and value vectors. Then, the feature with the highest weight among the sample features is calculated based on the query vector, key vector, and value vector. The text corresponding to the feature with the highest weight calculated is the keyword to be replaced. The keyword to be replaced is replaced by the candidate keyword to obtain the target sample. The formula for calculating the weight can use the following formula (1). The weight of the feature is calculated by formula (1).

[0063]

[0064] Among them, Attention(Q,K,V) is the above weight, the above Q is the above query vector, the above K is the above key vector, the above V is the above value vector, and the above dk is the vector dimension of the above Q.

[0065] Combination Figure 4 Provide explanation. Figure 4 This is a flow chart of this embodiment.

[0066] In this embodiment, data collection is performed based on massive business data to obtain raw data, and a large model is used to perform preliminary division on the massive raw data. The prompt word prompt of the large model adopts a form similar to the following: "The above is a bullet comment I entered. Please judge the category to which the bullet comment belongs according to the following rules. The rules are as follows: Category 1, provoking a fight and cursing. Category 2, meaningless. Category 3, attack. Category 4, advertising. Category 5, passed: normal bullet comments that do not belong to the above categories, for example: good-looking; so nervous; there is an easter egg at the end of the film, no need to thank me; hahaha. According to the above rules, to determine which category the input bullet comment belongs to, you only need to answer the category name, for example: The male protagonist next door is more handsome, the answer is: provoking a fight and cursing".

[0067] The large model is used to complete the initial cleaning of massive data, complete the coarse-grained labeling work, and then organize the labeled data into categories, and use manual or model screening to further classify the labeled data in more detailed categories. For example, the category of inciting war and cursing can be divided into actor attack category, character attack category, plot attack category, staff attack category, etc. according to the different targets of attack. Since the coarse category automatic labeling has been achieved in the first step, the second step of screening and labeling only needs to use the coarse labeling results as a pre-labeling reference for detailed labeling. This form of labeling will save a lot of manual time and reduce labor costs. Samples that have been finely labeled can be used as original samples.

[0068] Based on the labeled original samples, dynamic mask generation rules are set according to different categories. For example, for the actor attack category, the actor's name can be replaced; the dynamic mask replacement position, replacement probability and replacement word are set for this subcategory. For this category, the replacement position is the first 20% of the sentence, the replacement probability is 50%, and the replacement word is the name of the star in the business artist database, so as to achieve data augmentation for this subcategory; similarly, there is also the attack on staff category; through retrieval, the job name in the text is identified, the position where the job appears is dynamically masked, replaced with other job words or directly deleted, to complete the data augmentation for this subcategory.

[0069] For the original sample, the attention mechanism can be used to further realize data augmentation, which can complement the rule augmentation and make the data richer. The process is: the transformer pressure model is used to extract the features of the text data of the original sample to obtain the 768-dimensional text features, and then the self-attention mechanism is used to calculate the attention mechanism weight. The calculation method is:

[0070] Input a 768-dimensional feature sequence and undergo three linear transformations (three trainable weights) to obtain the query vector Q, key vector K and value vector V. Use Q, K and V to calculate the attention mechanism weights, obtain the position with the highest self-attention weight score, and determine the keyword here. This keyword is the position where the subsequent dynamic mask is set, that is, the keyword to be replaced.

[0071] During the training process, at each epoch, i.e. the beginning of the traversal process, the data is augmented differently, and then the text classification model is trained to improve the model's ability to understand the text context, improve the model's generalization ability, and ultimately improve the model's recognition accuracy.

[0072] It should be noted that although Figure 4 The processes in are parallel processes, but in this embodiment, after replacing the keywords of the original samples in the left process, the samples obtained after the replacement can be subjected to data augmentation of the attention mechanism of the right process, and it is not limited to that the two cannot be integrated with each other. This embodiment uses a large model to achieve coarse-category labeling of data, which can greatly shorten the data labeling cycle and reduce the difficulty of data labeling. With the coarse-category pre-labeling results, fine labeling can be performed, which can improve the labeling efficiency and greatly reduce the labeling time and cost. In the data augmentation part, customized data augmentation methods are used for different fine categories. Two forms of rule augmentation and attention mechanism augmentation are adopted. The data obtained by integrating the two augmentations are used for model tuning, which improves the diversity of data augmentation and greatly improves the generalization ability of the model. Data augmentation is dynamically adjusted during the model training process, instead of setting fixed mask positions and probabilities, which can make the augmentation richer and more natural, and more in line with the diversity of barrages and comments in daily life.

[0073] Figure 5 This is a schematic diagram of the structure of a text review device provided in an embodiment of the present application. Figure 5 As shown, the above-mentioned text review device includes:

[0074] An input module 502 is used to input the target text to be published into a target recognition model after acquiring the target text to be published, wherein the target recognition model is a recognition model trained using the original sample and the target sample, and the target sample is a sample obtained by replacing the keywords of the original sample;

[0075] An acquisition module 504 is used to acquire the category of the target text output by the target recognition model;

[0076] The publishing module 506 is used to publish the target text when the category is the target category.

[0077] This embodiment can be applied in the process of automatically reviewing bullet screens and comments. During the playback of multimedia programs, users may post inappropriate bullet screens, comments, and other content, which are not suitable for other users to watch. Therefore, this part of the content should be blocked and not allowed to be published. How to accurately review this part of the bullet screen and comments and identify the inappropriate bullet screen and comments is a problem that needs to be solved.

[0078] This embodiment proposes the above-mentioned text review method, which uses the original text and the target text obtained by replacing the keywords of the original text to train the recognition model to identify whether the target text is appropriate and whether it is allowed to be published.

[0079] The target text mentioned above is the text to be reviewed, such as bullet comments, comments, etc. Since the contents of bullet comments, comments, etc. are real-time, this embodiment can display the bullet comments and comments posted by the user locally on the client after the user posts the bullet comments and comments through the client, and after the background obtains the bullet comments and comments, it will use the text review method of this embodiment to review the bullet comments and comments, so that after the review is passed, it can be published to other users. If the review fails, it cannot be published.

[0080] The above-mentioned original samples can be samples used to train the recognition model. The samples can include bullet comments, comments and corresponding types. The types can be one or more of the above-mentioned appropriate types or inappropriate types. A bullet comment or comment cannot include both appropriate types and inappropriate types. Keywords can be determined in the original samples. After the keywords are determined, the keywords are replaced to obtain the target samples.

[0081] The above-mentioned target recognition model is a model for training to identify the category of the target text. The categories in the present embodiment can be divided into multiple categories, for example, to identify whether the target text is an inappropriate text, inappropriate texts are also divided into many categories, such as abusive type, spoiler type, etc., and appropriate texts can also be divided into many categories, such as encouraging type, friendly type, etc. By classifying the text, multiple categories are obtained, and some of the categories are determined to be inappropriate categories. If the target text is identified as an inappropriate category, it cannot be published. And other categories are appropriate categories. If the target text is identified as an appropriate category, it is allowed to be published.

[0082] The solution provided by the embodiment of the present application, after obtaining the target text to be published, inputs the above target text into the target recognition model, wherein the above target recognition model is a recognition model obtained by training using the original sample and the target sample, and the above target sample is a sample obtained by replacing the keywords of the above original sample; the category of the above target text output by the above target recognition model is obtained; when the above category is the target category, the above target text is published, so that the target sample replaced by the original sample can be used together with the original sample to train the recognition model, and the recognition accuracy of the trained recognition model is higher. The target text to be published is recognized by the trained recognition model, thereby improving the review accuracy of the target text.

[0083] As an optional example, the above-mentioned device also includes: a first training module, which is used to obtain the above-mentioned original sample and the first category to which the above-mentioned original sample belongs before obtaining the target text to be published; determine the position of the to-be-replaced keyword in the above-mentioned original sample and the replacement probability of the above-mentioned keyword to be replaced according to the above-mentioned first category; perform a replacement operation on the above-mentioned original sample according to the position of the above-mentioned keyword to be replaced and the replacement probability of the above-mentioned keyword to be replaced to obtain the above-mentioned target sample.

[0084] The first category is the classification category of the original sample. There can be one or more classification categories. According to the different classification categories, the position of the keyword to be replaced in the original sample and the replacement probability of being replaced are also different. The position of the keyword to be replaced is the position of the keyword to be replaced. In this embodiment, the keyword to be replaced is determined according to the position. The position is different for different categories of different original samples, and the result is that the determined keywords to be replaced are different. For example, for an original sample of a certain category, according to the category, the first two words of the original sample are determined to be replaced. For another category, the middle two words of the original sample are replaced, and so on. The above replacement probability is the possibility of the keyword to be replaced being replaced. A replacement probability of 50% means that the keyword to be replaced may have a 50% chance of being replaced, while a replacement probability of 100% means that the keyword to be replaced must be replaced, and a replacement probability of 0% means that the keyword to be replaced cannot be replaced.

[0085] As an optional example, the first training module includes: a first replacement unit, used to determine the number of times a replaceable word appears in each position in the original sample of the first category; determine the position with the highest number of times as the position of the keyword to be replaced; and determine the ratio of the number of times to the total number of times as the replacement probability of the keyword to be replaced.

[0086] In this embodiment, when determining the position of the keyword to be replaced, the number of times the replaceable word appears in each position in the original sample can be determined according to the category of the original sample. For example, for an original sample, the original sample may include multiple replaceable words, each replaceable word is located at a different position, and the number of times the replaceable words in different positions appear in each position is 1. For different original samples, the number of times each replaceable word appears in each position is counted. Since there are many original samples, the number of times it appears in each position is not 1, and the counted number of times is large or small. Among the counted numbers, the position in the original sample corresponding to the highest number is the position of the keyword to be replaced, and the ratio of the counted number of times at this position to the total number of times is the replacement probability of the keyword to be replaced.

[0087] For example, take two original samples in a category as an example. The text content of the original samples is "123456" and "abcdef". Among them, "123456" includes two replaceable words, "12" and "56", and "abcdef" includes one replaceable word, "ab". It can be found that "12" is located in the first two words of the first original sample, so the first two words are recorded as a position with a frequency of 1, and "56" is the fifth and sixth words of the original sample, so the fifth and sixth words are recorded as a position with a frequency of 1, and "ab" is located in the first two words of the second original sample, so the first two words are recorded as a position with a frequency increase of 1, becoming 2, and there is no replaceable word in the position of the third and fourth words, so the record number is 0. Then, for this type of original sample, the first two words have the largest number of replacements, so the positions of the first two words are used as the positions of the keywords to be replaced, and the first two words are used as keywords to be replaced, and the replacement probability is 66.6%.

[0088] As an optional example, the first training module includes: a second replacement unit, configured to determine candidate words of the keyword to be replaced according to the first category; and to replace the keyword to be replaced with the candidate words according to the replacement probability.

[0089] In this embodiment, for each category of original samples, a candidate word set can be prepared, and the candidate words in the candidate word set can be different according to different categories. For each category of original samples, the candidate words can be used to replace the keywords to be replaced. When replacing, the replacement probability of the keywords to be replaced in the original samples must be followed. After the replacement, the target sample is obtained.

[0090] In this embodiment, for each original sample, each candidate word in the candidate word set can be used to replace it to obtain a target sample. If there are 3 original samples and 3 candidate words, then when the replacement probability is 1, 9 target samples will be obtained. Adding the 3 original samples, there will be a total of 12 samples used to train the recognition model. And the combination of the candidate words and the positions of the candidate words in the 12 samples is different. If the replacement probability is not 1, it means that it may be replaced by the candidate word or may not be replaced by the candidate word. If the replacement probability is 1 / 3, then the 3 original samples and the 3 candidate words will also obtain 9 target samples, but 6 of the 9 target samples are the same as the original samples and have not been replaced by the candidate words.

[0091] As an optional example, the first training module includes: a correction unit, used to obtain the original category obtained by initially dividing the original sample; and correct the original category to obtain the first category of the original sample.

[0092] In this embodiment, the original samples can be preliminarily divided by the classification model, which can be a neural network model for classification. The classification model classifies the batch of original samples. The purpose of classification is to classify similar original samples into one category, rather than outputting the classification of the original samples as in the above-mentioned recognition model. Then, for the batch of original samples, they can be divided into multiple original sample sets. The original samples in the original sample sets have high similarity, while the original samples in different original sample sets have low similarity. For the original sample sets obtained after classification, the engineering personnel participate in manual classification, and correct the categories of the original samples in the original sample sets, so as to obtain accurately classified original samples.

[0093] For example, 10,000 original samples are classified into 8 categories by the classification model. Each category includes at least one original sample. The engineering staff assigns categories to the original samples of the 8 categories and corrects the original samples in each category. If there are original samples that are incorrectly classified, new categories are opened or moved to other categories among the 8 categories.

[0094] After correction, 8 or more categories of original samples are obtained, and the category of each original sample is accurate.

[0095] As an optional example, the above-mentioned device also includes: a second training module, which is used to obtain the above-mentioned original sample and the sample features of the above-mentioned original sample before obtaining the target text to be published; linearly change the above-mentioned sample features to obtain the query vector, key vector and value vector of the above-mentioned sample features; determine the feature with the highest weight among the above-mentioned sample features according to the above-mentioned query vector, the above-mentioned key vector and the above-mentioned value vector; determine the text corresponding to the feature with the highest weight as the keyword to be replaced; use the candidate word to replace the above-mentioned keyword to be replaced to obtain the target sample.

[0096] In this embodiment, the sample features of the original samples mentioned above may be features extracted by a feature extraction model, and the features may have multiple dimensions, for example, may be sample features of 768 dimensions.

[0097] The above linear transformation of sample features can be used to map 768-dimensional sample features to a high-dimensional space, such as mapping 768-dimensional sample features to 2000-dimensional sample features, which include query vectors, key vectors, and value vectors. Then, the feature with the highest weight among the sample features is calculated based on the query vector, key vector, and value vector. The text corresponding to the feature with the highest weight calculated is the keyword to be replaced. The keyword to be replaced is replaced by the candidate keyword to obtain the target sample. The formula for calculating the weight can use the above formula (1).

[0098] Combination Figure 4 Provide explanation. Figure 4 This is a flow chart of this embodiment.

[0099] In this embodiment, data collection is performed based on massive business data to obtain raw data, and a large model is used to perform preliminary division on the massive raw data. The prompt word prompt of the large model adopts a form similar to the following: "The above is a bullet comment I entered. Please judge the category to which the bullet comment belongs according to the following rules. The rules are as follows: Category 1, provoking a fight and cursing. Category 2, meaningless. Category 3, attack. Category 4, advertising. Category 5, passed: normal bullet comments that do not belong to the above categories, for example: good-looking; so nervous; there is an easter egg at the end of the film, no need to thank me; hahaha. According to the above rules, to determine which category the input bullet comment belongs to, you only need to answer the category name, for example: The male protagonist next door is more handsome, the answer is: provoking a fight and cursing".

[0100] The large model is used to complete the initial cleaning of massive data and the coarse-grained labeling work. The labeled data is then sorted into categories, and manual screening is used to further classify the labeled data into more detailed categories. For example, the category of inciting war and cursing can be divided into actor attack category, character attack category, plot attack category, staff attack category, etc. according to the different targets of attack. Since the coarse category automatic labeling has been achieved in the first step, manual labeling only needs to use the coarse labeling results as a pre-labeling reference for detailed labeling. This form of labeling will save a lot of manual time and reduce labor costs. Samples that have been finely labeled can be used as original samples.

[0101] Based on the labeled original samples, dynamic mask generation rules are set according to different categories. For example, for the actor attack category, the actor's name can be replaced; the dynamic mask replacement position, replacement probability and replacement word are set for this subcategory. For this category, the replacement position is the first 20% of the sentence, the replacement probability is 50%, and the replacement word is the name of the star in the business artist database, so as to achieve data augmentation for this subcategory; similarly, there is also the attack on staff category; through retrieval, the job name in the text is identified, the position where the job appears is dynamically masked, replaced with other job words or directly deleted, to complete the data augmentation for this subcategory.

[0102] For the original sample, the attention mechanism can be used to further realize data augmentation, which can complement the rule augmentation and make the data richer. The process is: the transformer pressure model is used to extract the features of the text data of the original sample to obtain the 768-dimensional text features, and then the self-attention mechanism is used to calculate the attention mechanism weight. The calculation method is:

[0103] Input a 768-dimensional feature sequence and undergo three linear transformations (three trainable weights) to obtain the query vector Q, key vector K and value vector V. Use Q, K and V to calculate the attention mechanism weights, obtain the position with the highest self-attention weight score, and determine the keyword here. This keyword is the position where the subsequent dynamic mask is set, that is, the keyword to be replaced.

[0104] During the training process, at each epoch, i.e. the beginning of the traversal process, the data is augmented differently, and then the text classification model is trained to improve the model's ability to understand the text context, improve the model's generalization ability, and ultimately improve the model's recognition accuracy.

[0105] It should be noted that although Figure 4The processes in are parallel processes, but in this embodiment, after replacing the keywords of the original samples in the left process, the samples obtained after the replacement can be subjected to data augmentation of the attention mechanism of the right process, and it is not limited to that the two cannot be integrated with each other. This embodiment uses a large model to achieve coarse-category labeling of data, which can greatly shorten the data labeling cycle and reduce the difficulty of data labeling. With the coarse-category pre-labeling results, fine labeling can be performed, which can improve the labeling efficiency and greatly reduce the labeling time and cost. In the data augmentation part, customized data augmentation methods are used for different fine categories. Two forms of rule augmentation and attention mechanism augmentation are adopted. The data obtained by integrating the two augmentations are used for model tuning, which improves the diversity of data augmentation and greatly improves the generalization ability of the model. Data augmentation is dynamically adjusted during the model training process, instead of setting fixed mask positions and probabilities, which can make the augmentation richer and more natural, and more in line with the diversity of barrages and comments in daily life.

[0106] For other examples of this embodiment, please refer to the above examples and will not be described in detail here.

[0107] like Figure 6 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0108] Memory 113, used for storing computer programs;

[0109] In one embodiment of the present application, the processor 111 is used to implement the text review method provided by any one of the aforementioned method embodiments when executing the program stored in the memory 113.

[0110] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the text review method provided in any of the aforementioned method embodiments is implemented.

[0111] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0113] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.

[0114] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A text review method, characterized in that: include: After obtaining the target text to be published, the target text is input into a target recognition model, wherein the target recognition model is a recognition model trained using an original sample and a target sample, and the target sample is a sample obtained by replacing keywords in the original sample; Obtaining the category of the target text output by the target recognition model; In a case where the category is a target category, the target text is published.

2. The method according to claim 1, characterized in that: Before obtaining the target text to be published, the method further includes: Acquire the original sample and the first category to which the original sample belongs; Determining, according to the first category, the position of the keyword to be replaced in the original sample and the replacement probability of the keyword to be replaced; The target sample is obtained by performing a replacement operation on the original sample according to the position of the keyword to be replaced and the replacement probability of the keyword to be replaced.

3. The method according to claim 2, characterized in that The step of determining the position of the keyword to be replaced in the original sample and the replacement probability of the keyword to be replaced according to the first category includes: Determine the number of times the replaceable word appears at each position in the original sample of the first category; Determine the position with the highest number of times as the position of the keyword to be replaced; The ratio of the number of times to the total number of times is determined as the replacement probability of the keyword to be replaced.

4. The method according to claim 2, characterized in that: The replacing operation on the original sample according to the position of the to-be-replaced keyword and the replacement probability of the to-be-replaced keyword to obtain the target sample comprises: Determining candidate words for the keyword to be replaced according to the first category; According to the replacement probability, the candidate word is used to replace the keyword to be replaced to obtain the target sample.

5. The method according to claim 2, characterized in that: Acquiring the first category to which the original sample belongs includes: Obtaining original categories obtained by initially dividing the original samples; The original category is corrected to obtain the first category of the original sample.

6. The method according to claim 1, characterized in that Before obtaining the target text to be published, the method further includes: Acquire the original sample and sample features of the original sample; Performing a linear change on the sample feature to obtain a query vector, a key vector, and a value vector of the sample feature; Determining a feature with the highest weight among the sample features according to the query vector, the key vector and the value vector; The text corresponding to the feature with the highest weight is determined as the keyword to be replaced; The keyword to be replaced is replaced with a candidate word to obtain a target sample.

7. The method according to claim 6, characterized in that The determining, according to the query vector, the key vector, and the value vector, the feature with the highest weight among the sample features comprises: The feature with the highest weight among the sample features is determined by the following formula (1): Where Attention(Q,K,V) is the weight, Q is the query vector, K is the key vector, V is the value vector, and d k is the vector dimension of Q.

8. A text review device, characterized in that: include: An input module is used to input the target text to be published into a target recognition model after acquiring the target text to be published, wherein the target recognition model is a recognition model trained using an original sample and a target sample, and the target sample is a sample obtained by replacing keywords in the original sample; An acquisition module, used for acquiring the category of the target text output by the target recognition model; A publishing module is used to publish the target text when the category is a target category.

9. An electronic device, characterized in that: include: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein the memory stores a computer program, and the processor implements the text review method described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the text review method described in any one of claims 1 to 7 of the present application.