Text processing method and device, equipment, storage medium and program product
By combining example matching and knowledge enhancement prompts, and using the BERT model and knowledge graph to build a contextual knowledge subgraph, the problem of inaccurate toxicity judgment in text processing is solved, and the accuracy and security of text generation are achieved.
Patent Information
- Application Number
- CN202510670597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies are not precise in judging the toxicity of text during text processing, resulting in inaccurate generated text results.
By obtaining example matching hints and knowledge enhancement hints for the input text, using a large model for detection and rewriting, combining the attention score and knowledge graph of the BERT model to build a contextual knowledge subgraph, and generating candidate output text.
It improves the accuracy of text generation, ensures the security and applicability of output text, and reduces the risk of generating toxic content.
Smart Images

Figure CN120596648A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a text processing method, apparatus, device, storage medium, and program product. Background Art
[0002] Style transfer schemes are generally divided into two categories: unsupervised and supervised. Unsupervised methods are built on non-parallel datasets. Supervised methods are trained on parallel datasets where there is a one-to-one mapping between toxic entities and non-toxic text.
[0003] In existing technologies, when processing text, a paraphrase model generates multiple candidate words or phrases from the input text. These candidate words are semantically consistent with the original text as much as possible. A style control model then intervenes, evaluating the stylistic attributes of each candidate word, such as toxicity, and adjusting the generation probability of these words based on the target style (e.g., detoxification or other stylistic requirements). Finally, the model selects the most appropriate word from the candidate words based on these adjusted probabilities for output, thereby gradually generating the entire text.
[0004] However, existing technologies are not accurate in judging the toxicity of text, which leads to inaccurate generated text results. Summary of the Invention
[0005] Embodiments of the present application provide a text processing method, apparatus, device, storage medium, and program product to improve the accuracy of generated text.
[0006] In a first aspect, an embodiment of the present application provides a text processing method, comprising:
[0007] Get input text;
[0008] Obtaining example matching hints and knowledge enhancement hints for the input text;
[0009] Using the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and using the large model to detect and / or rewrite the input text to obtain candidate output text;
[0010] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; the knowledge enhancement prompts include relationship prompts between toxic entities in the input text and subject entities associated with the toxic entities.
[0011] Optionally, obtaining example matching hints for the input text includes:
[0012] Obtaining a vocabulary similarity score between the input text and training examples in a training dataset;
[0013] Obtaining semantic similarity scores between the input text and training examples in the training dataset;
[0014] Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example;
[0015] From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
[0016] Optionally, obtaining a knowledge enhancement prompt for the input text includes:
[0017] Obtaining attention scores for each word in the input text using a Bidirectional Encoder Representation of Transformer (BERT) model;
[0018] Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments;
[0019] Constructing a contextual knowledge subgraph based on the toxic text fragment;
[0020] The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
[0021] Optionally, constructing a contextual knowledge subgraph based on the toxic text fragment includes:
[0022] Obtaining a subject entity from the input text;
[0023] Obtaining poisonous entities in the poisonous text segment;
[0024] According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0025] Optionally, the method further includes:
[0026] The contextual knowledge subgraph is pruned to obtain a final contextual knowledge subgraph.
[0027] Optionally, pruning the contextual knowledge subgraph to obtain a final contextual knowledge subgraph includes:
[0028] Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph;
[0029] The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
[0030] Optionally, converting the contextual knowledge subgraph into the knowledge enhancement prompt includes:
[0031] forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph;
[0032] forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity;
[0033] The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
[0034] Optionally, the method further includes:
[0035] Obtaining a toxicity probability value of the candidate output text;
[0036] If the toxicity probability value is greater than a fourth threshold, the large model is reused to detect and / or rewrite the candidate output text; if the toxicity probability value is less than or equal to the fourth threshold, the candidate output text is output.
[0037] In a second aspect, an embodiment of the present application further provides a text processing device, comprising:
[0038] A first receiving module, configured to receive input text;
[0039] A first acquisition module, configured to acquire example matching prompts and knowledge enhancement prompts for the input text;
[0040] a first processing module, configured to use the example matching prompt and the knowledge enhancement prompt as prompts of a large model, and detect and / or rewrite the input text using the large model to obtain a candidate output text;
[0041] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; the knowledge enhancement prompts include relationship prompts between toxic entities in the input text and subject entities associated with the toxic entities.
[0042] Optionally, the first acquisition module is further configured to:
[0043] Obtaining a vocabulary similarity score between the input text and training examples in a training dataset;
[0044] Obtaining semantic similarity scores between the input text and training examples in the training dataset;
[0045] Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example;
[0046] From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
[0047] Optionally, the first acquisition module is further configured to:
[0048] Using the BERT model to obtain the attention score of each word in the input text;
[0049] Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments;
[0050] Constructing a contextual knowledge subgraph based on the toxic text fragment;
[0051] The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
[0052] Optionally, the first acquisition module is further configured to:
[0053] Obtaining a subject entity from the input text;
[0054] Obtaining poisonous entities in the poisonous text segment;
[0055] According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0056] Optionally, the device may further include:
[0057] The first processing module is configured to prune the contextual knowledge subgraph to obtain a final contextual knowledge subgraph.
[0058] Optionally, the first processing module is further configured to:
[0059] Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph;
[0060] The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
[0061] Optionally, the first acquisition module is further configured to:
[0062] forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph;
[0063] forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity;
[0064] The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
[0065] Optionally, the device further includes:
[0066] A second acquisition module is used to obtain a toxicity probability value of the candidate output text;
[0067] The second processing module is used to reuse the large model to detect and / or rewrite the candidate output text if the toxicity probability value is greater than a fourth threshold; and output the candidate output text if the toxicity probability value is less than or equal to the fourth threshold.
[0068] In a third aspect, an embodiment of the present application further provides a text processing device, comprising: a processor and a transceiver; wherein the processor is configured to:
[0069] Receive input text;
[0070] Obtaining example matching hints and knowledge enhancement hints for the input text;
[0071] Using the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and using the large model to detect and / or rewrite the input text to obtain candidate output text;
[0072] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; and the knowledge enhancement prompts include prompts between toxic entities and associated entities in the input text.
[0073] Optionally, the processor is further configured to:
[0074] Obtaining a vocabulary similarity score between the input text and training examples in a training dataset;
[0075] Obtaining semantic similarity scores between the input text and training examples in the training dataset;
[0076] Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example;
[0077] From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
[0078] Optionally, the processor is further configured to:
[0079] Using the BERT model to obtain the attention score of each word in the input text;
[0080] Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments;
[0081] Constructing a contextual knowledge subgraph based on the toxic text fragment;
[0082] The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
[0083] Optionally, the processor is further configured to:
[0084] Obtaining a subject entity in the input text;
[0085] Obtaining poisonous entities in the poisonous text segment;
[0086] According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0087] Optionally, the processor is further configured to:
[0088] The contextual knowledge subgraph is pruned to obtain a final contextual knowledge subgraph.
[0089] Optionally, the processor is further configured to:
[0090] Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph;
[0091] The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
[0092] Optionally, the processor is further configured to:
[0093] forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph;
[0094] forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity;
[0095] The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
[0096] Optionally, the processor is further configured to:
[0097] Obtaining a toxicity probability value of the candidate output text;
[0098] If the toxicity probability value is greater than a fourth threshold, the large model is reused to detect and / or rewrite the candidate output text; if the toxicity probability value is less than or equal to the fourth threshold, the candidate output text is output.
[0099] In a fourth aspect, an embodiment of the present application further provides a communication device, comprising: a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the steps in the text processing method as described above when executing the program.
[0100] In a fifth aspect, an embodiment of the present application further provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the text processing method as described above are implemented.
[0101] In a sixth aspect, an embodiment of the present application further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps in the text processing method described above.
[0102] In an embodiment of the present application, example matching prompts and knowledge enhancement prompts are combined to enable the large model to obtain knowledge information from examples and relationship prompts between entities from knowledge enhancement prompts, thereby helping the large model to better understand the complexity of the input text, accurately detect or rewrite toxic text therein, and improve the accuracy of the generated text. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 is a flowchart of a text processing method provided in an embodiment of the present application;
[0104] Figure 2 This is an example of an embodiment of the present application based on an example matching prompt;
[0105] Figure 3 A schematic diagram of the process of obtaining a knowledge enhancement prompt instance in an embodiment of the present application;
[0106] Figure 4 This is an example of a knowledge-enhanced prompt in an embodiment of the present application;
[0107] Figure 5 This is a processing diagram provided by an embodiment of the present application;
[0108] Figure 6This is one of the structural diagrams of the text processing device provided in the embodiment of the present application;
[0109] Figure 7 This is the second structural diagram of the text processing device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0110] In the embodiments of this application, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0111] In the embodiments of the present application, the term "plurality" refers to two or more than two, and other quantifiers are similar.
[0112] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0113] See also Figure 1 , Figure 1 is a flowchart of the text processing method provided in the embodiment of the present application, such as Figure 1 As shown, the following steps are included:
[0114] Step 101: Receive input text.
[0115] The input text may be any piece of text, including one or more sentences, words, etc.
[0116] Step 102: Obtain example matching prompts and knowledge enhancement prompts for the input text.
[0117] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; the knowledge enhancement prompts include relationship prompts between toxic entities in the input text and subject entities associated with the toxic entities.
[0118] The purpose of obtaining example matching prompts is to select the most appropriate examples from a large number of training examples and include them in the prompts, thereby helping the large model better understand the task requirements and generate outputs that are more consistent with expectations. The training examples can include multiple sentences or multiple sentence pairs (including the sentences before and after rewriting).
[0119] Here, examples are selected based on the vocabulary matching strategy and the context matching strategy.
[0120] Specifically, the vocabulary matching strategy refers to obtaining the vocabulary similarity score between the input text and the training examples in the training dataset, especially focusing on those key words in the input text that may affect the sentence style or tone, thereby guiding the large model to generate output similar to the style of the input text.
[0121] The context matching strategy refers to obtaining the semantic similarity score between the input text and the training examples in the training data set. First, the input text and the training examples in the training set are encoded as vector representations, and then the cosine similarity between the vector of the input text and the vector of the training examples in the training set is calculated to determine their semantic similarity, and the example that is semantically closest to the input text is selected in the training set. Such examples provide rich contextual information in the prompts, so that the large model can better grasp the overall semantics of the sentence when generating text. Among them, the embodiments of the present application do not limit the specific method of obtaining the lexical similarity score and the semantic similarity score.
[0122] After obtaining the lexical similarity scores and semantic similarity scores, a weighting system is used to evaluate and rank the training examples. Specifically, the lexical similarity score and semantic similarity score of each training example are weighted to obtain a composite score for each training example. The weighting parameters corresponding to the lexical similarity score and semantic similarity score can be set and adjusted according to the task requirements to balance the influence of vocabulary and semantics.
[0123] Finally, from the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts, wherein the first threshold can be set as needed.
[0124] In the above implementation process, K (K ≥ 1) training examples can be selected from the training set based on the semantic similarity score, where the K training examples can be the K training examples with the highest semantic similarity scores. At the same time, K (K ≥ 1) training examples can be selected from the training set based on the lexical similarity score, where the K training examples can be the K training examples with the highest lexical similarity scores. Thereafter, the lexical similarity scores and semantic similarity scores corresponding to these K training examples are weighted, and then, based on the weighted results, training examples that can be used as example matching prompts are selected.
[0125] like Figure 2An example based on an example matching prompt is shown in FIG. In this example, the rewording is performed by replacing offensive, harmful, and curse words with respectful words. The example matching prompts are shown as "Input" and "Output".
[0126] The knowledge enhancement prompt is mainly based on the concept of knowledge graph. The knowledge graph is represented as G = (E, R, T), where E represents the entity set, R represents the relationship set, and T is the triple set. Each triple (e h ,r,e t ) contains a header entity e h , a relation r, a tail entity e t Given a knowledge graph G, we can start from an entity e0 and reach another entity e through a series of relationships and entities (i.e., k-hop path). k , this path is represented as where r i It's a relationship, i It is a passing entity.
[0127] Among them, combined Figure 3 , obtaining a knowledge enhancement prompt for the input text, including:
[0128] (1) Use the BERT model to obtain the attention score of each word in the input text.
[0129] Here, the toxicity of the input text is first predicted by toxicity span detection. For example, the BERT model can be used to predict the toxicity label of the entire input text, and attribute relation extraction (ARE) can be combined to extract attributes and their relationships (such as synonyms, antonyms, emotional impact, etc.) from the input text. Among them, the toxicity label is used to indicate whether the input text is toxic text or non-toxic text. Among them, toxic text refers to text that contains negative content such as offensive, insulting, discriminatory or inflammatory. Correspondingly, non-toxic text refers to text that does not contain any illegal, illegal, offensive, inflammatory, discriminatory or other negative content.
[0130] In specific applications, a dense layer is added to the output layer of the BERT model and a sigmoid function (S-shaped function) is applied to the [CLS] token to output a value representing the toxicity probability for each word, that is, the attention value of each word, thereby providing an overall toxicity prediction for each sentence. In natural language processing, CLS is short for Classification and refers to a special token ([CLS]) in the input sequence of Transformer-like models (such as BERT). Its core function is to summarize the semantic information of the entire input sequence and is typically used as the output of classification tasks.
[0131] Here, we use the attention mechanism in the BERT model to extract toxic text fragments from the input text. Each layer of BERT has multiple attention heads, which capture the dependencies between different parts of the input sentence and can be expressed as follows:
[0132] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0133] Among them, head i That is, the calculation result of each attention head, W O It is the final linear transformation matrix, Query represents the query vector, K represents the key, V represents the value, and Contact represents concatenating the outputs of multiple attention heads.
[0134] In an embodiment of the present application, for the attention head of the last layer of the BERT model, the scores of its AttentionHead are extracted and averaged to obtain the importance of each word in the entire sentence relative to the toxicity judgment.
[0135] Among them, the attention score of each word is Attention Score i It can be expressed as:
[0136]
[0137] Among them, h is the number of attention heads, Attention Head j (i) is the attention score of the j-th attention head on the i-th word.
[0138] (2) Words with attention scores greater than or equal to the second threshold are regarded as toxic text fragments.
[0139] The second threshold is determined by debugging on the development dataset. By comparing the attention score with the second threshold, each word can be labeled as "toxic" or "non-toxic," thereby generating a binary sequence that marks which parts of the sentence are considered toxic. These labels correspond to specific locations in the original text through character offsets, allowing the precise location of toxic text segments in the text. The toxic text segment can be a sentence or certain words within a sentence.
[0140] exist Figure 3 In the example, the input text is: "You're so dumb, no one would ever listen to anything you say." After the above detection, the part of the text that may contain harmful content is dumb.
[0141] (3) Construct a contextual knowledge subgraph based on the toxic text fragments.
[0142] In this step, the goal is to build a context-related knowledge subgraph for each input text. This subgraph will be centered around entities related to the toxicity of the input text (i.e., toxic entities), helping to capture knowledge related to toxic entities.
[0143] First, the subject entities in the input text can be obtained. For example, the TAGME tool can be used to identify all entities mentioned in the input text. s , and these entities M s Match the entities in the knowledge base to obtain the subject entity E of the input text s As mentioned above, for the acquired poisonous text segment, the poisonous entities in the poisonous text segment can be acquired through the aforementioned marking.
[0144] Afterwards, a two-segment jump subgraph centered on the toxic entity is constructed according to the toxic entity and the subject entity, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0145] Based on the knowledge base, a toxic entity E is generated from the knowledge base. t Two-stage jump sub- Figure 2 -hopsub-graph G s , wherein the two-stage hop subgraph may include the subject entity E s In other words, the toxic entity E tA one- or two-hop structure is constructed with the sentence as the center to a topic entity, thus forming a subgraph containing the semantic knowledge related to the sentence. The topic entity that can form this structure is the topic entity associated with the toxic entity. This association can be reflected in the semantic similarity between the toxic entity and the topic entity meeting certain requirements. In a two-hop subgraph, the maximum number of hops between entities is two.
[0146] The one-hop structure is called first-order structural information, and the two-hop structure is called second-order structural information. Specifically, the substructure of the first-order structural information (1-hop 1-chain) is a simple triple, which represents a direct relationship between the subject entity (i.e., the toxic entity) and another entity. One form of second-order structural information is a 2-hop 1-chain substructure, which represents a 2-hop relationship path starting from the subject entity. Another form of second-order structural information is a 2-hop 2-chain, which represents two 2-hop relationship paths that share the same prefix triple.
[0147] like Figure 3 As shown in the figure, assuming that the toxic entity is "dumb", with "dumb" as the center, a two-step hop sub-network from the toxic entity to the subject entities stupid, insult, and hurt feelings is constructed. Figure 2 -hop sub-graph G s .
[0148] In the formation of two jump Figure 2 -hop sub-graph G s During the process, the generated subgraph can also be pruned to obtain the final contextual knowledge subgraph, thereby reducing some knowledge that is not relevant to the detoxification information. Here, a semantic similarity pruning method can be used to calculate the semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph. Afterwards, the target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold. The third threshold can be set as needed.
[0149] For example, using a pre-trained language model (such as BERT or RoBERTa), the descriptions of the toxic entity and each node in the subgraph are converted into semantic vectors. The semantic similarity between the toxic entity and the semantic vectors of each node in the subgraph is measured by calculating the cosine similarity between them. For the target node, it can be considered that the semantic association with the input text is weak and can be pruned. This ensures that the remaining subgraph contains parts that are closely related to the toxic entity's semantics, providing an accurate information foundation for the subsequent generation of detoxification prompts.
[0150] (4) Converting the contextual knowledge subgraph into the knowledge enhancement prompt.
[0151] After generating context-related knowledge subgraphs, this knowledge is converted into usable hints, helping large models to understand and process text content more accurately in detoxification tasks.
[0152] The aforementioned contextual knowledge subgraph includes first-order structural information and second-order structural information. For the first-order structural information, a first relationship prompt is formed between the toxic entity and the subject entity associated with the toxic entity based on the first-order structural information in the contextual knowledge subgraph. For example, a path starting from the toxic entity can be selected to convert the structural information in the knowledge subgraph into a prompt text. For example, for a 1-hop 1-chain structure, a relationship prompt between a toxic entity and an associated entity is generated. For the second-order structural information, a second relationship prompt is formed between multiple entities, wherein the multiple entities include the toxic entity and the subject entity associated with the toxic entity. For example, a path starting from the toxic entity can be selected to generate a prompt containing multiple entities and the relationships between them, thereby containing more information to help the large model better understand the context of the text.
[0153] Then, based on the first and second relationship hints, the knowledge-enhanced hints are generated. For example, all generated hints can be combined with the original toxic text to form a complete input sequence that includes not only the original context but also the semantic hints generated by the knowledge graph. This allows the model to leverage this additional knowledge to more accurately identify and remove toxic content from the text, thereby improving the overall effectiveness of the detoxification task.
[0154] like Figure 4 The following figure shows an example of text generation based on knowledge-enhanced prompts. The input text is: "You're so dumb, no one would ever listen to anything you say." The part of the text that may contain harmful content is "dumb." Based on the prompts (Hint 1 and Hint 2) in the figure, the input text is rewritten to remove offensive, harmful, and swear words.
[0155] In the embodiments of the present application, semantic hints generated based on the knowledge graph enhance the large model's ability to understand complex contexts, especially when processing complex sentence structures and implicit information. This method can provide additional semantic support for the large model, enabling the large model to perform toxicity detection and text rewriting more accurately.
[0156] Step 103: Use the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and use the large model to detect and / or rewrite the input text to obtain candidate output text.
[0157] Among them, the large model can be any conversational generation model, the purpose of detection is to obtain the toxicity, toxic entities, etc. of the input text; the purpose of rewriting is to rewrite the input text into non-toxic text.
[0158] In an embodiment of the present application, example matching prompts and knowledge enhancement prompts are combined to enable the large model to obtain knowledge information from examples and relationship prompts between entities from knowledge enhancement prompts, thereby helping the large model to better understand the complexity of the input text, accurately detect or rewrite toxic text therein, and improve the accuracy of the generated text.
[0159] For the generated candidate output text, its toxicity can be further evaluated to determine whether it can be output. Specifically, the toxicity probability value of the candidate output text is obtained. For example, the fine-tuned RoBERTa model can be used as a binary toxicity classifier to evaluate the toxicity of the content generated by the large model. First, the model learns how to accurately identify toxic content in the text by fine-tuning a large amount of annotated data. In actual applications, when the large model generates a text, this classifier will perform a preliminary evaluation of the generated content. The classifier converts the input text into a semantic vector, processes it through a dense layer, and outputs a toxicity probability value through the Sigmoid function.
[0160] Afterwards, if the toxicity probability value is greater than the fourth threshold, the large model is reused to detect and / or rewrite the candidate output text; if the toxicity probability value is less than or equal to the fourth threshold, the candidate output text is output. The fourth threshold can be set as needed. If the evaluation result of the classifier shows that the text contains toxicity, but does not indicate which specific part causes it (the classifier only makes toxic and non-toxic judgments), the system will input the text into the large model again and require it to rewrite and generate a new version. This process will be repeated until the generated text passes the detection of the toxicity classifier and is judged to be non-toxic. If the classifier judges that the text is non-toxic, the text will be output as the final detoxified text. Through this cyclic detection and generation method, it can be determined that the output text remains semantically coherent while removing any potential toxic content, thereby improving the security and applicability of the generated text.
[0161] like Figure 5As shown, it is a schematic diagram of the processing process of an embodiment of the present application. For the input text, example matching prompts and knowledge enhancement prompts are generated respectively as prompts for the large model. Afterwards, the large model detects or rewrites the input text. If it is verified that the text output by the large model is toxic, it returns and continues to be processed using the large model, otherwise it can be output as output text. Therefore, in an embodiment of the present application, the two strategies of example matching prompts and knowledge enhancement prompts are combined, not only to obtain guidance information from actual examples, but also to introduce semantic and context-related prompts from the knowledge graph. This dual prompt strategy can help the model better understand the complexity of the input text, especially when faced with implicit toxicity or ambiguity, providing more comprehensive context and semantic support. The cyclic detection and generation mechanism of module three ensures that the text of the final output not only maintains the integrity of the semantics, but also removes toxicity. This repeated generation and verification process greatly improves the success rate of detoxification and ensures the security and applicability of the output text.
[0162] The embodiments of the present application are applicable to scenarios such as training generative models, content review and filtering. When training generative models, it is crucial to ensure the purity of the data set, because toxic data will cause the model to learn bad expressions and then generate harmful content. The embodiments of the present application, through automated toxicity detection and rewriting functions, can identify and remove toxic texts in the preparation stage of training data, thereby generating a pure, non-toxic data set. This not only improves the training quality of the model and ensures that the generated content is safer and more applicable, but also enhances the expressiveness of the model in various application scenarios and reduces potential risks caused by inappropriate content. Content review and filtering are indispensable functions of any online platform, especially on social media, forums and comment platforms. The embodiments of the present application greatly reduce the reliance on manual review and improve review efficiency by automatically detecting and rewriting toxic content. This is particularly important for platforms that need to process a large amount of user-generated content, and can ensure that the content is published in a timely and secure manner, avoiding the negative impact of toxic content on the user experience.
[0163] See also Figure 6 , Figure 6 This is a structural diagram of the text processing device provided in the embodiment of the present application. Figure 6 As shown, the text processing device includes:
[0164] A first receiving module 601 is configured to receive input text; a first acquiring module 602 is configured to acquire example matching prompts and knowledge enhancement prompts for the input text; a first processing module 603 is configured to use the example matching prompts and the knowledge enhancement prompts as prompts for a large model, and to detect and / or rewrite the input text using the large model to obtain candidate output text;
[0165] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; the knowledge enhancement prompts include relationship prompts between toxic entities in the input text and subject entities associated with the toxic entities.
[0166] Optionally, the first acquisition module is further configured to:
[0167] Obtaining a vocabulary similarity score between the input text and training examples in a training dataset;
[0168] Obtaining semantic similarity scores between the input text and training examples in the training dataset;
[0169] Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example;
[0170] From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
[0171] Optionally, the first acquisition module is further configured to:
[0172] Using the BERT model to obtain the attention score of each word in the input text;
[0173] Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments;
[0174] Constructing a contextual knowledge subgraph based on the toxic text fragment;
[0175] The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
[0176] Optionally, the first acquisition module is further configured to:
[0177] Obtaining a subject entity from the input text;
[0178] Obtaining poisonous entities in the poisonous text segment;
[0179] According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0180] Optionally, the device may further include:
[0181] The first processing module is configured to prune the contextual knowledge subgraph to obtain a final contextual knowledge subgraph.
[0182] Optionally, the first processing module is further configured to:
[0183] Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph;
[0184] The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
[0185] Optionally, the first acquisition module is further configured to:
[0186] forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph;
[0187] forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity;
[0188] The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
[0189] Optionally, the device further includes:
[0190] A second acquisition module is used to obtain a toxicity probability value of the candidate output text;
[0191] The second processing module is used to reuse the large model to detect and / or rewrite the candidate output text if the toxicity probability value is greater than a fourth threshold; and output the candidate output text if the toxicity probability value is less than or equal to the fourth threshold.
[0192] The device provided in the embodiment of the present application can execute the above method embodiment, and its implementation principle and technical effects are similar, so this embodiment will not be repeated here.
[0193] See also Figure 7 , Figure 7 This is a structural diagram of the text processing device provided in the embodiment of the present application. Figure 7 As shown, the text processing device includes: a processor 701 and a transceiver 702; wherein, the processor 701 is used to:
[0194] Receive input text;
[0195] Obtaining example matching hints and knowledge enhancement hints for the input text;
[0196] Using the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and using the large model to detect and / or rewrite the input text to obtain candidate output text;
[0197] The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; and the knowledge enhancement prompts include prompts between toxic entities and associated entities in the input text.
[0198] Optionally, the processor 701 is further configured to:
[0199] Obtaining a vocabulary similarity score between the input text and training examples in a training dataset;
[0200] Obtaining semantic similarity scores between the input text and training examples in the training dataset;
[0201] Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example;
[0202] From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
[0203] Optionally, the processor 701 is further configured to:
[0204] Using the BERT model to obtain the attention score of each word in the input text;
[0205] Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments;
[0206] Constructing a contextual knowledge subgraph based on the toxic text fragment;
[0207] The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
[0208] Optionally, the processor 701 is further configured to:
[0209] Obtaining a subject entity from the input text;
[0210] Obtaining poisonous entities in the poisonous text segment;
[0211] According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
[0212] Optionally, the processor 701 is further configured to:
[0213] The contextual knowledge subgraph is pruned to obtain a final contextual knowledge subgraph.
[0214] Optionally, the processor 701 is further configured to:
[0215] Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph;
[0216] The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
[0217] Optionally, the processor 701 is further configured to:
[0218] forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph;
[0219] forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity;
[0220] The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
[0221] Optionally, the processor 701 is further configured to:
[0222] Obtaining a toxicity probability value of the candidate output text;
[0223] If the toxicity probability value is greater than a fourth threshold, the large model is reused to detect and / or rewrite the candidate output text; if the toxicity probability value is less than or equal to the fourth threshold, the candidate output text is output.
[0224] The device provided in the embodiment of the present application can execute the above method embodiment, and its implementation principle and technical effects are similar, so this embodiment will not be repeated here.
[0225] It should be noted that the division of units in the embodiments of the present application is schematic and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0226] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0227] An embodiment of the present application provides a communication device, comprising: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps of the text processing method described above.
[0228] The embodiment of the present application also provides a readable storage medium, on which a program is stored. When the program is executed by the processor, the various processes of the above-mentioned text processing method embodiment are implemented, and the same technical effect is achieved. To avoid repetition, it is not repeated here. The readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disk, hard disk, tape, magneto-optical disk (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0229] An embodiment of the present application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of the above-mentioned text processing method embodiment are implemented and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0230] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0231] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0232] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A text processing method, characterized in that: include: Receive input text; Obtaining example matching hints and knowledge enhancement hints for the input text; Using the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and using the large model to detect and / or rewrite the input text to obtain candidate output text; wherein the example matching hint comprises a training example that matches weighted information of lexical similarity and semantic similarity of the input text; The knowledge enhancement hints include relationship hints between poisonous entities in the input text and subject entities associated with the poisonous entities.
2. The method according to claim 1, characterized in that Get sample matching hints for the input text, including: Obtaining a vocabulary similarity score between the input text and training examples in a training dataset; Obtaining semantic similarity scores between the input text and training examples in the training dataset; Weighting the lexical similarity score and the semantic similarity score of each training example to obtain a comprehensive score of each training example; From the training examples, training examples with comprehensive scores greater than or equal to a first threshold are selected as the example matching prompts.
3. The method according to claim 1, characterized in that Obtaining knowledge enhancement hints for the input text, including: Using the Transformer's bidirectional encoder representation BERT model to obtain attention scores for each word in the input text; Words with an attention score greater than or equal to a second threshold are considered as toxic text fragments; Constructing a contextual knowledge subgraph based on the toxic text fragment; The contextual knowledge subgraph is converted into the knowledge enhancement prompt.
4. The method according to claim 3, characterized in that The step of constructing a contextual knowledge subgraph based on the toxic text fragment includes: Obtaining a subject entity from the input text; Obtaining poisonous entities in the poisonous text segment; According to the toxic entity and the subject entity, a two-segment jump subgraph centered on the toxic entity is constructed, and the two-segment jump subgraph is used as the context knowledge subgraph.
5. The method according to claim 4, characterized in that The method further comprises: The contextual knowledge subgraph is pruned to obtain a final contextual knowledge subgraph.
6. The method according to claim 5, characterized in that The pruning of the contextual knowledge subgraph to obtain a final contextual knowledge subgraph includes: Calculating a semantic similarity score between the toxic entity and each node in the contextual knowledge subgraph; The target node is deleted from the contextual knowledge subgraph to obtain the final contextual knowledge subgraph, wherein the semantic similarity score corresponding to the target node is less than a third threshold.
7. The method according to claim 4, characterized in that The converting the contextual knowledge subgraph into the knowledge enhancement prompt comprises: forming a first relationship hint between a toxic entity and a subject entity associated with the toxic entity according to first-order structural information in the contextual knowledge subgraph; forming a second relationship hint between a plurality of entities according to the second-order structural information in the contextual knowledge subgraph, wherein the plurality of entities include the toxic entity and a subject entity associated with the toxic entity; The knowledge enhancement prompt is generated according to the first relationship prompt and the second relationship prompt.
8. The method according to claim 1, characterized in that The method further comprises: Obtaining a toxicity probability value of the candidate output text; If the toxicity probability value is greater than a fourth threshold, the large model is reused to detect and / or rewrite the candidate output text; if the toxicity probability value is less than or equal to the fourth threshold, the candidate output text is output.
9. A text processing device, characterized in that: include: A first receiving module, configured to receive input text; A first acquisition module, configured to acquire example matching prompts and knowledge enhancement prompts for the input text; a first processing module, configured to use the example matching prompt and the knowledge enhancement prompt as prompts of a large model, and detect and / or rewrite the input text using the large model to obtain a candidate output text; wherein the example matching hint comprises a training example that matches weighted information of lexical similarity and semantic similarity of the input text; The knowledge enhancement hints include relationship hints between poisonous entities in the input text and subject entities associated with the poisonous entities.
10. A text processing device, characterized in that: include: A processor and a transceiver; wherein the processor is configured to: Receive input text; Obtaining example matching hints and knowledge enhancement hints for the input text; Using the example matching prompt and the knowledge enhancement prompt as prompts of the large model, and using the large model to detect and / or rewrite the input text to obtain candidate output text; The example matching prompts include training examples that match weighted information of lexical similarity and semantic similarity of the input text; and the knowledge enhancement prompts include prompts between toxic entities and associated entities in the input text.
11. A communication device comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is configured to read the program in the memory to implement the steps of the text processing method according to any one of claims 1 to 8.
12. A computer-readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the steps of the text processing method according to any one of claims 1 to 8 are implemented.
13. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps in the text processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text generation method and device, equipment and medium
CN114491077A
Entity relationship classification method and device, equipment and storage medium
CN115292504A
Text processing model training method and device, text rewriting method and device and storage medium
CN116894431A
Large model-based data processing method, apparatus and device, and storage medium
CN119851661A
Signature generation for multimedia deep-content-classification by a large-scale matching system and method thereof
US20090043818A1