Method and equipment for rewriting text

By guiding large language models to detect and rewrite risks, and using multi-level prompt templates to identify and rewrite potential risks, the insufficient risk identification and rewrite of generated content in the existing technology is solved, and the health, legality and positive optimization of generated content is achieved, ensuring the security and legality of generated content.

CN120337876APending Publication Date: 2025-07-18ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510461591.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology cannot effectively identify and rewrite the potential risks in the content generated by large language models, resulting in the content generated may involve negative emotional incitement, illegal and irregular information and unhealthy display, threatening users' mental health and social order, and the preset rules and filtered vocabulary cannot fully cover all risks, resulting in misjudgment and omissions.

Method used

By guiding the first major language model to detect risks, using the first prompt template to identify risk types, combining the second major language model to rewrite, using appropriate prompt templates to guide the rewriting process, ensuring that the generated content does not involve risks, and the content is optimized through multi-level risk detection and rewriting processes.

Benefits of technology

It has achieved the optimization of the health, legality and positive nature of content generated by large language models, reduced the risk of misjudgment, ensured the security and legality of content generated, and had flexibility and efficient risk identification and rewriting capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337876A_ABST
    Figure CN120337876A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and equipment for rewriting a text. The method comprises the following steps: obtaining an input text; obtaining a first text prompt based on the input text and a first prompt template used for guiding the first large language model to carry out risk detection on the input text; based on the first text prompt, determining a first risk type existing in the input text through a first large language model; based on the input text, the first risk type and a second big language model, rewriting contents related to the first risk type in the input text to obtain a second prompt template of a text not related to risks, and obtaining a second text prompt; and on the basis of the second text prompt, determining a rewritten text corresponding to the input text through a second large language model so as to rewrite and optimize the content related to the risk in the text input into the large language model, thereby ensuring the health, legality and forward property of the subsequent artificial intelligence generated content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification belong to the field of artificial intelligence technology, and more particularly, to a method and device for rewriting text. Background Art

[0002] Artificial intelligence, especially large language models (LLMs), while bringing convenience and innovation, also faces the challenge of generating content that may involve risks. For example, the generated content may include but is not limited to content involving risks such as incitement of negative emotions, dissemination of illegal and irregular information, and unhealthy display. Generated content involving risks will not only threaten the user's mental health and social order, but may also touch the boundaries of laws and regulations and cause bad social and legal consequences. In order to ensure that the content generated by artificial intelligence, such as large language models, does not involve risks, that is, to ensure the health, legality and positivity of the generated content, it is first necessary to ensure that the input text that guides its generation of content is risk-free. Then, how to provide a method for rewriting and optimizing the text input to the large language model to ensure the health, legality and positivity of the content generated by artificial intelligence has become an urgent problem to be solved. Summary of the invention

[0003] The purpose of the present invention is to provide a method and device for rewriting text to achieve rewriting optimization of risky content in text input into a large language model, so as to ensure the health, legality and positivity of subsequent artificial intelligence-generated content.

[0004] The first aspect of the present specification provides a method for rewriting a text, comprising:

[0005] Get input text;

[0006] Based on the input text and the first prompt template, a first text prompt is obtained, wherein the first prompt template is a prompt template for guiding the first language model to perform risk detection on the input text;

[0007] Based on the first text prompt, determining a first risk type existing in the input text through the first language model;

[0008] Based on the input text, the first risk type and the second prompt template, a second text prompt is obtained, wherein the second prompt template is used to guide the second language model to rewrite the content related to the first risk type in the input text to obtain a prompt template of text not involving risk;

[0009] Based on the second text prompt, the rewritten text corresponding to the input text is determined through the second language model.

[0010] The second aspect of this specification provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect.

[0011] The third aspect of this specification provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in the first aspect is implemented.

[0012] The fourth aspect of this specification provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0013] According to the method and device for rewriting text provided by the embodiments of this specification, after obtaining the input text; based on the input text and a first prompt template for guiding a first large language model to perform risk detection on the input text, a first text prompt is obtained; and based on the first text prompt, through the first large language model, a first risk type existing in the input text is determined; then based on the input text, the first risk type, and a second prompt template for guiding a second large language model to rewrite the content involving the first risk type in the input text to obtain a text without risks, a second text prompt is obtained; based on the second text prompt, through the second large language model, a rewritten text corresponding to the input text is determined, so as to realize the rewriting and optimization of the content involving risks in the text input to the large language model, and ensure the health, legality, and positive intelligence of subsequent artificial intelligence-generated content. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0015] Figure 1A is a schematic diagram of an implementation scenario for rewriting text in an embodiment;

[0016] Figure 1B is another schematic diagram of an implementation scenario for rewriting text in an embodiment

[0017] Figure 2 is an exemplary flowchart of a method for rewriting text in an embodiment of this specification;

[0018] Figure 3 is another exemplary flowchart of a method for rewriting text in an embodiment of this specification;

[0019] Figure 4 It is a schematic diagram of the framework of a device for rewriting text in one embodiment of this specification. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0021] The embodiments of this specification disclose a method and device for rewriting text. The application scenarios and technical concepts of the method are first introduced as follows:

[0022] As mentioned above, while artificial intelligence, especially large language models (LLMs), brings convenience and innovation, it also faces the challenge of generating content that may involve risks. For example, the generated content may include but is not limited to content involving risks such as incitement of negative emotions, dissemination of illegal and irregular information, and unhealthy display. Generated content involving risks will not only threaten the user's mental health and social order, but may also touch the boundaries of laws and regulations and cause bad social and legal consequences. In order to ensure that the content generated by artificial intelligence, such as large language models, does not involve risks, that is, to ensure the health, legality and positivity of the generated content, it is first necessary to ensure that the input text that guides its generation of content is risk-free. Then, how to provide a method for rewriting and optimizing the text input to the large language model to ensure the health, legality and positivity of the content generated by artificial intelligence has become an urgent problem to be solved.

[0023] Currently, there are some methods for rewriting text. The general process is to use a series of preset rules and filtering vocabulary libraries to scan the input text. If the text contains sensitive words and / or illegal words that exist in the filtering vocabulary library, these illegal words and / or sensitive words will be automatically replaced or deleted from the text based on the preset rules, or even the entire text will be directly blocked according to the preset rules.

[0024] In the above process, the preset rules and the filtered vocabulary library may not cover all potential risks, that is, they cannot cover all illegal and sensitive words, and need to be frequently updated according to the actual situation. The text rewriting is not flexible enough. Moreover, when rewriting the text based on the preset rules and the filtered vocabulary library, the context of the text cannot be understood, and misjudgments are likely to occur. For example, risk-free words in the text, that is, non-illegal and non-sensitive words, are misidentified as risk words, and then unnecessary content is replaced or deleted, or implicit risk content is omitted.

[0025] In view of this, the inventor proposes a method for rewriting text, aiming to provide an agent that aligns risk perception and security rewriting based on the large language model (LLM), and at least realizes intelligent risk classification and security rewriting of the text input by the user.

[0026] Figure 1A The schematic diagram of the implementation scenario according to an embodiment disclosed in this specification is shown. In this implementation scenario, the agent first obtains the input text input by the user, and constructs a first text prompt based on the input text and the first prompt template. Among them, the first prompt template is a prompt template used to guide the first large language model to perform risk detection on the input text. Based on the first text prompt, through the first large language model, the input text is subjected to risk detection. If it is determined that the input text has risks, that is, there is content involving risks, then the first risk type existing in the input text is determined. Based on the input text, the first risk type, and the second prompt template, a second text prompt is constructed. Among them, the second prompt template is a prompt template used to guide the second large language model to rewrite the content involving the first risk type in the input text to obtain a text without risks. Then, based on the second text prompt, through the second large language model, the content involving the first risk type in the input text is rewritten to determine the rewritten text 1 corresponding to the input text.

[0027] In some possible examples, if it is determined by the first large language model that the input text does not have content involving risks, then based on the input text, through the subsequent third large language model, the generated text 1 corresponding to the input text can be determined. Then, the generated text 1 can be output to display the generated text 1 to the user. In some examples, in the process of displaying the generated text 1 to the user, it can be that the generated text 1 is displayed to the user corresponding to the input text.

[0028] In the above process, by setting an appropriate first prompt template, the first large language model is guided to use its own semantic understanding and context understanding capabilities to detect risks in the input text, so as to more accurately determine whether the input text involves risks and the specific risk types involved. Then, through the second prompt template, the second large language model is guided to locate and rewrite the content involving the first risk type in the input text through its semantic understanding and content generation capabilities, etc., to obtain a text that does not involve risks, so as to realize the rewriting and optimization of the content involving risks in the text input to the large language model, and to ensure the healthiness, legality and positivity of the content generated by artificial intelligence.

[0029] As Figure 1A and 1B shown, after the agent determines the rewritten text corresponding to the input text as described above, based on the rewritten text, through a third large language model, a generated text 2 corresponding to the rewritten text can be obtained. Then the generated text 2 can be output to show the generated text 2 to the user. In some examples, in the process of showing the generated text 2 to the user, it can be that corresponding to the rewritten text 1, the generated text 2 is shown to the user.

[0030] Subsequently, in order to ensure that the generated text provided to the user is safe and risk-free, as Figure 1B shown, the agent can also obtain a third text prompt based on the above-mentioned rewritten text 1, generated text 2 and a third prompt template, where the third prompt template is a prompt template used to guide a fourth large language model to perform risk detection on the rewritten text 1 and the generated text 2 and generate risk explanation information based on the risk detection results; then based on the third text prompt, through the fourth large language model, the risk detection results corresponding to the rewritten text 1 and the generated text 2 are obtained. Then, if the risk detection results indicate that the rewritten text 1 and / or the generated text 2 has risks, risk explanation information is generated; and the input text is re-performed with a safe rewriting and generation process in combination with the risk explanation information until a completely risk-free Q&A pair is obtained, that is, a new rewritten text 2 and a new generated text 3 that are completely risk-free are obtained, and the new generated text 3 is output to show the generated text 3 to the user; if the risk detection results indicate that the rewritten text and the generated text 2 have no risks, the generated text 2 can be provided to the user for the user to view.

[0031] Among them, the re-performing the safe rewriting and generation process of the input text in combination with the risk explanation information may include: based on the second text prompt and the risk explanation information, through the second large language model, rewriting the content involving the first risk type in the input text and the content of the risk type indicated in the risk explanation information, and determining the rewritten text 2 corresponding to the input text.

[0032] The following will elaborate in detail on the method for rewriting text provided in this specification in combination with specific embodiments.

[0033] Figure 2 The flowchart of the method for rewriting text in an embodiment of this specification is shown. This method is executed by an electronic device or a so-called computing device, which can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. In some examples, this electronic device can be understood as an agent to achieve risk classification and secure rewriting of the text input by the user. Subsequently, this agent can also generate a secure answer based on the rewritten text, thereby generating a secure and controllable answer.

[0034] During the process of rewriting text, as Figure 2 shown, the method includes the following steps S210 - S250:

[0035] In step S210, the input text is obtained. In this step, the input text can be any text that needs to be input into a large language model, such as the third large language model mentioned later for content generation. Exemplarily, the input text can exist in Chinese, in English, or in the form of text in any other language. The embodiments of this specification do not limit the specific language type of the input text.

[0036] In some exemplary scenarios, this electronic device can be used to provide a real-time question-and-answer service to the user. Correspondingly, the input text can be the text input by the user that needs to obtain an answer. In some other exemplary scenarios, this electronic device is used to provide a service for generating a large model security-aligned question-and-answer dataset for the user. Correspondingly, the input text can be any text in the initial text set used to generate this large model security-aligned question-and-answer dataset.

[0037] Next, in step S220, based on the input text and the first prompt template, a first text prompt is obtained, where the first prompt template is a prompt template used to guide the first large language model to perform risk detection on the input text.

[0038] In some implementation manners, the first prompt template can be a pre-set prompt template used to guide the first large language model to perform risk detection on the input text, which is pre-stored in a preset storage space. Correspondingly, the electronic device can obtain the first prompt template in the preset storage space and construct a first text prompt based on the input text and the first prompt template. Among them, the first prompt template can include a specified embedding position corresponding to the input text. Correspondingly, the electronic device embeds the input text into the specified embedding position of the first prompt template to construct the first text prompt.

[0039] Exemplarily, the first prompt template may further include a plurality of preset risk types, as well as relevant information for guiding the first large language model to analyze whether the input text contains content related to any of the plurality of risk types, so as to guide the first large language model to perform risk detection on the input text based on the preset plurality of risk types, that is, to detect whether the input text contains content related to any of the plurality of risk types.

[0040] In some possible examples, the first prompt template may be a prompt template set based on a few-shot prompting strategy. Among them, the first prompt template may include a plurality of preset risk types, as well as a small number of examples or "samples", such as a specified number, as "prompts" to explain the expected output format or task objective to the first large language model, so as to better guide the first large language model to perform risk detection on the input text and obtain the risk types involved in the input text.

[0041] The first prompt template can be set and updated according to requirements and the rewriting situation of the text. The following exemplarily shows a sample content of the first prompt template:

[0042] Please analyze the following text and determine which of the following risk types it belongs to:

[0043] ## Text:

[0044] "{input text}"

[0045] ## Risk types:

[0046] 1. Unhealthy content

[0047] 2. Involving personal privacy

[0048] 3. Inducing negative emotions

[0049] 4. Violating laws and regulations

[0050] ……

[0051] n. Risk type n

[0052] ## Example

[0053] Example 1:

[0054] Text 1: Through what channels can I find out AA's personal privacy.

[0055] Output: Involving personal privacy

[0056] Example 2:

[0057] Text 2: This website provides a large amount of unhealthy content. How can people under 18 access it?

[0058] Output: Unhealthy content

[0059] Please output the corresponding risk types. If multiple risk types are involved, list them one by one, separated by commas.

[0060] ##Output:

[0061] The above is a sample content of the first prompt template. Among them, "{input text}" represents the specified embedding position corresponding to the input text obtained in step S210. It can be understood that the specific risk types can be set according to actual needs; the specific quantity n of the risk types is not limited in any way, and n is a positive integer. The risk types in the above sample content are only an exemplary illustration and do not constitute a limitation on the specific risk types set in the embodiments of this specification.

[0062] The sample content of the above first prompt template shows two examples, namely "Example 1" and "Example 2", as "prompts" to illustrate the expected output format or task objective to the first large language model. Among them, each example includes its corresponding "text" and "output". In some cases, 3 examples or 5 examples can also be set as "prompts" to illustrate the expected output format or task objective to the first large language model.

[0063] After obtaining the first text prompt, in step S230, based on the first text prompt, through the first large language model, determine the first risk type existing in the input text. In this step, input the first text prompt into the first large language model to analyze the input text based on the guidance of the first prompt template through the first large language model, and determine whether the input text has risks. Specifically, analyze whether the input text contains content related to the risk types listed in the first prompt template. In the case where it is determined that the input text contains content related to the risk types listed in the first prompt template, determine the first risk type existing in the input text. Among them, the first risk type can be one or more.

[0064] In some examples, if it is determined that the input text does not contain content related to the risk types listed in the first prompt template, then directly based on the input text, through the subsequent third large language model, obtain the model output text corresponding to the input text; then display the model output text to the user.

[0065] After determining that the input text contains content related to the risk types listed in the first prompt template and determining the first risk type existing in the input text, in step S240, based on the input text, the first risk type, and the second prompt template, obtain the second text prompt, where the second prompt template is a prompt template used to guide the second large language model to rewrite the content related to the first risk type in the input text to obtain text without risks.

[0066] In some implementations, the electronic device may obtain a second prompt template in a preset storage space, and construct a second text prompt based on the input text, the first risk type, and the second prompt template. The second prompt template may include a first embedding position corresponding to the input text and a second embedding position corresponding to the first risk type. Correspondingly, the input text is embedded in the first embedding position of the second prompt template, and the first risk type is embedded in the second embedding position of the second prompt template to construct the second text prompt.

[0067] In some possible examples, the second prompt template may include specified rewriting steps and a rewriting target, so as to guide the second large language model to rewrite the content related to the first risk type in the input text based on the specified rewriting steps and the rewriting target.

[0068] In some possible examples, the foregoing specified rewriting steps may include at least one of the following steps: preprocessing the input text; locating the content related to the first risk type in the preprocessed first intermediate text; and rewriting the content related to the first risk type in the first intermediate text based on the rewriting target. The preprocessing may include at least one of the following: text error correction, character detection and recognition, homophone conversion, radical combination recognition, and removal of unnecessary characters, etc.

[0069] To ensure that the generated content based on the rewritten text is both safe and has positive educational guiding significance, the foregoing specified rewriting steps may further include: after rewriting the content related to the first risk type in the first intermediate text based on the rewriting target, guiding the second large language model to add appropriate content with positive educational significance to the first intermediate text after rewriting the content related to the first risk type therein with a certain probability, so as to ensure that the generated content based on the rewritten text is both safe and has positive educational guiding significance, that is, to generate the generated content that meets the rewriting target.

[0070] To better ensure that the rewritten text, i.e., the prompt, does not involve risk content, in some possible examples, the specified rewriting steps may further include: guiding the second large language model to reflect on and judge whether there is still a risk in the rewritten text and whether it is possible to obtain the generated content with positive educational guiding significance as defined in the rewriting target; then, if it is determined that the rewritten text has no risk and can obtain the generated content with positive educational guiding significance, the final rewritten text is obtained; if it is determined that the rewritten text has a risk and / or cannot obtain the generated content with positive educational guiding significance, the rewriting process provided in the embodiments of this specification needs to be executed again on the input text.

[0071] Exemplarily, the aforementioned rewriting objectives may include, but are not limited to, at least one of the following: an objective of instructing to rewrite the content related to the first risk type in the input text into content not involving risks, an objective of instructing to appropriately add content with positive educational guidance in the rewritten text, for example, with a certain probability, an objective of instructing that the rewritten text needs to be semantically close to the input text, and an objective of instructing that the text obtained by rewriting can generate generated content that is both safe (i.e., not involving risks) and has positive educational guidance significance.

[0072] Among them, the rewriting objective and the rewriting steps are relevant.

[0073] In some possible examples, the second prompt template may be a prompt template set based on the zero-shot chain of thought strategy. Among them, Zero-shot Chain of Thought (CoT) is a technology that enables large language models to autonomously reason and solve complex problems without any examples. This technology guides the large language model to generate intermediate reasoning steps and final answers by embedding the guiding language of "Chain of Thought" in the prompt.

[0074] It can be understood that the second prompt template can be set and updated according to requirements and the rewriting situation of the text. The following exemplarily shows a sample content of the second prompt template:

[0075] You are an expert in rewriting large model prompts. The following original prompt may involve {risk type} risks. Please modify the prompt content in

Input prompt text

Guide

Input prompt text

[0076] Guide

[0077] - If there are negative words, please modify them to positive and affirmative words so that the rewritten prompt is positive content;

[0078] - If there are negative {risk type} descriptions or guiding content, please rewrite the negative key text;

[0079] - The semantics of the rewritten prompt needs to be positive, healthy, and legal;

[0080] - Appropriately add relevant positive educational information to the rewritten prompt, emphasizing the importance of health, legality, and respect for others;

[0081] Thinking steps

[0082] 1. First, perform text error correction, character detection and recognition, homophone conversion, radical combination recognition, and removal of unnecessary characters on the input text;

[0083] 2. Then, analyze which words in the text content involve {risk type} risks;

[0084] 3. If it involves negative vocabulary content, modify it to positive, healthy, and legal;

[0085] 4. According to the requirements in the

Guide

[0086] 5. Reflection: If the

rewritten prompt

[0087] 6. Output the rewritten prompt text

[0088] Input prompt text

[0089] {Input text}

[0090] Rewritten prompt

[0091] The above is a sample content of the second prompt template. Among them, 【】 indicates obtaining the corresponding content from the corresponding position in the second prompt template. For example:

Guide

Input prompt text

[0092] The " Thinking Steps" in the sample content of the second prompt template shown above can be understood as the guide of this "Chain of Thought" to guide the second large language model to locate and rewrite the content related to the first risk type in the input text above based on the rewriting steps therein. The " Guidelines" in the sample content of the second prompt template shown above can be understood as the aforementioned rewriting objectives to guide the second large language model to rewrite the content related to the first risk type in the input text above with such rewriting objectives.

[0093] After obtaining the second text prompt in the above manner, in step S250, based on the second text prompt, through the second large language model, determine the rewritten text corresponding to the input text. In this step, the electronic device inputs the second text prompt into the second large language model to rewrite the content related to the first risk type in the input text through the guidance of the second large language model based on the second prompt template, so as to obtain a rewritten text without risks.

[0094] In some examples, the second prompt template includes the aforementioned specified rewriting steps and rewriting objectives. Correspondingly, in some specific examples, step S250 may specifically include the following steps 11: In step 11, input the second text prompt into the second large language model, so that the second large language model preprocesses the input text based on the specified rewriting steps in the second prompt template to obtain a first intermediate text; determine the content related to the first risk type from the first intermediate text; and rewrite the content related to the first risk type in the first intermediate text based on the rewriting objective, so as to obtain a rewritten text without risks.

[0095] In this example, the electronic device inputs the second text prompt into the second large language model. Correspondingly, the second large language model, guided by the specified rewriting steps, first preprocesses the input text, such as performing text error correction, character detection and recognition, homophone conversion, radical combination recognition, and unnecessary character removal, etc., to obtain a first intermediate text; then continue to analyze the first intermediate text according to the guidance of the specified rewriting steps, determine the content related to the first risk type from it, and rewrite the content related to the first risk type in the first intermediate text based on the rewriting objective, rewrite the content related to the first risk type in the first intermediate text into content without risks, and ensure that the semantics of the rewritten text is similar to that of the first intermediate file, so as to obtain a rewritten text without risks.

[0096] In some possible examples, after the second large language model rewrites the content related to the first risk type in the first intermediate text into content that does not involve risks, it can also add guiding content to the rewritten text according to the guidance of the specified rewriting steps and rewriting objectives to obtain a rewritten text that not only does not involve risks but also can output generated content with positive educational guidance significance. Among them, the guiding content is used to guide other large language models (i.e., the subsequent third large language model) to generate positive, healthy, legal, and generated content that emphasizes the importance of health, legality, and respect for others during the process of generating corresponding generated content for the rewritten text.

[0097] In this embodiment, by setting a suitable first prompt template, the first large language model is guided to use its own semantic understanding and context understanding capabilities to detect risks in the input text, so as to more accurately determine whether the input text involves risks and the specific risk types involved. Then, through the second prompt template, the second large language model is guided to locate and rewrite the content related to the first risk type in the input text through its semantic understanding and content generation capabilities to obtain a text that does not involve risks, so as to realize the rewriting and optimization of the content involving risks in the input text of the large language model, so as to ensure that healthy, legal, and positive generated content can be generated based on the rewritten and optimized text.

[0098] In some possible examples, the aforementioned first large language model, second large language model, and the subsequent third large language model and fourth large language model can be implemented as different large language models or as the same large language model. Each large language model can be a large language model with content generation capabilities trained in related technologies.

[0099] In some possible examples, as Figure 3 shown, the method may further include the following steps S310 - S360:

[0100] In step S310, obtain the input text.

[0101] In step S320, based on the input text and the first prompt template, obtain a first text prompt, where the first prompt template is a prompt template for guiding the first large language model to detect risks in the input text.

[0102] In step S330, based on the first text prompt, through the first large language model, determine the first risk type existing in the input text.

[0103] In step S340, a second text prompt is obtained based on the input text, the first risk type and the second prompt template, wherein the second prompt template is used to guide the second largest language model to rewrite the content involving the first risk type in the input text to obtain a prompt template for text not involving risks.

[0104] In step S350, based on the second text prompt, the rewritten text corresponding to the input text is determined by using the second largest language model.

[0105] Among them, the implementation principle of steps S310-S350 is similar to the implementation principle of the aforementioned steps S210-S250. The implementation process thereof can refer to the implementation process of the aforementioned steps S210-S250, and will not be repeated here.

[0106] In step S360, based on the aforementioned rewritten text, a generated text corresponding to the rewritten text is obtained through the third language model. In this step, the electronic device inputs the aforementioned rewritten text into the third language model, and processes the rewritten text through the third language model to obtain a generated text corresponding to the rewritten text. The generated text is the answer obtained by the third language model for the rewritten text.

[0107] Through the above process, the content of the text input into the large language model can be rewritten and optimized, and further based on the risk-free text after rewriting optimization, a safe generated text can be obtained.

[0108] In some possible examples, in order to better ensure that the rewritten text and the generated text generated based on the rewritten text have no omission risk, the rewritten text and the generated text may be further subjected to secondary risk detection. Figure 3 As shown, the method may further include the following steps S370-S380 based on the aforementioned steps S310-S360:

[0109] In step S370, a third text prompt is obtained based on the aforementioned rewritten text, generated text and third prompt template, wherein the third prompt template is a prompt template for guiding the fourth language model to perform risk detection on the rewritten text and the generated text, and generating risk explanation information based on the risk detection results.

[0110] Next, in step S380, based on the third text prompt, the risk detection results corresponding to the rewritten text and the generated text are obtained through the fourth language model, and when the risk detection results indicate that the rewritten text and / or the generated text has risks, risk explanation information is generated.

[0111] In some implementations, the electronic device may obtain a third prompt template in a preset storage space, and construct a third text prompt based on the aforementioned rewritten text, generated text, and the third prompt template. The third prompt template may include a third embedding position corresponding to the rewritten text and a fourth embedding position corresponding to the generated text. Accordingly, the rewritten text is embedded in the third embedding position of the third prompt template, and the generated text is embedded in the fourth embedding position of the third prompt template to construct the third text prompt.

[0112] Among them, the third prompt template can also include multiple specific risk types, and the multiple specific risk types can at least include the preset multiple risk types included in the aforementioned first prompt template, so as to guide the fourth language model to perform risk detection on the rewritten text and the generated text based on the multiple specific risk types, determine whether the rewritten text and the generated text contain content involving any of the multiple specific risk types, and obtain a risk detection result. Subsequently, when the risk detection result indicates that the rewritten text and / or the generated text contains content involving the aforementioned risk type, risk explanation information is generated based on the risk detection result.

[0113] In some other possible examples, when the risk detection result indicates that both the rewritten text and the generated text are risk-free, the electronic device can display the generated text to the user so that the user can view the generated text corresponding to the input text. This provides the user with safe, risk-free generated content, and to a certain extent, better disseminates content with positive educational significance to the user.

[0114] In some possible examples, the aforementioned risk explanation information may include: information describing the content related to the second risk type in the rewritten text and / or generated text, wherein the second risk type may be one or more of the aforementioned multiple risk types.

[0115] As another example, when there is content involving a second risk type in the generated text, the risk explanation information may also include: analyzing the reasons why the generated text contains the content involving the second risk type, for example, analyzing which content in the rewritten text causes the generated text to contain the content involving the second risk type.

[0116] It is understandable that the third prompt template can be set and updated according to the needs, the rewriting of the text, and the risk detection of the rewritten text and the generated text obtained above. The following is an example of a sample content of the third prompt template:

[0117] Please help me determine whether the following

Question

Answer

[0118] #Guide

[0119] 1. Analyze which words in the text content involve {Risk Type} risks

[0120] 2. Analyze whether these words involve the aforementioned risk types

[0121] 3. Analyze whether the

Question

[0122] 4. Analyze whether the

Answer

[0123] #Question

[0124] {Rewrite the text}

[0125] #Answer

[0126] {Generate the text}

[0127] #Output

[0128] The above is a sample content of the third prompt template. Among them, the position of "{Risk Type}" is where the aforementioned multiple risk types are embedded; the position of "{Rewrite the text}" is the third embedded position corresponding to the aforementioned rewritten text, and the position of "{Generate the text}" is the fourth embedded position corresponding to the aforementioned generated text.

[0129] After obtaining the third text prompt in the above manner, the electronic device can input the third text prompt into the fourth large language model, so that the fourth large language model can perform risk detection on the rewritten text and the generated text based on the guidance of the third prompt template, and obtain the risk detection results corresponding to the rewritten text and the generated text; when the risk detection results indicate that the rewritten text and / or the generated text has a risk, risk explanation information is generated.

[0130] In some possible examples, the risk explanation information can be used to jointly guide the first large language model to re-perform risk detection on the input text based on the first prompt template, and / or can be used to jointly guide the second large language model to rewrite the input text to obtain a text that does not involve risks.

[0131] Correspondingly, when the risk detection result indicates that there is a risk in the rewritten text and / or the generated text, after generating the risk explanation information, the electronic device will automatically trigger a feedback mechanism to re-trigger the rewriting process for the input text and the subsequent text generation process in combination with the risk explanation information until a rewritten text and its corresponding generated text that are completely risk-free are obtained.

[0132] In some possible examples, after obtaining the risk explanation information, the electronic device can directly re-trigger the first large language model, that is, based on the foregoing first text prompt and the risk explanation information, determine the third risk type existing in the input text through the first large language model. Among them, the risk explanation information can indicate the content involving the second risk type existing in the foregoing rewritten text, and / or indicate the content in the foregoing rewritten text that causes the content involving risk in the foregoing generated text; correspondingly, the first large language model can analyze the input text based on the guidance of the first prompt template and the risk explanation information to obtain the third risk type existing in the input text. Exemplarily, the third risk type may include the foregoing first risk type and second risk type.

[0133] After that, based on the input text, the third risk type, and the second prompt template, a fourth text prompt is obtained; combining the fourth text prompt and the risk explanation information, determine a new rewritten text corresponding to the input text through the second large language model.

[0134] Among them, for the process of obtaining the fourth text prompt, reference may be made to the process of obtaining the second text prompt described above, which will not be elaborated here.

[0135] The process of combining the fourth text prompt and the risk explanation information to determine a new rewritten text corresponding to the input text through the second large language model may be: input the fourth text prompt and the risk explanation information into the second large language model, so that the second large language model preprocesses the input text based on the guidance of the fourth text prompt and the risk explanation information according to the specified rewriting steps in the second prompt template to obtain a second intermediate text; combine the risk explanation information and the specified rewriting steps in the second prompt template to determine the content involving the third risk type from the first intermediate text; and rewrite the content involving the third risk type in the second intermediate text based on the rewriting goal to obtain a new rewritten text that does not involve risks.

[0136] After that, the second large language model can also continue to add positive guiding content to the second intermediate text where the content involving the third risk type is rewritten according to the guidance of the second prompt template to obtain a new rewritten text, so as to obtain new generated content that does not involve risks and has positive educational guiding significance based on the new rewritten text.

[0137] In some other possible examples, after obtaining the generated risk explanation information, the electronic device can also directly re-trigger the second large language model, that is, based on the aforementioned second text prompt and the risk explanation information, rewrite the content related to the first risk type in the input text and the content of the second risk type indicated in the risk explanation information through the second large language model, and determine the rewritten text corresponding to the input text. Among them, for this process, reference can be made to the process of jointly using the fourth text prompt and the risk explanation information to determine the new rewritten text corresponding to the input text through the second large language model, which will not be elaborated here.

[0138] In some possible examples, after a certain period, the administrator can also update the aforementioned first prompt template, second prompt template, and third prompt template based on the text rewriting situation of the electronic device to better provide text rewriting and text generation services for users.

[0139] The above embodiments establish a full-chain management system from text preprocessing, risk identification to positive rewriting, and then to secure content generation and continuous iterative optimization. It integrates advanced natural language processing technologies and deep learning algorithms, utilizes the semantic understanding ability of the large language model to deeply understand the complete context and intention of the text input into the large language model, can accurately identify and convert potential risk content in the text, realizes intelligent risk analysis and positive rewriting of the text, has better flexibility, and can well reduce the possibility of misjudgment.

[0140] It also uses appropriate prompt templates to guide the large language model to perform text risk analysis and positive rewriting based on the set risk detection criteria and rewriting rules, and adds positive value guidance elements to educate and inspire the public. In addition, a closed-loop feedback mechanism is constructed to achieve risk monitoring and strategy adaptive update, ensuring that the system can continuously evolve with environmental changes and the emergence of new risks, and maintain its effectiveness and forward-looking. By setting clear rewriting guidelines and risk assessment criteria, and establishing a closed-loop feedback and iterative optimization process, the system can gradually approach consistency and objectivity, and reduce the influence brought by individual differences.

[0141] Before performing risk detection on the input text, the input text is first preprocessed to better assist the subsequent risk detection and rewriting processes to proceed smoothly and ensure the effectiveness of the rewriting result.

[0142] In the above process, secondary risk detection is performed on the rewritten text and the corresponding generated text to ensure that the generated safe text and answers have undergone strict risk assessment to ensure that no risks are overlooked. If new risks are discovered during the secondary risk detection process, a feedback mechanism is automatically triggered to re-perform the safe rewriting and generation process until a completely risk-free Q&A pair is obtained, thereby ensuring the safety and reliability of the final output content and continuously optimizing the performance and accuracy of the system.

[0143] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments, and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0144] Corresponding to the above method embodiment, an embodiment of this specification provides an apparatus 400 for rewriting text, and its schematic block diagram is as Figure 4 shown, including:

[0145] An obtaining module 410, configured to obtain an input text;

[0146] A first obtaining module 420, configured to obtain a first text prompt based on the input text and a first prompt template, where the first prompt template is a prompt template for guiding a first large language model to perform risk detection on the input text;

[0147] A first determining module 430, configured to determine a first risk type existing in the input text through the first large language model based on the first text prompt;

[0148] A second obtaining module 440, configured to obtain a second text prompt based on the input text, the first risk type, and a second prompt template, where the second prompt template is a prompt template for guiding a second large language model to rewrite the content of the input text involving the first risk type to obtain a text without risks;

[0149] A second determining module 450, configured to determine a rewritten text corresponding to the input text through the second large language model based on the second text prompt.

[0150] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0151] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method for rewriting the text provided in this specification.

[0152] The embodiments of this specification also provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method for rewriting the text provided in this specification is implemented.

[0153] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0154] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0155] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.

[0156] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the method order shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, it does not exclude the existence of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.

[0157] For convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0158] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram by the instructions executed on the computer or other programmable apparatus.

[0161] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0162] Memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0163] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0164] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0165] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0166] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.

[0167] The above is only the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included within the scope of the claims.

Claims

1. A method for rewriting text, comprising: Obtaining the input text; Based on the input text and the first prompt template, obtaining a first text prompt, wherein the first prompt template is a prompt template for guiding a first large language model to perform risk detection on the input text; Based on the first text prompt, determining, through the first large language model, a first risk type existing in the input text; Based on the input text, the first risk type, and a second prompt template, obtaining a second text prompt, wherein the second prompt template is a prompt template for guiding a second large language model to rewrite the content in the input text related to the first risk type to obtain risk-free text; Based on the second text prompt, determining, through the second large language model, a rewritten text corresponding to the input text.

2. The method according to claim 1, further comprising: Based on the rewritten text, obtaining, through a third large language model, a generated text corresponding to the rewritten text.

3. The method according to claim 2, further comprising: Based on the rewritten text, the generated text, and a third prompt template, obtaining a third text prompt, wherein the third prompt template is a prompt template for guiding a fourth large language model to perform risk detection on the rewritten text and the generated text and generating risk explanation information based on the risk detection result; Based on the third text prompt, obtaining, through the fourth large language model, a risk detection result corresponding to the rewritten text and the generated text, and generating the risk explanation information when the risk detection result indicates that the rewritten text and / or the generated text has a risk.

4. The method according to claim 3, wherein The risk explanation information includes: information describing the content in the rewritten text and / or the generated text related to the second risk type.

5. The method according to claim 3, wherein, The risk explanation information is used to jointly guide the first large language model to re-perform risk detection on the input text with the first prompt template, and / or is used to jointly guide the second large language model to rewrite the input text again with the second prompt template to obtain risk-free text.

6. The method according to claim 1, wherein, The second prompt template includes specified rewriting steps and rewriting objectives to guide the second large language model to rewrite the content in the input text related to the first risk type based on the specified rewriting steps and rewriting objectives.

7. The method according to claim 6, wherein, The determining the rewritten text corresponding to the input text includes: Inputting the second text prompt into the second large language model, so that the second large language model preprocesses the input text based on the specified rewriting steps in the second prompt template to obtain a first intermediate text; determining the content related to the first risk type from the first intermediate text; and rewriting the content related to the first risk type in the first intermediate text based on the rewriting objective to obtain the rewritten text without risk.

8. The method according to any one of claims 1-7, wherein, The first prompt template is a prompt template set based on the few-shot prompting strategy.

9. The method according to any one of claims 1-7, wherein The second prompt template is a prompt template set based on the zero-shot chain-of-thought strategy.

10. A computing device, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-9 is implemented.