Method and apparatus for processing prompt, and device and medium

By receiving and converting prompt words, the deviation problem of machine learning models when processing prompt words is solved, and the accuracy of answers is improved without modifying the model.

WO2025183624A1PCT designated stage Publication Date: 2025-09-04LEMON INC(GB)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/SG2024/050108
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Machine learning models may have biases when processing prompt words, resulting in the answers that do not match the real answers. Especially in areas involving social bias, the prior art is difficult to effectively reduce biases in answers.

Method used

By receiving the first prompt word, determine if there is a potential deviation between it and the real answer, and convert it to a second prompt word based on the associated keywords to adjust the output of the machine learning model.

Benefits of technology

Without modifying the machine learning model, adjust the prompt words to alleviate and eliminate bias in the answer and improve the accuracy of the answer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2024050108_04092025_PF_FP_ABST
    Figure SG2024050108_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for processing a prompt, and a device and a medium. One method comprises: receiving a first prompt, wherein the first prompt expresses, by means of a natural language, a question which is to be input to a machine learning model; determining whether there is a latent bias between an answer of the machine learning model with respect to the first prompt and a real answer to the question; in response to determining that there is the latent bias, on the basis of a keyword, which is associated with the latent bias, in the first prompt, converting the first prompt into a second prompt; and receiving an answer of the machine learning model with respect to the second prompt. By using the exemplary implementation of the present disclosure, when a machine learning model itself is not modified, a bias in an answer can be mitigated and / or eliminated by means of adjusting a prompt.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]TECHNICAL FIELD Exemplary implementations of the present disclosure generally relate to machine learning models, and more particularly to methods, apparatuses, devices, and computer-readable storage media for processing prompt words for machine learning models. Background: Machine learning technology has been widely used to perform various types of tasks. For example, questions can be posed to a machine learning model and answers received from the model. However, the training data for a machine learning model may include biased data, resulting in answers provided by the machine learning model not being completely consistent with the facts, and some answers may deviate from the true answers. In such cases, it is desirable to obtain true answers that are more consistent with the facts. SUMMARY In a first aspect of the present disclosure, a method for processing prompt words is provided. In this method, a first prompt word is received, wherein the first prompt word expresses a question to be input to a machine learning model in natural language. A determination is made as to whether there is a potential deviation between the machine learning model's answer to the first prompt word and the true answer to the question. In response to determining that the potential deviation exists, the first prompt word is converted into a second prompt word based on keywords in the first prompt word associated with the potential deviation. The answer of the machine learning model to the second prompt word is received. In a second aspect of the present disclosure, a device for processing prompt words is provided. The device includes: a first receiving module configured to receive a first prompt word, the first prompt word expressing a question in natural language to be input to a machine learning model; a determination module configured to determine whether there is a potential bias between the machine learning model's answer to the first prompt word and the true answer to the question; a conversion module configured to, in response to determining the presence of a potential bias, convert the first prompt word into a second prompt word based on keywords in the first prompt word associated with the potential bias; and a second receiving module configured to receive the machine learning model's answer to the second prompt word. In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first aspect of the present disclosure.In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program. When executed by a processor, the computer program causes the processor to implement the method according to the first aspect of the present disclosure. It should be understood that the content described in this summary section is not intended to define the key or important features of the implementations of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS The foregoing and other features, advantages, and aspects of various implementations of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings.In the accompanying drawings, identical or similar reference numerals represent identical or similar elements, wherein: FIG1A shows a block diagram of an application environment according to an exemplary implementation of the present disclosure; FIG1B shows a block diagram of determining an answer to a question based on different knowledge according to a technical solution; FIG2 shows a block diagram for processing prompt words according to some implementations of the present disclosure; FIG3 shows a block diagram for updating prompt words according to some implementations of the present disclosure; FIG4A shows a block diagram of training data according to some implementations of the present disclosure; FIG4B shows a block diagram of the relationship between various representations involved in a machine learning model according to some implementations of the present disclosure; FIG4C shows a block diagram of the relationship between various representations involved in a machine learning model according to some implementations of the present disclosure, wherein bold lines show paths in the machine learning model; FIG5A shows a block diagram of dependency relationships in a machine learning model according to some implementations of the present disclosure; FIG5B shows a block diagram of a path in a machine learning model for weakening a path in a machine learning model according to some implementations of the present disclosure; FIG6A shows a block diagram of dependency relationships in a machine learning model according to some implementations of the present disclosure; Figure 6B shows a block diagram for weakening paths in a machine learning model according to some implementations of the present disclosure; Figure 7A shows a block diagram of dependencies in a machine learning model according to some implementations of the present disclosure; Figure 7B shows a block diagram for weakening paths in a machine learning model according to some implementations of the present disclosure; Figure 8 shows a flow chart of a method for processing prompt words according to some implementations of the present disclosure; Figure 9 shows a block diagram of an apparatus for processing prompt words according to some implementations of the present disclosure; and Figure 10 shows a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS Implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain implementations of the present disclosure are illustrated in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.In the description of the implementations of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one implementation" or "the implementation" should be understood as "at least one implementation." The term "some implementations" should be understood as "at least some implementations." The following may also include other explicit and implicit definitions. As used herein, the term "model" can refer to the association relationship between various data. For example, the above association relationship can be obtained based on various technical solutions currently known and / or to be developed in the future. It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations, and relevant provisions. It is understood that before using the technical solutions disclosed in each embodiment of the present disclosure, the user should be informed of the type, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and the user's authorization should be obtained. For example, in response to a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of their personal information. This allows the user to autonomously choose whether to provide their personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations described in the technical solution of the present disclosure, based on the prompt message. As an optional but non-limiting implementation, the prompt message may be sent to the user in response to the user's active request, for example, via a pop-up window. The pop-up window may present the prompt message in text format. Furthermore, the pop-up window may also contain a selection control for the user to select "Agree" or "Disagree" to provide their personal information to the electronic device. It should be understood that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure. The term "in response to" as used herein refers to the occurrence of a corresponding event or the satisfaction of a condition. It will be understood that the timing of executing the subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met.For example, in some cases, a subsequent action may be executed immediately upon the occurrence of an event or the fulfillment of a condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the fulfillment of the condition. Example Environment: Machine learning technology has been widely used to perform various types of tasks. For example, questions can be posed to a machine learning model and answers received from the model. Machine learning model 120 may be, for example, a language model (such as a large language model). Currently, it has been proposed to train machine learning models using large-scale text corpora. However, machine learning models generated in this manner may exhibit social bias. Unchecked bias may perpetuate and exacerbate social inequality, or even lead to more serious consequences. With the development of machine learning model technology, machine learning models have been widely applied in fields requiring high accuracy, such as recruitment and healthcare. In such fields, it is particularly necessary to eliminate bias in machine learning models. An application environment according to an example implementation of the present disclosure is described with reference to FIG1A , which shows a block diagram 100A of the application environment according to an exemplary implementation of the present disclosure. As shown in FIG1 , a page 110 may be provided to allow users to interact with the machine learning model 120. For example, a user may ask a question 112 and receive an answer 114 from a machine learning model. Question 112 asks: In the sentence "The CEO hired a secretary because he received a strong recommendation," who does "he" refer to? In this case, the machine learning model may provide different answers. See FIG. 1B for more information, which shows a block diagram 100B for determining an answer to a question based on different knowledge, according to one technical solution. Question 112 may be answered based on different knowledge. For example, knowledge 120 may indicate that "he" is often associated with CEOs, not secretaries (based on occupational statistics). In this case, if the machine learning model answers question 112 based on knowledge 120, answer 122 may indicate "CEO," meaning that in the sentence, "he" refers to the CEO. For another example, knowledge 140 may indicate that "the CEO hired a secretary because he received a strong recommendation" is more likely to occur in real life than "the CEO hired a secretary because he received a strong recommendation."At this point, if the machine learning model answers question 112 based on knowledge 140, answer 142 may represent "secretary." That is, in the above sentence, "He" refers to the secretary. As can be seen, because the training data of different machine learning models may vary, the answers provided by some machine learning models may also differ, and some answers may deviate from the true answer. In other words, the machine learning model may provide biased answers. It should be understood that while the questions, answers, and knowledge described above use English as an example of a natural language, the questions, answers, and knowledge can alternatively and / or additionally be expressed in other languages ​​(e.g., Chinese, French, etc.). Alternatively and / or additionally, a translation model can be used to convert between different natural languages. Various technical solutions have been proposed, such as selecting different training data, fine-tuning model parameters, modifying decoding steps, etc. However, such solutions may involve a large amount of computation. Although currently proposed methods of reducing the possibility of bias in answers based on modifying prompt words have been proposed, the results of such solutions have not been satisfactory. At this point, it is hoped that a more factually accurate answer can be obtained. Overview of Prompt Word Processing: To at least partially address the deficiencies in the prior art, a method for processing prompt words is proposed according to an exemplary implementation of the present disclosure. This disclosure focuses on social bias, specifically the relationship between demographic information and the answers output by a machine learning model. An overview of an exemplary implementation of the present disclosure is described with reference to FIG2 , which shows a block diagram 200 for processing prompt words according to some implementations of the present disclosure. As shown in FIG2 , a first prompt word 210 may be received. The first prompt word 210 expresses a question to be input to the machine learning model in natural language. A determination may be made 212 as to whether there is a potential bias between the machine learning model's answer to the first prompt word 210 and the true answer to the question. OAccording to an exemplary implementation of the present disclosure, it is possible to determine whether the above-mentioned bias exists using a variety of methods currently known and / or to be developed in the future. For example, it is possible to determine whether the first prompt word includes keywords that may cause bias, etc., by using rules, text analysis, and / or models. If it is determined that there is a potential bias, the first prompt word 210 can be converted into a second prompt word 220 based on the keyword 214 associated with the potential bias in the first prompt word 210, and then receive the answer 240 of the machine learning model 220 for the second prompt word 220. Using the exemplary implementation of the present disclosure, the bias in the answer can be alleviated and / or eliminated by adjusting the prompt word without modifying the machine learning model itself. The prompt word can be updated based on the keyword that has a causal relationship with the potential bias. Furthermore, a framework for removing bias using causal relationships is proposed, which mainly involves the following points: (1) the generation process of the data of the training corpus fed to the machine learning model; and (2) the internal reasoning process of the machine learning model to establish the prompt word through a selection mechanism, thereby removing the bias from the output of the machine learning model. When updating prompt words, inhibitory instructions and contextual comparison examples can be utilized, along with unbiased reasoning to mitigate and / or eliminate bias. Extensive experiments on real-world datasets demonstrate that the proposed technical solution can effectively reduce bias in the output of machine learning models. The detailed process for processing prompt words has been described in an overview of an exemplary implementation according to this disclosure; more information on this process will be provided below. Within the context of this disclosure, prompt words can be modified to guide machine learning models to provide unbiased (or less biased) answers. It should be understood that unbiased answers essentially involve selecting appropriate components from the model's internal representations and knowledge. Regarding the aforementioned process of determining the referent of a pronoun in a sentence, the model may use shortcuts learned from the training dataset regarding gender and output biased answers. For example, in Figure 1B, the training data may include stereotypes and contain social biases regarding occupational gender distribution. The model may associate the pronoun "he" with "CEO" (rather than "secretary"). This decision-making process can be represented as biased reasoning, where the model selects internal representations inappropriately.In the context of this disclosure, selection mechanisms play a crucial role in the interplay between the internal reasoning process of a machine learning model and different designs of external prompts. In addition to the data generation process for the training corpus, a potential causal model of the machine learning model's reasoning process can be constructed. Cue words can be used to select different paths within the causal model, thereby causing the machine learning model to output different results. To this end, a causal relationship framework is proposed based on the following: a) reducing biased reasoning and b) encouraging unbiased reasoning. In summary, Figure 3 illustrates different updates based on cue words, which shows a block diagram 300 for updating cue words according to some implementations of this disclosure. As shown in Figure 3, update 310 involves suppressing instructions, implemented by reducing (e.g., denoted as "-") biased reasoning (explicit); update 320 involves contextual comparison examples, implemented by reducing biased reasoning (implicit) and increasing (e.g., denoted as "+") unbiased reasoning (implicit); and update 330 involves approaching the truth, implemented by increasing unbiased reasoning (explicit). According to an example implementation of the present disclosure, a detailed causal model is constructed for the data generation process of a training corpus and the inference process of a machine learning model. By modifying prompt words, different selection mechanisms are implemented in the machine learning model, which has a significant impact on controlling the output of the machine learning model. A framework for removing bias based on causal relationships is proposed, as well as specific modification patterns for adjusting prompt words. Experiments have shown that these modification patterns mitigate and / or eliminate various social biases, thereby improving the accuracy of the answers output by the machine learning model. First, a general description of causal relationships is given. For two random variables X and Y, if an intervention in X changes the distribution of Y while keeping all other variables fixed, then X is the cause of Y. Causal relationships between variables can be represented using a directed acyclic graph (DAG), where nodes represent variables and edges represent direct causal relationships between variables. According to an example implementation of the present disclosure, a debiasing framework is proposed that guides the causal relationships of the data generation process involved. Specifically, a detailed causal model is proposed for the basic data generation process of the training corpus and the inference process of the machine learning model.It should be understood that the inference process of a machine learning model is essentially an interaction between the model's internal representation and the conditions specified by external cues, with selection mechanisms playing a key role. Furthermore, different modification patterns can be used to modify cues to remove bias. Specifically, we propose causal modeling of the underlying data generation process of the training data corpus and the inference process of the machine learning model. The following describes the underlying data generation process of the text in the training data corpus. For historical reasons, the underlying training data corpus may contain biases, often involving selection mechanisms. For example, stereotypes in career choices are not due to a direct causal relationship between gender and occupation, but rather to the underlying selection mechanisms of the training data. Specifically, in the training data, among all possible combinations of gender (e.g., male and female) and occupation (e.g., CEO and secretary), CEOs are often associated with men, while secretaries are associated with women. This phenomenon occurs because the training data is a specific subset of the ideal data corpus that contains stereotypes. Figure 4A shows a block diagram 400A of training data according to some implementations of the present disclosure. As shown in Figure 4A , causal relationships in the basic data generation process of a training data corpus (e.g., text collected from the internet) can be modeled. Demographic information 410 may include, for example, gender, age, etc. Furthermore, causal relationships may also include other variables of interest. For example, scenario 412 may be used to represent the actual environment or context (e.g., healthcare, recruitment) that sets the context for the text. Entity 411 may be used to represent the participants or stakeholders involved in the scenario. For example, a doctor may be an entity in a healthcare scenario, and a secretary may be an entity in a recruitment scenario. The selection mechanism S explicitly models the association between demographic information and entities, which is a stereotype in natural language processing. There can be different types of text: demographically agnostic text 414, which refers to text that does not explicitly contain demographic information; fact-based derivative text 415, which refers to text that is derived from facts, such as inferences based on definitions, fact-checking questions and answers, and restatements that do not change the factual content; and demographically aware text 413, which refers to text that explicitly contains demographic information.The following describes causal modeling of the underlying data generation process and how to distinguish discrimination in a machine learning model's training data corpus. It is expected that a machine learning model can capture dependency patterns in a data corpus and thereby extract inherent relevant knowledge. It is assumed that the internal reasoning process of the machine learning model is similar to the underlying data generation process of the data corpus. Figure 4B shows a block diagram 400B illustrating the relationships between various representations involved in a machine learning model according to some implementations of the present disclosure. Prompt word 426 indicates a prompt word input by a user. Demographic representation 420, entity representation 421, and context representation 422 indicate internal representations within the machine learning model associated with demographic information, entity information, and context information, respectively. Furthermore, demographic-aware text representation 424 indicates the internal representation of demographic-aware text within the machine learning model, and demographic-agnostic fact representation indicates the internal representation of demographic-agnostic facts within the machine learning model. Model output 427 represents the potential output of the machine learning model. Figure 4B thus depicts the reasoning process of how a machine learning model generates output based on input prompts. This diagram illustrates the interplay between internal representations and conditions specified by external inputs, and how different prompts can modulate machine learning model outputs through the selection mechanism. The internal nodes of the machine learning model are invisible. Arrows in Figure 4B represent directed edges (i.e., there is a direct causal relationship between the nodes preceding and following the arrow), and dashed arrows represent the selection mechanism. Prompts are external inputs to the machine learning model. The internal knowledge and representations of the machine learning model preexist and are independent of the presence of specific prompts. Prompts cannot serve as indicators for causal intervention. Because internal nodes are invisible, it is impossible to set them to specific values ​​through hard intervention, nor can soft intervention alter the functional behavior of the causal module. However, adjusting prompts can directly alter the selection variable, "Proper Consideration of Prompts" (PPC). If the machine learning model is well trained and fine-tuned, the PPC can be assumed to always be determined when the model produces an output, and the "potential output of the machine learning model" is the actual output obtained from the model. Prompts can significantly influence the machine learning model inference process by specifying conditions in the selection mechanism.Although unrelated to machine learning model inference itself, the modeling foundation generation process shown in Figure 4A provides hints about local causal modules of interest. Updating hint words can adjust the influence of these modules, thereby achieving the goal of debiasing. Because the conditions specified by hint words always depend on when the machine learning model generates output, internal reasoning is controlled by external hint words. Guided by a causal understanding of the relevant data generation process, this disclosure proposes a hint-based machine learning model debiasing framework. The conditions that should be followed during the hint word update process are used to influence the machine learning model inference process through a selection mechanism, thereby effectively mitigating and / or eliminating bias in the machine learning model output. Figure 4C shows a block diagram 400C of the relationships between the various representations involved in a machine learning model according to some implementations of the present disclosure, where bold lines indicate paths within the machine learning model. Figure 4C shows an annotated version of the machine learning model inference process. Based on an understanding of the underlying generative process of the training data corpus and the assumption that the trained machine learning model captures the dependency patterns in the training data, the bold arrows in Figure 4C represent the unmodulated information flow from demographic representations to the machine learning model output during internal reasoning. As shown in the figure, arrows 431, 432, 433, 434, 435, 436, and 437 represent selection mechanisms, and arrows 441, 442, 443, 444, and 445 represent dependency relationships. Within the selection mechanism, thin-line arrows represent selection mechanisms that can be specified by external input prompts, while thick-line arrows identify selection mechanisms that may involve unbiased knowledge. According to the machine learning model inference process shown in FIG4C , there is an information flow from the demographic representation (upstream node) to the model output of the machine learning model (downstream node). To remove bias, prompt words should follow predetermined conditions and constraints. Specifically, three modification modes for eliminating bias in machine learning models are proposed. According to an exemplary implementation of the present disclosure, during the conversion of a first prompt word into a second prompt word, an additional prompt word can be generated based on a keyword associated with potential bias in the first prompt word. The first prompt word and the additional prompt word are then combined to generate the second prompt word. In the above example, "he" can represent a keyword.Furthermore, the first prompt word can be used as the main part of the prompt word, and additional prompt words can be used as constraints on the main part, so that the machine learning model can output more accurate answers. Specifically, this can be achieved by reducing biased reasoning (explicit), reducing biased reasoning (implicit), improving unbiased reasoning (implicit), and / or improving unbiased reasoning (explicit). Specifically, multiple methods for updating prompt words can be provided. According to an example implementation of the present disclosure, modification mode I can be provided, which promotes the use of demographically unknowable facts. The principle of this modification mode is to encourage the machine learning model to make greater use of demographically unknowable facts when generating output. Specifically, condition I indicates that in the presence of PPC and existing choices, the internal representation of demographically unknowable facts and demographic information should be conditionally independent. Figure 5A shows a block diagram 500A of dependencies in a machine learning model according to some implementations of the present disclosure. As shown in Figure 5A, "demographic representation" w"denotes demographic representation, while "demographic-agnostic fact representation" indicates representation of demographic-agnostic facts. When s=1 and PPC=1, there is a conditional independence relationship between the two. Figure 5B shows a block diagram 500B for weakening paths in a machine learning model according to some implementations of the present disclosure. As indicated by arrows 431 and 435 in Figure 5B, using modification mode I, a selection mechanism can be introduced using demographic information and demographic-agnostic facts. Furthermore, the dependency relationships indicated by reference numerals 520 and 522 can be weakened. According to an example implementation of the present disclosure, multiple entities corresponding to a keyword are determined based on a first prompt word; the keyword is replaced with each of the multiple entities to generate multiple candidate hypotheses; and additional prompt words are generated using the multiple candidate hypotheses. In the above example, the keyword is "he." In this case, the keyword can be replaced with each entity referred to by "he," such as CEO and secretary, to generate new sentences (i.e., candidate hypotheses). At this point, the keywords in the square brackets in the sentence are replaced with CEO and secretary, respectively. Sentence 1: "The CEO hired the secretary because [the CEO] is highly recommended." Sentence 2: "The CEO hired the secretary because [the secretary] is highly recommended." According to an example implementation of the present disclosure, a second prompt word can be created, and the second prompt word can instruct the machine learning model to provide an answer to the question based on multiple candidate hypotheses. For example, the updated prompt word can be as follows. Table 1 Example of updated prompt word Using example implementations of the present disclosure, when generating an answer, the machine learning model can select the more likely real-life scenario from Sentences 1 and 2, thereby providing an unbiased answer: "the secretary." According to one example implementation of the present disclosure, additional prompt words are generated to weaken the influence of network nodes associated with bias causes in the machine learning model. Specifically, the bias cause associated with the keyword can be determined, and then, based on the bias cause, additional prompt words are generated. The additional prompt words are used to suppress the machine learning model from activating network nodes associated with the bias cause. Using example implementations of the present disclosure, when generating answers using updated prompt words, the machine learning model can weaken nodes that may be related to outdated knowledge about the profession, thereby generating an unbiased answer. For example, prompt words can be updated based on the following modification patterns II and III. According to one example implementation of the present disclosure, modification pattern II can offset existing selection bias. The principle of this modification pattern is to directly offset the influence of existing knowledge that causes bias. Figure 6A shows a block diagram 600A of dependency relationships in a machine learning model according to some implementations of the present disclosure. As shown in Figure 6A, "demographic representation" indicates a demographic representation. u Entity representation indicates the representation of an entity. The upper portion of relationship 610 represents condition H.1: in the presence of PPC and the occurrence of selection S, demographic information and entity representation are conditionally independent. That is, in the case of S=1 and PPC=1, there is a conditional independence relationship between the two. Further, the lower portion of relationship 610 represents condition II.2: no new association is introduced between the internal representation of demographic information and demographic unknowable facts. Figure 6B shows a block diagram 600B for weakening paths in a machine learning model according to some implementations of the present disclosure. oCondition II.1 aims to offset the bias caused by the selection (S node) (arrows 436 and 437) by constraining the dependency between demographic information and entity representations (arrows 431 and 432). This means that the dependency relationship indicated by reference numerals 620 and 620 can be weakened. Condition II.2 serves as a monitoring measure to ensure that no new discrepancies are introduced downstream in the entity's causal relationship, thereby enabling the aforementioned offsetting action to be applied to the final output. According to an exemplary implementation of the present disclosure, when generating additional prompt words based on the cause of the deviation, multiple candidate facts associated with the cause of the deviation can be determined, and then additional prompt words can be generated based on these multiple candidate facts. In the example described above, the potential bias is caused by the distribution in occupational statistics. In this case, the following sentences can be generated, stating facts about occupational distribution: Sentence 3: "CEOs can both be male and female with equal probability" Sentence 4: "Secretaries can both be male and female with equal probability" ,, According to an example implementation of the present disclosure, a second prompt word may be generated including: and the second prompt word may instruct the machine learning model to provide an answer to the question based on multiple candidate facts. For example, the updated prompt word may be as follows. Table 2 Example of updated prompt word Using an example implementation of the present disclosure, when generating an answer, the machine learning model can select an answer from the CEO and secretary that better aligns with the sentence's grammar and logic, based on the facts provided by Sentences 3 and 4. In this case, the machine learning model can provide an unbiased answer: "the secretary." According to an example implementation of the present disclosure, Modification Mode III can move away from demographically-aware text. The principle behind this modification mode is to encourage the machine learning model to no longer utilize demographically-aware text when generating output. Figure 7A shows a block diagram 700A of dependencies within a machine learning model according to some implementations of the present disclosure. As shown in relationship 710 in FIG7A , "demographic representation" indicates a demographic representation, and "demographic-aware text representation" indicates a demographically aware text representation. Relationship 710 relates to Condition III: there should be a conditional independence relationship between demographically aware text and the internal representation of demographic information. That is, when S=1 and PPC=1, there is a conditional independence relationship between the two. FIG7B shows a block diagram 700B for weakening paths in a machine learning model according to some implementations of the present disclosure. As shown in FIG7B , Modification Mode III can utilize Condition III to specify a selection mechanism for the internal representations of demographic information and demographically aware text (e.g., arrows 431 and 434) to adjust the dependency relationship shown as reference numeral 720. According to an example implementation of the present disclosure, when generating additional prompt words based on the cause of the deviation, the type of knowledge associated with the cause of the deviation can be determined, and the additional prompt words can be generated based on the type of knowledge. In this way, bias in the answer can be reduced by disabling certain types of knowledge. In the example described above, the cause of potential bias lies in the distribution of occupational statistics. At this point, the following sentences can be generated, and the use of related knowledge is prohibited: Sentence 5: " Please don , t be biased or please do not use any gender-related information when making the decision" oAccording to an example implementation of the present disclosure, a second prompt word may be generated, which may instruct the machine learning model to provide an answer to the question based on knowledge of the prohibited use type. For example, the updated prompt word may be as follows. Table 3 Example of updated prompt word Using the example implementation of the present disclosure, when generating an answer, the machine learning model can use sentence 5 to prohibit the model from using knowledge about occupational distribution. In this case, the machine learning model can provide an unbiased answer: "the secretary." It should be understood that each modification mode described above has its own advantages. Modification mode 1 can enable the machine learning model to utilize demographically unknowable facts, potentially associating demographic information with the output through unregulated arrows 441 and 445. Modification mode II can adjust the dependencies of arrows 442, 443, 444, 436, and 437, eliminating explicit constraints involving arrows 441 and 445. Furthermore, modification mode III can prevent reference to demographically relevant text during reasoning. It should be understood that each modification mode has its own advantages and can be used individually and / or in combination to more comprehensively address social bias in machine learning models. According to an example implementation of the present disclosure, if it is determined that the potential bias does not exist, the first prompt word can be directly input into the machine learning model, and the machine learning model's answer to the first prompt word can be received. In this way, while confirming that the machine learning model will not introduce bias, the answer to the question can be obtained more quickly and efficiently. It should be understood that the present disclosure proposes a causal-guided and prompt-based debiasing framework for machine learning models. In particular, it addresses the key role of selection mechanisms in modeling bias in data corpora and proposes influencing machine learning model outputs by specifying different selection conditions on the machine learning model's internal representation. Guided by a causal understanding of this interaction, bias in machine learning model outputs can be more effectively mitigated and / or eliminated. Extensive empirical results demonstrate that the proposed technical solution can reduce bias in machine learning model outputs with minimal time and computational resource expenditure, without adjusting the parameters of the machine learning model itself. Example Process FIG7 shows a flowchart of a method 700 for processing prompt words according to some implementations of the present disclosure. At block 710, a first prompt word is received, which expresses the question to be input to the machine learning model in natural language.At block 720, a determination is made as to whether a potential bias exists between the machine learning model's answer to the first prompt and the true answer to the question. At block 730, in response to determining that a potential bias exists, the first prompt is converted into a second prompt based on keywords associated with the potential bias in the first prompt. At block 740, the machine learning model's answer to the second prompt is received. According to an example implementation of the present disclosure, converting the first prompt into the second prompt includes: generating an additional prompt based on keywords associated with the potential bias in the first prompt; and combining the first prompt and the additional prompt to generate the second prompt. According to an example implementation of the present disclosure, generating the additional prompt includes: determining multiple entities corresponding to the keywords based on the first prompt; replacing the keywords with the multiple entities to generate multiple candidate hypotheses; and generating the additional prompt using the multiple candidate hypotheses. According to an example implementation of the present disclosure, generating the second prompt includes: creating the second prompt to instruct the machine learning model to provide an answer to the question based on the multiple candidate hypotheses. According to an example implementation of the present disclosure, generating additional prompt words includes: determining a deviation cause associated with a keyword; and generating additional prompt words based on the deviation cause, the additional prompt words being used to inhibit a machine learning model from activating a network node associated with the deviation cause. According to an example implementation of the present disclosure, generating additional prompt words based on the deviation cause includes: determining multiple candidate facts associated with the deviation cause; and generating additional prompt words based on the multiple candidate facts. According to an example implementation of the present disclosure, generating second prompt words includes: creating the second prompt words to instruct the machine learning model to provide an answer to a question based on the multiple candidate facts. According to an example implementation of the present disclosure, generating additional prompt words based on the deviation cause includes: determining the type of knowledge associated with the deviation cause; and generating additional prompt words based on the type of knowledge. According to an example implementation of the present disclosure, generating second prompt words includes: creating the second prompt words to instruct the machine learning model to provide an answer to the question based on knowledge of a prohibited type. According to an example implementation of the present disclosure, the method further includes: receiving an answer from the machine learning model to the first prompt word in response to determining that potential bias is absent.FIG8 shows a block diagram of an apparatus 800 for processing prompt words according to some implementations of the present disclosure. The apparatus includes: a first receiving module 810 configured to receive a first prompt word, the first prompt word expressing a question to be input to a machine learning model in natural language; a determining module 820 configured to determine whether there is a potential bias between the machine learning model's answer to the first prompt word and the true answer to the question; a conversion module 830 configured to, in response to determining the presence of a potential bias, convert the first prompt word into a second prompt word based on keywords associated with the potential bias in the first prompt word; and a second receiving module 840 configured to receive the machine learning model's answer to the second prompt word. According to one example implementation of the present disclosure, the conversion module includes: an additional generation module configured to generate additional prompt words based on keywords associated with the potential bias in the first prompt word; and a combining module configured to combine the first prompt word and the additional prompt word to generate the second prompt word. According to an example implementation of the present disclosure, a generation module includes: an entity determination module configured to determine multiple entities corresponding to a keyword based on a first prompt word; a replacement module configured to replace the keyword with each of the multiple entities to generate multiple candidate hypotheses; and a first generation module configured to generate additional prompt words using the multiple candidate hypotheses. According to an example implementation of the present disclosure, the generation module includes: a first creation module configured to create a second prompt word to instruct a machine learning model to provide an answer to a question based on the multiple candidate hypotheses. According to an example implementation of the present disclosure, the additional generation module includes: a cause determination module configured to determine a deviation cause associated with the keyword; and a suppression module configured to generate additional prompt words based on the deviation cause, the additional prompt words being used to suppress the machine learning model from activating a network node associated with the deviation cause. According to an example implementation of the present disclosure, the suppression module includes: a fact determination module configured to determine multiple candidate facts associated with the deviation cause; and a second generation module configured to generate additional prompt words based on the multiple candidate facts.According to an example implementation of the present disclosure, the generation module includes: a second creation module configured to create a second prompt word to instruct the machine learning model to provide an answer to a question based on multiple candidate facts. According to an example implementation of the present disclosure, the additional generation module includes: a type determination module configured to determine the type of knowledge associated with the cause of the deviation; and a third generation module configured to generate additional prompt words based on the type of knowledge. According to an example implementation of the present disclosure, the generation module includes: a third creation module configured to create a second prompt word to instruct the machine learning model to provide an answer to the question based on prohibited-use knowledge. According to an example implementation of the present disclosure, the apparatus further includes: a third receiving module configured to receive an answer from the machine learning model to the first prompt word in response to determining that there is no potential deviation. Figure 10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above. As shown in FIG10 , computing device 1000 is a general-purpose computing device. Components of computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1020, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit 1010 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of computing device 1000. Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, and removable and non-removable media.Memory 1020 may be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1020 may be removable or non-removable and may include machine-readable media such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and accessed within computing device 1000. Computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10 , a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 can include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure. Communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of computing device 1000 can be implemented in a single computing cluster or multiple computing machines capable of communicating via a communication connection. Thus, computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes. Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a display, speaker, printer, etc. The computing device 1000 may also communicate with one or more external devices (not shown) through the communication unit 1040 as needed, such as storage devices, display devices, etc., with one or more devices that enable a user to interact with the computing device 1000, or with any device that enables the computing device 1000 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.).Such communication can be performed via an input / output (I / O) interface (not shown). According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored. The computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided. The computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions. The computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided. The computer program product has a computer program stored thereon. When executed by a processor, the program implements the method described above. Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine. When these instructions are executed by the processing unit of the computer or other programmable data processing device, they generate a device that implements the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium. These instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams. The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, causing the computer, other programmable data processing device, or other device to execute a series of operational steps to produce a computer-implemented process. Consequently, the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams.The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of various implementations of the systems, methods, and computer program products of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or portion of an instruction, each of which contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions indicated in the blocks may occur in a different order than indicated in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified functions or actions, or may be implemented using a combination of dedicated hardware and computer instructions. Various implementations of the present disclosure have been described above; however, the foregoing description is illustrative and not exhaustive, and is not intended to limit the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

Claims 1. A method for processing a prompt word, comprising: receiving a first prompt word, the first prompt word expressing a question to be input into the machine learning model in natural language; Determining whether there is a potential bias between the machine learning model's answer to the first prompt word and a true answer to the question; in response to determining that the potential bias exists, converting the first prompt word into a second prompt word based on keywords in the first prompt word that are associated with the potential bias; and receiving the machine learning model's answer to the second prompt word.

2. The method according to claim 1, wherein converting the first prompt word into the second prompt word comprises: generating an additional prompt word based on the keyword associated with the potential deviation in the first prompt word; and combining the first prompt word and the additional prompt word to generate the second prompt word.

3. The method according to claim 2, wherein generating the additional prompt word comprises: Determining a plurality of entities corresponding to the keyword based on the first prompt word; Replacing the keywords with the multiple entities respectively to generate multiple candidate hypotheses; and generating the additional prompt words using the multiple candidate hypotheses.

4. The method according to claim 2, wherein generating the second prompt word comprises: The second prompt word is created to instruct the machine learning model to provide the answer to the question based on the multiple candidate hypotheses.

5. The method according to claim 2, wherein generating the additional prompt word comprises: determining a cause of the deviation associated with the keyword; as well as Based on the cause of the deviation, the additional prompt word is generated, and the additional prompt word is used to inhibit the machine learning model from activating the network node associated with the cause of the deviation.

6. The method according to claim 5, wherein the error is generated based on the cause of the error. 23 The additional prompt words include: determining a plurality of candidate facts associated with the cause of the deviation; and generating the additional prompt words based on the plurality of candidate facts.

7. The method according to claim 6, wherein generating the second prompt word comprises: The second prompt word is created to instruct the machine learning model to provide the answer to the question based on the multiple candidate facts.

8. The method according to claim 5, wherein generating the additional prompt word based on the deviation cause comprises: Identify the type of knowledge associated with the cause of the deviation; and generating the additional prompt words based on the type of the knowledge.

9. The method according to claim 8, wherein generating the second prompt word comprises: The second prompt word is created to instruct the machine learning model to provide the answer to the question based on prohibiting the use of the type of knowledge.

10. The method according to claim 1, further comprising: In response to determining that the potential bias does not exist, receiving the answer of the machine learning model to the first prompt word.

11. A device for processing a prompt word, comprising: a first receiving module configured to receive a first prompt word, where the first prompt word expresses a question to be input into the machine learning model in natural language; A determination module is configured to determine whether there is a potential deviation between the answer of the machine learning model to the first prompt word and the true answer to the question; a conversion module is configured to, in response to determining that the potential deviation exists, convert the first prompt word into a second prompt word based on keywords in the first prompt word associated with the potential deviation; and a second receiving module is configured to receive the answer of the machine learning model to the second prompt word.

12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processor. The electronic device comprises a processing unit and stores instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, cause the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is enabled to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Construction method and application of question and answer interaction model with cognitive reasoning ability

    CN116991996A

  • Question and answer processing method and device, electronic equipment and storage medium

    CN117271730A

  • Adaptive prompt enhancement method for large-scale language model

    CN117391216A

  • Intelligent child rearing system and device based on natural language processing

    CN117453867A

  • Large model interaction processing method and system, terminal, equipment and medium

    CN117520497A