Illusion mitigation for generative transducer models
By introducing a natural language inference (NLI) scoring system into the natural language generation model, evaluating and adjusting the token sorting and confidence level, the hallucination problem when generating text is solved, achieving a more realistic and reliable output text.
Patent Information
- Application Number
- CN202380072502.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2023-09-19
- Publication Date
- 2025-05-16
AI Technical Summary
Existing natural language generation models are prone to hallucinations when generating text, resulting in untrue ‘facts’ being included in the output text.
Natural Language Inference (NLI) scoring system is used to evaluate the faithfulness of the generated tokens, and output text that is more faithful to the input content is generated by adjusting the token sorting and confidence level.
It effectively alleviates the hallucination problem, ensures that the generated output text is more realistic and reliable, and improves the accuracy of natural language processing.
Smart Images

Figure CN120019379A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to natural language processing. For example, aspects of the present disclosure relate to systems and techniques for generating and using natural language generation models that mitigate hallucinations or situations in which the natural language generation model becomes convinced of untrue facts and generates text or speech based on untrue facts. Background Art
[0002] Machine learning models (e.g., deep learning models such as neural networks) can be used to perform a variety of tasks, including depth estimation, detection and / or recognition (e.g., scene or object detection and / or recognition), pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, image processing, etc. Machine learning models can be general and can achieve high-quality results in a variety of tasks. Summary of the invention
[0003] This paper describes systems and techniques for generating output text based on input content using natural language generation. In some examples, the system and technology are configured to use greedy search, beam search, or a combination thereof to search for possible tokens (e.g., words or parts thereof) to be used in the output text, such as taking into account previously generated words in the output text and / or taking into account the input content, and sorting these possible tokens based on the probability that the token will be used. The system and technology are configured to include a natural language inference (NLI) scoring system that generates an NLI score for a given possible token to identify the fidelity of the token to the input content, such as determining whether the use of the token in the output text results in a true, false, or neutral (e.g., undetermined) statement based on the input content. The system and technology can reorder the possible token based on the NLI score, or can factor the NLI score into the sorting of the possible token in other ways. The system and technology can select a token based on the sorting to generate the output text based on the sorting. By using the NLI scoring system, the systems and techniques are configured to mitigate hallucinations (eg, "facts" in the output text that are not true based on the input content).
[0004] Systems and techniques for natural language processing are provided. The system generates a plurality of tokens (e.g., words or parts thereof) based on input content (e.g., text and / or speech). The system searches the plurality of tokens to generate a first ordering of the plurality of tokens based on probability. The system generates a natural language inference (NLI) score for the plurality of tokens to generate a second ordering of the plurality of tokens based on fidelity to the input content (e.g., whether the tokens generate true statements based on the input content). The system generates an output text comprising at least one token selected from the plurality of tokens based on the first ordering and the second ordering.
[0005] According to at least one example, a method for natural language processing is provided. The method implemented by the processor includes: generating a token sequence based on input content; determining a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; generating a complete sentence including the token sequence; generating a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and adjusting the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence.
[0006] In another example, a device for natural language processing is provided, the device comprising at least one memory and at least one processor, the at least one processor being coupled to the at least one memory. The at least one processor is configured to: generate a token sequence based on input content; determine a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; generate a complete sentence including the token sequence; generate a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and adjust the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence.
[0007] In another example, a non-transitory computer-readable medium is provided, which has instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: generate a token sequence based on input content; determine a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; generate a complete sentence including the token sequence; generate a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and adjust the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level for the token sequence.
[0008] In another example, a device for natural language processing is provided. The device includes: a component for generating a token sequence based on input content; a component for determining a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; a component for generating a complete sentence including the token sequence; a component for generating a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and a component for adjusting the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence.
[0009] In some aspects, one or more of the methods, devices, and computer-readable media described above further include: generating the token sequence using a beam search based on the input content. In some aspects, one or more of the methods, devices, and computer-readable media described above further include: generating the complete sentence using a greedy search based on the token sequence.
[0010] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include: limiting the candidate token for use in generating the complete sentence based on whether the corresponding significance value of the candidate token exceeds a significance threshold. In some aspects, the significance threshold is based on an average of the corresponding significance values of the candidate token.
[0011] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: sorting the token sequence against the second token sequence based on the confidence level associated with the token sequence and the second confidence level associated with the second token sequence. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: re-sorting the token sequence against the second token sequence based on the updated confidence level associated with the token sequence and the second updated confidence level associated with the second token sequence, wherein the second updated confidence level is based on the second NLI score of the second complete sentence, and the second complete sentence is generated based on the second token sequence. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: re-sorting the token sequence against the second token sequence, selecting the highest ranked token sequence from at least the token sequence and the second token sequence based on the re-sorting of the token sequence against the second token sequence; and generating output text including the highest ranked token sequence. In some aspects, the output text is configured to generate a summary of the input content.
[0012] In some aspects, one or more of the methods, devices, and computer-readable media described above further include: generating an output text including the token sequence based on the updated confidence level of the token sequence exceeding the second updated confidence level of the second token sequence. In some aspects, one or more of the methods, devices, and computer-readable media described above further include: generating the second token sequence based on the input content; determining the second confidence level associated with the second token sequence based on the secondary corresponding confidence level associated with each token in the second token sequence; generating a second complete sentence including the second token sequence; generating a second NLI score for the second complete sentence based on the fidelity of the second complete sentence to the input content; and adjusting the second confidence level of the second token sequence based on the second NLI score of the second complete sentence to generate the second updated confidence level of the second token sequence. In some aspects, the output text is configured to perform summary generation on the input content.
[0013] In some aspects, the NLI score identifies whether at least a portion of the complete sentence is true, false, or neutral.
[0014] In some aspects, the input content comprises input text.In some aspects, each token in the sequence of tokens is at least a portion of a corresponding word.
[0015] In some aspects, the token sequence is configured to follow a previously determined token sequence in the complete sentence, wherein the complete sentence includes the previously determined token sequence, the token sequence, and at least one additional token.
[0016] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further comprises: generating the token sequence using a greedy search based on the input content.
[0017] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: outputting output text including the token sequence. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: causing a display to display the output text including the token sequence. In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: causing a communication interface to send the output text including the token sequence to a recipient device.
[0018] In some aspects, one or more of the devices described herein are and / or include and / or are part of an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device or wireless communication device (e.g., a mobile phone or other mobile device), a wearable device (e.g., a networked watch or other wearable device), a camera, a personal computer, a laptop, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-mentioned device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).
[0019] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0020] The foregoing and other features and embodiments will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Exemplary embodiments of the present application are described in detail below with reference to the following drawings:
[0022] Figure 1 is a conceptual diagram illustrating natural language processing (NLP) system technology according to some examples;
[0023] Figure 2 is a conceptual diagram illustrating an example of hallucination in a chatbot using natural language generation (NLG) according to some examples;
[0024] Figure 3A is a block diagram of a natural language generation (NLG) system according to some examples;
[0025] Figure 3Bis a block diagram of a natural language generation (NLG) system with a natural language inference (NLI) scoring system indicating faithfulness to input text according to some examples;
[0026] Figure 4A is a conceptual diagram of a greedy search decoding algorithm for a natural language generation (NLG) system according to some examples;
[0027] Figure 4B is a conceptual diagram of a beam search decoding algorithm for a natural language generation (NLG) system according to some examples;
[0028] Figure 5 is a conceptual diagram illustrating a histogram of entailment scores or natural language inference (NLI) scores indicating faithfulness to input content for output text with and without hallucinations according to some examples;
[0029] Figure 6 is a block diagram of a decoder with beam search and a natural language inference (NLI) scorer for a natural language generation (NLG) system according to some examples;
[0030] Fig. 7A is a block diagram of a decoder with greedy expansion and a natural language inference (NLI) scorer for a natural language generation (NLG) system according to some examples;
[0031] Figure 7B is a block diagram of a decoder with saliency-enhanced greedy expansion and a natural language inference (NLI) scorer for a natural language generation (NLG) system according to some examples;
[0032] Fig. 8A is a conceptual diagram illustrating examples of different output text strings having different natural language inference (NLI) scores according to some examples;
[0033] Figure 8B is a conceptual diagram illustrating examples of different output text strings having different natural language inference (NLI) scores according to some examples;
[0034] Fig. 9 is a conceptual diagram illustrating a model for generating a summary of input content using a natural language generation (NLG) system according to some examples;
[0035] Fig.10 is a flow chart illustrating an example process for natural language generation (NLG) according to aspects of the present disclosure;
[0036] Fig.11 is a block diagram illustrating an example of a deep learning network according to some examples; and
[0037] Fig.12 is a diagram illustrating an example system architecture for implementing certain aspects described herein. DETAILED DESCRIPTION
[0038] Some aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, which is obvious to those skilled in the art. In the following description, specific details are set forth for explanation purposes in order to provide a thorough understanding of each embodiment of the application. However, it will be apparent that each embodiment can be put into practice without these specific details. Each drawing and description are not intended to be restrictive.
[0039] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the subsequent description of the exemplary embodiments will provide an enabling description for implementing the exemplary embodiments to those skilled in the art. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the scope of the present application as set forth in the appended claims.
[0040] As described above, machine learning systems (e.g., deep neural network systems or models) can be used to perform various tasks, such as (for example, but not limited to) detection and / or recognition (e.g., scene or object detection and / or recognition, face detection and / or recognition, etc.), depth estimation, pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, and image processing, etc. In addition, machine learning models can be general and can achieve high-quality results in a variety of tasks.
[0041] In some examples, machine learning systems can be used for natural language processing (NLP) tasks, such as natural language understanding (NLU) and / or natural language generation (NLG). Examples of natural language generation include systems that use trained machine learning models to generate summaries of articles or other input content, chatbots, auto-completion systems, etc. In some cases, the NLG model may generate text that contains hallucinations or situations where the NLG model becomes convinced of untrue facts and generates text or speech based on untrue facts. For example, when trying to generate a summary of a news article about a car accident involving multiple people, the NLG model may produce hallucinations and incorrectly indicate in the output text that someone died in the accident, when in fact no one died in the accident.
[0042] This paper describes systems and techniques for generating output text based on input content using natural language generation. In some examples, the system and technology are configured to use greedy search, beam search, or a combination thereof to search for possible tokens (e.g., words or parts thereof) to be used in the output text, such as taking into account previously generated words in the output text and / or taking into account the input content, and sorting these possible tokens based on the probability that the token will be used. The system and technology are configured to include a natural language inference (NLI) scoring system that generates an NLI score for a given possible token to identify the fidelity of the token to the input content, such as determining whether the use of the token in the output text results in a true, false, or neutral (e.g., undetermined) statement based on the input content. The system and technology can reorder the possible token based on the NLI score, or can factor the NLI score into the sorting of the possible token in other ways. The system and technology can select a token based on the sorting to generate the output text based on the sorting. By using the NLI scoring system, the systems and techniques are configured to mitigate hallucinations (eg, "facts" in the output text that are not true based on the input content).
[0043] Systems and techniques for natural language processing are provided. The system generates a plurality of tokens (e.g., words or parts thereof) based on input content (e.g., text and / or speech). The system searches the plurality of tokens to generate a first ordering of the plurality of tokens based on probability. The system generates a natural language inference (NLI) score for the plurality of tokens to generate a second ordering of the plurality of tokens based on fidelity to the input content (e.g., whether the tokens generate true statements based on the input content). The system generates an output text comprising at least one token selected from the plurality of tokens based on the first ordering and the second ordering.
[0044] Figure 11 is a conceptual diagram 100 illustrating natural language processing (NLP) system technology. Natural language processing (NLP) 102 is useful in a variety of fields, such as the Internet of Things (IoT), wearable devices, cloud computing, software as a service, search engines, data querying, or a combination thereof. NLP 102 includes natural language understanding (NLU) 104 and natural language generation (NLG) 106. NLU 104 refers to understanding the meaning of written and / or spoken language (e.g., text, speech, or a combination thereof). Examples of NLU 104 include text reasoning or email classification. NLG 106 refers to the task of generating written and / or spoken language (e.g., text, speech, or a combination thereof) based on structured data, unstructured data, or a combination thereof. Examples of NLG 106 include query-centric summary generation, story generation, news summary generation, conversational artificial intelligence (AI), or a combination thereof. In some examples, an NLP system may include a combination of NLU 104 and NLG 106, such as question answering, interpreting and then summarizing content (e.g., a news article or story), or a combination thereof. In some examples, NLG 106 may include a transformer-based NLG 106 .
[0045] Figure 2 200 is a conceptual diagram illustrating an example of an illusion 202 in a chatbot using natural language generation (NLG). An illusion may refer to a situation where an NLG model becomes convinced of untrue facts and generates text or speech based on untrue facts. An illusion may also refer to text that is meaningless or unfaithful to the input content on which the text is based. For example, the chatbot in the chat illustrated in the conceptual diagram exhibits an illusion 202, and the chatbot outputs a factually incorrect statement "Yes, I am a person" when answering the query "So you're a person?" The chatbot in the chat illustrated in the conceptual diagram again exhibits an illusion 202, and for the query "Not a machine?", the chatbot answers: "Nope definitely not a machine, but sometimes it feels like people treat me like one when they ask me questions like that lol". Illusions similar to the illusion 202 may hinder the performance of the system and may cause safety issues, especially in the case of relying on the system to provide accurate medical data, news summaries, driving instructions, or other data that users may rely on to make decisions.
[0046] Another illustrative example of hallucinations is provided herein in the context of news summary generation. An exemplary news article discusses a car accident involving car A driven by person A and car B driven by person B, in which person B died in the car accident. An exemplary summary including hallucinations generated by an NLG system reads: "Person A has died investigated by police in Florida after a car crashed into her man car". The summary includes an illusion that indicates that person A died, when in fact person B died. The summary also includes additional hallucinations in the form of meaningless text, such as "has died investigated by police" or "car crashed into her man car". The various systems and techniques described herein for mitigating hallucinations produce an improved summary in the illustrative example, "Person A is being investigated by police in Florida after her car crashed into to Person B while she was driving", which does not contain any hallucinations.
[0047] Figure 3A 1 is a block diagram of a natural language generation (NLG) system 300. The NLG system 300 receives input text 302 at an encoder 304, which may tokenize the input text 302 to divide the input text 302 into tokens (e.g., words or parts thereof) and thereby understand the input text 302 through the NLU 104. The NLG system 300 includes a decoder 306 that generates an output text 308 by selecting tokens (e.g., words or parts thereof) from a set of possible tokens to include in the output text 308. The decoder 306 generates the set of possible tokens for the output text 308 and / or the selection of tokens from the set of possible tokens may be based on the input text 302 and / or tokens read from the input text 302 by the encoder 304. In some examples, the decoder 306 may select a token for the output text 308 from the set of possible tokens based on which token is most likely to occur next, given any previously selected tokens and / or given the input text 302.
[0048] Figure 3B1 is a block diagram of a natural language generation (NLG) system 350 with a natural language inference (NLI) scoring system indicating fidelity to input text. Similar to the NLG system 300, the NLG system 350 receives input text 302 at an encoder 304, which can tokenize the input text 302 to divide the input text 302 into tokens and thereby understand the input text 302 through the NLU 104.
[0049] The NLG system 350 includes a decoder 310 with hallucination mitigation that generates an output text 312 by selecting tokens (e.g., words or parts thereof) from a set of possible tokens to include in the output text 312. The decoder 310 with hallucination mitigation generates the set of possible tokens for the output text 312 and / or the selection of tokens from the set of possible tokens may be based on the input text 302 and / or tokens read from the input text 302 by the encoder 304. The decoder 310 with hallucination mitigation may select a token for the output text 312 from the set of possible tokens based on which token is most likely to occur next, taking into account any previously selected tokens and / or taking into account the input text 302. The decoder 310 with hallucination mitigation may select tokens for the output text 312 from the set of possible tokens based in part on which token(s) are most faithful to the input text 302 (or make the output text 312 most faithful to the input text 302), which token(s) have the highest factual accuracy (or make the output text 312 have the highest factual accuracy), which token(s) have the lowest factual inaccuracy (or make the output text 312 have the lowest factual inaccuracy), which token(s) are the most meaningful (or make the output text 312 the most meaningful), which token(s) are the least meaningless (or make the output text 312 the least meaningless), which token(s) have the highest contextual implication (or make the output text 312 have the highest contextual implication), which token(s) are the least contradictory with respect to the input content (or make the output text 312 the least contradictory with respect to the input content), or a combination thereof. In this way, the decoder 310 with hallucination mitigation may mitigate hallucinations in the output text 312 compared to the output text 308 because the decoder 306 may lack hallucination mitigation.
[0050] Figure 4A is a conceptual diagram of a greedy search decoding algorithm 400 for a natural language generation (NLG) system. In some examples, the greedy search decoding algorithm 400 uses the following equation:
[0051]
[0052] Given the words (y1, ..., yt-1) generated in the past and the activity report c also generated at each step, the greedy search decoding algorithm 400 can select a token (e.g., a word or part thereof) from a set of possible tokens at each branch based on which word is most likely to be used next. Figure 4A As shown, the selected tokens are indicated by the thicker lines between the tokens. Figure 4A The greedy search decoding algorithm 400 selects the token with the highest probability (or confidence value) at each level. For example, Figure 4A In the illustrated example, the greedy search decoding algorithm 400 outputs the phrase "The nice woman" based on the fact that "nice" (50% probability) is more likely to follow "The" than "dog" (40% probability) or "car" (10% probability), and based on the fact that "woman" (40% probability) is more likely to follow "nice" than "house" (30% probability) or "guy" (30% probability).
[0053] Figure 4B is a conceptual diagram of a beam search decoding algorithm 450 for a natural language generation (NLG) system. The beam search decoding algorithm 450 explores the N tokens with the highest probability at each step, taking into account the words and activity reports generated in the past, and selects the best overall sentence or phrase (e.g., the sentence or phrase with the highest probability overall). For example, taking into account the words, sentences, phrases, and / or activity reports generated in the past, the beam search decoding algorithm 450 may select the sentence or phrase with the highest probability. In an illustrative example, the beam search decoding algorithm 450 may generate several sentences or phrases, including "The nice woman", "The nice guy", "The dog has", and "The dog and". In the illustrative example, the beam search decoding algorithm 450 may select the sentence or phrase "The nice woman" because the overall sentence or phrase has a higher probability of use (e.g., taking into account the words, sentences, phrases, and / or activity reports generated in the past) than other generated sentences or phrases (e.g., "The nice guy", "The dog has", and "The dog and").
[0054] Advances in pre-trained large language models have significantly improved their performance for conditional language generation tasks, including summary generation, despite consistent hallucinations. To reduce hallucinations, some systems may improve beam search or use fact checkers as a post-processing step. The systems and techniques described herein use natural language inference (NLI) entailment metrics to detect and prevent hallucinations in summary generation. The NLI-assisted bundle reordering mechanism is implemented by calculating the entailment probability score between the input context and the bundle generated by the summary generation model during saliency-enhanced greedy decoding. In addition, a diversity metric is introduced to compare its effectiveness with the basic beam search. For the Xsum and CNN / DM datasets, the decoder 700 and decoder 750 discussed herein are significantly better than the basic bundle decoding in terms of this metric and other metrics.
[0055] Pre-trained sequence-to-sequence transformer models such as BART or Pegasus have shown substantial improvements in the performance of NLP tasks such as summary generation, story generation, generative question answering, etc. Hallucinations are a problem that can be observed during the generation process in some cases (especially when pre-training is mainly done on unlabeled data). During the pre-training phase, the model learns the language and its grammar inaccurately and can generate words that are irrelevant to the given input during inference time.
[0056] Some systems or techniques can mitigate or suppress hallucinations during decoding using beam search modifications that constrain the decoding step to focus on tokens that support the input. In some examples, for NLP-based summary generation, inaccuracies in summaries provided to the ML model as training data can cause inconsistencies (e.g., hallucinations) in the text generated by the ML model for NLG. In some examples, the relationship between hallucinations and prediction uncertainty can be exploited by modifying the beam search to prefer low prediction uncertainty.
[0057] While constraining the beam search using a heuristic function may be slightly more effective in alleviating hallucinations, it may still (in some examples) benefit from manual inspection of the initialization of the beam search hyperparameters using sophisticated knowledge of the dataset, task, and model. For example, PINOCCHIO may use cosine distance at each decoding step to measure the consistency of the generated word with the context. As the dataset becomes more generative, relying solely on cosine distance and simple word-level heuristics to control beam decoding at the fact level may become less effective.
[0058] The NLG systems and techniques described herein for alleviating hallucinations based on natural language inference (NLI) scores (e.g., decoder 310 with hallucination mitigation of NLG system 350) can overcome the limitations of heuristics and cosine distances by reordering the top N predictions of a model using semantically matching NLP tasks of natural language inference (NLI). The NLG systems and techniques can calculate NLI entailment scores at each bundle decoding step to provide the model with opportunities to steer bundle trajectories toward regions, tokens, or words that are less hallucinatory. Each intermediate bundle can be generated using greedy unroll decoding while focusing on salient contextual portions. In some examples, the bundles can be ordered at a sentence-level granularity using a SummaC score metric.
[0059] The NLI score can be used to detect hallucinations in generative summary generation, as discussed later in Figure 5 The NLG systems and techniques described herein for alleviating hallucinations based on Natural Language Inference (NLI) scores include a hallucination mitigation component for beam search that can modify cumulative beam probabilities at a token level using an NLI metric or score, and can compute re-ranking performance using a diversity and summary consistency (SummaC) score metric on an extreme summary generation (Xsum) and / or Cable News Network / Daily Mail (CNN / DM) dataset.
[0060] NLI scores can be used to measure and / or improve the fidelity of the output text to the input content. Fidelity can refer to how consistent the generated output text is with respect to the input content. For example, terms, phrases, or sentences that are inconsistent at the factual level in the generated output text compared to the input content can be examples of hallucinated text. Other types of hallucinations in the generated output text (such as nonsense text) may also be unfaithful compared to the input content. NLI scores can be applied to alleviate the hallucinations of different NLG-based generative summary generators, such as GPT fine-tuned Seq2Seq based on recurrent neural networks (RNNs) and Seq2Seq (BertS2S) based on bidirectional encoder representations of transformers. In some examples, the Spearman correlation coefficient between the textual entailment score and the faithful summary is the highest compared to other automatic metrics such as recall-oriented key point evaluation substitute (ROUGE)-1, ROUGE-2, and BertScore (e.g., using a large model of bidirectional encoder representations (BERT) based on transformers fine-tuned on a multi-genre natural language inference (MNLI) dataset). Therefore, NLI scores that measure textual entailment can be used to reduce hallucinations.
[0061] To measure factual inconsistency, a trained factual consistency checking model (FACTCC), i.e., a BERT-based model, is fine-tuned on synthetically hallucinated summaries using semantically varying / invariant transformations such as entity swapping, sentence negation, paraphrase, and noise injection. However, in some examples, such models may lack explanations and / or have low generalizability to other datasets, e.g., being good at finding certain hallucinations. Improvements to the loss function components may improve the overall factual accuracy. For example, truncating the loss by adaptively removing high log loss examples may improve factual accuracy in the model.
[0062] Hallucinations exist in various NLP downstream tasks and can be measured using various metrics. A generated summary may be defined as hallucinating if it has any text span that is not semantically supported by the input content on which the generated summary is based. Hallucinations can be divided into two main types - intrinsic and extrinsic. Intrinsic hallucinations refer to inconsistencies in the generated summary about the input content. For example, intrinsic hallucinations may include the use of incorrect pronouns, swapping names and verbs, etc. Models similar to FACTCC (e.g., trained on slight text transformations) can be used to detect intrinsic hallucinations. Extrinsic hallucinations may refer to unsupported text spans that exist in the generated summary and cannot be verified using only the input content. Extrinsic hallucinations may occur due to the presence of extrinsic hallucinations in the human-written summaries in the training data on which the model is trained (e.g., may be overfitted) during the training process. For example, in Seq2Seq models such as GPT2, the percentage of hallucinations can be amplified or reduced by modifying the training data.
[0063] Natural language inference (NLI) may refer to the task of determining whether a natural language hypothesis can be inferred from a given premise. Given the premise and the hypothesis, NLI calculates the relationship between the premise and the hypothesis in the form of three probabilities (implication, contradiction, and neutrality). In some examples, the NLI algorithm may focus on one, two, or all three of these probabilities. For example, in the illustrative examples, the NLI system may focus on the implication. For example, in the illustrative examples, the NLI system may focus on the implication. For example, if the premise is "the sky looks cloudy today." And the hypothesis is "it may rain today", the NLI model will assign a greater probability to the implication because the hypothesis implies the premise. Natural language inference (NLI) can be used to detect hallucinations.
[0064] Figure 55 is a conceptual diagram illustrating a histogram of entailment scores or natural language inference (NLI) scores indicating faithfulness to input content for output text with and without hallucinations. Text entailment can be used to detect hallucinations in generative summary generation tasks. Intrinsic hallucinations can be difficult to detect because their detection may require more than just lexical matching to infer the relevance of a given word to the context.
[0065] The histogram includes a histogram 502 of textual entailment scores for training data with hallucinations and a histogram 504 of textual entailment scores for training data without hallucinations. Figure 5 In the context of , entity-based hallucinations are counted for analysis purposes. This histogram illustrates the experimental results of analyzing the correlation between entailment scores and entity hallucinations on 2000 randomly selected training samples from the Xsum dataset. Figure 5 It is evident from the data that, although there is a high frequency of low entailment scores for both data with / without hallucinations, the distinction between them becomes clearer at higher entailment scores. In fact, higher entailment scores are associated with a low probability of entity hallucinations. This is also reflected in the average entailment scores in Table 1. This analysis illustrates that entity-based hallucinations can be detected by the NLI metric. Therefore, introducing NLI during bundle decoding can be used to mitigate hallucinations.
[0066] Dataset / Bucket Output Illusion No hallucinations Xsum 0.24347 0.43320
[0067] Table 1: Average entailment scores of Xsum training data on 2000 samples .
[0068] Figure 6 6 is a block diagram of a decoder 600 with a beam search and a natural language inference (NLI) scorer for a natural language generation (NLG) system. An encoded representation 602 is input into a transformer block 604 to identify a set of possible tokens. A beam search 606 is used to sort the tokens based on the probability of use, thereby generating an intermediate beam 612 that is input into an NLI scorer 608. The NLI scorer 608 in turn generates a reordered intermediate beam 614, taking into account the intermediate beam 612 and the context activity report 610, so as to be input back into the beam search 606, thereby producing a final beam 616 that is ultimately used to generate an output text 618. The NLI scorer 608 is introduced into the beam search 606 decoding process. At each token generation step, the model considers the NLI score from the NLI scorer 608 and the predicted score from the beam search 606.
[0069] Fig. 7A608 is a block diagram of a decoder 700 for a natural language generation (NLG) system with greedy expansion 704 and a natural language inference (NLI) scorer 608. Natural language inference (NLI) may refer to the task of determining whether a hypothesis is true (implied), false (contradicted), or undetermined (neutral, or neither contradictory nor implied) given a "premise". In some examples, the respective probabilities of contradiction, implication, or neutrality add up to 1. Thus, if the probability of implication is high, the probability of contradiction and / or neutrality may be low. Similarly, if the probability of contradiction is high, the probability of implication and / or neutrality may be low. If the probability of neutrality is high, the probability of contradiction and / or implication may be low.
[0070] Figure 7B 7 is a block diagram of a decoder 750 for a natural language generation (NLG) system with a significance-enhanced greedy expansion 712 and a natural language inference (NLI) scorer 608. In some examples, the decoder 700 and / or the decoder 750 may use a transformer-based bidirectional encoder and autoregressive decoder representation (BART) base model fine-tuned on a given dataset for an NLI-assisted beam search reranker. A BART-like architecture may have an autoregressive decoder that generates output word by word conditioned on the input text and the words generated so far. The beam search may perform an extensive first search with limited branches, where the beam size starts at a BOS (start of sentence) token and ends the search at an EOS (end of sentence) token. Each path from BOS to EOS may be referred to as a hypothesis.
[0071]
[0072] An intermediate bundle or partial hypothesis is a sequence of subpaths of a hypothesis that starts at the BOS and ends before the EOS. Figure 4A Examples of intermediate bundles in the context of include "the nice woman", "the nice guy", "the dog has", and "the dog and". In decoder 700, greedy expansion 704 focuses on the important part of the context related to intermediate bundle 702 (e.g., as in intermediate bundle 612) and completes the bundle until EOS. In decoder 750, saliency-enhanced greedy expansion 712 focuses on the important part of the context related to intermediate bundle 702 (e.g., as in intermediate bundle 612) and completes the bundle until EOS.
[0073] Intermediate bundle 702 may be decoded by decoder 700 and / or decoder 750 based on the probability of each word (e.g., using Figure 4A Greedy search decoding algorithm 400 and / or Figure 4BThe decoder 700 and / or the decoder 750 may sort the intermediate bundles 702 based on a cumulative probability based on the probability of each word in each intermediate bundle 702. For example, the decoder 700 and / or the decoder 750 may sort the intermediate bundles 702 based on a cumulative probability based on the probability of each word in each intermediate bundle 702. Figure 7B indicates that the middle bundle 702 is selected and sorted, wherein the first is "The death of", the second is "Tennis star Venus", and the third is "Venus Williams is".
[0074] Fig. 7A The greedy expansion 704 of the decoder 700 uses a greedy search (e.g., as in Figure 4A In the greedy search decoding algorithm 400 of ), a word is added to each intermediate bundle in the intermediate bundle 702 until each intermediate bundle in the intermediate bundle 702 is completed into a corresponding complete sentence. Similarly, Figure 7B The saliency-enhanced greedy expansion 712 of the decoder 750 uses a greedy search (e.g., as in Figure 4A ) adds words to each of the intermediate bundles 702 until each of the intermediate bundles 702 is completed as a corresponding complete sentence, but compared to the greedy expansion 704, the greedy expansion model's field of view is limited to only the most important or most significant words. For example, the words determined to be the most important words or the most significant words may be words with a significance level or significance level that exceeds a significance threshold. In some examples, the significance threshold may be based on an average significance value and / or a standard deviation significance value of the corresponding significance values of the candidate words, such that a word having a significance above the average significance or having a significance that exceeds the average significance plus a standard deviation (e.g., multiplied by a multiplier) may be considered to exceed the significance threshold. Regardless of which type of greedy expansion is used, the NLI scorer 608 then scores each of the complete sentences to generate an NLI score for each of the complete sentences. For example, NLI scorer 608 generates NLI score 706 for the complete sentence generated by greedy expansion 704 , and generates NLI score 716 for the complete sentence generated by saliency-enhanced greedy expansion 712 .
[0075] The NLI score for the complete sentence (e.g., NLI score 706 or NLI score 716) is passed to a bundle reorderer 708 with weighted NLI scores and model probabilities to reorder the intermediate bundles 702 to generate reordered intermediate bundles. Thus, the reordered intermediate bundles are reordered based on the complete sentence that each of the intermediate bundles is most likely to produce, thereby essentially allowing the decoder to use a greedy search to quickly look ahead in time to see what each of the intermediate bundles 702 might turn into, saving time and computational resources compared to performing a more exhaustive search (e.g., a beam search). If NLI scorer 608 indicates (e.g., via NLI score 706 and / or NLI score 716) that a complete sentence corresponding to a particular intermediate bundle (e.g., generated using greedy expansion 704 or saliency-enhanced greedy expansion 712) contains hallucinations, factual inaccuracies, contradictions, and / or other errors, then bundle re-ranker 708 may lower the ranking of the intermediate bundle to a lower ranking because this indicates that the complete sentence generated using the intermediate bundle may contain hallucinations, factual inaccuracies, contradictions, and / or other errors. On the other hand, if NLI scorer 608 indicates (e.g., via NLI score 706 and / or NLI score 716) that a complete sentence corresponding to a particular intermediate bundle (e.g., generated using greedy expansion 704 or saliency-enhanced greedy expansion 712) contains no hallucinations, factual inaccuracies, contradictions, and / or other errors (or contains fewer hallucinations, factual inaccuracies, contradictions, and / or other errors than complete sentences corresponding to other intermediate bundles), then bundle reorderer 708 may lower the ranking of the intermediate bundle to a higher ranking because this indicates that the complete sentence generated using the intermediate bundle is likely to contain no hallucinations, factual inaccuracies, contradictions, and / or other errors. For example, the decoder 750 is shown in FIG. Figure 7B Indicates that the reordered intermediate bundles 720 reordered by the bundle reorderer 708 have demoted the intermediate bundle "The death of" from 1st to 3rd place (e.g., based on the high hallucination level in the corresponding complete sentence as indicated in the NLI score 716), have raised the intermediate bundle "Venus Williams is" from 3rd to 1st place (e.g., based on the low hallucination level (or no hallucination) in the corresponding complete sentence as indicated in the NLI score 716), and have maintained the intermediate bundle "Tennis star Venus" at 2nd place (e.g., based on the medium hallucination level in the corresponding complete sentence as indicated in the NLI score 716).
[0076] In some examples, at each bundle step, the decoder 750 progressively reorders additional bundle steps. Each bundle step may include a set number of additional words. For example, Figure 7BEach of the intermediate bundles 705 illustrated in includes 3 words. Once the words are reordered by the bundle reorderer 708, the system can select the highest reordered bundle and continue to generate text (e.g., a summary) by adding another 3 words and then using the same hallucination mitigation process (e.g., using greedy expansion 704 or greedy expansion with saliency enhancement 712, NLI scorer 608, and bundle reorderer 708) for the new set of intermediate bundles of the next 3 words. For example, if "Venus Williams is" is selected, the next set of intermediate bundles for the next round of hallucination mitigation may be "Venus Williams is being investigated by", "Venus Williams is under investigation for", and "Venus Williams is involved in an". If, among these bundles, the bundle reorderer 708 ranks "Venus Williams is being investigated by" highest, then the next set of intermediate bundles for the next round of hallucination relief may be "Venus Williams is being investigated bypolice in Florida", "Venus Williams is being investigated by authorities foran", and "Venus Williams is being investigated by United States police". Among these bundles, the bundle reorderer 708 may rank "Venus Williams is being investigated by police inFlorida" highest, and may generate the next set of intermediate bundles for the next round of hallucination relief for the next set of three additional words, as discussed above. This process may continue until a complete sentence is generated.
[0077] As indicated above, the intermediate bundles 702 are passed to the greedy expansion 704 and / or the saliency-enhanced greedy expansion 712 as a look-ahead mechanism for completing these bundles. The completed candidate bundles (e.g., complete sentences 714) are configured to be scored using the entailment probabilities of the NLI scorer 608 model (e.g., NLI score 706 and / or NLI score 716). The intermediate bundles are then reordered based on the weighted probabilities between the entailment and model probabilities using the bundle reorderer 708 with weighted NLI scores and model probabilities. The detailed steps are provided in Pseudo-Code 1. In some examples, the bundle reorderer 708 with weighted NLI scores and model probabilities may reorder according to the following equation:
[0078] Bundle reorderer = log(a*cumulative bundle probability + b*NLI implied probability)
[0079] Equation 3: Bundle Reorderer
[0080] In some examples, the sum of parameters a and b in Equation 3 is 1, so if one of these parameters increases, the other decreases. In some examples, increasing parameter b can improve the level of hallucination mitigation, and thus can improve the fidelity of the resulting generated text (e.g., the generated summary). In some examples, increasing parameter a can reduce hallucination mitigation, which may be helpful in situations where it is desired that the resulting generated text is neutral or generative with little risk of hallucination. In the case of greedy expansion 704 of decoder 700, greedy search is used to complete the bundle (e.g., as in Figure 4A Greedy search decoding algorithm 400).
[0081] Regarding greedy expansion 704 and / or saliency-enhanced greedy expansion 712—When the NLI model has been trained with complete sentences, it may be difficult to perform NLI tasks on partial hypotheses. Therefore, the decoder 700 and / or the decoder 750 may complete 2B intermediate bundles 702 as an initial step (e.g., using greedy expansion 704 and / or saliency-enhanced greedy expansion 712), where B is the bundle size. The decoder 700 and / or the decoder 750 may use greedy search (e.g., greedy expansion 704 and / or saliency-enhanced greedy expansion 712) on the intermediate bundles 702 in order to generate the remaining words and complete the partial hypothesis. In pseudocode 1, the saliency-enhanced greedy expansion (SGR) function takes a concatenated input of context, intermediate bundles, and the next word separated by a sentence separation token ([SEP] token), and generates a complete bundle. During the greedy search, similar words may be used to complete the bundles regardless of the words in the intermediate bundles. This may be because the pre-trained transformer has a long context and a shorter attention span. Therefore, in some examples, the model may not effectively focus on the parts of the context related to the words in the intermediate bundle. To address this issue, the decoder 700 and / or the decoder 750 may take two steps. First, the decoder 700 and / or the decoder 750 may enhance the effectiveness and diversity of the greedy search by introducing saliency on the context relative to the intermediate bundle using an attention head mask (e.g., a saliency-enhanced greedy expansion 712). The decoder 750 may calculate a saliency score for each word or token in the context by averaging the cosine distance of each word or token with each word in the intermediate bundle. Using a threshold as a hyperparameter, the decoder 750 calculates a mask matrix m (see Equation 4 below) to selectively focus on words in the context that are related to the completion of the current intermediate bundle.
[0082]
[0083] Second, decoder 700 and / or decoder 750 may perform the proposed reordering only if the hypothesis has the fewest words so that the bundles do not converge to the same space during the greedy search. This is because if the hypothesis has very few words, the bundles may not have the necessary entities suitable for measuring hallucinations. In some examples, decoder 700 and / or decoder 750 may automatically identify the appropriate time step suitable for reordering the hypothesis to avoid hallucinations. In some examples, the minimum number of time steps to perform the reordering is a hyperparameter of decoder 700 and decoder 750.
[0084]
[0085] Pseudocode 1
[0086] As a next step, the decoder 700 and / or the decoder 750 may pass the greedy expansion bundle to the NLI scorer 608. The decoder 700 and / or the decoder 750 takes the context as an premise and the bundle as an assumption to obtain the entailment probability, as illustrated in Equation 5. The NLI function takes the context C as an premise, the expanded bundle R as an assumption, and calculates their relationship as an entailment score. In some examples, the entailment probability may be inversely proportional to the hallucination content of the bundle. In order to quantify whether the bundle is able to explore different areas of the text space, the decoder 700 and / or the decoder 750 may use the diversity metric Diversity (see Equation 5 below) to measure the average frequency of novel words across bundles. In some examples, a set intersection operation may merge the semantic representations of words.
[0087]
[0088] In Equation 5 above, n is the beam size, and b i is the set of unique words in bundle i.
[0089] To incorporate the NLI score into the overall cumulative bundle probability, the decoder 700 and / or decoder 750 takes a weighted average of the entailment and model probabilities for each decoding step and adds it to the cumulative bundle probability. The bundle reorderer 708 with weighted NI scores and model probabilities then reorders the bundles based on the modified cumulative probabilities and selects the top B candidates as the reordered intermediate bundles 710. When we add two random variables, the weights need to be normalized. As mentioned in Equation 6, the decoder 700 and / or decoder 750 treats the weight (α) as a hyperparameter that can be increased to 1.0 depending on the necessity of fidelity in the text generated for a given task.
[0090] P 蕴含 :=NLI(C,R)
[0091] P 加权 : =αP i +(1-α)P 蕴含
[0092] Equation 6: Weighted average of implicature and model probabilities
[0093] In an illustrative example, two data sets (i.e., CNNDM and Xsum) are utilized to test decoder 700 and / or decoder 750 to evaluate model performance. The CNNDM corpus is generated from multi-line summaries written manually based on news articles from CNN and the Daily Mail. It consists of more than 285k training pairs, 13,368 validation pairs, and 11,487 test pairs. The Xsum data set consists of articles from BBC and corresponding single-line summaries. It contains more than 90k training samples and is more generative than CNN / DM because it contains more than 18.6% of new words. The system and method described herein consistently handle summaries of generative types and summaries of extractive types.
[0094] In an illustrative example, decoder 700 and / or decoder 750 may be implemented using a pytorch implementation of the Bidirectional Encoder and Autoregressive Decoder Representation (BART) based model from the huggingface library. In an illustrative example, decoder 700 and / or decoder 750 may be implemented using a 4e with linear decay -5 The learning rate is trained for 6 epochs. For the decoding process, the decoder 700 and / or the decoder 750 may use a beam search with a beam size of 5 and a maximum length of 125 tokens after byte pair encoding (BPE) tokenization. In some examples, early stopping is set to true and the repetition penalty is set to 3.0. For NLI, the decoder 700 and / or the decoder 750 may use a BART large model fine-tuned on the MNLI dataset.
[0095] In an illustrative example, decoder 700 and / or decoder 750 may use the SummaC model to measure summary consistency. The NLI model can measure the similarity between each sentence in the context and the summary by creating an NLI pair matrix. Two methods can be used: SummaConv and Summac ZS, which differ in the way the final score is calculated. SummacZS takes the direct maximum of the columns in the NLI pair matrix, while the former uses 1-D convolution to reach a single score. In an illustrative example, decoder 700 and / or decoder 750 may use SummaConv and diversity scores as their evaluation metrics. Both decoder 700 and decoder 750 provide technical improvements that are superior to separate beam search and separate greedy search, such as reducing hallucinations, improving accuracy and / or improving reliability. In Table 2, decoder 750 is measured using SummaConv scores and diversity scores against beam search. In Table 2, decoder 750 is benchmarked against 6 consistency data sets including FactCC, SummEval, and non-NLI consistency metrics (such as DAE and FACTCC).
[0096] The beam search modifications described herein reduce hallucinations during inference time compared to beam search lacking the beam search modifications described herein. In Table 2, the improvement in the SummaConv score of decoder 750 for both the Xsum and CNNDM datasets relative to beam search alone confirms that NLI helps reduce hallucinations and aligns the generated text with facts from the context. The relatively high diversity score of decoder 750 relative to beam search alone suggests that the beams generated by decoder 750 explore a more diverse range of text to avoid hallucinations (compared to beam search alone).
[0097] Dataset Decoding algorithm Diversity score Summa Conv Score XSum Beam search only 2.27 0.315 Decoder 750 (α=0.8) 2.53 0.526 CNN / DM Beam search only 1.66 0.872 Decoder 750 (α=0.6) 1.69 0.898
[0098] Table 2: Single beam search vs. Figure 7B Performance comparison of decoder 750
[0099] The hyperparameter α plays a role in guiding the beam towards fact generation by varying its value across the spectrum. Fig. 8A 800 is a conceptual diagram illustrating examples of different output text strings with different natural language inference (NLI) scores. Fig. 8A In , different generated summaries are illustrated for different sets of parameters and / or sets of weights. The parameters and / or weights may be input to the bundle reorderer 708 and may be used to generate the summaries of the different sets of parameters and / or weights. Fig. 8A3. The parameters and / or weights a and b may be the same as those indicated in Equation 3. In some examples, the sum of parameters a and b is 1, so if one of these parameters increases, the other decreases. In some examples, increasing parameter b may increase the level of hallucination relief, and thus may increase the fidelity of the resulting generated text (e.g., a generated summary). In some examples, increasing parameter a may reduce hallucination relief, which may be helpful when the resulting generated text is desired to be neutral or generative with little risk of hallucinations.
[0100] Table 3 below illustrates the impact of NLI probabilities of implications (E) and contradictions (C) on the overall performance of decoder 700 and / or decoder 750 .
[0101]
[0102] Table 3: Impact of NLI probabilities of entailment (E) and contradiction (C) on overall performance
[0103] Table 3 illustrates an example of combining the contradiction probability and the entailment probability as illustrated using Equation 7. For the example in Table 3, α1 = 0.6, and α2 = 0.2.
[0104] P 概率 : =α1·P 蕴含 +(1-α1)·P 矛盾
[0105] P 加权 : =α2·P i +(1-α2)·P 概率
[0106] Equation 7: Weighted average of implicature and model probabilities
[0107] Table 4 below illustrates an analysis of different decoding strategies for the expansion components of decoder 700 and / or decoder 750 (e.g., greedy expansion 705 and / or saliency-enhanced greedy expansion 712). In Table 4, it can be seen that random sampling expansion increases by 0.212 compared to greedy expansion. Since XSum is usually generative, random sampling helps explore less frequent faithful words that may be ignored by other methods. Since CNN / DM is mostly extractive, greedy search is able to select the words that are most likely to appear in the context.
[0108]
[0109] Table 4: Analysis of different decoding strategies for expanding components
[0110] Table 5 below illustrates the impact of different NLI datasets, such as Multi-Genre Natural Language Inference (MNLI) and Stanford Natural Language Inference (SNLI), on the overall SummaConv score of the decoder 700 and / or the decoder 750:
[0111]
[0112] Table 5: Impact of the NLI dataset on the overall SummaConv score
[0113] exist Fig. 8A In Figure 1, a golden summary is illustrated. The golden summary is a summary of a manually generated news article and reads "US tennis star Venus Williams has been involved in a car accident that led to the death of a 78-year-old man". An exemplary bad summary generated by a bad model is illustrated as "Tennis star Venus Williams has died investigated by police in Florida after a car crashed into her man's car", which is factually inaccurate and inconsistent with the article and the corresponding golden summary.
[0114] exist Fig. 8A , ten summaries of articles generated using the systems and methods described herein (e.g., using decoder 750) are illustrated. Of the ten summaries, the first set of five summaries were written with parameters set to a=0.8 and b=0.2. The first set of five summaries all have factual inaccuracies, such as indicating that Venus Williams is dead. Of the ten summaries, the second set of five summaries were written with parameters set to a=0.0 and b=1.0. The second set of five summaries includes three summaries containing factual inaccuracies (again indicating that Venus Williams is dead) and two factually accurate summaries (labeled #3 and #4 and outlined in black rounded rectangles). Each of the ten summaries is followed by a corresponding confidence value generated by decoder 750 (e.g., by bundle reorderer 708), indicating the confidence that the summary is accurate. The two factually accurate summaries have the highest confidence values of the ten summaries, 0.98 and 0.99, respectively.
[0115] Figure 8B 850 is a conceptual diagram illustrating examples of different output text strings with different natural language inference (NLI) scores. Figure 8B For parameters α = 0.0 and α = 0.2, Figure 8BAll generated bundles illustrated in are factually incorrect. Figure 8B , the frequency of generated bundles that are consistent at the fact level grows steadily from α=0.4 to 1.0, indicating that the greater the impact of the NLI, the more consistent the generated bundles are at the fact level. The parameter α can be an input into the bundle reorderer 708, for example, as in Equation 7. Each of these summaries is followed by a corresponding confidence value generated by the decoder 750 (e.g., by the bundle reorderer 708), indicating the confidence that the summary is accurate. Figure 8B In , the four summaries that are factually accurate are outlined with rounded rectangles and have the highest confidence values among these summaries, where each summary has a confidence value of 0.98 or 0.99.
[0116] In some examples, using the entailment probability and the contradiction probability may produce different effects on the NLI scorer. In some examples, decoder 700 and / or decoder 750 may take a weighted average of the entailment probability and the contradiction probability and combine the weighted average with the token probability. weighted The following equation 7 can be used to modify:
[0117] P 概率 : =α1P 蕴含 +(1-α1)P 矛盾
[0118] P 加权 : =α2P i +(1-α2)P 概率
[0119] Equation 7: Combination of implicature probability and contradiction probability
[0120] In some examples, decoder 700 and / or decoder 750 may be influenced by the correlation of the saliency attention between the intermediate bundle and the context. In some examples, each word in the intermediate bundle may affect the saliency to a greater extent to establish the importance of the cross attention between the two components.
[0121] By analyzing the correlation between the implication score and the entity hallucination, the NLI model can be used as a reliable guide to alleviate the hallucination during reasoning time. Decoder 600, decoder 700 and / or decoder 750 show modifications to the beam search decoding algorithm, which guides beam generation to avoid falling into the hallucination area by reordering the beam based on the NLI implication score, which is calculated based on the partial hypothesis greedily expanded based on the significance enhancement. In some examples, the NLI-based reorderer can consistently improve the SummaConv score. In some examples, the NLI-based reorderer can further improve other NLP downstream tasks, such as story generation with prompts, question answering, and query-oriented summary generation. In some examples, NLI can be incorporated as a guidance mechanism for the decoding algorithm. In some examples, NLI can be extended to other NLG tasks, such as question answering.
[0122] Fig. 9 900 is a conceptual diagram illustrating a model for generating a summary of input content using a natural language generation (NLG) system. In some examples, the system may generate output text that conforms to the authenticity of the source input text. In the word-by-word text generation process, the system may help keep the model on the right track. If available, the system may provide a summary of objective measures of performance. For example, Summac Conv may calculate factuality scores by segmenting the input text and the output text into sentence units and aggregating natural language inference (NLI) scores between sentence pairs. Rouge (R-1, R-2, RL) may compare the overlap of words / phrases between the generated summary and the golden summary (e.g., a predetermined summary written by a person). Diversity scores may be used to calculate how different their bundles are by comparing the word overlap of the generated bundles. In some examples, model 900 may be used to check the factuality of the summary as a quality check after the summary is generated. In some examples, model 900 may be part of NLI scorer 608.
[0123] Fig.10 is a flow chart illustrating an example process 1000 for language generation (NLG) using one or more of the techniques described herein. The process 1000 can be performed using an NLG system, which can include, for example, the NLG system 300, the NLG system 350, the encoder 304, the decoder 306, the decoder with hallucination mitigation 310, the decoder 600, the transformer block 604, the beam search 606, the NLI scorer 608, the decoder 700, the decoder 750, the greedy expansion 704, the beam reranker with weighted NLI scores and model probabilities 708, the greedy expansion with saliency enhancement 712, the model 900, the NN 1100, the computing system 1200, or a combination thereof.
[0124] At operation 1005, the NLG system (or at least one of its subsystems) is configured and capable of generating a token sequence based on input content. In some examples, the input content includes input text (e.g., input text 302), input speech, or a combination thereof. The token sequence may correspond to the intermediate bundle 702.
[0125] At operation 1010, the NLG system (or at least one subsystem thereof) is configured and capable of determining a confidence level associated with the sequence of tokens based on the respective confidence levels associated with each token in the sequence of tokens. The confidence level may correspond to Figure 7B Initial ordering of intermediate bundles 702 prior to illustrated phantom relief.
[0126] At operation 1015 , the NLG system (or at least one subsystem thereof) is configured and capable of generating a complete sentence including the sequence of tokens, for example, using greedy expansion 704 or saliency-enhanced greedy expansion 712 .
[0127] In some aspects, the NLG system (or at least one subsystem thereof) is configured and capable of using a beam search (e.g., as in Figure 4B beam search and / or beam search 606), using a greedy search based on the input content (e.g., as in Figure 4A In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of using a greedy search (e.g., as in Figure 4A , greedy expansion 704, and / or saliency-enhanced greedy expansion 712), using a beam search based on the token sequence (e.g., as in Figure 4B The complete statement is generated by beam search and / or beam search 606) or a combination thereof.
[0128] In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of limiting the candidate token for use in generating the complete sentence based on whether the corresponding significance value of the candidate token exceeds a significance threshold. In some aspects, the significance threshold is based on an average value of the corresponding significance value of the candidate token. For example, the threshold can be an average value (e.g., mean, median, mode) of the corresponding significance value, an average value of the corresponding significance value offset by an offset value (e.g., the product of a standard deviation and a multiplier), a product of the offset average value of the corresponding significance value and a multiplier, or a combination thereof.
[0129] In some aspects, the token sequence is configured to follow a previously determined token sequence in the complete sentence, and the complete sentence includes the previously determined token sequence, the token sequence, and at least one additional token.
[0130] At operation 1020, the NLG system (or at least one subsystem thereof) is configured and capable of generating a natural language inference (NLI) score (e.g., one of the NLI scores 706 or one of the NLI scores 716) for the complete sentence based on the fidelity of the complete sentence to the input content (e.g., based on the context activity report 610).
[0131] In some aspects, an NLI score in the NLI scores identifies whether at least a portion of the complete sentence (e.g., a token in the output text or a resulting statement) (e.g., relative to the input content) is true, false, or neutral (e.g., as in Fig. 7A exemplified).
[0132] At operation 1025, the NLG system (or at least one subsystem thereof) is configured and capable of adjusting the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence. The updated confidence level may correspond to the reordering of the intermediate bundle 702 by the bundle reorderer 708 after hallucination relief and / or the ordering of the reordered intermediate bundle (e.g., the reordered intermediate bundle 710 or the reordered intermediate bundle 720), as shown in FIG. FIG. 7A to FIG. 7B exemplified.
[0133] For example, in some aspects, the NLG system (or at least one of its subsystems) is configured and capable of sorting the token sequence against the second token sequence based on the confidence level associated with the token sequence and the second confidence level associated with the second token sequence. In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of re-sorting the token sequence against the second token sequence based on the updated confidence level associated with the token sequence and the second updated confidence level associated with the second token sequence. The second updated confidence level is based on the second NLI score of the second complete sentence, which is generated based on the second token sequence. In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of re-sorting the token sequence based on the comparison of the second token sequence, selecting the highest ranked token sequence from at least the token sequence and the second token sequence. The NLG system (or at least one of its subsystems) is capable of generating an output text including the highest ranked token sequence. In some aspects, the output text is configured to generate a summary of the input content.
[0134] In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of generating an output text including the token sequence based on the updated confidence level of the token sequence exceeding the second updated confidence level of the second token sequence. In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of generating the second token sequence based on the input content. The NLG system (or at least one of its subsystems) can determine the second confidence level associated with the second token sequence based on the secondary corresponding confidence level associated with each token in the second token sequence. The NLG system (or at least one of its subsystems) can generate a second complete sentence including the second token sequence. The NLG system (or at least one of its subsystems) can generate a second NLI score for the second complete sentence based on the fidelity of the second complete sentence to the input content. The NLG system (or at least one of its subsystems) can adjust the second confidence level of the second token sequence based on the second NLI score of the second complete sentence to generate the second updated confidence level of the second token sequence. In some aspects, the output text is configured to generate a summary of the input content.
[0135] In some aspects, the output text is configured to generate a summary of the input content (e.g., as in FIG. 8A to FIG. 8B ’s news article summary generator).
[0136] In some aspects, the input content includes input text. In some aspects, the at least one token is at least a portion of a word (e.g., such as Figure 4A , Figure 4B , Fig. 8A or Figure 8B In some aspects, each token in the sequence of tokens is at least a portion of a corresponding word.
[0137] In some aspects, the plurality of tokens is also based on at least one previously generated output token of the output text. Figure 4A In , “nice” may be a previously generated output token for “woman”, and “woman” may be generated or selected based on the previously generated output token “nice”. Similarly, in Figure 4A , “The” may be a previously generated output token for “nice”, and “nice” may be generated or selected based on the previously generated output token “The”.
[0138] In some aspects, searching the plurality of tokens to generate the first ordering includes using a beam search (e.g., as in Figure 4B and Figure 6 In some aspects, searching the plurality of tokens to generate the first ordering includes using a greedy search (e.g., as in Figure 4A , Fig. 7A and Figure 7B middle).
[0139] In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of outputting the output text. In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of causing a display to display the output text. In some aspects, the NLG system (or at least one of its subsystems) is configured and capable of causing a communication interface to send the output text to a recipient device.
[0140] In some examples, the NLG system includes: a component for generating a plurality of tokens based on input content; a component for searching the plurality of tokens to generate a first ordering of the plurality of tokens based on probability; a component for generating a natural language inference (NLI) score for the plurality of tokens to generate a second ordering of the plurality of tokens based on faithfulness to the input content; and a component for generating an output text, the output text including at least one token selected from the plurality of tokens based on the first ordering and the second ordering. The components for performing these operations may include, for example, NLG system 300, NLG system 350, encoder 304, decoder 306, decoder with hallucination mitigation 310, decoder 600, transformer block 604, beam search 606, NLI scorer 608, decoder 700, decoder 750, greedy expansion 704, beam reranker with weighted NLI score and model probability 708, greedy expansion with saliency enhancement 712, model 900, NN 1100, computing system 1200, or a combination thereof.
[0141] In some examples, the processes described herein (e.g., process 1000 and / or any other process described herein) may be performed by a computing device or apparatus. In one example, process 1000 may be performed by: NLG system 300, NLG system 350, encoder 304, decoder 306, decoder with hallucination mitigation 310, decoder 600, transformer block 604, beam search 606, NLI scorer 608, decoder 700, decoder 750, greedy expansion 704, beam reranker with weighted NLI score and model probability 708, greedy expansion with saliency enhancement 712, model 900, NN 1100, computing system 1200, or a combination thereof. For example, with Fig.12 The computing device of the computing device architecture of the computing system 1200 shown may implement Fig.10 The operation and / or this article about Figure 3A , Figure 3B , Figure 6 , Fig. 7A , Figure 7B , Fig. 9 , Fig.11 and / or Fig.12 The components and / or operations described in any diagrams.
[0142] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, an XR device (e.g., a VR headset, an AR headset, AR glasses, etc.), a wearable device (e.g., a connected watch or smart watch or other wearable device), a server computer, a vehicle (e.g., an autonomous vehicle) or a computing device of a vehicle, a robotic device, a laptop, a smart TV, a camera, and / or any other computing device having resource capabilities to perform the processes described herein (including process 1000 and / or any other processes described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.
[0143] The components of the computing device may be implemented in circuits. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0144] Process 1000 is illustrated as a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer-executable instructions include routines, programs, objects, components, and data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0145] Additionally, 1000 and / or any other process described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, through hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes multiple instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0146] As described in this article, Fig.11 The neural network system 1100 can be implemented using one neural network or multiple neural networks. Fig.11 It can be Fig.11 1. An illustrative example of a deep learning neural network 1100 used by a neural network system 1100 of the present invention. An input layer 1120 includes input data. In an illustrative example, the input layer 1120 may include data representing pixels of an input video frame. The neural network 1100 includes a plurality of hidden layers 1122a, 1122b to 1122n. The hidden layers 1122a, 1122b to 1122n include "n" hidden layers, where "n" is an integer greater than or equal to one. The plurality of hidden layers may include as many layers as required for a given application. The neural network 1100 also includes an output layer 1124 that provides outputs resulting from processing performed by the hidden layers 1122a, 1122b to 1122n. In an illustrative example, the output layer 1124 may provide a classification of an object in an input video frame. The classification may include a category that identifies a type of object (e.g., a person, dog, cat, or other object).
[0147] Neural network 1100 is a multi-layer neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information as it processes it. In some cases, neural network 1100 may include a feedforward network, in which case there is no feedback connection in which the output of the network is fed back into itself. In some cases, neural network 1100 may include a recurrent neural network, which may have loops that allow information to be carried across nodes when reading in input.
[0148] Information can be exchanged between nodes by node-to-node interconnection between each layer. The nodes of input layer 1120 can activate the set of nodes in the first hidden layer 1122a. For example, as shown in the figure, each input node in the input node of input layer 1120 is connected to each node in the node of the first hidden layer 1122a. The nodes of hidden layers 1122a, 1122b to 1122n can transform the information by applying activation function to the information of each input node. The information derived from the transformation can then be passed to the node of the next hidden layer 1122b and can activate the node of the next hidden layer, and the node of the next hidden layer can perform their own specified functions. Example functions include convolution, upsampling, data conversion and / or any other suitable function. Then, the output of hidden layer 1122b can activate the node of the next hidden layer, and so on. The output of last hidden layer 1122n can activate one or more nodes of output layer 1124, and output is provided at the one or more nodes. In some cases, although a node in neural network 1100 (e.g., node 1126) is shown as having multiple output lines, the node has a single output, and all lines shown as output from the node represent the same output value.
[0149] In some cases, each node or interconnection between nodes may have a weight, which is a set of parameters derived from the training of the neural network 1100. Once the neural network 1100 is trained, it may be referred to as a trained neural network, which may be used to classify one or more objects. For example, an interconnection between nodes may represent a piece of information learned about the interconnected nodes. The interconnection may have a tunable digital weight that may be tuned (e.g., based on a training data set), thereby allowing the neural network 1100 to adapt to the input and be able to learn as more and more data is processed.
[0150] The neural network 1100 is pre-trained to process features from the data in the input layer 1120 using different hidden layers 1122a, 1122b to 1122n to provide an output through the output layer 1124. In examples where the neural network 1100 is used to identify objects in images, the neural network 1100 may be trained using training data that includes both images and labels. For example, training images may be input into the network, where each training image has a label indicating the class of one or more objects in each image (essentially, indicating to the network what the objects are and what features they have). In one illustrative example, the training images may include an image of the number 2, in which case the label for the image may be [00 1 0 0 0 0 0 0 0].
[0151] In some cases, the neural network 1100 may use a training process known as back propagation to adjust the weights of the nodes. Back propagation may include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. For each set of training images, this process may be repeated a certain number of iterations until the neural network 1100 is trained well enough so that the weights of each layer are accurately tuned.
[0152] For the example of identifying an object in an image, a forward pass may include passing a training image through the neural network 1100. Prior to training the neural network 1100, the weights are initially randomized. The image may include, for example, an array of numbers representing pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 digital array having 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or a brightness and two chrominance components, etc.).
[0153] For the first training iteration of the neural network 1100, the output will likely include values that do not prioritize any particular class due to the random selection of weights at initialization. For example, if the output is a vector with probabilities that an object includes different classes, the probability values for each of the different classes may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 1100 is unable to determine low-level features and therefore cannot make an accurate determination of what the classification of an object may be. A loss function may be used to analyze the errors in the output. Any suitable loss function definition may be used. An example of a loss function includes mean squared error (MSE). It calculates the sum of the ground truth output (e.g., actual answer) minus one-half the square of the predicted output (e.g., predicted answer). The loss can be set equal to E 总 The value of .
[0154] For the first training images, the loss (or error) will be high because the actual values will be very different from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. The neural network 1100 can perform a backward pass by determining which inputs (weights) contribute most to the network's loss, and the weights can be adjusted so that the loss is reduced and ultimately minimized.
[0155] The derivative of the loss with respect to the weight (expressed as dL / dW, where W is the weight at a particular layer) can be calculated to determine the weight that contributes most to the loss of the network. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be expressed as Where w represents the weight, w i Represents the initial weight, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes a larger weight update, while a lower value indicates a smaller weight update.
[0156] In some cases, neural network 1100 may be trained using self-supervised learning.
[0157] Neural network 1100 may include any suitable deep network. One example includes a convolutional neural network (CNN) that includes an input layer and an output layer with multiple hidden layers between the input layer and the output layer. Fig.12 An example of a CNN is described. The hidden layer of the CNN includes a series of convolutional layers, nonlinear layers, pooling layers (for downsampling), and fully connected layers. The neural network 1100 may include any other deep network other than a CNN, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.
[0158] Fig.12 is a diagram illustrating an example of a system for implementing certain aspects of the present disclosure. Specifically, Fig.12 An example of a computing system 1200 is illustrated, which may be, for example, any computing device that constitutes a computing system, a camera system, or any component thereof, where the components of the system communicate with each other using a connection 1205. The connection 1205 may be a physical connection using a bus, or a direct connection into the processor 1210, such as in a chipset architecture. The connection 1205 may also be a virtual connection, a networked connection, or a logical connection.
[0159] In some examples, computing system 1200 is a distributed system, where the functionality described in the present disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more of the described system components represent a number of such components, each component performing some or all of the functionality for which the component is described. In some examples, each component may be a physical or virtual device.
[0160] The example system 1200 includes at least one processing unit (CPU or processor) 1210 and connections 1205 that couple various system components including system memory 1215, such as read only memory (ROM) 1220 and random access memory (RAM) 1225, to the processor 1210. The computing system 1200 may include a cache 1212 of high-speed memory directly connected to, close to, or integrated as part of the processor 1210.
[0161] Processor 1210 may include any general purpose processor and hardware or software services, such as services 1232, 1234, and 1236 stored in storage device 1230, that are configured to control processor 1210 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 1210 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0162] To enable user interaction, the computing system 1200 includes an input device 1245 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 1200 may also include an output device 1235, which may be one or more of a plurality of output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 1200. The computing system 1200 may include a communication interface 1240, which may generally govern and manage user input and system output.
[0163] The communication interface can perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using audio jacks / plugs, microphone jacks / plugs, universal serial bus (USB) ports / plugs, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Dedicated wired ports / plugs, Wireless signal transmission, Low energy (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 1202.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.
[0164] The communication interface 1240 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1200 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's Beidou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no limitation to operating on any particular hardware arrangement, and thus the underlying features herein may be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0165] The storage device 1230 may be a non-volatile and / or non-transitory and / or computer-readable memory device, and may be a hard disk or other type of computer-readable medium that can store data that can be accessed by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disk-read only memory (CD-ROM) optical disk, a rewritable compact disk (CD) optical disk, a digital video disk (DVD) optical disk, a Blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.
[0166] Storage device 1230 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 1210, causes the system to perform functions. In some examples, hardware services that perform specific functions may include software components stored in a computer-readable medium connected to necessary hardware components (such as processor 1210, connection 1205, output device 1235, etc.) to perform functions. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media that can store data and does not include carrier waves and / or transient electronic signals that propagate wirelessly or over a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer readable media may store thereon code and / or machine executable instructions, which may represent a process, function, subprogram, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.
[0167] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals including bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0168] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that embodiments may be practiced without these specific details. For the sake of clarity of explanation, in some cases, the present technology may be presented as including a separate functional block, which includes a device, device component, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein may be used. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid these embodiments becoming incomprehensible in unnecessary details. In other cases, known circuits, processes, algorithms, structures and techniques may be shown without necessary details to avoid making each embodiment become incomprehensible.
[0169] Individual embodiments may be described above as processes or methods depicted as flow charts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although flow charts may describe operations as sequential processes, many operations in the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operations of the process are completed, but the process may have additional steps not included in the accompanying drawings. The process may correspond to a method, function, process, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or main function.
[0170] The processes and methods according to the above examples can be implemented using stored computer executable instructions or computer executable instructions otherwise obtained from computer readable media. Such instructions may include, for example, instructions and data that enable or configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions in other ways. Parts of the computer resources used can be accessed through a network. Computer executable instructions can be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0171] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of additional examples, such functionality may also be implemented on different chips or circuit boards between different processes executed on a single device.
[0172] The instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0173] In the foregoing description, various aspects of the present application are described with reference to the specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and adopted in various other ways, and the appended claims are intended to be interpreted as including such variations, unless limited by the prior art. Various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, without departing from the broader essence and scope of this specification, the embodiments can be used in any number of environments and applications beyond the environment and application described herein. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For the purpose of illustration, each method is described in a specific order. It should be understood that in an alternative embodiment, each method can be performed in a different order than described.
[0174] It should be understood by those of ordinary skill in the art that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.
[0175] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0176] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component is directly or indirectly in communication with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0177] Claim language or other language in this disclosure that states "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language that states "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0178] Various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints proposed for the entire system. The technician may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of the present application.
[0179] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices with multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium including a program code, which includes instructions for executing one or more of the above methods, algorithms, and / or operations when executed. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0180] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in the present disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the techniques described herein.
[0181] Exemplary aspects of the present disclosure include:
[0182] Aspect 1. A device for natural language processing, the device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory, the at least one processor being configured to: generate a token sequence based on input content; determine a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; generate a complete sentence including the token sequence; generate a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and adjust the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level for the token sequence.
[0183] Aspect 2. According to the apparatus of aspect 1, the at least one processor is configured to: generate the token sequence using a beam search based on the input content.
[0184] Aspect 3. According to the apparatus according to any one of aspects 1 to 2, the at least one processor is configured to: generate the complete sentence using a greedy search based on the token sequence.
[0185] Aspect 4. According to the apparatus according to any one of Aspects 1 to 3, the at least one processor is configured to: limit the candidate tokens used to generate the complete sentence based on whether the corresponding significance value of the candidate token exceeds a significance threshold.
[0186] Aspect 5. An apparatus according to Aspect 4, wherein the significance threshold is based on an average of the corresponding significance values of the candidate tokens.
[0187] Aspect 6. According to the device described in any one of Aspects 1 to 5, the at least one processor is configured to: sort the token sequence against the second token sequence based on the confidence level associated with the token sequence and the second confidence level associated with the second token sequence.
[0188] Aspect 7. According to the apparatus of Aspect 6, the at least one processor is configured to: reorder the token sequence against the second token sequence based on the updated confidence level associated with the token sequence and the second updated confidence level associated with the second token sequence, wherein the second updated confidence level is based on a second NLI score of a second complete sentence, and the second complete sentence is generated based on the second token sequence.
[0189] Aspect 8. According to the apparatus of Aspect 7, the at least one processor is configured to: reorder the token sequence based on comparison with the second token sequence, select the highest-ranked token sequence from at least the token sequence and the second token sequence; and generate an output text including the highest-ranked token sequence.
[0190] Aspect 9. The apparatus according to aspect 8, wherein the output text is configured to generate a summary of the input content.
[0191] Aspect 10. According to the apparatus of any one of Aspects 1 to 9, the at least one processor is configured to: generate an output text including the token sequence based on the updated confidence level of the token sequence exceeding the second updated confidence level of the second token sequence.
[0192] Aspect 11. According to the apparatus of Aspect 10, the at least one processor is configured to: generate the second token sequence based on the input content; determine the second confidence level associated with the second token sequence based on the secondary corresponding confidence level associated with each token in the second token sequence; generate a second complete sentence including the second token sequence; generate a second NLI score for the second complete sentence based on the fidelity of the second complete sentence to the input content; and adjust the second confidence level of the second token sequence based on the second NLI score of the second complete sentence to generate the second updated confidence level of the second token sequence.
[0193] Aspect 12. The apparatus according to any one of Aspects 10 to 11, wherein the output text is configured to generate a summary of the input content.
[0194] Aspect 13. An apparatus according to any one of Aspects 1 to 12, wherein the NLI score identifies whether at least a portion of the complete sentence is true, false, or neutral.
[0195] Aspect 14. The apparatus according to any one of aspects 1 to 13, wherein the input content comprises input text.
[0196] Aspect 15. An apparatus according to any one of aspects 1 to 14, wherein each token in the sequence of tokens is at least a part of a corresponding word.
[0197] Aspect 16. An apparatus according to any one of Aspects 1 to 15, wherein the token sequence is configured to follow a previously determined token sequence in the complete sentence, wherein the complete sentence includes the previously determined token sequence, the token sequence and at least one additional token.
[0198] Aspect 17. According to the apparatus of any one of aspects 1 to 16, the at least one processor is configured to: generate the token sequence using a greedy search based on the input content.
[0199] Aspect 18. An apparatus according to any one of Aspects 1 to 17, wherein the at least one processor is configured to: output an output text including the token sequence.
[0200] Aspect 19. An apparatus according to any one of aspects 1 to 18, wherein the at least one processor is configured to: cause a display to display output text including the token sequence.
[0201] Aspect 20. An apparatus according to any one of aspects 1 to 19, the apparatus further comprising: a communication interface configured to send an output text including the token sequence to a recipient device.
[0202] Aspect 21. An apparatus according to any one of aspects 1 to 20, wherein the apparatus comprises at least one of a head mounted display (HMD), a mobile phone, or a wireless communication device.
[0203] Aspect 22. A method for natural language processing, the method comprising: generating a token sequence based on input content; determining a confidence level associated with the token sequence based on a corresponding confidence level associated with each token in the token sequence; generating a complete sentence including the token sequence; generating a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; and adjusting the confidence level of the token sequence based on the NLI score of the complete sentence to generate an updated confidence level for the token sequence.
[0204] Aspect 23. The method according to aspect 22, further comprising: generating the token sequence using a beam search based on the input content.
[0205] Aspect 24. The method according to any one of aspects 22 to 23, further comprising: generating the complete sentence using a greedy search based on the token sequence.
[0206] Aspect 25. The method according to any one of aspects 22 to 24, further comprising: limiting the candidate tokens for generating the complete sentence based on whether the corresponding significance value of the candidate token exceeds a significance threshold.
[0207] Aspect 26. The method according to aspect 25, wherein the significance threshold is based on an average of the corresponding significance values of the candidate tokens.
[0208] Aspect 27. The method according to any one of Aspects 22 to 26, further comprising: sorting the token sequence against the second token sequence based on the confidence level associated with the token sequence and a second confidence level associated with the second token sequence.
[0209] Aspect 28. The method according to Aspect 27 further includes: reordering the token sequence against the second token sequence based on the updated confidence level associated with the token sequence and a second updated confidence level associated with the second token sequence, wherein the second updated confidence level is based on a second NLI score of a second complete sentence, and the second complete sentence is generated based on the second token sequence.
[0210] Aspect 29. The method according to Aspect 28 further includes: reordering the token sequence based on comparison with the second token sequence, selecting the highest-ranked token sequence from at least the token sequence and the second token sequence; and generating an output text including the highest-ranked token sequence.
[0211] Aspect 30. The method according to aspect 29, wherein the output text is configured to generate a summary of the input content.
[0212] Aspect 31. A method according to any one of Aspects 22 to 30, further comprising: generating an output text including the token sequence based on the updated confidence level of the token sequence exceeding the second updated confidence level of a second token sequence.
[0213] Aspect 32. The method according to Aspect 31 further includes: generating the second token sequence based on the input content; determining a second confidence level associated with the second token sequence based on a secondary corresponding confidence level associated with each token in the second token sequence; generating a second complete sentence including the second token sequence; generating a second NLI score for the second complete sentence based on the fidelity of the second complete sentence to the input content; and adjusting the second confidence level of the second token sequence based on the second NLI score of the second complete sentence to generate the second updated confidence level of the second token sequence.
[0214] Aspect 33. The method according to any one of Aspects 31 to 32, wherein the output text is configured to generate a summary of the input content.
[0215] Aspect 34. A method according to any one of Aspects 22 to 33, wherein the NLI score identifies whether at least a portion of the complete sentence is true, false, or neutral.
[0216] Aspect 35. A method according to any one of Aspects 22 to 34, wherein the input content comprises input text.
[0217] Aspect 36. A method according to any one of Aspects 22 to 35, wherein each token in the sequence of tokens is at least a part of a corresponding word.
[0218] Aspect 37. A method according to any one of Aspects 22 to 36, wherein the token sequence is configured to follow a previously determined token sequence in the complete sentence, wherein the complete sentence includes the previously determined token sequence, the token sequence and at least one additional token.
[0219] Aspect 38. The method according to any one of aspects 22 to 37, further comprising: generating the token sequence using a greedy search based on the input content.
[0220] Aspect 39. The method according to any one of Aspects 22 to 38, further comprising: outputting an output text including the token sequence.
[0221] Aspect 40. The method according to any one of aspects 22 to 39, further comprising: causing a display to display output text including the token sequence.
[0222] Aspect 41. The method according to any one of aspects 22 to 40, further comprising: causing the communication interface to send an output text including the token sequence to a recipient device.
[0223] Aspect 42. A method according to any one of aspects 22 to 41, wherein the method is performed using an apparatus comprising at least one of a head mounted display (HMD), a mobile phone, or a wireless communication device.
[0224] Aspect 43. A non-transitory computer-readable storage medium, comprising instructions stored thereon, which, when executed by at least one processor, cause the at least one processor to perform operations according to any one of Aspects 1 to 42.
[0225] Aspect 44. An apparatus comprising means for performing the operations according to any one of aspects 1 to 42.
Claims
1. A device for natural language processing, the device comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory, the at least one processor configured to: Generate a token sequence based on the input content; determining a confidence level associated with the sequence of tokens based on a respective confidence level associated with each token in the sequence of tokens; generating a complete sentence including the token sequence; generating a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; as well as The confidence level of the token sequence is adjusted based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence.
2. The apparatus of claim 1 , wherein the at least one processor is configured to: The token sequence is generated using a beam search based on the input content.
3. The apparatus of claim 1 , wherein the at least one processor is configured to: The complete sentence is generated using a greedy search based on the token sequence.
4. The apparatus of claim 1 , wherein the at least one processor is configured to: The candidate tokens are restricted for use in generating the complete sentence based on whether their corresponding salience values exceed a salience threshold.
5. An apparatus according to claim 4, wherein the significance threshold is based on an average of the corresponding significance values of the candidate tokens.
6. The apparatus of claim 1 , wherein the at least one processor is configured to: The sequence of tokens is ranked against a second sequence of tokens based on the confidence level associated with the sequence of tokens and a second confidence level associated with a second sequence of tokens.
7. The apparatus of claim 6, wherein the at least one processor is configured to: The token sequence is reordered against the second token sequence based on the updated confidence level associated with the token sequence and a second updated confidence level associated with the second token sequence, wherein the second updated confidence level is based on a second NLI score of a second complete sentence generated based on the second token sequence.
8. The apparatus of claim 7, wherein the at least one processor is configured to: selecting a highest ranked token sequence from at least the token sequence and the second token sequence based on reordering the token sequence against the second token sequence; and An output text is generated including the highest ranked token sequence. 9 . The apparatus according to claim 8 , wherein the output text is configured to generate a summary of the input content.
10. The apparatus of claim 1, wherein the at least one processor is configured to: Based on the updated confidence level for the sequence of tokens exceeding a second updated confidence level for a second sequence of tokens, output text is generated that includes the sequence of tokens.
11. The apparatus of claim 10, wherein the at least one processor is configured to: generating the second token sequence based on the input content; determining a second confidence level associated with the second sequence of tokens based on the secondary corresponding confidence level associated with each token in the second sequence of tokens; generating a second complete sentence including the second token sequence; generating a second NLI score for the second complete sentence based on the fidelity of the second complete sentence to the input content; as well as The second confidence level of the second token sequence is adjusted based on the second NLI score of the second complete sentence to generate the second updated confidence level of the second token sequence. The apparatus according to claim 10 , wherein the output text is configured to generate a summary of the input content.
13. The apparatus of claim 1, wherein the NLI score identifies whether at least a portion of the complete sentence is true, false, or neutral. The apparatus according to claim 1 , wherein the input content comprises input text.
15. The apparatus of claim 1, wherein each token in the sequence of tokens is at least a portion of a corresponding word.
16. The apparatus of claim 1, wherein the token sequence is configured to follow a previously determined token sequence in the complete sentence, wherein the complete sentence comprises the previously determined token sequence, the token sequence, and at least one additional token.
17. The apparatus of claim 1, wherein the at least one processor is configured to: The token sequence is generated using a greedy search based on the input content.
18. The apparatus of claim 1, wherein the at least one processor is configured to: Outputs the output text including the sequence of tokens.
19. The apparatus of claim 1, wherein the at least one processor is configured to: The display is caused to display output text including the token sequence.
20. The apparatus according to claim 1, further comprising: A communication interface is configured to send an output text including the token sequence to a recipient device.
21. The apparatus of claim 1, wherein the apparatus comprises at least one of a head mounted display (HMD), a mobile handset, or a wireless communication device.
22. A method for natural language processing, the method comprising: Generate a token sequence based on the input content; determining a confidence level associated with the sequence of tokens based on a respective confidence level associated with each token in the sequence of tokens; generating a complete sentence including the token sequence; generating a natural language inference (NLI) score for the complete sentence based on the fidelity of the complete sentence to the input content; as well as The confidence level of the token sequence is adjusted based on the NLI score of the complete sentence to generate an updated confidence level of the token sequence.
23. The method according to claim 22, further comprising: The token sequence is generated using a beam search based on the input content.
24. The method according to claim 22, further comprising: The complete sentence is generated using a greedy search based on the token sequence.
25. The method according to claim 22, further comprising: The candidate tokens are restricted for use in generating the complete sentence based on whether their corresponding salience values exceed a salience threshold.
26. The method according to claim 22, further comprising: The sequence of tokens is ranked against a second sequence of tokens based on the confidence level associated with the sequence of tokens and a second confidence level associated with a second sequence of tokens.
27. The method according to claim 26, further comprising: The token sequence is reordered against the second token sequence based on the updated confidence level associated with the token sequence and a second updated confidence level associated with the second token sequence, wherein the second updated confidence level is based on a second NLI score of a second complete sentence generated based on the second token sequence.
28. The method of claim 22, further comprising: Based on the updated confidence level for the sequence of tokens exceeding a second updated confidence level for a second sequence of tokens, output text is generated that includes the sequence of tokens.
29. The method of claim 22, further comprising: The token sequence is generated using a greedy search based on the input content.
30. The method of claim 22, further comprising: Outputs the output text including the sequence of tokens.