Automatically generate diverse texts

Through reinforcement learning algorithm combined with problem generator and evaluation manager, the generator is dynamically adjusted to generate synthesis problems of diversity and accuracy, solving the problem of insufficient diversity and accuracy in the existing technology, and improving the quality of generated text.

CN115298659BActive Publication Date: 2025-07-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180022536.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2021-03-08
Publication Date
2025-07-22
Estimated Expiration
2041-03-08

AI Technical Summary

Technical Problem

When the prior art generates synthetic questions corresponding to the basic fact paragraphs, it is difficult to ensure the diversity and accuracy of the problems at the same time, resulting in high repetition of the generated text or the answers are irrelevant.

Method used

Using reinforcement learning algorithms combined with question generator, answer manager and evaluation manager, we evaluate the diversity and accuracy of synthetic questions through reward functions, and dynamically adjust the question generator to improve text diversity and maintain the accuracy of answers.

Benefits of technology

It realizes the generation of diverse and accurate synthesis questions, reduces text repetition and improves the relevance and quality of answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115298659B_ABST
    Figure CN115298659B_ABST
Patent Text Reader

Abstract

Embodiments relate to an artificial intelligence (AI) computer platform for combining synthetic data and ground truth data and for facilitating the generation of diversity and accuracy of synthetic data. A synthetic question is generated by a question generator in response to semantically related ground truth paragraphs and answer data. Each generated question is presented to an answer generator together with the semantically related ground truth paragraph. The diversity of each synthetic question is evaluated relative to previous synthetic questions generated for the same ground truth paragraphs and answer data. Each synthetic question is also evaluated relative to the accuracy of the answer generated by the answer generator. A reward function that captures both the accuracy and diversity of each synthetic question is used to selectively modify the question generator, where the selective modification is aimed at increasing text diversity and maintaining the accuracy of the generated synthetic questions.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This embodiment relates to an artificial intelligence platform and an optimization method for automatically generating diverse texts corresponding to received input paragraphs and input answers. More specifically, this embodiment aims to evaluate the generated texts in terms of quality (especially diversity and accuracy) to selectively adjust one or more parameters of a reinforcement learning algorithm to support the diversity and quality characteristics of the generated texts. Summary of the Invention

[0002] The embodiment includes a computer system, a computer program product, and a method for reinforcement learning.

[0003] In one aspect, there is provided a computer system having a processor and a memory for use with an artificial intelligence (AI) platform and corresponding AI tools. The processor is operably coupled to the memory and communicates with the AI platform. The AI platform tools include a question manager, an answer manager, an evaluation manager, and a director. The question manager is configured to generate a first synthetic question corresponding to a ground truth passage and a ground truth answer using a question generator. The answer manager is configured to present the synthetic question to a trained neural network and obtain a first answer from the ground truth passage. The first answer is semantically related to the first synthetic question and the ground truth passage. The evaluation manager is configured to automatically evaluate the first synthetic question against the first answer, where the evaluation incorporates a reward function that captures the diversity and accuracy of the first synthetic question. The director is configured to selectively modify the question generator in response to the evaluation using reinforcement learning, where the modified question generator produces one or more synthetic questions having increased text diversity compared to before the modification.

[0004] In another aspect, there is provided a computer program device supporting reinforcement learning. The computer program product includes a computer-readable storage medium having program code embodied therewith. The program code is executable by a processor to generate synthetic questions to further produce textually diverse synthetic questions and accurate answers. The program code uses a question generator to generate a first synthetic question corresponding to a ground truth passage and a ground truth answer. The program code submits the synthetic question to a trained neural network and obtains a first answer from the ground truth passage, where the first answer is semantically related to the first synthetic question and the ground truth passage. The program code is provided to automatically evaluate the first synthetic question against the first answer, where the evaluation incorporates a reward function that captures the diversity and accuracy of the first synthetic question. The program code uses reinforcement learning to selectively modify the question generator in response to the evaluation, where the modified question generator produces one or more synthetic questions having increased text diversity compared to before the modification.

[0005] In another aspect, a method for supporting reinforcement learning is provided. A first synthetic question corresponding to a ground truth passage and a ground truth answer is generated from a question generator. The synthetic question is presented to a trained neural network, and a first answer from the ground truth passage is obtained from the neural network. The first answer is semantically related to the first synthetic question and the ground truth passage. The first synthetic question is automatically evaluated against the first answer. The evaluation incorporates a reward function that captures the diversity and accuracy of the first synthetic question. Reinforcement learning is used to selectively modify the question generator based on the evaluation, where the modified question generator produces one or more synthetic questions having increased text diversity compared to before the modification.

[0006] These and other features and advantages will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate some, but not all embodiments of the present disclosure. The features shown in the drawings are merely illustrative of some embodiments and not all embodiments.

[0008] Figure 1 A system diagram depicting a computer system is shown.

[0009] Figure 2 A block diagram depicting tools from a computing system and their associated application programming interfaces is shown.

[0010] Figure 3 A flowchart depicting the application of reinforcement learning to support and enable the generation of diverse and accurate questions is shown.

[0011] Figure 4 A block diagram depicting an example of a computer system / server of a cloud-based support system for implementing the systems and processes described above with respect to Figures 1-3 is described.

[0012] Figure 5 A block diagram depicting a cloud computing environment is shown.

[0013] Figure 6 A block diagram depicting a set of functional abstraction model layers provided by a cloud computing environment is shown. DETAILED DESCRIPTION

[0014] It will be readily understood that the components of the present embodiments, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following detailed description of the embodiments of the apparatus, systems, methods, and computer program products of the present embodiments, as presented in the figures, is not intended to limit the scope of the claimed embodiments, but is merely representative of selected embodiments.

[0015] Throughout the specification, references to "selected embodiments", "an embodiment", or "embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment. Thus, the phrases "selected embodiments", "in an embodiment", or "in embodiments" appearing throughout this specification are not necessarily referring to the same embodiment. The various embodiments may be combined with each other.

[0016] The illustrated embodiments will be better understood by reference to the accompanying drawings, in which like parts are always designated by the same reference numerals. The following description is intended only as an example and only shows certain selected embodiments of the devices, systems, and processes consistent with the embodiments.

[0017] Artificial intelligence (AI) relates to the field of computer science of computers and computer behavior related to humans. AI refers to the intelligence when a machine can make decisions based on information to maximize the chance of success in a given subject. More specifically, AI can learn from data sets to solve problems and provide relevant recommendations. For example, in the field of artificial intelligence computer systems, natural language systems (such as IBM artificial intelligence computer systems or other natural language question-and-answer systems) process natural language based on the knowledge obtained by the system. To process natural language, the system can be trained with data derived from a database or knowledge base, but for various reasons, the results may be incorrect or inaccurate.

[0018] Machine learning (ML) is a subset of artificial intelligence (AI) that uses algorithms to learn from data and create a forward look based on that data. More specifically, ML is an AI application that creates a neural network that can demonstrate learning behavior by performing tasks not explicitly programmed. ML requires data to be analyzed, formatted, and conditioned to build a machine learning model and train a machine learning algorithm. It can be understood in the art that an ML algorithm is a computerized process that generates an ML model when trained on data. Selecting an ML algorithm is necessary for the successful application of ML. Examples of ML include, but are not limited to, regression algorithms, decision trees, instance-based algorithms, and clustering algorithms. After the data is prepared and the algorithm is trained, the ML model can make determinations or predictions about the data. The larger the amount of data provided, the more the model learns and the more accurate its predictions become.

[0019] ML models fall into the following basic categories: supervised machine learning, unsupervised machine learning, reinforcement learning, and deep learning. Supervised learning algorithms learn a mapping function for a dataset with existing classifications, where unsupervised learning algorithms can classify unlabeled datasets based on some hidden features in the data. Reinforcement learning can learn a strategy for making decisions in an environment through iterative exploration of an uncertain environment. Deep learning combines neural networks in successive layers to learn from data iteratively. A neural network is a model of the way the nervous system operates. The basic units are called neurons, which are typically organized into layers. A neural network works by simulating interconnected processing units that are abstract versions of a large number of similar neurons. There are usually three parts in a neural network, including an input layer with units representing the input field, one or more hidden layers, and an output layer with one or more units representing the target field. These units are connected with different connection strengths or weights. Input data is provided to the first layer, and values propagate from each neuron to each neuron in the next layer. Eventually, the result is delivered from the output layer. Deep learning complex neural networks are designed to mimic how the human brain works, so computers can be trained to support ill-defined abstractions and problems. Neural networks and deep learning are often used in image recognition, speech, and computer vision applications.

[0020] A language model is a mechanism that takes a sequence of words such as phrases or partial phrases as input and produces an output in the form of a probability distribution of the next word. For each word in the vocabulary, the language model predicts the probability that the next word in the sequence is that particular word. The variable Y is used to represent the vocabulary filled with words, where the words Y0, Y1, Y2, ..., Y n-1 are the first n - 1 words in the sequence, and Y n is the next unknown word. The language model produces a probability assessment for each word y in the vocabulary Y, where the probability is expressed as follows:

[0021] Pr(Y n = y|y0, y1, y2, ..., y n-1 )

[0022] Similar to a language model, a sequence-to-sequence model maps a fixed-length input to a fixed-length output, where the fixed lengths of the input and output may be different. The sequence-to-sequence model takes a sequence x0, ..., x n , and produces a probability distribution over the output sequence y0, ..., y m . For example, for machine translation, the sequence X0, ..., X n represents a sentence with n words in the first language, and Y0, ..., Y mRepresents a sentence in a second language different from the first language with a potentially different number of words. A sequence-to-sequence model can be used for various tasks, including generating the next utterance produced by a chatbot, or the question answer for a given question, along with the corresponding passage as input.

[0023] The model typically generates text sequentially, for example, one word at a time. Generally, the probability distribution over the next word to be generated depends on the input and the words that have already been generated. The "next" word to be generated affects the generated sentence and, in one embodiment, changes the sentence to be generated. As shown and described herein, the model is modified to create diversity in the output. The sequence-to-sequence model takes a passage and an answer as input and outputs the corresponding question. In one embodiment, this model with question output is referred to as a question generator (QG). As shown and described here, the embodiments are intended to teach the QG to generate various questions and questions with desired semantics. Diversity is intended to mitigate or avoid repeatedly generating the same question, while mitigating or avoiding generating text that is irrelevant to the corresponding question and answer. Desired semantics is intended to generate questions that a reasonable mechanism would answer as desired.

[0024] The systems, computer program products, and methods shown and described herein are intended to teach a question generator (QG) to generate synthetic questions that address the characteristics of semantics and diversity. More specifically, the embodiments are intended to teach the QG to automatically enable the generation of diverse questions, limit the generation of reasonable text, evaluate the diversity of the generated questions, and evaluate the quality of the generated questions, i.e., accuracy, diversity, fluency, or other relevant attributes, through reinforcement learning.

[0025] Reference Figure 1, depicts a schematic diagram of a computing system (100). As shown, a server (110) is provided, which communicates with a plurality of computing devices (180), (182), (184), (186), (188) and (190) via links (102) and (104) through a computer network (105). The server (110) is configured with a processor (112) that communicates with a memory (116) via a bus (114). The server (110) is shown to have an artificial intelligence (AI) platform (150) for performing reinforcement learning from one or more of the computing devices (180), (182), (184), (186), (188) and (190) via the computer network (105). More specifically, the computing devices (180), (182), (184), (186), (188) and (190) communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, where each communication link can include one or more of wires, routers, switches, transmitters, receivers, etc. The server (110) can be used with components, systems, subsystems and / or devices other than those described here.

[0026] The AI platform (150) is shown here as being configured with tools to support machine learning, and more specifically, reinforcement learning with respect to synthetic data (e.g., machine-generated data) related to corresponding ground truth (GT) data (e.g., verified or labeled data). It can be understood in the art that synthetic data is artificially created, and in one embodiment, the synthetic data is created as the output from a corresponding neural model. The tools include but are not limited to a problem manager (152), an answer manager (154), an evaluation manager (156) and a director (158). The AI platform (150) utilizes and in one embodiment modifies a trained neural network. The AI platform (150) can receive inputs from the computer network (105) and utilize a data source (170) (also referred to here as a corpus or knowledge base) to selectively process data. The processing includes but is not limited to generating synthetic data and using the synthetic data to optimize the output from the selected or identified neural model.

[0027] As shown, the data source (170) is configured with a library (172) having a plurality of neural models, shown here as models A (174 A )、model B (174 B )、...、model N (174 N)。However, only three models are shown, and this number should not be construed as restrictive. Similarly, in an embodiment, the data source (170) may be understood to have one or more additional libraries, each library having one or more neural models. Each neural model in the neural models may combine the functions of the question generator QG and the answer generator AG into a single neural model. In an embodiment, the neural model may embody QG and AG as components of the model, such as embodied in the model A (174 A ) as the first question generator QG A,0 (176 A,0 ), the second question generator QG A,1 (176 A,1 ) and AG A (178 A ), embodied in the model B (174 B ) as QG B (176 B ) and AG B (178 B ), and embodied in the model N (172 N ) as QG N (176 N ) and AG N (178 N ). Details of how to utilize the model and how to modify the model in one embodiment are described in detail below.

[0028] It is understood that supervised learning utilizes the data reflected in one of the models in the data source. As shown here, the data source is referred to as the knowledge base (170) and is configured with a logical grouping model. The question manager (152) is used to identify the appropriate library and model that is semantically related to the corresponding GT data. After being identified, the question manager presents or inputs the GT first paragraph (160 A (174 A )) and the corresponding ground truth first answer (160 P ) to the identified model (e.g., the model A,0 ), which for the purpose of description includes a QG component, herein referred to as the first QG component QG A,0 (176 A,0 ) and an AG component AG A (178 A ). The output from the first QG component (e.g., QG A,0 (176 A,0 )) is in the form of a synthesized first question (160 Q’,0 ), and the synthesized first question (160 Q’,0 ) has the same relationship with the GT first paragraph (160P ) and the semantics corresponding to the GT first answer (160 A,0 ). Accordingly, in response to the submission presenting GT data, the question manager (152) receives or obtains the semantically corresponding first synthetic question (160 Q’,0 ) as the first GQ component QG A,0 (176 A,0 ) output.

[0029] The answer manager (154) (shown here operatively coupled to the question manager (152)) manages the processing of the synthetic question (160 Q’,0 ). More specifically, the answer manager (154) presents and inputs the synthetic question (160 A ) to the AG component (e.g., AG A )(178 Q’,0 ), corresponding to the identified or topic neural model (e.g., model A (174 A ). The AG component AG A (178 A ) generates the first answer (160 A,1 ), which is presented to the answer manager (154). The first answer (160 A ) generated by AG (AG A )(178 A,1 ) is semantically related to the synthetic question (160 Q’,0 ) and the GT passage (160 P ). Correspondingly, QG generates the synthetic question (160 Q’,0 ) and AG generates the answer (160 A,1 ), where the two generated components are semantically related to the GT passage (160 P ) and the GT answer (160 A,0 ).

[0030] The evaluation manager (156) is shown here operatively coupled to the question manager (152) and the answer manager (154). The evaluation manager (156) serves to automatically evaluate the generated synthetic question (160 A,1 ) against the generated first answer (160 Q’,0 ). The goal is to generate synthetic questions that meet the requirements of the corresponding semantic relationships and are also diverse, in order to enhance the output from AG while reducing the repetition of the same question. The evaluation manager (156) employs a reward function to capture the diversity and quality of the synthetic question (160 Q’,0 ). Thus, the evaluation manager (156) evaluates the diversity and quality and reflects the evaluation in the corresponding reward function.

[0031] The director (158) operably coupled to the evaluation manager (156) as shown herein is for selectively modifying the first question generator QG Q’,0 ) based on the evaluation of the first synthetic question (160 A,0 (176 A,0 ). The modification is selective with the aim of increasing the text diversity of the generated synthetic questions. As shown, the second question generator QG A,1 (176 A,1 ) is created by the question manager (152) and operably coupled to the model (the model A (174 A ). The second question generator QG A,1 (176 A,1 ) is a modification of the first question generator QG A,0 (176 A,0 ) and is used to generate a second synthetic question (160 Q’,1 ), which is different from the first synthetic question (160 Q’,0 ) while being accurate with respect to the corresponding GT passage (160 P ) and GT answer (160 A,0 ). For example, in an embodiment, the evaluation can direct and utilize the question manager (152) for processing, and the question manager (152) utilizes the modified QG A,1 (176 A,1 ) to generate the second synthetic question (160 Q’,1 ), and the answer manager (154) generates a second answer (160 Q’,1 ) based on the corresponding second synthetic question (160 A,2 ). In an embodiment, the optimization involves modifying the first question generator QG A,0 (176 A,0 ) for its parameters using a policy gradient algorithm based on a reward function. When generating the first synthetic question (160 Q’,0 ) and the second synthetic question (160 Q’,1 ), the director (158) employs sampling with text generation to promote the diversity of the synthetic questions. The sampling can be nucleus sampling, or in another embodiment, top-k sampling. The optimization of the question generator includes the director (158) adopting a diversity score to represent the first synthetic question (160 Q’,0 ) and modifying one or more parameters of the first question generator QG A,0 (176 A,0 ) so as to effectively create the second question generator QG A,1 (176 A,1 ). The optimization employed by the director (158) alleviates the problem of the model (the modelA (174 A ) for the second or subsequent synthesis problems of the output (160 Q’,1 ) that has a context similar to the first or previous synthesis problem (160 Q’,0 ). It can be understood that in an embodiment, the evaluation may indicate that the first synthesis problem (160 Q’,0 ) and the corresponding answer (160 A,0 ) have met the threshold or parameter of the reward function. Thus, the director (158) is interfaced with the evaluation manager (156) to selectively generate one or more additional synthesis problems.

[0032] As briefly described above, the synthesis problems are evaluated for diversity. In addition, the evaluation manager (156) evaluates or assesses the accuracy of the synthesis problems to ensure that the GT first paragraph (160 P ) can support the synthesis problems. The synthesis problems are provided to the answer generator AG A (178 A ), and the generated answers associated with each synthesis problem are compared with the GT answer (160 A,0 ). The evaluation requires the evaluation manager (156) to evaluate the output in the form of the first answer (160 A ) in response to the first synthesis problem (160 A ) from AG (AG Q’,0 )(178 A,0 ) and the second answer (160 Q’,1 ) in response to the second synthesis problem (160 A,1 ) with the GT answer (160 A,0) comparison. The evaluation manager (156) also quantifies the accuracy of the evaluation with an accuracy score. Based on the value assigned to the accuracy score, the question manager (152) will selectively modify one or more parameters of the generated one or more synthetic questions. The accuracy score and the diversity score described herein can be combined by the evaluation manager (156) to calculate a reward, and communicate with the question manager (152) to selectively modify one or more parameters of the generated synthetic questions based on the calculated reward, so as to mitigate or eliminate the problem that the question manager (152) generates synthetic questions that are semantically similar to one or more previously generated synthetic questions. In one embodiment, the reward is a convex combination of the accuracy and diversity scores. Similarly, in an embodiment, the evaluation manager (156) can selectively prioritize one or more components of the reward function by selectively adjusting the corresponding components or parameters of the convex combination. The response output (132) in the form of answer data identified in the paragraph by increasing the diversity in the generated one or more synthetic questions can be presented on an operatively coupled visual display (130) or transmitted to one or more of the computing devices (180), (182), (184), (186), (188), and (190).

[0033] As shown, in various embodiments, the computer network (105) can include local network connections and remote connections, such that the AI platform (150) can operate in environments of any size, including local and global, such as the Internet. Additionally, the AI platform (150) serves as a front-end system that can make available various knowledge extracted from network-accessible sources and / or structured data sources or represented in network-accessible sources and / or structured data sources. In this way, some processes utilize the artificial intelligence platform (150) to populate the AI platform (150), and the artificial intelligence platform (150) also includes an input interface to receive requests and respond accordingly.

[0034] The knowledge base (160) is configured with a library and one or more neural models, which are populated or logically grouped therein for use by the AI platform (150). Various computing devices (180), (182), (184), (186), (188), and (190) communicating with the computer network (105) can include access points for the logically grouped models. Some computing devices can include devices for a database that stores a corpus of data as a body of information, and the AI platform (150) uses this body of information to generate a response output (132) and transmit the response output (132) to a corresponding network device, such as a visual display (130), which is operatively coupled to the server (110) or one or more computing devices (180), (182), (184), (186), (188), and (190) on the computer network (105).

[0035] In some illustrative embodiments, the server (110) may be an IBM system available from International Business Machines Corporation of Armonk, New York, enhanced with the mechanisms of the illustrative embodiments described below. The managers (152), (154), (156), and (158), hereinafter collectively referred to as AI tools, are shown as embodied in or integrated into the AI platform (150) of the server (110). The AI tools may be implemented in a separate computing system (such as 190), or in one embodiment, the AI tools may be implemented in one or more systems connected to the server (110) via a computer network (105). Wherever embodied, the AI tools are used to dynamically optimize the generation of diverse and accurate synthetic questions to optimize the generated answers. The types of devices and corresponding systems that may use the artificial intelligence platform (150) range from small handheld devices such as handheld computers / mobile phones (180) to large mainframe systems such as large computers (182). Examples of handheld computers (180) include personal digital assistants (PDAs), personal entertainment devices such as MP4 players, portable TVs, and CD players. Other examples of information processing systems include pen or tablet computers (184), laptop or notebook computers (186), personal computer systems (188), and servers (190). As shown, the various devices and systems may be networked together using a computer network (105). The types of computer networks (105) that may be used to interconnect the various devices and systems include local area networks (LANs), wireless local area networks (WLANs), the Internet, public switched telephone networks (PSTNs), other wireless networks, and any other network topologies that may be used to interconnect devices and systems. Many devices and systems include non-volatile data storage, such as hard disk drives and / or non-volatile memory. Some devices and systems may use separate non-volatile data memories (e.g., the server (190) uses non-volatile data memory (190A), and the large computer (182) uses non-volatile data memory (182A). The non-volatile data memory (182A) may be a component external to the various devices and systems, or may be a component internal to one of the devices and systems.

[0036] The devices and systems used to support the artificial intelligence platform (150) may take many forms, some of which are

[0037] described below. Figure 1is shown. For example, an information processing system may take the form of a desktop computer, a server, a portable computer, a laptop computer, a notebook computer, or a computer or data processing system of other form factors. Additionally, the devices and systems may take other form factors, such as a personal digital assistant (PDA), a gaming device, an ATM machine, a portable telephone device, a communication device, or other devices including a processor and a memory.

[0038] An application programming interface (API) is understood in the art as a software intermediary between two or more applications. Regarding Figure 1 the AI platform (150) shown and described herein, one or more APIs may be utilized to support one or more of the tools (152), (154), (156), and (158) and their associated functions. Referring to Figure 2 , a block diagram (200) is provided showing at least some of the tools (252), (254), (256), and (258) and their associated APIs. As shown, a plurality of tools are embedded within the AI platform (205), where the tools include a question manager (152) associated with API0 (212) shown herein as (252), an answer manager (154) associated with API1 (222) shown herein as (254), an evaluation manager (156) associated with API2 (232) shown herein as (256), and a director (158) associated with API3 (242) shown herein as (258). Each API may be implemented in one or more languages and interface specifications. API0 (212) provides functional support to generate one or more synthetic questions, where each generated synthetic question semantically corresponds to the GT segment and the GT first answer. API1 (222) provides functional support for generating answers semantically related to the corresponding synthetic questions and GT paragraphs.

[0039] API2 (232) provides functional support to evaluate the generated answers related to the GT answers, which includes incorporating a reward function in the evaluation, and the reward function captures the diversity and accuracy of the first synthetic question, where the accuracy is measured using the first generated answer. API3 (242) provides functional support to selectively modify the question generator in response to answer evaluation, where the modification aims to increase text diversity and maintain the accuracy of the synthetic question(s). As shown, each of the APIs (212), (222), (232), and (242) is operatively coupled to an API coordinator (260), or a coordination layer, which is understood in the art as an abstraction layer used to transparently thread together separate APIs. In an embodiment, the functions of the separate APIs can be combined or integrated. Thus, the configuration of the APIs shown herein should not be considered restrictive. Accordingly, as shown herein, the functions of the tool can be embodied or supported by the corresponding APIs of the tool.

[0040] Reference Figure 3 , a flowchart (300) is provided to illustrate the process for applying reinforcement learning to a question generator to support and enable the generation of diverse and accurate synthetic questions. GT is a term used in machine learning, which refers to the training set for machine learning techniques, and it contains trusted reference facts. For example, the GT training set will contain the generated questions and answers that are known to be true by people. For example, the labels of the dataset classified as GT are considered accurate. As shown and described herein, and for the purpose of training the QG model, GT (302) is provided in the form of a paragraph P, a question Q, and a corresponding answer A. As shown, the paragraph P and the answer A are presented as inputs (304) to a trained neural network question generator (QG). In one embodiment, a second neural network is provided to generate an answer from the question. Similarly, in one embodiment, a single neural network is constructed for both the question generator and the answer system. The QG creates an output corresponding to the received input, that is, generates a synthetic question Q0′ (306) aligned with the presented paragraph P and answer A. The process of generating the synthetic question Q0′ sequentially generates the words q that form the synthetic question Q0′ j , and performs a diversity evaluation on the word q j . In an embodiment, as shown herein, the word q j is sampled, where at each word generation step, the possible next words are ranked. In the case of nucleus sampling, a probability mass is established, such as a probability limit, and the word q j. In an embodiment, the probability mass is a configurable parameter. Words forming and including the synthetic question are generated and evaluated until the end word of the synthetic question of characters is reached. After that, a synthetic question Q0′ is generated from the QG model using the words selected from the sampling. Thus, a synthetic question Q0′ generated from the QG model is created to align with the GT passage P and the GT answer A.

[0041] The synthetic questions generated in step (306) are further evaluated and refined in terms of diversity and accuracy. As shown, an Answer Generator (AG) model is used to generate a second output, herein referred to as answer A0′, which corresponds to the synthetic question Q0′ and the passage P (308). The accuracy reward of the answer is calculated. In an embodiment, the accuracy reward R acc is as follows:

[0042] |A ∩ A0′| / |A ∪ A0′|

[0043] The accuracy reward evaluates the similarity between the answer A0′ generated by the AG model and corresponding to the synthetic question and the GT answer A. In addition, a diversity reward r is also calculated for the generated synthetic question(s) div (310). In an embodiment, the diversity reward r div evaluates a diversity metric to quantify the novelty of the generated synthetic question(s) relative to the GT question Q and all the questions {Q′1,..., Q′ i-1} previously generated by QG for the GT passage P and answer A. The following is the representation of the quality reward R calculation:

[0044] R = w1 · r acc (A, A0′) + w2 · R div (Q i ′, Q, {Q′1,..., Q′ i-1})

[0045] where w1 and w2 are adjustable weights for the accuracy and diversity rewards respectively. Thus, the shown reward R is a weighted average of accuracy and diversity.

[0046] A loss is obtained based on the reward R and provided to the QG model to optimize the quality of one or more synthetic questions to be generated in the future (312). The parameters of the QG model are updated using this loss (314), and QG is adjusted based on the loss obtained in step (312) using an optimization algorithm (316). The optimizations in steps (314) and (316) use reinforcement learning during training. Figure 3The process shown can be repeated over all instances of the training examples present in the training set, and this process can be repeated multiple times over all such instances. Reinforcement learning is utilized to obtain and apply a loss to learn how to generate synthetic questions as output from the QG model, where the synthetic questions represent diverse content or context regarding other generated synthetic questions and accurate content or context regarding the GT passage P and answer A.

[0047] The tool and its corresponding functionality utilize reinforcement learning and a corresponding reinforcement learning algorithm to optimize the question generator based on the (one or more) synthetic questions and the corresponding (one or more) generated answers. The output of the reinforcement learning algorithm learns the values of different states, such as diversity and accuracy, and evaluates the corresponding rewards that incorporate these states. Reinforcement learning supports selectively adjusting the components of the reward, e.g., prioritizing one or more components, to enhance the diversity and / or accuracy of the synthetic questions and corresponding answers from the GT passage. Thus, the reinforcement learning shown and described herein dynamically learns the values of different states, which can then be applied to create the corresponding neural model.

[0048] The embodiments shown and described herein can be in the form of a computer system for use with an intelligent computer platform to provide and support reinforcement learning aimed at generating diverse synthetic questions. Aspects of the tools (152), (154), (156), and (158) and their associated functionality can be embodied in a computer system / server at a single location, or in an embodiment, aspects of the tools (152), (154), (156), and (158) and their associated functionality can be configured in a cloud-based system sharing computing resources. Referring Figure 4 to, a block diagram (400) is provided showing an example of a computer system / server (402), hereinafter referred to as the host (402), which communicates with a cloud-based support system (410) to implement the systems, tools, and processes described above with reference Figures 1-3 to. The host (402) can operate with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the host (402) include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and file systems (e.g., distributed storage environments and distributed cloud computing environments) including any one of the foregoing systems, devices, and their equivalents.

[0049] The host (402) may be described in the general context of computer system executable instructions, such as program modules executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The host (402) may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.

[0050] As Figure 4 shown, the host (402) is shown in the form of a general purpose computing device such as a computer system and / or a server. Components of the host (402) may include, but are not limited to, one or more processors or processing units (404), such as a hardware processor, a system memory (406), and a bus (408) that couples various system components including the system memory (406) to the processor (404). The bus (408) represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, these architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. The host (402) generally includes a variety of computer system readable media. Such media may be any available media accessible by the host (402), and it includes both volatile and nonvolatile media, removable and non-removable media.

[0051] The memory (406) may include computer system readable media in the form of volatile memory, such as random access memory (RAM) (430) and / or cache memory (432). By way of example only, a storage system (434) may be provided for reading from and writing to a non-removable, nonvolatile magnetic medium (not shown and typically referred to as a "hard disk drive"). Although not shown, a disk drive for reading from and writing to a removable, nonvolatile disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, nonvolatile optical disk such as a CD-ROM, DVD-ROM or other optical media may be provided. In such cases, each may be connected to the bus (408) by one or more data media interfaces.

[0052] A program / utility (440) having a set (at least one) of program modules (442), as well as an operating system, one or more application programs, other program modules, and program data can be stored in the memory (406), by way of example and not limitation. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, can include an implementation of a networked environment. The program modules (442) generally execute the functions and / or methods of the embodiments to dynamically coordinate activities across one or more domains, thereby minimizing risk. For example, the set of program modules (442) can include tools (152), (154), (156), and (158) as described in Figure 1 above.

[0053] The host (402) can also communicate with one or more external devices (414), such as a keyboard, pointing device, etc.; a display (424); one or more devices that enable a user to interact with the host (402); and / or any device that enables the host (402) to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface (422). Additionally, the host (402) can communicate with one or more networks via a network adapter (420), such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet). As described, the network adapter (420) communicates with other components of the host (402) via a bus (408). In one embodiment, multiple nodes of a distributed file system (not shown) communicate with the host (402) via the I / O interface (422) or via the network adapter (420). It should be understood that other hardware and / or software components can be used in conjunction with the host (402) although not shown. Examples include but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0054] In this document, the terms “computer program medium,” “computer usable medium,” and “computer readable medium” are used generally to refer to media such as main memory (406), including RAM (430), cache (432), and storage systems (434), such as removable storage drives and hard disks installed in hard disk drives.

[0055] A computer program (also known as computer control logic) is stored in a memory (406). The computer program can also be received via a communication interface, such as a network adapter (420). When run, the computer program enables the computer system to perform the features of the present embodiment as discussed herein. In particular, when run, the computer program enables the processing unit (404) to perform the features of the computer system. Thus, the computer program represents the controller of the computer system.

[0056] A computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage medium includes the following: a portable computer disk, a hard disk, a dynamic or static random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a magnetic storage device, a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0057] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0058] The computer-readable program instructions for performing the operations of this embodiment may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server cluster. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the embodiment.

[0059] In an embodiment, the host (402) is a node of a cloud computing environment. As is well known in the art, cloud computing is a service delivery model for enabling convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models. Examples of these characteristics are as follows:

[0060] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed, without human interaction with the service provider.

[0061] Wide area network access: The capabilities are available over a network and are accessed through a standard mechanism that facilitates use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptop computers, and PDAs).

[0062] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated according to demand. There is a location-independent meaning because consumers generally do not control or know the exact location of the resources provided, but are able to specify a location at a higher level of abstraction (e.g., country, state, or data center).

[0063] Rapid elasticity: In some cases, the ability to rapidly scale out and rapidly scale in can be provided quickly and elastically. To consumers, the available capacity often appears unlimited and can be purchased in any quantity at any time.

[0064] Measured service: The cloud system automatically controls and optimizes resource use by leveraging metering capabilities at an abstract level appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.

[0065] The service models are as follows:

[0066] Software as a Service (SaaS): The ability provided to the consumer is to use the provider's applications running on the cloud infrastructure. The applications can be accessed from various client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0067] Platform as a Service (PaaS): The ability provided to the consumer is to deploy the applications created or acquired by the consumer onto the cloud infrastructure, where the applications are created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but has control over the deployed applications and possibly the configuration of the application hosting environment.

[0068] Infrastructure as a Service (IaaS): The ability provided to the consumer is to provide processing, storage, networking, and other basic computing resources on which the consumer can deploy and run arbitrary software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0069] The deployment models are as follows:

[0070] Private cloud: The cloud infrastructure is operated only for an organization. It can be managed by the organization or a third party and can be located either inside or outside the building.

[0071] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can be located either on-premises or off-premises.

[0072] Public cloud: The cloud infrastructure is available for the general public or large industrial groups and is owned by an organization that sells cloud services.

[0073] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public), which remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0074] The cloud computing environment is service-oriented, with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure of a network including interconnected nodes.

[0075] Now refer to Figure 5 , an illustrative cloud computing network (500). As shown, the cloud computing network (500) includes a cloud computing environment (550) having one or more cloud computing nodes (510), and local computing devices used by cloud consumers can communicate with the cloud computing nodes. Examples of these local computing devices include, but are not limited to, a personal digital assistant (PDA) or cellular phone (554A), a desktop computer (554B), a laptop computer (554C), and / or an in-vehicle computer system (554N). Each of the nodes within the nodes (510) can also communicate with each other. The nodes can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud or a combination thereof as described above. This allows the cloud computing environment (500) to provide infrastructure, platform, and / or software as a service, and cloud consumers do not need to maintain resources on local computing devices for it. It should be understood that Figure 5 the types of computing devices (554A-N) shown in

[0076] Now refer to Figure 6 , a set of functional abstraction layers (600) provided by the Figure 5 cloud computing network is shown. It can be pre-understood that Figure 6 the components, layers, and functions shown in

[0077] are only for illustration, and the embodiments are not limited thereto. As depicted, the following layers and corresponding functions are provided: a hardware and software layer (610), a virtualization layer (620), a management layer (630), and a workload layer (640). System; a server based on the RISC (Reduced Instruction Set Computer) architecture, in one example IBM System; IBM System; IBM System; storage devices; networks and network components. Examples of software components include web application server software, in one example IBM Web application server software; and database software, in one example IBM Database software. (IBM, zSeries, pSeries, xSeries, BladeCerter, WebSphere, and DB2 are trademarks of International Business Machines Corporation registered in many jurisdictions worldwide).

[0078] The virtualization layer (620) provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers; virtual storage; virtual networks, including virtual private networks; virtual applications and operating systems; and virtual clients.

[0079] In one example, the management layer (630) can provide the following functions: resource provisioning, metering and pricing, user portal, service layer management, and SLA planning and fulfillment. Resource provisioning provides the dynamic procurement of computing resources and other resources used to perform tasks within a cloud computing environment. Metering and pricing provides cost tracking for resource utilization within a cloud computing environment, as well as billing or invoicing for the consumption of those resources. In one example, these resources can include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. The user portal provides access to the cloud computing environment for consumers and system administrators. Service layer management provides cloud computing resource allocation and management such that the required service layer is met. Service level agreement (SLA) planning and fulfillment provides the pre-arrangement and procurement of cloud computing resources, where future requirements are projected based on the SLA.

[0080] The workload layer (640) provides examples of functions that a cloud computing environment can be utilized for. Examples of workloads and functions that can be provided from this layer include but are not limited to: mapping and navigation; software development and lifecycle management; virtual classroom education delivery; data analysis processing; transaction processing; and synthetic data generation and reinforcement learning.

[0081] Although specific embodiments of the present invention have been shown and described, it will be apparent to those skilled in the art that, based on the teachings herein, changes and modifications can be made without departing from the present invention and its broader aspects. Accordingly, the appended claims will cover all such changes and modifications that fall within the true spirit and scope of the embodiments. In addition, it should be understood that the embodiments are defined only by the appended claims. Those skilled in the art will understand that if the specific number of elements introduced into the claims is intentional, such intention will be explicitly recited in the claims, and in the absence of such recitation, there is no such limitation. For non-limiting examples, for the purpose of aiding understanding, the appended claims include the use of the introductory phrases "at least one" and "one or more" to introduce claim elements. However, the use of such phrases should not be construed as implying that the introduction of a claim element by the indefinite article "a" or "an" limits any particular claim containing such introduced claim element to an embodiment containing only one such element, even when the same claim includes the introductory phrases "one or more" or "at least one" and the indefinite article such as "a" or "an"; this also applies to the use of the definite article in the claims.

[0082] Embodiments of the present invention may be a system, a method, and / or a computer program product. Additionally, selected aspects of the present embodiment may take the form of a full hardware embodiment, a full software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and / or hardware aspects, which may all be collectively referred to herein as "circuitry", "module", or "system". Further, aspects of embodiments of the present invention may take the form of a computer program product implemented in a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of embodiments of the present invention. The disclosed system, method, and / or computer program product so implemented is operable to improve the functionality and operation of an artificial intelligence platform to build a federated learning framework.

[0083] Aspects of the present embodiment have been described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0084] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions that implement aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0085] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0086] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two boxes shown in succession may in fact be executed substantially concurrently, or the boxes may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each box of the block diagrams and / or flowchart illustrations, and combinations of boxes in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.

[0087] It should be understood that, although specific embodiments have been described herein for purposes of illustration, various modifications may be made without departing from the spirit and scope of the embodiments. Accordingly, the scope of the embodiments is defined only by the appended claims and their equivalents.

Claims

1. A computer system, comprising: A processor operatively coupled to a memory; An artificial intelligence (AI) platform in communication with the processor to support reinforcement learning, the AI platform having tools, the tools including: A question manager for inputting a ground truth first paragraph and a ground truth first answer to a trained neural network and obtaining a first synthetic question generated by a question generator in the trained neural network, the generated first synthetic question having a semantic correspondence with the ground truth first paragraph and the first answer; An answer manager for inputting the first synthetic question to the trained neural network and obtaining a first answer generated by the trained neural network from the ground truth paragraph, the first answer being semantically related to the generated first synthetic question and the ground truth paragraph; An evaluation manager for automatically evaluating the generated first synthetic question against the generated first answer, the evaluation incorporating a reward function that captures the diversity and accuracy of the generated first synthetic question; and A director operatively coupled to the evaluation manager to selectively modify the question generator in response to the evaluation using reinforcement learning, wherein the modified question generator produces one or more synthetic questions having increased text diversity compared to before the modification.

2. The system according to claim 1, wherein The selective modification of the question generator employs nucleus sampling to generate the one or more synthetic questions different from the first synthetic question.

3. The system according to claim 2, wherein, Optimization of variants of the generated first synthetic question includes the director representing the generated first synthetic question with a diversity score and modifying one or more parameters of the trained neural network based on the diversity score, and wherein the optimization mitigates the generation of a second synthetic question having a context similar to the first synthetic question as an output from the trained neural network.

4. The system according to claim 3, further comprising the evaluation manager for evaluating the accuracy of the generated first synthetic question by comparing the generated first answer with the ground truth answer and for quantifying the accuracy evaluation with an accuracy score, and the question manager for selectively modifying one or more question generation parameters in response to the accuracy score.

5. The system according to claim 4, further comprising the evaluation manager calculating the reward function as a convex combination of the accuracy score and the diversity score, the question manager selectively modifying one or more question generation parameters in response to the calculated reward function, the modification mitigating the generation of a second synthetic question semantically similar to the generated first synthetic question.

6. The system according to claim 5, further comprising the evaluation manager adjusting the convex combination to selectively prioritize one or more components of the reward function.

7. A computer program product that supports reinforcement learning, the computer program product including a computer-readable storage medium having program code embodied therewith, the program code executable by a processor to: Input a ground truth first paragraph and a ground truth first answer into a trained neural network, and obtain a first synthetic question generated by a question generator in the trained neural network, where the generated first synthetic question has a semantic correspondence with the ground truth first paragraph and the first answer; Input the first synthetic question into the trained neural network, and obtain a first answer generated by the trained neural network from the ground truth paragraph, where the first answer is semantically related to the generated first synthetic question and the ground truth paragraph; Automatically evaluate the generated first synthetic question for the generated first answer, where the evaluation incorporates a reward function that captures the diversity and accuracy of the generated first synthetic question; And Employ reinforcement learning to selectively modify the question generator in response to the evaluation, where the modified question generator produces one or more synthetic questions having increased text diversity compared to before the modification.

8. The computer program product according to claim 7, wherein, The selective modification of the question generator includes the program code using nucleus sampling to generate the one or more synthetic questions different from the first synthetic question.

9. The computer program product according to claim 8, wherein, The optimization of variants of the generated first synthetic question includes the program code representing the generated first synthetic question using a diversity score and modifying one or more parameters of the trained neural network based on the diversity score, and where the optimization mitigates the generation of a second synthetic question having a context similar to the first synthetic question as an output from the trained neural network.

10. The computer program product according to claim 9, further comprising the program for evaluating the accuracy of the generated first synthetic question by comparing the generated first answer with the ground truth answer and for quantifying the accuracy evaluation using an accuracy score, and the program code for selectively modifying one or more synthetic question generation parameters in response to the accuracy score.

11. The computer program product according to claim 10, further comprising the program code for calculating the reward function as a convex combination of the accuracy score and the diversity score, and selectively modifying one or more synthetic question generation parameters in response to the calculated reward function, the modification being for mitigating the generation of a second synthetic question semantically similar to the generated first synthetic question.

12. The computer program product according to claim 11, further comprising the program code for adjusting the convex combination to selectively prioritize one or more components of the reward function.

13. A computer-implemented method, comprising: Input a ground truth first paragraph and a ground truth first answer into a trained neural network, and obtain a first synthetic question generated by a question generator in the trained neural network, where the generated first synthetic question has a semantic correspondence with the ground truth first paragraph and the first answer; Input the first synthesis problem into the trained neural network, and obtain a first answer generated by the trained neural network from the ground truth passage, where the first answer is semantically related to the generated first synthesis problem and the ground truth passage; Automatically evaluate the generated first synthesis problem for the generated first answer, where the evaluation combines a reward function that obtains the diversity and accuracy of the generated first synthesis problem; And Adopt reinforcement learning and selectively modify the question generator in response to the evaluation, where the modified question generator generates one or more synthesis problems with increased text diversity compared to before the modification.

14. The computer-implemented method according to claim 13, wherein, The selective modification of the question generator includes using nucleus sampling to generate the one or more synthesis problems different from the first synthesis problem.

15. The computer-implemented method according to claim 14, wherein, The optimization of the variation of the generated first synthesis problem includes using a diversity score to represent the generated first synthesis problem and modifying one or more parameters of the trained neural network based on the diversity score, and where the optimization reduces the generation of a second synthesis problem having a context similar to the first synthesis problem as an output from the trained neural network.

16. The computer-implemented method according to claim 15, further comprising evaluating the accuracy of the generated first synthesis problem by comparing the generated first answer with the ground truth answer and using an accuracy score to quantify the accuracy evaluation, and selectively modifying one or more synthesis problem generation parameters in response to the accuracy score.

17. The computer-implemented method according to claim 16, further comprising calculating the reward function as a convex combination of the accuracy score and the diversity score, and selectively modifying one or more synthesis problem generation parameters in response to the calculated reward function, the modification being for reducing the generation of a second synthesis problem semantically similar to the generated first synthesis problem.

18. The computer-implemented method according to claim 17, further comprising adjusting the convex combination to selectively prioritize one or more components of the reward function.

Citation Information

Patent Citations

  • Problem generation method based on progressive multi-discriminator

    CN109271483A

  • Dialogue type problem generation method based on enhanced dynamic reasoning

    CN109992657A