Method, system and computer program (training data identification and model selection)
The method and system address the challenge of identifying training data for language learning models by calculating a source score based on n-gram groupings, enabling the detection of bias and intellectual property issues and improving user experience through model selection.
Patent Information
- Application Number
- JP2024196434
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-27
AI Technical Summary
Conventional language learning models (LLMs) do not reveal the datasets on which they are trained, making it difficult to identify potential problems such as bias or intellectual property infringement associated with the training data.
A method and system that query a language learning model with a query, generate a grouping of the response into n-grams, and calculate a source score based on first and second training data to identify the training data used by the model. This allows for the selection of an appropriate LLM based on user preferences and the ranking of training data.
Enables the identification of training data used by language learning models, allowing for the detection of potential issues such as bias and intellectual property infringement, and improves user experience by selecting models trained with data aligned with user preferences.
Smart Images

Figure 2025081253000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a language learning model (LLM), and more specifically, to identifying training data used to train an LLM and selecting an LLM to be used based on the identified training data.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Conventional LLMs are designed to understand and reproduce human language by analyzing training data to learn language-based patterns, semantics, and context cues. LLMs can process training data using neural network architectures and deep learning techniques to establish a correlation or causal relationship between input data and the output of human language. However, LLMs generally do not reveal the datasets on which they are trained, thereby making it difficult to identify potential problems associated with the training data.
Means for Solving the Problems
[0003] According to an embodiment of the present disclosure, a method is provided. The method includes querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams; generating a source score of the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score.
[0004] According to an embodiment of the present disclosure, a system is provided. The system includes a processor; and procedures that, when executed by the processor, query a first language learning model with a first query, where the first language learning model generates a text output response to the first query; generate a grouping of the text output response, where the grouping includes a plurality of n-grams; generate a source score of the grouping based on first training data and second training data; and based on the source score, identify the first training data as the training data of the first language learning model, and includes a memory or storage having an algorithm or computer instructions for performing the operations.
[0005] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium having computer-readable program code embodied therein, where the computer-readable program code is executable by one or more computer processors for performing operations. The operations include querying a first language learning model with a first query, where the first language learning model generates a text output response to the first query; generating a grouping of the text output response, where the grouping includes a plurality of n-grams; generating a source score of the grouping based on first training data and second training data; and based on the source score, identifying the first training data as the training data of the first language learning model.
[0006] According to an embodiment of the present disclosure, a system is provided. The system includes a processor; and procedures that, when executed by the processor, identify first training data as training data for a first language learning model based on a source score; generate a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes a ranking of the first training data; select a second language learning model from the set of language learning models based on the ranking of the training data and the query; and generate a response to the query based on the second language learning model, and a memory or storage having an algorithm or computer instructions for performing operations including these procedures.
[0007] According to an embodiment of the present disclosure, there is provided a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors for performing operations. The operations include procedures that identify first training data as training data for a first language learning model based on a source score; generate a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes a ranking of the first training data; select a second language learning model from the set of language learning models based on the ranking of the training data and the query; and generate a response to the query based on the second language learning model.
Brief Description of the Drawings
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
DETAILED DESCRIPTION OF THE INVENTION
[0012] According to one embodiment of the present disclosure, a method is provided. The method includes querying a first language learning model with a first query, where the first language learning model generates a text output response to the first query; generating a grouping of the text output responses, where the grouping includes a plurality of n-grams; generating a source score for the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score. Advantageously, this enables the identification of the training data for the language learning model, thereby enabling the identification of potential problems associated with the training data (e.g., bias associated with the training data or potential infringement of intellectual property rights).
[0013] According to another embodiment of the present disclosure, the method further includes generating a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of training data includes a first ranking of training data; selecting a second language learning model from the set of language learning models based on the ranking of training data and a second query; and generating a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model according to the user's preferences to respond to the query, thereby improving the user experience when participating in the response of the machine learning model.
[0014] According to another embodiment of the present disclosure, the source score represents the likelihood that the first training data was used to train the first language learning model, and the step of generating the source score includes generating a first score for grouping based on the first training data, where the first training data represents a potential source of training data used by the first language learning model; and generating a second score for grouping based on the second training data, where the second training data represents standard communication data of a given population. Further, according to another embodiment, the source score is determined as score s = max(0, (score 1 - score 2 ))), where score S represents the source score, score 1 represents the first score for grouping, and score 2 represents the second score for grouping. Advantageously, this enables a comparison of the number of grouping occurrences in the first training data and the expected number of grouping occurrences in the baseline second training data, thereby enabling a determination of whether more groupings occur in the first training data.
[0015] According to another embodiment of the present disclosure, the first score is [Number] is determined as, and the score 1 represents the first score, and the n-gram 重複 represents the number of occurrences of each n-gram and the duplicates in the first training data. The n-gram 一意 represents the number of unique n-grams in the text output response, and the first training data バイト represents the size of the first training data in bytes. Advantageously, this makes it possible to normalize the number of grouped occurrences in the first training data with respect to the judgment scale, thereby enabling an appropriate comparison of the number of grouped occurrences in the first training data and the number of grouped occurrences in the baseline second training data.
[0016] According to another embodiment of the present disclosure, the second score is [Number] is determined as, and the score 2 represents the second score, and the n-gram 重複 represents the number of occurrences of each n-gram and the duplicates in the second training data. The n-gram 一意 represents the number of unique n-grams in the text output response, and the second training data バイト represents the size of the second training data in bytes. Advantageously, this makes it possible to normalize the number of grouped occurrences in the baseline second training data with respect to the judgment scale, thereby enabling an appropriate comparison of the number of grouped occurrences in the second training data and the number of grouped occurrences in the first training data.
[0017] According to an embodiment of the present disclosure, a system is provided. The system includes a processor; and, when executed by the processor, procedures to query a first language learning model with a first query, where the first language learning model generates a text output response to the first query; procedures to generate a grouping of the text output response, where the grouping includes a plurality of n-grams; procedures to generate a source score for the grouping based on first training data and second training data; and procedures to identify the first training data as training data for the first language learning model based on the source score. Advantageously, this enables the identification of training data for the language learning model, thereby enabling the identification of potential problems associated with the training data (e.g., bias associated with the training data or potential infringement of intellectual property rights).
[0018] According to another embodiment of the present disclosure, the operations further include procedures to generate a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes the ranking of the first training data; procedures to select a second language learning model from the set of language learning models based on the ranking of the training data and a second query; and procedures to generate a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model according to a user's preference to respond to a query, thereby improving the user experience when engaging with the response of the machine learning model.
[0019] According to another embodiment of the present disclosure, the source score represents the likelihood that the first training data was used to train the first language learning model, and the step of generating the source score includes generating a first score for grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for grouping based on the second training data, where the second training data represents standard communication data of a given population. Further, according to another embodiment, the source score is determined as s = max(0, (score 1 - score 2 ))), where score S represents the source score, score 1 represents the first score for grouping, and score 2 represents the second score for grouping. Advantageously, this enables comparison of the number of grouping occurrences in the first training data and the expected number of grouping occurrences in the baseline second training data, thereby enabling determination of whether more groupings occur in the first training data.
[0020] According to another embodiment of the present disclosure, the first score is
Number
[0021] According to another embodiment of the present disclosure, the second score is [Number] judged as, and the score 2 represents the second score, n-gram 重複 represents the number of duplicates of each n-gram and the second training data, and n-gram 一意 represents the number of unique n-grams in the text output response, and the second training data バイト represents the size of the second training data in bytes. Advantageously, this enables the number of grouping occurrences in the baseline second training data to be normalized against a judgment scale, thereby enabling an appropriate comparison of the number of grouping occurrences in the second training data and the number of grouping occurrences in the first training data.
[0022] There is provided a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations. The operation includes a procedure of querying a first language learning model with a first query, where the first language learning model generates a text output response to the first query; a procedure of generating a grouping of the text output response, where the grouping includes a plurality of n-grams; a procedure of generating a source score of the grouping based on first training data and second training data; and a procedure of identifying the first training data as training data of the first language learning model based on the source score. Advantageously, this enables the identification of the training data of the language learning model, thereby enabling the identification of potential problems related to the training data (e.g., bias related to the training data or potential infringement of intellectual property rights).
[0023] According to another embodiment of the present disclosure, the operation further includes a procedure of generating a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes the ranking of the first training data; a procedure of selecting a second language learning model from the set of language learning models based on the ranking of the training data and a second query; and a procedure of generating a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model according to the user's preference to respond to the query, thereby improving the user experience when participating in the response of the machine learning model.
[0024] According to another embodiment of the present disclosure, the source score represents the likelihood that the first training data was used to train the first language learning model, and the procedure for generating the source score includes a procedure for generating a first score for grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and a procedure for generating a second score for grouping based on the second training data, where the second training data represents standard communication data of a given population. Further, according to another embodiment, the source score is determined as score s = max(0, (score 1 - score 2 ))), where score S represents the source score, score 1 represents the first score for grouping, and score 2 represents the second score for grouping. Advantageously, this enables a comparison of the number of grouping occurrences in the first training data and the expected number of grouping occurrences in the baseline second training data, thereby enabling a determination of whether more groupings occur in the first training data.
[0025] According to another embodiment of the present disclosure, the first score is
Number
[0026] According to another embodiment of the present disclosure, the second score is [Number] judged as, and the score 2 represents the second score, n-gram 重複 represents the number of each n-gram and the number of duplicates in the second training data, and n-gram 一意 represents the number of unique n-grams in the text output response, and the second training data バイト represents the size of the second training data in bytes. Advantageously, this enables the number of grouping occurrences in the baseline second training data to be normalized against a judgment scale, thereby allowing an appropriate comparison of the number of grouping occurrences in the second training data and the number of grouping occurrences in the first training data.
[0027] According to one embodiment of the present disclosure, a system is provided. The system includes a processor; and an algorithm or computer instructions having operations that, when executed by the processor, identify first training data as training data for a first language learning model based on a source score; generate a ranking of the training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes a ranking of the first training data; select a second language learning model from the set of language learning models based on the ranking of the training data and the query; and generate a response to the query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model according to a user's preference to respond to a query, thereby improving the user experience when participating in the response of the machine learning model.
[0028] According to another embodiment of the present disclosure, the operations further include querying the first language learning model with a first query, where the first language learning model generates a text output response to the first query; generating a grouping of the text output response, where the grouping includes a plurality of n-grams; and generating a source score of the grouping based on the first training data and the second training data. Advantageously, this enables the identification of training data for the language learning model, thereby enabling the identification of potential problems associated with the training data (e.g., bias associated with the training data or potential infringement of intellectual property rights).
[0029] According to another embodiment of the present disclosure, the source score represents the likelihood that the first training data was used to train the first language learning model, and the procedure for generating the source score includes a procedure for generating a first score for grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and a procedure for generating a second score for grouping based on the second training data, where the second training data represents standard communication data of a given population. Further, according to another embodiment, the source score is determined as score s = max(0, (score 1 - score 2 ), where score S represents the source score, score 1 represents the first score of the grouping, score 2 represents the second score of the grouping; the first score is
Number
Number
[0030] A computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations, is provided in accordance with an embodiment of the present disclosure. The operations include a procedure for identifying first training data as training data for a first language learning model based on a source score; a procedure for generating a ranking of training data based on a query and ratings of corresponding responses of a set of language learning models, where the ranking of the training data includes a ranking of the first training data; a procedure for selecting a second language learning model from the set of language learning models based on the ranking of the training data and the query; and a procedure for generating a response to the query based on the second language learning model. Advantageously, this enables selection and use of a machine learning model according to a user's preference to respond to a query, thereby improving the user experience when participating in the response of the machine learning model.
[0031] According to another embodiment of the present disclosure, the operation includes a procedure of querying the first language learning model with a first query, the first language learning model generating a text output response to the first query; a procedure of generating a grouping of the text output response, the grouping including a plurality of n-grams; and a procedure of generating a source score of the grouping based on first training data and second training data. Advantageously, this enables the identification of the training data of the language learning model, thereby enabling the identification of potential problems associated with the training data (e.g., bias associated with the training data or potential infringement of intellectual property rights).
[0032] According to another embodiment of the present disclosure, the source score represents the likelihood that first training data was used to train the first language learning model, and the procedure of generating the source score includes a procedure of generating a first score of the grouping based on the first training data, the first training data representing a potential source of the training data used by the first language learning model; and a procedure of generating a second score of the grouping based on the second training data, the second training data representing standard communication data of a given population. Advantageously, this enables a comparison of the number of occurrences of the grouping in the first training data and the expected number of occurrences of the grouping in the baseline second training data, thereby enabling a determination of whether the grouping occurs more frequently in the first training data.
[0033] Embodiments of the present disclosure improve the identification technique of training data by providing a training data identification and model selection (TDIMS) module that determines the likelihood that given training data was used to train a language learning model. In one embodiment, the TDIMS module queries an LLM using a query designed to elicit a long-form response. The TDIMS module then generates n-gram scores for the response and uses these scores to determine which training data was used to train the LLM. The TDIMS module can also rank the determined training data based on the topic, user rating, and user demographics. The TDIMS module can then use this ranking to select an LLM that generates an optimal response to the query.
[0034] One advantage of the disclosed embodiments is identifying potential problems associated with the training data (e.g., bias associated with the training data or potential infringement of intellectual property rights). Further, the embodiments of the present disclosure can improve the user experience with language learning models by selecting and using models trained with data according to the user's preferences.
[0035] Various aspects of the present disclosure are illustrated by descriptions, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be executed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be executed in reverse order, as a single integrated step, simultaneously, or at least partially overlapping in time.
[0036] An embodiment of a computer program product (referred to as "CPP embodiment" or "CPP") is a term used in the present disclosure to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices, which collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. Without limitation, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices including these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the main surface of a disk), or any suitable combination of the foregoing. A computer-readable storage medium shall not be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals communicated via a wire, and / or other transmission media, when the term is used in the present disclosure. As will be understood by those skilled in the art, data is normally moved at some irregular points during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but since the data is not transient while being stored, the above does not make the storage device transient.
[0037] FIG. 1 shows a computing environment 100 according to one embodiment. The computing environment 100 includes an example of an environment for executing at least some of the computer code necessary to perform the method of the present invention, such as a novel training data identification and model selection (TDIMS) module 150 shown in block 190. In addition to block 190, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including a processing circuit 120 and a cache 121), a communication fabric 111, a volatile memory 112, a persistent storage 113 (including an operating system 122 and block 190 as shown above), a set of peripheral devices 114 (including a user interface (UI) device set 123, a storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0038] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or hereafter developed that is capable of executing programs, accessing a network, or querying a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the execution of computer-implemented methods may be distributed among multiple computers and / or between multiple locations. On the other hand, in this description of computing environment 100, for the sake of simplicity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although not shown within the cloud in FIG. 1, computer 101 may be located within the cloud. On the other hand, computer 101 need not exist within the cloud, except within any range that may be affirmatively shown.
[0039] Processor set 110 includes one or more computer processors of any type now known or hereafter developed. Processing circuitry 120 may be distributed among multiple packages, e.g., multiple conditioned integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within a processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, processor set 110 may be designed to operate using qubits and perform quantum computing.
[0040] Computer-readable program instructions cause a set of operation steps to be executed by a processor set 110 of computer 101, thereby realizing a computer-implemented method that is normally loaded onto computer 101, and as a result, the instructions thus executed instantiate the method specified in the flowchart and / or description of the computer-implemented method (collectively referred to as "the method of the present invention") included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 121 and other storage media discussed below. The program instructions and related data are accessed by processor set 110 to control and direct the execution of the method of the present invention. In computing environment 100, at least some of the instructions for executing the method of the present invention may be stored in block 190 within persistent storage 113.
[0041] Communication fabric 111 is a signal conduction path that enables various components of computer 101 to communicate with each other. Typically, this fabric is made up of switches and conductive paths such as buses, bridges, physical input / output ports, and the like. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.
[0042] Volatile memory 112 is any type of volatile memory known currently or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly stated. In computer 101, volatile memory 112 is located within a single package and exists inside computer 101, but alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0043] The persistent storage 113 is any form of non-volatile storage for a computer, known currently or developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is supplied directly to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a portion of the persistent storage enables writing of data, deletion of data, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems, or an open-source portable operating system interface type of operating system using a kernel. The code included in block 190 typically includes at least some of the computer code involved in the execution of the method of the present invention.
[0044] The peripheral device set 114 includes a set of peripheral devices of the computer 101. The data communication connections between the peripheral devices of the computer 101 and other components may be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near-Field Communication (NFC) connection, a connection via a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (for example, a Secure Digital (SD) card), a connection via a local area communication network, and even a connection via a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touch pad, a game controller, and a haptic device. The storage 124 is an external storage such as an external hard drive or an insertable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, when the computer 101 locally stores and manages a large-scale database), this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 125 is composed of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer, and another sensor may be a motion detector.
[0045] The network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers via WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for executing the method of the present invention can typically be downloaded to computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.
[0046] WAN 102 is any wide area network (e.g., the Internet) that can communicate computer data over a non-local distance by any technology for communicating computer data that is currently known or developed in the future. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0047] An end-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating computer 101), and can take any of the forms discussed above in relation to computer 101. The EUD 103 typically receives useful and beneficial data from the operation of computer 101. For example, in a hypothetical case where computer 101 is designed to provide recommendations to an end user, this recommendation would typically be communicated from the network module 115 of computer 101, via the WAN 102, to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, a thick client, a mainframe computer, and a desktop computer, etc.
[0048] A remote server 104 is any computer system that provides at least some data and / or functions to computer 101. The remote server 104 can be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores data that is useful and beneficial for use by other computers such as computer 101. For example, in a hypothetical case where computer 101 is designed and programmed to provide recommendations based on past data, this past data can be provided from the remote database 130 of the remote server 104 to computer 101.
[0049] The public cloud 105 is any computer system that provides on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing capabilities, for use by multiple entities without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available in the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from a virtual machine set 143 and / or containers from a container set 144. It is understood that these VCEs can be stored as images and transferred either as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of the VCE, and manages the active instantiation of VCE deployments. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate via the WAN 102.
[0050] Here, some further explanations of virtual computing environments (VCEs) are provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to the container, and this feature is known as containerization.
[0051] The private cloud 106 is similar to the public cloud 105, except that computing resources are only available for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability among the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0052] FIG. 2 shows a training data identification and model selection (TDIMS) environment 200 according to one embodiment. In the illustrated embodiment, the TDIMS environment 200 includes a computer 101, a language learning model 202, training data 206, and a set of language learning models 208. The features of the TDIMS environment 200 are communicatively coupled via the WAN 102.
[0053] As shown, computer 101 can include a TDIMS module 150, which further includes a query module 152, a grouping module 154, a scoring module 156, and a selection module 158. In one embodiment, the TDIMS module 150, and the modules included therein, represent one or more algorithms, instruction sets, software applications, or other computer-readable program code that can be executed by the processor set 110 of computer 101 to perform the functions, operations, or processes described herein.
[0054] In one embodiment, query module 152 queries language learning model (LLM) 202 using a predetermined query designed to elicit a long-form response 204 from LLM 202. Grouping module 154 can create n-grams of response 204. Thereafter, scoring module 156 can generate a score for the n-gram indicating the likelihood that the n-gram is found within various training data. The TDIMS module 150 can then use this score to determine whether a given training data (e.g., training data 206) was used to train LLM 202. These processes are further considered in FIG. 3 below.
[0055] After the training data 206 used to train LLM 202 is identified, selection module 158 can rank the training data 206 based on the user rating of response 204 and the first query. This process can be repeated for different LLMs, and corresponding queries and responses, to generate a set of ranked training data. The selection module 158 can then select (from the set of LLMs 208) an LLM trained with the highest-ranked training data to respond to user queries similar to the query used to generate response 204. These processes are further considered in FIG. 4 below.
[0056] Figure 3 shows a flowchart of a method 300 for determining that given training data has been used to train a language learning model, according to one embodiment. In one embodiment, method 300 is performed by a Training Data Identification and Model Selection (TDIMS) module 150.
[0057] As previously discussed, TDIMS module 150 can include a query module 152, a grouping module 154, a scoring module 156, and a selection module 158. Method 300 begins at block 302.
[0058] In block 304, query module 152 queries language learning model 202 with a first query, where language learning model 202 generates a text output response to the first query. In one embodiment, query module 152 queries language learning model (LLM) 202 using one of a plurality of predetermined queries designed to elicit a long-form response 204 from LLM202. The long-form response 204 can include an explanation directed to the query, as opposed to a single-word or binary response.
[0059] For example, the query module can input a prompt such as "What is an electron?" into LLM202. The response 204 of LLM202 can be "An electron is a subatomic particle that carries a negative electric charge".
[0060] In block 306, the grouping module 154 generates a grouping of the text output responses, where the grouping includes a plurality of n-grams. In one embodiment, an n-gram is a consecutive sequence (equal to "n" text elements) of text elements such as, for example, words, characters, punctuation, spaces, etc. For example, a 3-gram can include three text elements ("An electron is"), while a 4-gram can include four text elements ("An electron is a").
[0061] In one embodiment, the grouping module 154 generates overlapping unique 5-grams of the response 204 of the LLM 202. One advantage of using 5-grams is to capture the identifier of the source of the expression, as empirical evidence suggests that five consecutive words in the text are sufficient to represent the specificity that can be used as an identifier.
[0062] The grouping module 154 can generate a number of n-grams of the response 204 that collectively cover all the words of the response 204. Continuing with the previous example, the grouping module 154 may generate the following eight 5-grams: 1. "An electron is a subatomic", 2. "An electron is a subatomic particle", 3. "Is a subatomic particle that", 4. "A subatomic particle that carries", 5. "A subatomic particle that carries a", 6. "Particle that carries a negative", 7. "That carries a negative electric", and 8. "Carries a negative electric charge".
[0063] In block 308, the scoring module 156 generates a source score for grouping based on the first training data and the second training data. In one embodiment, the source score represents the likelihood that the first training data was used to train the language learning model 202.
[0064] The first training data can represent a potential source of the training data used by the language learning model. The first training data can be, for example, a corpus such as an online database, a patent database, a code repository, a book, a web-based text reference, etc.
[0065] The second training data can represent standard or general communication data of a given population. For example, the second training data can be a corpus such as a collection of text data that crawls websites in a given country and covers various topics and domains. The second training data provides a representative sample of language use in that country as reflected on the web, thereby allowing the evaluation of the groupings that appear in the first training data to be compared with the baseline expected number of occurrences of the groupings in the second training data.
[0066] In one embodiment, the source score is determined as follows. Score s = max(0, (score 1 - score 2 ), where score S represents the source score, score 1 represents the first score of the grouping for the first training data, and score 2 represents the second score of the grouping for the second training data.
[0067] The first score may be generated as follows.
[0068]
Number
[0069] Continuing with the previous example, assume that response 204 includes further explanations of electrons for academic personnel, and the grouping module 154 may generate 20,000 unique 5-grams of response 204. Assume that all text elements of a given 5-gram appear 5000 times in the first training data, then there will be 5000 n-grams 重複 existing. Therefore, log 10 (5000 + 1) / 20000 will represent the total first element. Similar totals are done for each n-gram of the unique 5-grams 重複 . Assume that the first training data is 100GB, then the total sum is then divided by log 10 (1e11 bytes).
[0070] The second score may be generated as follows.
[0071]
Number
[0072] In block 310, the TDIMS module 150 identifies first training data as the training data for the language learning model 202 based on the source score. In one embodiment, the larger the source score, the more likely the first training data was used to train the language learning model 202. Thus, the training data with the highest source score is identified as the primary training data for the language learning model 202. Method 300 ends at block 312.
[0073] FIG. 4 shows a flowchart of a method 400 for selecting and using a language learning model according to one embodiment. In one embodiment, method 400 is performed by a training data identification and model selection (TDIMS) module 150.
[0074] As previously discussed, the TDIMS module 150 can include a query module 152, a grouping module 154, a scoring module 156, and a selection module 158. Method 400 begins at block 402.
[0075] In block 404, the selection module 158 generates a ranking of the training data based on the query and the ratings of the corresponding responses in the set 208 of language learning models. In one embodiment, the selection module 158 classifies the query according to the topic of the query. The selection module 158 can retrieve the ratings of the responses from a database that includes user ratings collected from a survey of the responses. In one embodiment, the ranking of the training data includes the classification of the topic of the first query and the ranking of the first training data.
[0076] The user ratings may also include the demographics of the users that are capable of providing the context of the ratings. For example, academic users may give higher ratings to responses that reflect data and jargon from academic journals.
[0077] In block 406, selection module 158 selects a language learning model from the set 208 of language learning models based on the ranking of the training data and the second query. In one embodiment, selection module 158 can use a natural language processing engine to extract features from the second query and use these features to determine the topic of the second query.
[0078] Selection module 158 can also use these features to determine the user demographics of the user who generated the second query. For example, selection module 158 can consider the selection of words in the second query to determine the user's age, or consider the topic of the second query to determine the user's occupation, etc.
[0079] In one embodiment, selection module 158 filters the LLM in the set of LLM208 based on the match between the topic of the second query and the classification of the topic in the ranking of the training data. Selection module 158 can further filter the LLM based on the match between the user demographics of the user who generated the second query and the user demographics of the ranking of the training data. Then, selection module 158 can select the LLM trained using the highest-ranked training data (which can be determined via a process similar to the process discussed in FIG. 3 above) among the filtered LLM.
[0080] In block 408, selection module 158 generates a response to the second query based on the selected language learning model. Thus, TDIMS module 150 can provide an optimized response to the second query as judged by the user's preferences with similar user demographics. Method 400 ends at block 412.
[0081] Although the above is directed to embodiments of the present invention, other and further embodiments of the present invention may be devised without departing from the basic scope of the present invention, and the scope of the present invention is determined by the following claims.
Claims
1. querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of said text output responses, said grouping including a plurality of n-grams; generating a source score for the grouping based on the first training data and the second training data; and identifying the first training data as training data for the first language learning model based on the source scores; A method for providing the above.
2. generating a ranking of training data based on the queries and ratings of the corresponding responses of the set of language learning models, where the ranking of training data includes a ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of the training data and a second query; and generating a response to the second query based on the second language learning model. The method of claim 1 further comprising:
3. The source score represents a likelihood that the first training data was used to train the first language learning model, and generating the source score comprises: generating a first score for the grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for the grouping based on the second training data, the second training data representing typical communication data for a given population; The method of claim 1 , comprising:
4. The source score is score s = max(0,(score 1 -Score 2 )) and the score is S represents the source score, and the score 1 represents the first score of the grouping, and score 2 The method of claim 3 , wherein: denotes the second score for the grouping.
5. The first score is ##EQU00011## The score is as follows: 1 represents the first score, and 重複 represents the number of overlaps of each of the n-grams and the first training data, 一意 represents the number of unique n-grams in the text output response, and the first training data バイト The method of claim 4 , wherein: represents a size of the first training data in bytes.
6. The second score is ##EQU00012## The score is as follows: 2 represents the second score, and 重複 represents the number of overlaps of each of the n-grams and the second training data, and 一意 represents the number of unique n-grams in the text output response, and the second training data バイト The method of claim 4 or 5, wherein represents the size of the second training data in bytes.
7. A processor; and When executed by the processor, querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of said text output response, said grouping including a plurality of n-grams; generating a source score for the grouping based on the first training data and the second training data; and identifying the first training data as training data for the first language learning model based on the source scores; A memory or storage having algorithms or computer instructions for performing operations including A system comprising:
8. The operation includes: generating a ranking of training data based on the queries and ratings of the corresponding responses of the set of language learning models, where the ranking of training data includes a ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of the training data and a second query; and generating a response to the second query based on the second language learning model. The system of claim 7 further comprising:
9. The source score represents a likelihood that the first training data was used to train the first language learning model, and generating the source score comprises: generating a first score for the grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for the grouping based on the second training data, the second training data representing typical communication data for a given population; The system of claim 7 , comprising:
10. The source score is score s = max(0,(score 1 -Score 2 )) and the score is S represents the source score, and the score 1 represents the first score of the grouping, and score 2 The system of claim 9 , wherein: denotes the second score for the grouping.
11. The first score is ##EQU00013## The score is as follows: 1 represents the first score, and 重複 represents the number of overlaps of each of the n-grams and the first training data, 一意 represents the number of unique n-grams in the text output response, and the first training data バイト The system of claim 10 , wherein represents a size of the first training data in bytes.
12. The second score is ##EQU14## The score is as follows: 2 represents the second score, and 重複 represents the number of overlaps of each of the n-grams and the second training data, and 一意 represents the number of unique n-grams in the text output response, and the second training data バイト The system of claim 10 or 11, wherein represents a size in bytes of the second training data.
13. A computer program having computer readable program code embodied therein, the computer readable program code comprising: querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of said text output response, said grouping including a plurality of n-grams; generating a source score for the grouping based on the first training data and the second training data; and identifying the first training data as training data for the first language learning model based on the source scores; A computer program executable by one or more computer processors to perform operations including:
14. The operation includes: generating a ranking of training data based on the queries and ratings of the corresponding responses of the set of language learning models, where the ranking of training data includes a ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of the training data and a second query; and generating a response to the second query based on the second language learning model. The computer program product of claim 13 , further comprising:
15. The source score represents a likelihood that the first training data was used to train the first language learning model, and generating the source score comprises: generating a first score for the grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for the grouping based on the second training data, the second training data representing typical communication data for a given population; 14. The computer program of claim 13, comprising:
16. The source score is score s = max(0,(score 1 -Score 2 )) and the score is S represents the source score, and the score 1 represents the first score of the grouping, and score 2 The computer program product of claim 15 , wherein:
17. The first score is ##EQU00015## The score is as follows: 1 represents the first score, and 重複 represents the number of overlaps of each of the n-grams and the first training data, 一意 represents the number of unique n-grams in the text output response, and the first training data バイト The computer program product of claim 16 , wherein: represents a size in bytes of the first training data.
18. The second score is ##EQU00016## The score is as follows: 2 represents the second score, and 重複 represents the number of overlaps of each of the n-grams and the second training data, and 一意 represents the number of unique n-grams in the text output response, and the second training data バイト The computer program product of claim 16 or 17, wherein: represents a size in bytes of the second training data.
19. A processor; and When executed by the processor, identifying the first training data as training data for a first language learning model based on the source scores; generating a ranking of training data based on the queries and ratings of corresponding responses of the set of language learning models, where the ranking of training data includes a ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of the training data and the query; and generating a response to the query based on the second language learning model. A memory or storage having algorithms or computer instructions for performing operations including A system comprising:
20. The operation includes: querying the first language learning model with a first query, where the first language learning model generates a text output response to the first query; generating a grouping of the text output response, the grouping including a plurality of n-grams; and generating the source scores for the groupings based on first training data and second training data; 20. The system of claim 19, further comprising:
21. The source score represents a likelihood that the first training data was used to train the first language learning model, and generating the source score comprises: generating a first score for the grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for the grouping based on the second training data, the second training data representing typical communication data for a given population; The system of claim 20, comprising:
22. The source score is score s = max(0,(score 1 -Score 2 )) and the score is S represents the source score, and the score 1 represents the first score of the grouping, and score 2 represents the second score for the grouping; The first score is ##EQU00017## The score is as follows: 1 represents the first score, and 重複 represents the number of overlaps of each of the n-grams and the first training data, 一意 represents the number of unique n-grams in the text output response, and the first training data バイト represents the size of the first training data in bytes; and The second score is [0018] The score is as follows: 2 represents the second score, and 重複 represents the number of overlaps of each of the n-grams and the second training data, and 一意 represents the number of unique n-grams in the text output response, and the second training data バイト 22. The system of claim 21, wherein represents a size of the second training data in bytes.
23. A computer program having computer readable program code embodied therein, the computer readable program code comprising: identifying the first training data as training data for a first language learning model based on the source scores; generating a ranking of training data based on the queries and ratings of corresponding responses of the set of language learning models, where the ranking of training data includes a ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of the training data and the query; and generating a response to the query based on the second language learning model. A computer program executable by one or more computer processors to perform operations including:
24. The operation includes: querying the first language learning model with a first query, where the first language learning model generates a text output response to the first query; generating a grouping of the text output response, the grouping including a plurality of n-grams; and generating the source scores for the groupings based on first training data and second training data; 24. The computer program of claim 23, further comprising:
25. The source score represents a likelihood that the first training data was used to train the first language learning model, and generating the source score comprises: generating a first score for the grouping based on the first training data, where the first training data represents a potential source of the training data used by the first language learning model; and generating a second score for the grouping based on the second training data, the second training data representing typical communication data for a given population; 25. The computer program of claim 24, comprising: