Training data identification and model selection
By querying the language learning model and generating source scores, the problem of difficult to identify the language learning model training data in the prior art is solved, and the identification and resolution of potential problems are realized, improving user experience and security.
Patent Information
- Application Number
- CN202411617046.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-16
AI Technical Summary
Existing language learning models (LLMs) are difficult to effectively identify potential problems associated with training data, such as bias or potential infringement of intellectual property rights.
By querying the first language learning model using the first query, a packet of text output response is generated, and a source score of the packet is generated based on the first training data and the second training data, the first training data is identified as training data of the first language learning model.
Identification of the training data of the language learning model is implemented, enabling the identification of potential problems associated with the training data, improving the user experience and improving the security of using the language learning model.
Smart Images

Figure CN120011801A_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates to language learning models (LLMs), and more particularly, to identifying training data for training LLMs and selecting LLMs for use based on the identified training data.
[0002] Traditional LLMs are designed to understand and reproduce human language by analyzing training data to learn language-based patterns, semantics, and contextual clues. LLMs can employ neural network architectures and deep learning techniques to process training data and establish correlations or causal relationships between input data and human language outputs. However, LLMs typically do not reveal the datasets on which they were trained, complicating efforts to identify potential issues related to the training data. Summary of the invention
[0003] According to one embodiment of the present disclosure, a method is provided. The method includes querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams (n-gram continuous terms); generating a source score for the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score.
[0004] According to one embodiment of the present disclosure, a system is provided. The system includes a processor; and a memory or storage device, including an algorithm or computer instructions, which, when executed by the processor, performs operations including querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams; generating a source score for the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score.
[0005] According to one embodiment of the present disclosure, a computer-readable storage medium is provided, which has a computer-readable program code implemented therewith, and the computer-readable program code can be executed by one or more computer processors to perform operations. A first language learning model is queried with a first query, wherein the first language learning model generates a text output response to the first query; a grouping of the text output responses is generated, wherein the grouping includes a plurality of n-grams; a source score is generated for the grouping based on first training data and second training data; and the first training data is identified as training data for the first language learning model based on the source score.
[0006] According to one embodiment of the present disclosure, a system is provided. The system includes a processor; and a memory or storage device including an algorithm or computer instruction, the algorithm or computer instruction performs operations when executed by the processor, the operations including identifying first training data as training data of a first language learning model based on a source score; generating a ranking of training data based on a query and ratings of corresponding responses of a language learning model set, wherein the ranking of training data includes a ranking of the first training data; selecting a second language learning model of the language learning model set based on the ranking of training data and the query; and generating a response to the query based on the second language learning model.
[0007] According to one embodiment of the present disclosure, a computer-readable storage medium is provided, which has a computer-readable program code implemented therewith, and the computer-readable program code can be executed by one or more computer processors to perform operations. Identify first training data as training data for a first language learning model based on a source score; generate a ranking of the training data based on a query and a rating of a corresponding response of a language learning model set, the ranking of the training data including the ranking of the first training data; select a second language learning model of the language learning model set based on the ranking of the training data and the query; and generate a response to the query based on the second language learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A computing environment according to one embodiment is shown.
[0009] Figure 2 A training data identification and model selection environment is shown according to one embodiment.
[0010] Figure 3 A flow chart illustrating a method of determining that given training data is used to train a language learning model according to one embodiment.
[0011] Figure 4 A flow chart illustrating a method of selecting and using a language learning model according to one embodiment. DETAILED DESCRIPTION
[0012] According to one embodiment of the present disclosure, a method is provided. The method includes querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-gram continuous terms; generating a source score for the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score. Advantageously, this enables identification of the training data for the language learning model, thereby enabling identification of potential issues associated with the training data (e.g., bias or potential infringement of intellectual property rights associated with the training data).
[0013] According to another embodiment of the present disclosure, the method further includes generating a ranking of training data based on queries and ratings of corresponding responses of the language learning model set, wherein the ranking of the training data includes a ranking of first training data; selecting a second language learning model in the language learning model set based on the ranking of the training data and the second query; and generating a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model that satisfies user preferences to respond to queries, which improves the user's experience in participating in the response machine learning mode.
[0014] According to another embodiment of the present disclosure, the source score indicates the possibility that the first training data is used to train the first language learning model, and generating the source score involves generating a first score for the group based on the first training data, the first training data indicating a potential source of the training data used by the first language learning model; and generating a second score for the group based on the second training data, the second training data indicating standard communication data for a given group. In addition, according to another embodiment, the source score is determined as Score S =max(0,(Score1-Score2)),Score S represents the source score, Score1 represents the first score of the grouping, and Score2 represents the second score of the grouping. Advantageously, this enables a comparison between the number of occurrences of the grouping in the first training data and the expected number of occurrences of the grouping in the baseline, second training data, thereby allowing a determination to be made whether the grouping is more present in the first training data.
[0015] According to another embodiment of the present disclosure, the first score is determined as Score1 indicates the first score, n-grams overlaps represents the amount of overlap between each of the n-grams and the first training data, n-grams unique Represents the number of unique n-grams in the text output response, as well as the first training data Bytes (First Training Data Bytes) represents the byte size of the first training data. Advantageously, this enables the number of packet occurrences in the first training data to be normalized to a certain scale, thereby enabling appropriate comparisons to be made between the number of packet occurrences in the first training data and the number of packet occurrences in the baseline, second training data.
[0016] According to another embodiment of the present disclosure, the second score is determined as Score2 represents the second score, n-grams overlaps represents the amount of overlap between each of the n-grams and the second training data, n-gramsunique Represents the number of unique n-grams in the text output response and the second training data Bytes (Second Training Data Bytes) represents the size in bytes of the second training data. Advantageously, this enables the number of packet occurrences in the baseline second training data to be normalized to the determined scale, thereby enabling an appropriate comparison to be made between the number of packet occurrences in the second training data and the number of packet occurrences in the first training data.
[0017] According to one embodiment of the present disclosure, a system is provided. The system includes a processor; and a memory or storage device, including an algorithm or computer instruction, which, when executed by the processor, performs operations including querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-gram continuous terms; generating a source score for the grouping based on first training data and second training data; and identifying the first training data as training data for the first language learning model based on the source score. Advantageously, this enables identification of the training data for the language learning model, thereby enabling identification of potential issues associated with the training data (e.g., bias or potential infringement of intellectual property rights associated with the training data).
[0018] According to another embodiment of the present disclosure, the operation further includes generating a ranking of training data based on the query and rating of the corresponding response of the language learning model set, wherein the ranking of the training data includes the ranking of the first training data; selecting a second language learning model in the language learning model set based on the ranking of the training data and the second query; and generating a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model that satisfies the user's preferences to respond to the query, which improves the user's experience in participating in the response machine learning mode.
[0019] According to another embodiment of the present disclosure, the source score indicates the possibility that the first training data is used to train the first language learning model, and generating the source score involves generating a first score for the group based on the first training data, the first training data indicating a potential source of the training data used by the first language learning model; and generating a second score for the group based on the second training data, the second training data indicating standard communication data for a given group. In addition, according to another embodiment, the source score is determined as Score S =max(0,(Score1-Score2)),Score Srepresents the source score, Score1 represents the first score of the grouping, and Score2 represents the second score of the grouping. Advantageously, this enables a comparison between the number of occurrences of the grouping in the first training data and the expected number of occurrences of the grouping in the baseline, second training data, thereby allowing a determination to be made whether the grouping is more present in the first training data.
[0020] According to another embodiment of the present disclosure, the first score is determined as Score1 indicates the first score, n-grams overlaps represents the amount of overlap between each of the n-grams and the first training data, n-grams unique represents the number of unique n-grams in the text output response, and first training data bytes represents the byte size of the first training data. Advantageously, this enables the number of group occurrences in the first training data to be normalized to a certain scale, thereby enabling appropriate comparisons to be made between the number of group occurrences in the first training data and the number of group occurrences in the baseline, second training data.
[0021] According to another embodiment of the present disclosure, the second score is determined as Score2 represents the second score, n-grams overlaps represents the amount of overlap between each of the n-grams and the second training data, n-grams unique Represents the number of unique n-grams in the text output response, as well as the second training data Bytes (Second Training Data Bytes) represents the size in bytes of the second training data. Advantageously, this enables the number of packet occurrences in the baseline second training data to be normalized to the determined scale, thereby enabling an appropriate comparison to be made between the number of packet occurrences in the second training data and the number of packet occurrences in the first training data.
[0022] According to one embodiment of the present disclosure, a computer-readable storage medium is provided, which has a computer-readable program code implemented therewith, and the computer-readable program code can be executed by one or more computer processors to perform operations. A first language learning model is queried with a first query, wherein the first language learning model generates a text output response to the first query; a grouping of the text output responses is generated, wherein the grouping includes a plurality of n-gram continuous terms; a source score for the grouping is generated based on first training data and second training data; and the first training data is identified as training data for the first language learning model based on the source score. Advantageously, this enables identification of the training data for the language learning model, thereby enabling identification of potential issues associated with the training data (e.g., bias or potential infringement of intellectual property rights associated with the training data).
[0023] According to another embodiment of the present disclosure, the operation further includes generating a ranking of training data based on the query and rating of the corresponding response of the language learning model set, wherein the ranking of the training data includes the ranking of the first training data; selecting a second language learning model in the language learning model set based on the ranking of the training data and the second query; and generating a response to the second query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model that satisfies the user's preferences to respond to the query, which improves the user's experience in participating in the response machine learning mode.
[0024] According to another embodiment of the present disclosure, the source score indicates the possibility that the first training data is used to train the first language learning model, and generating the source score involves generating a first score for the group based on the first training data, the first training data indicating a potential source of the training data used by the first language learning model; and generating a second score for the group based on the second training data, the second training data indicating standard communication data for a given group. In addition, according to another embodiment, the source score is determined as Score S =max(0,(Score1-Score2)),Score S represents the source score, Score1 represents the first score of the grouping, and Score2 represents the second score of the grouping. Advantageously, this enables a comparison between the number of occurrences of the grouping in the first training data and the expected number of occurrences of the grouping in the baseline, second training data, thereby allowing a determination to be made whether the grouping is more present in the first training data.
[0025] According to another embodiment of the present disclosure, the first score is determined as Score1 indicates the first score, n-grams overlaps represents the amount of overlap between each of the n-grams and the first training data, n-grams unique represents the number of unique n-grams in the text output response, and first training data bytes represents the byte size of the first training data. Advantageously, this enables the number of group occurrences in the first training data to be normalized to a certain scale, thereby enabling appropriate comparisons to be made between the number of group occurrences in the first training data and the number of group occurrences in the baseline, second training data.
[0026] According to another embodiment of the present disclosure, the second score is determined as Score2 represents the second score, n-grams overlaps represents the amount of overlap between each of the n-grams and the second training data, n-grams uniqueRepresents the number of unique n-grams in the text output response, as well as the second training data Bytes (Second Training Data Bytes) represents the size in bytes of the second training data. Advantageously, this enables the number of packet occurrences in the baseline second training data to be normalized to the determined scale, thereby enabling an appropriate comparison to be made between the number of packet occurrences in the second training data and the number of packet occurrences in the first training data.
[0027] According to one embodiment of the present disclosure, a system is provided. The system includes a processor; and a memory or storage device including an algorithm or computer instruction, which performs operations when executed by the processor, including identifying the first training data as training data for the first language learning model based on the source score; generating a ranking of the training data based on the ratings of the corresponding responses of the query and the language learning model set, wherein the ranking of the training data includes the ranking of the first training data; selecting a second language learning model of the language learning model set based on the ranking of the training data and the query; and generating a response to the query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model that satisfies the user's preferences to respond to the query, which improves the user's experience in participating in the response machine learning mode.
[0028] According to another embodiment of the present disclosure, the operation also includes querying a first language learning model using a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-gram contiguous terms; and generating a source score for the grouping based on the first training data and the second training data. Advantageously, this enables identification of the training data for the language learning model, thereby enabling identification of potential issues associated with the training data (e.g., bias or potential infringement of intellectual property rights associated with the training data).
[0029] According to another embodiment of the present disclosure, the source score indicates the possibility that the first training data is used to train the first language learning model, and generating the source score involves generating a first score for the group based on the first training data, the first training data indicating a potential source of the training data used by the first language learning model; and generating a second score for the group based on the second training data, the second training data indicating standard communication data for a given group. In addition, according to another embodiment, the source score is determined as Score S =max(0,(Score1-Score2)),Score S represents the source score, Score1 represents the first score of the group, and Score2 represents the second score of the group; the first score is determined as Score1 indicates the first score, n-grams overlapsrepresents the amount of overlap between each n-grams and the first training data, n-grams unique Represents the number of unique n-grams in the text output response, as well as the first training data Bytes (first training data byte) represents the byte size of the first training data; and the second score is determined as Score2 represents the second score n-grams unique represents the amount of overlap between each n-grams and the second training data, n-grams unique Represents the number of unique n-grams in the text output response, and the second training data Bytes (Second Training Data Bytes) represents the size in bytes of the second training data. Advantageously, this enables a comparison between the number of occurrences of the packet in the first training data and the expected number of occurrences of the packet in the baseline, second training data, thereby allowing a determination whether the packet is more present in the first training data. Furthermore, this enables the number of occurrences of the packet in the first training data to be normalized to a determined scale, thereby enabling an appropriate comparison to be made between the number of occurrences of the packet in the first training data and the number of occurrences of the packet in the baseline, second training data.
[0030] According to one embodiment of the present disclosure, a computer-readable storage medium is provided, which has a computer-readable program code implemented therewith, and the computer-readable program code can be executed by one or more computer processors to perform operations. Identify first training data as training data for a first language learning model based on a source score; generate a ranking of training data based on the ratings of the query and the corresponding responses of the language learning model set, the ranking of training data including the ranking of first training data; select a second language learning model of the language learning model set based on the ranking of training data and the query; and generate a response to the query based on the second language learning model. Advantageously, this enables the selection and use of a machine learning model that satisfies the user's preferences to respond to the query, which improves the user's experience in participating in the response machine learning mode.
[0031] According to another embodiment of the present disclosure, the operation also includes querying a first language learning model using a first query, the first language learning model generating a text output response to the first query; generating a grouping of the text output response, the grouping including a plurality of n-gram contiguous terms; and generating a source score for the grouping based on the first training data and the second training data. Advantageously, this enables identification of the training data for the language learning model, thereby enabling identification of potential issues associated with the training data (e.g., bias or potential infringement of intellectual property rights associated with the training data).
[0032] According to another embodiment of the present disclosure, the source score indicates the possibility that the first training data is used to train the first language learning model, and generating the source score involves generating a first score for the group based on the first training data, the first training data indicating a potential source of the training data used by the first language learning model; and generating a second score for the group based on the second training data, the second training data indicating standard communication data for a given group. Advantageously, this enables a comparison between the number of occurrences of the group in the first training data and the expected number of occurrences of the group in the baseline, second training data, thereby allowing a determination of whether the group appears more in the first training data.
[0033] Embodiments of the present disclosure improve training data identification techniques by providing a training data identification and model selection (TDIMS) module that determines the likelihood that given training data is used to train a language learning model. In one embodiment, the TDIMS module queries the LLM using a query designed to elicit a long-form response. The TDIMS module can then generate scores for the n-ary continuations of the response, and use the scores to determine which training data is used to train the LLM. The TDIMS module can also rank the determined training data based on topic, user rating, and user demographics. The TDMIS module can then use the ranking to select the LLM that generates the best response to the query.
[0034] One benefit of the disclosed embodiments is the identification of potential issues associated with training data (e.g., bias or potential infringement of intellectual property rights associated with the training data). In addition, embodiments of the present disclosure can improve the user experience of using language learning models by selecting and using models trained on data that suits the user's preferences.
[0035] Various aspects of the present disclosure are described by narrative text, flow charts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flow chart, depending on the technology involved, the operations may be performed in an order different from the order shown in a given flow chart. For example, again depending on the technology involved, two operations shown in consecutive flow chart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0036] Computer program product embodiments ("CPP embodiments" or "CPP") are terms used in this disclosure to describe any collection of one or more storage media (also referred to as "media") collectively included in a collection of one or more storage devices that collectively include machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. Without limitation, a computer-readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), static random access memories (SRAM), compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / land formed in a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, should not be construed as storing in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, light pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As will be appreciated by those skilled in the art, data is typically moved at certain occasional points in time during normal operation of the storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device transitory because the data is not transitory while it is stored.
[0037] Figure 1A computing environment 100 according to one embodiment is shown. The computing environment 100 includes an example of an environment for executing at least some of the computer codes involved in performing the methods of the present invention, such as the new training data identification and model selection (TDIMS) module 150 shown in block 190. In addition to block 190, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuits 120 and caches 121), a communication structure 111, a volatile memory 112, a permanent storage device 113 (including an operating system 122 and block 190, as described above), a peripheral device group 114 (including a user interface (UI) device group 123, a storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140 , a cloud coordination module 141 , a host physical machine group 142 , a virtual machine group 143 , and a container group 144 .
[0038] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or developed in the future that is capable of running programs, accessing a network, or querying a database such as remote database 130. As is well known in the art of computer technology, and depending on the technology, the performance of computer-implemented methods may be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, particularly computer 101, to keep the presentation as simple as possible. Computer 101 may be located in the cloud, even though it is not yet fully understood. Figure 1 1 is not shown in the cloud, on the other hand, computer 101 need not be in the cloud unless it can be positively indicated to any extent.
[0039] Processor set 110 includes one or more computer processors of any type now known or developed in the future. Processing circuit 120 can be distributed over multiple packages, such as multiple coordinated integrated circuit chips. Processing circuit 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located in the processor chip package and is generally used for data or code that should be quickly accessed by threads or cores running on processor set 110. Cache memory is generally organized into multiple levels according to relative proximity to the processing circuit. Alternatively, some or all of the caches of the processor group may be located "off chip". In some computing environments, processor set 110 may be designed to work with qubits and perform quantum computing.
[0040] Computer readable program instructions are typically loaded onto computer 101 to cause processor set 110 of computer 101 to perform a series of operating steps to implement a computer-implemented method, such that the instructions so executed will instantiate the method specified in the flow chart and / or the narrative description of the computer-implemented method included in this document (collectively referred to as "the present method"). These computer readable program instructions are stored in various types of computer readable storage media, such as cache 121 and other storage media discussed below. The program instructions and related data are accessed by processor set 110 to control and direct the execution of the present method. In computing environment 100, at least some of the instructions for executing the present method may be stored in permanent storage device 113 in box 190.
[0041] The communication fabric 111 is the signal conduction paths that allow the various components of the computer 101 to communicate with each other. Typically, the fabric is made up of switches and conductive paths, such as those that make up a bus, a bridge, physical input / output ports, etc. Other types of signal communication paths may be used, such as fiber optic communication paths and / or wireless communication paths.
[0042] The volatile memory 112 is any type of volatile memory now known or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, the volatile memory 112 is characterized by random access, but this is not required unless expressly stated. In the computer 101, the volatile memory 112 is located in a single package and is internal to the computer 101, but, alternatively or additionally, the volatile memory can be distributed in multiple packages and / or located externally relative to the computer 101.
[0043] Permanent storage 113 is any form of non-volatile storage for computers known now or developed in the future. The non-volatility of the memory means that the stored data is maintained regardless of whether power is supplied to the computer 101 and / or directly to the permanent storage 113. Permanent storage 113 can be a read-only memory (ROM), but usually at least a portion of the permanent storage allows the writing of data, the deletion of data, and the rewriting of data. Some common forms of persistent storage include disks and solid-state storage devices. Operating system 122 can take several forms, such as various known proprietary operating systems or operating systems of the open source portable operating system interface type using a kernel. The code included in block 190 is generally included in at least some of the computer codes involved in executing the method of the present invention.
[0044] The peripheral device group 114 includes a set of peripheral devices of the computer 101. The data communication connection between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as a Bluetooth connection, a near field communication (NFC) connection, a connection made by a cable (such as a universal serial bus (USB) type cable), a plug-in type connection (e.g., a secure digital (SD) card), a connection made through a local area communication network, and even a connection made through a wide area network such as the Internet. In various embodiments, the user interface (UI) device group 123 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smart watches), keyboards, mice, printers, touch pads, game controllers, and tactile devices. The storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. The storage device 124 can be permanent and / or volatile. In some embodiments, the storage device 124 can take the form of a quantum computing storage device for storing data in the form of quantum bits. In embodiments where the computer 101 needs to have a large amount of storage (e.g., where the computer 101 locally stores and manages a large database), the storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The Internet of Things (IoT) sensor set 125 is comprised of sensors that may be used in IoT applications. For example, one sensor may be a thermometer, while another sensor may be a motion detector.
[0045] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software defined networks (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separated devices so that the control function manages several different network hardware devices. Computer-readable program instructions for executing the method of the present invention can typically be downloaded to the computer 101 from an external computer or an external storage device via a network adapter card or a network interface included in the network module 115.
[0046] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any technology now known or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include computer hardware, such as copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0047] End-user device (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of an enterprise operating computer 101), and may take any of the forms discussed above in connection with computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to an end-user, the recommendations would typically be transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this manner, EUD 103 may display or otherwise present the recommendations to the end-user. In some embodiments, EUD 103 may be a client device, such as a thin client, a heavy client, a mainframe computer, a desktop computer, etc.
[0048] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0049] The public cloud 105 is any computer system that can be used by multiple entities, which provides on-demand availability of computer system resources and / or other computer capabilities, particularly data storage (cloud storage) and computing capabilities, without direct active management by users. Cloud computing typically utilizes the sharing of resources to achieve consistency and economy of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud coordination module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that constitute the host physical machine group 142, which is the universe of physical computers in the public cloud 105 and / or available for the public cloud. The virtual computing environment (VCE) is typically in the form of a virtual machine from the virtual machine group 143 and / or a container from the container group 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after instantiation of the VCE. The cloud coordination module 141 manages the transmission and storage of images, deploys new instantiations of VCEs, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allows public cloud 105 to communicate over WAN 102 .
[0050] Some further explanation of a virtualized computing environment (VCE) will now be provided. A VCE can be stored as an "image". A new active instance of the VCE can be instantiated from the image. Two common types of VCEs are virtual machines and containers. Containers are VCEs that use operating system-level virtualization. This refers to an operating system feature in which the kernel allows the existence of multiple isolated user space instances, called containers. From the perspective of the programs running in them, these isolated user space instances typically behave like actual computers. Computer programs running on a normal operating system can utilize all of the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices assigned to the container, a feature known as containerization.
[0051] The private cloud 106 is similar to the public cloud 105, except that the computing resources are only available to a single enterprise. Although the private cloud 106 is depicted as communicating with the WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and can only be accessed through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable coordination, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, the public cloud 105 and the private cloud 106 are both part of a larger hybrid cloud.
[0052] Figure 2 A training data identification and model selection (TDIMS) environment 200 is shown according to one embodiment. In the illustrated embodiment, the TDIMS environment 200 includes a computer 101, a language learning model 202, training data 206, and a language learning model set 208. Features of the TDIMS environment 200 are communicatively coupled via the WAN 102.
[0053] As shown, the computer 101 may include a TDIMS module 150, which further includes a query module 152, a grouping module 154, a scoring module 156, and a selection module 158. In one embodiment, the TDIMS module 150 and the modules included therein represent one or more algorithms, sets of instructions, software applications, or other computer-readable program codes that may be executed by the processor set 110 of the computer 101 to perform the functions, operations, or processes described herein.
[0054] In one embodiment, query module 152 interrogates language learning model (LLM) 202 using a predetermined query designed to elicit a long-form response 204 from LLM 202. Grouping module 154 may formulate response 204 of n-grams. Scoring module 156 may then generate a score for the n-grams that indicates the likelihood of finding the n-grams within various training data. TDIMS module 150 may then use the score to determine whether given training data (e.g., training data 206) should be used to train LLM 202. These processes are described below. Figure 3 Further discussion in.
[0055] After identifying the training data 206 for training the LLM 202, the selection module 158 can rank the training data 206 based on the user ratings of the response 204 and the first query. This process can be repeated for different LLMs and corresponding queries and responses to produce a ranked set of training data. The selection module 158 can then select (from the set of LLMs 208) the LLM trained on the training data with the highest ranking to respond to user queries similar to the query used to generate the response 204. These processes are described below. Figure 4 Further discussion in.
[0056] Figure 3 A flow chart of a method 300 for determining that given training data is used to train a language learning model is shown according to one embodiment. In one embodiment, the method 300 is performed by the training data identification and model selection (TDIMS) module 150 .
[0057] As previously discussed, the TDIMS module 150 may include a query module 152 , a grouping module 154 , a scoring module 156 , and a selection module 158 . The method 300 begins at block 302 .
[0058] At block 304, the query module 152 queries the language learning model 202 with a first query, wherein the language learning model 202 generates a text output response to the first query. In one embodiment, the query module 152 queries the language learning model (LLM) 202 with one of a plurality of predetermined queries designed to elicit a long-form response 204 from the LLM 202. As opposed to a single-word or binary response, the long-form response 204 may include an interpretation for the query.
[0059] For example, the query module may input to the LLM 202 the statement "What is an electron?", and the response 204 of the LLM 202 may be, "An electron is a subatomic particle that carries a negative electric charge.".
[0060] At block 306, the grouping module 154 generates groups of text output responses, wherein the groups include multiple n-grams. In one embodiment, an n-gram is a continuous sequence of text elements (a total of "n" text elements), such as words, characters, punctuation, spaces, etc. For example, a 3-gram may include 3 text elements ("An electron is"), and a 4-gram may include 4 text elements ("An electron is a").
[0061] In one embodiment, the grouping module 154 generates overlapping, unique 5-grams of the responses 204 of the LLM 202. One benefit of using 5-grams is to capture identifiers of the source of an expression, as empirical evidence suggests that 5 consecutive words of a text are sufficient to express a quality that can be used as an identifier.
[0062] The grouping module 154 may generate the response 204 of n-grams of contiguous terms that together cover all words of the response 204. Continuing with the previous example, the grouping module 154 may generate the following eight 5-grams: 1. "electrons are subatomic"; 2. "electrons are subatomic particles"; 3. "are subatomic particles, i.e."; 4. "carry subatomic particles"; 5. "carry a subatomic particle"; 6. "negatively charged particles"; 7. "i.e. carry negative electricity"; 8. "i.e. carry a negative charge".
[0063] At block 308 , the scoring module 156 generates a source score for the group based on the first training data and the second training data. In one embodiment, the source score indicates the likelihood that the first training data was used to train the language learning model 202 .
[0064] The first training data may represent a potential source of training data used by the language learning model. The first training data may be, for example, a corpus such as an online database, a patent database, a code library, a book, a web-based text reference, or the like.
[0065] The second training data may represent standard or common communication data for a given group. For example, the second training data may be a corpus, such as a collection of text data crawled from websites in a given country, covering a wide range of topics and domains. The second training data may provide a representative sample of language usage in the country as reflected on the web, so that an assessment of the occurrence of the groupings in the first training data may be compared to a baseline, expected number of occurrences of the groupings in the second training data.
[0066] In one embodiment, the source score is determined as follows: S =max(0,(Score1-Score2)): where Score S represents the source score, Score1 represents the first score of the group relative to the first training data, and Score2 represents the second score of the group relative to the second training data.
[0067] The first score can be generated as follows:
[0068] Among them, Score1 represents the first score of the group, n-grams overlaps represents the number of overlaps between each of the n-grams and the first training data, n-gramsunique represents the number of unique n-grams in the response 204, and the first training data bytes represents the size of the first training data in bytes.
[0069] Continuing with the previous example, assuming that the response 204 contains a further explanation of the academic community on electronics, the grouping module 154 may generate 20,000 unique 5-grams of the response 204. Assuming that all text elements of a given 5-gram appear 5000 times in the first training data, there will be 5000 n-grams. overlaps Therefore, log 10 (5000+1) / 20000 will represent the first element of the sum. Perform a similar summation for each unique overlap of 5-grams. Assuming the first training data is 100GB, then divide the sum by log 10 (1e11 bytes).
[0070] The second score can be generated as follows:
[0071] Among them, Score2 represents the second score of the group, n-grams overlaps represents the number of overlaps between each of the n-grams and the second training data, n-grams unique represents the number of unique n-grams in the response 204, and the second training data bytes represents the size of the second training data in bytes.
[0072] At block 310, the TDIMS module 150 identifies the first training data as training data for the language learning model 202 based on the source score. In one embodiment, a larger source score indicates a greater likelihood that the first training data is used to train the language learning model 202. Thus, the training data with the largest source score is identified as the primary training data for the language learning model 202. The method 300 ends at box 312.
[0073] Figure 4 A flow chart of a method 400 of selecting and using a language learning model according to one embodiment is shown. In one embodiment, the method 400 is performed by the training data identification and model selection (TDIMS) module 150.
[0074] As previously discussed, the TDIMS module 150 may include a query module 152, a grouping module 154, a scoring module 156, and a selection module 158. The method 400 begins at block 402.
[0075] At block 404, the selection module 158 generates a ranking of the training data based on the queries and the ratings of the corresponding responses of the language learning model set 208. In one embodiment, the selection module 158 categorizes the queries according to the subject matter of the queries. The selection module 158 may retrieve the ratings of the responses from a database including user ratings collected from response surveys. In one embodiment, the ranking of the training data includes the categorization of the subject matter of the first query and the ranking of the first training data.
[0076] User ratings may also include user demographics that may provide context for the ratings. For example, academic users may give higher ratings to responses that reflect data and jargon from academic journals.
[0077] At block 406, the selection module 158 selects a language learning model from the language learning model set 208 based on the ranking of the training data and the second query. In one embodiment, the selection module 158 may use a natural language processing engine to extract features from the second query and use the features to determine a topic of the second query.
[0078] The selection module 158 may also use these features to determine the demographics of the user generating the second query. For example, the selection module 158 may consider the selection of words in the second query to determine the user's age, or consider the subject matter of the second query to determine the user's occupation, etc.
[0079] In one embodiment, the selection module 158 filters the LLMs in the set of LLMs 208 based on a match between the subject matter of the second query and the subject classifications in the ranking of the training data. The selection module 158 may also filter the LLMs based on a match between the user demographics that generated the second query and the user demographics that ranked the training data. The selection module 158 may then select the highest ranked training data to use (which may be obtained by comparing the above Figure 3 A process similar to the process discussed in is used to determine the LLM in the trained filtered LLM.
[0080] At block 408, the selection module 158 generates a response to the second query based on the selected language learning model. In this way, the TDIMS module 150 can provide an optimized response to the second query, as determined by the preferences of users with similar user demographics. The method 400 ends at box 412.
[0081] While the foregoing is directed to embodiments of the present invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof, and the scope of the invention is determined by the claims that follow.
Claims
1. A method comprising: querying a first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams of consecutive terms; generating a source score for the grouping based on the first training data and the second training data; as well as The first training data is identified as training data for the first language learning model based on the source score.
2. The method according to claim 1, further comprising: generating a ranking of training data based on the queries and the ratings of the corresponding responses of the set of language learning models, wherein the ranking of training data includes the ranking of the first training data; selecting a second language learning model from the set of language learning models based on the ranking of training data and the second query; and A response to the second query is generated based on the second language learning model.
3. The method according to claim 1, wherein: The source score indicates the likelihood that the first training data is used to train the first language learning model, and wherein generating the source score comprises: generating a first score for the grouping based on the first training data, wherein the first training data represents a potential source of the training data used by the first language learning model; and A second score for the grouping is generated based on the second training data, wherein the second training data represents standard communication data for a given population.
4. The method according to claim 3, wherein the source score is determined as Score S =max(0,(Score1-Score2)), where Score S represents the source score, wherein Score1 represents the first score of the grouping, and wherein Score2 represents the second score of the grouping.
5. The method according to claim 4, wherein: The first score is determined as Among them, Score1 represents the first score, where n-grams overlaps represents the amount of overlap between each of the n-grams and the first training data, where n-grams unique represents the number of unique n-gram contiguous items in the text output response, and wherein the first training data Bytes The size of the first training data is expressed in bytes.
6. The method according to claim 4, wherein: The second score is determined as Where Score2 represents the second score, where n-grams overlaps represents the number of overlaps between each of the n-grams and the second training data, where n-grams unique represents the number of unique n-gram contiguous items in the text output response, and wherein the second training data Bytes The size of the second training data is expressed in bytes.
7. A system comprising: processor; as well as A memory or storage device comprising an algorithm or computer instructions which, when executed by the processor, perform a method according to any one of claims 1-6.
8. A computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code being executable by one or more computer processors to perform the method of any one of claims 1-6.
9. A system comprising modules respectively used to perform the steps of the method according to any one of claims 1 to 6.
10. A system comprising: processor; as well as a memory or storage device including an algorithm or computer instructions that, when executed by the processor, performs operations comprising: identifying the first training data as training data for a first language learning model based on the source score; generating a ranking of training data based on a plurality of queries and ratings of corresponding responses of a set of language learning models, wherein the ranking of training data comprises a ranking of first training data; selecting a second language learning model from the set of language learning models based on the ranking and querying of the training data; and A response to the query is generated based on the second language learning model.
11. The system of claim 10, wherein the operations further comprise: querying the first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams of consecutive terms; as well as The source score for the packet is generated based on first training data and second training data.
12. The system of claim 11, wherein the source score represents a likelihood that the first training data is used to train the first language learning model, and wherein generating the source score comprises: generating a first score for the grouping based on the first training data, wherein the first training data represents a potential source of the training data used by the first language learning model; as well as A second score for the grouping is generated based on the second training data, wherein the second training data represents standard communication data for a given population.
13. The system of claim 12, wherein the source score is determined as Score S =max(0,(Score1-Score2)), where Score S represents the source score, wherein Score1 represents the first score of the grouping, and wherein Score2 represents the second score of the grouping; Wherein the first score is determined as in, Score1 represents the first score, where n-gramso verlaps represents the number of overlaps between each of the n-grams and the first training data, where n-grams unique represents the number of unique n-gram contiguous items in the text output response, and wherein the first training data Bytes representing the size of the first training data in bytes; and Wherein the second score is determined as Where Score2 represents the second score, where n-grams overlaps represents the number of overlaps between each of the n-grams and the second training data, where n-grams unique represents the number of unique n-gram contiguous items in the text output response, and wherein the second training data Bytes The size of the second training data is expressed in bytes.
14. A computer readable storage medium having computer readable program code embodied therewith, the computer readable program code being executable by one or more computer processors to perform operations comprising: identifying the first training data as training data for a first language learning model based on the source score; generating a ranking of training data based on the query and the ratings of the corresponding responses of the set of language learning models, wherein the ranking of training data includes a ranking of first training data; selecting a second language learning model from the set of language learning models based on the ranking and querying of the training data; and A response to the query is generated based on the second language learning model.
15. The computer-readable storage medium of claim 14, the operations further comprising: querying the first language learning model with a first query, wherein the first language learning model generates a text output response to the first query; generating a grouping of the text output response, wherein the grouping includes a plurality of n-grams of consecutive terms; as well as The source score for the packet is generated based on first training data and second training data.
16. The computer-readable storage medium of claim 15, wherein: The source score indicates the likelihood that the first training data is used to train the first language learning model, and wherein generating the source score comprises: generating a first score for the grouping based on the first training data, wherein the first training data represents a potential source of the training data used by the first language learning model; and A second score for the grouping is generated based on the second training data, wherein the second training data represents standard communication data for a given population.