Data generation method, training method and device based on deep learning model
The deep learning model enhances intelligent systems by determining the need for external components and generating intermediate queries, improving response quality and adaptability.
Patent Information
- Application Number
- JP2023170081
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-10
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2043-09-29
AI Technical Summary
Current intelligent systems have weak processing capabilities for user input data, resulting in poor quality of generated response content.
A deep learning model is used to determine whether to call external functional components, generate intermediate queries, and adjust parameters based on intermediate results to enhance response generation.
Improves the quality of generated answers by adapting to user intentions and utilizing external components for more accurate and timely responses.
Smart Images

Figure 0007710012000007 
Figure 0007710012000008 
Figure 0007710012000009
Abstract
Description
Detailed Description of the Invention
[0001]
Technical Field
[0002] The present disclosure relates to the technical field of artificial intelligence, and in particular, to the technical fields of natural language processing and deep learning. Specifically, it relates to a data generation method based on a deep learning model, a training method for a deep learning model, a data generation device based on a deep learning model, a training device for a deep learning model, an electronic device, and a computer-readable storage medium.
Background Art
[0003] Artificial intelligence is a subject that studies how to simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) on a computer. There are both hardware technologies and software technologies. The hardware technologies of artificial intelligence generally include technologies such as sensors, artificial intelligence dedicated chips, cloud computing, distributed storage, and big data processing. The artificial intelligence software technologies mainly include several directions such as natural language processing technology, computer vision technology, speech recognition technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0004] The methods described in this section are not necessarily the methods previously assumed or adopted. Unless otherwise specified, none of the methods described in this section should be considered as prior art just because they are included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered as those approved by the prior art.
Summary of the Invention
[0005] The present disclosure provides a data generation method based on a deep learning model, a training method for a deep learning model, a data generation device based on a deep learning model, a training device for a deep learning model, an electronic device, and a computer-readable storage medium.
[0006] According to one aspect of the present disclosure, a data generation method based on a deep learning model is provided. The deep learning model can generate response data based on input data of a user. The data generation method includes determining an initial input used for the deep learning model based on the input data from the user, obtaining a first output of the deep learning model, where, in response to determining that the deep learning model needs to call a first functional component different from the deep learning model to generate a response based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component determined based on the initial input, obtaining a first intermediate result determined by the first functional component based on the first intermediate query, determining a second input used for the deep learning model based on at least the initial input and the first intermediate result, and obtaining a second output of the deep learning model to generate a response to the initial input.
[0007] According to another aspect of the present disclosure, a method for training a deep learning model is provided. The deep learning model is used to generate response data based on user input data. The training method includes obtaining first sample data, where the first sample data includes a first sample initial input and a first sample output. Here, the first sample initial input includes an expression of intention to call a first preset functional component different from the deep learning model, and the first sample output includes a first token for calling the first preset functional component and a first sample intermediate input identifiable by the first preset functional component; obtaining second sample data, where the second sample data includes a second sample initial input and a second sample output. Here, the second sample initial input does not include an expression of intention to call any preset functional component different from the deep learning model, and the second sample output does not include a corresponding token for calling any preset functional component; using the deep learning model to process the first sample initial input to obtain a first predicted output; adjusting the parameters of the deep learning model based on a comparison between the first sample output and the first predicted output; using the deep learning model to process the second sample initial input to obtain a second predicted output; and adjusting the parameters of the deep learning model based on a comparison between the second sample output and the second predicted output.
[0008] According to another aspect of the present disclosure, there is provided a data generation device based on a deep learning model. The deep learning model can generate response data based on input data of a user. The data generation device includes a first determination unit configured to determine an initial input used for the deep learning model based on the input data from the user, a first acquisition unit configured to acquire a first output of the deep learning model, and in response to determining that the deep learning model needs to call a first functional component different from the deep learning model to generate a response based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component determined based on the initial input, a second acquisition unit configured to acquire a first intermediate result determined by the first functional component based on the first intermediate query, a second determination unit configured to determine a second input used for the deep learning model based on at least the initial input and the first intermediate result, and a third acquisition unit configured to acquire a second output of the deep learning model to generate a response to the initial input.
[0009] According to another aspect of the present disclosure, a training device for a deep learning model is provided. The deep learning model is used to generate response data based on user input data. The training device includes a fourth acquisition unit configured to acquire first sample data, where the first sample data includes a first sample initial input and a first sample output. Here, the first sample initial input includes an intention expression for calling a first preset functional component different from the deep learning model, and the first sample output includes a first token for calling the first preset functional component and a first sample intermediate input identifiable by the first preset functional component; a fifth acquisition unit configured to acquire second sample data, where the second sample data includes a second sample initial input and a second sample output. Here, the second sample initial input does not include an intention expression for calling any preset functional component different from the deep learning model, and the second sample output is configured not to include a corresponding token for calling any preset functional component; a first processing unit configured to process the first sample initial input using the deep learning model to obtain a first predicted output; a first parameter adjustment unit configured to adjust the parameters of the deep learning model based on a comparison between the first sample output and the first predicted output; a second processing unit configured to process the second sample initial input using the deep learning model to obtain a second predicted output; and a second parameter adjustment unit configured to adjust the parameters of the deep learning model based on a comparison between the second sample output and the second predicted output.
[0010] According to one or more embodiments of the present disclosure, the present disclosure uses a deep learning model to determine whether it is necessary to call a first functional component different from the deep learning model. When it is determined that it is necessary to call the first functional component, a first intermediate query that can be identified by the first functional component is generated using the deep learning model. Further, in order to obtain a first intermediate result, the first functional component is called using the first intermediate query, and finally, based on the first intermediate result, the deep learning model is used to generate a result for the user's initial input.
[0011] As described above, for a deep learning model that can perform tasks such as understanding and generation by itself, further ability enhancement is realized, thereby improving the quality of the finally generated answer. Further, by directly generating an intermediate query that can be identified by an external functional component using the deep learning model, the acquisition of the intermediate query and the intermediate result is adapted to the potential intention in the user's initial input, and thus enables the model to output an answer that meets the user's needs.
[0012] It should be understood that the content described in this part is not intended to identify the key points or important features of the embodiments of the present disclosure, nor is it intended to limit the protection scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
Brief Description of the Drawings
[0013] The drawings illustrate embodiments by way of example and constitute a part of the specification. Together with the written description of the specification, they are used to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to elements that are similar but not necessarily the same.
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Best Mode for Carrying Out the Invention
[0015] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. The various details in the embodiments of the present disclosure included therein are for helping understanding, and they should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described in this specification without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] In the present application, unless otherwise specified, terms such as "first", "second", etc. for explaining various elements are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are used only to distinguish one element from another. In some examples, the first element and the second element may refer to the same example of the element, and in some cases, they may refer to different examples based on the description of the context.
[0017] The terms used in the description of various examples of the present disclosure are for the sole purpose of describing specific examples and are not intended to be limiting. Unless otherwise explicitly indicated in the context, elements may be singular or plural if the number of elements is not particularly limited. Note that the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0018] In the related art, an intelligent system can generate corresponding response content based on user input data. However, current intelligent systems have weak processing capabilities for user input data and the quality of the generated response content is poor.
[0019] To solve the above problems, the present disclosure uses a deep learning model to determine whether it is necessary to call a first functional component different from the deep learning model. If it is determined that it is necessary to call the first functional component, a first intermediate query that can be identified by the first functional component is generated using the deep learning model. Furthermore, to obtain a first intermediate result, the first functional component is called using the first intermediate query. Finally, based on the first intermediate result, the deep learning model is used to generate a result for the user's initial input.
[0020] As described above, for a deep learning model that can perform tasks such as understanding and generation by itself, further ability enhancement is realized, thereby improving the quality of the finally generated answer. Furthermore, by directly generating an intermediate query that can be identified by an external functional component using the deep learning model, the acquisition of the intermediate query and the intermediate result is adapted to the potential intention in the user's initial input. Therefore, it enables the model to output an answer that meets the user's needs.
[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Figure 1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented, according to an embodiment of the present disclosure. Referring to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0022] In an embodiment of the present disclosure, the server 120 operates to be able to execute one or more services or software applications of the data generation method or the deep learning model training method of the present disclosure. In an exemplary embodiment, the server can deploy a deep learning model that supports an intelligent system.
[0023] In some embodiments, the server 120 can also provide other services or software applications that can include a non-virtual environment and a virtual environment. In some embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 in a software as a service (SaaS) model.
[0024] In the arrangement shown in FIG. 1, server 120 may include one or more assemblies that implement the functions executed by server 120. These assemblies may include software assemblies, hardware assemblies, or combinations thereof that can be executed by one or more processors. A user operating client devices 101, 102, 103, 104, 105, and / or 106 can interact with server 120 using one or more client applications to utilize the services provided by these assemblies. It should be understood that various different system arrangements are possible and may differ from system 100. Thus, FIG. 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0025] A user can input into the intelligent system using client devices 101, 102, 103, 104, 105, and / or 106. The client device can provide an interface for a user of the client device to interact with the client device. The client device can also output information to the user via this interface, for example, output to the user an answer generated by the intelligent system for a user input. Although only six client devices are illustrated in FIG. 1, as will be understood by those skilled in the art, the present disclosure can support any number of client devices.
[0026] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices such as portable handheld devices, general-purpose computers (e.g., personal computers and notebook computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors, or other detection devices. These computer devices can run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (e.g., GOOGLE Chrome OS), and can include various mobile operating systems such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include mobile phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Game systems may include various handheld game devices, Internet-connected game devices, etc. Client devices can run, for example, Internet-related applications, communication applications (e.g., email applications), short message service (SMS) applications, and can execute various applications and use various communication protocols.
[0027] Network 110 can be any type of network known to those skilled in the art, and it can use any one of a plurality of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0028] Server 120 can include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines that execute a virtual operating system, or other computing architectures related to virtualization (e.g., one or more flexible pools of virtualized logical memory devices to maintain virtual memory devices of the server). In various embodiments, server 120 can execute one or more services or software applications that provide the functions described below.
[0029] The computing unit in server 120 can execute one or more operating systems including any of the above-described operating systems and any commercial server operating systems. Server 120 can also execute any one of various additional server applications and / or middleware applications, such as an HTTP server, an FTP server, a CGI server, a JAVA server, a database server, etc.
[0030] In some embodiments, server 120 may include one or more applications for analyzing and integrating data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0031] In some embodiments, server 120 may be a server of a distributed system or a server incorporating a blockchain. Server 120 may be a cloud server, or an intelligent cloud computing server or an intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, which solves the defects of high management difficulty and weak business scalability existing in conventional physical hosts and virtual private server (VPS) services.
[0032] System 100 may also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can be located at various positions. For example, the database used by server 120 may be local to server 120, or may communicate with server 120 remotely via a network or a dedicated connection. Databases 130 can be of various types. In some embodiments, the database used by server 120 may be a relational database. One or more of these databases can store, update, and retrieve data from the database in response to instructions.
[0033] In some embodiments, one or more of the databases 130 can be used by an application and can also store the data of the application. The database used by the application can be various types of databases, such as a key-value repository, an object repository, a general-purpose repository supported by a file system, etc.
[0034] The system 100 of FIG. 1 can be configured and operated in various ways so as to apply the various methods and apparatuses described based on the present disclosure. According to one aspect of the present disclosure, there is provided a data generation method based on a deep learning model. The deep learning model can generate response data based on user input data. As shown in FIG. 2, the data generation method includes: step S201 of determining an initial input used for the deep learning model based on input data from a user; obtaining a first output of the deep learning model, and here, in response to determining that the deep learning model needs to call a first functional component different from the deep learning model to generate a response based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component determined based on the initial input, step S202; step S203 of obtaining a first intermediate result determined by the first functional component based on the first intermediate query; step S204 of determining a second input used for the deep learning model based on at least the initial input and the first intermediate result; and step S205 of obtaining a second output of the deep learning model to generate a response to the initial input.
[0035] Therefore, as described above, for a deep learning model that can perform tasks such as understanding and generation on its own, further ability enhancement is realized, thereby improving the quality of the finally generated answer. Further, by directly generating an intermediate query that can be identified by an external functional component using the deep learning model, the acquisition of the intermediate query and the intermediate result is adapted to the potential intention in the user's initial input, and thus enables the model to output an answer that meets the user's needs.
[0036] In the present disclosure, the deep learning model is also referred to as an understanding generation integrated interactive large-scale model (abbreviated as an understanding generation large-scale model or an integrated large-scale model). The understanding generation large-scale model has an end-to-end characteristic and can directly generate answer data based on the user's input data without passing through functional components other than the understanding generation large-scale model and other inputs. In other words, the understanding generation large-scale model itself has a generation function. Further, the system in which the understanding generation large-scale model is arranged can be called an intelligent system. The intelligent system may also include an interactive module for receiving input data from the user and providing the finally generated answer to the user. In a single conversation between the user and the intelligent system, the intelligent system can use the understanding generation large-scale model arranged therein to have multiple conversations with the user.
[0037] The understanding and generation large-scale model can adopt, for example, an N-layer Transformer network structure having an encoder and a decoder, or a Unified pre-trained Language Model (UniLM) network structure. It should be understood that the understanding and generation large-scale model may also be a neural network model based on other Transformer network structures, and is not limited herein. Both the input and output of the understanding and generation large-scale model are composed of tokens. Each token can correspond to a single word, character, word, special symbol, or a certain external functional component as described below.
[0038] It should be understood that the deep learning model used in the data generation method described in the present disclosure may also be trained by the training method of the deep learning model described later in the present disclosure.
[0039] Before step S201, the user's input data may first be obtained. The user's input data may be, for example, a user input to an intelligent system, and may include, for example, text input, voice input, image input, etc. It should be understood that the user's input data may also have other data forms and is not limited in this specification. The user's input data may be a factual question, an instruction to execute a specific task, or casual conversation content. For different types of user inputs, the intelligent system can generate appropriate answers.
[0040] According to some embodiments, the first functional component may be an external memory bank capable of storing a first set of data groups related to the user. Each data group in the first set of data groups may include at least a historical input data item and a historical response item generated by a deep learning model for the historical input data item. It should be understood that the historical input data item and the corresponding historical response item may include, for example, conversations generated in the historical conversation between the user and the intelligent system, and may also include conversations generated by the user and the intelligent system in the current conversation. Thereby, by installing the external memory bank, the long-term historical conversation between the user and the intelligent system is stored, the memory capacity of the intelligent system is improved, and by obtaining the historical conversation related to the user input, the deep learning model can refer to the historical conversation to generate a more targeted, richer and more specific response to the user, thereby improving the quality of the response, improving the intelligence of the conversation, and improving the user experience.
[0041] According to some embodiments, each data group in the first set of data groups may further include an entry time item (or timestamp) corresponding to the historical input data item and the historical response item in the set. Thereby, by providing the entry time item, when searching or deleting the historical conversation in the external memory bank, more rich operations can be realized according to the entry time of the historical conversation, and the effectiveness of the memory is improved.
[0042] According to some embodiments, each data group in the first data group set can further include a theme item corresponding to the historical input data item and the historical answer item in the set. In one exemplary embodiment, when acquiring memory, a historical dialogue having the same theme as the current dialogue can be directly acquired, or the theme item can be used as one of the bases for similarity calculation so that a more efficient historical dialogue can be acquired more efficiently. Thereby, by providing a theme item, a specific memory can be converted into an abstract memory, and in the search and deletion of historical dialogues in an external memory bank, richer operations can be realized according to the theme of the historical dialogue.
[0043] In one exemplary embodiment, the data group in the external memory bank can be shown in Table 1 below.
[0044]
Table 1
[0045] According to some embodiments, the first intermediate query can be based on the input data. The first intermediate query may match the user's input data, may include the user's input data and context information, or may be a rewrite of the initial input determined based on the input data by a deep learning model. The context information can include a plurality of dialogues conducted between the user and the intelligent system before the acquired user input data.
[0046] According to some embodiments, the first intermediate result may be a historical answer item corresponding to a historical input data item in the first data group set whose similarity to the input data is higher than the first threshold. Therefore, in order to obtain the first intermediate result, by acquiring historical answer items related to the current user input from an external memory bank, the deep learning model can refer to the historical conversation between the user and the intelligent system to generate an answer to the user's current round of input, thereby improving the quality of the answer finally output by the intelligent system.
[0047] In some embodiments, the first intermediate result may also include the historical input data item itself whose similarity to the input data is higher than the first threshold. In some embodiments, historical conversation information related to the user's input data can be obtained by calculating the dense vector similarity. The dense vector similarity can be expressed as follows:
[0048]
Number
[0049] Here,
[0050]
Number
[0051] It should be understood that the above-described similarity calculation process may be realized by a neural network. The similarity between the user's input data (or both the user's input data and the context information, or the first intermediate query obtained based on the user's input data) and each historical input data item (or both the historical input data item and the corresponding historical answer item) in the external memory bank can be calculated, and the historical answer item (and optionally, the historical input data item) in one or more data groups that satisfy the condition that the similarity s is greater than the preset first threshold δ can be returned to the understanding generation large-scale model. In some embodiments, the historical answer items that need to be returned based on the similarity may be determined by other methods such as Top K, which is not limited here.
[0052] In some embodiments, the external memory bank may be optimized in association with the understanding generation large-scale model as described below. According to some embodiments, the first intermediate query may be based on the input data, and the first intermediate result may be the historical answer item corresponding to the historical input data item whose similarity with the input data in the first data group set is higher than the first threshold and whose time stamp is the latest. Thereby, when a plurality of historical answer items related to the input data are obtained, by returning the historical answer item with the latest time stamp, the deep learning model generates an answer based on the latest correlation memory and fully utilizes the timeliness of the memory.
[0053] In some embodiments, the historical input data item itself in the first data group set whose similarity with the input data is higher than the first threshold and whose time stamp is the latest may be returned to the deep learning model.
[0054] In some embodiments, as shown in FIG. 3, the user and the intelligent system 310 have experienced a conversation regarding going out with a pet named Beibei twice historically. The intelligent system 310 may be, for example, the system described above that deploys an understanding and generation large-scale model and can interact with the user. In the current conversation, the intelligent system 310 obtains a user input of "Recently, I want to take Beibei to play with a friend I knew before", and based on this user input, performs a memory acquisition in the external memory bank 320 to retrieve a historical input data item with a timestamp of 20XX0812, "Recently, I want to take Beibei to the pet park, but is there any recommended place?", and the corresponding historical response item "You can walk to XX Land, there are many pet attractions", and a historical input data item with a timestamp of 20XX0817, "Tomorrow, I want to go to the suburbs with Beibei and breathe some fresh air", and the corresponding historical response item "YY Park is a good choice". Furthermore, the intelligent system can return the historical conversation with the latest timestamp to the deep learning model, and the deep learning model generates a response of "Are you going to YY Park? You can meet a lot of friends there" based on this historical conversation. It should be understood that the intelligent system can also provide both of the two retrieved historical conversations to the model for generating a response by the model.
[0055] Through the above embodiments, by using the external memory bank, it is possible to record the historical conversations generated between the user and the intelligent system in previous conversations (e.g., one week ago, one month ago, or earlier), improve the memory capacity of the intelligent system, and use relevant historical conversations as a reference when generating a response to the user's current input, thereby generating a more targeted, richer, and more specific response to the user. As a result, it can be seen that the response quality is improved, the intelligence of the conversation is enhanced, and the user experience is improved.
[0056] The foregoing embodiments have described the search operation of the external memory bank. Hereinafter, operations such as adding and deleting data groups in the external memory bank will be described. FIG. 4 is a schematic diagram showing operations such as adding and deleting data groups in the external memory bank 420 according to an exemplary embodiment. The intelligent system 410 may be, for example, the system described above that deploys the understanding generation large-scale model and can interact with the user. Note that the query operation of the external memory bank is performed in the process of generating answer data for the user input data using the deep learning model, and operations such as addition and deletion are performed after the generation of the answer data by the deep learning model.
[0057] According to some embodiments, the data generation method may further include entering the first data group into the first data group set in response to determining that the similarity between the first data group based on the input data and the answer and any data group in the first data group set is less than a second threshold.
[0058] In some embodiments, for the user input data u t‐1 and the answer data r t‐1 of the deep learning model in the (t-1)-th round, if the similarity between the first data group m t‐1 =(u t‐1 , r t‐1 ) and the data groups in the external memory bank M is also lower than a preset second threshold, m t‐1 =(u t‐1 , r t‐1 ) is added to the external memory bank M.
[0059] According to some embodiments, the data generation method may further include entering a first data group into a first data group set and deleting a second data group from the first data group set in response to determining that the similarity between the first data group based on input data and an answer and a second data group in the first data group set is higher than a third threshold and the first data group and the second data group conflict with each other.
[0060] In some embodiments, for the input data u of the user in the (t - 1)-th round t‐1 and the answer data r of the deep learning model t‐1 , when it is determined that the similarity between the first data group m t‐1 =(u t‐1 , r t‐1 ) and a second data group m i ∈M in the external memory bank M is higher than the third threshold and the consistency between m t‐1 and m i collides, m i is deleted and m t‐1 is added to M. In one exemplary embodiment, the consistency judgment (e.g., collision detection) between m t‐1 and m i may be performed using a neural network based on both semantic vectors, or may be implemented in other ways, and is not limited herein.
[0061] Thereby, by the above method, adding and deleting data groups to and from the external memory bank is realized, the flexibility of data group operations in the external memory bank is improved, and the timeliness and content accuracy of the data groups in the external memory bank are improved.
[0062] In some embodiments, as shown in FIG. 4, after the deep learning model generates an answer to the user input, the current conversation (including the user input and the answer generated by the model) can be added to the external memory bank. If the current conversation content conflicts with the historical conversation in the external memory bank, the historical conversation in the external memory bank can be deleted.
[0063] According to some embodiments, the data generation method can further include deleting, from the external memory bank, data groups with old timeliness based on the entry time item. In some exemplary embodiments, a retention period for the data group can be set, and data groups exceeding that period can be deleted. Timeliness checks can be performed periodically or irregularly based on the content of the data group, and data groups that fail the check can be deleted. It is also possible to implement deleting data groups with old timeliness from the external memory bank in other ways. Thereby, by the above method, it is guaranteed that not all data groups in the external memory bank will become old, and the timeliness of memory is improved.
[0064] In some embodiments, at the stage of constructing the initial input of the deep learning model (i.e., before using the deep learning model to process the initial input), the intelligent system can directly obtain, from the external memory bank, the historical conversation information corresponding to the input data of the user in the current round, and determine the initial input of the deep learning model based on the historical conversation information.
[0065] According to some embodiments, as shown in FIG. 5, step S201 of determining the initial input used in the deep learning model may include step S501 of obtaining, based on the input data, a historical response item corresponding to a historical input data item whose similarity to the input data is higher than a first threshold value from an external memory bank, and step S502 of determining the initial input based on the input data and the historical response item. The operation of step S501 can refer to the above description regarding the acquisition of the first intermediate result, and it should be understood that it will not be described here. Thereby, it can be guaranteed that every time the deep learning model generates an answer, it can refer to the historical conversation information obtained from the external memory bank.
[0066] In some embodiments, the input data of the user and the historical response item can be directly stitched to obtain the initial input of the deep learning model, and the input data of the user and the historical response item can also be processed in other ways to obtain the initial input of the deep learning model, but it is not limited here.
[0067] The effect of memory capacity reinforcement for the deep learning model and the intelligent system will be further described below in relation to some exemplary embodiments. In one exemplary embodiment, as shown in FIG. 6, the dialogue system 610 without an external memory bank cannot form long-term memory. Therefore, when the user queries about the content of the historical conversation, the system can only mechanically answer. The intelligent system 620 with the external memory bank described in the present disclosure can obtain the corresponding historical conversation from the external memory bank 630 for the user input, thereby generating an answer that meets the user's needs, embodying the reinforcement of the memory capacity of the deep learning model and the intelligent system.
[0068] In some embodiments, the first functional component may be other functional components such as an external search engine, a search model, an application programming interface, etc. These different functional components each have corresponding tokens. In step S202, the deep learning model determines whether to call an external functional component (and / or which functional component to call), and the determination result is embodied in whether the result output by the deep learning model contains a token corresponding to the call of the external functional component (and / or which token corresponding to which functional component is specifically included in the result). Note that external functional components such as an external search engine, a search model, and an application programming interface do not necessarily require context information and / or an external memory bank. In other words, these external functional components can be called by the deep learning model alone.
[0069] In some embodiments, when a deep learning model based on a Transformer network structure makes a prediction, the model first receives an initial input and generates a first output token, token_1. Next, the model receives token_1 and generates a second output token, token_2. The loop call to the deep learning model is repeated until the token_n output by the model indicates the completion of the model output. Each token output by the model can correspond to a specific external functional component, embody the determination result of whether to call the external functional component, and may also be in the form of a specific markup, or a specific single word, character, or word, so as to generate an answer to the user input, or may also be a special symbol indicating that the current content has already been generated. Therefore, it is realized to automatically make a decision using the model and then determine the task that needs to be executed next (for example, calling an external functional component or generating an answer).
[0070] Figure 7 shows a schematic diagram of a deep learning model generating an answer based on an initial input according to an exemplary embodiment. The structure of the understanding generation large-scale model 710 (i.e., the deep learning model) may be UniLM. First, the initial input of the model based on the user's input data (and optionally, context information) is input into the deep learning model to obtain the first token output by the model, and the corresponding content is <api1>is. This token reflects the model's decision that it is necessary to call the functional component API1. The model can continue to output to generate a first intermediate query input_1 that can be identified by API1. This process can also be understood as rewriting the user's input data in order to obtain call information that can be identified by API1 and from which the desired result can be obtained from API1. After outputting input_1, the model marks up < / api1> capable of outputting the corresponding token, indicating that the first intermediate query for API 1 has already been generated. The first output can include the complete <api1>input_1< / api1> .
[0071] In some embodiments, the first intermediate query input_1 corresponding to API 1 may be generated word by word by repeatedly calling the deep learning model, that is, each time, the user's input data and the part generated in input_1 are input into the model to obtain the next single word, character, or markup in input_1. input_1 may be obtained by decoding a single token output by the deep learning model. input_1 can also be obtained from the tokens output by the model in other ways, which is not limited here.
[0072] After the first intermediate query input_1 is obtained, input_1 is used to call API1 to obtain the first intermediate result <api1-r>result_1< / api1-r> . Further, by combining the user's input data and the first intermediate result, a second input used in the deep learning model can be obtained to obtain the next token output by the model. In some embodiments, when determining the second input, the first intermediate query (or the complete first output) can also be incorporated. As shown in Figure 7, the downward dashed arrow of the first output <api1>input_1< / api1> and the left dashed block of the first intermediate result <api1-r>result_1< / api1-r> are shown. This dashed block may be the first intermediate query input_1 or the complete first output <api1>input_1< / api1>It may also be. In one exemplary embodiment, the second input is the stitching of the initial input of the model, the first output, and the first intermediate result.
[0073] According to some embodiments, step S204 of determining the second input used in the deep learning model based on at least the initial input and the first intermediate result may include determining the second input used in the deep learning model based on the initial input, the first intermediate result, and the first intermediate query. In this way, by using the first intermediate query as a reference factor for the deep learning model to generate the second output, the accuracy of model determination can be further improved, and the quality of the finally generated answer can be improved.
[0074] The second token generated by the deep learning model based on the second input can continue to output tokens corresponding to the corresponding content <api2>and this token reflects the model's decision that it is necessary to call the functional component API 2. The model has a second intermediate query input_2 and markup < / api2> and continue to output tokens corresponding to the corresponding content. Further, API 2 can be called using input_2 to obtain the second intermediate result <api2-r>result_2< / api2-r> and combine the user input data with the second intermediate result (and optionally the second intermediate query) to obtain the third input used in the deep learning model. In one exemplary embodiment, the third input is the stitching of the initial input of the model, the first output, the first intermediate result, the second output, and the second intermediate result.
[0075] The third token generated by the deep learning model based on the third input does not correspond to any external functional component. Therefore, this third token can be used to instruct the model to start generating an answer to the initial input of the model (which can also be understood as the input data to the user). In some embodiments, the third token may be the first single word, character, or word in the answer, or a special symbol without semantic information for indicating that the model generates the answer from the next token. Next, the model generates the answer word by word and finally generates a special symbol indicating that the answer has been generated.
[0076] Note that the calls to different external function components are independent of each other and there is no pre-set order relationship. The tokens output by the model determine which external function component needs to be called. Therefore, in some exemplary embodiments, the model may determine whether to call the same function component multiple times or to call multiple function components in a specific logical order based on an understanding of the user input to perform a specific task.
[0077] In this way, by causing the understanding generation large-scale model to output tokens with different meanings, the model can automatically determine the task (e.g., calling a specific external function component or directly generating an answer) and the execution order that needs to be performed based on an understanding of the user input (and optionally, context information), realizing automated understanding, reasoning, decision-making, and generation using a single deep learning model, and improving the intelligence of the system.
[0078] In some embodiments, the UniLM model has only one input. Therefore, in step S204, the initial input and the first intermediate result can be combined by means such as stitching to obtain the second input of the user deep learning model.
[0079] In some embodiments, when adopting an N-layer Transformer network structure having an encoder and a decoder, the input of the encoder may be the initial input of the model, and the output of the encoder may be the encoding result for the initial input. The two inputs of the decoder are, respectively, the encoding result for the initial input output by the encoder and all the tokens that the model has already generated. The output of the decoder is the next token to be predicted. Therefore, in step S204, the first intermediate result and the encoding result for the initial input can be used as the two inputs to the decoder, respectively.
[0080] According to some embodiments, the first functional component may be an external search engine. The external search engine may be a general-purpose search engine, a knowledge engine or a specialized knowledge library customized for a specific field, or a private database, thereby obtaining different types of knowledge and updating the knowledge in real time.
[0081] The first intermediate query generated by the deep learning model may be, for example, a search query, whereby the external search engine can be utilized to perform a search based on this search query in order to obtain one or more search results. In some embodiments, one or more search results returned by the search engine may be directly used as the first intermediate result, or these search results may be processed to obtain the first intermediate result. Subsequently, a second input for being processed by the deep learning model can be determined based on the initial input (e.g., the user input data and, optionally, context information) of the deep learning model and the first intermediate result (e.g., one or more search results). For the second input, the deep learning model may determine that it is necessary to further call a second functional component, and as will be described below, it may also determine that it is not necessary to call other functional components and directly generate an answer to the initial input.
[0082] In some embodiments, the initial input and the first intermediate result can be combined by means such as stitching to obtain the second input. First, each search result can be processed by means such as content extraction, rewriting, semantic vector calculation, or other methods, and subsequently, the initial input and the processed search results can be combined by means such as stitching to obtain the second input, but it is not limited herein.
[0083] In some embodiments, through training, data can be fully internalized into the model in a parameterized manner, and such a model can be used to directly generate an answer to user input. In this mechanism, for relatively unpopular factual information, since the frequency of occurrence in the training data is low, the learning of the model may not be solid, so it may "forget" or "have disrupted memory".
[0084] Thereby, by obtaining search results from an external search engine, various types of accurate knowledge, information, and timely data are accurately and promptly transmitted to the upper-level understanding generation large-scale model, and the understanding generation large-scale model combines the explicitly searched information and the knowledge internalized in the model to jointly complete the satisfaction and answer to the user's needs. Also, the understanding generation model generates the final answer based on one or more search results included in the second input, realizes the alignment and processing of the searched information, and thereby can output an answer that conforms to the user's intention, improving the quality of the answer data.
[0085] According to some embodiments, the first functional component is a search model trained in association with a deep learning model. The search model may be a large-scale model based on an end-to-end Transformer structure that can further include a recall model and a sorting model. The search model can also be realized by a single neural network model (for example, a large-scale model based on an end-to-end Transformer structure). The joint training of the deep learning model and the search model will be described later.
[0086] The first intermediate query generated by the deep learning model may be, for example, a search query, whereby in order to obtain one or more search results, it can be searched using a search model trained in association with the deep learning model. The processing of the search results can refer to the above-mentioned processing of the search results returned by the search engine, and it should be understood that it will not be described here.
[0087] In this way, by using an external search model, the above advantages of using an external search engine can be realized. On the other hand, since the external search model and the understanding and generation large-scale model are jointly optimized, the two cooperate, and the external search model can provide the understanding and generation large-scale model with more accurate and more appropriate content for answer generation. The understanding and generation large-scale model can better integrate and process the search results, thereby generating high-quality answers that more conform to the user's intention. Therefore, by using an external search engine or an external search model, knowledge reinforcement for deep learning models and intelligent systems can be realized.
[0088] Hereinafter, the effect of knowledge reinforcement on deep learning models and intelligent systems will be further described in relation to several exemplary embodiments. In one exemplary embodiment, as shown in FIG. 8, in the dialogue system 810 without knowledge reinforcement, the internalized knowledge is limited and cannot provide accurate answers when encountering more knowledge-intensive queries. Furthermore, the dialogue system 810 cannot update knowledge in real time, and therefore the results it outputs may become outdated or incorrect. The intelligent system 820 with knowledge reinforcement described in the present disclosure can perform a search on the user input using an external search engine / search model 830, thereby obtaining accurate knowledge content and improving the accuracy of the knowledge. For the question from the user, "What is the famous poem written by the son of the monarch of Wei during the Three Kingdoms period?", the search engine / search model 830 returns two relevant results. One of them shows that the monarch of Wei during the Three Kingdoms period was Cao Cao and he had sons Cao Pi and Cao Zhi. The other shows that the poem "Seven-Step Poem" by Cao Zhi, the son of Cao Cao, is famous. The deep learning model combines these two search results obtained from the outside with its own internalized knowledge and then gives an accurate answer.
[0089] In addition, since the databases, knowledge bases, and resource repositories behind external search engines and search models are updated in real time, the knowledge obtained through searches is more up-to-date. This shows knowledge enhancement for deep learning models and intelligent systems.
[0090] According to some embodiments, the first functional component is at least one application programming interface (API) that can be called by a deep learning model. Different APIs each have a corresponding markup format, that is, a token for calling this API. When the deep learning model makes a prediction and outputs a token / markup corresponding to a specific API, the intelligent system recognizes that it needs to trigger this API. Next, the model continues to output an intermediate query that can be identified by this API (that is, the input used for this API, also called the rewritten query). Furthermore, based on the intermediate result obtained by calling this API with the intermediate query, a second input for re-inputting into the deep learning model can be determined, and the prediction by the model can be continued. Regarding the second input, the decision of the deep learning model may further require calling a second functional component (search engine, search model, or other API), or may not require calling other functional components and may directly generate an answer to the initial input.
[0091] As described above, in the process of generating an answer by the model for a single round, all APIs (or all external functional modules) may be called, or only some APIs may be called, and both the call order and the number of calls of these APIs are determined by the model.
[0092] In some embodiments, the APIs used in the intelligent system can include scientific calculators, form processing tools, smart home controls, etc. By calling the APIs that can execute various tasks, the intelligent system can be enhanced in capabilities. By using external functional components such as scientific calculators, the problem that the logical calculation ability of the deep learning model is weak can be solved, and the logical reasoning ability of the entire intelligent system can be improved. Compared with the method of calling the API using the mapping table of keywords and API call instructions, the deep learning model is used to directly generate intermediate queries that can be identified by the API, and the acquisition of the intermediate queries and intermediate results is adapted according to the potential intention in the user's initial input, ultimately improving the quality of the generated answer and the intelligence of the system. Also, by combining the understanding generation large-scale model and the API, the intelligent system is given the ability to execute automated operations, realizing the enhancement of capabilities for the deep learning model and the intelligent system.
[0093] The effects of enhancing the capabilities of the deep learning model and the intelligent system in relation to some exemplary embodiments will be further described below. In one exemplary embodiment, as shown in FIG. 9, the dialogue system 910 without the ability to enhance capabilities (e.g., the ability to call external APIs) has limited tasks that can be completed and cannot handle tasks that require the invocation of external functional components such as weather inquiries and mathematical calculations. The intelligent system 920 with the ability to enhance capabilities described in this disclosure can determine the API 930 that needs to be called for the user input, further call this API 930, and process the returned results to generate an answer that meets the user's needs, demonstrating the enhancement of capabilities for the deep learning model and the intelligent system.
[0094] According to some embodiments, the second output can include a second token for invoking a second functional component and a second intermediate query obtained based on the second input and identifiable by the second functional component. The second functional component may be the same as the first functional component (i.e., the same functional component may be invoked multiple times), or may be different from the first functional component, and it should be understood that there is no limitation here.
[0095] According to some embodiments, as shown in FIG. 10, in order to generate an answer to the initial input, step S205 of obtaining the second output of the deep learning model is step S1001 of performing a corresponding function call operation on the second output, where the function call operation obtains a second intermediate result determined by the second functional component based on the second intermediate query, determines a third input used in the deep learning model based on at least the second input and the second intermediate result, and obtains the third output of the deep learning model, and in response to including in the Nth output of the deep learning model a token for invoking an arbitrary functional component different from the deep learning model and an Nth intermediate query identifiable by the Nth functional component obtained based on the Nth token for invoking the Nth functional component and the Nth input, performing a function call operation corresponding to the Nth output until it is determined that the corresponding token for invoking an arbitrary functional component different from the deep learning model is not included in the (N + 1)th output, and using the (N + 1)th output as the answer to the initial input, where N is an integer greater than 2, and step S1002.
[0096] Therefore, by the above method, the deep learning model can invoke external functional components multiple times until the model determines that the invocation of external functional components is no longer necessary.
[0097] According to some embodiments, the second functional component and the Nth functional component may each be one of a group of functional components including an external search engine, a search model trained in association with a deep learning model, at least one application programming interface that can be called by the deep learning model, and an external memory bank. A first set of data groups related to the user is stored in the external memory bank. Here, each data group in the first set of data groups includes at least a historical input data item and a historical answer item generated by the deep learning model for the historical input data item.
[0098] According to some embodiments, the second output may not include a corresponding token for calling any functional component different from the deep learning model. To generate an answer to the initial input, step S205 of obtaining the second output of the deep learning model may include using the second output as the answer to the initial input. Thereby, when the second output generated by the model does not include any token corresponding to any functional component, the final answer output by the model for the initial input can be obtained.
[0099] The effects of reinforcing multiple capabilities of the deep learning model and the intelligent system in relation to some exemplary embodiments will be further described below. In one exemplary embodiment, as shown in FIG. 11, the dialogue system 1110 without ability reinforcement has simple answer content generated based on the knowledge internalized in the model and cannot complete the task described in the user input, and thus cannot meet the user needs. The intelligent system 1120 with ability reinforcement described in the present disclosure can accurately understand the intention indicated by the user input, and further utilize external components such as the external memory bank 1130, the search engine / search model 1140, and the API 1150 to accurately complete many tasks such as historical memory query, text generation, and email sending by API call, and can execute the above tasks with accurate logic.
[0100] Also, when generating text, the model can use an external search engine / search model to obtain explicit information as the material for the text, and use the internalized knowledge to extract, integrate, and modify these materials, and generate the beginning, end, and transitional paragraphs to form a complete text. As shown in FIG. 11, among the texts generated by the intelligent system 1120, the two texts "City X is a beautiful city" and "If there is a chance to travel to City X, you will surely like this city" are generated based on the knowledge internalized in the model. The three intermediate contents regarding the travel season, gourmet food, and travel methods are extracted from three search results respectively and generated after being modified based on the search results. Thus, high-quality answer content can be generated by the above method.
[0101] In one exemplary embodiment, as shown in FIG. 12, the dialogue system 1210 without ability enhancement cannot obtain the historical dialogue with the user, thus cannot complete the task described in the user input, and therefore cannot meet the user's needs. In contrast, the intelligent system 1220 with ability enhancement described in the present disclosure can accurately understand the intention indicated by the user input, and use external components such as the external memory bank 1230, API 1240, and search engine / search model 1250 to accurately complete many tasks such as historical memory query, music playback by API call, and lyric search, and can execute the above tasks with accurate logic. This shows the enhancement of multiple capabilities of the deep learning model and the intelligent system.
[0102] Return to step S201. According to some embodiments, the initial input can include the context information of the input data. The context information can include a plurality of dialogues conducted between the user and the intelligent system before the obtained user input data.
[0103] In some embodiments, the context information includes a plurality of conversations that the user has with the intelligent system in the current conversation with the intelligent system, but does not include the conversations transmitted in the historical conversations between the user and the intelligent system. In other words, after the user shuts down the application or service of the intelligent system, the context information is cleared accordingly, and when the user starts the application or service of the intelligent system again, the recording of the context information is resumed.
[0104] Furthermore, it is limited by the upper limit of the input length of the deep learning model, and the context information usually has a preset maximum encodable length and limited memory capacity. Therefore, when the user has multiple conversations with the intelligent system or the content is long, some of the context information may be discarded.
[0105] According to some embodiments, when obtaining historical conversation information from an external memory bank, the context information may be used as a reference based on the user's input data. In addition to the historical answer items, the corresponding historical input data items may also be obtained. As shown in FIG. 13, the step S201 of determining the initial input used in the deep learning model may include step S1301 of obtaining at least one pair of historical input data items and historical answer items whose similarity between the input data and the context information from the external memory bank meets a fourth threshold value, and step S1302 of determining the initial input used in the deep learning model based on the input data, the context information, and at least one pair of historical input data items and historical answer items. Thereby, by performing similarity calculation using both the user's input data and the context information, more effective historical conversation information can be obtained from the external memory bank. On the other hand, by using the input data, the context information, and the corresponding at least one pair of historical input data items and historical answer items, the quality of the answer generated by the deep learning model can be further improved.
[0106] In some embodiments, for other external functional components, both the user input data and the context information may be used as references when generating the corresponding first intermediate query.
[0107] It should be understood that when implementing the method of the present disclosure, if necessary, the first threshold, the second threshold, the third threshold, and the fourth threshold can be set. The values of these preset thresholds may be the same or different and are not limited herein.
[0108] The intelligent system and the understanding generation large-scale model arranged therein can richly present the generated answers and can interact with the user to improve the user experience.
[0109] In some embodiments, the dialogue system may generate a final answer from a single search result, and there may be a possibility of generating an incomplete answer or an incorrect answer. As shown in FIG. 14, the intelligent system of the present disclosure can implement an answer aggregation presentation method (both single-answer aggregation and multi-answer aggregation are achievable) by performing online calculation after searching or retrieval.
[0110] In some embodiments, as shown in FIG. 15, in addition to aggregating and presenting the retrieved content, the intelligent system can generate answers by itself, such as mathematical reasoning and common sense reasoning related to disciplines, in addition to writing poems, novels, emails, summary reports, compositions, marketing documents, etc. For these results, the intelligent system can perform a structured presentation.
[0111] In some embodiments, the intelligent system can interact with the user multiple times to achieve interactive presentation, including clarification, active guidance, in-depth topic Q&A, and execution of certain instructions. In some exemplary embodiments, as shown in part A of FIG. 16, the intelligent system can actively clarify the theme and content of the conversation to the user and generate content that meets the user's desires. As shown in part B of FIG. 16, the intelligent system can actively guide the user and explore the user's specific needs.
[0112] According to another aspect of the present disclosure, a method for training a deep learning model is provided. The deep learning model is used to generate response data based on user input data. As shown in FIG. 17, the training method includes: acquiring first sample data, where the first sample data includes a first sample initial input and a first sample output. Here, the first sample initial input includes an intention expression for invoking a first preset functional component different from the deep learning model, and the first sample output includes a first token for invoking the first preset functional component and a first sample intermediate input identifiable by the first preset functional component (step S1701); acquiring second sample data, where the second sample data includes a second sample initial input and a second sample output. Here, the second sample initial input does not include an intention expression for invoking any preset functional component different from the deep learning model, and the second sample output does not include the corresponding token for invoking any preset functional component (step S1702); using the deep learning model to process the first sample initial input to obtain a first predicted output (step S1703); adjusting the parameters of the deep learning model based on the comparison between the first sample output and the first predicted output (step S1704); using the deep learning model to process the second sample initial input to obtain a second predicted output (step S1705); and adjusting the parameters of the deep learning model based on the comparison between the second sample output and the second predicted output (step S1706).
[0113] Therefore, by training the deep learning model as described above, when the trained deep learning model needs to call a specific preset function component, it can output a token corresponding to the preset function component and an intermediate input that can be identified from this preset function component, and when there is no need to call any function component, it can generate output content that does not include a token and an intermediate input corresponding to any preset function component. Thereby, the model is provided with the ability to execute tasks such as understanding, decision-making, and generation, and the deep learning model can be enhanced in ability by using external function components, improving the quality of the generated response data.
[0114] In some embodiments, before step S1701, first, hybrid training of language text and a priori knowledge may be performed on the understanding and generation large-scale model.
[0115] The understanding and generation large-scale model can be trained with a large amount of text data (for example, Internet data), knowledge maps, and weakly supervised data. In addition, it is also important to add artificially compiled knowledge to the model. The artificially compiled a priori knowledge helps the model better understand language, generate language, make decisions, and enables the model to interact with humans efficiently and smoothly. The specific steps include the following.
[0116] 1) Collect text data on the Internet and perform low-quality and noise removal processing on it to remove invalid and redundant information in the big data. 2) Integrate a priori knowledge, mainly including three types of knowledge: A. An enormous Internet-based knowledge map: including <entity-attribute-attribute value> or <entity-relationship-entity 2>; for example, <Star A-height-172>, <Star A-spouse-Star B>; B. High-quality manual a priori annotation data: Manually label various types of tasks. For example, for classification label data, "XX was elected as the new men's basketball chairman" is labeled as <"XX was elected as the new men's basketball chairman" - "Sports"; or, question-and-answer data: <"Will eating chocolate for a long time cause diabetes?" "No">; C. Industry knowledge: For example, dictionaries in the medical, safety, transportation, finance, and energy industries, and structured knowledge of the industry; 3) As shown in FIG. 18, in the knowledge fusion technology, the above three types of structured knowledge 1810 are converted into a natural language description form (i.e., natural language-formatted data 1830) by the verbalization template 1820, and then mixed and learned with Internet text data. In one exemplary embodiment, the structured knowledge <Star A - couple - Star B> can be converted into data in the natural language form of "The wife of Star A is Star B" by the verbalization template. Through the method of mixed learning, the model can better understand natural language, thereby having basic dialogue and interaction capabilities.
[0117] In some embodiments, for the first sample data obtained in step S1701 and the second sample data obtained in step S1702, the first sample initial input and the second sample initial input may be true user data or constructed data, and may include input data (and optionally, context information). The first sample initial input includes an intentional expression for calling a first preset functional component different from the deep learning model, that is, the content described by the first sample initial input requires or desires the model to call the first preset functional component. The second sample initial input does not include an intentional expression for calling any preset functional component different from the deep learning model, that is, the content described by the second sample initial input does not require or desire the model to call any preset functional component. The first sample output and the second sample output may be results desired to be output by the deep learning model, that is, ground truth.
[0118] In some embodiments, the first token included in the first sample output corresponds to the corresponding first preset functional component, whereby the trained deep learning model needs to call the first preset functional component by this token. In some embodiments, the first token output by the model can be encoded in a markup format corresponding to this first preset functional component, and the API call result can be converted into a string, whereby the trained model can perform determination, call information generation, and understanding of the call result in a text processing manner.
[0119] In some embodiments, the first sample intermediate input included in the first sample output can be processed by an external first preset function component to obtain the result returned by this first preset function component. When the first preset function component is an external memory bank, the first sample intermediate input may be the input data of the user (and optionally context information) for which the external memory bank can calculate the similarity. When the first preset function component is a search engine, the first sample intermediate input may be a search expression that can be identified by the search engine. When the first preset function component is a search model, the first sample intermediate input may be a search query that can be processed by the search model. When the first preset function component is a specific API, the first sample intermediate input can be encoded to have a markup format corresponding to this API. In this way, the trained model can have the ability to output intermediate inputs that can be identified by these preset function components.
[0120] In some embodiments, the first prediction output output by the deep learning model obtained in step S1703 may be close to or completely different from the first sample output, but the goal of training the deep learning model, that is, the first prediction output generated by the trained model, includes a token for calling the first preset function component and a predicted intermediate input that can be identified by the first preset function component and that matches the function or meaning of the first sample intermediate input.
[0121] In some embodiments, the second sample output does not include a corresponding token for invoking any preset functional component, and thus, the second sample output should be the answer of the deep learning model to the second sample initial input. The second predicted output output by the deep learning model obtained in step S1705 may be close to the second sample output or may be completely different, but the goal of training the deep learning model is to make the second predicted output generated by the trained model not include a token for invoking any preset functional component and include high-quality answer data for the second sample initial input.
[0122] In some embodiments, in step S1704 and step S1706, a corresponding loss function is determined based on the demand, a loss value describing the difference between the sample output and the predicted output is calculated, and further, based on the loss value, the parameters of the deep learning model are adjusted.
[0123] In some embodiments, the first sample data can further include a first sample target input and a first sample answer. The first sample target input includes the first sample initial input and a first sample intermediate result obtained from a first preset functional component based on the first sample intermediate input. In some embodiments, the first sample target input can further include the first sample intermediate input. The first sample answer is the true (ground truth) answer to the first sample initial input constructed using the first sample intermediate result. The training method can include using the deep learning model to process the first sample target input to obtain a first predicted answer and adjusting the parameters of the deep learning model based on the comparison between the first sample answer and the first predicted answer.
[0124] As a result, the deep learning model after training, combined with the results obtained from external functional components and the knowledge internalized in the model, can complete the satisfaction and response to the user's needs, and finally obtain high-quality response content.
[0125] According to some embodiments, as shown in FIG. 19, the training method includes obtaining third sample data including a third sample initial input, a sample search query, a plurality of sample search results, and a third sample response of the deep learning model to the third sample initial input, where the sample search query is a sample intermediate input generated by the deep learning model based on the third sample initial input, the sample intermediate input is distinguishable by a search model different from the deep learning model, and here, the plurality of sample search results are the results output by the search model based on the sample search query, step S1907; performing a sorting operation on the plurality of sample search results based on the degree of match between each of the plurality of sample search results and the third sample response, step S1908; and further including training the search model based on the sorted plurality of sample search results, step S1909. Since steps S1901 to S1906 in FIG. 19 are respectively the same as steps S1701 to S1706 in FIG. 17, it should be understood that the description here is omitted.
[0126] Thereby, by determining the sorting result of the plurality of sample search results in the third sample data, and using the sorting result as supervision to train the search model, the joint optimization of the understanding generation large-scale model and the search model is realized, enabling the two to cooperate, and the external search model can provide the understanding generation large-scale model with more accurate and more answer-generation suitable content, so that the understanding generation large-scale model can generate answers that are more in line with the user's intention and of higher quality.
[0127] In some embodiments, the sample search query included in the third sample data is, for example, the search query query, and the plurality of sample search results are, for example, a plurality of contents that match the needs of the third sample initial input in the search library used by the search model and are aligned to generate a third sample answer for the third sample initial input. The third sample answer may be the content obtained after manually performing steps such as selection, modification, and refinement on the plurality of sample search results. In some embodiments, referring to steps S1701, S1703 to S1704 in FIG. 17, by training a deep learning model using the third sample data, the deep learning model has the ability to automatically perform the above-mentioned steps such as selection, modification, and refinement.
[0128] In some embodiments, in step S1908, the content consistency between the plurality of sample search results and the third sample answer may be calculated based on, for example, similarity calculation based on semantic vectors.
[0129] According to some embodiments, as shown in FIG. 20, based on the consistency between each of the plurality of sample search results and the third sample answer, step S1908 of performing a sorting operation on the plurality of sample search results may include step S2001 of screening the first sample search result with the highest current consistency from the plurality of sample search results, step S2002 of deleting the overlapping content between the third sample answer and the first sample search result and updating the third sample answer, and step S2003 of repeating the sorting operation on the remaining part until the sorting of all the sample search results in the plurality of sample search results is completed, based on the consistency between each of the remaining parts of the plurality of sample search results and the updated third sample answer.
[0130] In this way, sorting of a plurality of sample search results for generating a third sample answer is realized, whereby, the joint optimization of the understanding generation large-scale model and the search model can be realized.
[0131] According to some embodiments, the search model may include a sorting sub-model and a recall sub-model. The step S1909 of training the search model based on the sorted plurality of sample search results may include training the sorting sub-model of the search model based on the sorted plurality of sample search results, and using the trained sorting sub-model as a teacher model to train the recall sub-model. Thereby, the joint optimization among the understanding generation large-scale model, the sorting sub-model in the search model, and the recall sub-model can be realized by the above method.
[0132] In some embodiments, the sorting sub-model is a cross-encoder model for end-to-end search. The input of the cross-encoder model consists of a query (q) and a passage (p), and the output is the similarity sim(q, p) between the two. A listwise loss can be used as supervision, whereby the sorting result output by the cross-encoder model is approximated or made to coincide with the sorting result generated for the plurality of sample search results.
[0133] In some embodiments, the recall sub-model may be a bi-encoder model. Here, one encoder is used to generate the feature vector of query q, and the other encoder is used to generate the feature vector of document p. From these two feature vectors, the similarity between them can be calculated. After the sorting model is trained, in the way of model distillation, the sorting model is used as the teacher model to construct training samples for the recall model, so that the optimization target of the recall model is consistent with that of the sorting model, and further realize the joint optimization of the understanding generation large-scale model and the retrieval model. In one exemplary embodiment, KL-divergence can be used as the supervision to train the recall model by using the sorting model as the teacher model.
[0134] In some embodiments, the end-to-end retrieval model can be trained alone before joint training. One exemplary embodiment can jointly train the recall sub-model and the sorting sub-model.
[0135] According to some embodiments, as shown in FIG. 21, the training method includes obtaining fourth sample data, where the fourth sample data includes a fourth sample initial input, a fourth sample intermediate input identifiable by an external memory bank, a sample memory result, and a fourth sample answer, and the fourth sample intermediate input is determined based on the fourth sample initial input, step S2107; obtaining a predicted memory result determined based on the fourth sample intermediate input by the external memory bank, step S2108; adjusting the parameters of the external memory bank based on a comparison between the predicted memory result and the sample memory result, step S2109; determining a fourth sample target input used for the deep learning model based on at least the fourth sample initial input and the sample memory result, step S2110; processing the fourth sample target input using the deep learning model to obtain a fourth predicted answer, step S2111; and adjusting the parameters of the deep learning model based on a comparison between the fourth sample answer and the fourth predicted answer, step S2112. It should be understood that the operations of steps S2101 to S2106 in FIG. 21 are the same as the operations of steps S1701 to S1706 in FIG. 17 respectively, so the description here is omitted. Thereby, the joint training of the external memory bank and the understanding generation large-scale model is realized.
[0136] It should be understood that the external memory bank obtained as described above can be used in the data generation method described above as an external functional component for the acquisition of the external memory bank.
[0137] In some embodiments, the training objective of the joint training of the memory query and the understanding generation large-scale model may be to maximize the answer generation probability of memory enhancement.
[0138]
Number
[0139] Here, M is the external memory bank, and ct is a sample intermediate input corresponding to an external memory bank, and may include a sample initial input and context information, m i is a queried history conversation (i.e., a data group), and r is an answer generated by a deep learning model. Correspondingly
[0140]
Number
[0141] is a memory query process,
[0142]
Number
[0143] is a memory-augmented answer generation process. By performing joint optimization on the external memory bank and the understanding and generation large-scale model based on the training target, the correlation between the external memory bank after joint optimization and the user input is higher, and a history conversation useful for answer generation is provided. The understanding and generation large-scale model after joint optimization can generate high-quality answer content for the user input based on the obtained history conversation.
[0144] In some embodiments, as described above, historical conversation information regarding the user input can be obtained from the external memory bank by calculating the dense vector similarity, and can be specifically realized by using a neural network. In step S2109, the parameters of the neural network for calculating the dense vector similarity are adjusted to increase the similarity between the fourth sample intermediate input determined based on the fourth sample initial input and the sample memory result. Thereby, the optimized neural network (external memory bank) can return the sample memory result for the fourth sample intermediate input. The parameter adjustment to the deep learning model in step S2112 can refer to step S1704 or step S1706 in FIG. 17, and it should be understood that it will not be described here.
[0145] According to another aspect of the present disclosure, a data generation device based on a deep learning model is provided. The deep learning model can generate response data based on input data of a user. As shown in FIG. 22, the data generation device 2200 includes a first determination unit 2210 configured to determine an initial input used for the deep learning model based on input data from the user, and obtain a first output of the deep learning model. Here, in response to determining that the deep learning model needs to call a first functional component different from the deep learning model to generate a response based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component determined based on the initial input. A first acquisition unit 2220 configured to obtain, a second acquisition unit 2230 configured to obtain a first intermediate result determined by the first functional component based on the first intermediate query, a second determination unit 2240 configured to determine a second input used for the deep learning model based on at least the initial input and the first intermediate result, and a third acquisition unit 2250 configured to obtain a second output of the deep learning model to generate a response to the initial input. It should be understood that the operations of units 2210-2250 in device 2200 are respectively similar to the operations of steps S201-S205 in FIG. 2 and will not be described here.
[0146] According to some embodiments, the first functional component may be an external memory bank capable of storing a first set of data groups related to the user. Each data group in the first set of data groups can include at least a historical input data item and a historical response item generated by the deep learning model for the historical input data item.
[0147] According to some embodiments, each data group in the first data group set may further include an entry time item corresponding to the historical input data item and the historical answer item in the set.
[0148] According to some embodiments, the first intermediate query can be based on the input data. The first intermediate result may be a historical answer item corresponding to a historical input data item in the first data group set whose similarity to the input data is higher than a first threshold.
[0149] According to some embodiments, the first intermediate query can be based on the input data. The first intermediate result may be a historical answer item corresponding to a historical input data item in the first data group set whose similarity to the input data is higher than a first threshold and whose timestamp is the latest.
[0150] According to some embodiments, in response to determining that the similarity between the first data group based on the input data and the answer and any data group in the first data group set is less than a second threshold, the data generation device may further include a first entry unit configured to enter the first data group into the first data group set.
[0151] According to some embodiments, in response to determining that the similarity between the first data group based on the input data and the answer and a second data group in the first data group set is higher than a third threshold and that the first data group and the second data group conflict with each other, the data generation device may enter the first data group into the first data group set and further include a second entry unit configured to delete the second data group from the first data group set.
[0152] According to some embodiments, the data generation device may further include a deletion unit configured to delete an old data group from the external memory bank based on the entry time item.
[0153] According to some embodiments, the first determination unit may include a first acquisition subunit configured to acquire a history response item corresponding to a history input data item having a similarity with the input data higher than a first threshold from the external memory bank based on the input data, and a first determination subunit configured to determine an initial input based on the input data and the history response item. A first data group set related to the user may be stored in the external memory bank. Each data group in the first data group set may include at least a history input data item and a history response item generated by a deep learning model for the history input data item.
[0154] According to some embodiments, the second determination unit may include a third determination subunit configured to determine a second input used for the deep learning model based on the initial input, the first intermediate result, and the first intermediate query.
[0155] According to some embodiments, the first functional component may be an external search engine. According to some embodiments, the first functional component may be a search model trained in association with the deep learning model.
[0156] According to some embodiments, the first functional component may be at least one application programming interface that can be called by the deep learning model.
[0157] According to some embodiments, the second output may include a second token for invoking a second functional component and a second intermediate query identifiable by the second functional component obtained based on the second input. The third acquisition unit is a third acquisition subunit configured to perform a corresponding function invocation operation on the second output, where the function invocation operation includes obtaining a second intermediate result determined by the second functional component based on the second intermediate query, determining a third input used in the deep learning model based on at least the second input and the second intermediate result, and obtaining a third output of the deep learning model, and in response to including, in the Nth output of the deep learning model, an Nth token for invoking an Nth functional component and an Nth intermediate query identifiable by the Nth functional component obtained based on the Nth input, performing a function invocation operation corresponding to the Nth output until it is determined that the corresponding token for invoking any functional component different from the deep learning model is not included in the (N + 1)th output, and using the (N + 1)th output as an answer to the initial input, where N is an integer greater than 2, and may include a call subunit configured as such.
[0158] According to some embodiments, the second functional component and the Nth functional component may each be one of a group of functional components including an external search engine, a search model trained in association with a deep learning model, at least one application programming interface callable by the deep learning model, and an external memory bank, where the external memory bank stores a first set of data groups related to the user, and here, each data group in the first set of data groups includes at least a historical input data item and a historical answer item generated by the deep learning model for the historical input data item.
[0159] According to some embodiments, the second output may not include a corresponding token for invoking any functional component different from the deep learning model. The third acquisition unit may include a response subunit configured to use the second output as a response to the initial input.
[0160] According to some embodiments, the initial input may include context information of the input data. According to some embodiments, the first determination unit may include a second acquisition subunit configured to acquire at least a pair of historical input data items and historical response items whose similarity between the input data and the context information from the external memory bank meets a fourth threshold value, and a second determination subunit configured to determine an initial input used for the deep learning model based on the input data, the context information, and at least the pair of historical input data items and historical response items. A first set of data groups related to the user can be stored in the external memory bank. Each data group in the first set of data groups may include at least a historical input data item and a historical response item generated by the deep learning model for the historical input data item.
[0161] According to another aspect of the present disclosure, a training apparatus for a deep learning model is provided. The deep learning model is used to generate response data based on user input data. As shown in FIG. 23, the training apparatus 2300 includes a fourth acquisition unit 2310 configured to acquire first sample data, where the first sample data includes a first sample initial input and a first sample output. Here, the first sample initial input includes an expression of intention to call a first preset functional component different from the deep learning model. Here, the first sample output is configured to include a first token for calling the first preset functional component and a first sample intermediate input identifiable by the first preset functional component; a fifth acquisition unit 2320 configured to acquire second sample data, where the second sample data includes a second sample initial input and a second sample output. Here, the second sample initial input does not include an expression of intention to call any preset functional component different from the deep learning model. Here, the second sample output is configured not to include a corresponding token for calling any preset functional component; a first processing unit 2330 configured to process the first sample initial input using the deep learning model to obtain a first predicted output; a first parameter adjustment unit 2340 configured to adjust the parameters of the deep learning model based on a comparison between the first sample output and the first predicted output; a second processing unit 2350 configured to process the second sample initial input using the deep learning model to obtain a second predicted output; and a second parameter adjustment unit 2360 configured to adjust the parameters of the deep learning model based on a comparison between the second sample output and the second predicted output. It should be understood that the operations of units 2310-2360 in apparatus 2300 are respectively similar to the operations of steps S1701-S1706 in FIG. 17 and will not be described here.
[0162] According to some embodiments, the training device obtains third sample data including a third sample initial input, a sample search query, a plurality of sample search results, and a third sample answer of the deep learning model for the third sample initial input. The sample search query is a sample intermediate input generated by the deep learning model based on the third sample initial input, and the sample intermediate input is distinguishable by a search model different from the deep learning model. Here, the plurality of sample search results are configured to be the results output by the search model based on the sample search query. A sorting unit configured to perform a sorting operation on the plurality of sample search results based on the degree of match between each of the plurality of sample search results and the third sample answer, and a training unit configured to train the search model based on the sorted plurality of sample search results.
[0163] According to some embodiments, the sorting unit includes a screening subunit configured to screen a first sample search result with the highest current degree of match from the plurality of sample search results, a deletion subunit configured to delete the overlapping content between the third sample answer and the first sample search result and update the third sample answer, and a sorting subunit configured to repeat the sorting operation on the remaining part until the sorting of all the sample search results in the plurality of sample search results is completed based on the degree of match between each of the remaining parts of the plurality of sample search results and the updated third sample answer.
[0164] According to some embodiments, the search model can include a sorting sub-model and a recall sub-model. The training unit can include a first training sub-unit configured to train the sorting sub-model of the search model based on a plurality of sorted sample search results, and a second training sub-unit configured to train the recall sub-model with the trained sorting sub-model as a teacher model.
[0165] According to some embodiments, the training device can further include a seventh acquisition unit configured to acquire fourth sample data, where the fourth sample data includes a fourth sample initial input, a fourth sample intermediate input identifiable by an external memory bank, a sample storage result, and a fourth sample answer, and the fourth sample intermediate input is configured to be determined based on the fourth sample initial input; an eighth acquisition unit configured to acquire a predicted storage result determined based on the fourth sample intermediate input by the external memory bank; a third parameter adjustment unit configured to adjust parameters of the external memory bank based on a comparison between the predicted storage result and the sample storage result; a third determination unit configured to determine a fourth sample target input used for a deep learning model based on at least the fourth sample initial input and the sample storage result; a third processing unit configured to process the fourth sample target input using the deep learning model to obtain a fourth predicted answer; and a fourth parameter adjustment unit configured to adjust parameters of the deep learning model based on a comparison between the fourth sample answer and the fourth predicted answer.
[0166] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of related user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0167] According to the embodiments of the present disclosure, an electronic device, a readable storage medium, and a computer program product are further provided. Referring to FIG. 24, here, a configuration block diagram of an electronic device 2400 that can be used as a server or a client of the present disclosure, which is an example of a hardware device applicable to various aspects of the present disclosure, will be described. The electronic device represents various forms of digital electronic computers, such as laptop computers, desktop computers, tablets, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown in this specification, their connection relationships, and their functions are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0168] As shown in FIG. 24, the electronic device 2400 includes a computing unit 2401, which can execute various appropriate operations and processes by a computer program stored in a read-only memory (ROM) 2402 or a computer program loaded from a storage unit 2408 into a random access memory (RAM) 2403. In the RAM 2403, various programs and data necessary for operating the electronic device 2400 may be further stored. The computing unit 2401, the ROM 2402, and the RAM 2403 are connected to each other via a bus 2404. An input / output (I / O) interface 2405 is also connected to the bus 2404.
[0169] A plurality of components in the electronic device 2400 are connected to the I / O interface 2405 and include an input unit 2406, an output unit 2407, a storage unit 2408, and a communication unit 2409. The input unit 2406 may be any type of device capable of inputting information into the electronic device 2400. The input unit 2406 can generate input numerical or character information and key signal inputs related to user settings and / or function controls of the electronic device, and may include, but is not limited to, a mouse, a keyboard, a touch screen, a track board, a track ball, an operation lever, a microphone, and / or a remote control. The output unit 2407 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 2408 may include, but is not limited to, a magnetic disk and an optical disk. The communication unit 2409 enables the electronic device 2400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0170] The computing unit 2401 may be various general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 2401 may include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2401 executes each method and process described above, such as a data generation method or a training method of a deep learning model. For example, in some embodiments, the data generation method or the training method of the deep learning model may be implemented as a computer software program tangibly included in a machine-readable medium, such as the storage unit 2408. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 2400 via the ROM 2402 and / or the communication unit 2409. When the computer program is loaded into the RAM 2403 and executed by the computing unit 2401, one or more steps of the data generation method or the training method of the deep learning model described above can be executed. Alternatively, in other embodiments, the computing unit 2401 may be configured to execute the data generation method or the training method of the deep learning model in any other suitable manner (e.g., by firmware).
[0171] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be implemented in one or more computer programs, and the one or more computer programs may be executed and / or interpreted in a programmable system including at least one programmable processor, and the programmable processor may be a dedicated or general-purpose programmable processor, receiving data and instructions from a memory system, at least one input device, and at least one output device, and transmitting the data and instructions to the memory system, the at least one input device, and the at least one output device, which may also be included.
[0172] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations defined in the flow chart and / or block diagram are implemented. The program code may be executed entirely by a machine, partially by a machine, partially by a machine as an independent software package and partially by a remote machine, or entirely by a remote machine or server.
[0173] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may comprise or store a program for use in or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections through one or more leads, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0174] To provide for interaction with a user, a computer may implement the systems and techniques described herein, the computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user may provide input to the computer. Other kinds of devices may be further provided for interacting with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including acoustic input, speech input, or tactile input).
[0175] The systems and techniques described herein may be implemented in a computing system that includes backend members (e.g., as a data server), a computing system that includes middleware members (e.g., an application server), a computing system that includes frontend members (e.g., a user computer having a graphical user interface or a web browser through which a user can realize interactions with embodiments of those systems and techniques), or a computing system consisting of any combination of those backend members, middleware members, or frontend members. The members of the system may be interconnected by digital data communication in any form or medium (e.g., a communication network). An example of a communication network includes a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0176] A computer system may include a client side and a server. The client side and the server are generally far apart from each other and usually interact via a communication network. The relationship between the client side and the server is generated by operating a computer program corresponding to a computer having a client - server relationship on each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0177] It should be understood that the steps may be reordered, increased, or deleted again using the various forms of flow described above. For example, each step described in the present disclosure may be executed in parallel, sequentially, or in a different order, and the text is not limited to this as long as the technical solutions disclosed in the present disclosure can achieve the desired results.
[0178] Embodiments or examples of the present disclosure have been described with reference to the drawings. However, the above methods, systems, and apparatuses are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but it should be understood that it is limited only by the scope of the claims after authorization and their equivalent scope. Various elements of the embodiments or examples may be omitted or replaced by their equivalent elements. Note that each step may be executed in an order different from the order described in the present disclosure. Furthermore, various elements of the embodiments or examples may be combined in various ways. What is important is that with the evolution of technology, many of the elements described here can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A data generation method based on a deep learning model, wherein the deep learning model can generate response data based on user input data, and the data generation method includes: determining an initial input used for the deep learning model based on input data from a user; obtaining a first output of the deep learning model, where, in response to determining that the deep learning model needs to call a first functional component different from the deep learning model to generate a response based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component and determined based on the initial input; obtaining a first intermediate result determined by the first functional component based on the first intermediate query; determining a second input used for the deep learning model based on at least the initial input and the first intermediate result; obtaining a second output of the deep learning model to generate the response to the initial input. A data generation method based on a deep learning model, characterized by including the above.
2. The first functional component is an external memory bank that stores a first data group set related to the user, where each data group in the first data group set includes at least a historical input data item and a historical response item generated by the deep learning model for the historical input data item. The data generation method according to Claim 1, characterized by this.
3. The first intermediate query is based on the input data, where the first intermediate result is a historical response item corresponding to a historical input data item in the first data group set that has a similarity with the input data higher than a first threshold. The data generation method according to Claim 2, characterized by this.
4. The data generation method includes: In response to determining that the similarity between the first data group based on the input data and the response and any data group in the first data group set is less than a second threshold value, further including entering the first data group into the first data group set, the data generation method according to claim 2, characterized in that.
5. The data generation method includes: In response to determining that the similarity between the first data group based on the input data and the response and a second data group in the first data group set is higher than a third threshold value and the first data group and the second data group conflict with each other, entering the first data group into the first data group set and further including deleting the second data group from the first data group set, the data generation method according to claim 2, characterized in that.
6. Each data group in the first data group set further includes an entry time item corresponding to the historical input data item and the historical response item in that data group, the data generation method according to claim 2, characterized in that.
7. The first intermediate query is based on the input data, where the first intermediate result is a historical response item corresponding to the historical input data item in the first data group set that has a similarity higher than a first threshold value with the input data and has the latest timestamp, the data generation method according to claim 6, characterized in that.
8. The data generation method includes: Further including deleting data groups with old timeliness from the external memory bank based on the entry time item, the data generation method according to claim 6, characterized in that.
9. Determining the initial input used in the deep learning model includes: Obtaining, based on the input data, a historical response item corresponding to a historical input data item whose similarity with the input data is higher than a first threshold value from an external memory bank; and Determining the initial input based on the input data and the historical response item, where: The external memory bank stores a first set of data groups related to the user, where each data group in the first set of data groups includes at least a history input data item and a history answer item generated by the deep learning model for the history input data item. The data generation method according to claim 1 is characterized in that it includes the above.
10. The data generation method according to claim 1 is characterized in that the initial input includes context information of the input data.
11. Determining the initial input used in the deep learning model includes: Obtaining at least a pair of history input data items and history answer items whose similarity between the input data and the context information from an external memory bank meets a fourth threshold; Determining the initial input used in the deep learning model based on the input data, the context information, and the at least a pair of history input data items and history answer items. Here, The external memory bank stores a first set of data groups related to the user, where each data group in the first set of data groups includes at least a history input data item and a history answer item generated by the deep learning model for the history input data item. The data generation method according to claim 10 is characterized in that it includes the above.
12. The data generation method according to any one of claims 9 to 11 is characterized in that the first functional component is an external search engine.
13. The data generation method according to any one of claims 9 to 11 is characterized in that the first functional component is a search model trained in association with the deep learning model.
14. The data generation method according to any one of claims 9 to 11 is characterized in that the first functional component is at least one application programming interface that can be called by the deep learning model.
15. Determining the second input used in the deep learning model based on at least the initial input and the first intermediate result includes: Determining a second input used in the deep learning model based on the initial input, the first intermediate result, and the first intermediate query, characterized in that it comprises the method for generating data according to any one of claims 1 to 11.
16. The second output does not include a corresponding token for invoking any functional component different from the deep learning model, where Obtaining the second output of the deep learning model for generating the answer to the initial input Characterized in that it comprises using the second output as the answer to the initial input, the method for generating data according to any one of claims 1 to 11.
17. The second output includes a second token for invoking a second functional component and a second intermediate query identifiable by the second functional component obtained based on the second input, where Obtaining the second output of the deep learning model for generating the answer to the initial input Performing a corresponding function call operation on the second output, where the function call operation Obtaining a second intermediate result determined by the second functional component based on the second intermediate query Determining a third input used in the deep learning model based on at least the second input and the second intermediate result And obtaining a third output of the deep learning model In response to including a token for invoking any functional component different from the deep learning model in the N+1th output until it is determined that the Nth output does not include a corresponding token for invoking any functional component different from the deep learning model, performing a function call operation corresponding to the Nth output, and using the N+1th output as the answer to the initial input, where N is an integer greater than 2, characterized in that it comprises the method for generating data according to any one of claims 1 to 11.
18. The second functional component and the Nth functional component are respectively An external search engine A search model trained in association with the deep learning model At least one application programming interface that can be called by the deep learning model, One of the functional component groups including an external memory bank, wherein a first data group set related to the user is stored in the external memory bank, and here, each data group in the first data group set includes at least historical input data items and historical answer items generated by the deep learning model for the historical input data items. The data generation method according to claim 17, characterized in that it comprises.
19. A method for training a deep learning model, wherein the deep learning model is used to generate answer data based on user input data, and the training method comprises Obtaining first sample data, the first sample data including a first sample initial input and a first sample output, where the first sample initial input includes an intention expression for calling a first preset functional component different from the deep learning model, and the first sample output includes a first token for calling the first preset functional component and a first sample intermediate input identifiable by the first preset functional component; Obtaining second sample data, the second sample data including a second sample initial input and a second sample output, where the second sample initial input does not include an intention expression for calling any preset functional component different from the deep learning model, and the second sample output does not include a corresponding token for calling any preset functional component; Processing the first sample initial input using the deep learning model to obtain a first predicted output; Adjusting the parameters of the deep learning model based on a comparison between the first sample output and the first predicted output; Processing the second sample initial input using the deep learning model to obtain a second predicted output; Adjusting the parameters of the deep learning model based on a comparison between the second sample output and the second predicted output. A method for training a deep learning model, characterized by comprising.
20. The training method comprises Obtain third sample data including a third sample initial input, a sample search query, a plurality of sample search results, and a third sample answer of the deep learning model for the third sample initial input, where the sample search query is a sample intermediate input generated by the deep learning model based on the third sample initial input, the sample intermediate input is identifiable by a search model different from the deep learning model, and here, the plurality of sample search results are results output by the search model based on the sample search query, Perform a sorting operation on the plurality of sample search results based on the degree of match between each of the plurality of sample search results and the third sample answer, The training method according to claim 19, further comprising training the search model based on the sorted plurality of sample search results.
21. The performing a sorting operation on the plurality of sample search results based on the degree of match between each of the plurality of sample search results and the third sample answer includes: Screening a first sample search result with the highest current degree of match from the plurality of sample search results, Deleting the overlapping content between the third sample answer and the first sample search result and updating the third sample answer, Repeating the sorting operation on the remaining part until the sorting of all sample search results in the plurality of sample search results is completed, based on the degree of match between each of the remaining part of the plurality of sample search results and the updated third sample answer. The training method according to claim 20 is characterized by including this.
22. The search model includes a sorting sub-model and a recall sub-model. The training the search model based on the sorted plurality of sample search results includes: Training the sorting sub-model of the search model based on the sorted plurality of sample search results, The training method according to claim 20 or 21, characterized by including training the recall sub-model using the trained sorting sub-model as a teacher model.
23. The training method includes: obtaining fourth sample data, where the fourth sample data includes a fourth sample initial input, a fourth sample intermediate input identifiable by an external memory bank, a sample storage result, and a fourth sample answer, and the fourth sample intermediate input is determined based on the fourth sample initial input; obtaining a predicted storage result determined based on the fourth sample intermediate input by the external memory bank; adjusting parameters of the external memory bank based on a comparison between the predicted storage result and the sample storage result; determining a fourth sample target input used for the deep learning model based on at least the fourth sample initial input and the sample storage result; processing the fourth sample target input using the deep learning model to obtain a fourth predicted answer; further including adjusting parameters of the deep learning model based on a comparison between the fourth sample answer and the fourth predicted answer, the training method according to any one of claims 19 to 21.
24. A data generation device based on a deep learning model, where the deep learning model can generate answer data based on input data of a user, and the data generation device includes: a first determination unit configured to determine an initial input used for the deep learning model based on input data from a user; a first acquisition unit configured to acquire a first output of the deep learning model, where in response to determining that a first functional component different from the deep learning model needs to be called for the deep learning model to generate an answer based on the initial input, the first output includes a first token for calling the first functional component and a first intermediate query identifiable by the first functional component and determined based on the initial input; a second acquisition unit configured to acquire a first intermediate result determined by the first functional component based on the first intermediate query; A second determination unit configured to determine a second input used in the deep learning model based on at least the initial input and the first intermediate result; A data generation device based on a deep learning model, comprising: a third acquisition unit configured to acquire a second output of the deep learning model to generate the answer for the initial input. **Claim 25** The first functional component is an external memory bank that stores a first set of data groups related to the user, wherein each data group in the first set of data groups includes at least a historical input data item and a historical answer item generated by the deep learning model for the historical input data item. The data generation device according to claim 24. **Claim 26** The first intermediate query is based on the input data, wherein the first intermediate result is a historical answer item corresponding to a historical input data item in the first set of data groups, the similarity of which to the input data is higher than a first threshold. The data generation device according to claim 25. **Claim 27** The data generation device Further comprising a first entry unit configured to enter the first data group into the first set of data groups in response to determining that the similarity between the first data group based on the input data and the answer and any data group in the first set of data groups is less than a second threshold. The data generation device according to claim 25. **Claim 28** The data generation device Further comprising a second entry unit configured to enter the first data group into the first set of data groups and delete the second data group from the first set of data groups in response to determining that the similarity between the first data group based on the input data and the answer and a second data group in the first set of data groups is higher than a third threshold and the first data group and the second data group conflict with each other. The data generation device according to claim 25. **Claim 29** The data generation device according to claim 25, wherein each data group in the first data group set further includes an entry time item corresponding to a history input data item and a history answer item in the data group.
30. The first intermediate query is based on the input data, and here, the first intermediate result is a history answer item corresponding to a history input data item in the first data group set, which has a similarity with the input data higher than a first threshold value and the latest time stamp. The data generation device according to claim 29.
31. The data generation device The data generation device according to claim 29, further comprising a deletion unit configured to delete data groups with old timeliness from the external memory bank based on the entry time item.
32. The first determination unit A first acquisition subunit configured to acquire a history answer item corresponding to a history input data item having a similarity with the input data higher than a first threshold value from an external memory bank based on the input data; A first determination subunit configured to determine the initial input based on the input data and the history answer item, where A first data group set related to the user is stored in the external memory bank, and here, each data group in the first data group set includes at least a history input data item and a history answer item generated by the deep learning model for the history input data item. The data generation device according to claim 24.
33. The data generation device according to claim 24, wherein the initial input includes context information of the input data.
34. The first determination unit A second acquisition subunit configured to acquire at least a pair of a history input data item and a history answer item whose similarity between the input data and the context information conforms to a fourth threshold value from an external memory bank A second determination subunit configured to determine the initial input used in the deep learning model based on the input data, the context information, and the at least one pair of historical input data items and historical answer items, where The external memory bank stores a first set of data groups related to the user, where each data group in the first set of data groups includes at least a historical input data item and a historical answer item generated by the deep learning model for the historical input data item. The data generation device according to claim 33. **Claim 35** The data generation device according to any one of claims 32 to 34, wherein the first functional component is an external search engine. **Claim 36** The data generation device according to any one of claims 32 to 34, wherein the first functional component is a search model trained in association with the deep learning model. **Claim 37** The data generation device according to any one of claims 32 to 34, wherein the first functional component is at least one application programming interface that can be called by the deep learning model. **Claim 38** The second determination unit The data generation device according to any one of claims 24 to 34, further comprising a third determination subunit configured to determine a second input used in the deep learning model based on the initial input, the first intermediate result, and the first intermediate query. **Claim 39** The second output does not include a corresponding token for calling any functional component different from the deep learning model, where The third acquisition unit The data generation device according to any one of claims 24 to 34, further comprising an answer subunit configured to use the second output as the answer to the initial input. **Claim 40** The second output includes a second token for calling a second functional component and a second intermediate query identifiable by the second functional component obtained based on the second input, where The third acquisition unit including a third acquisition subunit configured to execute a corresponding function call operation for the second output, the function call operation being obtaining a second intermediate result determined by the second functional component based on the second intermediate query; determining a third input used for the deep learning model based on at least the second input and the second intermediate result; obtaining a third output of the deep learning model; and In response to including in the Nth output of the deep learning model an Nth intermediate query identifiable by the Nth functional component, obtained based on an Nth token and an Nth input for calling the Nth functional component, until it is determined that the corresponding token for calling any functional component different from the deep learning model is not included in the (N + 1)th output, executing a function call operation corresponding to the Nth output, and using the (N + 1)th output as the answer to the initial input, where N is an integer greater than 2, the calling subunit is configured as described above. The data generation device according to any one of claims 24 to 34, characterized by including
41. The second functional component and the Nth functional component are each an external search engine; a search model trained in association with the deep learning model; at least one application programming interface that can be called by the deep learning model; one of a group of functional components including an external memory bank, and a first data group set related to the user is stored in the external memory bank. Here, each data group in the first data group set includes at least a history input data item and a history answer item generated by the deep learning model for the history input data item. The data generation device according to claim 40, characterized by
42. A training device for a deep learning model, wherein the deep learning model is used to generate answer data based on user input data, and the training device is Obtain first sample data, where the first sample data includes a first sample initial input and a first sample output. Here, the first sample initial input includes an intention expression for invoking a first preset functional component different from the deep learning model, and the first sample output is configured to include a first token for invoking the first preset functional component and a first sample intermediate input identifiable by the first preset functional component. A fourth acquisition unit; Obtain second sample data, where the second sample data includes a second sample initial input and a second sample output. Here, the second sample initial input does not include an intention expression for invoking any preset functional component different from the deep learning model, and the second sample output is configured not to include a corresponding token for invoking any preset functional component. A fifth acquisition unit; A first processing unit configured to process the first sample initial input using the deep learning model to obtain a first predicted output; A first parameter adjustment unit configured to adjust the parameters of the deep learning model based on a comparison between the first sample output and the first predicted output; A second processing unit configured to process the second sample initial input using the deep learning model to obtain a second predicted output; A training apparatus for a deep learning model, comprising a second parameter adjustment unit configured to adjust the parameters of the deep learning model based on a comparison between the second sample output and the second predicted output.
43. The training apparatus is Obtain third sample data including a third sample initial input, a sample search query, a plurality of sample search results, and a third sample answer of the deep learning model for the third sample initial input, where the sample search query is a sample intermediate input generated by the deep learning model based on the third sample initial input, the sample intermediate input is identifiable by a search model different from the deep learning model, and here, the plurality of sample search results are configured to be the results output by the search model based on the sample search query, a sixth acquisition unit; A sorting unit configured to perform a sorting operation on the plurality of sample search results based on the degree of match between each of the plurality of sample search results and the third sample answer; The training device according to claim 42, further comprising a training unit configured to train the search model based on the sorted plurality of sample search results.
44. The sorting unit A screening subunit configured to screen the first sample search result with the highest current degree of match from the plurality of sample search results; A deletion subunit configured to delete the overlapping content between the third sample answer and the first sample search result and update the third sample answer; The training device according to claim 43, further comprising a sorting subunit configured to repeat the sorting operation on the remaining part until the sorting of all the sample search results in the plurality of sample search results is completed, based on the degree of match between each of the remaining parts of the plurality of sample search results and the updated third sample answer.
45. The search model includes a sorting sub-model and a recall sub-model, and here, the training unit A first training subunit configured to train the sorting sub-model of the search model based on the sorted plurality of sample search results; A second training subunit configured to train the recall submodel using the trained sorting submodel as a teacher model, the training apparatus according to claim 43 or 44.
46. The training apparatus includes a seventh acquisition unit configured to acquire fourth sample data, the fourth sample data including a fourth sample initial input, a fourth sample intermediate input identifiable by an external memory bank, a sample storage result, and a fourth sample answer, the fourth sample intermediate input being configured to be determined based on the fourth sample initial input; an eighth acquisition unit configured to acquire a predicted storage result determined based on the fourth sample intermediate input by an external memory bank; a third parameter adjustment unit configured to adjust parameters of the external memory bank based on a comparison between the predicted storage result and the sample storage result; a third determination unit configured to determine a fourth sample target input used for the deep learning model based on at least the fourth sample initial input and the sample storage result; a third processing unit configured to process the fourth sample target input using the deep learning model to obtain a fourth predicted answer; and a fourth parameter adjustment unit configured to adjust parameters of the deep learning model based on a comparison between the fourth sample answer and the fourth predicted answer, the training apparatus according to any one of claims 42 to 44.
47. An electronic device, the electronic device including at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by at least one processor, the instructions being executed by the at least one processor such that the at least one processor can execute the method according to any one of claims 1 to 11, the electronic device.
48. A non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause a computer to execute the method according to any one of claims 1 to 11, the computer-readable storage medium. Claim 49 A computer program, which when executed by a processor, is used to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Information presentation device and information presentation method
JP2021077268A
Techniques for Interaction Processing Using Context Data
JP2022547598A
Data-Informed Decision Making Through a Domain-General Artificial Intelligence Platform
US20220343903A1