Answer prediction method, device and equipment based on distributed learning
By determining the main prediction node and the auxiliary prediction node in the distributed learning system, generating the main prompt words and the auxiliary prompt words, and combining the target large language model to predict answers, the problem of inaccurate answer prediction in distributed learning is solved, and the accuracy of answer prediction is improved.
Patent Information
- Application Number
- CN202510667786.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing distributed learning technology has the problem of inaccurate answer prediction.
By determining the primary prediction node and the secondary prediction node between the first node and the second node, the primary prediction node generates the primary prompt word, the secondary prediction node generates the secondary prompt word, and combining the pre-trained target large language model for answer prediction.
It improves the accuracy of answer prediction, makes the generation of words more comprehensive and accurate, and provides a more accurate basis for subsequent answer prediction.
Smart Images

Figure CN120196729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and particularly to an answer prediction method, device, and equipment based on distributed learning. Background Art
[0002] Distributed Machine Learning (abbreviated as DML in English), also known as distributed learning, refers to algorithms and systems that use multiple computing nodes for machine learning or deep learning, aiming to improve performance, protect privacy, and scale to larger-scale training data and larger models. However, existing distributed learning has the problem of inaccurate answer prediction. Summary of the Invention
[0003] In view of this, the present invention provides an answer prediction method, device, and equipment based on distributed learning, mainly aiming to solve the problem of inaccurate answer prediction currently existing.
[0004] To solve the above problems, the present application provides an answer prediction method based on distributed learning, which is applied to the first node and includes: Receiving second data directories sent by each second node, and matching the target question with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; When the first node is not determined as the main prediction node or the auxiliary prediction node, sending the target question to the main prediction node and each auxiliary prediction node; When the first node is determined as the auxiliary prediction node, constructing an auxiliary prompt word based on the target question, sending the auxiliary prompt word to the main prediction node, and sending the target question to the main prediction node and the remaining auxiliary prediction nodes; When the first node is determined as the main prediction node, constructing a main prompt word based on the target question, sending the target question to each auxiliary prediction node, and receiving the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predicting and outputting the target answer corresponding to the target question by using the pre-trained target large language model based on the main prompt word and each auxiliary prompt word.
[0005] To solve the above problems, the present application provides an answer prediction method based on distributed learning, which is applied to the second node and includes: Sending the local second data directory to the first node, for the first node to perform matching based on the first data directory and the second data directories of each second node, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; When the second node is determined as the secondary prediction node, receive the target question sent by the first node, construct a secondary prompt word based on the target question, and send the secondary prompt word to the primary prediction node; When the second node is determined as the primary prediction node, receive the target question sent by the first node, construct a primary prompt word based on the target question, and receive the secondary prompt words constructed based on the target question sent by each secondary prediction node, and predict and output the target answer corresponding to the target question by using the pre-trained target large language model based on the primary prompt word and each secondary prompt word.
[0006] To solve the above problems, the present application provides an answer prediction device based on distributed learning, including: A first receiving module, configured to receive the second data directories sent by each second node, and match the target question with the first data directory corresponding to the first node and each second data directory, so as to determine the primary prediction node and each secondary prediction node from the first node and each second node; A first sending module, configured to send the target question to the primary prediction node and each secondary prediction node when the first node is not determined as the primary prediction node or the secondary prediction node; A first constructing module, configured to construct a secondary prompt word based on the target question when the first node is determined as the secondary prediction node, send the secondary prompt word to the primary prediction node, and send the target question to the primary prediction node and the remaining secondary prediction nodes; A first predicting module, configured to construct a primary prompt word based on the target question when the first node is determined as the primary prediction node, send the target question to each secondary prediction node, and receive the secondary prompt words constructed based on the target question sent by each secondary prediction node, and predict and output the target answer corresponding to the target question by using the pre-trained target large language model based on the primary prompt word and each secondary prompt word.
[0007] To solve the above problems, the present application provides an answer prediction device based on distributed learning, including: A second sending module, configured to send the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directories of each second node, and determines the primary prediction node and each secondary prediction node from the first node and each second node; A second constructing module, configured to receive the target question sent by the first node when the second node is determined as the secondary prediction node, construct a secondary prompt word based on the target question, and send the secondary prompt word to the primary prediction node; The second prediction module is used to receive the target question sent by the first node when the second node is determined as the main prediction node, construct the main prompt word based on the target question, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predict the target answer corresponding to the target question and output it by using the target large language model obtained by pre-training based on the main prompt word and each auxiliary prompt word.
[0008] To solve the above problems, the present application provides an electronic device, which at least includes a memory and a processor. A computer program is stored on the memory, and the processor implements the steps of the answer prediction method based on distributed learning described in any one of the above when executing the computer program on the memory.
[0009] In the answer prediction method, device and equipment based on distributed learning in the present application, by determining the main prediction node for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, the main prompt word can be generated by using the main prediction node and the auxiliary prompt words can be generated by using the auxiliary prediction nodes subsequently, making the generation of the prompt words more comprehensive and accurate, laying a foundation for accurately predicting the answer based on the main prompt word and each auxiliary prompt word by using the target large language model, and being beneficial to improving the accuracy of answer prediction.
[0010] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Description of the Drawings
[0011] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a flowchart of an answer prediction method based on distributed learning according to an embodiment of the present application; Figure 2 It is a flowchart of an answer prediction method based on distributed learning according to another embodiment of the present application; Figure 3 It is a structural block diagram of an answer prediction device based on distributed learning according to another embodiment of the present application; Figure 4 It is a structural block diagram of an answer prediction device based on distributed learning according to another embodiment of the present application; Figure 5 It is a structural block diagram of an electronic device according to another embodiment of the present application. Detailed implementation manners
[0012] Reference is made herein to the accompanying drawings to describe various solutions and features of the present application.
[0013] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.
[0014] The accompanying drawings included in the specification and forming a part of the specification illustrate embodiments of the present application, and together with the general description of the present application given above and the detailed description of the embodiments given below are used to explain the principles of the present application.
[0015] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting examples with reference to the accompanying drawings.
[0016] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.
[0017] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.
[0018] Hereinafter, specific embodiments of the present application will be described with reference to the accompanying drawings; however, it should be understood that the embodiments applied are only examples of the present application, and it can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details applied herein are not intended to be limiting, but only as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0019] This specification may use the phrase "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may all refer to one or more of the same or different embodiments according to the present application.
[0020] The embodiments of the present application provide a method for answer prediction based on distributed learning, which can be specifically applied to the first node / master computing node / leading party of federated learning, such as Figure 1 As shown, the method in this embodiment includes the following steps: Step S101: Receive the second data directories sent by each second node, and match the target problem with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; In this step, the first node will also build a first knowledge base based on the predetermined local service data, and then create a first data directory corresponding to the first knowledge base. Similarly, each second node can pre-build a second knowledge base based on the predetermined service data, create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node. Thus, the first node will determine the nodes that match the target problem from each second node and the first node participating in the federated learning as the main prediction node and each auxiliary prediction node according to the directories of the local data of each node (the first data directory and each second data directory).
[0021] In the specific implementation process, the main prediction node and the auxiliary prediction nodes can be determined from each node according to the matching degree between the data directory in each node and the target problem. That is to say, when the matching degree between the data directory and the target problem is greater than the predetermined threshold, this node can be used as a prediction node. When multiple prediction nodes are determined, the node with the highest matching degree can be determined as the main prediction node, and the remaining prediction nodes are used as auxiliary prediction nodes.
[0022] Step S102: When the first node is not determined as the main prediction node or the auxiliary prediction node, send the target problem to the main prediction node and each auxiliary prediction node; In the specific implementation process of this step, if the first node determines that the local first data directory does not match the target problem, it can be determined that the local area is not suitable as a prediction node (that is, not suitable as the main prediction node or the auxiliary prediction node). Thus, the first node can send the target problem to the corresponding main prediction node and each auxiliary prediction node, so that each auxiliary prediction node can generate / build auxiliary prompt words based on the target problem and the local database of the auxiliary prediction node, and each main prediction node can generate / build main prompt words based on the target problem and the local database of the main prediction node, and enable the main prediction node to predict the target answer corresponding to the target problem based on the main prompt and the auxiliary prompt words sent by each auxiliary prediction node and output it.
[0023] Step S103: When the first node is determined as the auxiliary prediction node, build an auxiliary prompt word based on the target problem, send the auxiliary prompt word to the main prediction node, and send the target problem to the main prediction node and the remaining auxiliary prediction nodes; In this step, when the first node is determined as the secondary prediction node, the first node can directly generate / build a secondary prompt word based on the target question according to the local database of the first node. At the same time, the first node can send the secondary prompt word to the primary prediction node, and send the target question to the primary prediction node and each remaining secondary prediction node, so that each remaining secondary prediction node can build a secondary prompt word based on the local database of the secondary prediction node, and the primary prediction node can build a primary prompt word based on the local database of the primary prediction node, so that the primary prediction node can predict and output the target answer corresponding to the target question based on the primary prompt word and each received secondary prompt word by using the pre-trained target large language model.
[0024] Step S104, when the first node is determined as the primary prediction node, build a primary prompt word based on the target question, send the target question to each secondary prediction node, and receive the secondary prompt words sent by each secondary prediction node and built based on the target question. Predict and output the target answer corresponding to the target question based on the primary prompt word and each secondary prompt word by using the pre-trained target large language model.
[0025] In this step, when the first node is determined as the primary prediction node, that is, the answer prediction can be performed locally at the first node. Therefore, the first node can generate / build a primary prompt word based on the target question and the local database. At the same time, send the target question to each secondary prediction node, and then receive the secondary prompt words generated / built by the secondary prediction node for the target question. Furthermore, the first node can predict and output the target answer corresponding to the target question based on the primary prompt word and each secondary prompt word by using the pre-trained target large language model.
[0026] In the answer prediction method based on distributed learning in this embodiment, by determining the primary prediction node for answer prediction and the secondary prediction nodes for auxiliary prediction from each second node and the first node, the primary prompt word can be generated by the primary prediction node and the secondary prompt words can be generated by the secondary prediction nodes subsequently, making the generation of the prompt words more comprehensive and accurate, laying a foundation for accurately predicting the answer based on the primary prompt word and each secondary prompt word by using the target large language model, which is beneficial to improving the accuracy of answer prediction.
[0027] Based on the above embodiments, another embodiment of the present application provides an answer prediction method based on distributed learning, which is applied to the first node. The overall process of answer prediction is as follows: Step S201, installation and initialization of the distributed large language model.
[0028] In this step, the leading party of the joint modeling (financial infrastructure, the first node) deploys the components of the multi-center distributed large language model collaboration network in its local data center, and sends the initial large language model and the corresponding components to each second node (slave computing node). Thus, the participating parties of the joint modeling (financial institutions, the second nodes) will also deploy the components of the multi-center distributed large language model collaboration network in their local data centers. The framework of distributed machine learning described in this embodiment can be implemented using other software products with the same functions such as Apache Spark MLlib, PyTorchtorch.distributed, and TensorFlowtf.distribute.
[0029] Step S202, generation of the total data directory of the collaboration network; In this step, the first node can construct a first knowledge base based on the predetermined business data and create a first data directory corresponding to the first knowledge base; at the same time, the first node will receive the second data directories sent by each second node; the second data directory corresponds to the second knowledge base constructed locally by the second node. Thus, the first node can subsequently directly match the target problem based on the first data directory and each second data directory, or generate a total data directory for matching with the target problem based on the first data directory and each second data directory.
[0030] Specifically, the first node and each second node are connected to their local bond business databases, and the local business data is segmented into documents and vectorized using natural language processing technology as the knowledge base. The knowledge base of node i is marked as K nodei , and the file names of the internal data files used by K nodei are sorted into the node data directory D nodei (including the first data directory and each second data directory). The second node sends the second data directory to the first node for synchronization. The first node collects the data directories fed back by all nodes and generates the total data directory D fl ={D node1 ,D node2 ,……}. That is, after receiving each second data directory, the first node can construct the total data directory based on each second data directory and the first data directory, and then match the target problem with the total data directory, so as to determine the first data directory and / or several second data directories that match the target problem from the total data directory.
[0031] Step S203, construction of the prompt engineering of the multi-center distributed large language model to construct the original first sample prompt set; In this step, the first node can pre-construct the original first sample prompt set. The specific process is as follows: Based on each prompting method, several original first sample prompt templates corresponding to each prompting method are generated respectively; where the prompting methods include any one or more of the following: zero-shot prompting method, few-shot prompting method, knowledge base prompting method, and chain of thought prompting method; Based on each original first sample prompt template, the original first sample prompt set is constructed.
[0032] That is, the first node can design the prompting engineering related to its own business. For the needs of different tasks, the following several prompting methods are used to complete the sample preparation of several types of prompting engineering, so as to construct and obtain the original first sample prompt set. 1. Zero-shot prompting method: The institution does not provide demonstrations related to the task results, and directly prompts the language model to give answers related to the task. For example, the input is [Question] What materials are needed to open a settlement account in the inter-bank bond market. 2. Few-shot prompting method: The institution provides a small number of prompting examples, such as task descriptions, etc. For example, the input is [Question] Translate Chinese into English. Bond -> bond; Valuation -> assessment; Management -> management; The European Central Bank has kept its three main interest rates unchanged for the fourth consecutive time, and traders have increased their expectations for the European Central Bank to cut interest rates, expecting a 100-basis-point cut this year. 3. Knowledge base prompting method: The institution provides a knowledge base related to the question and its keywords, and uses the context content of the knowledge base to form a prompt word. For example, the input is [Question]: What materials are needed to open a settlement account in the inter-bank bond market; [Knowledge Base]: Documents such as the "Notice on Matters Concerning Counter Business in the Inter-bank Bond Market". After the computing node performs text extraction and semantic understanding on the knowledge base, it calculates the relevance between the vectorized knowledge base document fragments and the input using a vector distance measurement index, and forms a prompt word jointly with the most similar several document fragments and the input question. The vector distance measurement index described in this application can be implemented using cosine similarity, dot product, Hamming distance, and other vector distance measurement methods. 4. Chain of thought prompting method: For relatively complex and logic-intensive questions, the institution introduces a chain of thought step in the prompting engineering to enhance the large model's processing ability for complex questions, and uses the following steps: 4.1. Construct examples: Manually write a small number of examples of tasks to be solved, which include the input question and its answer, and at the same time provide the chain of thought process in the solution process. 4.2. Provide chain of thought examples: Add chain of thought examples corresponding to the answers to the examples, so that the model can learn how to first analyze the question, then gradually perform chain of thought, and finally obtain the answer. 4.3. Prompt the model for chain-of-thought: Provide chain-of-thought answers in the examples to prompt the model to solve problems in a similar way, gradually presenting the analysis ideas of the problems, and finally giving the answers. 4.4. Evaluate the effect of chain-of-thought: Compare the effect of directly giving answers with the effect after adding chain-of-thought prompts on the test set, and prove that the former can only answer simple questions, while the latter can gradually analyze more complex questions through chain-of-thought, thus obtaining better results.
[0033] In this step, by adopting the above four prompting methods, the original first-sample prompt templates corresponding to each prompting method can be constructed, and thus the original first-sample prompt set can be obtained based on each original first-sample prompt template.
[0034] Step S204, based on the original first-sample prompt set, construct the target first-sample prompt set to complete the construction of the automatic prompting project; In the specific implementation process of this step, based on the original first-sample prompt set on the first node, the target first-sample prompt set can be generated using the target model. The target model can be a pre-trained distributed large language model. That is, the original first-sample prompt set can be used as the input, and the distributed large language model is used to automatically generate prompt instructions, and the best candidate prompt instructions are selected through filtering and searching methods to obtain the target first-sample prompt set PE nodei , the steps are as follows: 1. Based on the instruction examples in the original first-sample prompt set, let the distributed large language model generate automatic instructions. For example, for the following prompt template, the input and output examples are given to guide the model to generate instructions: [Question] Generate input-output instruction pairs according to the instruction examples. The instruction is... 2. Given a training set Dtrain, set a scoring function, and select the instruction in the automatically generated candidate instructions that can make the samples in the training set have the highest score. In this embodiment, the scoring function can be implemented using accuracy, language model likelihood value, and other machine learning evaluation metrics. 3. Use the Monte Carlo method for iterative search: Prompt the model to generate variants of the second-step instructions to further search for better instructions. For example, this prompt is used to generate approximate instructions: [Question] Generate a variant of the following instruction while keeping the semantic meaning unchanged. Input:... Output:...
[0035] Step S205, receive the target second-sample prompt sets sent by each second node to synchronize the prompting project.
[0036] In the specific implementation process of this step, the first node can receive the target second-sample prompt sets PE sent by each second node nodei, and based on the target second sample prompt set of each second node and the target first sample prompt set of the first node, determine the prompt subset PE for model training fl 'Prompt set directory D_PE fl '; Then, based on the prompt set directory D_PE fl 'From the target first sample prompt set PE nodei Determine the target first sample prompt subset PE nodei ', and set the prompt set directory D_PE fl 'Sent to the second node for the second node to use based on the prompt set directory D_PE fl 'From the target second sample prompt set PE nodei Determine the target second sample prompt subset PE nodei '.
[0037] Specifically, the second node i forms a prompt engineering instruction set / target second sample prompt set P Enodei After that, the target second sample prompt set can be sent to the first node, so that the first node can receive the target second sample prompt sets sent by each second node, and then generate a total sample prompt set PE based on each target second sample prompt set and the local target first sample prompt set of the first node fl ={PE node1 ,PE node2 ,……}. Then from the total sample prompt set PE fl Select the prompt subset PE used in this training fl ', prompt subset PE fl 'The corresponding prompt set directory is D_PE fl '. Finally, the first node will prompt the set directory as D_PE fl 'And the pre-trained model used in this training and other information are sent to each second node synchronously, so that each second node can use the prompt set directory D_PE fl 'From the target second sample prompt set PE node_i Determine the target second sample prompt subset PE nodei '.
[0038] In this embodiment, all nodes involved in this task will be based on the prompt set directory D_PE fl ', save the local instruction set PE nodei Selected PE nodei 'As training data, training is carried out in batches. The advantage of this process is that the training data from the computing nodes meets the screening criteria of the dominant party and has high security. Since the first node feedback process transmits the instruction set directory, the communication pressure is small and the possibility of attack and theft is low; the distributed computing method flexibly uses the computing resources of each node, and the system efficiency is high.
[0039] Step S206: Perform model training on the initial large language model. In this step, the first node can perform model training on the initial large language model based on the target first sample prompt subset PE nodeii ’ to obtain the initial first training parameters. In this step, for the initial large language model / pre-trained model of the large language model in this embodiment, other Chinese base models with the same function such as Chinese-LLaMA, OpenChineseLLaMA, and Ziya-LLaMA can be used, and Chinese LLaMA is used by default.
[0040] Step S207: Receive the initial second training parameters sent by each second node, and perform parameter aggregation based on each initial second training parameter and the initial first training parameter to obtain the initial aggregation parameters. In this step, all nodes involved in this task complete a local training respectively and send the training results to the first node, and the first node reads all the parameters and performs parameter aggregation. In this embodiment, the parameter aggregation method can use arithmetic mean, instruction number weighted mean, node weighted mean, etc., and arithmetic mean is used by default. Step S208: Send the initial aggregation parameters to each second node for the next round of model of each second node, and receive the current second training parameters sent by each second node. When the predetermined training condition is met, stop the model training, use the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.
[0041] In this step, the predetermined training condition can be that the number of training rounds reaches a predetermined round threshold, or the current aggregation parameter error is less than a predetermined error threshold.
[0042] Step S209: Adaptive selection of prediction nodes. In this step, the first node can match the target problem with the first data directory corresponding to the first node and the second data directories corresponding to each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node. Specifically, the first node can match the target problem with the total data directory to determine the matching first data directory and / or each second data directory from the total data directory, and thus determine the corresponding nodes as the main prediction node and each auxiliary prediction node based on the matching first data directory and / or each second data directory.
[0043] That is, when the customer inputs the target question Q and sends it to a certain computing node of the distributed large language model, first, according to the key elements in the customer input, the capabilities of the distributed large language model are used to match with the node data directory, and the computing node to be called in this conversation is adaptively selected. For example, the input is [Question] Now there is an input: Q. Please determine which knowledge in the total data directory D fl of the collaborative network is required, so as to determine the main prediction node and the auxiliary prediction nodes from each node based on the matching situation between the target question and the total data directory D fl , thus realizing the adaptive selection of the prediction nodes.
[0044] Step S210, answer prediction; In this step, when the first node performs answer prediction, it is divided into the following three situations: Situation 1: When the first node is not determined as the main prediction node or the auxiliary prediction node, the target question is sent to the main prediction node and each auxiliary prediction node; Situation 2: When the first node is determined as the auxiliary prediction node, an auxiliary prompt word is constructed based on the target question, the auxiliary prompt word is sent to the main prediction node, and the target question is sent to the main prediction node and the remaining auxiliary prediction nodes; Situation 3: When the first node is determined as the main prediction node, a main prompt word is constructed based on the target question, the target question is sent to each auxiliary prediction node, and the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node are received. Based on the main prompt word and each auxiliary prompt word, the target answer corresponding to the target question is predicted and output by using the pre-trained target large language model.
[0045] In this embodiment, the main prediction node can adaptively select the corresponding target prompting method according to the characteristics of the input question Q by using the capabilities of the distributed large language model. That is, the main prediction node can generate a total prompt word based on the main prompt word and each auxiliary prompt word, and then generate a prompt template for answer prediction based on the total prompt word and the target prompting method. For example, the input is [Question] Now there is an input: Q. For this input Q, it can be determined which prompting method among the zero-shot prompting method, the few-shot prompting method, the knowledge base prompting method, and the chain of thought prompting method is the target prompting scheme, and then combined with the total prompt word, the target answer is output by using the target large language model. For example, the input is [Question], now there is an input: total prompt word Q k , and please use the way of thinking tree to answer.
[0046] In the answer prediction method based on distributed learning in this embodiment, by determining the main prediction nodes for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, subsequent main prompt words can be generated using the main prediction nodes and auxiliary prompt words can be generated using the auxiliary prediction nodes, making the generation of the prompt words more comprehensive and accurate, laying a foundation for subsequent accurate answer prediction based on the main prompt words, each auxiliary prompt word, and using the target large language model, and facilitating the improvement of the accuracy of answer prediction.
[0047] Another embodiment of this application provides an answer prediction method based on distributed learning, which can be specifically applied to the second node / slave computing node / participant in federated learning, such as Figure 2 As shown, the method in this embodiment includes the following steps: Step S301: Send the local second data directory to the first node for the first node to perform matching based on the first data directory and the second data directories of each second node, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; In the specific implementation process of this step, the second node can pre-construct a second knowledge base based on the predetermined service data, create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node; at the same time, the first node will also construct a first knowledge base based on the predetermined local service data. Thus, when the first node receives each second data directory, it can perform matching with the target question based on the first data directory and each second data directory to determine the nodes matching the target question from each second node and the first node participating in the federated learning as the main prediction node and each auxiliary prediction node.
[0048] Step S302: When the second node is determined as an auxiliary prediction node, receive the target question sent by the first node, construct an auxiliary prompt word based on the target question, and send the auxiliary prompt word to the main prediction node; In this step, when the second node is determined as an auxiliary prediction node, the second node can receive the target question sent by the first node, and then generate / build an auxiliary prompt word based on the target question according to the second database of the second node locally. Thus, the main prediction node can predict and output the target answer corresponding to the target question based on the main prompt word constructed by itself, each auxiliary prompt word, and using the target large language model obtained through pre-training.
[0049] Step S303: When the second node is determined as the main prediction node, receive the target question sent by the first node, construct a main prompt word based on the target question, and receive the auxiliary prompt words sent by each auxiliary prediction node based on the target question, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word, and using the target large language model obtained through pre-training.
[0050] In this step, when the second node is determined as the main prediction node, that is, it is determined to perform answer prediction locally at the second node. Thus, the second node can receive the target question sent by the first node, and then generate / build the main prompt word based on the target question according to the second database locally at the second node, while receiving the auxiliary prompt words sent by each auxiliary prediction node. Furthermore, the second node can predict the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word, and output it by using the pre-trained target large language model.
[0051] In the answer prediction method based on distributed learning in this embodiment, by determining the main prediction node for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, the main prompt word can be generated by using the main prediction node and the auxiliary prompt words can be generated by using the auxiliary prediction nodes subsequently, making the generation of the prompt words more comprehensive and accurate, laying a foundation for accurately performing answer prediction based on the main prompt word and each auxiliary prompt word by using the target large language model, and being beneficial to improving the accuracy of answer prediction.
[0052] Based on the above embodiment, another embodiment of the present application provides an answer prediction method based on distributed learning, which is applied to the second node. The overall process of answer prediction in this embodiment is as follows: Step S401, installation and initialization of the distributed large language model.
[0053] In this step, the leading party of the joint modeling (financial infrastructure, the first node) deploys the components of the multi-center distributed large language model collaboration network in its local data center, and sends the initial large language model and the corresponding components to each second node (slave computing node). Thus, the participating parties of the joint modeling (financial institutions, the second node) will also deploy the components of the multi-center distributed large language model collaboration network in their local data centers. The framework of the distributed machine learning in this embodiment can be implemented by using other software products with the same functions such as Apache Spark MLlib, PyTorchtorch.distributed, and TensorFlowtf.distribute.
[0054] Step S402, sending the local second data directory to the first node for the first node to generate the total data directory of the collaboration network; In this step, the first node can construct a first knowledge base based on predetermined business data and create a first data directory corresponding to the first knowledge base. Similarly, each second node can also construct a second knowledge base based on the predetermined business data, create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node. Thus, the first node can generate a total data directory for determining the main prediction node and each auxiliary prediction node based on the first data directory and each second data directory.
[0055] Specifically, the first node and each second node are connected to their local bond business databases, and the local business data is segmented into documents and vectorized into text using natural language processing technology as the knowledge base. The knowledge base of node i is marked as K nodei , and the file names of the internal data files used by K nodei are sorted into the node data directory D nodei (including the first data directory and each second data directory). The second node sends the second data directory to the first node for synchronization. The first node collects the data directories fed back by all nodes and generates the collaborative network total data directory D fl = {D node1 , D node2 ,...}. That is, after receiving each second data directory, the first node can construct the total data directory based on each second data directory and the first data directory, and then match the target problem with the total data directory, so as to determine the first data directory and / or several second data directories that match the target problem from the total data directory, and then use the corresponding nodes as the main prediction node or auxiliary prediction nodes according to the matching results.
[0056] Step S403, constructing the prompt engineering of the multi-center distributed large language model to obtain the original second sample prompt set; In this step, the second node can pre-construct the original second sample prompt set. The specific process is as follows: Based on each prompting method, several original second sample prompt templates corresponding to each prompting method are generated respectively; where the prompting methods include any one or several of the following: zero-shot prompting method, few-shot prompting method, knowledge base prompting method, and chain of thought prompting method; based on each original first sample prompt template, the original second sample prompt set is constructed.
[0057] That is, the second node can design the prompt engineering related to its own business, and for the needs of different tasks, adopt the following several prompting methods to complete the sample preparation of several types of prompt engineering to obtain the original second sample prompt set. 1. Zero-shot prompting method: The institution does not provide demonstrations related to the task results and directly prompts the language model to give answers related to the task. For example, the input is [Question] What materials are needed to open a settlement account in the inter-bank bond market. 2. Small sample prompting method: The institution provides a small number of prompt examples, such as task descriptions, etc. For example, the input is [Question] Translate the Chinese into English. Bond -> bond; Valuation -> assessment; Management -> management; The European Central Bank has kept its three main interest rates unchanged for the fourth consecutive time, and traders have increased their expectations for the ECB to cut interest rates, expecting a 100-basis-point cut this year ->. 3. Knowledge base prompting method: The institution provides a knowledge base related to the question and its keywords, and uses the context content of the knowledge base to form prompt words. For example, the input is [Question]: What materials are needed to open a settlement account in the inter-bank bond market; [Knowledge base]: Documents such as the "Notice on Matters Concerning Over-the-Counter Business in the Inter-bank Bond Market". After the computing node extracts and semantically understands the knowledge base, it calculates the relevance between the vectorized knowledge base document fragments and the input using a vector distance measurement index, and forms prompt words together with the input question for the several most similar document fragments. The vector distance measurement index described in this application can be implemented using cosine similarity, dot product, Hamming distance, and other vector distance measurement methods. 4. Chain of thought prompting method: For relatively complex and logically strong questions, the institution introduces chain-of-thought steps in the prompt engineering to enhance the large model's ability to handle complex questions, using the following steps: 4.1. Construct examples: Manually write a small number of examples of tasks to be solved, including the input question and its answer, and at the same time provide the chain-of-thought process in the solution process. 4.2. Provide chain-of-thought examples: Add chain-of-thought examples corresponding to the answers in the examples so that the model can learn how to first analyze the question, then gradually carry out chain-of-thought, and finally obtain the answer. 4.3. Prompt the model to carry out chain-of-thought: Provide a chain-of-thought answer in the example to prompt the model to solve the problem in a similar way, gradually showing the analysis idea of the question, and finally giving the answer. 4.4. Evaluate the effect of chain-of-thought: Compare the effect of directly giving answers in the test set with the effect after adding chain-of-thought prompts, and prove that the former can only answer simple questions, while the latter can gradually analyze more complex questions through chain-of-thought, thus obtaining better results.
[0058] In this step, by adopting the above 4 prompting methods, the original second sample prompt templates corresponding to each prompting method can be constructed, and thus the original second sample prompt set can be constructed based on each original second sample prompt template.
[0059] Step S404, based on the original second sample prompt set, construct a target second sample prompt set to complete the construction of the automatic prompt engineering; In the specific implementation process of this step, based on the original second sample prompt set local to the second node, the target model can be used to generate the target second sample prompt set. The target model can be a pre-trained distributed large language model. That is, the original first sample prompt set can be used as the input, and the distributed large language model can be used to automatically generate prompt instructions. The best candidate prompt instructions can be selected through filtering and searching methods to obtain the target second sample prompt set PE nodei , and the steps are as follows: 1. Based on the instruction examples in the original second sample prompt set, let the distributed large language model generate automatic instructions. For example, for the following prompt template, input and output examples are given to guide the model to generate instructions: [Question] Generate input-output instruction pairs according to the instruction examples. The instruction is... 2. Given a training set Dtrain, set a scoring function, and select the instruction that can make the samples in the training set have the highest score from the automatically generated candidate instructions. In this embodiment, the scoring function can be implemented using accuracy, language model likelihood value, and other machine learning evaluation metrics. 3. Use the Monte Carlo method for iterative search: Let the model generate variants of the instructions in the second step through prompts to further search for better instructions. For example, this prompt is used to generate approximate instructions: [Question] Generate a variant of the following instruction while keeping the semantic meaning unchanged. Input:... Output:...
[0060] In step S405, send the target second sample prompt set to the first node for prompt engineering synchronization.
[0061] In the specific implementation process of this step, the second node can send the target second sample prompt set PE nodei to the first node for the first node to determine the prompt set directory of the prompt subset used for initial model training based on the target second sample prompt sets of each second node and the target first sample prompt set of the first node; receive the prompt set directory sent by the first node, and based on the prompt set directory D_PE fl ' determine the target second sample prompt subset PE nodei from the local target second sample prompt set PE fl '.
[0062] That is, after the second node i forms the prompt engineering instruction set / target second sample prompt set P Enodei , it can send the target second sample prompt set P Enodei to the first node. Thus, the first node can receive the target second sample prompt sets sent by each second node, and then can generate the total sample prompt set PE fl={PE node1 , PE node2 , ……}. Then, select the prompt subset PE fl for this training from the total sample prompt set PE fl ’, and determine the prompt set directory D_PE fl ’ corresponding to the prompt subset PE fl ’. Finally, the first node synchronously sends the information such as the prompt set directory D_PE fl ’ and the pre-trained model used in this training to each second node, so that each second node can determine the target second sample prompt subset PE fl ’ from the target second sample prompt set PE nodei according to the prompt set directory D_PE nodei ’.
[0063] In this embodiment, all nodes involved in this task will, according to the prompt set directory D_PE fl ’, use the selected part PE nodei ’ of the instruction set PE nodei ’ saved locally as training data and conduct training in batches. The advantage of this process is that the training data of the computing nodes meets the screening criteria of the leading party, with relatively high security. Since the first node transmits the instruction set directory during the feedback process, the communication pressure is small and the possibility of being attacked and stolen is low; the distributed computing method flexibly utilizes the computing resources of each node, and the system efficiency is high.
[0064] Step S406: Perform model training on the initial large language model; In this step, the second node can perform model training on the initial large language model based on the target second sample prompt subset PE nodeii ’ to obtain the initial second training parameters; In this step, in this embodiment, the initial large language model / pre-trained model of the large language model can use other Chinese base models with the same functions such as Chinese-LLaMA, OpenChineseLLaMA, and Ziya-LLaMA, and Chinese LLaMA is used by default.
[0065] Step S407: Send the initial second training parameters to the first node for the first node to perform parameter aggregation based on the initial second training parameters of each second node and the initial first training parameters of the first node to obtain the initial aggregation parameters; In this step, all nodes involved in this task complete a local training respectively and send the training results to the first node, and the first node reads all the parameters and performs parameter aggregation. In this embodiment, the parameter aggregation method can use methods such as arithmetic mean, instruction quantity weighted average, and node weighted average, and arithmetic mean is used by default. Step S408: Receive the initial aggregation parameters sent by the first node, perform the next round of model training based on the initial aggregation parameters, send the currently obtained second training parameters to the first node, and stop the model training until the predetermined training conditions are met. Take the received current aggregation parameters as the target aggregation parameters to obtain the target large language model.
[0066] In this step, the predetermined training conditions can be that the number of training rounds reaches a predetermined round threshold, or the current aggregation parameter error is less than a predetermined error threshold.
[0067] Step S409: Answer prediction; Before answer prediction, the first node determines the primary prediction node and several secondary prediction nodes from the first node and each second node. That is, the first node can match the target question with the first data directory corresponding to the first node and the second data directories corresponding to each second node to determine the primary prediction node and each secondary prediction node from the first node and each second node. Specifically, the first node can match the target question with the total data directory to determine the matching first data directory and / or each second data directory from the total data directory, and then determine the corresponding nodes as the primary prediction node and each secondary prediction node based on the matching first data directory and / or each second data directory. That is, when the customer inputs the target question Q to a certain computing node of the distributed large language model, first, according to the key elements in the customer input, use the capabilities of the distributed large language model to match with the node data directory, and adaptively select the computing nodes to be called for this conversation. For example, the input is [Question] Now there is an input: Q. Please determine which knowledge in the total data directory D fl of the collaboration network is required, and then based on the matching situation between the target question and the total data directory D fl determine the primary prediction node and secondary prediction nodes from each node, so as to achieve the adaptive selection of prediction nodes.
[0068] In this step, when the second node is determined as the secondary prediction node, the second node can receive the target question sent by the first node, construct a secondary prompt word based on the target question, and send the secondary prompt word to the primary prediction node. That is, the second node can generate / construct a secondary prompt word based on the target question according to the second database of the second node locally, and then send the secondary prompt word to the primary prediction node. Thus, the primary prediction node can predict the target answer corresponding to the target question based on the primary prompt word constructed by itself and each secondary prompt word, and output it using the pre-trained target large language model.
[0069] In this step, when the second node is determined as the main prediction node, the second node can receive the target question sent by the first node, construct the main prompt word based on the target question, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node. Then, based on the main prompt word and each auxiliary prompt word, use the pre-trained target large language model to predict the target answer corresponding to the target question and output it.
[0070] In this embodiment, the main prediction node can adaptively select the corresponding target prompting method according to the characteristics of the input question Q, that is, the main prediction node can generate the total prompt word based on the main prompt word and each auxiliary prompt word, and then generate the prompt template for answer prediction based on the total prompt word and the target prompting method. For example, the input is [Question] Now there is an input: Q. For this input Q, it can be determined which of the zero-shot prompting method, few-shot prompting method, knowledge base prompting method, and chain of thought prompting method is the target prompting scheme, and then combined with the total prompt word, use the target large language model to output the target answer. For example, the input is [Question], now there is an input: total prompt word Q k , and please use the way of thinking tree to answer.
[0071] In the answer prediction method based on distributed learning in this embodiment, by determining the main prediction node for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, subsequently, the main prediction node can be used to generate the main prompt word and the auxiliary prediction nodes can be used to generate the auxiliary prompt words, making the generation of the prompt words more comprehensive and accurate, laying a foundation for accurately predicting the answer based on the main prompt word and each auxiliary prompt word using the target large language model, which is beneficial to improving the accuracy of answer prediction.
[0072] Another embodiment of this application provides an answer prediction device based on distributed learning, as Figure 3 shown, including: The first receiving module is used to receive the second data directories sent by each second node, and match the target question with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; The first sending module is used to send the target question to the main prediction node and each auxiliary prediction node when the first node is not determined as the main prediction node or the auxiliary prediction node; The first construction module is used to construct the auxiliary prompt word based on the target question when the first node is determined as the auxiliary prediction node, send the auxiliary prompt word to the main prediction node, and send the target question to the main prediction node and the remaining auxiliary prediction nodes; The first prediction module is used to, when the first node is determined as the main prediction node, construct a main prompt word based on the target question, send the target question to each auxiliary prediction node, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node. Then, based on the main prompt word and each auxiliary prompt word, use the pre-trained target large language model to predict the target answer corresponding to the target question and output it.
[0073] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes a first creation module, and the first creation module is used for: Construct a first knowledge base based on the predetermined service data and create a first data directory corresponding to the first knowledge base.
[0074] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes: a first generation module, an aggregation module, and a first training module; The first generation module is used for: generating a target first sample prompt subset based on the original first sample prompt set on the local side of the first node; The first training module is used for: performing model training on the initial large language model based on the target first sample prompt subset to obtain initial first training parameters; The aggregation module is used for: receiving the initial second training parameters sent by each second node, and performing parameter aggregation based on each initial second training parameter and the initial first training parameter to obtain initial aggregation parameters; The first sending module is further used for: sending the initial aggregation parameters to each second node for each second node to perform the next round of model, and receiving the current second training parameters sent by each second node based on the aggregation module until the predetermined training condition is met, then stopping the model training, taking the current aggregation parameters as the target aggregation parameters, and obtaining the target large language model.
[0075] In the specific implementation process of this embodiment, the first generation module specifically includes: The first generation unit is used for: generating a target first sample prompt set based on the original first sample prompt set on the local side of the first node by using the target model; The prompt set receiving unit is used for: receiving the target second sample prompt sets sent by each second node, and determining the prompt set directory of the prompt subset for model training based on the target second sample prompt sets of each second node and the target first sample prompt set of the first node; The first determination unit is used for: determining the target first sample prompt subset from the target first sample prompt set based on the prompt set directory, and sending the prompt set directory to the second node for the second node to determine the target second sample prompt subset from the target second sample prompt set based on the prompt set directory.
[0076] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes: a first prompt set construction module for pre-constructing an original first sample prompt set, and the prompt set construction module is used to: generate a number of original first sample prompt templates corresponding to each prompt method based on each prompt method; where the prompt method includes any one or several of the following: zero-shot prompt method, few-shot prompt method, knowledge base prompt method, and chain of thought prompt method; construct an original first sample prompt set based on each original first sample prompt template.
[0077] In the answer prediction device based on distributed learning in this application, by determining the main prediction node for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, subsequently, the main prediction node can be used to generate main prompt words and the auxiliary prediction nodes can be used to generate auxiliary prompt words, making the generation of prompt words more comprehensive and accurate, laying a foundation for accurately predicting answers based on the main prompt words and each auxiliary prompt word using the target large language model, and being beneficial to improving the accuracy of answer prediction.
[0078] Another embodiment of this application provides an answer prediction device based on distributed learning, as Figure 4 shown, including: A second sending module 21, configured to send the local second data directory to the first node, for the first node to perform matching based on the first data directory and the second data directories of each second node, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; A second construction module 22, configured to receive the target question sent by the first node when the second node is determined as an auxiliary prediction node, construct an auxiliary prompt word based on the target question, and send the auxiliary prompt word to the main prediction node; A second prediction module 23, configured to receive the target question sent by the first node when the second node is determined as the main prediction node, construct a main prompt word based on the target question, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word using the pre-trained target large language model.
[0079] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes: a second creation module, and the second creation module is used to: construct a second knowledge base based on the predetermined service data; create a second data directory corresponding to the second knowledge base based on the second knowledge base.
[0080] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes: a second generation module, a second training module, and a second receiving module; The second generation module is used to generate a target second sample prompt subset based on the original second sample prompt set local to the second node; The second training module is used to perform model training on the initial large language model based on the target second sample prompt subset to obtain initial second training parameters; The second sending module is further used to send the initial second training parameters to the first node for the first node to perform parameter aggregation based on the initial second training parameters of each second node and the initial first training parameters of the first node to obtain initial aggregation parameters; The second receiving module is further used to receive the initial aggregation parameters sent by the first node, perform the next round of model training using the second training module based on the initial aggregation parameters, and send the currently obtained second training parameters to the first node based on the second sending module. When the predetermined training condition is met, stop the model training, and use the received current aggregation parameters as the target aggregation parameters to obtain the target large language model.
[0081] In the specific implementation process of this embodiment, the second generation module specifically includes: A second generation unit is used to generate a target second sample prompt set based on the original second sample prompt set local to the second node by using a target model; A prompt set sending unit is used to send the target second sample prompt set to the first node for the first node to determine a prompt set directory for the prompt subset used for initial model training based on the target second sample prompt sets of each second node and the target first sample prompt set of the first node; A second determination unit is used to receive the prompt set directory sent by the first node and determine the target second sample prompt subset from the local target second sample prompt set based on the prompt set directory.
[0082] In the specific implementation process of this embodiment, the answer prediction device based on distributed learning further includes: a second prompt set construction module for pre - constructing the original second sample prompt set, and the second prompt set construction module is used for: Based on each prompting method, respectively generate a number of original second sample prompt templates corresponding to each prompting method; where the prompting methods include any one or more of the following: zero - shot prompting method, few - shot prompting method, knowledge - base prompting method, and chain - of - thought prompting method; construct the original second sample prompt set based on each original second sample prompt template.
[0083] In the answer prediction device based on distributed learning in this application, by determining the main prediction nodes for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, subsequently, the main prediction nodes can be used to generate main prompt words and the auxiliary prediction nodes can be used to generate auxiliary prompt words, making the generation of prompt words more comprehensive and accurate, laying a foundation for subsequent accurate answer prediction based on the main prompt words, each auxiliary prompt word, and using the target large language model, and being beneficial to improving the accuracy of answer prediction.
[0084] Another embodiment of this application provides an electronic device, as Figure 5 shown, at least including a memory 1 and a processor 2. A computer program is stored on the memory 1, and when the processor 2 executes the computer program on the memory 1, the following method steps are implemented: Step 1: Receive the second data directories sent by each second node, and match the target question with the first data directory corresponding to the first node and each second data directory to determine the main prediction node and each auxiliary prediction node from the first node and each second node; Step 2: When the first node is not determined as the main prediction node or the auxiliary prediction node, send the target question to the main prediction node and each auxiliary prediction node; When the first node is determined as the auxiliary prediction node, construct an auxiliary prompt word based on the target question, send the auxiliary prompt word to the main prediction node, and send the target question to the main prediction node and the remaining auxiliary prediction nodes; When the first node is determined as the main prediction node, construct a main prompt word based on the target question, send the target question to each auxiliary prediction node, receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word using the pre-trained target large language model.
[0085] Alternatively, the following method steps are implemented: Step 1: Send the local second data directory to the first node for the first node to perform matching based on the first data directory and the second data directories of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node; Step 2: When the second node is determined as the auxiliary prediction node, receive the target question sent by the first node, construct an auxiliary prompt word based on the target question, and send the auxiliary prompt word to the main prediction node; When the second node is determined as the main prediction node, receive the target question sent by the first node, construct the main prompt word based on the target question, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node. Then, based on the main prompt word and each auxiliary prompt word, use the pre-trained target large language model to predict the target answer corresponding to the target question and output it.
[0086] For the specific implementation process of the above method steps, reference can be made to the embodiments of any of the above answer prediction methods based on distributed learning or image processing methods. This embodiment will not be repeated here.
[0087] In the device of this application, by determining the main prediction node for answer prediction and the auxiliary prediction nodes for auxiliary prediction from each second node and the first node, subsequently, the main prediction node can be used to generate the main prompt word and the auxiliary prediction nodes can be used to generate the auxiliary prompt words, making the generation of the prompt words more comprehensive and accurate. This lays a foundation for accurately predicting the answer based on the main prompt word and each auxiliary prompt word using the target large language model, which is beneficial to improving the accuracy of answer prediction.
[0088] The above embodiments are only exemplary embodiments of this application and are not used to limit this application. The protection scope of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of this application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of this application.
Claims
1. A method for answer prediction based on distributed learning, applied to a first node, characterized in that, Including: Receiving the second data directories sent by each second node, and matching the target problem with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; When the first node is not determined as the main prediction node or the auxiliary prediction node, sending the target problem to the main prediction node and each auxiliary prediction node; When the first node is determined as the auxiliary prediction node, constructing an auxiliary prompt word based on the target problem, sending the auxiliary prompt word to the main prediction node, and sending the target problem to the main prediction node and the remaining auxiliary prediction nodes; When the first node is determined as the main prediction node, constructing a main prompt word based on the target problem, sending the target problem to each auxiliary prediction node, and receiving the auxiliary prompt words constructed based on the target problem sent by each auxiliary prediction node, and predicting and outputting the target answer corresponding to the target problem by using the pre-trained target large language model based on the main prompt word and each auxiliary prompt word.
2. The method according to claim 1, wherein Before matching the target problem with the first data directory corresponding to the first node and the second data directories corresponding to each second node, the method further includes: Constructing a first knowledge base based on predetermined service data, and creating a first data directory corresponding to the first knowledge base.
3. The method according to claim 1, wherein The method further includes: performing model training on the initial large language model, specifically including: Generating a target first sample prompt subset based on the original first sample prompt set local to the first node; Performing model training on the initial large language model based on the target first sample prompt subset to obtain initial first training parameters; Receiving the initial second training parameters sent by each second node, and performing parameter aggregation based on each initial second training parameter and the initial first training parameters to obtain initial aggregation parameters; Sending the initial aggregation parameters to each second node for each second node to perform the next round of model, and receiving the current second training parameters sent by each second node, and stopping the model training until the predetermined training conditions are met, taking the current aggregation parameters as the target aggregation parameters, and obtaining the target large language model.
4. The method according to claim 3, characterized in that, The generating the target first sample prompt subset based on the original first sample prompt set local to the first node specifically includes: Generating a target first sample prompt set by using a target model based on the original first sample prompt set local to the first node; Receiving the target second sample prompt sets sent by each second node, and determining a prompt set directory of the prompt subset for model training based on the target second sample prompt sets of each second node and the target first sample prompt set of the first node; Determining a target first sample prompt subset from the target first sample prompt set based on the prompt set directory, and sending the prompt set directory to the second node for the second node to determine a target second sample prompt subset from the target second sample prompt set based on the prompt set directory.
5. The method according to claim 3, characterized in that The method further includes: pre-constructing the original first sample prompt set, specifically including: Based on each prompting method, several original first-sample prompting templates corresponding to each prompting method are generated respectively; where the prompting methods include any one or more of the following: zero-shot prompting method, few-shot prompting method, knowledge base prompting method, and chain of thought prompting method; Based on each original first-sample prompting template, an original first-sample prompting set is constructed.
6. A method for answer prediction based on distributed learning, applied to the second node, characterized in that, Including: Send the local second data directory to the first node for the first node to perform matching based on the first data directory and the second data directories of each second node, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; When the second node is determined as an auxiliary prediction node, receive the target problem sent by the first node, construct an auxiliary prompt word based on the target problem, and send the auxiliary prompt word to the main prediction node; When the second node is determined as the main prediction node, receive the target problem sent by the first node, construct a main prompt word based on the target problem, and receive the auxiliary prompt words sent by each auxiliary prediction node based on the target problem, and predict the target answer corresponding to the target problem and output it by using the pre-trained target large language model based on the main prompt word and each auxiliary prompt word.
7. The method according to claim 6, characterized in that, Before sending the local second data directory to the first node, the method further includes: Construct a second knowledge base based on the predetermined business data; Create a second data directory corresponding to the second knowledge base based on the second knowledge base.
8. An answer prediction device based on distributed learning, characterized in that, Including: A first receiving module, configured to receive the second data directories sent by each second node, and match the target problem with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; A first sending module, configured to send the target problem to the main prediction node and each auxiliary prediction node when the first node is not determined as the main prediction node or the auxiliary prediction node; A first construction module, configured to construct an auxiliary prompt word based on the target problem when the first node is determined as an auxiliary prediction node, send the auxiliary prompt word to the main prediction node, and send the target problem to the main prediction node and the remaining auxiliary prediction nodes; A first prediction module, configured to construct a main prompt word based on the target problem when the first node is determined as the main prediction node, send the target problem to each auxiliary prediction node, and receive the auxiliary prompt words sent by each auxiliary prediction node based on the target problem, and predict the target answer corresponding to the target problem and output it by using the pre-trained target large language model based on the main prompt word and each auxiliary prompt word.
9. An answer prediction device based on distributed learning, characterized in that, Including: A second sending module, configured to send the local second data directory to the first node for the first node to perform matching based on the first data directory and the second data directories of each second node, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node; A second construction module, configured to receive the target problem sent by the first node when the second node is determined as an auxiliary prediction node, construct an auxiliary prompt word based on the target problem, and send the auxiliary prompt word to the main prediction node; The second prediction module is used to receive the target question sent by the first node when the second node is determined as the main prediction node, construct the main prompt word based on the target question, and receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node. Then, it predicts the target answer corresponding to the target question by using the main prompt word and each auxiliary prompt word and the pre-trained target large language model, and outputs the result.
10. An electronic device, characterized in that, It includes at least a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program on the memory, it implements the steps of the answer prediction method based on distributed learning described in any one of claims 1-5 or 6-7 above.
Citation Information
Patent Citations
Search cue word determination method and system and information processing method
CN113032819A
Large model cue word intelligent routing method, device and equipment and storage medium
CN119691137A
Data query method and device based on large model and electronic equipment
CN119829602A
Multilevel data analysis
US12210839B1
Method for pre-training model, device, and storage medium
US20230040095A1