An answer prediction method, device and equipment based on distributed learning

By identifying the primary and secondary prediction nodes in a distributed learning system, generating primary and secondary prompt words, and combining them with a target large language model for answer prediction, the problem of inaccurate answer prediction in distributed machine learning is solved, and the accuracy of answer prediction is improved.

CN120196729BActive Publication Date: 2025-11-21ZHONGJINKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510667786.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-11-21
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The problem of inaccurate answer prediction in existing distributed machine learning.

Method used

By determining the primary prediction node and secondary prediction node in a distributed learning system, the primary prediction node generates primary prompt words, and the secondary prediction node generates secondary prompt words. The answer is then predicted by combining these with the target large language model.

Benefits of technology

It improves the accuracy of answer prediction, making the generation of prompt words more comprehensive and accurate, thus enhancing the precision of answer prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196729B_ABST
    Figure CN120196729B_ABST
Patent Text Reader

Abstract

The application discloses an answer prediction method, device and equipment based on distributed learning. The method comprises the following steps: receiving second data directories sent by second nodes, and determining a main prediction node and each auxiliary prediction node based on a target question, a first data directory and each second data directory; sending the target question to the main prediction node and each auxiliary prediction node; or, constructing an auxiliary prompt word based on the target question, sending the auxiliary prompt word to the main prediction node, and sending the target question to the main prediction node and each remaining auxiliary prediction node; or, constructing a main prompt word based on the target question, sending the target question to each auxiliary prediction node, receiving auxiliary prompt words sent by each auxiliary prediction node and constructed based on the target question, and predicting and outputting a target answer corresponding to the target question by using a target large language model obtained by pre-training based on the main prompt word and each auxiliary prompt word. The application can provide answer prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, in particular to an answer prediction method and device based on distributed learning and equipment. BACKGROUND

[0002] Distributed machine learning (DML) is also known as distributed learning, which refers to an algorithm and system for machine learning or deep learning using multiple computing nodes, aiming to improve performance, protect privacy, and be scalable to larger training data and larger models. However, the existing distributed learning has the problem of inaccurate answer prediction. SUMMARY

[0003] Therefore, the present application provides an answer prediction method and device based on distributed learning and equipment, which mainly aims to solve the problem of inaccurate answer prediction.

[0004] To solve the above problems, the present application provides an answer prediction method based on distributed learning, applied to a first node, comprising:

[0005] receiving second data directories sent by each second node, and matching the target question with a first data directory corresponding to the first node and each second data directory to determine a main prediction node and each auxiliary prediction node from the first node and each second node;

[0006] when the first node is not determined as the main prediction node or the auxiliary prediction node, sending the target question to the main prediction node and each auxiliary prediction node;

[0007] when the first node is determined as the auxiliary prediction node, constructing an auxiliary prompt word based on the target question, sending the auxiliary prompt word to the main prediction node, and sending the target question to the main prediction node and each remaining auxiliary prediction node;

[0008] when the first node is determined as the main prediction node, constructing a main prompt word based on the target question, sending the target question to each auxiliary prediction node, receiving auxiliary prompt words sent by each auxiliary prediction node based on the target question, and predicting a target answer corresponding to the target question using a target large language model obtained by pre-training based on the main prompt word and each auxiliary prompt word and outputting.

[0009] To solve the above problems, the present application provides an answer prediction method based on distributed learning, applied to a second node, comprising:

[0010] send the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directory of each second node to determine the primary prediction node and each secondary prediction node from the first node and each second node;

[0011] when the second node is determined as the secondary prediction node, receive the target question sent by the first node, construct the secondary prompt word based on the target question, and send the secondary prompt word to the primary prediction node;

[0012] when the second node is determined as the primary prediction node, receive the target question sent by the first node, construct the primary prompt word based on the target question, and receive the secondary prompt word sent by each secondary prediction node based on the target question, and predict the target answer corresponding to the target question using the target large language model trained in advance based on the primary prompt word and each secondary prompt word and output.

[0013] To solve the above problems, the application provides an answer prediction device based on distributed learning, comprising:

[0014] The first receiving module is configured to receive the second data directory sent by each second node, and match the target question with the first data directory corresponding to the first node and the second data directory of each second node to determine the primary prediction node and each secondary prediction node from the first node and each second node.

[0015] The first sending module is configured to send the target question to the primary prediction node and each secondary prediction node when the first node is not determined as the primary prediction node or the secondary prediction node.

[0016] The first constructing module is configured to, when the first node is determined as the secondary prediction node, construct the secondary prompt word based on the target question, send the secondary prompt word to the primary prediction node, and send the target question to the primary prediction node and each remaining secondary prediction node.

[0017] The first prediction module is configured to, when the first node is determined as the primary prediction node, construct the primary prompt word based on the target question, send the target question to each secondary prediction node, receive the secondary prompt word sent by each secondary prediction node based on the target question, and predict the target answer corresponding to the target question using the target large language model trained in advance based on the primary prompt word and each secondary prompt word and output.

[0018] To solve the above problems, the application provides an answer prediction device based on distributed learning, comprising:

[0019] The second sending module is configured to send the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directory of each second node to determine the primary prediction node and each secondary prediction node from the first node and each second node.

[0020] a second construction module, configured to, when the second node is determined as the auxiliary prediction node, receive the target question sent by the first node, construct an auxiliary prompt word based on the target question, and send the auxiliary prompt word to the main prediction node;

[0021] a second prediction module, configured to, when the second node is determined as the main prediction node, receive the target question sent by the first node, construct a main prompt word based on the target question, and receive the auxiliary prompt words sent by the auxiliary prediction nodes based on the target question, and predict a target answer corresponding to the target question based on the main prompt word and the auxiliary prompt words by using a target large language model trained in advance and output the target answer.

[0022] To solve the above problems, the present application provides an electronic device, at least comprising a memory and a processor, the memory has a computer program stored thereon, and the processor implements the steps of the answer prediction method based on distributed learning according to any one of the above embodiments when executing the computer program stored on the memory.

[0023] The answer prediction method, device and equipment based on distributed learning in the present application determine the main prediction node for predicting answers and the auxiliary prediction node for auxiliary prediction from the second nodes and the first node, and then generate the main prompt word by using the main prediction node and the auxiliary prompt word by using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction by using the target large language model based on the main prompt word and the auxiliary prompt words, and helps to improve the accuracy of answer prediction.

[0024] The above description is only a summary of the technical scheme of the present application. In order to more clearly understand the technical means of the present application, the content of the specification can be implemented, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0025] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present application. Moreover, the same reference numerals in the attached drawings indicate the same or similar components. In the drawings:

[0026] Figure 1 A flowchart of an answer prediction method based on distributed learning according to an embodiment of the present application;

[0027] Figure 2 A flowchart of an answer prediction method based on distributed learning according to another embodiment of the present application;

[0028] Figure 3 A structure block diagram of an answer prediction apparatus based on distributed learning according to another embodiment of the present application;

[0029] Figure 4 A structure block diagram of an answer prediction apparatus based on distributed learning according to another embodiment of the present application;

[0030] Figure 5 A structure block diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION

[0031] Various aspects and features of the present application are described in the specification set forth below. The advantages of the present application will become apparent.

[0032] It is to be understood that the embodiments of the application which are being applied can be subjected to many modifications. Therefore, the description set forth above is not to be considered as limiting, but merely as a basis- example for the practice of the embodiments. Other modifications that are within the scope and spirit of the present application will occur to those skilled in the art upon reading this description.

[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present application and, together with the general description of the application given above, and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0034] These and other characteristics of the present application will become apparent from the following description of the preferred forms given, by way of non-limiting example only, with reference to the attached drawings.

[0035] It is also to be understood that, although a certain number of specific embodiments of the present application are described herein, many other modifications, which will come readily to mind of one skilled in the art, can be made thereto without deviating from the spirit of the present application.

[0036] The above and other aspects, features, and advantages of the present application will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, when properly considered together.

[0037] Specific embodiments of the present application are described hereinafter, with reference to the drawings; however, it will be understood that the application is not limited to the specific embodiments described and shown herein, but encompasses many other embodiments, which will come to mind to one skilled in the art upon reading the description and drawings. Well-known and / or repeated functions and structures are not described in detail to avoid obscuring the present application in unnecessary or redundant detail. Therefore, specific structural and functional details disclosed herein are not to be interpreted in a limiting manner, but merely as a basis for the claims and representative basis for teaching one skilled in the art to employ the present application in virtually any appropriate detailed structure.

[0038] The specification can use phrases such as "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which can refer to one or more of the same or different embodiments of the application.

[0039] The embodiment of the present application provides an answer prediction method based on distributed learning, which can be applied to a first node / main computing node / leader in federated learning, as shown in the following figure. Figure 1 The method in the embodiment includes the following steps:

[0040] In step S101, the first node receives the second data directory sent by each second node, and matches the target question with the first data directory corresponding to the first node and each second data directory, to determine the main prediction node and each auxiliary prediction node from the first node and each second node.

[0041] In this step, the first node also constructs a first knowledge base based on predetermined local business data, and then creates a first data directory corresponding to the first knowledge base. Similarly, each second node can construct a second knowledge base based on predetermined business data in advance, and create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node. Therefore, the first node can determine the node matching the target question as the main prediction node and each auxiliary prediction node from each second node participating in federated learning and the first node according to the directory of the local data of each node (the first data directory and each second data directory).

[0042] In the specific implementation process, the main prediction node and the auxiliary prediction node can be determined from each node according to the matching degree of the data directory and the target question in each node, that is, when the matching degree of the data directory and the target question is greater than a predetermined threshold, the node can be determined as a prediction node, and when there are multiple prediction nodes, the node with the highest matching degree can be determined as the main prediction node from the prediction nodes, and the remaining prediction nodes are auxiliary prediction nodes.

[0043] In step S102, when the first node is not determined as the main prediction node or the auxiliary prediction node, the target question is sent to the main prediction node and each auxiliary prediction node.

[0044] In the specific implementation process of this step, if the first node determines that the local first data directory does not match the target question, it can be determined that the local node is not suitable for being a prediction node (i.e., not suitable for being a main prediction node or an auxiliary prediction node). Therefore, the first node can send the target question to the corresponding main prediction node and each auxiliary prediction node, so that each auxiliary prediction node generates / constructs auxiliary prompt words based on the target question and the database of the auxiliary prediction node, each main prediction node generates / constructs main prompt words based on the target question and the database of the main prediction node, and the main prediction node uses the target large language model obtained by pre-training to predict the target answer corresponding to the target question based on the main prompt and the auxiliary prompt words sent by each auxiliary prediction node and outputs the target answer.

[0045] Step S103, when the first node is determined as an auxiliary prediction node, constructing an auxiliary prompt word based on the target question, sending the auxiliary prompt word to the main prediction node, and sending the target question to the main prediction node and each remaining auxiliary prediction node;

[0046] In this step, when the first node is determined as an auxiliary prediction node, the first node can directly generate / construct an auxiliary prompt word based on the target question according to the database locally stored in the first node. Meanwhile, the first node can send the auxiliary prompt word to the main prediction node, and send the target question to the main prediction node and each remaining auxiliary prediction node, so that each remaining auxiliary prediction node can construct an auxiliary prompt word based on the database locally stored in the auxiliary prediction node, and the main prediction node can construct a main prompt word based on the database locally stored in the main prediction node. Thus, the main prediction node can predict and output a target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word, using a target large language model trained in advance.

[0047] Step S104, when the first node is determined as a main prediction node, constructing a main prompt word based on the target question, sending the target question to each auxiliary prediction node, receiving auxiliary prompt words constructed based on the target question and sent by each auxiliary prediction node, and predicting and outputting a target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word, using a target large language model trained in advance.

[0048] In this step, when the first node is determined as a main prediction node, that is, the first node can locally predict an answer. Thus, the first node can generate / construct a main prompt word based on the target question and the database locally stored in the first node, and send the target question to each auxiliary prediction node, and then receive auxiliary prompt words generated / constructed for the target question and sent by each auxiliary prediction node. Furthermore, the first node can predict and output a target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word, using a target large language model trained in advance.

[0049] The answer prediction method based on distributed learning in the embodiment can determine a main prediction node for predicting an answer and auxiliary prediction nodes for assisting in predicting an answer from each second node and the first node. Subsequently, the main prediction node can generate a main prompt word, and the auxiliary prediction nodes can generate auxiliary prompt words. This makes the generation of prompt words more comprehensive and accurate, lays a foundation for predicting an answer accurately based on a target large language model using a main prompt word and each auxiliary prompt word, and is beneficial to improving the accuracy of answer prediction.

[0050] On the basis of the above embodiment, another embodiment of the present application provides an answer prediction method based on distributed learning, applied to a first node. The overall process of answer prediction is as follows:

[0051] Step S201, installation and initialization of the distributed large language model.

[0052] In this step, the leading party of joint modeling (financial infrastructure, first node) deploys the components of the multi-center distributed large language model collaboration network in its local data center, and sends the initial large language model and the corresponding components to each second node (slave computing node). As a result, the participants of joint modeling (financial institutions, second nodes) will also deploy the components of the multi-center distributed large language model collaboration network in their local data centers. The distributed machine learning framework in this embodiment can be implemented using Apache Spark MLlib, PyTorchtorch.distributed, and TensorFlowtf.distribute, and other software products with the same function.

[0053] Step S202, generation of the total data directory of the collaboration network;

[0054] In this step, the first node can construct a first knowledge base based on the predetermined business data, and create a first data directory corresponding to the first knowledge base. At the same time, the first node will receive the second data directory sent by each second node. The second data directory corresponds to the second knowledge base constructed by the second node locally. Therefore, the first node can subsequently directly match the target problem based on the first data directory and each second data directory, or generate a total data directory for matching the target problem based on the first data directory and each second data directory.

[0055] Specifically, the first node and each second node are connected with the local bond business database, and the local business data is segmented and vectorized using natural language processing technology to serve as a knowledge base. The knowledge base of node i is marked as K nodei , and K nodei The file name of the internal data file used is arranged as the node data directory D nodei (including the first data directory and each second data directory), and the second node sends the second data directory to the first node for synchronization. The first node collects the data directories fed back by all nodes to generate the total data directory D fl of the collaboration network = {D node1 , D node2 , …}. That is, after receiving each second data directory, the first node can construct a total data directory based on each second data directory and the first data directory, and then match the target problem with the total data directory, so as to determine the first data directory and / or several second data directories matching the target problem from the total data directory.

[0056] Step S203, prompt engineering of the multi-center distributed large language model is constructed to construct the original first sample prompt set;

[0057] In this step, the first node can pre-construct the original first sample prompt set. The specific process is: based on each prompt method, a plurality of original first sample prompt templates corresponding to each prompt method are generated; wherein the prompt method includes any one or several of the following: zero sample prompt method, small sample prompt method, knowledge base prompt method and thought chain prompt method; based on each original first sample prompt template, the original first sample prompt set is constructed.

[0058] That is, the first node can design a prompt engineering related to its own business, and according to the needs of different tasks, use the following several prompt methods to complete the sample preparation of several types of prompt engineering to construct the original first sample prompt set.

[0059] 1. Zero sample prompt method: the institution does not provide task result related demonstration, and directly prompts the language model to give task related answers. For example, the input is [question] what materials are needed to open a bank bond market settlement account.

[0060] 2. Small sample prompt method: the institution provides a small amount of prompt examples, such as task instructions, etc. For example, the input is [question] translate Chinese into English. Bond -> bond; assessment -> assessment; management -> management; The European Central Bank has maintained the three interest rates unchanged for the fourth consecutive time, and traders have increased their expectations for the European Central Bank to cut interest rates by 100 basis points this year.

[0061] 3. Knowledge base prompt method: the institution provides a knowledge base related to the question and the key words of the question, and uses the context content of the knowledge base to form prompt words. For example, the input is [question]: what materials are needed to open a bank bond market settlement account; [knowledge base]: documents such as "Notice on Matters Related to Bank Bond Market Counter Business". After the computing node performs text extraction and semantic understanding on the knowledge base, the most similar document fragments are calculated with the input question to form prompt words. The vector distance measurement index described in the present application can be realized using cosine similarity, dot product, Hamming distance and other vector distance measurement methods.

[0062] 4. Thought chain prompt method: the institution introduces chain thinking steps in the prompt engineering for complex and logical problems to enhance the processing capacity of the large model for complex problems, using the following steps:

[0063] 4.1. Constructing samples: manually writing a small number of task samples that need to be solved, which contain input questions and their answers, as well as the chain thinking process in the solution process.

[0064] 4.2. Provide chain-of-thought examples: In the examples, add chain-of-thought examples corresponding to the answers, so that the model learns how to analyze the problem first, then gradually carry out chain-of-thought, and finally get the answer.

[0065] 4.3. Prompt the model to carry out chain-of-thought: Provide chain-of-thought answers in the examples to prompt the model to solve problems in a similar way, gradually show the analysis of the problem, and finally give the answer.

[0066] 4.4. Evaluate the effect of chain-of-thought: Compare the effect of directly giving the answer prompt and adding the chain-of-thought prompt on the test set, to prove that the former can only answer simple questions, while the latter can gradually analyze more complex problems through chain-of-thought, thus achieving better results.

[0067] In this step, by using the above 4 prompting methods, the original first sample prompt templates corresponding to each prompting method can be constructed, so that the original first sample prompt set can be constructed based on each original first sample prompt template.

[0068] Step S204, based on the original first sample prompt set, construct the target first sample prompt set to complete the construction of the automatic prompting project;

[0069] In the specific implementation process of this step, the target first sample prompt set can be generated based on the original first sample prompt set in the local of the first node using the target model. The target model can be a pre-trained distributed large language model. That is, the original first sample prompt set can be used as input to automatically generate prompt instructions using the distributed large language model, and the best candidate prompt instructions can be selected through filtering and searching methods to obtain the target first sample prompt set PE nodei , the steps are as follows:

[0070] 1. Based on the instruction examples in the original first sample prompt set, let the distributed large language model generate automatic instructions. For example, for the following prompt template, give the input and output examples to guide the model to generate instructions: [Question] Generate input and output instruction pairs according to instruction examples. The instruction is...

[0071] 2. Given a training set Dtrain, set the scoring function to select the instruction that can get the highest score in the training set among the automatically generated candidate instructions. In this embodiment, the scoring function can be implemented using accuracy, language model likelihood value, and other machine learning evaluation indicators.

[0072] 3. Iterative search using Monte Carlo method: further search for better instructions by prompting the model to generate variants of the second-step instructions. For example, this prompt is used to generate approximate instructions: [Prompt] Generate a variant of the following instruction while keeping the semantic meaning unchanged. Input:... Output:...

[0073] In step S205, the target second sample prompt set sent by each second node is received for synchronization of the prompting engineering.

[0074] In the implementation process, the first node can receive the target second sample prompt set PE nodei sent by each second node, and determine the prompt subset PE fl for model training based on the target second sample prompt set of each second node and the target first sample prompt set of the first node, and the prompt set directory D_PE fl ’; then, the target first sample prompt subset PE nodei ’ is determined from the target first sample prompt set PE nodei based on the prompt set directory D_PE fl ’, and the prompt set directory D_PE fl ’ is sent to the second node, so that the second node determines the target second sample prompt subset PE nodei ’ from the target second sample prompt set PE nodei based on the prompt set directory D_PE fl ’.

[0075] Specifically, after the second node i forms the prompting engineering instruction set / target second sample prompt set P Enodei , the target second sample prompt set can be sent to the first node, so that the first node can receive the target second sample prompt set sent by each second node, and then generate the total sample prompt set PE fl ={PE node1 , PE node2 , …} based on each target second sample prompt set and the target first sample prompt set of the first node. Then, the prompt subset PE fl ’ used for this training is selected from the total sample prompt set PE fl , and the prompt set directory corresponding to the prompt subset PE fl ’ is D_PE fl ’. Finally, the first node synchronously sends the prompt set directory D_PE fl ’ and the pre-trained model used for this training to each second node, so that each second node determines the target second sample prompt subset PE nodei from the target second sample prompt set PE node_i based on the prompt set directory D_PE fl ’.'.

[0076] In this embodiment, all nodes involved in this task will be prompted according to the instruction set directory D_PE fl ’. The selected part PE nodei ’ of the instruction set PE nodei ’ saved locally is used as training data, and batch training is performed. The advantage of this process is that the training data from the computing nodes meets the screening criteria of the dominant party, and the security is higher. Since the first node feedback process transmits the instruction set directory, the communication pressure is smaller, and the possibility of being attacked and stolen is lower; the distributed computing mode flexibly uses the computing resources of each node, and the system efficiency is high.

[0077] Step S206, model training is performed on the initial large language model;

[0078] In this step, the first node can perform model training on the initial large language model based on the target first sample prompt subset PE nodeii ’ to obtain the initial first training parameter;

[0079] In this step, the initial large language model / pre-training model of the large language model in this embodiment can use Chinese-LLaMA, OpenChineseLLaMA, and Ziya-LLaMA, and other Chinese base models with the same function. By default, Chinese LLaMA is used.

[0080] Step S207, receiving the initial second training parameters sent by each second node, and performing parameter aggregation based on each initial second training parameter and the initial first training parameter to obtain the initial aggregation parameter;

[0081] In this step, all nodes involved in this task complete a local training and send the training result to the first node, and the first node reads all parameters and performs parameter aggregation. In this embodiment, the parameter aggregation method can use arithmetic mean, instruction quantity weighted average, and node weighted average, etc. By default, the arithmetic mean is used.

[0082] Step S208, sending the initial aggregation parameter to each second node for the next round of model, and receiving the current second training parameter sent by each second node until the predetermined training condition is met, stopping model training, taking the current aggregation parameter as the target aggregation parameter, and obtaining the target large language model.

[0083] In this step, the predetermined training condition can be that the number of training rounds reaches a predetermined round threshold, or the error of the current aggregation parameter is less than a predetermined error threshold.

[0084] Step S209, adaptive selection of nodes;

[0085] In this step, the first node can match the target problem with the first data directory corresponding to the first node and the second data directory corresponding to each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node. Specifically, the first node can match the target problem with the total data directory to determine the matched first data directory and / or each second data directory from the total data directory, so as to determine the corresponding node as the main prediction node and each auxiliary prediction node based on the matched first data directory and / or each second data directory.

[0086] That is, when the customer inputs the target problem Q to a certain computing node of the distributed large language model, first, according to the key elements in the customer input, the ability of the distributed large language model is matched with the node data directory to adaptively select the computing node to be called in this dialogue. For example, the input is [question] now there is an input: Q. Please judge which knowledge in the cooperative network total data directory D fl is needed to use, so as to determine the main prediction node and the auxiliary prediction node from each node based on the matching of the target problem and the total data directory D fl , so as to realize adaptive selection of the prediction node.

[0087] Step S210, answer prediction;

[0088] In this step, when the first node predicts the answer, it is divided into the following three cases:

[0089] Case one, when the first node is not determined as the main prediction node or the auxiliary prediction node, the target problem is sent to the main prediction node and each auxiliary prediction node;

[0090] Case two, when the first node is determined as the auxiliary prediction node, the auxiliary prompt word is constructed based on the target problem, the auxiliary prompt word is sent to the main prediction node, and the target problem is sent to the main prediction node and each remaining auxiliary prediction node;

[0091] Case three, when the first node is determined as the main prediction node, the main prompt word is constructed based on the target problem, the target problem is sent to each auxiliary prediction node, and the auxiliary prediction node sends the auxiliary prompt word constructed based on the target problem, the target answer corresponding to the target problem is predicted and output by using the target large language model obtained by pre-training based on the main prompt word and each auxiliary prompt word.

[0092] In this embodiment, the main prediction node can adaptively select the corresponding target prompt mode using the capabilities of the distributed large language model according to the characteristics of the input question Q, that is, the main prediction node can generate a total prompt word based on the main prompt word and each auxiliary prompt word, and then generate a prompt template for answer prediction based on the total prompt word and the target prompt mode. For example, the input is [question] now there is an input: Q. For this input Q, determine which of the zero-shot prompt mode, small-sample prompt mode, knowledge base prompt mode, and thought chain prompt mode is the target prompt scheme, and then combine the total prompt word to output the target answer using the target large language model. For example, the input is [question], now there is an input: total prompt word Q k Please use the method of thought tree to answer.

[0093] The answer prediction method based on distributed learning in this embodiment determines the main prediction node for performing answer prediction and the auxiliary prediction node for performing auxiliary prediction from each second node and the first node, and then generates a main prompt word using the main prediction node and an auxiliary prompt word using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word using the target large language model, and helps to improve the accuracy of answer prediction.

[0094] Another embodiment of the present application provides an answer prediction method based on distributed learning, which can be applied to a second node / participant / computing node of federated learning, as shown in Figure 2 The method in this embodiment includes the following steps:

[0095] Step S301, send the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directory of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node;

[0096] In the specific implementation process, the second node can pre-build a second knowledge base based on predetermined business data, and create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node. At the same time, the first node also builds a first knowledge base based on predetermined local business data. Therefore, when the first node receives each second data directory, it can match the target question based on the first data directory and each second data directory to determine the nodes matching the target question from each second node and the first node participating in federated learning as the main prediction node and each auxiliary prediction node.

[0097] Step S302, when the second node is determined as an auxiliary prediction node, receiving the target question sent by the first node, constructing an auxiliary prompt word based on the target question, and sending the auxiliary prompt word to the main prediction node.

[0098] In this step, when the second node is determined as an auxiliary prediction node, the second node can receive the target question sent by the first node, and then generate / construct an auxiliary prompt word based on the target question according to the second database locally stored in the second node. Thus, the main prediction node can predict the target answer corresponding to the target question by using the target large language model pre-trained based on the main prompt word constructed by the main prediction node and each auxiliary prompt word, and output the target answer.

[0099] Step S303, when the second node is determined as a main prediction node, receiving the target question sent by the first node, constructing a main prompt word based on the target question, and receiving auxiliary prompt words sent by each auxiliary prediction node based on the target question, predicting the target answer corresponding to the target question by using the target large language model pre-trained based on the main prompt word and each auxiliary prompt word, and outputting the target answer.

[0100] In this step, when the second node is determined as a main prediction node, that is, the answer prediction is performed locally in the second node. Thus, the second node can receive the target question sent by the first node, and then generate / construct a main prompt word based on the target question according to the second database locally stored in the second node, and receive auxiliary prompt words sent by each auxiliary prediction node. Further, the second node can predict the target answer corresponding to the target question by using the target large language model pre-trained based on the main prompt word and each auxiliary prompt word, and output the target answer.

[0101] The answer prediction method based on distributed learning in the embodiment can determine a main prediction node for performing answer prediction and auxiliary prediction nodes for performing auxiliary prediction from each second node and the first node. Subsequently, the main prediction node can generate a main prompt word and the auxiliary prediction nodes can generate auxiliary prompt words, so that the generation of the prompt words is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word by using the target large language model, and is conducive to improving the accuracy of answer prediction.

[0102] On the basis of the above embodiment, another embodiment of the present application provides an answer prediction method based on distributed learning, applied to a second node. The overall process of answer prediction in the embodiment is as follows:

[0103] Step S401, installation and initialization of a distributed large language model.

[0104] In this step, the leading party of joint modeling (financial infrastructure, first node) deploys the components of the multi-center distributed large language model collaboration network in its local data center, and sends the initial large language model and the corresponding components to each second node (slave computing node). Thus, the participants of joint modeling (financial institutions, second nodes) will also deploy the components of the multi-center distributed large language model collaboration network in their local data centers. The distributed machine learning framework in this embodiment can be implemented using Apache Spark MLlib, PyTorchtorch.distributed, and TensorFlowtf.distribute and other software products with the same function.

[0105] In step S402, the local second data directory is sent to the first node for the first node to generate a collaboration network total data directory;

[0106] In this step, the first node can construct a first knowledge base based on predetermined business data and create a first data directory corresponding to the first knowledge base. Similarly, each second node can construct a second knowledge base based on predetermined business data and create a second data directory corresponding to the second knowledge base, and then send the second data directory to the first node. Thus, the first node can generate a total data directory for determining the main prediction node and each auxiliary prediction node based on the first data directory and each second data directory.

[0107] Specifically, the first node and each second node are connected with their local bond business database, and the local business data is segmented and vectorized using natural language processing technology as a knowledge base, and the knowledge base of node i is marked as K nodei , and K nodei The file name of the internal data file used is arranged as a node data directory D nodei (including the first data directory and each second data directory), and the second node sends the second data directory to the first node for synchronization. The first node collects all the data directories fed back by the nodes to generate a collaboration network total data directory D fl ={D node1 ,D node2 ,……}. That is, after receiving each second data directory, the first node can construct a total data directory based on each second data directory and the first data directory, and then match the target problem with the total data directory, so as to determine the first data directory and / or several second data directories matching the target problem from the total data directory, and then determine the corresponding node as the main prediction node or auxiliary prediction node according to the matching result.

[0108] In step S403, the prompt engineering of the multi-center distributed large language model is constructed to obtain an original second sample prompt set;

[0109] In this step, the second node can pre-build the original second sample prompt set. The specific process is: based on each prompt method, generate a plurality of original second sample prompt templates corresponding to each prompt method; wherein the prompt method includes any one or several of the following: zero sample prompt method, small sample prompt method, knowledge base prompt method and thinking chain prompt method; based on each original first sample prompt template, build an original second sample prompt set.

[0110] That is, the second node can design its own business-related prompt project, and use the following several prompt methods to complete several types of sample preparation for the prompt project according to the needs of different tasks, in order to build and obtain the original second sample prompt set.

[0111] 1. Zero sample prompt method: the institution does not provide task result related demonstration, directly prompts the language model to give task related answers. For example, the input is [question] what materials are needed to open a bank bond market settlement account.

[0112] 2. Small sample prompt method: the institution provides a small amount of prompt examples, such as task instructions, etc. For example, the input is [question] translate Chinese into English. Bond -> bond; assessment -> assessment; management -> management; The European Central Bank has maintained the three key interest rates unchanged for the fourth consecutive time, and traders have increased their expectations for the European Central Bank to cut interest rates by 100 basis points this year ->.

[0113] 3. Knowledge base prompt method: the institution provides a knowledge base related to the question and the key words of the question, and uses the context content of the knowledge base to form prompt words. For example, the input is [question]: what materials are needed to open a bank bond market settlement account; [knowledge base]: documents such as “Notice on Matters Related to Bank Bond Market Counter Business”. After the computing node performs text extraction and semantic understanding on the knowledge base, the most similar document fragments are calculated with the input question to form prompt words. The vector distance measurement index described in this application can be realized using cosine similarity, dot product, Hamming distance and other vector distance measurement methods.

[0114] 4. Thinking chain prompt method: the institution introduces chain thinking steps in the prompt project for complex and logical problems to enhance the processing capacity of the large model for complex problems, using the following steps:

[0115] 4.1. Build samples: manually write a small number of samples of tasks that need to be solved, which contain input questions and their answers, as well as the chain thinking process in the solution process.

[0116] 4.2. Provide examples of chain thinking: Add examples of chain thinking for the corresponding answers to the sample questions so that the model can learn how to analyze the problem first, then carry out chain thinking step by step, and finally arrive at the answer.

[0117] 4.3. Prompt the model to use chain thinking: Provide chain thinking answers in the examples to prompt the model to solve the problem in a similar way, gradually showing the analytical thinking of the problem, and finally giving the answer.

[0118] 4.4. Evaluate the effectiveness of chain thinking: Compare the effects of directly providing the answer to the chain thinking hints on the test set. The results show that the former can only answer simple questions, while the latter can gradually analyze more complex questions through chain thinking, thus achieving better results.

[0119] In this step, by adopting the above four prompting methods, we can construct the original second sample prompting template corresponding to each prompting method, and thus construct the original second sample prompting set based on each original second sample prompting template.

[0120] Step S404: Based on the original second sample prompt set, construct the target second sample prompt set to complete the construction of the automatic prompting project;

[0121] In this step, the target second sample prompt set (PE) can be generated based on the original second sample prompt set located locally on the second node, using the target model. The target model can be a pre-trained distributed large language model. That is, the original first sample prompt set can be used as input, and the distributed large language model can automatically generate prompt commands. The best candidate prompt commands are selected through filtering and searching methods to obtain the target second sample prompt set (PE). nodei The steps are as follows:

[0122] 1. Based on the instruction examples in the original second sample prompt set, enable the distributed large language model to generate automatic instructions. For example, given the following prompt template with input and output examples, guide the model to generate instructions: [Problem] Generate an input-output instruction pair based on the instruction example. The instruction is...

[0123] 2. Given a training set Dtrain, define a scoring function to select the instruction that gives the highest score to the samples in the training set from the automatically generated candidate instructions. In this embodiment, the scoring function can be implemented using accuracy, language model likelihood, and other machine learning evaluation metrics.

[0124] 3. Iterative search using Monte Carlo method: Prompt the model to generate a variant of the second-step instruction, further searching for a better instruction. For example, this prompt can be used to generate an approximate instruction: [Problem] Generate a variant of the following instruction while maintaining its semantic meaning. Input: ... Output: ...

[0125] Step S405: Send the target second sample cue set to the first node to synchronize the cue project.

[0126] In the specific implementation of this step, the second node can use the target second sample hint set PE nodei Send to the first node so that the first node can determine a hint set directory for the hint subset used for initial model training based on the target second sample hint set of each second node and the target first sample hint set of the first node; receive the hint set directory sent by the first node, and based on the hint set directory D_PE fl 'From the local target second sample cue set PE nodei The second sample hint subset PE of the target is determined in the middle. fl '.

[0127] That is, the second node i forms the prompt engineering instruction set / target second sample prompt set P. Enodei Then, the target second sample cue set P can be... Enodei The first node receives the target second sample hint sets sent by each second node, and then generates the total sample hint set PE based on each target second sample hint set and the target first sample hint set locally on the first node. fl ={PE node1 PE node2 ,……}. Then from the total sample cue set PE fl Select the subset of prompts PE used in this training. fl ', and determine the subset PE of the prompts. fl The corresponding prompt set directory D_PE fl The first node will ultimately suggest the directory as D_PE. fl The information, including the pre-trained model used in this training, is simultaneously sent to each second node so that each second node can use the prompt set directory D_PE. fl 'From the target second sample cue set PE nodei The second sample hint subset PE of the target is determined in the middle. nodei '.

[0128] In this embodiment, all nodes involved in this task will be based on the prompt set directory D_PE fl ', will store the locally saved instruction set PE nodei Selected PE nodeiAs training data, training is carried out in batches. The advantage of this process is that the training data from the computing nodes meets the screening criteria of the dominant party, and the security is higher. Since the first node feedback process transmits a set of instructions, the communication pressure is small, and the possibility of attack and theft is low; the distributed computing mode flexibly uses the computing resources of each node, and the system efficiency is high.

[0129] Step S406, model training is performed on the initial large language model;

[0130] In this step, the second node can prompt the second sample subset PE nodeii The initial large language model is trained to obtain the initial second training parameter;

[0131] In this step, the initial large language model / pre-training model of the large language model in the embodiment can use Chinese-LLaMA, OpenChineseLLaMA, and Ziya-LLaMA and other Chinese base models with the same function. By default, Chinese LLaMA is used.

[0132] Step S407, the initial second training parameter is sent to the first node, so that the first node performs parameter aggregation based on the initial second training parameter of each second node and the initial first training parameter of the first node, and obtains the initial aggregation parameter;

[0133] In this step, all nodes involved in this task complete a local training and send the training result to the first node, and the first node reads all parameters and performs parameter aggregation. In the embodiment, the parameter aggregation method can use arithmetic mean, instruction quantity weighted average, and node weighted average, and by default, the arithmetic mean is used.

[0134] Step S408, receiving the initial aggregation parameter sent by the first node, performing next round of model training based on the initial aggregation parameter, sending the current second training parameter obtained by training to the first node, until the predetermined training condition is met, stopping the model training, taking the received current aggregation parameter as the target aggregation parameter, and obtaining the target large language model.

[0135] In this step, the predetermined training condition can be that the training round number reaches a predetermined round number threshold, or the current aggregation parameter error is less than a predetermined error threshold.

[0136] Step S409, answer prediction;

[0137] Before the answer prediction is performed, the first node determines the primary prediction node and the secondary prediction nodes from the first node and the second nodes. That is, the first node can match the target question with the first data directory corresponding to the first node and the second data directory corresponding to the second nodes to determine the primary prediction node and the secondary prediction nodes from the first node and the second nodes. Specifically, the first node can match the target question with the total data directory to determine the matched first data directory and / or the second data directory from the total data directory, so as to determine the corresponding node as the primary prediction node and the secondary prediction nodes based on the matched first data directory and / or the second data directory. That is, when the customer inputs the target question Q to a certain computing node of the distributed large language model, first, according to the key elements in the customer input, the computing node is adaptively selected by matching the capabilities of the distributed large language model with the node data directory. For example, the input is [question] now there is an input: Q. Please judge which knowledge in the cooperative network total data directory D fl is needed to use, so as to determine the primary prediction node and the secondary prediction nodes from the nodes based on the matching of the target question with the total data directory D fl , so as to realize adaptive selection of the prediction node.

[0138] In this step, when the second node is determined as the secondary prediction node, the second node can receive the target question sent by the first node, construct the secondary prompt word based on the target question, and send the secondary prompt word to the primary prediction node. That is, the second node can generate / construct the secondary prompt word based on the target question according to the second database of the second node, and then send the secondary prompt word to the primary prediction node. Thus, the primary prediction node can predict and output the target answer corresponding to the target question by using the target large language model obtained by pre-training based on the primary prompt word constructed by itself and the secondary prompt word.

[0139] In this step, when the second node is determined as the primary prediction node, the second node can receive the target question sent by the first node, construct the primary prompt word based on the target question, and receive the secondary prompt word constructed based on the target question and sent by the secondary prediction node, and predict and output the target answer corresponding to the target question by using the target large language model obtained by pre-training based on the primary prompt word and the secondary prompt word.

[0140] In this embodiment, the main prediction node can adaptively select the corresponding target prompting mode using the capability of the distributed large language model according to the characteristics of the input question Q, that is, the main prediction node can generate a total prompt word based on the main prompt word and each auxiliary prompt word, and then generate a prompt template for answer prediction based on the total prompt word and the target prompting mode. For example, the input is [question] now there is an input: Q. For this input Q, determine which of the zero-shot prompting mode, small-sample prompting mode, knowledge base prompting mode, and thought chain prompting mode is the target prompting scheme, and then combine the total prompt word to output the target answer using the target large language model. For example, the input is [question], now there is an input: total prompt word Q k Please answer in the way of a thought tree.

[0141] The answer prediction method based on distributed learning in this embodiment determines the main prediction node for performing answer prediction and the auxiliary prediction node for performing auxiliary prediction from the second nodes and the first node, and then generates the main prompt word using the main prediction node and generates the auxiliary prompt word using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, laying a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word using the target large language model, and improving the accuracy of answer prediction.

[0142] Another embodiment of the present application provides an answer prediction device based on distributed learning, as shown in Figure 3 The device comprises:

[0143] The first receiving module is configured to receive the second data directory sent by each second node, and match the target question with the first data directory corresponding to the first node and each second data directory, so as to determine the main prediction node and each auxiliary prediction node from the first node and each second node;

[0144] The first sending module is configured to send the target question to the main prediction node and each auxiliary prediction node when the first node is not determined as the main prediction node or the auxiliary prediction node;

[0145] The first constructing module is configured to, when the first node is determined as the auxiliary prediction node, construct an auxiliary prompt word based on the target question, send the auxiliary prompt word to the main prediction node, and send the target question to the main prediction node and the remaining auxiliary prediction nodes;

[0146] The first prediction module is configured to, when the first node is determined as the main prediction node, construct a main prompt word based on the target question, send the target question to each auxiliary prediction node, receive the auxiliary prompt word constructed based on the target question and sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word using the target large language model obtained by pre-training.

[0147] In the implementation process of the embodiment, the answer prediction device based on distributed learning further comprises a first creating module, which is configured to:

[0148] construct a first knowledge base based on predetermined business data, and create a first data directory corresponding to the first knowledge base.

[0149] In the implementation process of the embodiment, the answer prediction device based on distributed learning further comprises a first generating module, an aggregating module, and a first training module.

[0150] The first generating module is configured to generate a target first sample prompt subset based on a first node local original first sample prompt set.

[0151] The first training module is configured to perform model training on an initial large language model based on the target first sample prompt subset to obtain an initial first training parameter.

[0152] The aggregating module is configured to receive the initial second training parameters sent by the second nodes, and perform parameter aggregation based on the initial second training parameters and the initial first training parameter to obtain an initial aggregated parameter.

[0153] The first sending module is further configured to send the initial aggregated parameter to the second nodes for next round model training, and receive current second training parameters sent by the second nodes based on the aggregating module until a predetermined training condition is met, stop the model training, take the current aggregated parameter as a target aggregated parameter, and obtain a target large language model.

[0154] In the implementation process of the embodiment, the first generating module specifically comprises:

[0155] The first generating unit is configured to generate a target first sample prompt set based on a first node local original first sample prompt set by using a target model.

[0156] The prompt set receiving unit is configured to receive target second sample prompt sets sent by the second nodes, and determine a prompt set directory of a prompt subset for model training based on the target second sample prompt sets of the second nodes and the target first sample prompt set of the first node.

[0157] The first determining unit is configured to determine a target first sample prompt subset from the target first sample prompt set based on the prompt set directory, and send the prompt set directory to the second nodes so that the second nodes determine a target second sample prompt subset from the target second sample prompt set based on the prompt set directory.

[0158] In the implementation process of the embodiment, the answer prediction device based on distributed learning further includes a first prompt set construction module for pre-constructing the original first sample prompt set, and the prompt set construction module is configured to generate a plurality of original first sample prompt templates corresponding to each prompt mode based on the prompt mode, wherein the prompt mode includes any one or several of the following: zero sample prompt mode, small sample prompt mode, knowledge base prompt mode, and thinking chain prompt mode; and construct the original first sample prompt set based on the original first sample prompt templates.

[0159] The answer prediction device based on distributed learning in the application determines the main prediction node for predicting answers and the auxiliary prediction node for assisting in prediction from the second nodes and the first node, and then generates the main prompt word using the main prediction node and generates the auxiliary prompt word using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word using the target large language model, and helps to improve the accuracy of answer prediction.

[0160] Another embodiment of the application provides an answer prediction device based on distributed learning, as shown in Figure 4 , which includes

[0161] The second sending module 21 is configured to send the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directory of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node.

[0162] The second construction module 22 is configured to, when the second node is determined as an auxiliary prediction node, receive the target question sent by the first node, construct an auxiliary prompt word based on the target question, and send the auxiliary prompt word to the main prediction node.

[0163] The second prediction module 23 is configured to, when the second node is determined as a main prediction node, receive the target question sent by the first node, construct a main prompt word based on the target question, and receive the auxiliary prompt word constructed based on the target question sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question using the target large language model trained in advance based on the main prompt word and each auxiliary prompt word.

[0164] In the implementation process of the embodiment, the answer prediction device based on distributed learning further includes a second creation module, and the second creation module is configured to construct a second knowledge base based on predetermined business data, and create a second data directory corresponding to the second knowledge base based on the second knowledge base.

[0165] In the implementation process of the embodiment, the answer prediction device based on distributed learning further includes a second generation module, a second training module, and a second receiving module.

[0166] The second generation module is configured to generate a target second sample prompt subset based on a local original second sample prompt set of the second node.

[0167] The second training module is configured to perform model training on the initial large language model based on the target second sample prompt subset to obtain initial second training parameters.

[0168] The second sending module is further configured to send the initial second training parameters to the first node, so that the first node performs parameter aggregation based on the initial second training parameters of each second node and the initial first training parameters of the first node to obtain initial aggregated parameters.

[0169] The second receiving module is further configured to receive the initial aggregated parameters sent by the first node, perform next round of model training based on the initial aggregated parameters using the second training module, send the current second training parameters obtained by training to the first node based on the second sending module, and stop the model training when a predetermined training condition is met, take the received current aggregated parameters as target aggregated parameters, and obtain a target large language model.

[0170] In the implementation process of the embodiment, the second generation module specifically includes:

[0171] The second generation unit is configured to generate a target second sample prompt set based on a local original second sample prompt set of the second node using a target model.

[0172] The prompt set sending unit is configured to send the target second sample prompt set to the first node, so that the first node determines a prompt set directory of a prompt subset used for initial model training based on the target second sample prompt set of each second node and the target first sample prompt set of the first node.

[0173] The second determination unit is configured to receive the prompt set directory sent by the first node, and determine a target second sample prompt subset from the local target second sample prompt set based on the prompt set directory.

[0174] In the implementation process of the embodiment, the answer prediction device based on distributed learning further includes a second prompt set construction module for pre-construction of an original second sample prompt set.

[0175] The original second sample prompt templates corresponding to the prompt modes are generated respectively based on the prompt modes, wherein the prompt modes include any one or several of the following: zero sample prompt mode, small sample prompt mode, knowledge base prompt mode, and thought chain prompt mode; and the original second sample prompt set is constructed based on the original second sample prompt templates.

[0176] The answer prediction device based on distributed learning in the present application determines the main prediction node for performing answer prediction and the auxiliary prediction node for performing auxiliary prediction from the second nodes and the first node, and then generates the main prompt word using the main prediction node and generates the auxiliary prompt word using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word using the target large language model, and helps to improve the accuracy of answer prediction.

[0177] Another embodiment of the present application provides an electronic device, as shown in Figure 5 The electronic device includes at least a memory 1 and a processor 2, the memory 1 stores a computer program, and the processor 2 implements the following method steps when executing the computer program on the memory 1:

[0178] Step one, receiving the second data directory sent by each second node, and matching the target question with the first data directory corresponding to the first node and each second data directory to determine the main prediction node and each auxiliary prediction node from the first node and each second node;

[0179] Step two, when the first node is not determined as the main prediction node or the auxiliary prediction node, sending the target question to the main prediction node and each auxiliary prediction node;

[0180] When the first node is determined as the auxiliary prediction node, constructing the auxiliary prompt word based on the target question, sending the auxiliary prompt word to the main prediction node, and sending the target question to the main prediction node and each remaining auxiliary prediction node;

[0181] When the first node is determined as the main prediction node, constructing the main prompt word based on the target question, sending the target question to each auxiliary prediction node, receiving the auxiliary prompt word constructed based on the target question sent by each auxiliary prediction node, and predicting the target answer corresponding to the target question using the target large language model trained in advance based on the main prompt word and each auxiliary prompt word and outputting.

[0182] Alternatively, the following method steps are implemented:

[0183] Step one, sending the local second data directory to the first node, so that the first node matches based on the first data directory and the second data directory of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node;

[0184] Step two, when the second node is determined as an auxiliary prediction node, receiving the target question sent by the first node, constructing an auxiliary prompt word based on the target question, and sending the auxiliary prompt word to the main prediction node;

[0185] When the second node is determined as a main prediction node, receiving the target question sent by the first node, constructing a main prompt word based on the target question, and receiving the auxiliary prompt word sent by each auxiliary prediction node based on the target question, predicting the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word using the target large language model obtained by pre-training and outputting.

[0186] The specific implementation process of the above method steps can be referred to the embodiments of any of the above answer prediction methods based on distributed learning image processing methods, which will not be repeated here.

[0187] The device in the present application determines the main prediction node for answer prediction and the auxiliary prediction node for auxiliary prediction from each second node and the first node, and then generates a main prompt word using the main prediction node and an auxiliary prompt word using the auxiliary prediction node, so that the generation of the prompt word is more comprehensive and accurate, which lays a foundation for subsequent accurate answer prediction based on the main prompt word and each auxiliary prompt word using the target large language model, and helps to improve the accuracy of answer prediction.

[0188] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application, the protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the spirit and protection scope of the present application, and such modifications or equivalent replacements are also regarded as falling within the protection scope of the present application.

Claims

1. A method for answer prediction based on distributed learning, applied to the first node, characterized in that, include: Receive the second data directory sent by each second node, and match the target problem with the first data directory corresponding to the first node and each second data directory to determine the main prediction node and each auxiliary prediction node from the first node and each second node. If the first node is not determined to be the primary or secondary prediction node, the target problem is sent to the primary prediction node and each secondary prediction node. When the first node is identified as an auxiliary prediction node, auxiliary prompt words are constructed based on the target question, and the auxiliary prompt words are sent to the main prediction node. The target question is also sent to the main prediction node and the remaining auxiliary prediction nodes. When the first node is determined to be the main prediction node, a main prompt word is constructed based on the target question. The target question is sent to each auxiliary prediction node, and the auxiliary prompt words constructed based on the target question are received from each auxiliary prediction node. Based on the main prompt word and each auxiliary prompt word, the target answer corresponding to the target question is predicted and output using the pre-trained target large language model. The step of receiving the second data directories sent by each second node and matching the target problem with the first data directory corresponding to the first node and each second data directory to determine the main prediction node and each auxiliary prediction node from the first node and each second node specifically includes: The first node constructs a first knowledge base based on predetermined business data and creates a first data directory corresponding to the first knowledge base; The first node receives the second data directory sent by each second node, and the second data directory corresponds to the second knowledge base built locally by the second node; The first node determines the main prediction node and the auxiliary prediction node from each node based on the degree of matching between the first data directory and the second data directory in each node and the target problem.

2. The method as described in claim 1, characterized in that, Before matching the target problem with the first data directory corresponding to the first node and the second data directories corresponding to each second node, the method further includes: A first knowledge base is constructed based on predetermined business data, and a first data directory corresponding to the first knowledge base is created.

3. The method as described in claim 1, characterized in that, The method further includes: training the initial large language model, specifically including: Based on the original first sample hint set local to the first node, generate a target first sample hint subset; The initial large language model is trained based on the target first sample hint subset to obtain the initial first training parameters; Receive the initial second training parameters sent by each second node, and aggregate the parameters based on each initial second training parameter and the initial first training parameter to obtain the initial aggregated parameters; The initial aggregation parameters are sent to each second node for each second node to perform the next round of modeling. The current second training parameters sent by each second node are received until the predetermined training conditions are met. Then, the model training is stopped, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

4. The method as described in claim 3, characterized in that, The process of generating a target first sample hint subset based on the original first sample hint set local to the first node specifically includes: Based on the original first sample hint set local to the first node, the target first sample hint set is generated using the target model; Receive the target second sample cue set sent by each second node, and determine the cue set directory for the cue subset used for model training based on the target second sample cue set of each second node and the target first sample cue set of the first node; Based on the prompt set directory, a subset of the target first sample prompts is determined from the target first sample prompt set, and the prompt set directory is sent to the second node so that the second node can determine the subset of the target second sample prompts from the target second sample prompt set based on the prompt set directory.

5. The method as described in claim 3, characterized in that, The method further includes: pre-constructing an original first sample cue set, specifically including: Based on each prompting method, several original first sample prompting templates corresponding to each prompting method are generated; wherein the prompting methods include any one or more of the following: zero sample prompting method, small sample prompting method, knowledge base prompting method, and mind chain prompting method; Based on each original first sample prompt template, construct the original first sample prompt set.

6. A method for predicting answers based on distributed learning, applied to the second node, characterized in that, include: The local second data directory is sent to the first node so that the first node can match the first data directory and the second data directories of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node. When the second node is identified as the auxiliary prediction node, it receives the target question sent by the first node, constructs auxiliary prompts based on the target question, and sends the auxiliary prompts to the main prediction node. When the second node is determined to be the main prediction node, it receives the target question sent by the first node, constructs the main prompt word based on the target question, and receives the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node. Based on the main prompt word and each auxiliary prompt word, it uses the pre-trained target large language model to predict the target answer corresponding to the target question and outputs it. The step of sending the local second data directory to the first node, so that the first node can match the first data directory with the second data directories of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node, specifically includes: The second node pre-builds a second knowledge base based on predetermined business data and creates a second data directory corresponding to the second knowledge base. It then sends the second data directory to the first node, which builds the first knowledge base based on predetermined local business data. After the first node receives each second data directory, it matches the first data directory and each second data directory with the target question, and determines the node that matches the target question from each second node and the first node participating in federated learning as the main prediction node and each auxiliary prediction node.

7. The method as described in claim 6, characterized in that, Before sending the local second data directory to the first node, the method further includes: A second knowledge base is built based on pre-defined business data; A second data directory corresponding to the second knowledge base is created based on the second knowledge base.

8. An answer prediction device based on distributed learning, used to implement the answer prediction method based on distributed learning as described in claim 1, characterized in that, include: The first receiving module is used to receive the second data directory sent by each second node, and match the target problem with the first data directory corresponding to the first node and each second data directory to determine the main prediction node and each auxiliary prediction node from the first node and each second node. The first sending module is used to send the target problem to the main prediction node and each auxiliary prediction node when the first node is not determined to be the main prediction node or the auxiliary prediction node. The first construction module is used to construct auxiliary prompt words based on the target question when the first node is determined to be an auxiliary prediction node, send the auxiliary prompt words to the main prediction node, and send the target question to the main prediction node and the remaining auxiliary prediction nodes. The first prediction module is used to construct a main prompt word based on the target question when the first node is determined to be the main prediction node, send the target question to each auxiliary prediction node, receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word using a pre-trained target large language model.

9. An answer prediction device based on distributed learning, used to implement the answer prediction method based on distributed learning as described in claim 6, characterized in that, include: The second sending module is used to send the local second data directory to the first node, so that the first node can match the first data directory and the second data directories of each second node to determine the main prediction node and each auxiliary prediction node from the first node and each second node. The second construction module is used to receive the target question sent by the first node when the second node is determined to be the auxiliary prediction node, construct auxiliary prompt words based on the target question, and send the auxiliary prompt words to the main prediction node. The second prediction module is used to receive the target question sent by the first node when the second node is determined to be the main prediction node, construct the main prompt word based on the target question, receive the auxiliary prompt words constructed based on the target question sent by each auxiliary prediction node, and predict and output the target answer corresponding to the target question based on the main prompt word and each auxiliary prompt word and using the pre-trained target large language model.

10. An electronic device, characterized in that, It includes at least a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the answer prediction method based on distributed learning as described in any one of claims 1-5 or 6-7.

Citation Information

Patent Citations

  • Search cue word determination method and system and information processing method

    CN113032819A

  • Large model cue word intelligent routing method, device and equipment and storage medium

    CN119691137A