Data processing method, device and equipment

By training a large language model through multiple tasks and combining historical interaction data and thought chains, the problem of the model generating illusions in complex tasks was solved, thereby improving the credibility of feedback data and the accuracy of business processing.

CN120806151APending Publication Date: 2025-10-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510914359.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Large language models are prone to hallucinations when processing complex tasks, resulting in low credibility of the generated feedback data and affecting the accuracy of subsequent business processing.

Method used

By acquiring thought chains from historical interaction data and feedback data, multi-task training is performed using a large language model, including determining the prediction feedback data and thought chains, until the model converges, forming a pre-defined large language model.

Benefits of technology

It improves the reasoning ability of large language models and the accuracy of feedback data, enhances the accuracy of subsequent business processing, and avoids increasing the time spent on reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806151A_ABST
    Figure CN120806151A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, device and equipment, and the method comprises the steps: obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; determining first prediction feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information by using the large language model; determining a prediction thinking chain based on the historical interaction data and second preset prompt information by using the large language model; and according to the historical feedback data, the first prediction feedback data, the historical thinking chain and the prediction thinking chain, determining whether the large language model is converged or not, and when it is determined that the large language model is not converged, continuing to train the large language model until the large language model is converged. And obtaining a preset large language model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a data processing method, device and equipment. BACKGROUND

[0002] With the rapid development of artificial intelligence, a large language model (LLM) is widely applied to various business processing scenarios due to its strong understanding and logical reasoning capability. In an interaction scenario, for example, the large language model can be used to generate feedback data for the interaction data input by a user, and then the feedback data can be used for business processing (such as risk detection and processing of subsequent transactions according to the feedback data to protect the privacy data security of the user).

[0003] However, for a relatively complex task (such as a complex data problem in the interaction data input by the user), the large language model is prone to hallucination, and thus gives an incorrect answer (such as low credibility of the generated feedback data), resulting in poor subsequent business processing effect. Therefore, the technical solution provided by the embodiments of the present application improves the processing effect of the large language model to improve the accuracy of subsequent business processing. SUMMARY

[0004] The technical solution provided by the embodiments of the present application improves the processing effect of the large language model to improve the accuracy of subsequent business processing.

[0005] In order to achieve the above technical solution, the embodiments of the present application are implemented as follows: The data processing method provided by the embodiments of the present application comprises the following steps: obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; determining, by using the large language model, first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; determining, by using the large language model, a predicted thinking chain based on the historical interaction data and second preset prompt information, the predicted thinking chain being used to represent a thinking reasoning process of the large language model for generating feedback data corresponding to the historical interaction data; determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, and continuing to train the large language model until the large language model converges to obtain a preset large language model, the preset large language model being used to determine feedback data corresponding to interaction data.

[0006] The embodiment of the specification provides a data processing device, the device comprises: a data acquisition module, used for acquiring historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; a first processing module, used for facilitating the large language model to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; a second processing module, used for utilizing the large language model to determine a predicted thinking chain based on the historical interaction data and second preset prompt information, the predicted thinking chain being used for representing a thinking reasoning process of the large language model generating feedback data corresponding to the historical interaction data; a model training module, used for determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, continuing to train the large language model until the large language model converges in the case of determining that the large language model does not converge, obtaining a preset large language model, the preset large language model being used for determining feedback data corresponding to interaction data.

[0007] The embodiment of the specification provides a data processing device, the device comprises: a data acquisition module, used for acquiring historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; a first processing module, used for facilitating the large language model to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; a second processing module, used for utilizing the large language model to determine a predicted thinking chain based on the historical interaction data and second preset prompt information, the predicted thinking chain being used for representing a thinking reasoning process of the large language model generating feedback data corresponding to the historical interaction data; a model training module, used for determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, continuing to train the large language model until the large language model converges in the case of determining that the large language model does not converge, obtaining a preset large language model, the preset large language model being used for determining feedback data corresponding to interaction data.

[0008] The embodiment of the specification also provides a storage medium for storing computer executable instructions, which, when executed by a processor, implement the following processes: obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; determining, based on the historical interaction data and first preset prompt information, first predicted feedback data corresponding to the historical interaction data by using the large language model; determining, based on the historical interaction data and second preset prompt information, a predicted thinking chain by using the large language model, the predicted thinking chain being used to represent a thinking reasoning process of the large language model for generating feedback data corresponding to the historical interaction data; determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, and continuing to train the large language model until the large language model converges to obtain a preset large language model, the preset large language model being used to determine feedback data corresponding to interaction data.

[0009] The embodiment of the specification also provides a computer program product, comprising a computer program which, when executed by a processor, implements the following processes: obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; determining, based on the historical interaction data and first preset prompt information, first predicted feedback data corresponding to the historical interaction data by using the large language model; determining, based on the historical interaction data and second preset prompt information, a predicted thinking chain by using the large language model, the predicted thinking chain being used to represent a thinking reasoning process of the large language model for generating feedback data corresponding to the historical interaction data; determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, and continuing to train the large language model until the large language model converges to obtain a preset large language model, the preset large language model being used to determine feedback data corresponding to interaction data. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the specification or the prior art, brief introductions will be given to the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments described in the specification, and other drawings can be obtained by those skilled in the art without creative labor under the premise of not paying creative labor. Figure 1 A schematic diagram of a data processing method of the specification; Figure 2 A schematic diagram of a training process of a large language model of the present specification; Figure 3 A schematic diagram of a business processing based on first preset prompt information of the present specification; Figure 4 A schematic diagram of a business processing based on second preset prompt information of the present specification; Figure 5 A schematic diagram of a training process of another large language model of the present specification; Figure 6 A schematic diagram of a determination process of a thought chain prediction of the present specification; Figure 7 A schematic diagram of a data processing apparatus of the present specification; Figure 8 A schematic diagram of a data processing apparatus of the present specification; Figure 9 A schematic diagram of a data processing apparatus of the present specification. DETAILED DESCRIPTION

[0011] The embodiments of the present specification provide a data processing method, apparatus and device.

[0012] In order to enable personnel in the technical field to better understand the technical solutions in the present specification, the technical solutions in the embodiments of the present specification will be described clearly and completely in conjunction with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the scope of protection of the present specification.

[0013] The embodiment of the specification provides a technical solution for improving the processing effect of a large language model to improve the accuracy of subsequent business processing. The large language model (LLM) is widely used in various business processing scenarios due to its powerful understanding and logical reasoning capabilities. For example, in an interactive scenario, the large language model can be used to generate feedback data for user input interactive data. However, for complex tasks (such as complex data problems in user input interactive data), the large language model is prone to hallucination, which may result in incorrect answers (such as low credibility of generated feedback data), leading to poor subsequent business processing results. For example, in a transaction risk detection scenario, due to the low credibility of feedback data generated by the large language model for resource transfer interactive data, the subsequent transaction cannot be accurately detected for risk based on the feedback data, and the data security of the user cannot be guaranteed. In this solution, the historical interactive data, the historical feedback data corresponding to the historical interactive data, and the historical thinking chain corresponding to the historical feedback data are obtained, and the large language model is used to determine the first predicted feedback data corresponding to the historical interactive data based on the historical interactive data and the first preset prompt information. The large language model is used to determine the predicted thinking chain based on the historical interactive data and the second preset prompt information, wherein the predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interactive data. The convergence of the large language model is determined based on the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain. If the large language model is not converged, the large language model is continuously trained until the large language model is converged, and a preset large language model is obtained, wherein the preset large language model can be used to determine feedback data corresponding to interactive data. In this way, through the multi-task learning method (i.e., through the first preset prompt information to determine the first predicted feedback data, and through the second preset prompt information to determine the predicted thinking chain), the model can organically integrate the thinking chain process during the training process, learn clear and correct reasoning logic, and enhance the reasoning ability of the model. In actual application, the first preset prompt information can be used to directly reason using the preset large language model, which can increase the reasoning time of the model compared to the thinking chain reasoning method. That is, the thinking chain process can be learned implicitly to improve the reasoning ability of the preset large language model, while not increasing the reasoning time. On the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved. The specific processing can be referred to the specific content in the following embodiments.

[0014] As Figure 1As shown, the embodiment of the present specification provides a data processing method, the execution subject of the method can be a server, wherein the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server of a financial service or a network shopping service, etc., or a background server of an application program, etc. In the embodiment, the execution subject is taken as the server as an example for detailed description. The method can specifically include the following steps: In step S102, historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thinking chains corresponding to the historical feedback data are obtained.

[0015] The historical interaction data can be interaction data input by a user on an arbitrary interaction platform within a preset model training period (such as the last month, the last three months, etc.), and the historical thinking chains corresponding to the historical feedback data can be thinking chains determined by a preset auditing party for the historical interaction data and the corresponding historical feedback data, or the historical thinking chains can also be thinking chains determined by a preset thinking chain generation model for the historical interaction data and the corresponding historical feedback data, which are passed by the preset auditing party, etc.

[0016] In implementation, taking the historical interaction data as an example, which is interaction data input by a user for a resource transfer service on an interaction platform within the last month, the server can obtain historical feedback data corresponding to the historical interaction data, and then the server can use a preset thinking chain generation model to generate corresponding thinking chains for the historical interaction data and the corresponding historical feedback data, wherein the preset thinking chain generation model can be a model constructed based on a preset deep learning algorithm.

[0017] The server can determine the thinking chains generated by the preset thinking chain generation model as the historical thinking chains corresponding to the historical feedback data, or the server can also send the generated thinking chains, the historical interaction data and the corresponding historical feedback data to a preset auditing party for screening, and determine the thinking chains screened by the preset auditing party as the historical thinking chains corresponding to the historical feedback data.

[0018] Alternatively, the server can also send the historical interaction data and the corresponding historical feedback data to multiple auditing parties, and receive thinking chains determined by each auditing party for the historical interaction data and the corresponding historical feedback data. The server can determine the historical thinking chains corresponding to the historical feedback data based on the received thinking chains, such as the server can perform clustering processing on the received thinking chains, and determine the historical thinking chains corresponding to the historical feedback data according to the clustering result, etc.

[0019] In addition, the above-mentioned method for determining the historical thinking chain is an optional and feasible determination method. In actual application scenarios, there can be multiple different determination methods. Different determination methods can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.

[0020] In step S104, first predicted feedback data corresponding to the historical interaction data is determined based on the historical interaction data and the first preset prompt information, taking advantage of the large language model.

[0021] Among them, the first preset prompt information can be used to prompt the large language model to generate corresponding feedback data for historical interaction data. The large language model can also be a multimodal large language model, and the historical interaction data can include multimodal data such as text data, table data, and graph data.

[0022] In step S106, the large language model is used to determine a predicted thought chain based on the historical interaction data and the second preset prompt information.

[0023] Among them, the Chain of Thought (CoT) refers to the process of breaking down logically complex problems and forming a complete thinking process through a series of logically related thinking. That is, the predicted chain of thought can be used to represent the thinking and reasoning process of the large language model generating feedback data corresponding to historical interaction data. The second preset prompt information can be used to prompt the large language model to generate corresponding feedback data for historical interaction data, and provide the thinking chain for generating feedback data.

[0024] In step S108, whether the large language model has converged is determined based on the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain. If it is determined that the large language model has not converged, the large language model continues to be trained until the large language model converges to obtain the preset large language model.

[0025] The preset large language model may be used to determine feedback data corresponding to the interaction data.

[0026] In implementation, Figure 2 As shown, the server can construct two tasks. Task 1 can require the large language model to "directly answer the user's question" (i.e., the first preset prompt information promote1), that is, output the first predicted feedback data corresponding to the historical interaction data. Task 2 can require the large language model to "give the thinking process of answering the user's question" (i.e., the second preset prompt information promote2), that is, output the specific thinking chain reasoning process (i.e., the predicted thinking chain).

[0027] Then, in the process of training the large language model, the two task samples corresponding to the same question (i.e., the same historical interaction data) can be put into a batch for gradient update, so that the large language model considers the influence of the thinking chain in the learning process and does not take a shortcut to avoid large language model illusion.

[0028] After obtaining the preset large language model, the interaction data input by the user for a certain business can be processed using the preset large language model to obtain feedback data corresponding to the interaction data input by the user. The embodiment of the present specification provides a data processing method. By obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data, a large language model is used to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information, and determine a predicted thinking chain based on the historical interaction data and second preset prompt information, wherein the predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data. Whether the large language model converges is determined according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain. In the case where the large language model does not converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained. The preset large language model can be used to determine feedback data corresponding to interaction data. In this way, through the multi-task learning mode (i.e., through the task of determining the first predicted feedback data by the first preset prompt information, and through the task of determining the predicted thinking chain by the second preset prompt information), the model can be organically integrated into the thinking chain process in the training process, and the clear and correct reasoning logic can be learned to enhance the reasoning ability of the model. In actual application, the first preset prompt information can be used to directly reason by the preset large language model, which can increase the model reasoning time compared with the thinking chain reasoning mode, i.e., the thinking chain process can be learned implicitly to improve the reasoning ability of the preset large language model, while not increasing the reasoning time, on the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved.

[0029] In actual application, the feedback data corresponding to the interaction data determined by the preset large language model can also be used to process the business corresponding to the interaction data. The specific processing mode of business processing can be various. The following provides an optional processing mode, as shown in Figure 3 The specific processing mode of business processing can be various. The following provides an optional processing mode, as shown in

[0030] In step S302, the interaction data input by the user for the target business is received.

[0031] The target service can be any service, for example, the target service can be a resource transfer service, a logic detection service, a teaching assistance service, etc., and the interaction data input by the user can include text data, table data, graph data, and other multi-modal data.

[0032] In implementation, for example, taking the target service as a teaching assistance service, the interaction data input by the user for the teaching assistance service can include mathematical logic test data, specifically, the interaction data input by the user can be “Xiaoming has 13 candies, gives 5 candies to Xiaohong, then buys 6 more, how many candies does Xiaoming have left?”.

[0033] Or, for example, taking the target service as a resource transfer service, the interaction data input by the user for the resource transfer service can include resource transfer relationship data (such as resource transfer timing data, resource transfer graph, etc.), specifically, the interaction data input by the user can be “The current account has 1000 yuan, after transferring 150 yuan to user A, the account balance is how many yuan after receiving 200 yuan transferred by user B?”.

[0034] Or, for example, taking the target service as a logic detection service, the interaction data input by the user for the logic detection service can include logic data to be detected (such as data relationship graph, table data containing data relationship, etc.), specifically, the interaction data input by the user can be “After 5 minus 6, the result of multiplying by any number is always negative, right?”.

[0035] In step S304, the first feedback data corresponding to the interaction data is determined based on the interaction data and the first preset prompt information using a preset large language model.

[0036] In step S306, the target service is processed based on the first feedback data.

[0037] In implementation, the server can input the interaction data and the first preset prompt information into the preset large language model, so that the preset large language model can directly output the first feedback data under the action of the first preset prompt information. Since the large language model learns the thinking chain process in the training process and implicitly learns the reasoning logic inside the model, it helps to alleviate the illusion problem of the model, improve the reasoning logic of the model, ensure the accuracy of the output first feedback data, and at the same time reduce the reasoning time of the model, which can meet the dual requirements of improving the reasoning logic of the model and not increasing the reasoning time in actual application.

[0038] In actual application, the feedback data corresponding to the interaction data determined by the preset large language model can also be used to process the service corresponding to the interaction data, and the specific processing method of service processing can be various, and one optional processing method is provided as follows: Figure 4As shown, the specific process can include the following steps S402-S406.

[0039] In step S402, the interaction data input by the user for the target service is received.

[0040] In implementation, the specific process of step S402 can participate in the specific process of step S302, which will not be repeated here.

[0041] In step S404, using a preset large language model, based on the interaction data and the second preset prompt information, the second feedback data corresponding to the interaction data and the target thinking chain corresponding to the second feedback data are determined.

[0042] In step S406, based on the second feedback data and the corresponding target thinking chain, the target service is processed.

[0043] In implementation, the server can output the second feedback data and the corresponding target thinking chain to the user as feedback data.

[0044] In addition, in actual application, the specific processing method of processing the target service based on the second feedback data and the corresponding target thinking chain in step S406 can be various, and an optional processing method is provided below, which can include the following steps A1-A2.

[0045] In step A1, the target thinking chain is subjected to logical detection processing.

[0046] In implementation, the server can use a pre-trained logical detection model to perform logical detection on the target thinking chain, wherein the logical detection model can be a model constructed based on a preset deep learning algorithm.

[0047] Alternatively, the server can also send the target thinking chain and the user input interaction data to an auditing party for logical detection processing.

[0048] Alternatively, the logical detection model can also output the detection result and confidence of the logical detection on the target thinking chain, and the server can classify the target thinking chain according to the confidence, such as classifying the target thinking chain according to the relationship between the confidence and a preset confidence threshold, to send the target thinking chain with a confidence not greater than the preset confidence threshold to the auditing party for secondary logical detection processing, to ensure the accuracy of logical detection.

[0049] In step A2, in the case where the target thinking chain passes the logical detection, the target service is processed according to the second feedback data.

[0050] In implementation, in the case that the target thought chain is detected by logic, the server can determine the second feedback data as the feedback data corresponding to the interaction data input by the user, and output the feedback data to the user.

[0051] In practical application, the specific processing manner of determining whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thought chain and the predicted thought chain in the above step S108 can be various, and an optional processing manner is provided as follows, which can specifically include the processing of the following steps S502-S506. Figure 5

[0052] In step S502, the first loss value is determined according to the historical feedback data and the first predicted feedback data.

[0053] In step S504, the second loss value is determined according to the historical thought chain and the predicted thought chain.

[0054] In step S506, whether the large language model converges is determined according to the first loss value and the second loss value.

[0055] In practical application, the specific processing manner of determining whether the large language model converges according to the first loss value and the second loss value in the above step S506 can be various, and an optional processing manner is provided as follows, which can specifically include the processing of the following steps B1-B3.

[0056] In step B1, the second predicted feedback data corresponding to the predicted thought chain is determined by using the large language model based on the historical interaction data and the second preset prompt information.

[0057] In step B2, the third loss value is determined according to the historical feedback data and the second predicted feedback data.

[0058] In step B3, whether the large language model converges is determined according to the first loss value, the second loss value and the third loss value.

[0059] In implementation, taking the historical interaction data "Xiaoming has 13 candies, gives 5 candies to Xiaohong, and then buys 6 candies, how many candies does Xiaoming have finally?" as an example, as shown in the following table, Figure 6 ​As shown, the server can input the first preset prompt information (i.e., Task1 Promot: directly answer the user's question) and the historical interaction data as input data of task 1 into the large language model to obtain first predicted feedback data (i.e., "Finally, Xiaoming has 14 candies left"). At the same time, the server can also input the second preset prompt information (i.e., Task2 Promot: first give the thinking process, and then answer the user's question) and the historical interaction data as input data of task 2 into the large language model to obtain second predicted feedback data (i.e., "Finally, Xiaoming has 14 candies left"), and a predicted thinking chain corresponding to the second predicted feedback data (i.e., "Xiaoming originally has 13 candies, gives 5 candies to Xiaohong, and has 13-5=8 candies left. Then buys 6 candies, so has 8+6=13 candies").

[0060] Then, the server can determine a first loss value, a second loss value, and a third loss value based on the historical feedback data, the first predicted feedback data, the second predicted feedback data, the historical thinking chain, and the predicted thinking chain. Further, according to the above three loss values, a target loss value is determined to determine whether the large language model converges according to the target loss value.

[0061] In actual application, the specific processing manner of determining the predicted thinking chain based on the historical interaction data and the second preset prompt information by using the large language model in the above step S106 can be various, and an optional processing manner is provided as follows. Figure 7 As shown, the specific processing can include the following steps S1062-S1064.

[0062] In step S1062, a target prediction task is constructed according to the historical interaction data, and the target prediction task is subjected to task splitting processing to obtain a plurality of subtasks.

[0063] Among the plurality of subtasks, there is a task dependency relationship.

[0064] In implementation, the server can perform task splitting processing on the target prediction task corresponding to the historical interaction data according to a pre-trained splitting model to obtain a plurality of subtasks with a task dependency relationship.

[0065] For example, assuming that the historical interaction data is "Xiaoming has 13 candies, gives 5 candies to Xiaohong, and then buys 6 candies, how many candies does Xiaoming have left?", then the plurality of subtasks corresponding to the historical interaction data can include subtask 1 and subtask 2, wherein subtask 1 can be "Xiaoming has 13 candies, gives 5 candies to Xiaohong, and how many candies does Xiaoming have left?", and subtask 2 can be "On the result of subtask 1, buys 6 more, and how many candies does Xiaoming have left?".

[0066] In step S1064, the plurality of sub-tasks, the task dependency relationship between the plurality of sub-tasks, and the second preset prompt information are input into a second processing module of the large language model to determine a predicted thinking chain.

[0067] In implementation, the large language model can sequentially process the sub-tasks according to the task dependency relationship between the plurality of sub-tasks, and respectively give a sub-thinking chain corresponding to each sub-task. Finally, the predicted thinking chain can be determined according to the sub-thinking chain corresponding to each sub-task and the task dependency relationship between the sub-tasks.

[0068] The embodiment of the present specification provides a data processing method. By obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data, a large language model is facilitated to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information, to determine a predicted thinking chain based on the historical interaction data and second preset prompt information using the large language model, wherein the predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data, to determine whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, and to continue training the large language model until the large language model converges to obtain a preset large language model, wherein the preset large language model can be used to determine feedback data corresponding to interaction data. In this way, through the multi-task learning mode (i.e., through the task of determining the first predicted feedback data by the first preset prompt information, and through the task of determining the predicted thinking chain by the second preset prompt information), the model can be organically integrated into the thinking chain process during the training process, learn clear and correct reasoning logic, and enhance the reasoning ability of the model. In actual application, the preset large language model can be used for direct reasoning through the first preset prompt information, which can increase the model reasoning time compared to the thinking chain reasoning mode, i.e., the thinking chain process can be learned through implicit learning, the reasoning ability of the preset large language model can be improved, and the accuracy of subsequent business processing can be improved on the basis of improving the processing effect of the large language model.

[0069] The above is the data processing method provided by the embodiment of the present specification. Based on the same idea, the embodiment of the present specification also provides a data processing apparatus, as shown in Figure 8 .

[0070] The data processing apparatus comprises a data acquisition module 801, a first processing module 802, a second processing module 803, and a model training module 804, wherein: The data acquisition module 801 is configured to acquire historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data. The first processing module 802 is configured to determine, based on the historical interaction data and first preset prompt information, first predicted feedback data corresponding to the historical interaction data, by using the large language model. The second processing module 803 is configured to determine, based on the historical interaction data and second preset prompt information, a predicted thinking chain by using the large language model, the predicted thinking chain being used to represent a thinking reasoning process of the large language model for generating feedback data corresponding to the historical interaction data. The model training module 804 is configured to determine, according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, whether the large language model converges, and continue to train the large language model until the large language model converges, to obtain a preset large language model, the preset large language model being used to determine feedback data corresponding to interaction data.

[0071] In the embodiments of the present specification, the apparatus further includes: The first receiving module is configured to receive interaction data input by a user for a target service. The third processing module is configured to determine, based on the interaction data and the first preset prompt information, first feedback data corresponding to the interaction data, by using the preset large language model. The fourth processing module is configured to process the target service based on the first feedback data.

[0072] In the embodiments of the present specification, the apparatus further includes: The second receiving module is configured to receive interaction data input by a user for a target service. The fifth processing module is configured to determine, based on the interaction data and the second preset prompt information, second feedback data corresponding to the interaction data and a target thinking chain corresponding to the second feedback data, by using the preset large language model. The sixth processing module is configured to process the target service based on the second feedback data and the target thinking chain.

[0073] In the embodiments of the present specification, the sixth processing module is configured to: perform logical detection processing on the target thinking chain. In a case where the target thinking chain passes the logical detection, process the target service according to the second feedback data.

[0074] In the embodiments of the present specification, the sixth processing module is configured to: perform logic detection on the target thought chain by using a pre-trained logic detection model, wherein the logic detection model is a model constructed based on a preset deep learning algorithm.

[0075] In the embodiments of the present specification, the model training module 804 is configured to: determine the first loss value according to the historical feedback data and the first predicted feedback data; determine a second loss value according to the historical thought chain and the predicted thought chain; determine whether the large language model converges according to the first loss value and the second loss value.

[0076] In the embodiments of the present specification, the model training module 804 is configured to: obtain second predicted feedback data corresponding to the predicted thought chain determined by the large language model based on the historical interaction data and the second preset prompt information; determine a third loss value according to the historical feedback data and the second predicted feedback data; determine whether the large language model converges according to the first loss value, the second loss value, and the third loss value.

[0077] In the embodiments of the present specification, the second processing module 803 is configured to: construct a target prediction task according to the historical interaction data, and perform task splitting processing on the target prediction task to obtain a plurality of subtasks, the plurality of subtasks having a task dependency relationship; input the plurality of subtasks, the task dependency relationship between the plurality of subtasks, and the second preset prompt information into a second processing module of the large language model to determine the predicted thought chain.

[0078] The embodiment of the present specification provides a data processing apparatus. By obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data, a large language model is beneficial. Based on the historical interaction data and the first preset prompt information, first predicted feedback data corresponding to the historical interaction data is determined. The large language model is used to determine a predicted thinking chain based on the historical interaction data and the second preset prompt information. The predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data. According to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, it is determined whether the large language model converges. In the case where it is determined that the large language model does not converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained. The preset large language model can be used to determine feedback data corresponding to the interaction data. In this way, through the multi-task learning mode (i.e., through the task of determining the first predicted feedback data by the first preset prompt information, and through the task of determining the predicted thinking chain by the second preset prompt information), the model can be organically integrated into the thinking chain process during the training process, and the clear and correct reasoning logic can be learned to enhance the reasoning ability of the model. In actual application, the preset large language model can be used for direct reasoning through the first preset prompt information. Compared with the reasoning mode through the thinking chain, the model reasoning time can be increased. That is, the thinking chain process can be learned implicitly to improve the reasoning ability of the preset large language model, while not increasing the reasoning time. On the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved.

[0079] The above is the data processing apparatus provided by the embodiment of the present specification. Based on the same idea, the embodiment of the present specification also provides a data processing device, as shown in Figure 9 .

[0080] The data processing device can be a terminal device or a server provided by the above embodiment.

[0081] The data processing device can have a large difference due to different configurations or performances, and can include one or more processors 901 and memories 902, and one or more storage applications or data can be stored in the memories 902. Among them, the memories 902 can be temporary storage or persistent storage. The application stored in the memory 902 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the data processing device. Further, the processor 901 can be configured to communicate with the memory 902 to execute a series of computer executable instructions in the memory 902 on the data processing device. The data processing device can also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, and one or more keyboards 906.

[0082] In particular, in the present embodiment, the data processing device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the data processing device, and the one or more processors are configured to execute the one or more programs include the following computer executable instructions: Obtain historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thinking chain corresponding to the historical feedback data; Based on the historical interaction data and the first preset prompt information, a first predicted feedback data corresponding to the historical interaction data is determined for the large language model; Using the large language model, a predicted thinking chain is determined based on the historical interaction data and the second preset prompt information, and the predicted thinking chain is used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data; According to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, it is determined whether the large language model converges, and in the case where the large language model does not converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained, which is used to determine feedback data corresponding to interaction data.

[0083] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment.

[0084] The embodiment of the present specification provides a data processing device. By obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data, a large language model is facilitated to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information, determine a predicted thinking chain based on the historical interaction data and second preset prompt information using the large language model, wherein the predicted thinking chain can be used to represent a thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data, and determine whether the large language model converges according to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain. In the case where the large language model is determined not to converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained, wherein the preset large language model can be used to determine feedback data corresponding to interaction data. In this way, through the multi-task learning mode (i.e., the task of determining the first predicted feedback data through the first preset prompt information, and the task of determining the predicted thinking chain through the second preset prompt information), the model can be organically integrated into the thinking chain process during the training process, learn clear and correct reasoning logic, and enhance the reasoning ability of the model. In actual application, the preset large language model can be used for direct reasoning through the first preset prompt information. Compared with the way of reasoning through the thinking chain, the model reasoning time can be increased, that is, the reasoning ability of the preset large language model can be improved through implicit learning of the thinking chain process, while the reasoning time is not increased. On the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved.

[0085] Further, based on the above Figures 1 to 7 One or more embodiments of the present specification also provide a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc. The computer executable instruction information stored in the storage medium can realize the following processes when executed by a processor. obtain historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; facilitate the large language model to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; determine a predicted thinking chain based on the historical interaction data and second preset prompt information using the large language model, wherein the predicted thinking chain is used to represent a thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data; According to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, it is determined whether the large language model converges, and in a case where it is determined that the large language model does not converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained, which is used to determine feedback data corresponding to interaction data.

[0086] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment.

[0087] The embodiment of the specification provides a storage medium. By obtaining historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data, a large language model is used to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information. The large language model is used to determine a predicted thinking chain based on the historical interaction data and second preset prompt information, wherein the predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data. According to the historical feedback data, the first predicted feedback data, the historical thinking chain, and the predicted thinking chain, it is determined whether the large language model converges. In a case where it is determined that the large language model does not converge, the training of the large language model is continued until the large language model converges, and a preset large language model is obtained, which can be used to determine feedback data corresponding to interaction data. In this way, through the multi-task learning mode (i.e., the task of determining the first predicted feedback data through the first preset prompt information, and the task of determining the predicted thinking chain through the second preset prompt information), the model can organically integrate the thinking chain process in the training process, learn clear and correct reasoning logic, and enhance the reasoning ability of the model. In actual application, the preset large language model can be used for direct reasoning through the first preset prompt information. Compared with the thinking chain reasoning mode, the model reasoning time can be increased, that is, the thinking chain process can be learned implicitly to improve the reasoning ability of the preset large language model, while the reasoning time is not increased. On the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved.

[0088] Further, based on the above Figures 1 to 7 One or more embodiments of the specification also provide a computer program product, which includes a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented. obtain historical interaction data, historical feedback data corresponding to the historical interaction data, and a historical thinking chain corresponding to the historical feedback data; For the large language model, based on the historical interaction data and the first preset prompt information, first predicted feedback data corresponding to the historical interaction data is determined; Using the large language model, based on the historical interaction data and the second preset prompt information, a predicted thinking chain is determined, which is used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data; According to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, it is determined whether the large language model converges, and in the case where the large language model does not converge, the training of the large language model is continued until the large language model converges, obtaining a preset large language model, which is used to determine feedback data corresponding to interaction data.

[0089] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0090] The embodiment of the specification provides a computer program product, by acquiring historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thinking chain corresponding to the historical feedback data, a large language model is beneficial, based on the historical interaction data and the first preset prompt information, first predicted feedback data corresponding to the historical interaction data is determined, the large language model is used, based on the historical interaction data and the second preset prompt information, a predicted thinking chain is determined, wherein the predicted thinking chain can be used to represent the thinking reasoning process of the large language model to generate feedback data corresponding to the historical interaction data, whether the large language model converges is determined according to the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, in the case of determining that the large language model does not converge, the large language model is continuously trained until the large language model converges, and a preset large language model is obtained, wherein the preset large language model can be used to determine feedback data corresponding to the interaction data. In this way, through the multi-task learning mode (that is, through the task of determining the first predicted feedback data through the first preset prompt information, and through the task of determining the predicted thinking chain through the second preset prompt information), the model can be organically integrated into the thinking chain process in the training process, and clear and correct reasoning logic is learned, and the reasoning ability of the model is enhanced. In actual application, the preset large language model can be used for direct reasoning through the first preset prompt information, which can increase the model reasoning time consumption compared with the thinking chain reasoning mode, that is, the thinking chain process can be learned through implicit learning, the reasoning ability of the preset large language model is improved, and at the same time, the reasoning time consumption is not increased, on the basis of improving the processing effect of the large language model, the accuracy of subsequent business processing can be improved.

[0091] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in an order different than the order in the embodiments, and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.

[0092] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it themselves, without having to ask a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDL, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is very easy to obtain a hardware circuit that implements a logical method flow by simply logically programming the method flow in one of the above-mentioned hardware description languages and programming it into an integrated circuit.

[0093] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the microprocessor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the same functionality in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. The controller can thus be considered a hardware component, and the means included therein for implementing the various functions can be considered structures within the hardware component. Alternatively, or even additionally, the means for implementing the various functions can be considered both software modules implementing the method and structures within the hardware component.

[0094] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0095] For the sake of description, the above apparatuses are described in various units with functions respectively. Of course, the functions of each unit can be implemented in one or more software and / or hardware in implementing one or more embodiments of the present specification.

[0096] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] The embodiments of the present specification are described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable electronic devices to produce a machine, so that the instructions executed by the computer or other programmable electronic devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0098] These computer program instructions can also be stored in a computer readable memory that can direct the computer or other programmable electronic devices to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0099] These computer program instructions can also be loaded into a computer or other programmable electronic devices, so that a series of operation steps are performed on the computer or other programmable electronic devices to produce a computer implemented process, so that the instructions executed on the computer or other programmable electronic devices provide steps for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0100] In a typical configuration, the computing device includes one or more processors (CPU), input / output interface, network interface and memory.

[0101] The memory can include non-persistent memory in the computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0102] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0103] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0104] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] One or more embodiments of the present specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One or more embodiments of the present specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including storage devices.

[0106] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0107] The above only describes the embodiments of the specification and is not used to limit the file. The specification can have various changes and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the specification shall be included in the scope of claims of the specification.

Claims

1. A data processing method, comprising: Acquire historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thought chains corresponding to the historical feedback data; Facilitating the large language model to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; Determining a predicted thought chain using the large language model based on the historical interaction data and second preset prompt information, wherein the predicted thought chain is used to represent a thought reasoning process by which the large language model generates feedback data corresponding to the historical interaction data; Based on the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, determine whether the large language model has converged. If it is determined that the large language model has not converged, continue to train the large language model until the large language model converges to obtain a preset large language model. The preset large language model is used to determine the feedback data corresponding to the interaction data.

2. The method according to claim 1, further comprising: Receive interaction data input by users for target services; Determining, using the preset large language model, first feedback data corresponding to the interaction data based on the interaction data and the first preset prompt information; The target service is processed based on the first feedback data.

3. The method according to claim 1, further comprising: Receive interaction data input by users for target services; Determining, using the preset large language model, second feedback data corresponding to the interaction data and a target thought chain corresponding to the second feedback data based on the interaction data and the second preset prompt information; The target business is processed based on the second feedback data and the corresponding target thinking chain.

4. The method according to claim 3, wherein processing the target business based on the second feedback data and the corresponding target thought chain comprises: Performing logic detection on the target thought chain; When the target thought chain passes the logic detection, the target business is processed according to the second feedback data.

5. The method according to claim 4, wherein the logic detection of the target thought chain comprises: The target thought chain is logically checked using a pre-trained logic detection model, wherein the logic detection model is a model constructed based on a preset deep learning algorithm.

6. The method according to claim 1, wherein determining whether the large language model has converged based on the historical feedback data, the first predicted feedback data, the historical thought chain, and the predicted thought chain comprises: determining the first loss value according to the historical feedback data and the first predicted feedback data; determining a second loss value according to the historical thought chain and the predicted thought chain; Determine whether the large language model converges according to the first loss value and the second loss value.

7. The method according to claim 6, wherein determining whether the large language model has converged based on the first loss value and the second loss value comprises: Obtaining second predicted feedback data corresponding to the predicted thought chain, determined using the large language model based on the historical interaction data and the second preset prompt information; determining a third loss value based on the historical feedback data and the second predicted feedback data; Determine whether the large language model converges according to the first loss value, the second loss value, and the third loss value.

8. The method according to claim 1, wherein the determining the predicted thought chain based on the historical interaction data and the second preset prompt information using the large language model comprises: Constructing a target prediction task based on the historical interaction data, and performing task splitting processing on the target prediction task to obtain multiple subtasks, wherein the multiple subtasks have task dependency relationships; The multiple subtasks, the task dependencies between the multiple subtasks, and the second preset prompt information are input into the second processing module of the large language model to determine the predicted thinking chain.

9. A data processing device comprising: A data acquisition module is used to acquire historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thought chains corresponding to the historical feedback data; a first processing module, configured to determine, using the large language model, first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; a second processing module, configured to use the large language model to determine a predicted thought chain based on the historical interaction data and second preset prompt information, the predicted thought chain being used to represent a thought and reasoning process by which the large language model generates feedback data corresponding to the historical interaction data; The model training module is used to determine whether the large language model has converged based on the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain. If it is determined that the large language model has not converged, continue to train the large language model until the large language model converges to obtain a preset large language model. The preset large language model is used to determine the feedback data corresponding to the interaction data.

10. A data processing device, comprising: processor; as well as a memory arranged to store computer-executable instructions which, when executed, cause the processor to: Acquire historical interaction data, historical feedback data corresponding to the historical interaction data, and historical thought chains corresponding to the historical feedback data; Facilitating the large language model to determine first predicted feedback data corresponding to the historical interaction data based on the historical interaction data and first preset prompt information; Determining a predicted thought chain using the large language model based on the historical interaction data and second preset prompt information, wherein the predicted thought chain is used to represent a thought reasoning process by which the large language model generates feedback data corresponding to the historical interaction data; Based on the historical feedback data, the first predicted feedback data, the historical thinking chain and the predicted thinking chain, determine whether the large language model has converged. If it is determined that the large language model has not converged, continue to train the large language model until the large language model converges to obtain a preset large language model. The preset large language model is used to determine the feedback data corresponding to the interaction data.