Data processing method and related device
Through the interaction between the first agent and the second agent, the training process of the neural network is optimized, and the dependence on manual annotated data in large-scale language model training is solved, and the self-evolution and performance improvement of the neural network is achieved.
Patent Information
- Application Number
- PCT/CN2025/071258
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
The prior art relies on manual annotation data in the training of large language models, resulting in model evolution and improvement limited by the amount of available annotation data, requiring a lot of manpower and resources, making it difficult to completely solve the problem of false or harmful information in content output.
Through the interaction between the first agent and the second agent, the evolution of the neural network is continued, and does not rely entirely on human supervision. The first agent obtains the request information and optimizes the response information through the evaluation information of the second agent, and trains the neural network to improve the reasoning effect.
The self-improvement and continuous evolution of neural networks have been achieved, the dependence on manual annotation data has been reduced, and the generalization ability and robustness of neural networks have been improved.
Smart Images

Figure CN2025071258_17072025_PF_FP_ABST
Abstract
Description
A data processing method and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 12, 2024, with application number 202410051824.5 and invention name “A data processing method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a data processing method and related equipment. Background Art
[0003] Artificial intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and basic AI theory.
[0004] Currently, supervised fine-tuning (SFT) is a common and effective strategy in the training practice of large language models. It starts with pre-training a complex model, such as GPT, on a large unlabeled dataset. The main goal is to enable the model to predict the next word that appears in the text. Next, for specific tasks, corresponding labeled datasets are collected or prepared, such as using text sets with category labels in text classification tasks. Subsequently, these labeled data are used to fine-tune the pre-trained model, adjust the model weights and possibly use different optimizers and learning rate settings to adapt to the new task requirements. Finally, the performance of the model is evaluated on an independent test set to ensure that it not only performs well on the training data, but can also effectively generalize to unknown data.
[0005] However, despite its simplicity and efficiency, the SFT method relies heavily on manually labeled data. This reliance means that the continued evolution and improvement of the model is often limited by the amount of available labeled data, which typically requires a significant investment of manpower and resources. Summary of the Invention
[0006] The embodiments of the present application provide a data processing method and related equipment, which can continue the evolution of a neural network through interaction between a first intelligent agent and a second intelligent agent, without relying entirely on human supervision.
[0007] The first aspect of the present application provides a data processing method, which is executed by a data processing device, or the method is executed by some components in the data processing device (such as a processor, chip or chip system, etc.), or the method can also be implemented by a logic module or software that can realize all or part of the functions of the data processing device. In the first aspect and its possible implementation, the method is described as being executed by a data processing device. The method includes: step 1: obtaining first request information; step 2: obtaining first response information corresponding to the first request information based on the first intelligent agent; step 3: obtaining evaluation information of the first response information based on the second intelligent agent; step 4: optimizing the first response information based on the first intelligent agent and the evaluation information to obtain second response information; step 5: training a neural network based on the first request information and the second response information.
[0008] In this embodiment of the present application, a first agent responds to a first request to obtain a first response, and a second agent evaluates the first response. Specifically, the first agent can use the second agent's evaluation of the first response to optimize the first response, and use the information from the optimization process to train a neural network, thereby improving the neural network's reasoning performance. Alternatively, the interaction between the first and second agents can sustain the evolution of the neural network without relying entirely on human oversight.
[0009] Optionally, in a possible implementation of the first aspect, the above method also includes: Step 6: Based on the second intelligent agent, multiple second request information corresponding to the first request information is obtained, and the similarity between each second request information and the first request information is greater than or equal to a first preset threshold; Step 7: Treat each second request information as the first request information and repeat steps 2 to 6 until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0010] In this possible implementation, on the one hand, a second agent can provide multiple similar requests and use them as the first request information to repeat steps 2 through 6, achieving a process similar to the "learning by analogy" process in human learning and thinking. On the other hand, through the interaction between the first and second agents, comprehensive and high-quality training data can be constructed from a horizontal dimension (generalization to increase robustness) combined with self-improvement ideas for neural network training. Furthermore, the optimized neural network can generate better data, achieving iterative evolution.
[0011] Optionally, in a possible implementation of the first aspect, the above method also includes: Step 8: obtaining multiple third request information corresponding to the multiple second request information based on the second intelligent agent and the first preset instruction, and the first preset instruction is used to increase the complexity of each second request information; Step 9: repeating step 7 for each third request information as each second request information.
[0012] In this possible implementation, on the one hand, a second agent can provide multiple similar and complex requests, and repeat step 7 using the complex request as the first request information, achieving a process similar to the "learning by analogy" and "deliberate learning" process in human learning and thinking. On the other hand, through the interaction between the first and second agents, comprehensive and high-quality training data can be constructed from both horizontal (generalization to increase robustness) and vertical (gradually learning complex requests) dimensions, combined with self-improvement ideas, for neural network training. Furthermore, the optimized neural network can generate better data and achieve iterative evolution.
[0013] Optionally, in a possible implementation of the first aspect, the above method also includes: Step 10: Obtaining fourth request information corresponding to the first request information based on the second intelligent agent and the second preset instruction, where the second preset instruction is used to increase the complexity of the first request information; Step 11: Executing steps 2 to 5 and 10 with the fourth request information as the first request information until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0014] In this possible implementation, on the one hand, a second agent can provide a complex request and use it as the first request information to repeat steps 2 through 5 and 10, achieving a "deliberate learning" process similar to the human learning and thinking process. On the other hand, through the interaction between the first and second agents, comprehensive and high-quality training data can be constructed from a vertical dimension (gradually learning complex requests) combined with self-improvement ideas for neural network training. Furthermore, the optimized neural network can generate better data and achieve iterative evolution.
[0015] Optionally, in a possible implementation manner of the first aspect, the above step of: training the neural network based on the first request information and the second response information includes: training the neural network based on the first request information, the second response information and the evaluation information.
[0016] In this possible implementation, the evaluation information can also be used as part of the training data, thereby increasing the dimensionality of the training data and improving the processing power of the neural network. This process can then be carried out continuously and interactively, allowing the neural network to continuously and interactively evolve.
[0017] Optionally, in a possible implementation of the first aspect, the first agent, the second agent, and the neural network are the same neural network or different neural networks.
[0018] This possible implementation approach, on the one hand, can combine self-improvement ideas to build comprehensive and high-quality training data for self-training. On the other hand, the optimized neural network can generate better data and achieve iterative evolution.
[0019] Optionally, in a possible implementation of the first aspect, the first agent, the second agent, and the neural network are located in the same computing device or different computing devices.
[0020] This possible implementation can be applied to one or more device scenarios, such as cloud-terminal interaction scenarios, or pure device-side scenarios, thereby increasing the diversity of applicable scenarios.
[0021] The second aspect of the present application provides a data processing device, which is a data processing device, or the device is a partial component in the data processing device (such as a processor, chip or chip system, etc.), or the device is a logic module or software that can realize all or part of the functions of the data processing device. The data processing device includes a transceiver unit and a processing unit. The transceiver unit is used to execute step 1: obtain first request information; the processing unit is used to execute step 2: obtain first response information corresponding to the first request information based on the first intelligent agent; the processing unit is also used to execute step 3: obtain evaluation information of the first response information based on the second intelligent agent; the processing unit is also used to execute step 4: optimize the first response information based on the first intelligent agent and the evaluation information to obtain the second response information; the processing unit is also used to execute step 5: train the neural network based on the first request information and the second response information.
[0022] Optionally, in a possible implementation of the second aspect, the above-mentioned processing unit is also used to execute step 6: obtaining multiple second request information corresponding to the first request information based on the second intelligent agent, and the similarity between each second request information and the first request information is greater than or equal to a first preset threshold; the processing unit is also used to execute step 7: treating each second request information as the first request information and repeatedly executing steps 2 to 6 until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0023] Optionally, in a possible implementation of the second aspect, the above-mentioned processing unit is also used to execute step 8: obtaining multiple third request information corresponding to multiple second request information based on the second intelligent agent and the first preset instruction, and the first preset instruction is used to increase the complexity of each second request information; the processing unit is also used to execute step 9: repeating step 7 for each third request information as each second request information.
[0024] Optionally, in a possible implementation of the second aspect, the above-mentioned processing unit is also used to execute step 10: obtain fourth request information corresponding to the first request information based on the second intelligent agent and the second preset instruction, and the second preset instruction is used to increase the complexity of the first request information; the processing unit is also used to execute step 11 and treat the fourth request information as the first request information to execute steps 2 to 5 and step 10 until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0025] Optionally, in a possible implementation manner of the second aspect, the above-mentioned processing unit is specifically used to train a neural network based on the first request information, the second response information and the evaluation information.
[0026] Optionally, in a possible implementation of the second aspect, the first agent, the second agent, and the neural network are the same neural network or different neural networks.
[0027] Optionally, in a possible implementation of the second aspect, the above-mentioned data processing device includes a first agent, a second agent and a neural network.
[0028] In a third aspect, the present application provides a data processing device, comprising at least one processor coupled to a memory; the memory is used to store programs or instructions; and the at least one processor is used to execute the program or instructions so that the device implements a method of any possible implementation of the first aspect described above.
[0029] In a fourth aspect, the present application provides a data processing device, comprising at least one processor, wherein the at least one processor is coupled to a memory; the memory is used to store programs or instructions; and the at least one processor is used to execute the program or instructions so that the device implements a method of any possible implementation method of the aforementioned second aspect.
[0030] In a fifth aspect, the present application provides a data processing device comprising at least one logic circuit and an input / output interface; the logic circuit is used to execute the method described in any possible implementation of the first aspect.
[0031] In a sixth aspect, the present application provides a data processing device comprising at least one logic circuit and an input / output interface; the logic circuit is used to execute a method as in any possible implementation of the second aspect described above.
[0032] A seventh aspect of the present application provides a communication system, comprising at least one of the following: a data processing device, a first device, and a second device. The data processing device is configured to execute the method described in any possible implementation of any aspect of the first aspect. The first device stores a first agent described in any possible implementation of the first aspect. The second device stores a second agent described in any possible implementation of the first aspect.
[0033] In an eighth aspect, the present application provides a communication system, the communication system comprising a first agent and a second agent;
[0034] The first agent is configured to obtain first response information of the first request information;
[0035] The second agent is configured to evaluate the first response information to obtain evaluation information;
[0036] The first intelligent agent is further used to optimize the first response information based on the evaluation information to obtain second response information; the first request information and the second response information are used to train a neural network.
[0037] In a ninth aspect, the present application provides a computer-readable storage medium for storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method described in any possible implementation of any aspect of the first aspect.
[0038] The tenth aspect of the present application provides a computer program product (or computer program). When the computer program in the computer program product is executed by the processor, the processor executes the method described in any possible implementation of any aspect of the first aspect.
[0039] In an eleventh aspect, the present application provides a chip system comprising at least one processor for supporting a data processing device to implement the method described in any possible implementation of any aspect of the first aspect.
[0040] In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the data processing device. The chip system may be composed of a chip alone or may include a chip and other discrete components. Optionally, the chip system may also include an interface circuit that provides program instructions and / or data to at least one processor.
[0041] Among them, the technical effects brought about by any design method in the second aspect to the eleventh aspect can refer to the technical effects brought about by the different design methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] FIG1 is a schematic diagram of the structure of the system architecture provided in an embodiment of the present application;
[0043] FIG2 is a schematic diagram of a chip hardware structure provided in an embodiment of the present application;
[0044] FIG3A is a schematic structural diagram of a data processing system provided in an embodiment of the present application;
[0045] FIG3B is another schematic diagram of the structure of the data processing system provided in an embodiment of the present application;
[0046] FIG3C is another schematic diagram of the structure of the data processing system provided in an embodiment of the present application;
[0047] FIG4 is a flow chart of a data processing method provided in an embodiment of the present application;
[0048] FIG5 is a flowchart of an example of a data processing method provided in an embodiment of the present application;
[0049] FIG6 is another flow chart of the data processing method provided in an embodiment of the present application;
[0050] FIG7 is another example flow chart of the data processing method provided in an embodiment of the present application;
[0051] FIG8 is another example flow chart of a data processing method provided in an embodiment of the present application;
[0052] FIG9 is a schematic diagram of a process for obtaining neural network training data according to an embodiment of the present application;
[0053] FIG10 is another flow chart of the data processing method provided in an embodiment of the present application;
[0054] FIG11 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0055] FIG12 is another structural diagram of a data processing device provided in an embodiment of the present application;
[0056] FIG13 is another structural diagram of the data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The following describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] To facilitate understanding, the following first introduces the relevant terms and concepts mainly involved in the embodiments of this application.
[0059] 1. Neural Networks
[0060] A neural network can be composed of neural units, which can be represented by X s and intercept b as inputs, the output of the operation unit can be:
[0061] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s For X s The weight of the neural unit, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0062] 2. Loss Function
[0063] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value, and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the neural network can predict the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function or objective function, which is an important equation used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Therefore, training a deep neural network becomes a process of minimizing this loss as much as possible.
[0064] 3. Large language model (LLM)
[0065] LLM is an AI algorithm that uses neural network-based language models with extremely large parameter sizes pre-trained on massive amounts of text data. For example, OpenAI's GPT series has demonstrated remarkable capabilities in a variety of tasks, including translation, question answering, and text generation. Its pre-training process consists of two main phases: initial pre-training and subsequent fine-tuning. In the initial phase, the model learns linguistic structure by training on large amounts of unlabeled text data, using techniques such as "next word prediction." This process enables the model to master the basics of natural language and broad common sense, laying a solid foundation for subsequent applications.
[0066] 4. Supervised learning
[0067] Supervised learning refers to a method of training a model using labeled data, and the goal of training is to enable the model to predict the corresponding label after receiving the data. Specifically, the model obtained after determining the parameters of the initial AI model based on the data in a given training data set and the labels corresponding to each data in the training data set, wherein the process of determining the parameters of the initial AI model using the data in the training data set and the labels corresponding to the data is also called supervised learning (or supervised training). The labels of the data in the training data set are usually manually annotated to identify the correct answer to the data for a specific task. Typical supervised learning models include: support vector machines, neural network models, logistic regression models, decision trees, naive Bayes models, Gaussian discriminant models, etc. Supervised learning models are usually used for classification or regression.
[0068] 5. Supervised finetuning (SFT)
[0069] SFT is a machine learning method in which a pre-trained model is fine-tuned using labeled data to adapt to a specific task or dataset, such as natural language question answering tasks.
[0070] 6. Alignment
[0071] Systems based on large language models are prone to generating false, harmful, or unhelpful content. The process of aligning the model's output with human preferences and values through various training methods is called alignment.
[0072] Currently, SFT begins by pre-training a complex model, such as GPT, on a large unlabeled dataset. The primary goal is to enable the model to predict the next word in a text. Next, a corresponding labeled dataset is collected or prepared for the specific task, such as a set of text with category labels for text classification tasks. This labeled data is then used to fine-tune the pre-trained model, adjusting the model weights and potentially using different optimizer and learning rate settings to adapt to the new task requirements. Finally, the model's performance is evaluated on an independent test set to ensure that it not only performs well on the training data but also generalizes effectively to unseen data.
[0073] However, despite its simplicity and efficiency, the SFT method relies heavily on manually annotated data. This reliance means that the continued evolution and improvement of the model is often limited by the amount of annotated data available, which typically requires significant human and resource investment. Furthermore, given the vast range of content generated by large language models, relying solely on traditional supervised learning methods is unlikely to fully address the potential for false or harmful information in this output.
[0074] In order to solve the above technical problems, an embodiment of the present application provides a data processing method and related equipment, which obtains first response information based on the first agent's answer to the first request, and evaluates the first response information based on the second agent. That is, the first agent can use the evaluation information of the second agent for the first response information to optimize, and use the information in the optimization process to train the neural network, thereby improving the reasoning effect of the neural network. Or it can be understood that the evolution of the neural network can be sustained through the interaction between the first agent and the second agent, without relying entirely on human supervision. For example, when the neural network includes a first agent and a second agent, it is equivalent to evolving through the interaction of the two agents, so that the neural network can improve itself and continue to evolve.
[0075] Before introducing the data processing method and related equipment of the embodiment of the present application in conjunction with the accompanying drawings, the system architecture provided by the embodiment of the present application is first described.
[0076] Referring to FIG. 1 , an embodiment of the present invention provides a system architecture 100. As shown in the system architecture 100, a data acquisition device 160 is used to collect training data. In the embodiment of the present application, the training data includes: request information and response information to the request information. Alternatively, the training data includes request information, response information to the request information, and evaluation information of the request information. The request information can be in the form of text, images, audio, video, etc., which are not specifically limited here. The evaluation information is used to evaluate the quality or accuracy of the response information. The training data is stored in a database 130, and the training device 120 trains the target model / rule 101 based on the training data maintained in the database 130. The following will describe in more detail how the training device 120 obtains the target model / rule 101 based on the training data. The target model / rule 101 can be used to implement the data processing method provided in the embodiment of the present application. The target model / rule 101 in the embodiment of the present application may specifically include a neural network. It should be noted that in actual applications, the training data maintained in the database 130 does not necessarily come from the data acquisition device 160, but may also be received from other devices. It should also be noted that the training device 120 does not necessarily train the target model / rule 101 entirely based on the training data maintained by the database 130. It is also possible to obtain training data from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0077] The target model / rule 101 obtained through training with the training device 120 can be applied to different systems or devices, such as the execution device 110 shown in FIG1 . The execution device 110 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR / VR, a vehicle-mounted terminal, etc., or a server or a cloud, etc. In FIG1 , the execution device 110 is configured with an I / O interface 112 for data interaction with an external device. The user can input data to the I / O interface 112 through the client device 140. The input data can include request information, etc. in the embodiment of the present application. In addition, the input data can be input by the user, uploaded by the user through a shooting device, output by a second agent, or from a database, which is not limited here.
[0078] The pre-processing module 113 is used to perform pre-processing (eg, splitting, grouping, format conversion, etc.) on the input data received by the I / O interface 112 .
[0079] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing, the execution device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.
[0080] Finally, the I / O interface 112 returns the processing result, such as the response information obtained above, to the client device 140 to provide it to the user.
[0081] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different goals or different tasks. The corresponding target models / rules 101 can be used to achieve the above goals or complete the above tasks, thereby providing users with the desired results.
[0082] For example, if the neural network, the first agent, and the second agent belong to the same neural network, or if the neural network and the second agent belong to the same neural network, the neural network may have multiple tasks or goals. For example, two tasks may be performed: one task is to obtain response information to request information; the other task is to obtain evaluation information of the response information.
[0083] In the scenario shown in FIG. 1 , the user can manually input data, which can be performed through the interface provided by I / O interface 112. Alternatively, client device 140 can automatically send input data to I / O interface 112. If user authorization is required for client device 140 to automatically send input data, the user can set the corresponding permissions in client device 140. The user can view the results output by execution device 110 on client device 140, which can be presented in a display, sound, action, or other specific form. Client device 140 can also serve as a data acquisition terminal, collecting input data input into I / O interface 112 and output results from I / O interface 112 as new sample data and storing them in database 130. Of course, collection can also be performed without client device 140, with I / O interface 112 directly storing the input data input into I / O interface 112 and output results from I / O interface 112 as new sample data in database 130.
[0084] It is worth noting that Figure 1 is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 1, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0085] The following describes a chip hardware structure provided by an embodiment of the present application.
[0086] FIG2 shows a chip hardware structure provided by an embodiment of the present invention, which includes a neural network processor 20. The chip can be provided in the execution device 110 shown in FIG1 to complete the computational work of the computation module 111. The chip can also be provided in the training device 120 shown in FIG1 to complete the training work of the training device 120 and output the target model / rule 101.
[0087] The neural network processor 20 can be a neural network processing unit (NPU), a tensor processing unit (TPU), or a graphics processing unit (GPU), any processor suitable for large-scale XOR operation processing. Taking the NPU as an example: the neural network processor 20 is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and the main CPU assigns tasks. The core part of the NPU is the operation circuit 203, and the controller 204 controls the operation circuit 203 to extract data from the memory (weight memory or input memory) and perform operations.
[0088] In some implementations, arithmetic circuit 203 includes multiple processing engines (PEs). In some implementations, arithmetic circuit 203 is a two-dimensional systolic array. Arithmetic circuit 203 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 203 is a general-purpose matrix processor.
[0089] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 201 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 208.
[0090] The vector calculation unit 207 can further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 207 can be used for network calculations of non-convolutional / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.
[0091] In some implementations, the vector calculation unit 207 stores the processed output vector in the unified buffer 206. For example, the vector calculation unit 207 can apply a nonlinear function to the output of the operation circuit 203, such as a vector of accumulated values, to generate an activation value. In some implementations, the vector calculation unit 207 generates a normalized value, a merged value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 203, for example, for use in a subsequent layer in a neural network.
[0092] The unified memory 206 is used to store input data and output data.
[0093] The weight data is directly transferred from the external memory to the input memory 201 and / or the unified memory 206 through the direct memory access controller 205 (DMAC), the weight data in the external memory is stored in the weight memory 202, and the data in the unified memory 206 is stored in the external memory.
[0094] The bus interface unit (BIU) 210 is used to implement interaction between the main CPU, DMAC and instruction fetch memory 209 through the bus.
[0095] An instruction fetch buffer 209 connected to the controller 204 is used to store instructions used by the controller 204 .
[0096] The controller 204 is used to call the instructions cached in the memory 209 to control the working process of the computing accelerator.
[0097] Generally, the unified memory 206, the input memory 201, the weight memory 202 and the instruction fetch memory 209 are all on-chip memories, and the external memory is a memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM) or other readable and writable memory.
[0098] Next, several data processing system scenarios involved in this application are introduced.
[0099] FIG3A is a schematic diagram of the structure of a data processing system provided in an embodiment of the present application, wherein the data processing system includes a terminal device (in FIG3A , only a mobile phone is used as an example of the terminal device) and a data processing device. It is understandable that, in addition to being a mobile phone, the terminal device may also be a tablet computer (pad), a portable game console, a personal digital assistant (PDA), a notebook computer, an ultra mobile personal computer (UMPC), a handheld computer, a netbook, a vehicle-mounted media player, a wearable electronic device, a virtual reality (VR) terminal device, an augmented reality (AR), a vehicle, a vehicle-mounted terminal, an aircraft terminal, an intelligent robot, and other terminal devices. The terminal device is the initiator of data processing, and as the initiator of a data processing request, the request is usually initiated by a user through the terminal device.
[0100] The aforementioned data processing devices can be devices or servers with data processing capabilities, such as cloud servers, network servers, application servers, and management servers. The data processing devices receive data processing requests from terminal devices via interactive interfaces and then perform data processing methods such as machine learning, deep learning, search, reasoning, and decision-making through the memory used to store data and the processors used for data processing. The memory in a data processing device is a general term that includes local storage and databases that store historical data. The database can be located on the data processing device or on other network servers.
[0101] In the data processing system shown in Figure 3A, a terminal device can receive user instructions. For example, the terminal device can obtain request information input / selected by the user and then initiate a request to the data processing device, causing the data processing device to execute a data processing application (e.g., computer vision tasks such as classification, segmentation, detection, and image generation) based on the request information received by the terminal device, thereby obtaining a processing result corresponding to the request information. For example, the terminal device can obtain a question input by the user and then initiate a request to the data processing device, causing the data processing device to reason about the question, thereby obtaining an answer to the question and displaying the answer for the user to view and use.
[0102] In FIG3A , the data processing device may execute the data processing method according to the embodiment of the present application.
[0103] Figure 3B is another structural diagram of the data processing system provided in an embodiment of the present application. In Figure 3B, the terminal device (Figure 3B only takes the terminal device as a mobile phone as an example) directly serves as a data processing device. The terminal device can directly obtain the request information and directly process it by the hardware of the terminal device itself. The specific process is similar to that of Figure 3A. Please refer to the above description and will not repeat it here.
[0104] Optionally, in the data processing system shown in Figure 3B, the terminal device can receive instructions from the user. For example, the terminal device can obtain the request information entered by the user in the terminal device, and then the terminal device itself executes a data processing application (for example, computer vision tasks such as classification, segmentation, detection, and image generation) for the request information, thereby obtaining response information corresponding to the request information, and displaying the response information for the user to view and use.
[0105] In FIG3B , the terminal device itself can execute the data processing method of the embodiment of the present application.
[0106] The terminal device in the above Figures 3A and 3B can specifically be the client device 140 or the execution device 110 in Figure 1, and the data processing device in Figure 3A can specifically be the execution device 110 in Figure 1, wherein the data storage system 150 can store the data to be processed of the execution device 110, and the data storage system 150 can be integrated on the execution device 110, or it can be set on the cloud or other network servers.
[0107] The processors in Figures 3A and 3B can perform data training / machine learning / deep learning through a neural network model or other models (such as an attention model, MLP, etc.), and use the model finally trained or learned from the data to execute data processing applications on multiple data to obtain corresponding processing results.
[0108] Figure 3C is another structural diagram of the data processing system provided in an embodiment of the present application. The data processing system may include: a first intelligent agent, a second intelligent agent, and a neural network.
[0109] The data processing system shown in FIG3C may include one or more loops. For example, the data processing system includes loop 1. Another example, the data processing system includes loop 1 and loop 2. Another example, the data processing system includes loop 1 and loop 3. Another example, the data processing system includes loop 1, loop 2, and loop 3. Of course, the data processing system may also include loop 2, loop 3, loop 2 and loop 3, etc. The specifics are not limited here.
[0110] Loop 1 can also be called the initial question-answering and optimization process. Loop 2 can also be called the similar question-answering and optimization process. Loop 3 can also be called the complex question-answering and optimization process.
[0111] In Loop 1, the second agent issues a request (or request information), the first agent processes the request and obtains a response (or response information), the second agent evaluates the response, and the first agent optimizes its response based on the evaluation (or evaluation information). Alternatively, in Loop 1, the student agent responds to the teacher agent's initial request and continuously optimizes its response based on the teacher's feedback until it passes the teacher's evaluation. Loop 1 can also be understood as a self-improving response process (or a reflective optimization process).
[0112] In loop 2, the second agent issues a similar request. The first agent processes the request and obtains a response. The second agent evaluates the response, and the first agent optimizes its response based on the evaluation. Alternatively, in loop 2, the teacher agent issues multiple requests similar to the initial request. The student agent then follows the same process as loop 1, responding to similar requests from the teacher agent until the teacher agent's requirements are met. Loop 2 can also be understood as a process of "learning from one example and applying it to other situations."
[0113] In loop 3, the second agent gives a more difficult request (more difficult or more complex than the request in loop 1 or the similar request in loop 2), the first agent processes the more difficult request and gets a response, the second agent evaluates the response, and the first agent optimizes the response based on the evaluation. Or it can be understood that in loop 3, the teacher agent gives the student agent several more difficult requests starting from the initial request. The student agent then refers to the process of loop 1 and answers the difficult requests given by the teacher agent until the requirements of the teacher agent are met. The process of loop 3 can also be understood as a process of "deliberate training" (or from easy to difficult). Through the ability to evolve complex instructions, more difficult requests are gradually constructed, allowing the neural network reasoning ability to gradually advance.
[0114] The data generated during the above-mentioned multiple cycles (for example, including at least one of the following: initial request and initial response, similar request and corresponding response, difficult request and corresponding response, evaluation of each response, etc.) can be used to train the neural network to improve the reasoning ability and alignment effect of the neural network.
[0115] In addition, the first agent, the second agent, and the neural network may belong to the same model or different models. For example, the first agent, the second agent, and the neural network belong to the same model. For another example, the second agent and the neural network belong to the same model, and the first agent is another model. For another example, the first agent and the second agent belong to the same model, and the neural network is another model, and so on. The embodiment of the present application does not limit the relationship between the first agent, the second agent, and the neural network. It can be understood that when the first agent, the second agent, and the neural network belong to the same model, the data processing system can also be understood as a self-improving question-answering system.
[0116] The following describes the first agent and the second agent by taking the data processing system including loop 1, loop 2 and loop 3 as an example.
[0117] The first agent, also known as the student agent or student model, is primarily responsible for: 1. processing requests and obtaining responses (or answering questions); and 2. optimizing responses based on evaluations.
[0118] The second agent, also called the teacher agent or teacher model, is primarily responsible for: 1. making requests; 2. evaluating responses; 3. making similar requests; and 4. making more difficult requests.
[0119] It is understandable that the above request may be given by the second agent, or by the user, or may be extracted from a database or memory, etc., and the specifics are not limited here.
[0120] The neural network is trained using the information in the above cycle. Alternatively, the information obtained from the interaction between the first agent and the second agent is used as training data to train the neural network.
[0121] The aforementioned data processing system draws on the principles of "learning by analogy" and "deliberate practice" in human learning and thinking, enabling interactive dialogue among multiple intelligent agents. On the one hand, it combines self-improvement strategies to build comprehensive, high-quality training data for self-training, both horizontally (generalization increases robustness) and vertically (gradually learning complex requests). On the other hand, the optimized neural network can generate better data, enabling iterative evolution.
[0122] The data processing method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0123] Referring to FIG4 , an embodiment of a data processing method provided in an embodiment of the present application is shown. The method can be executed by a data processing device (terminal device / cloud server) or by a component of the data processing device (e.g., a processor, chip, or chip system). The method includes steps 1 to 5. The method can be applied to model training scenarios.
[0124] Step 1: Get the first request information.
[0125] In the embodiment of the present application, there are multiple ways for the data processing device to obtain the first request information, which can be through collection / photography, user input, receiving information sent by other devices, or selecting from a database, etc., and the specific methods are not limited here.
[0126] The first request information in the embodiments of the present application can also be understood as request information input by the user, and can specifically be a question (also referred to as a question) input / selected by the user, a text processing request (for example, a summary and analysis request) input / selected by the user, a translation request input / selected by the user, etc., which are not specifically limited here. In addition, in the embodiments of the present application, the request information (for example, the first request information, the second request information, etc.) can be carried in the form of text, images, voice, etc., which are not specifically limited here.
[0127] It should be noted that the first request information may be an initial request input by the user, or a similar request given in response to the initial request, or a complex request given in response to the initial request or a similar request, etc., and is not specifically limited here.
[0128] In a possible implementation, the user may directly input the first request information to the data processing device through a peripheral device (such as a USB flash drive, a keyboard, a mouse, etc.).
[0129] In another possible implementation, the data processing device may determine the first request information based on a user operation. For example, the data processing device may display an interface to the user, the interface including multiple options, and then determine the first request information based on the user's operation on the multiple options (e.g., a click selection operation, a voice operation, etc.).
[0130] In another possible implementation manner, the second agent provides the first request information.
[0131] In Example 1, the first request information is a question (or a question) and is carried in text. The first request information is: "A sold cookies to 36 friends in March and sold twice as many in April. How many cookies did A sell in total?"
[0132] Step 2: Obtain first response information corresponding to the first request information based on the first agent.
[0133] After the data processing device obtains the first request information, it obtains the first response information corresponding to the first request information based on the first agent. The first response information can also be understood as the response information obtained by the first agent after processing the first request information.
[0134] In the embodiment of the present application, the first agent, the second agent, and the neural network are the same neural network or different neural networks. In addition, the first agent, the second agent, and the neural network can be located in the same computing device or different computing devices.
[0135] In one possible implementation, the first agent is located in the first device, and the first device and the data processing device are different devices. This step includes: the data processing device sends a first request message to the first device. Accordingly, the first device receives the first request message sent by the data processing device. The first device uses the first agent to process the first request message to obtain a first response message. The first device sends a first response message to the data processing device. Accordingly, the data processing device receives the first response message sent by the first device. It can be understood that the way in which the first device sends the first response message to the data processing device can be direct sending or through an intermediate device (such as a subsequent second device, etc.), which is not specifically limited here.
[0136] In another possible implementation, the first agent is located in a data processing device. Then this step includes: the data processing device uses the first agent to process the first request information to obtain the first response information.
[0137] Among them, the first intelligent agent can be understood as a model that can process the request information to obtain response information.
[0138] The above process of using the first agent to process the first request information to obtain the first response information can also be understood as inputting the first request information into the first agent to obtain the first response information.
[0139] For example, continuing with the above example, a first request message is input to the first agent: "A sold cookies to 36 friends in March and sold twice as many in April. How many cookies did A sell in total?" The first response message is obtained: "36 (March) + 72 (April) = 108 cookies."
[0140] Step 3: Obtain first evaluation information of the first response information based on the second agent.
[0141] After the data processing device obtains the first response information, it can obtain first evaluation information of the first response information based on the second agent.
[0142] In one possible implementation, the second agent is located in a second device, which is different from the data processing device. This step includes: the data processing device sends a first response message to the second device. In response, the second device receives the first response message sent by the data processing device. The second device processes the first response message using the second agent to obtain first evaluation information. The second device sends the first evaluation information to the data processing device. In response, the data processing device receives the first evaluation information sent by the second device.
[0143] In another possible implementation, the second agent is located in a data processing device. Then this step includes: the data processing device uses the second agent to process the first response information to obtain the first evaluation information.
[0144] The second agent can be understood as a model that can evaluate the response information to obtain evaluation information. For example, the second agent can evaluate the quality or accuracy of the response information.
[0145] In the embodiment of the present application, the evaluation information may include at least one of the following: a score for the response information, reasons for the score, analysis of the response information, suggestions for improvement, etc. The score may be expressed in the form of a numerical score or a grade of goodness or badness (e.g., excellent, good, average, poor, bad, etc.).
[0146] The above process of using the second agent to process the first response information to obtain the first evaluation information can also be understood as inputting the first response information into the second agent to obtain the first evaluation information.
[0147] It is understandable that in order to improve the ability of the second agent to understand the first response information, the first request information and / or the first preset template can be added. The first preset template is used for the second agent to clarify the reasoning task. For example, the first preset template includes: "Please evaluate the first response information corresponding to the first request information, and give an evaluation analysis and improvement suggestions." For another example, the first preset template includes: "Please analyze and evaluate the quality of the first response information and give a score (1-10)." For another example, the preset template includes: The first preset template includes: "Please analyze and evaluate the quality of the first response information and give a score (1-10). If the score is greater than 9, the reply does not need to be improved, otherwise please give an evaluation."
[0148] For example, continuing the above example, the second agent is fed the first response {36 (March) + 72 (April) = 108}, which yields a first evaluation: "The answer is clear, but since we don't know how many cookies each friend gave, we need to set variables and equations. The score is 6, which needs to be revised."
[0149] Step 4: Optimize the first response information based on the first agent and the first evaluation information to obtain the second response information.
[0150] After obtaining the first evaluation information, the data processing device can optimize the first response information based on the first agent and the first evaluation information to obtain the second response information.
[0151] In one possible implementation, the first agent is located in the first device, and the first device and the data processing device are different devices. This step includes: the data processing device sends the first evaluation information to the first device. Accordingly, the first device receives the first evaluation information sent by the data processing device. The first device uses the first agent and the first evaluation information to optimize the first response information to obtain the second response information. The first device sends the second response information to the data processing device. Accordingly, the data processing device receives the second response information sent by the first device. Similarly, the way in which the first device sends the second response information to the data processing device can be direct sending or through an intermediate device (such as a second device, etc.), which is not limited here.
[0152] In another possible implementation, the first agent is located in a data processing device. This step includes: the data processing device uses the first agent and the first rating information to optimize the first response information to obtain the second response information.
[0153] The above process of using the first agent and the first rating information to optimize the first response information to obtain the second response information can also be understood as inputting the first request information, the first response information and the first evaluation information into the first agent to obtain the second response information.
[0154] It is understood that in order to improve the first agent's understanding of the optimization process, a first request message and / or a second preset template may be added. The second preset template is used by the first agent to clarify the optimization task. For example, the second preset template may include: "Please optimize the first response message based on the first evaluation information for the first response message to obtain the second response message, and improve the accuracy of the second response message."
[0155] Exemplarily, continuing with Example 1 above, the first request information, the first response information and the first evaluation information are input to the first agent to obtain the second response information. The second response information is: "Suppose each friend is given x cookies: 36x (3 months) + 72x (4 months) = 108x. Since the number of x is unknown, a specific answer cannot be given." Of course, the second agent can also evaluate the optimized second response information. If the score does not meet the threshold, the response information can also be updated through the first agent and the new rating. Until the score of the response information meets the threshold. The overall example process can be shown in Figure 5.
[0156] It should be noted that the evaluation process of step 3 and the optimization process of step 4 can be executed once or multiple times. Alternatively, it can be understood that the number of optimization iterations can be one or more times. The stopping condition can be that the number of optimization iterations reaches a preset threshold, or that the score of the optimized response information is higher than a preset threshold, etc., which are not specifically limited here. For example, after the first agent obtains the optimized response information, the response information can be updated again based on the evaluation information of the second agent on the optimized response information until the stopping condition is met.
[0157] Step 5: Train a neural network based on the first request information and the second response information.
[0158] After the data processing device obtains the first request information and the second response information, it can train the neural network based on the first request information and the second response information. The relationship or description between the neural network, the first agent, and the second agent can be referred to the description in FIG. 3C , and will not be repeated here.
[0159] Optionally, the first request information is used as training data, and the second response information is used as a label for the training data. The neural network is trained with the goal of ensuring that the value of a loss function is less than a threshold. The loss function is used to represent the difference between the output of the neural network and the label.
[0160] The above-mentioned loss function can be L1, L2, L3, cross entropy, KL discreteness, etc. The threshold value can be set according to actual needs. The structure of the loss function and the value of the threshold value are not limited here. In addition, the neural network obtained before the above-mentioned training can be a neural network selected by the user, or a preset neural network, etc., which is not specifically limited here. The neural network can be a convolutional neural network (CNN), a feedforward neural network (FNN), a recursive neural network (RNN) (such as a long short-term memory network, a gate-controlled recurrent unit, an attention network, etc.), Transformers, a generative adversarial network (GAN), LLM, etc. The embodiment of this application does not limit the specific structure and type of the neural network.
[0161] Optionally, after the iterative interaction process, multi-dimensional model training data is generated. This data can include at least one of the following: initial request and response data, liberalization process data, and evaluation information. This data is treated as multi-dimensional training data and incorporated into model training to enhance model capabilities. This process can then continue interactively, allowing the model to continuously evolve.
[0162] In an embodiment of the present application, a first response message is obtained based on the first agent's answer to the first request, and the first response message is evaluated by the second agent. That is, the first agent can use the evaluation information of the second agent for the first response message to optimize, and use the information in the optimization process to train the neural network, thereby improving the reasoning effect of the neural network. Alternatively, it can be understood that the evolution of the neural network can be sustained through the interaction between the first agent and the second agent, without relying entirely on human supervision. For example, when the neural network includes a first agent and a second agent, it is equivalent to evolving through the interaction of the two agents, allowing the neural network to self-improve and continuously evolve.
[0163] The data processing method provided by the embodiment of the present application is described above with reference to Figure 4. The embodiment shown in Figure 4 can also be understood as the process of loop 1 shown in Figure 3C. Several other data processing methods provided by the present application are described below.
[0164] Please refer to Figure 6, which shows an embodiment of a data processing method provided in an embodiment of the present application. The method can be executed by a data processing device (terminal device / cloud server) or by a component of the data processing device (such as a processor, chip, or chip system). The method includes steps 1 to 9. This method can be applied to model training scenarios.
[0165] Step 1: Get the first request information.
[0166] Step 2: Obtain first response information corresponding to the first request information based on the first agent.
[0167] Step 3: Obtain first evaluation information of the first response information based on the second agent.
[0168] Step 4: Optimize the first response information based on the first agent and the first evaluation information to obtain the second response information.
[0169] Step 5: Train a neural network based on the first request information and the second response information.
[0170] The description of steps 1 to 5 in this embodiment can refer to the description of the embodiment shown in Figure 4 above, and will not be repeated here.
[0171] Step 6: Based on the second agent, obtain multiple second request information corresponding to the first request information.
[0172] After obtaining the first request information, the data processing device may further obtain multiple second request information corresponding to the first request information based on the second agent, wherein the similarity between each of the multiple second request information and the first request information is greater than or equal to a first preset threshold.
[0173] This step can also be understood as using the second agent to provide multiple second request information similar to the first request information.
[0174] The second agent can also be used to provide multiple similar request information for the request information. Alternatively, it can be understood that the second agent can expand or generalize the request information. The similarity between the first request information and the second request information can include at least one of the following: the first request information and the second request information have the same or similar mathematical expressions (for example, both are addition operations, subtraction operations, multiplication operations, division operations, etc.), the first request information and the second request information have the same or similar mathematical units, etc.
[0175] In one possible implementation, the second agent is located in a second device, and the second device is a different device from the data processing device. This step includes: the data processing device sends a first request message to the second device. Accordingly, the second device receives the first request message sent by the data processing device. The second device uses the second agent to process the first request message to obtain multiple similar request messages (i.e., multiple second request messages, which can also be called multiple similar questions). The second device sends multiple second request messages to the data processing device. Accordingly, the data processing device receives the multiple second request messages sent by the second device.
[0176] In another possible implementation, the second agent is located in a data processing device. Then this step includes: the data processing device uses the second agent to process the first request information to obtain multiple similar request information.
[0177] Optionally, the processing of the first request information to obtain the plurality of second request information may include replacement, modification, deletion, etc. For example, if the request information is text, numerical values, years, months, and days, units of magnitude, nouns, verbs, adjectives, adverbs indicating degree, colors, etc. in the text may be modified or replaced.
[0178] It is understandable that in order for the second agent to understand the similarity expansion process, the first request information and / or the third preset template may be added. The third preset template is used for the second agent to refer to the previous request information for similarity expansion. For example, the third preset template includes: "Please refer to the first request information to give similar second request information." For another example, the third preset template includes: "Please refer to the first request information to give similar second request information, and the second request information has a similar solution to the first request information." For another example, the third preset template includes: "Please refer to the first request information to give similar second request information, and the second request information has a similar solution to the first request information, but can be composed of completely different numbers and expressions."
[0179] For example, continuing with the example of the first request message in the embodiment shown in FIG4 , the first request message is: "A sold cookies to 36 friends in March and doubled the number in April. How many cookies did A sell in total?" The second agent performs similarity expansion on the first request message, resulting in a second request message: "A sold stickers to 24 friends in March, 1 / 3 of the number sold in February. How many stickers did A sell in total?"
[0180] Step 7: Repeat steps 2 to 6 for each second request message as the first request message until the preset condition is met.
[0181] After obtaining multiple pieces of second request information, the data processing device may treat each piece of second request information as the first request information and repeatedly perform steps 2 to 6 until a preset condition is met.
[0182] Among them, the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0183] The processes of step 6 and step 7 can be understood as the "learning from one example and applying it to other cases" process of loop 2 in the embodiment shown in FIG. 3C .
[0184] It is understood that in order for the first agent to refer to the previous first request message and second response message to conduct the inference process of the second request message, the first request message, the second response message, and / or a fourth preset template may be added. The fourth preset template is used by the first agent to refer to the previous question and answer to make the current answer. For example, the fourth preset template may include: "Please refer to the first request message and the second response message to answer the current request message (i.e., the second request message)."
[0185] Exemplarily, continuing the above example, the example process of “learning from one example and applying it to other cases” in steps 6 and 7 can be shown in FIG7 .
[0186] Step 8: Based on the second agent and the first preset instruction, obtain multiple third request information corresponding to the multiple second request information. This step is optional.
[0187] Optionally, the data processing device may further obtain a plurality of third request information corresponding to the plurality of second request information based on the second agent and a first preset instruction. The first preset instruction is used to increase the complexity of each second request information.
[0188] This step can also be understood as using the second agent and the first preset instruction to provide a plurality of third request information that is more complex or difficult than the first request information.
[0189] It is understandable that in order for the second agent to understand the process of providing complex request information, the first request information and / or the fourth preset template can be added. The fourth preset template is used for the second agent to refer to the previous request information to expand the complex request information. For example, the fifth preset template includes: "Please refer to the first request information to give a more difficult third request information." For another example, the fifth preset template includes: "Please refer to the first request information to give a more difficult third request information. Please add more conditions and restrictions to increase the difficulty of mathematical calculation or solution of the third request information."
[0190] For example, the first request is: "A sold cookies to 36 friends in March and doubled the number in April. How many cookies did A sell in total?" The second agent then performs a complex expansion on the first request, resulting in a third request: "A sold cookies to 36 friends in March and doubled the number in April, and donated an additional 15 cookies to charity. How many cookies did A lose in total?"
[0191] Step 9: Treat each third request message as each second request message and repeat step 7. This step is optional.
[0192] Optionally, after obtaining multiple pieces of third request information, the data processing device may further treat each piece of third request information as each piece of second request information and repeatedly perform step 7.
[0193] The process of step 8 and step 9 can be understood as the "deliberate learning" process of loop 3 in the embodiment shown in FIG. 3C .
[0194] Exemplarily, continuing the above example, the example process of “deliberate learning” in steps 8 and 9 can be shown in FIG8 .
[0195] It is understood that there are multiple scenarios in the embodiment shown in FIG5 . In one possible implementation, the embodiment shown in FIG5 includes steps 1 to 9, which can be understood as the solution of cycle 1+cycle 2+cycle 3 shown in FIG3C . In another possible implementation, the embodiment shown in FIG5 includes steps 1 to 7, which can be understood as the solution of cycle 1+cycle 2 shown in FIG3C .
[0196] As shown in Figure 9, after the iterative interaction of the first three cycles, multi-dimensional model training data is generated. This data can include at least one of the following: initial request and response, similar and self-improving request and response, complex and self-improving request and response, multiple rounds of inference request and response, multiple rounds of easy-to-difficult request and response, liberalization process data, and complex instruction evolution data. This data is treated as multi-dimensional training data and incorporated into model training to enhance model capabilities. This process can then be continued interactively, allowing the model to continuously evolve interactively.
[0197] This example draws on the principles of "learning by analogy" and "deliberate practice" in human learning and thinking, enabling multiple agents to engage in interactive dialogue. On the one hand, by combining self-improvement strategies from both horizontal (generalization to increase robustness) and vertical (gradually learning complex requests) dimensions, comprehensive and high-quality training data can be constructed for self-training. On the other hand, the optimized neural network can generate better data, enabling iterative evolution.
[0198] Please refer to Figure 10, which shows an embodiment of a data processing method provided in an embodiment of the present application. The method can be executed by a data processing device (terminal device / cloud server) or by a component of the data processing device (such as a processor, chip, or chip system). The method includes steps 1 to 5, step 10, and step 11. This method can be applied to model training scenarios.
[0199] Step 1: Get the first request information.
[0200] Step 2: Obtain first response information corresponding to the first request information based on the first agent.
[0201] Step 3: Obtain first evaluation information of the first response information based on the second agent.
[0202] Step 4: Optimize the first response information based on the first agent and the first evaluation information to obtain the second response information.
[0203] Step 5: Train a neural network based on the first request information and the second response information.
[0204] The description of steps 1 to 5 in this embodiment can refer to the description of the embodiment shown in Figure 4 above, and will not be repeated here.
[0205] Step 10: Obtain fourth request information corresponding to the first request information based on the second agent and the second preset instruction.
[0206] After the data processing device obtains the first request information, it can obtain fourth request information corresponding to the first request information based on the second agent and the second preset instruction. The second preset instruction is used to increase the complexity of the first request information.
[0207] Step 11: Execute steps 2 to 5 and 10 by treating the fourth request information as the first request information until the preset conditions are met.
[0208] After obtaining the fourth request information, the data processing device may treat the fourth request information as the first request information and execute steps 2 to 5 and 10 until a preset condition is met.
[0209] Among them, the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0210] The embodiment shown in Figure 10 can be understood as the solution of loop 1 + loop 3 shown in Figure 3C. The process of steps 10 and 11 can be understood as the "deliberate learning" process of loop 3 in the embodiment shown in Figure 3C.
[0211] Similarly, after the iterative interaction of loops 1 and 3, multi-dimensional model training data is generated. This data can include at least one of the following: initial request and response, complex and self-improving request and response, multiple rounds of easy-to-difficult request and response, data on the liberalization process, and data on the evolution of complex instructions. This data is treated as multi-dimensional training data and incorporated into model training to enhance model capabilities. This process can then continue interactively, allowing the model to continuously evolve.
[0212] This example, on the one hand, draws on the concept of "deliberate practice" in the human learning and thinking process to enable multiple intelligent agents to engage in interactive dialogue. On the other hand, from a vertical dimension (gradually learning complex requests), combined with self-improvement principles, comprehensive and high-quality training data can be constructed for self-training. Furthermore, the optimized neural network can generate better data, achieving iterative evolution.
[0213] The communication method in the embodiment of the present application is described above. The data processing device in the embodiment of the present application is described below. Please refer to Figure 11. An embodiment of a data processing device 1100 in the embodiment of the present application is shown. The data processing device 1100 can implement the functions of the first device or the second device in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment. In the embodiment of the present application, the data processing device 1100 can be a data processing device, or it can be an integrated circuit or component within the data processing device, such as a chip. The data processing device 1100 includes: a transceiver unit 1101 and a processing unit 1102.
[0214] The transceiver unit 1101 is configured to execute step 1: obtaining first request information;
[0215] The processing unit 1102 is configured to execute step 2: obtaining, based on the first agent, first response information corresponding to the first request information;
[0216] The processing unit 1102 is further configured to execute step 3: obtaining evaluation information of the first response information based on the second agent;
[0217] The processing unit 1102 is further configured to execute step 4: optimizing the first response information based on the first agent and the evaluation information to obtain second response information;
[0218] The processing unit 1102 is further configured to execute step 5: training a neural network based on the first request information and the second response information.
[0219] Optionally, the processing unit 1102 is also used to execute step 6: obtaining multiple second request information corresponding to the first request information based on the second intelligent agent, and the similarity between each second request information and the first request information is greater than or equal to a first preset threshold; the processing unit 1102 is also used to execute step 7: treating each second request information as the first request information and repeatedly executing steps 2 to 6 until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0220] Optionally, the processing unit 1102 is also used to execute step 8: obtaining multiple third request information corresponding to the multiple second request information based on the second agent and the first preset instruction, and the first preset instruction is used to increase the complexity of each second request information; the processing unit 1102 is also used to execute step 9: repeating step 7 with each third request information as each second request information.
[0221] Optionally, the processing unit 1102 is also used to execute step 11: obtain fourth request information corresponding to the first request information based on the second intelligent agent and the second preset instruction, and the second preset instruction is used to increase the complexity of the first request information; the processing unit 1102 is also used to execute step 11: treat the fourth request information as the first request information to execute steps 2 to 5 and step 10 until the preset conditions are met; the preset conditions include at least one of the following: the number of repeated executions meets the second preset threshold, the number of request information meets the third preset threshold, and the evaluation information obtained after repeated execution meets the preset requirements.
[0222] Optionally, the processing unit is specifically configured to train a neural network based on the first request information, the second response information, and the evaluation information.
[0223] Optionally, the first agent, the second agent and the neural network are the same neural network or different neural networks.
[0224] Optionally, the data processing device includes a first agent, a second agent, and a neural network.
[0225] In this embodiment, the operations performed by each unit in the data processing device are similar to the description of the data processing device in the embodiments shown in Figures 1 to 10 above, and will not be repeated here.
[0226] This example, on the one hand, draws on the principles of "learning by analogy" and "deliberate practice" in the human learning and thinking process to enable interactive dialogue among multiple intelligent agents. Furthermore, it combines self-improvement strategies to build comprehensive, high-quality training data for self-training, both horizontally (generalization increases robustness) and vertically (gradually learning complex requests). Furthermore, the optimized neural network can generate better data, enabling iterative evolution.
[0227] Please refer to Figure 12, which is another schematic structural diagram of a data processing device 1200 provided in this application. The data processing device 1200 includes a logic circuit 1201 and an input / output interface 1202. The data processing device 1200 may be a chip or an integrated circuit.
[0228] The transceiver unit 1101 shown in FIG11 may be a communication interface, which may be the input / output interface 1202 shown in FIG12 . The input / output interface 1102 may include an input interface and an output interface. Alternatively, the communication interface may be a transceiver circuit, which may include an input interface circuit and an output interface circuit. The processing unit 1102 shown in FIG11 may be the logic circuit 1201 shown in FIG12 .
[0229] Optionally, the logic circuit 1201 is configured to obtain first response information to the first request information, obtain evaluation information of the first response information, optimize the first response information based on the evaluation information, and train a neural network based on the first request information and the second response information. The input / output interface 1202 is configured to obtain the first request information.
[0230] The logic circuit 1201 and the input / output interface 1202 may also execute other steps executed by the data processing device in any embodiment and achieve corresponding beneficial effects, which will not be described in detail here.
[0231] Optionally, the logic circuit 1201 may be a processing device, and the functions of the processing device may be partially or entirely implemented by software. The functions of the processing device may be partially or entirely implemented by software.
[0232] Optionally, the processing device may include a memory and a processor, wherein the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to perform corresponding processing and / or steps in any one of the method embodiments.
[0233] Alternatively, the processing device may include only a processor. A memory for storing the computer program is located outside the processing device, and the processor is connected to the memory via circuits / wires to read and execute the computer program stored in the memory. The memory and processor may be integrated or physically separate.
[0234] Optionally, the processing device may be one or more chips, or one or more integrated circuits. For example, the processing device may be one or more field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system-on-chips (SoCs), central processor units (CPUs), network processors (NPs), digital signal processors (DSPs), microcontroller units (MCUs), programmable logic devices (PLDs), or other integrated chips, or any combination of the above chips or processors.
[0235] Please refer to FIG. 13 , which shows a data processing device 1300 involved in the above embodiments provided in an embodiment of the present application. Specifically, the data processing device 1300 may be the data processing device in the above embodiments.
[0236] Herein, a possible logical structure diagram of the data processing device 1300 is shown. The data processing device 1300 may include but is not limited to at least one processor 1301 and a communication port 1302 .
[0237] The transceiver unit 1101 shown in FIG11 may be a communication interface, which may be the communication port 1302 in FIG13 , which may include an input interface and an output interface. Alternatively, the communication port 1302 may be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0238] It is understandable that the communication port 1302 in FIG. 13 may be used to obtain request information.
[0239] Further optionally, the apparatus may also include at least one of a memory 1303 and a bus. In an embodiment of the present application, the at least one processor 1301 is used to control and process the actions of the data processing device 1300 .
[0240] In addition, the processor 1301 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0241] It should be noted that the data processing device 1300 shown in Figure 13 can be specifically used to implement the steps implemented by the data processing device in the aforementioned method embodiment and to achieve the corresponding technical effects of the data processing device. The specific implementation methods of the data processing device shown in Figure 13 can refer to the description in the aforementioned method embodiment and will not be repeated here.
[0242] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0243] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0244] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in whole or in part through software, hardware, firmware, or any combination thereof.
[0245] When software is used to implement the integrated unit, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0246] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
Claims
1. A data processing method, characterized in that, The method includes: Step 1: Obtain first request information; Step 2: Based on a first agent, obtain first response information corresponding to the first request information; Step 3: Based on a second agent, obtain evaluation information of the first response information; Step 4: Optimize the first response information based on the first agent and the evaluation information to obtain second response information; Step 5: Train a neural network based on the first request information and the second response information.
2. The method according to claim 1, wherein The method further includes: Step 6: Based on the second agent, obtain multiple second request information corresponding to the first request information, and the similarity between each second request information in the multiple second request information and the first request information is greater than or equal to a first preset threshold; Step 7: Treat each second request information as the first request information and repeat Steps 2 to 6 until a preset condition is met; the preset condition includes at least one of the following: the number of repetitions meets a second preset threshold, the number of request information meets a third preset threshold, and the evaluation information obtained after repetition meets a preset requirement.
3. The method according to claim 2, wherein The method further includes: Step 8: Based on the second agent and a first preset instruction, obtain multiple third request information corresponding to the multiple second request information, and the first preset instruction is used to increase the complexity of each second request information; Step 9: Treat each third request information as each second request information and repeat Step 7.
4. The method according to claim 1, characterized in that, The method further includes: Step 10: Based on the second agent and a second preset instruction, obtain fourth request information corresponding to the first request information, and the second preset instruction is used to increase the complexity of the first request information; Step 11: Treat the fourth request information as the first request information and execute Steps 2 to 5 and Step 10 until a preset condition is met; the preset condition includes at least one of the following: the number of repetitions meets a second preset threshold, the number of request information meets a third preset threshold, and the evaluation information obtained after repetition meets a preset requirement.
5. The method according to any one of claims 1 to 4, characterized in that The training of the neural network based on the first request information and the second response information includes: Training the neural network based on the first request information, the second response information, and the evaluation information.
6. The method according to any one of claims 1 to 5, characterized in that The first agent, the second agent, and the neural network are the same neural network or different neural networks.
7. The method according to any one of claims 1 to 6, characterized in that, The first agent, the second agent, and the neural network are located in the same computing device or different computing devices.
8. A data processing device, characterized in that, The data processing device includes a transceiver unit and a processing unit: The transceiver unit is used to execute Step 1: Obtain first request information; The processing unit is used to execute Step 2: Based on a first agent, obtain first response information corresponding to the first request information; The processing unit is further used to execute Step 3: Based on a second agent, obtain evaluation information of the first response information; The processing unit is further used to execute Step 4: Optimize the first response information based on the first agent and the evaluation information to obtain second response information; The processing unit is further configured to perform step 5: training a neural network based on the first request information and the second response information.
9. The data processing device according to claim 8, wherein the processing unit is further configured to perform step 6: obtaining, based on the second agent, a plurality of second request information corresponding to the first request information, and a similarity between each second request information in the plurality of second request information and the first request information being greater than or equal to a first preset threshold; the processing unit is further configured to perform step 7: repeating steps 2 to 6 with each second request information regarded as the first request information until a preset condition is met; the preset condition includes at least one of the following: the number of repetitions meets a second preset threshold, the number of request information meets a third preset threshold, and the evaluation information obtained after repetition meets a preset requirement.
10. The data processing device according to claim 9, wherein the processing unit is further configured to perform step 8: obtaining, based on the second agent and a first preset instruction, a plurality of third request information corresponding to the plurality of second request information, the first preset instruction being used to increase the complexity of each second request information; the processing unit is further configured to perform step 9: repeating step 7 with each third request information regarded as each second request information.
11. The data processing device according to claim 8, wherein the processing unit is further configured to perform step 10: obtaining, based on the second agent and a second preset instruction, a fourth request information corresponding to the first request information, the second preset instruction being used to increase the complexity of the first request information; the processing unit is further configured to perform step 11: repeating steps 2 to 5 and step 10 with the fourth request information regarded as the first request information until a preset condition is met; the preset condition includes at least one of the following: the number of repetitions meets a second preset threshold, the number of request information meets a third preset threshold, and the evaluation information obtained after repetition meets a preset requirement.
12. The data processing device according to any one of claims 8 to 11, characterized in that The processing unit is specifically configured to train the neural network based on the first request information, the second response information, and the evaluation information.
13. The data processing device according to any one of claims 8 to 12, characterized in that, The first agent, the second agent, and the neural network are the same neural network or different neural networks.
14. The data processing device according to any one of claims 8 to 13, characterized in that The data processing device includes the first agent, the second agent, and the neural network.
15. A data processing device, characterized in that, Comprising: a processor, the processor being coupled to a memory, the memory being configured to store programs or instructions, and when the programs or instructions are executed by the processor, causing the data processing device to perform the method according to any one of claims 1 to 7.
16. The data processing device according to claim 15, wherein The data processing device is a chip.
17. A communication system, characterized in that, The communication system includes a first agent and a second agent; The first agent is configured to obtain a first response information of the first request information; The second agent is configured to evaluate the first response information to obtain evaluation information; The first intelligent agent is further configured to optimize the first response message based on the evaluation information to obtain a second response message; the first request message and the second response message are used to train a neural network.
18. The communication system according to claim 17, characterized in that, The communication system further includes the neural network.
19. A computer storage medium, characterized in that, It includes computer instructions that, when running on a terminal device, cause the terminal device to execute the method according to any one of claims 1 to 7.
20. A computer program product, characterized in that, When the computer program product runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and related equipment
CN120317312A
Training method and device of neural network model for question-answer matching
CN110427466A
Neural network model compression method and device, storage medium and chip
CN112446476A
Market making method based on teacher-student model and reinforcement learning
CN114049223A
Intelligent agent training method and device, medium and computing equipment
CN114676847A
Cited By
Data analysis method and device based on multi-agent cooperation and storage medium
CN121073371A