Data processing method and related product
By leveraging multi-node collaboration and training data optimization, the problem of insufficient accuracy in reward prediction for large-scale artificial intelligence models has been solved, resulting in more accurate reward prediction and a simplified data processing workflow.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2023-10-18
- Publication Date
- 2026-04-28
AI Technical Summary
Existing large-scale artificial intelligence models are not accurate enough in reward prediction and struggle to provide operational instructions that match the characteristics of the environment.
By working collaboratively across multiple nodes and utilizing training and feedback data, the parameters of child nodes are optimized, weights and weighted sums are determined, and more accurate reward prediction results are provided.
It improves the accuracy of reward prediction, reduces node processing pressure and signaling overhead, and simplifies the data processing flow.
Smart Images

Figure CN121941997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology in data processing, and more particularly to data processing methods and related products. Background Technology
[0002] Large-scale artificial intelligence (AI) models with millions or billions of parameters have shown enormous potential in various AI applications, such as Large Language Model Meta AI (LLaMA), and are considered one of the most promising AI technologies for the future. In the training of large-scale AI models, reinforcement learning from human feedback (RLHF) techniques have proven to play a crucial role in the fine-tuning of these models.
[0003] The purpose of providing this background information is to disclose information that the applicant believes may be relevant to the present invention. It is not necessarily an admission that such information is relevant to the present invention, nor should it be construed as constituting prior art in relation to the present invention. Summary of the Invention
[0004] In a first aspect, one embodiment of the present invention provides a data processing method, comprising:
[0005] The first node acquires a first observation, wherein the first observation is in response to the operation of the second node indicating the state of the environment;
[0006] The first node obtains a first prediction of the operation for the second node based on the first observation, wherein the first prediction is related to a first feature information of a third node in the environment, and the first prediction is used to update the second node.
[0007] By obtaining a first prediction related to the first feature information of the third node in the environment, which is used to update the second node, the first node is able to provide a more accurate reward prediction, thereby enabling the second node to generate an operation that matches the first feature information of the third node.
[0008] In one possible implementation of the first aspect, the first node is associated with multiple feature information, which are included in the multiple feature information.
[0009] In this implementation, since the first node is associated with multiple feature information including the first feature information, the pre-trained first node can be used to provide more accurate reward prediction results for customers with different features (such as expertise categories).
[0010] In one possible implementation of the first aspect, the first node obtaining the first prediction of the operation for the second node based on the first observation includes:
[0011] The first node determines the first prediction based on the first observation and a first indication from the fourth node, wherein the first indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
[0012] In this implementation, since the first indication indicates the first feature information and is determined by the fourth node based on the feedback data about the first observation provided by the third node, the feedback data provided by the third node can be used to help identify the first feature information of the third node.
[0013] In one possible implementation of the first aspect, the first node is trained using a first training dataset, which includes multiple training data sets, each of which includes a first data portion and a second data portion, the second data portion indicating feature information of the first data portion.
[0014] In this implementation, since the first node is trained using training data, which includes a first data portion and a second data portion containing feature information indicating the first data portion, the first node can provide a more accurate reward prediction result based on the first feature information indicated by the first indication and feedback data about the first observation.
[0015] In one possible implementation of the first aspect, the first node includes a plurality of child nodes, each of the plurality of child nodes being associated with one of the plurality of feature information.
[0016] In this implementation, the first node includes multiple child nodes. The parameters of each child node can be optimized / trained to be associated with one of the multiple feature information. Thus, each child node can provide a reward prediction result based on the feature information, and all child nodes can provide a more accurate reward prediction result.
[0017] In one possible implementation of the first aspect, the first node obtaining the first prediction of the operation for the second node based on the first observation includes:
[0018] For each of the plurality of child nodes, the child node determines a second prediction of the operation for the second node based on the first observation, and the child node determines a weight corresponding to the second prediction based on a second indication from a fourth node, wherein the second indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
[0019] The first child node among the plurality of child nodes obtains a weighted sum of the second predictions of the plurality of child nodes as the first prediction, based on the second prediction of each of the plurality of child nodes and the weight corresponding to the second prediction.
[0020] In this implementation, each of the multiple child nodes determines a second prediction for the operation of the second node based on the first observation, and determines the weight corresponding to the second prediction, so that the first child node among the multiple child nodes can determine the first prediction as a weighted sum of the second predictions of the multiple child nodes, thereby enabling the first child node to provide a more accurate reward prediction result.
[0021] In one possible implementation of the first aspect, the child node determining the weight corresponding to the second prediction includes:
[0022] The child node determines the weight corresponding to the second prediction based on the second indication and the feature information associated with the child node.
[0023] In this implementation, for each of the multiple child nodes, a second indication indicating the first feature information of the third node is sent to the multiple child nodes to determine the weight corresponding to the second prediction. Therefore, the determined weight can be used to calibrate the initial second prediction, thereby enabling the first child node to provide a more accurate reward prediction result based on the calibrated second prediction.
[0024] In one possible implementation of the first aspect, for each of the plurality of child nodes, the second indication further indicates a weight corresponding to the child node.
[0025] In this implementation, the fourth node determines a second instruction corresponding to the weight of each of the multiple child nodes and sends the second instruction to the multiple child nodes, so that the multiple child nodes do not need to determine the weight corresponding to the second prediction based on the second instruction. Instead, the first child node among the multiple child nodes can directly determine the weighted sum of the second prediction based on the weight indicated by the second instruction, thereby enabling the first child node to provide a more accurate reward prediction result.
[0026] In one possible implementation of the first aspect, the method further includes:
[0027] For each of the plurality of child nodes, the child node determines a third prediction for the operation of the second node based on the first observation, and the child node sends the third prediction to a fourth node so that the fourth node determines a third indication indicative of the first prediction based on the third prediction and feedback data about the first observation provided by the third node.
[0028] The first prediction obtained by the first node based on the first observation for the operation against the second node includes:
[0029] The first child node among the plurality of child nodes receives the third indication, wherein the third indication indicates the first prediction sent by the fourth node.
[0030] In this implementation, each of the multiple child nodes determines a third prediction and sends it to a fourth node. This allows the fourth node to determine a third instruction based on the third prediction and feedback data, instructing the first prediction (which would be a more accurate reward prediction result for the second node's operation). The fourth node then sends this third instruction to the first child node, enabling the first child node to provide a more accurate reward prediction result. Furthermore, this calculation is performed by the fourth node, reducing the processing load on the first node. Since the first prediction can be sent from the first child node to the second node, the fourth node does not need to interact directly with the second node, thus simplifying its operation.
[0031] In one possible implementation of the first aspect, the first child node is predefined or selected by the fourth node.
[0032] In this implementation, the first child node, which serves as the anchor node, undertakes more interaction tasks than other child nodes. This first child node can be predefined or selected by the fourth node, thereby simplifying the operation of the data processing method.
[0033] In one possible implementation of the first aspect, the first prediction includes a third prediction determined by the plurality of child nodes respectively.
[0034] The first prediction obtained by the first node based on the first observation for the operation against the second node includes:
[0035] For each of the plurality of child nodes, the child node determines a third prediction of the operation for the second node based on the first observation;
[0036] The method further includes:
[0037] For each of the plurality of child nodes, each of the plurality of child nodes sends the third prediction to the fourth node, so that the fourth node determines a fourth prediction for the operation of the second node based on the third prediction and feedback data provided by the third node regarding the first observation, wherein the fourth prediction corresponds to the first feature information of the third node in the environment.
[0038] In this implementation, each of the multiple child nodes determines a third prediction and sends it to a fourth node. This allows the fourth node to determine a fourth prediction (which would be a more accurate reward prediction for the operation of the second node) based on the third prediction and feedback data about the first observation provided by the third node. This enables the fourth node to provide a more accurate reward prediction. Furthermore, since this calculation is performed by the fourth node, the processing load on the first node is reduced, and the first prediction can be sent directly from the fourth node to the second node, thus reducing signaling overhead.
[0039] In one possible implementation of the first aspect, the first node is trained using a second training dataset, which includes multiple sets of training data, each set of training data corresponding to predefined feature information among the multiple feature information.
[0040] In this implementation, since the first node is trained using training data corresponding to predefined feature information from multiple feature information, the reward prediction results provided by the first node will be more accurate.
[0041] In one possible implementation of the first aspect, the method further includes:
[0042] The first node sends the first prediction to the second node to implement the update of the second node.
[0043] In this implementation, the first prediction sent by the first node to the second node is used to update the second node, making the parameters / one or more models of the second node updated according to the first prediction more accurate.
[0044] In one possible implementation of the first aspect, the first node acquiring the first observation includes:
[0045] The first node receives the first observation from the third node; or
[0046] The first node receives the first observation broadcast by the fourth node.
[0047] In this implementation, the first node receives the first observation from the third node or receives the first observation broadcast by the fourth node, thereby diversifying the ways in which the first node receives the first observation.
[0048] Secondly, one embodiment of the present invention provides a data processing method, including:
[0049] The fourth node receives feedback data from the third node regarding the first observation, wherein the first observation is in response to the state of the environment indicated by the operation of the second node.
[0050] By receiving feedback data about the first observation from the third node, the fourth node can determine the first feature information of the third node based on the feedback data, thereby obtaining a reward prediction result that matches the first feature information of the third node.
[0051] In one possible implementation of the second aspect, the first node includes multiple child nodes, each of which is associated with one of the multiple feature information.
[0052] The method further includes:
[0053] The fourth node receives a third prediction from each of the plurality of child nodes;
[0054] The fourth node determines a final prediction for the operation of the second node based on the third prediction and the feedback data from the third node regarding the first observation, wherein the final prediction corresponds to the first feature information of the third node in the environment.
[0055] In this implementation, the first node includes multiple child nodes. The parameters of each child node can be optimized / trained to be associated with one of multiple feature information, thereby providing a more accurate reward prediction result. Furthermore, each child node determines a third prediction and sends it to a fourth node. The fourth node determines the final prediction for the operation on the second node based on the third prediction and feedback data from the third node regarding the first observation, which further contributes to providing a more accurate reward prediction result.
[0056] In one possible implementation of the second aspect, the fourth node determines the final prediction by:
[0057] The fourth node determines the similarity between the feedback data and the third prediction from each of the plurality of child nodes;
[0058] The fourth node determines the final prediction from each of the plurality of child nodes based on the third prediction and the similarity between the feedback data and the third prediction.
[0059] In this implementation, the final prediction is determined from each of the multiple child nodes based on the third prediction and the similarity between the feedback data and the third prediction, which helps to provide more accurate reward prediction results.
[0060] In one possible implementation of the second aspect, the fourth node determines the final prediction from each of the plurality of child nodes based on the third prediction and the similarity between the feedback data and the third prediction, including:
[0061] If there is at least one third prediction with a similarity greater than or equal to a preset threshold, the fourth node will determine the third prediction with the highest similarity as the final prediction.
[0062] In the absence of a third prediction with a similarity greater than or equal to the preset threshold, the fourth node determines the weight of the third prediction corresponding to each of the plurality of child nodes based on the similarity between the feedback data and the third prediction, and the fourth node calculates the weighted sum of the third predictions of the plurality of child nodes as the final prediction based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction.
[0063] In this implementation, the final prediction can be determined in multiple ways based on the third prediction, thereby providing a more accurate reward prediction result.
[0064] In one possible implementation of the second aspect, the final prediction is the first prediction;
[0065] The method further includes:
[0066] The fourth node selects a first child node from the plurality of child nodes based on the similarity of the third prediction of the plurality of child nodes;
[0067] The fourth node sends a third indication to the first child node, indicating the first prediction.
[0068] In this implementation, a first child node is selected from multiple child nodes based on the similarity of the third predictions of multiple child nodes, and its role (as an anchor child node) is notified to the first child node in a more flexible way by indicating the third indication of the first prediction.
[0069] In one possible implementation of the second aspect, the final prediction is a fourth prediction;
[0070] The method further includes:
[0071] The fourth node sends the fourth prediction to the second node.
[0072] In this implementation, the final prediction is sent directly from the fourth node to the second node, thereby avoiding the forwarding of the final prediction and reducing the signaling required to implement the data processing method.
[0073] In one possible implementation of the second aspect, the method further includes:
[0074] The fourth node determines the first feature information of the third node based on the feedback data.
[0075] In this implementation, the first feature information of the third node is determined based on the feedback data from the third node regarding the first observation. The feedback data provided by the third node can be used to help identify the first feature information of the third node.
[0076] In one possible implementation of the second aspect, the method further includes:
[0077] The fourth node sends a first indication to the first node, wherein the first indication indicates the first feature information of the third node.
[0078] In this implementation, since the first instruction indicates the first feature information of the third node and is sent to the first node by the fourth node, the reward prediction result subsequently determined by the first node based on the first instruction is more accurate. In other words, the first instruction can be used to help provide a more accurate reward prediction result.
[0079] In one possible implementation of the second aspect, the first node includes multiple child nodes, each of which is associated with one of the multiple feature information.
[0080] The method further includes:
[0081] The fourth node sends a second indication to each of the plurality of child nodes, wherein the second indication indicates the first feature information of the third node; or
[0082] For each of the plurality of child nodes, the fourth node determines a weight corresponding to the feature information associated with the child node, and the fourth node sends a second indication to the child node, wherein the second indication indicates the first feature information of the third node and the weight corresponding to the child node.
[0083] In this implementation, the first node includes multiple child nodes. The parameters of each child node can be optimized / trained to be associated with one of multiple feature information, thereby providing a more accurate reward prediction result. Furthermore, the fourth node sends a first feature information indicating the third node or a second indication corresponding to the weight of that child node to each of the multiple child nodes. This allows the first child node to determine the final prediction based on the second indication, thus enabling it to provide a more accurate reward prediction result.
[0084] In one possible implementation of the second aspect, the first node includes multiple child nodes, each of which is associated with one of the multiple feature information.
[0085] The method further includes:
[0086] The fourth node receives a third prediction from each of the plurality of child nodes;
[0087] For each of the plurality of child nodes, the fourth node determines the weight of the third prediction corresponding to the child node based on the first feature information of the third node and the feature information associated with the child node;
[0088] The fourth node calculates a weighted sum of the third predictions of the plurality of child nodes as the final prediction based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction, wherein the final prediction is the fourth prediction;
[0089] The fourth node sends the fourth prediction to the second node.
[0090] In this implementation, the first node includes multiple child nodes. The parameters of each child node can be optimized / trained to be associated with one of multiple feature information, thereby providing a more accurate reward prediction result. Furthermore, the fourth node receives a third prediction from each of the multiple child nodes and, for each child node, determines the weight of the third prediction corresponding to that child node based on the first feature information of the third node and the feature information associated with that child node. It then calculates the weighted sum of the third predictions from the multiple child nodes as the fourth prediction sent to the second node, thus providing a more accurate reward prediction result.
[0091] In one possible implementation of the second aspect, the fourth node determines the first feature information of the third node based on the feedback data by:
[0092] The fourth node determines the first feature information of the third node based on the feedback data and the preset classification algorithm.
[0093] In this implementation, the first feature information of the third node is determined based on the feedback data from the third node regarding the first observation and a preset classification algorithm. The feedback data provided by the third node can be used to help identify the first feature information of the third node.
[0094] In one possible implementation of the second aspect, the method further includes:
[0095] The fourth node receives the first observation from the third node;
[0096] The fourth node broadcasts the first observation to each of the plurality of child nodes.
[0097] In this implementation, the first observation is received by the fourth node from the third node and broadcast to each of the first node's multiple child nodes. Since the third node no longer needs to send the first observation to each of these child nodes, its processing load is reduced, and the first node is able to determine a more accurate reward prediction based on the first observation.
[0098] Thirdly, one embodiment of the present invention provides a data processing method, including:
[0099] The third node sends feedback data about the first observation to the fourth node, so that the fourth node determines the first feature information of the third node or the final prediction of the operation performed by the second node on the environment, wherein the first observation is in response to the operation of the second node indicating the state of the environment; the final prediction corresponds to the first feature information, and the final prediction is either the first prediction determined by the first node or the fourth prediction determined by the fourth node.
[0100] By receiving feedback data about the first observation from the third node, the fourth node can determine the first feature information of the third node based on the feedback data, thereby obtaining a reward prediction result that matches the first feature information of the third node.
[0101] In one possible implementation of the third aspect, the method further includes:
[0102] The third node sends the first observation to the first node or the fourth node.
[0103] Fourthly, one embodiment of the present invention provides a data processing method, comprising:
[0104] The second node receives a final prediction of the operation performed by the second node on the environment, wherein the final prediction corresponds to first feature information of a third node in the environment; the final prediction is based on a first observation and feedback data about the first observation, the first observation being in response to the operation of the second node indicating the state of the environment, and the final prediction is either a first prediction determined by the first node or a fourth prediction determined by the fourth node.
[0105] The second node is updated based on the final prediction.
[0106] By receiving the final prediction of the operation performed by the second node on the environment and updating it based on the final prediction, the second node is able to generate an operation that matches the first feature information.
[0107] Fifthly, one embodiment of the present invention provides a data processing method, comprising:
[0108] The fifth node sends a first request to the sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information;
[0109] The fifth node receives a first response from the sixth node, wherein the first response is determined by the sixth node based on the first request and instructs the sixth node to support the ability to train the first node;
[0110] The fifth node configures the first node and the sixth node to perform the training based on the first response.
[0111] Through the interaction between the fifth and sixth nodes, the fifth node can be configured based on the capabilities of the sixth node, thereby preparing the first and sixth nodes for training, which in turn enables the first node after training to determine more accurate reward prediction results.
[0112] In one possible implementation of the fifth aspect, the fifth node configuring the first node and the sixth node to perform the training based on the first response includes:
[0113] The fifth node determines the first configuration information and the second configuration information of the first node based on the first response, wherein the first configuration information indicates the type of the first node, and the second configuration information indicates the parameters used to train the first node.
[0114] The fifth node configures the first node based on the first configuration information and the second configuration information;
[0115] The fifth node notifies the sixth node of the configuration of the first node.
[0116] In this implementation, the fifth node uses the type of the first node and the parameters used to train the first node to configure the first node, thereby enabling the trained first node to determine a more accurate reward prediction result.
[0117] In one possible implementation of the fifth aspect, the fifth node determines the first configuration information of the first node by including:
[0118] The fifth node determines the type of the first node based on the first response, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information;
[0119] The fifth node notifies the sixth node of the configuration of the first node, including:
[0120] The fifth node informs the sixth node of the type of the first node.
[0121] In this implementation, the type of the first node is determined based on the first response and notified to the sixth node, which is beneficial for configuring the first node and thus helps improve the accuracy of the first node in determining the reward prediction result.
[0122] In one possible implementation of the fifth aspect, the first response includes an indication of feature information supported by the sixth node, and an indication of the sixth node including data required to train the first node;
[0123] The fifth node further determines the first configuration information of the first node by including:
[0124] The fifth node determines the number of candidate nodes to implement the first node based on the indication of the feature information supported by the sixth node;
[0125] The fifth node determines the second configuration information of the first node, including:
[0126] The fifth node determines the second configuration information based on the indication from the sixth node, which includes the data required to train the first node.
[0127] In this implementation, the fifth node configures the first node using the indications of the feature information supported by the sixth node and the indications of the sixth node including the data required to train the first node. Therefore, the configured first node matches well with the sixth node, thereby enabling the trained first node to determine a more accurate reward prediction result.
[0128] In one possible implementation of the fifth aspect, the second configuration information includes parameters of a first type and parameters of a second type, wherein the parameters of the first type remain unchanged during the training of the first node, and the parameters of the second type are updated based on training data provided by the sixth node during the training of the first node.
[0129] In this implementation, the parameters used to configure the first node are divided into different types: parameters that remain unchanged during the training of the first node and parameters that are updated during the training of the first node. Therefore, the configuration can be implemented in a more accurate way.
[0130] In one possible implementation of the fifth aspect, the parameters of the first type include at least one of the structure of the neural network for training the first node and the method for training the first node;
[0131] The parameters of the second type include at least one of the initial weights and biases of the neural network.
[0132] In one possible implementation of the fifth aspect, the fifth node configuring the first node based on the first configuration information and the second configuration information includes:
[0133] The fifth node configures network elements for implementing the first node and the connection between the sixth node and the first node based on the first configuration information and the second configuration information.
[0134] In this implementation, the first configuration information and the second configuration information are used to configure the network elements used to implement the first node and the connection between the sixth node and the first node. This is beneficial for configuring the first node, thereby improving the accuracy of the first node in determining the reward prediction result.
[0135] In one possible implementation of the fifth aspect, the fifth node notifying the sixth node of the configuration of the first node includes:
[0136] The fifth node sends a third message to the sixth node, wherein the third message indicates at least one path between the first node and the sixth node.
[0137] In this implementation, the fifth node sends information to the sixth node indicating at least one path between the first node and the sixth node. This facilitates sending training data to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0138] In one possible implementation of the fifth aspect, the method further includes:
[0139] The fifth node sends a fourth message to the sixth node, wherein the fourth message indicates the training dataset of the first node.
[0140] In this implementation, the fifth node sends information to the sixth node indicating the training dataset for the first node. This helps determine the training data to be provided to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0141] In one possible implementation of the fifth aspect, the fourth information further indicates the format of the data used to train the first node.
[0142] In this implementation, the format of the data used to train the first node is sent from the fifth node to the sixth node. This facilitates sending training data to the first node for training purposes, thereby improving the accuracy of the first node in determining the reward prediction result.
[0143] In one possible implementation of the fifth aspect, the first information includes at least one indication of at least one feature information, and the second information includes a minimum amount of data associated with each of the at least one feature information.
[0144] In this implementation, the fifth node sends at least one indication of at least one feature information and the minimum amount of data associated with each of the at least one feature information to the sixth node, which helps the training of the first node and thus improves the accuracy of the first node in determining the reward prediction result.
[0145] Sixthly, one embodiment of the present invention provides a data processing method, comprising:
[0146] The sixth node receives a first request from the fifth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information;
[0147] The sixth node determines a first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node;
[0148] The sixth node sends the first response to the fifth node.
[0149] Through the interaction between the fifth and sixth nodes, the sixth node can demonstrate its capabilities to the fifth node, and the fifth node can configure itself based on the sixth node's capabilities. This prepares the first and sixth nodes for training, enabling the first node to determine more accurate reward predictions after training.
[0150] In one possible implementation of the sixth aspect, the first response includes an indication of feature information supported by the sixth node, and an indication of the sixth node including data required to train the first node.
[0151] The sixth node determines the first response based on the first request, including:
[0152] The sixth node determines the indication of the feature information supported by the sixth node based on the first information;
[0153] The sixth node determines, based on the second information, the indication that includes the data required to train the first node.
[0154] In this implementation, the first response includes an indication of the feature information supported by the sixth node, determined based on the first information and the second information, and an indication of the sixth node including the data required to train the first node. This facilitates the configuration of the first node and thus improves the accuracy of the first node in determining the reward prediction result.
[0155] In one possible implementation of the sixth aspect, the method further includes:
[0156] The sixth node receives third information from the fifth node, wherein the third information indicates at least one path between the first node and the sixth node.
[0157] In this implementation, the fifth node sends information to the sixth node indicating at least one path between the first node and the sixth node. This facilitates sending training data to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0158] In one possible implementation of the sixth aspect, the method further includes:
[0159] The sixth node receives a training request from the first node, wherein the training request indicates the feature information of the first node;
[0160] The sixth node determines the training data for the first node based on the type of the first node and the training request.
[0161] The sixth node sends the training data to the first node based on the third information, so that the first node performs the training based on the training data.
[0162] In this implementation, the training data of the first node corresponding to the feature information of the first node is determined by the sixth node based on the type of the first node and the training request received from the first node, and sent to the first node based on the third information, which is beneficial to the training of the first node and thus helps to improve the accuracy of the first node in determining the reward prediction result.
[0163] In one possible implementation of the sixth aspect, the method further includes:
[0164] The sixth node receives fourth information from the fifth node, wherein the fourth information indicates the training dataset of the first node.
[0165] In this implementation, the fifth node sends information to the sixth node indicating the training dataset for the first node. This facilitates sending training data to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0166] In one possible implementation of the sixth aspect, the method further includes:
[0167] When the first node is determined to be of the first type and includes an integrated node associated with multiple feature information, the sixth node determines the first training dataset based on the fourth information, and the sixth node sends the training data of the first training dataset to the first node based on the third information, so that the first node performs the training based on the training data of the first training dataset.
[0168] When it is determined that the first node is of the second type and includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information, the sixth node determines the second training dataset based on the fourth information, and the sixth node sends the training data of the second training dataset to each of the multiple child nodes based on the third information, so that the first node performs the training based on the training data of the second training dataset.
[0169] The first training dataset includes multiple training data sets, each training data set including a first data portion and a second data portion, the second data portion indicating the feature information of the first data portion; the second training dataset includes multiple sets of training data, each set of training data being associated with predefined feature information among the multiple feature information sets.
[0170] In this implementation, the training dataset is determined by the sixth node based on the type of the first node and the third information, and sent to the first node based on the fourth information, so that the first node can perform training based on the training data of the training dataset, thereby improving the accuracy of the first node in determining the reward prediction result.
[0171] In one possible implementation of the sixth aspect, the type of the first node is predefined or notified by the sixth node.
[0172] In this implementation, the type of the first node is predefined or notified by the sixth node, which simplifies the signaling used to determine the training dataset and thus simplifies the training process of the first node.
[0173] In a seventh aspect, one embodiment of the present invention provides a data processing method, comprising:
[0174] The first node receives training data from the sixth node, wherein the training data corresponds to the feature information of the first node;
[0175] The first node performs training based on the training data from the sixth node.
[0176] In this implementation, the training data of the first node, which corresponds to the feature information of the first node, is received from the first node and used to perform training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0177] In one possible implementation of the seventh aspect, the method further includes:
[0178] The first node sends a training request to the sixth node, wherein the training request indicates the feature information of the first node.
[0179] In this implementation, the first node sends a training request to the sixth node to determine the training data for the first node to perform training, thereby helping to improve the accuracy of the first node in determining the reward prediction result.
[0180] In one possible implementation of the seventh aspect, the first node is of a first type and includes an integrated node associated with multiple feature information;
[0181] The first node receives the training data from the sixth node, including:
[0182] The first node receives training data from the sixth node in the first training dataset, wherein the first training dataset includes multiple training data, each training data includes a first data part and a second data part, and the second data part indicates the feature information of the first data part.
[0183] In this implementation, since the first node is trained using training data, which includes a first data part and a second data part indicating the feature information of the first data part, the reward prediction results provided by the first node are more accurate.
[0184] In one possible implementation of the seventh aspect, the first node is of the second type and includes multiple child nodes, each of the multiple child nodes being associated with one of the multiple feature information;
[0185] The first node receives the training data from the sixth node, including:
[0186] The first node receives training data from the sixth node in the second training dataset, wherein the second training dataset includes multiple sets of training data, and each set of training data is associated with predefined feature information in the multiple feature information.
[0187] In this implementation, since the first node is trained using a training dataset that includes multiple sets of training data, and each set of training data is associated with predefined feature information from multiple feature information, the reward prediction results provided by the first node are more accurate.
[0188] Eighthly, one embodiment of the present invention provides a data processing method, comprising:
[0189] The fifth node obtains a fourth indication from the third node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node;
[0190] The fifth node is configured based on the fourth instruction to execute the target task initiated by the third node.
[0191] Since the configuration is performed based on the fourth instruction indicating the type of feedback data about the first observation provided by the third node to execute the target task initiated by the third node, the feedback data provided by the third node can be used to help configure the system for executing the target task.
[0192] In one possible implementation of the eighth aspect, the fifth node obtaining the fourth instruction from the third node includes:
[0193] The fifth node sends a second request to the third node, wherein the second request is used to request the type of the feedback data;
[0194] The fifth node receives the fourth instruction from the third node.
[0195] In this implementation, the fifth node sends a second request to the third node to request the type of feedback data, in order to obtain the type of feedback data of the first observation provided by the third node, thereby configuring the system to execute the target task initiated by the third node, and thus enabling the system to be well configured to execute the target task.
[0196] In one possible implementation of the eighth aspect, the second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein...
[0197] The fourth instruction also indicates at least one of the items requested by the fifth node.
[0198] In this implementation, the fifth node sends a second request to the third node requesting at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, so as to configure the system for executing the target task initiated by the third node based on these items in a more flexible manner, thereby enabling the system to be well configured for executing the target task.
[0199] In one possible implementation of the eighth aspect, the method further includes:
[0200] The fifth node receives a third request from the third node, wherein the third request indicates description information corresponding to the target task.
[0201] In this implementation, sending a third request from the third node to the fifth node, which indicates the description information corresponding to the target task, is beneficial to the execution of the target task.
[0202] In one possible implementation of the eighth aspect, the fifth node performs the configuration based on the fourth instruction to execute the target task initiated by the third node, including:
[0203] The fifth node determines the first node and the second node to execute the target task based on the description information corresponding to the target task;
[0204] The fifth node, based on the type of the feedback data and the first node, determines the operations to be performed by the first node, the second node, and the fourth node to execute the target task, wherein the fourth node is used to provide predictions for the observations provided by the third node;
[0205] The fifth node configures network elements and connections for implementing the operations performed by the first node, the second node, and the fourth node, based on the operations performed by the first node, the second node, and the fourth node.
[0206] The fifth node, based on the network elements and the connections used to implement the operations performed by the first node, the second node, and the fourth node, notifies the third node of the configuration to perform the target task.
[0207] In this implementation, by determining the first and second nodes to execute the target task, determining the operations to be performed by the first, second, and fourth nodes, configuring the network elements and connections to implement the operations performed by the first, second, and fourth nodes, and notifying the third node of the configuration to execute the target task, the system for executing the target task can be well configured.
[0208] In one possible implementation of the eighth aspect, the fifth node determining the first node to execute the target task based on the description information corresponding to the target task includes:
[0209] The fifth node determines at least one candidate node that has performed a historical task as the first node based on the description information corresponding to the target task, wherein the similarity between the historical task and the target task is above a preset threshold.
[0210] Since the first node is identified as at least one candidate node that has executed a historical task, and the similarity between the historical task and the target task is above a preset threshold, reusing the first node makes configuration easier.
[0211] In one possible implementation of the eighth aspect, the fifth node determining the first node to execute the target task based on the description information corresponding to the target task includes:
[0212] The fifth node determines the type of the first node based on the description information corresponding to the target task, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information;
[0213] When the first node is determined to be of the second type, the fifth node determines all available candidate nodes as child nodes of the first node.
[0214] In this implementation, by determining the type of the first node based on the description information corresponding to the target task, the system for executing the target task can be well configured.
[0215] In one possible implementation of the eighth aspect, the fifth node notifies the third node of the configuration for performing the target task based on the network elements and the connection used to implement the operations performed by the first node, the second node, and the fourth node, including:
[0216] The fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the first node and the connection between the third node and the first node; or
[0217] The fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the fourth node and the connection between the third node and the fourth node.
[0218] In this implementation, the fifth node sends a third configuration notification to the third node. This third configuration notification indicates the network elements used to implement the operations performed by the first or fourth node, as well as the connection between the third node and the first or fourth node, and enables the third node to be well configured for the system to perform the target task.
[0219] In one possible implementation of the eighth aspect, the method further includes:
[0220] The fifth node sends a first configuration notification to the first node, wherein the first configuration notification indicates the operation to be performed by the first node;
[0221] The fifth node sends a second configuration notification to the second node, wherein the second configuration notification indicates the operation to be performed by the second node;
[0222] The fifth node sends a fourth configuration notification to the fourth node, wherein the fourth configuration notification indicates the operation to be performed by the fourth node.
[0223] In this implementation, the fifth node sends different configuration notifications to the first, second, and fourth nodes of the system used to execute the target task, respectively, in order to better configure the first, second, and fourth nodes and enable these nodes to perform corresponding operations.
[0224] In one possible implementation of the eighth aspect, the type of the feedback data includes a first feedback type and a second feedback type, wherein the first feedback type of the feedback data indicates the third node's evaluation of the observation, and the second feedback type of the feedback data indicates the feature information of the third node.
[0225] In this implementation, the feedback data has two different types, so the operation to be performed by the nodes of the system used to perform the target task can be determined based on the type of feedback data, thereby better configuring the nodes of the system used to perform the target task.
[0226] Ninthly, one embodiment of the present invention provides a data processing method, comprising:
[0227] The first node receives a first configuration notification from the fifth node, wherein the first configuration notification indicates an operation to be performed by the first node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node;
[0228] The first node is configured based on the first configuration notification.
[0229] The configuration of the first node is based on a first configuration notification, which instructs the first node to perform operations to execute a target task initiated by the third node. These operations are determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node. The feedback data provided by the third node can be used to help configure the system for executing the target task.
[0230] In a tenth aspect, one embodiment of the present invention provides a data processing method, comprising:
[0231] The second node receives a second configuration notification from the fifth node, wherein the second configuration notification indicates an operation to be performed by the second node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node;
[0232] The second node is configured based on the second configuration notification.
[0233] The configuration of the second node is based on a first configuration notification, which instructs the second node to perform operations to execute a target task initiated by the third node. These operations are determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node. The feedback data provided by the third node can be used to help configure the system for executing the target task.
[0234] Eleventhly, one embodiment of the present invention provides a data processing method, comprising:
[0235] The third node sends a fourth instruction to the fifth node so that the fifth node can be configured based on the fourth instruction to execute the target task initiated by the third node, wherein the fourth instruction indicates the type of feedback data provided by the third node.
[0236] Since the configuration is performed based on the fourth instruction indicating the type of feedback data about the first observation provided by the third node to execute the target task initiated by the third node, the feedback data provided by the third node can be used to help configure the system for executing the target task.
[0237] In one possible implementation of the eleventh aspect, the method further includes:
[0238] The third node receives a second request from the fifth node, wherein the second request is used to request the type of the feedback data.
[0239] In this implementation, the third node receives a second request from the fifth node for the type of feedback data to obtain the type of feedback data from the first observation provided by the third node, thereby configuring it to execute the target task initiated by the third node, and thus enabling the system to be well configured to execute the target task.
[0240] In one possible implementation of the eleventh aspect, the second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein...
[0241] The fourth instruction also indicates at least one of the items requested by the fifth node.
[0242] In this implementation, the third node receives a second request from the fifth node for requesting at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, thereby configuring the system to execute the target task initiated by the third node based on these items, and thus enabling the system to be well configured to execute the target task.
[0243] In one possible implementation of the eleventh aspect, the method further includes:
[0244] The third node sends a third request to the fifth node, wherein the third request indicates description information corresponding to the target task.
[0245] In this implementation, sending a third request from the third node to the fifth node, which indicates the description information corresponding to the target task, is beneficial to the execution of the target task.
[0246] In one possible implementation of the eleventh aspect, the method further includes:
[0247] The third node receives a third configuration notification from the fifth node, wherein the third configuration notification indicates a network element for implementing an operation performed by the first node and a connection between the third node and the first node; or the third configuration notification indicates a network element for implementing an operation performed by the fourth node and a connection between the third node and the fourth node;
[0248] The third node is configured based on the third configuration notification.
[0249] In this implementation, the fifth node sends a third configuration notification to the third node. This third configuration notification indicates the network elements used to implement the operations performed by the first or fourth node, as well as the connection between the third node and the first or fourth node, and enables the third node to be well configured for the system to perform the target task.
[0250] In a twelfth aspect, one embodiment of the present invention provides a data processing method, comprising:
[0251] The fourth node receives a fourth configuration notification from the fifth node, wherein the fourth configuration notification indicates an operation to be performed by the fourth node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node;
[0252] The fourth node is configured based on the fourth configuration notification.
[0253] The configuration of the fourth node is based on a first configuration notification, which instructs the fourth node to perform operations to execute a target task initiated by the third node. These operations are determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node. The feedback data provided by the third node can be used to help configure the system for executing the target task.
[0254] In a thirteenth aspect, one embodiment of the present invention provides a data processing method, comprising:
[0255] The first node acquires a first observation, wherein the first observation is in response to the operation of the second node indicating the state of the environment;
[0256] The third node sends feedback data about the first observation to the fourth node;
[0257] The first node obtains a first prediction for the operation of the second node based on the first observation, wherein the first prediction is related to the first feature information of the third node in the environment, the first prediction is a final prediction or a third prediction, the third prediction is generated by the fourth node as the final prediction, the final prediction corresponds to the first feature information, and the final prediction is based on the first observation and the feedback data.
[0258] The fourth node receives the feedback data from the third node;
[0259] The second node receives the final prediction of the operation performed by the second node on the environment from the first node or the fourth node;
[0260] The second node is updated based on the final prediction.
[0261] By obtaining the final prediction related to the first feature information of the third node in the environment, which is used to update the second node, the first or fourth node can provide a reward prediction result, thereby enabling the second node to generate an operation that matches the first feature information of the third node.
[0262] In a fourteenth aspect, one embodiment of the present invention provides a data processing method, comprising:
[0263] The fifth node sends a first request to the sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information;
[0264] The sixth node receives the first request from the fifth node;
[0265] The sixth node determines a first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node;
[0266] The sixth node sends the first response to the fifth node;
[0267] The fifth node receives the first response from the sixth node;
[0268] The fifth node configures the first node and the sixth node based on the first response to perform the training;
[0269] The sixth node sends training data to the first node, wherein the training data corresponds to the feature information of the first node;
[0270] The first node receives the training data from the sixth node;
[0271] The first node performs the training based on the training data from the sixth node.
[0272] After configuring the first node and the sixth node based on the first response indicating the sixth node's ability to support training the first node, training the first node based on training data corresponding to the feature information of the first node improves the accuracy of the first node in determining the reward prediction result.
[0273] In a fifteenth aspect, one embodiment of the present invention provides a data processing method, comprising:
[0274] The third node sends a fourth indication to the fifth node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node;
[0275] The fifth node obtains the fourth instruction from the third node;
[0276] The fifth node is configured based on the fourth instruction to execute the target task initiated by the third node.
[0277] The configuration is performed based on a fourth instruction indicating the type of feedback data provided by the third node regarding the first observation, in order to execute the target task initiated by the third node. The feedback data provided by the third node can be used to help configure the system for executing the target task.
[0278] In a sixteenth aspect, one embodiment of the present invention provides a data processing system, comprising:
[0279] The first node is used to execute the data processing method according to the first aspect or any possible implementation thereof;
[0280] The second node is used to execute the data processing method described in the fourth aspect;
[0281] The third node is used to execute the data processing method described in accordance with the third aspect or a possible implementation thereof;
[0282] The fourth node is used to execute the data processing method described in accordance with the second aspect or any possible implementation thereof.
[0283] In a seventeenth aspect, one embodiment of the present invention provides a data processing system, comprising:
[0284] The first node is used to execute the data processing method according to the seventh aspect or any possible implementation thereof;
[0285] The fifth node is used to execute the data processing method described in accordance with the fifth aspect or any possible implementation thereof;
[0286] The sixth node is used to execute the data processing method described in accordance with the sixth aspect or any possible implementation thereof.
[0287] In an eighteenth aspect, one embodiment of the present invention provides a data processing system, comprising:
[0288] The first node is used to execute the data processing method described in the ninth aspect;
[0289] The second node is used to execute the data processing method according to the tenth aspect;
[0290] The third node is used to execute the data processing method according to the eleventh aspect or any possible implementation thereof;
[0291] The fourth node is used to execute the data processing method described in aspect 12;
[0292] The fifth node is used to execute the data processing method described in accordance with the eighth aspect or any possible implementation thereof.
[0293] In a nineteenth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing a data processing method according to the first aspect or any possible implementation thereof, or according to the seventh aspect or any possible implementation thereof, or according to the data processing method described in the ninth aspect.
[0294] In a twentieth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing the data processing method according to the fourth or tenth aspect.
[0295] In a twentieth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing the data processing method according to the third aspect or any possible implementation thereof, or according to the eleventh aspect or any possible implementation thereof.
[0296] In a twentieth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing the data processing method according to the second aspect or any possible implementation thereof, or according to the twelfth aspect.
[0297] In a twentieth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing the data processing method according to the fifth aspect or any possible implementation thereof, or according to the eighth aspect or any possible implementation thereof.
[0298] In a twentieth aspect, one embodiment of the present invention provides a data processing apparatus comprising various modules for performing the data processing method according to the sixth aspect or any possible implementation thereof.
[0299] In a twentieth aspect, one embodiment of the present invention provides a first node including a processing circuit for performing a data processing method according to the first aspect or any possible implementation thereof, or according to the seventh aspect or any possible implementation thereof and the ninth aspect.
[0300] In a twentieth aspect, one embodiment of the present invention provides a second node including processing circuitry for performing the data processing method according to the fourth aspect or the tenth aspect.
[0301] In a twentieth aspect, one embodiment of the present invention provides a third node including processing circuitry for performing the data processing method according to the third aspect or any possible implementation thereof, or according to the eleventh aspect or any possible implementation thereof.
[0302] In a twentieth aspect, one embodiment of the present invention provides a fourth node including processing circuitry for performing the data processing method according to the second aspect or any possible implementation thereof, or according to the twelfth aspect.
[0303] In a twentieth aspect, one embodiment of the present invention provides a fifth node including a processing circuit for performing a data processing method according to the fifth aspect or any possible implementation thereof, or according to the eighth aspect or any possible implementation thereof.
[0304] In a thirtieth aspect, one embodiment of the present invention provides a sixth node including processing circuitry for performing the data processing method according to the sixth aspect or any possible implementation thereof.
[0305] In a thirty-first aspect, one embodiment of the present invention provides a computer-readable medium storing computer-executable instructions, which, when executed by a processor, cause the processor to perform a data processing method according to the first aspect or any possible implementation of the first to fifteenth aspects.
[0306] In a thirty-second aspect, one embodiment of the present invention provides a computer program product including computer-executable instructions, which, when executed by a processor, cause the processor to perform a data processing method according to the first aspect or any possible implementation of the first to fifteenth aspects.
[0307] This invention provides a data processing method and related products. By obtaining a final prediction for updating a second node related to a first feature information of a third node in the environment, a first or fourth node can provide a reward prediction result, thereby enabling the second node to generate an operation matching the first feature information of the third node. Furthermore, after configuring the first node and the sixth node based on a first response indicating the ability of a sixth node to support training the first node, training the first node based on training data corresponding to the feature information of the first node improves the accuracy of the first node in determining the reward prediction result. Moreover, by performing configuration based on a fourth instruction indicating the type of feedback data provided by the third node regarding a first observation to execute a target task initiated by the third node, the feedback data provided by the third node can be used to help configure the system for executing the target task. Attached Figure Description
[0308] The following figures illustrate exemplary embodiments of the present invention by way of example, in which:
[0309] Figure 1 This is a simplified schematic diagram of a communication system according to one or more embodiments of the present invention.
[0310] Figure 2This is a schematic diagram of an exemplary communication system according to one or more embodiments of the present invention.
[0311] Figure 3 This is a schematic diagram of the basic component structure of a communication system according to one or more embodiments of the present invention.
[0312] Figure 4 This is a block diagram of a device in a communication system according to one or more embodiments of the present invention.
[0313] Figure 5 It is a schematic framework of a traditional RLHF system.
[0314] Figure 6 This is an illustrative framework for an RLCF system provided by the present invention.
[0315] Figure 7A This is an illustrative framework of an RLCF system α with multiple ERPs according to one or more embodiments of the present invention.
[0316] Figure 7B This is an illustrative framework of an RLCF system β having only one integrated ERP (I-ERP) according to one or more embodiments of the present invention.
[0317] Figure 8 This invention provides a workflow for an RLCF system.
[0318] Figure 9A This is a flowchart of a first data processing method according to one or more embodiments of the present invention.
[0319] Figure 9B It corresponds to Figure 9A One possible implementation of the process shown.
[0320] Figure 10A This is a flowchart of a second data processing method according to one or more embodiments of the present invention.
[0321] Figure 10B It corresponds to Figure 10A One possible implementation of the process shown.
[0322] Figure 11 This is a flowchart of a third data processing method according to one or more embodiments of the present invention.
[0323] Figure 12A This is a schematic diagram of a fourth data processing method according to one or more embodiments of the present invention.
[0324] Figure 12B It corresponds to Figure 12A One possible implementation of the process shown.
[0325] Figure 13A This is a schematic diagram of a fifth data processing method according to one or more embodiments of the present invention.
[0326] Figure 13B It corresponds to Figure 13A One possible implementation of the process shown.
[0327] Figure 14A This is a schematic diagram of a sixth data processing method according to one or more embodiments of the present invention.
[0328] Figure 14B It corresponds to Figure 14A One possible implementation of the process shown.
[0329] Figure 15A This is a schematic diagram of a seventh data processing method according to one or more embodiments of the present invention.
[0330] Figure 15B It corresponds to Figure 15A One possible implementation of the process shown.
[0331] Figure 16 This is a schematic diagram of a data processing apparatus according to one or more embodiments of the present invention.
[0332] Figure 17 This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention.
[0333] Figure 18 This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention.
[0334] Figure 19 This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention.
[0335] Figure 20 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0336] Figure 21 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0337] Figure 22 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0338] Figure 23 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0339] Figure 24This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0340] Figure 25 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0341] Figure 26 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention.
[0342] Figure 27 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. Detailed Implementation
[0343] In the following description, reference is made to the accompanying drawings, which form part of this invention, and which illustrate by way of description specific aspects of embodiments of the invention or aspects in which embodiments of the invention may be used. It should be understood that embodiments of the invention can be used in other aspects and include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.
[0344] To aid in understanding the present invention, examples of wireless communication systems and devices are described below.
[0345] Exemplary communication systems and devices
[0346] refer to Figure 1 A simplified schematic diagram of a communication system is provided as an illustrative, not limiting, example. Communication system 100 includes a radio access network 120. Radio access network 120 can be a next-generation (e.g., sixth-generation, 6G, or later) radio access network or a traditional (e.g., 5G, 4G, 3G, or 2G) radio access network. In radio access network 120, one or more electric devices (EDs) 110a to 120j (generally referred to as 110) can be interconnected with each other or connected to one or more network nodes (170a, 170b, generally referred to as 170). Core network 130 can be part of the communication system and can depend on or be independent of the radio access technology used in communication system 100. Furthermore, communication system 100 includes a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160.
[0347] Figure 2An exemplary communication system 100 is illustrated. Generally, the communication system 100 enables multiple wireless or wired components to transmit data and other content. The purpose of the communication system 100 may be to provide content such as voice, data, video, and / or text via broadcast, multicast, and unicast. The communication system 100 can operate by sharing resources such as carrier spectrum bandwidth among its constituent components. The communication system 100 may include terrestrial communication systems and / or non-terrestrial communication systems. The communication system 100 can provide a wide range of communication services and applications (e.g., earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery, and mobility). The communication system 100 can provide high availability and robustness through the joint operation of terrestrial and non-terrestrial communication systems. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system can create a heterogeneous network that can be considered as comprising multiple layers. Compared to traditional communication networks, heterogeneous networks can achieve better overall performance through efficient multi-link joint operation, more flexible function sharing, and faster physical layer link switching between terrestrial and non-terrestrial networks.
[0348] Terrestrial and non-terrestrial communication systems can be considered as subsystems of a communication system. In the example shown, communication system 100 includes electronic devices (EDs) 110a to 110d (generally referred to as ED 110), radio access networks (RANs) 120a and 120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. RANs 120a and 120b include corresponding base stations (BSs) 170a and 170b, which can generally be referred to as terrestrial transmit and receive points (T-TRPs) 170a and 170b. The non-terrestrial communication network 120c includes an access node 120c, which can generally be referred to as a non-terrestrial transmit and receive point (NT-TRP) 172.
[0349] Any ED 110 can alternatively or additionally be used to connect to, access, or communicate with any other T-TRP 170a and 170b, NT-TRP 172, Internet 150, core network 130, PSTN 140, other network 160, or any combination thereof. In some examples, ED 110a can communicate uplink and / or downlink with T-TRP 170a via interface 190a. In some examples, ED 110a, 110b, and 110d can also communicate directly with each other via one or more sidelink air interfaces 190b. In some examples, ED 110d can communicate uplink and / or downlink with NT-TRP 172 via interface 190c.
[0350] Air interfaces 190a and 190b can use similar communication technologies, such as any suitable wireless access technology. For example, communication system 100 can implement one or more channel access methods in air interfaces 190a and 190b, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), or single-carrier FDMA (SC-FDMA). Air interfaces 190a and 190b can utilize other higher-dimensional signal spaces, which may involve combinations of orthogonal and / or non-orthogonal dimensions.
[0351] The air interface 190c enables communication between the ED 110d and one or more NT-TRP 172s via a wireless link or simply via a link. In some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection for multicast transmission between a group of EDs and one or more NT-TRPs.
[0352] RANs 120a and 120b communicate with the core network 130 to provide various services, such as voice, data, and other services, to EDs 110a, 110b, and 110c. RANs 120a and 120b and / or the core network 130 can communicate directly or indirectly with one or more other RANs (not shown), which may or may not be directly served by the core network 130, and may or may not use the same radio access technology as RANs 120a and / or RAN 120b. The core network 130 can also serve as a gateway access between (i) RANs 120a and 120b or EDs 110a, 110b, and 110c or both RANs and EDs and (ii) other networks (e.g., PSTN 140, Internet 150, and other networks 160). Additionally, some or all of EDs 110a, 110b, and 110c may include the ability to communicate with different wireless networks via different radio links using different radio technologies and / or protocols. ED 110a, 110b, and 110c can communicate with a service provider or exchange (not shown) via a wired communication channel and with the Internet 150, but not wirelessly (or also wirelessly). PSTN 140 may include a circuit-switched telephone network for providing plain old telephone service (POTS). The Internet 150 may include a network of computers and subnets (intranets) or both, and also includes protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), and User Datagram Protocol (UDP). ED 110a, 110b, and 110c may be multimode devices capable of operating under various wireless access technologies and may include multiple transceivers required to support such operation.
[0353] Basic component structure
[0354] Figure 3Another example of the ED 110 and base stations 170a, 170b, and / or 170c is shown. The ED 110 is used to connect people, objects, machines, etc. The ED 110 can be widely used in various scenarios, such as cellular communication, device-to-device (D2D), vehicle-to-everything (V2X), peer-to-peer (P2P), machine-to-machine (M2M), machine-type communications (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery, and mobility.
[0355] Each ED 110 represents any suitable end-user equipment for wireless operation and may include (or be referred to as): user equipment / device (UE), wireless transmit / receive unit (WTRU), mobile station, fixed or mobile subscriber unit, cellular phone, station (STA), machine-type communication (MTC) device, personal digital assistant (PDA), smartphone, laptop, computer, tablet, wireless sensor, consumer electronics device, smart book, vehicle, automobile, truck, bus, train, IoT device, or industrial equipment or apparatus of the foregoing (e.g., communication module, modem, or chip), etc. Future generations of ED 110 may be referred to using other terms. Base stations 170a and 170b are T-TRPs, referred to below as T-TRP 170. Alternatively... Figure 3 As shown, NT-TRP is referred to as NT-TRP 172 below. Each ED 110 connected to T-TRP 170 and / or NT-TRP 172 can be dynamically or semi-statically turned on (i.e., established, activated, or enabled), turned off (i.e., released, deactivated, or disabled), and / or configured in response to one or more of connectivity availability and connectivity necessity.
[0356] ED 110 includes a transmitter 201 and a receiver 203 coupled to one or more antennas 204. Only one antenna 204 is shown in the figure. One, part, or all of the antennas may also be panels. The transmitter 201 and receiver 203 may be integrated as a transceiver, etc. The transceiver is used to modulate data or other content for transmission through at least one antenna 204 or a network interface controller (NIC). The transceiver is also used to demodulate data or other content received through at least one antenna 204. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or for processing signals received wirelessly or wiredly. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals.
[0357] ED 110 includes at least one memory 208. Memory 208 stores instructions and data used, generated, or collected by ED 110. For example, memory 208 may store software instructions or modules executed by one or more processing units 210 for implementing some or all of the functions and / or embodiments described herein. Each memory 208 includes any suitable one or more volatile and / or non-volatile storage and retrieval devices. Any suitable type of memory can be used, such as random access memory (RAM), read-only memory (ROM), hard disk, optical disk, subscriber identity module (SIM) card, memory stick, secure digital (SD) memory card, on-processor cache, etc.
[0358] ED 110 may also include one or more input / output devices (not shown) or interfaces (e.g., Figure 1 (Wired interface of Internet 150 in the network). Input / output devices support interaction with the user or other devices in the network. Each input / output device includes any suitable structure for providing or receiving information from the user, such as a speaker, microphone, keypad, keyboard, display, or touchscreen, including network interface communication.
[0359] ED 110 also includes a processor 210 for performing operations related to: operations related to preparing uplink transmissions to NT-TRP 172 and / or T-TRP 170; operations related to processing downlink transmissions received from NT-TRP 172 and / or T-TRP 170; and operations related to processing sidelink transmissions to and from another ED 110. Processing operations related to preparing uplink transmissions may include operations such as encoding, modulation, transmit beamforming, and generating symbols for transmission. Processing operations related to processing downlink transmissions may include operations such as receive beamforming, demodulation, and decoding of received symbols. According to an embodiment, receiver 203 may receive downlink transmissions (possibly using receive beamforming), and processor 210 may extract signaling from the downlink transmissions (e.g., by detecting and / or decoding signaling). Examples of signaling may be reference signals transmitted by NT-TRP 172 and / or T-TRP 170. In some embodiments, processor 276 performs transmit beamforming and / or receive beamforming based on beam direction indications (e.g., beam angle information (BAI)) received from T-TRP 170. In some embodiments, processor 210 may perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as operations related to detecting synchronization sequences, decoding, and acquiring system information. In some embodiments, for example, processor 210 may use reference signals received from NT-TRP 172 and / or T-TRP 170 to perform channel estimation.
[0360] Although not shown, processor 210 may be part of transmitter 201 and / or receiver 203. Although not shown, memory 208 may be part of processor 210.
[0361] The processor 210 and the processing components of the transmitter 201 and receiver 203 may each be implemented by the same or different one or more processors for executing instructions stored in memory (e.g., memory 208). Alternatively, some or all of the processing components in the processor 210 and the transmitter 201 and receiver 203 may be implemented using dedicated circuitry, such as a programmable field-programmable gate array (FPGA), a graphics processing unit (GPU), or an application-specific integrated circuit (ASIC).
[0362] In some implementations, T-TRP 170 may be referred to by other names, such as base station, base transceiver station (BTS), wireless base station, network node, network device, network-side device, transmit / receive node, NodeB, evolved NodeB (eNodeB or eNB), home eNodeB, generation NodeB (gNB), transmission point (TP), site controller, access point (AP) or wireless router, relay station, remote radio head, ground node, ground network device or ground base station, base band unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), location node, etc. T-TRP 170 can be a macro BS, pico BS, relay node, host node, or a combination thereof. T-TRP 170 may refer to the aforementioned equipment or to a component within the aforementioned equipment (e.g., a communication module, modem, or chip).
[0363] In some embodiments, the various parts of T-TRP 170 may be distributed. For example, some modules in T-TRP 170 may be located remotely from the device housing the antenna of T-TRP 170 and may be coupled to the device housing the antenna via a communication link (not shown) sometimes referred to as the fronthaul (e.g., a common public radio interface (CPRI)). Therefore, in some embodiments, the term "T-TRP 170" may also refer to modules on the network side that perform processing operations such as ED 110 location determination, resource allocation (scheduling), message generation, and encoding / decoding, which are not necessarily part of the device housing the antenna of T-TRP 170. These modules may also be coupled to other T-TRPs. In some embodiments, T-TRP 170 may actually be multiple T-TRPs operating together to serve ED 110 through cooperative multicast or similar methods.
[0364] T-TRP 170 includes at least one transmitter 252 and at least one receiver 254 coupled to one or more antennas 256. Only one antenna 256 is shown in the figure. One, some, or all of the antennas may also be panels. The transmitter 252 and receiver 254 may be integrated as a transceiver. T-TRP 170 also includes a processor 260 for performing operations related to: preparing downlink transmissions to ED 110, processing uplink transmissions received from ED 110, preparing backhaul transmissions to NT-TRP 172, and processing transmissions received from NT-TRP 172 via backhaul. Processing operations related to preparing downlink or backhaul transmissions may include operations such as encoding, modulation, precoding (e.g., MIMO precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the uplink or backhaul may include operations such as receive beamforming, demodulation, and decoding of received symbols. Processor 260 can also perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as generating the contents of a synchronization signal block (SSB), generating system information, etc. In some embodiments, processor 260 also generates a beam direction indication, such as a BAI, which can be scheduled for transmission by scheduler 253. Processor 260 can perform other network-side processing operations described herein, such as determining the location of ED 110, determining the location for deploying NT-TRP 172, etc. In some embodiments, processor 260 can generate signaling to configure one or more parameters of ED 110 and / or one or more parameters of NT-TRP 172, etc. Any signaling generated by processor 260 is transmitted by transmitter 252. It should be noted that the term "signaling" used herein can also be referred to as control signaling. Dynamic signaling can be transmitted in control channels such as the physical downlink control channel (PDCCH), while static or semi-static higher-layer signaling can be included in data packets that are transmitted in data channels such as the physical downlink shared channel (PDSCH).
[0365] Scheduler 253 may be coupled to processor 260. Scheduler 253 may be included within or operate separately from T-TRP 170. T-TRP 170 may schedule uplink, downlink, and / or backlink transmissions, including issuing scheduling authorizations and / or configuring schedule-free (“configuration authorization”) resources. T-TRP 170 also includes memory 258 for storing information and data. Memory 258 stores instructions and data used, generated, or collected by T-TRP 170. For example, memory 258 may store software instructions or modules executed by processor 260 for implementing some or all of the functions and / or embodiments described herein.
[0366] Although not shown, processor 260 may be part of transmitter 252 and / or receiver 254. Furthermore, although not shown, processor 260 may implement scheduler 253. Although not shown, memory 258 may be part of processor 260.
[0367] The processor 260, scheduler 253, and processing components of transmitter 252 and receiver 254 may each be implemented by the same or different one or more processors for executing instructions stored in memory (e.g., memory 258). Alternatively, some or all of the processing components of processor 260, scheduler 253, and transmitter 252 and receiver 254 may be implemented using dedicated circuitry, such as FPGA, GPU, or ASIC.
[0368] Although the NT-TRP 172 is shown as a drone for example only, it can be implemented in any suitable non-terrestrial form. Furthermore, in some implementations, the NT-TRP 172 may have other names, such as a non-terrestrial node, a non-terrestrial network device, or a non-terrestrial base station. The NT-TRP 172 includes a transmitter 272 and a receiver 274 coupled to one or more antennas 280. Only one antenna 280 is shown in the figure. One, some, or all of the antennas may also be panels. The transmitter 272 and receiver 274 may be integrated as a transceiver. The NT-TRP 172 also includes a processor 276 for performing operations related to: preparing downlink transmissions to be sent to ED 110, processing uplink transmissions received from ED 110, preparing backhaul transmissions to be sent to T-TRP 170, and processing transmissions received from T-TRP 170 via backhaul. Processing operations related to preparing downlink or backhaul transmissions may include operations such as encoding, modulation, precoding (e.g., MIMO precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing receive transmissions in the uplink or backhaul may include operations such as receive beamforming, demodulation, and decoding of received symbols. In some embodiments, processor 276 performs transmit beamforming and / or receive beamforming based on beam direction information (e.g., BAI) received from T-TRP 170. In some embodiments, processor 276 may generate signaling to configure one or more parameters of ED 110. In some embodiments, NT-TRP 172 implements physical layer processing but does not implement higher-level functions such as medium access control (MAC) or radio link control (RLC) layer functions. Since this is only an example, in general, NT-TRP 172 may implement higher-level functions in addition to physical layer processing.
[0369] The NT-TRP 172 also includes a memory 278 for storing information and data. Although not shown, a processor 276 may be part of the transmitter 272 and / or receiver 274. Although not shown, the memory 278 may be part of the processor 276.
[0370] The processor 276 and the processing components of the transmitter 272 and receiver 274 may each be implemented by the same or different one or more processors for executing instructions stored in memory (e.g., memory 278). Alternatively, some or all of the processing components of the processor 276 and the transmitter 272 and receiver 274 may be implemented using dedicated circuitry, such as a programmable FPGA, GPU, or ASIC. In some embodiments, the NT-TRP 172 may actually be multiple NT-TRPs operating together to serve ED 110 via cooperative multicast or similar methods.
[0371] T-TRP 170, NT-TRP 172 and / or ED 110 may include other components, but these components have been omitted for clarity.
[0372] Basic module structure
[0373] One or more steps of the methods in the embodiments provided herein can be derived from... Figure 4 The corresponding unit or module is executed. Figure 4 Units or modules in devices such as ED 110, T-TRP 170, or NT-TRP 172 are illustrated. For example, signals can be transmitted by a transmitting unit or transmitting module. Signals can be received by a receiving unit or receiving module. Signals can be processed by a processing unit or processing module. Other steps can be performed by an artificial intelligence (AI) module or a machine learning (ML) module. The corresponding units or modules can be implemented using hardware, one or more components or devices executing software, or a combination thereof. For example, one or more of these units or modules can be integrated circuits, such as programmable FPGAs, GPUs, or ASICs. It should be understood that if these modules are implemented using software executed by a processor, etc., then these modules can be retrieved by the processor, wholly or partially, individually or collectively, for processing, in one or more instances, and these modules themselves can include instructions for further deployment and instantiation.
[0374] Further details regarding ED 110, T-TRP 170, and NT-TRP 172 are known to those skilled in the art. Therefore, these details are omitted herein.
[0375] 6G Smart Air Interface
[0376] An air interface typically includes multiple components and associated parameters that collectively specify how transmissions are sent and / or received over a wireless communication link between two or more communication devices. For example, an air interface may include one or more waveforms, one or more frame structures, one or more multiple access schemes, one or more protocols, one or more coding schemes, and / or one or more modulation schemes that define the transmission of information (e.g., data) over the wireless communication link. The wireless communication link may support links between a radio access network and user equipment (e.g., a Uu link), and / or it may support links between devices, such as links between two user equipments (e.g., a sidelink), and / or it may support links between non-terrestrial (NT) communication networks and user equipment (UE). Below are some examples of the components mentioned above:
[0377] Waveform components can specify the shape and form of the signal being transmitted. Waveform options can include orthogonal multiple access (OFDM) and non-orthogonal multiple access (NOA) waveforms. Non-limiting examples of such waveform options include orthogonal frequency division multiplexing (OFDM), filtered OFDM (f-OFDM), time-domain windowed OFDM, filter bank multicarrier (FBMC), universal filtered multicarrier (UFMC), generalized frequency division multiplexing (GFDM), wavelet packet modulation (WPM), faster than Nyquist (FTN) waveforms, and low peak-to-average power ratio (LPPR) waveforms (WF).
[0378] The frame structure component can specify the configuration of a frame or a group of frames. The frame structure component can indicate one or more of the following parameters for a frame or group of frames: time, frequency, pilot signature, code, or other parameters. Further details about the frame structure will be discussed below.
[0379] Multiple access scheme components can specify multiple access technology options, including technologies that limit how communication devices share the common physical channel, such as: time division multiple access (TDMA), frequency division multiple access (FDMA), code division multiple access (CDMA), single carrier frequency division multiple access (SC-FDMA), low density signature multicarrier code division multiple access (LDS-MC-CDMA), non-orthogonal multiple access (NOMA), pattern division multiple access (PDMA), lattice partition multiple access (LPMA), resource spread multiple access (RSMA), and sparse code multiple access (SCMA). In addition, multiple access technology options may include: scheduled access and unscheduled access (also known as unlicensed access); non-orthogonal multiple access and orthogonal multiple access, such as via dedicated channel resources (e.g., not shared among multiple communication devices); contention-based shared channel resources and non-contention-based shared channel resources; and access based on sensing radio.
[0380] The Hybrid Automatic Repeat Request (HARQ) protocol component can specify how transmissions and / or retransmissions are performed. Non-limiting examples of transmission and / or retransmission mechanism options include mechanisms for specifying the size of the scheduled data pipeline, signaling mechanisms for transmission and / or retransmission, and retransmission mechanisms themselves.
[0381] Encoding and modulation components can specify how the information being transmitted can be encoded / decoded and modulated / demodulated for transmitting / receiving purposes. Encoding can refer to methods of error detection and forward error correction. Non-limiting examples of encoding options include turbo lattice codes, turbo product codes, fountain codes, low-density parity-check codes, and polar codes. Modulation can simply refer to constellations (e.g., including modulation techniques and orders), or more specifically to various types of advanced modulation methods, such as layered modulation and low PAPR modulation.
[0382] In some embodiments, the air interface can be a "one-size-fits-all" concept. For example, once the air interface is defined, the components within it cannot be changed or adapted. In some implementations, only a limited number of parameters or modes of the air interface can be configured, such as cyclic prefix (CP) length or multiple input multiple output (MIMO) mode. In some embodiments, the air interface design can provide a unified or flexible framework to support frequency bands below 6 GHz and frequency bands above 6 GHz (e.g., millimeter wave) for licensed and unlicensed access. For example, the flexibility of a configurable air interface provided by scalable parameter sets (numerology) and symbol durations can enable optimization of transmission parameters for different spectrum bands and different services / devices. As another example, a unified air interface can be self-contained in the frequency domain, and a frequency-domain self-contained design can support more flexible radio access network (RAN) slicing by sharing channel resources between different services in terms of frequency and time.
[0383] Terminal type
[0384] The data processing method provided by the embodiments of the present invention can be applied to various communication scenarios, such as one or more of the following communication scenarios: enhanced mobile broadband (eMBB), ultra-reliable low latency communication (URLLC), machine type communication (MTC), Internet of Things (IoT), narrow band Internet of Things (NB-IoT), customer front-end equipment (CPE), augmented reality (AR), virtual reality (VR), mass machine type communications (mMTC), device to device (D2D), vehicle to everything (V2X), vehicle to vehicle (V2V), etc.
[0385] It should be noted that, in the embodiments of the present invention, the Internet of Things (IoT) may include one or more of NB-IoT, MTC, mMTC, etc. This is not a limitation.
[0386] eMBB can be a high-bandwidth mobile broadband service, such as three-dimensional (3D) or ultra-high-definition video. Specifically, eMBB can also improve the performance of mobile broadband services, such as network speed and user experience. For example, when a user watches 4K HD video, the peak network speed can reach 10 Gbit / s.
[0387] URLLC can refer to services with high reliability, low latency, and extremely high availability. Specifically, URLLC can include the following communication scenarios and applications: industrial applications and control, traffic safety and control, remote manufacturing, remote training, remote surgery, autonomous driving, industrial automation, and the security industry.
[0388] MTC can refer to low-cost and enhanced coverage services, also known as M2M, while mMTC refers to large-scale IoT services.
[0389] NB-IoT can be a range of services characterized by wide coverage, massive connectivity, low data rates, low cost, low power consumption, and a superior architecture. Specifically, NB-IoT can include smart water meters, smart parking, smart pet tracking, smart bicycles, smart smoke detectors, smart toilets, smart vending machines, and more.
[0390] CPE can refer to a mobile signal access device that receives mobile signals and forwards them using wireless fidelity (Wi-Fi) signals, or it can refer to a device that converts high-speed 4G or 5G signals into WiFi signals. It can support a large number of mobile terminals accessing the Internet simultaneously. CPEs can be widely used in rural areas, towns, hospitals, workplaces, factories, and residential areas for wireless network access, thereby reducing the cost of wired network deployment.
[0391] V2X enables communication between vehicles, between vehicles and network devices, and between network devices to obtain a range of traffic information, such as real-time traffic conditions, road information, and pedestrian information, and to provide in-vehicle entertainment information, thereby improving driving safety, reducing congestion, and increasing traffic efficiency.
[0392] For example, terminal types include eMBB devices, URLLC devices, NB-IoT devices, and CPE devices. eMBB devices are primarily used for transmitting large data packets, but can also be used for small data packets, and are typically in a mobile state. Requirements for transmission latency and reliability are generally moderate, and both uplink and downlink communication are present. The channel environment is relatively complex and variable, and indoor or outdoor communication can be used. For example, an eMBB device can be a mobile phone. URLLC devices are primarily used for transmitting small data packets, but can also transmit medium-sized data packets. Generally, URLLC devices are in a stationary state, but can also move along fixed routes. URLLC devices have high requirements for transmission latency and reliability, requiring low latency and high reliability, and both uplink and downlink communication are present. The channel environment is stable. For example, a URLLC device can be factory equipment. NB-IoT devices are primarily used for transmitting small data. NB-IoT devices are typically in a stationary state, with a known location, moderate requirements for transmission latency and reliability, relatively high uplink traffic, and a relatively stable channel environment. For example, an NB-IoT device can be a smart water meter or a sensor. CPE devices are primarily used for transmitting large data packets. They are typically in a stationary state or can move over very short distances. They have moderate requirements for transmission latency and reliability, and handle both uplink and downlink communication in a relatively stable channel environment. For example, CPE devices can be terminal devices in smart homes, AR / VR systems, etc. When determining the terminal type, it can be based on the terminal device's service type, mobility, transmission latency requirements, reliability requirements, channel environment, and communication scenario. The terminal type corresponding to the terminal device can be determined as an eMBB device, URLLC device, NB-IoT device, or CPE device.
[0393] Large-scale artificial intelligence (AI) models with millions or billions of parameters have shown great potential in various AI applications, such as Large Language Model Meta AI (LLaMA), and are considered one of the most promising AI technologies for the future. In the training of large-scale AI models, reinforcement learning from human feedback (RLHF) techniques have proven to play a crucial role in the fine-tuning of large-scale AI models. However, traditional RLHF systems in related technologies may have some drawbacks, which will be combined with... Figure 5 Provide a detailed description.
[0394] Figure 5 This is a schematic framework for a traditional RLHF system. For example... Figure 5As shown, the system includes a reward predictor, an RL algorithm, and an environment. The RL algorithm here can be implemented within an AI model. Generally, the AI model interacts with the environment to produce a set of actions / outputs. During the RL training process of the RL algorithm, the AI model's RL algorithm receives an observation from the environment; in response, the RL algorithm generates and outputs an action to the environment. Then, when the environment is affected by the RL algorithm's action, the environment generates a response observation and sends it to the reward predictor. This response observation may include information about the new state of the environment changed due to the RL algorithm's action. The reward predictor then generates a predicted reward based on the response observation and sends the predicted reward to the RL algorithm so that the RL algorithm updates one or more of its parameters based on the predicted reward. The parameters of the RL algorithm are updated using conventional RL methods (e.g., Q-learning) to maximize the predicted reward generated by the reward predictor.
[0395] Here, the reward predictor can be trained based on human feedback data; that is, one or more parameters of the reward predictor can be optimized / trained through supervised learning to fit the human feedback data. The reward predictor can be trained offline before performing the RL training process described above. Human feedback data can be prepared by selecting (historical) outputs of the RL algorithm, randomly generating multiple pairs of outputs, and sending them to humans for comparison to determine which operation is appropriate. The human comparison results can be used to form the human feedback dataset.
[0396] However, in traditional RLHF systems, reward predictors are trained using generic human feedback data collected from diverse groups with varying preferences or characteristics. Since generic human feedback data only represents common / general human preferences or behaviors—such as identifying targets in a graph or providing reasonable answers to questions without violating common sense—traditional RLHF cannot tailor the RL algorithm's output to the specific needs / characteristics / expertise / preferences (or simply expertise) of the RLHF's client (also known as the RL client). This results in inaccurate RL algorithm outputs. An RL client can be defined as a device / application operated / used by a human. The expertise of an RL client refers to the expertise of the person operating / using that client. For example, an RL client with financial / business expertise and a client with physics research expertise may have different requirements for the RL algorithm's output, even though they might ask the same questions, such as, "What were the main achievements of Thomas Edison in his lifetime?" or "What are the main obstacles to fully utilizing fusion energy?"
[0397] Furthermore, in traditional RLHF systems, because RLHF systems do not focus on distinguishing the source of human feedback data (from people with different expertise), there is no mechanism to obtain prior knowledge of the client's expertise. Additionally, in certain scenarios, collecting client expertise information may be difficult or prohibited for specific reasons (such as protecting client privacy).
[0398] Furthermore, RL training and the execution of the entire system in related technologies are usually carried out in the same network, so there are no systems and methods that support RLHF as a network-native service.
[0399] The objective of this invention is to solve the aforementioned problems through a reinforcement learning from customer feedback (RLCF) system with several functional nodes and related data processing methods. These functional nodes can be distributed across different network elements. Customer feedback data can be collected from a group of other customers with the same expertise as the customer, or from the customer itself. This customer feedback data can be used to help identify customer expertise and train a reward predictor, enabling the predictor to provide customized reward predictions. This allows the RL algorithm to generate outputs that match the customer's expertise. Furthermore, the data processing method and related products proposed in this invention incorporate novel network functions and corresponding signaling / procedures to support RLCF as a network-local service.
[0400] To achieve the above objectives, this invention proposes an RLCF system that can both interact with the environment to execute traditional RLHF processes and customize the final predicted reward based on customer feedback data, thereby enabling customized RL algorithm outputs. The environment considered in the RLCF system can be the same concept defined in a traditional RLHF system, whose state can change based on different RL algorithm outputs. In response to each RL algorithm output, the environment can provide observations to the RLCF system. These observations can include information about the new state of the environment changed due to the operation of the RL algorithm, and (in some embodiments) the output information of the RL algorithm, such as the operation of the RL algorithm or a triple consisting of the previous state of the environment, the operation of the RL algorithm, and the new state of the environment. RL clients can receive / collect observations from the environment to determine customer feedback data sent to the RLCF system. In some embodiments, the RL client can be part (or all) of the environment. In some embodiments, the RL client may not be part of the environment. In the case where the RL client is not part (or all) of the environment, the RL client can obtain the observation from the RL algorithm.
[0401] Before describing the system and related processes proposed in this invention, we will first introduce several functional nodes involved (including the first node, the second node, the third node, the fourth node, the fifth node, and the sixth node).
[0402] Throughout this description, the first node can refer to the expert reward predictor, the second node to the RL algorithm, the third node to the RL client, the fourth node to the expert classifier (EC), the fifth node to the RLCF controller, and the sixth node to the controller of the human feedback dataset (here, the controller can refer to any network element responsible for the human feedback dataset so that other entities can extract data or interact with the human feedback dataset). The names of these functional nodes can be used interchangeably with the elements they refer to.
[0403] It should be noted that some or all of these functional nodes can be integrated into a single network element or distributed across different network elements. When the arrangement of these functional nodes changes, the interactions between them will also change accordingly, which will be explained in detail below with reference to the method flowchart. In the case of integration, these functional nodes can be used as different functional modules within a single network element.
[0404] Specifically, the RL algorithm implemented in the RLCF system can be the same as or different from the RL algorithm in the traditional RLHF system. However, the specific operations (especially the interaction with other functional nodes) will differ from those in related technologies.
[0405] In an RLHF system, there will be one or more ERPs (Reward Advisors), which can generate predicted rewards for RL algorithms. The parameters of one or more ERPs can be optimized / trained (e.g., through supervised learning) to fit human feedback data with expertise information. Therefore, one or more ERPs in an RLCF system can provide customized reward predictions corresponding to the expertise of RL clients, achieving more accurate reward predictions than the general-purpose reward predictors used in traditional RLHF systems. One or more ERPs can be trained offline before executing / training the RL algorithm in the RLCF system.
[0406] Human feedback datasets are used to provide training data for RL (Research-Based Learning) algorithms. They can include multiple human feedback data samples. Each data sample in a human feedback dataset can be associated with a pre-defined area of expertise, corresponding to the expertise of the client who provided the data sample. For example, a data sample provided by a client such as a doctor / computer engineer might be associated with the expertise category of a doctor / computer engineer. For each data sample in a human feedback dataset, its associated area of expertise category information is defined as the expertise information of the data sample. It should be noted that human feedback datasets can also include data samples that are not associated with pre-defined areas of expertise. Because human feedback datasets are used for RL training, the data samples within them may take different forms depending on the specific RLHF (Research-Based Learning) system, particularly the type of one or more ERPs involved in the RLHF system.
[0407] The EC (Expertise Capability) can categorize RL (Reliance, Intelligence, and Professional) clients into one of several predefined expertise categories (e.g., physician, computer engineer) based on client feedback data. The expertise category categorized as an RL client is identified as the RL client's expertise. The EC can send the RL client's expertise information to one or more ERPs (Enterprises for Resource Planning) to generate one or more customized reward predictions. The EC can be implemented as one or more AI / ML-based algorithms, such as clustering algorithms. In some embodiments, the EC should interact with one or more ERPs to determine the client's expertise. In some embodiments, the EC can receive predicted rewards generated by one or more ERPs, then generate a final predicted reward and send it to the RL algorithm.
[0408] The RLCF controller is the logical controller that configures the nodes in the RLCF system to execute each RLCF task. An RLCF task is the process of training an RL algorithm within the RLCF system. RLCF task execution can be requested by an RL client. For example, an RL client raises a question, which is considered initiating an RLCF task. Then, after the RLCF system executes the task, it provides an answer to the RL client; this process might be considered executing an RLCF task. To complete an RLCF task, the RLCF system needs to perform operations (e.g., AI inference, AI training, and data processing). During RLCF task execution, the RLCF controller can configure other network functions within the RLCF system. Configuring these network functions can instruct them to perform corresponding operations to complete the RLCF task.
[0409] RL customers can receive observations from the environment, generate customer feedback data based on these observations, and send the customer feedback data to EC.
[0410] The following is combined Figures 6 to 8 The basic concepts of this invention are described. Wherein, Figure 6 For a general RLCF system, and Figure 7 and Figure 8 This describes the specific implementation of this general RLCF system.
[0411] like Figure 6 As shown, the RLCF system includes functional nodes, such as one or more ERPs, RL algorithms, RL customers, ECs, and RLCF controllers. The figure also shows a human feedback dataset, which can be implemented as a database.
[0412] Generally, an RL algorithm can obtain an observation from the environment, and in response, it generates and outputs an operation to the environment (in...). Figure 6 (Shown as decision / output). Then, when the environment is acted upon by the RL algorithm, the environment provides the observation to the RL client, which generates client feedback data based on the observation and sends it to the EC. The EC then interacts with the ERP to generate the final predicted reward for the RL algorithm, which can be sent to the RL algorithm by the EC or ERP, and the parameters of the RL algorithm are then updated.
[0413] From an ERP perspective, there are two types of RLCF systems: Figure 6 Different implementations of the system are shown. One is called an RLCF system α (such as...). Figure 7A As shown), there are multiple ERPs that are trained using data samples unrelated to pre-determined expertise; another type is called RLCF system β (as shown). Figure 7B As shown, one of these systems is an integrated ERP, which is trained using data samples associated with pre-defined expertise. Both systems can be used to complete RLCF tasks, but the functional nodes involved in the RLCF systems can perform different operations, which will be described below in conjunction with the method flowchart.
[0414] like Figure 7AAs shown, the RLCF system α includes multiple ERPs, and the parameters of each ERP can be optimized / trained (e.g., through supervised learning) to fit a subset (or group) of the human feedback dataset. Samples in the human feedback data within this subset can be categorized into specific expertise (e.g., doctor, computer engineer), and these expertise are identified as the associated expertise of the ERP. An ERP whose associated expertise is the same as the expertise of the RL client is defined as the corresponding ERP, which provides the most accurate reward prediction compared to other ERPs with different associated expertise. That is, the human feedback dataset can include multiple sets of training data (training samples), each set corresponding to a predefined expertise, with different sets of training data corresponding to different predefined expertise. Different subsets of the human feedback dataset can be stored on different network elements or on the same network element; this invention does not limit this.
[0415] In the RLCF system α, EC can be used to determine customer expertise based on (1) customer feedback data or (2) customer feedback data and predicted rewards from multiple ERPs. In some embodiments, EC can be used to generate a final predicted reward based on customer feedback data and predicted rewards from multiple ERPs, and then send the final predicted reward to the RL algorithm.
[0416] like Figure 7B As shown, the RLCF system α includes an integrated ERP (called I-ERP), whose parameters can be optimized / trained (e.g., through supervised learning) to fit a general human feedback dataset. Unlike traditional RLHF systems and RLCF system α, each training data point used to train the I-ERP is appended with an additional fragment indicating its associated expertise, i.e., an additional expertise fragment. Therefore, the input data for the I-ERP consists of observation fragments and additional expertise fragments. In contrast, the input data for the ERP in RLCF system α only has observation fragments. With the help of additional expertise fragments, even given the same observations, the I-ERP can adjust the predicted reward based on different customer expertise.
[0417] In the RLCF system β, EC can be used to generate data consisting of observation segments and additional expertise segments based on customer feedback data, and send this data to I-ERP as input data for I-ERP.
[0418] Using the aforementioned functional nodes, the ERP is first trained based on human feedback data provided by the human feedback dataset. Then, the RLCF controller can configure the functional nodes involved in executing RLCF tasks, and each functional node will then function according to its configuration.
[0419] like Figure 8As shown, the proposed RLCF system workflow consists of three steps: ERP training, RLCF configuration, and RLCF execution. Among them:
[0420] (a) ERP training: The RLCF controller trains one or more ERPs (i.e., optimizes / trains the parameters of each ERP by a predetermined method (e.g., supervised learning)) to fit the training data provided by the human feedback dataset, wherein one or more trained ERPs in this step can be reused by different RLCF tasks;
[0421] (b) RLCF Configuration: Given a request from an RL client, the RLCF controller configures one or more trained ERP, EC, and RL algorithms to execute the RLCF task requested by the RL client;
[0422] (c) RLCF Execution: The RLCF system and RL clients execute RLCF tasks according to the execution logic configured by the RLCF controller. This invention provides a data processing method for executing the three steps of the proposed RLCF system workflow, and also provides a corresponding data processing system.
[0423] The data processing method and related products provided by the present invention will be described in detail below with reference to the accompanying drawings.
[0424] Figure 9A This is a flowchart of a first data processing method according to one or more embodiments of the present invention. Figure 9B It corresponds to Figure 9A One possible implementation of the process shown differs in that the subjects are represented as the human feedback dataset (i.e., the controller of the human feedback dataset), the RLCF controller, and the ERP, and the information exchanged between the different entities is described using different names. However, the principles shown in both diagrams are similar. This method can be implemented by one or more ERPs (third nodes), the RLCF controller (fifth node), and the controller of the human feedback dataset (sixth node) in the RLCF system. Figure 9A As shown, the method may include the following steps.
[0425] S901: The fifth node sends a first request to the sixth node, and the sixth node receives the first request from the fifth node. The first request includes first information and second information. The first information indicates feature information used to train the first node, and the second information indicates data attributes associated with the feature information.
[0426] Specifically, the fifth node can send a first request to the sixth node to initiate training for the first node. In one possible implementation of the invention, the first request may be... Figure 9B The example shown is a request for human feedback training data.
[0427] In one possible implementation of the invention, the first information includes at least one indication of at least one feature information. The feature information is used to characterize the features of the first node and may be a professional knowledge category required by the RLCF controller to train the ERP. Accordingly, at least one indication of the at least one feature information may be one or more identities (IDs) / one or more names of one or more professional knowledge categories required by the fifth node to train the first node; or, if the fifth node does not have prior knowledge of one or more IDs / one or more names of potential one or more professional knowledge categories, it may be a description of the required one or more professional knowledge categories. When the fifth node knows one or more IDs / one or more names of one or more professional knowledge categories required by the fifth node to train the first node, the one or more IDs / one or more names will be used as an indication of the feature information and may be included in the first request; when the fifth node does not know the one or more IDs / one or more names but still needs to train the first node, the description of the required one or more professional knowledge categories may be used as an indication of the feature information and may be included in the first request. The following explanation uses a professional knowledge category as an example of feature information, but it should be understood that this approach is equally applicable when the feature information is other types of information.
[0428] In this implementation, the fifth node sends at least one indication of at least one feature information and the minimum amount of data associated with each of the at least one feature information to the sixth node, which helps the training of the first node and thus improves the accuracy of the first node in determining the reward prediction result.
[0429] In one possible implementation of the invention, the second information includes a minimum amount of data associated with each of the at least one feature information. To complete the training of the first node, in addition to indicating one or more required expertise categories to the sixth node, the fifth node may also notify the amount of data used for training, one way being to carry a minimum amount of data in the first request so that the sixth node can verify whether it has sufficient data to support such training.
[0430] S902: The sixth node determines the first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node.
[0431] Upon receiving the first request, the sixth node can analyze the request to determine whether it can support the training required by the fifth node.
[0432] In one possible implementation of the invention, since the first request includes the aforementioned first and second information, the sixth node can then determine one or more required expertise categories based on the first information. The sixth node can then determine whether it has training data associated with the required one or more expertise categories. Furthermore, when determining that the sixth node has training data associated with the required one or more expertise categories, the sixth node can also determine whether the minimum amount of data in the first request is sufficient to support training. In response, the sixth node generates a first response based on the above determination. Therefore, supporting a required expertise category means that the sixth node not only has training data associated with that expertise category, but also that the amount of such training data is sufficient to complete the training.
[0433] In one possible implementation of the invention, the first response includes an indication of feature information supported by the sixth node, and an indication of the sixth node including data required to train the first node; accordingly, step S902 includes:
[0434] The sixth node determines the indication of the feature information supported by the sixth node based on the first information;
[0435] The sixth node determines, based on the second information, an indication of the data required to train the first node.
[0436] Since the first information indicates one or more professional knowledge categories requested by the fifth node, the sixth node can determine the indication of the feature information supported by the sixth node based on the first information. Let's continue with the example where the indication of the feature information supported by the sixth node is information about one or more professional knowledge categories supported by the sixth node (e.g., one or more IDs of one or more professional knowledge categories supported by the sixth node). Since the first information in the first request may, in some cases, only include the description of the required one or more professional knowledge categories, the sixth node can send one or more IDs corresponding to the description of the required one or more professional knowledge categories (i.e., the professional knowledge categories supported by the sixth node) in the first response. In this case, the indication of the feature information supported by the sixth node will be one or more IDs corresponding to the description of the required one or more professional knowledge categories.
[0437] The indication of the sixth node including the data required to train the first node could be, for example, an indication that the amount of training data for a particular category of expertise is sufficient to train the first node. The sixth node can determine the indication of the sixth node including the data required to train the first node based on second information, since the second information includes the minimum amount of data associated with each of the at least one feature.
[0438] It should be noted here that the sixth node may or may not support all the professional knowledge categories requested by the fifth node. If the sixth node supports all the professional knowledge categories requested by the fifth node, the sixth node may send an acknowledgment to the fifth node only in the first response. This acknowledgment can be considered both an indication of the feature information supported by the sixth node and an indication of the sixth node including the data required to train the first node; alternatively, the sixth node may send the amount of data associated with each required professional knowledge category to the fifth node in the first response; or, the sixth node may send one or more IDs of one or more required professional knowledge categories in the first response, which is not limited in this invention. If the sixth node only supports some of the professional knowledge categories requested by the fifth node, the sixth node may also send one or more IDs of the supported one or more professional knowledge categories and the amount of data associated with the supported one or more professional knowledge categories as indications of the feature information supported by the sixth node and indications of the sixth node including the data required to train the first node, respectively, to the fifth node in the first response.
[0439] For example, if the fifth node sends four IDs—ID1, ID2, ID3, and ID4—in the first request, each ID indicating a professional knowledge category and having its corresponding minimum amount of data for training (represented as data amount 1, data amount 2, data amount 3, and data amount 4), then the sixth node, based on the first request, determines that it has the data required for the professional knowledge categories corresponding to ID3 and ID4. The sixth node can then use ID3 and ID4 as indicators of the feature information it supports; and / or use data amount 3 and data amount 4 as indicators of the data required by the sixth node for training the first node, and then send one or more of these indicators to the fifth node. It should be noted that the first response is used to clarify the sixth node's ability to support the training of the first node, and therefore can take various forms, which are not limited in this invention.
[0440] exist Figure 9B In this context, the first response represents the supported professional knowledge category information, and step S902 corresponds to... Figure 9B Step 1.2 is shown in the diagram. Figure 9B In this dataset, the controller of the human feedback dataset can analyze human feedback and determine whether one or more expertise categories required in the human feedback training data request can be supported by the human feedback dataset. Specifically, for each required expertise category, the human feedback dataset contains sufficient (i.e., greater than or equal to the minimum amount of data indicated in the human feedback training data request or a predetermined threshold) training data associated with the required expertise category. If the human feedback training data request includes a description of the required expertise, the controller of the human feedback dataset can parse these descriptions and determine one or more expertise categories that match the description of the required expertise.
[0441] S903: The sixth node sends a first response to the fifth node, and the fifth node receives the first response from the sixth node.
[0442] After determining the first response, the sixth node can send the first response to the fifth node.
[0443] This step corresponds to Figure 9B In step 1.3, the controller of the human feedback dataset sends the supported professional knowledge category information to the RLCF controller. The supported professional knowledge category information may include: ID (an indication corresponding to the feature information supported by the sixth node mentioned above) and the amount of training data that the human feedback dataset can support for each required professional knowledge category (an indication corresponding to the data required to train the first node mentioned above).
[0444] S904: The fifth node configures the first and sixth nodes to perform training based on the first response.
[0445] Specifically, the fifth node configures the first node and the sixth node based on the first response (i.e., based on information about one or more professional knowledge categories supported by the sixth node and an indication that the amount of training data for each required professional knowledge category that the sixth node can support is sufficient to train the first node included in the first response to perform training).
[0446] In one possible implementation of the present invention, step S904 includes:
[0447] The fifth node determines the first configuration information and the second configuration information of the first node based on the first response, wherein the first configuration information indicates the type of the first node and the second configuration information indicates the parameters used to train the first node;
[0448] The fifth node configures the first node based on the first and second configuration information;
[0449] The fifth node notifies the sixth node of the configuration of the first node.
[0450] Specifically, the fifth node determines first configuration information indicating the type of the first node and second configuration information indicating the parameters used to train the first node based on the first response, configures the first node based on the first and second configuration information, and then notifies the sixth node of the configuration of the first node.
[0451] In this implementation, the fifth node uses the type of the first node and the parameters used to train the first node to configure the first node, thereby enabling the trained first node to determine a more accurate reward prediction result.
[0452] In one possible implementation of the present invention, determining the first configuration information of the first node may include:
[0453] The fifth node determines the type of the first node based on the first response;
[0454] The configuration for notifying the first node from the sixth node includes:
[0455] The fifth node informs the sixth node of the type of the first node.
[0456] Specifically, the first configuration information indicates the type of the first node. Correspondingly, the fifth node determines the first configuration information of the first node by determining the type of the first node and notifying the sixth node of the first node's type. In one possible implementation of the invention, the type of the first node can be predefined, therefore the fifth node may not need to notify the sixth node of the first node's type.
[0457] The type of the first node can include a first type and a second type. The first type of node includes an integrated node associated with multiple feature information, while the second type of node includes multiple child nodes, each of which is associated with one feature information among the multiple feature information. For example, the first type of node could be the I-ERP in the aforementioned RLCF system β, while the second type of node has multiple child nodes (…). Figure 7A The ERPs (ERP A, ERP B, ..., ERP N) shown are multiple ERPs in the aforementioned RLCF system α. It should be noted that the specific implementation of the first node is not limited, as long as its functionality is achieved. For example, when the first node is of type 1, it can be implemented as a distributed functional module (e.g., distributed across different network elements); when the first node is of type 2, it can be implemented as different functional modules integrated within a single network element.
[0458] In this implementation, the type of the first node is determined based on the first response and notified to the sixth node, which is beneficial for configuring the first node and thus helps improve the accuracy of the first node in determining the reward prediction result.
[0459] In one possible implementation of the present invention, the first response sent by the sixth node includes an indication of feature information supported by the sixth node and an indication of the sixth node including data required for training the first node; determining the first configuration information of the first node may further include:
[0460] The fifth node determines the number of candidate nodes to implement the first node based on the feature information supported by the sixth node.
[0461] The fifth node determines the second configuration information of the first node, including:
[0462] The fifth node determines the second configuration information based on the sixth node's indication of the data required to train the first node.
[0463] Specifically, the information in the first response regarding one or more professional knowledge categories supported by the sixth node, and the indication that the amount of training data for each required professional knowledge category supported by the sixth node is sufficient to train the first node, are used by the fifth node to determine the number of candidate nodes for implementing the first node and the parameters for training the first node, respectively. As described above, the sixth node can return to the fifth node the indication of the supported professional knowledge categories in the first response and the sixth node's indication of the data required to train the first node. Therefore, taking the example above, when the fifth node requests ID1 to ID4, but the sixth node only supports ID3 and ID4, after receiving the indication of the supported feature information, the fifth node will know that the sixth node only supports two IDs. Therefore, for example, in the case of RLCF system α, the fifth node uses only two candidate nodes as child nodes for implementing the first node, instead of using four candidate nodes as child nodes for implementing the first node. Furthermore, since the amount of data can be a factor related to hyperparameters and initial parameters, the fifth node can determine the second configuration information based on the sixth node's indication of the data required to train the first node.
[0464] In this implementation, the fifth node configures the first node using the indications of the feature information supported by the sixth node and the indications of the sixth node including the data required to train the first node. Therefore, the configured first node matches well with the sixth node, thereby enabling the trained first node to determine a more accurate reward prediction result.
[0465] In one possible implementation of the invention, the second configuration information includes parameters of a first type and parameters of a second type, wherein the parameters of the first type remain unchanged during the training of the first node, and the parameters of the second type are updated based on training data provided by the sixth node during the training of the first node. For example, the parameters of the first type (also called hyperparameters) include at least one of the structure of the neural network (e.g., a deep neural network (DNN)) used to train the first node, or at least one of the methods used to train the first node (e.g., a detailed RL algorithm of an RLCF system). The parameters of the second type (also called initial parameters) include at least one of the initial weights and biases of the neural network. Furthermore, the parameters of the second type may also include links between one or more ERPs to be trained.
[0466] In this implementation, the parameters used to configure the first node are divided into different types: parameters that remain unchanged during the training of the first node and parameters that are updated during the training of the first node. Therefore, the configuration can be implemented in a more accurate way.
[0467] In one possible implementation of the present invention, configuring the first node based on the first configuration information and the second configuration information includes:
[0468] The fifth node configures network elements for implementing the first node and the connection between the sixth node and the first node based on the first and second configuration information.
[0469] Specifically, since the first configuration information includes the number of candidate nodes to implement the first node and the type of the first node, and the second configuration information includes determined hyperparameters and initial parameters, the fifth node configures the network elements to implement the first node and the connection between the sixth node and the first node based on the type of the first node, the number of candidate nodes to implement the first node, and the determined hyperparameters and initial parameters. Here, configuring the connection between the sixth node and the first node is to enable the sixth node to provide training data to the first node, thereby enabling the first node to perform training. If the first node is of the first type, this configuration may be configuring the connection between the sixth node and the first node (integration node); if the first node is of the second type, this configuration may be configuring the connection between the sixth node and all child nodes of the first node.
[0470] In this implementation, the first configuration information and the second configuration information are used to configure the network elements used to implement the first node and the connection between the sixth node and the first node. This is beneficial for configuring the first node, thereby improving the accuracy of the first node in determining the reward prediction result.
[0471] After configuring the network elements for the first node and the connection between the sixth node and the first node, the sixth node and the first node can exchange training data to perform training. There are several ways for the sixth node to send training data to the first node. One way is to send training data in response to a training request from the first node; another way is to send training data directly without triggering a request. These two methods are described in detail below.
[0472] Method 1: Training request is required
[0473] In this approach, the notification configuration for the first node includes:
[0474] The fifth node sends a third message to the sixth node, and the sixth node receives the third message, wherein the third message indicates at least one path between the first node and the sixth node.
[0475] Specifically, the fifth node can send third information (also known as routing information) indicating at least one path between the first node and the sixth node to the sixth node, thereby enabling data transmission between the first node and the sixth node based on at least one path indicated by the third information.
[0476] In this implementation, the fifth node sends information to the sixth node indicating at least one path between the first node and the sixth node. This facilitates sending training data to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0477] The method may also include:
[0478] The first node sends a training request to the sixth node, wherein the training request indicates the feature information of the first node;
[0479] The sixth node receives the training request from the first node;
[0480] The sixth node determines the training data for the first node based on the type and training request of the first node;
[0481] The sixth node sends training data to the first node based on the third information, so that the first node can perform training based on the training data.
[0482] Specifically, the first node sends a training request to the sixth node. After receiving the training request from the first node, the sixth node determines the training data of the first node based on the type of the first node (which may be predefined or notified by the fifth node) and the feature information of the first node indicated by the training request. Then, it sends the training data to the first node based on the third information so that the first node can perform training based on the training data.
[0483] In this implementation, the training data of the first node corresponding to the feature information of the first node is determined by the sixth node based on the type of the first node and the training request received from the first node, and sent to the first node based on the third information, which is beneficial to the training of the first node and thus helps to improve the accuracy of the first node in determining the reward prediction result.
[0484] Method 2: No training request
[0485] The first node can also be trained without a training request.
[0486] Specifically, in this approach, the notification configuration for the first node includes:
[0487] The fifth node sends a fourth message to the sixth node, and the sixth node receives the fourth message, which indicates the training dataset of the first node.
[0488] For a description of the third message, please refer to the first method. In method 2, in addition to sending the third message, the fifth node can also send a fourth message to the sixth node. Figure 9B (Training data requirements).
[0489] Specifically, for the first type of node, since the first node is an integrated node associated with multiple professional knowledge categories supported by the sixth node, the fourth information may include one or more IDs of multiple professional knowledge categories, enabling the sixth node to locate the training datasets corresponding to the multiple professional knowledge categories based on the fourth information and send the training data in these training datasets to the first node. For the second type of node, since the first node includes multiple child nodes, and each of the multiple child nodes is associated with a professional knowledge category, the fourth information may also include the IDs of the corresponding professional knowledge categories associated with the multiple child nodes. That is, the sixth node knows the relationship between the child nodes and the professional knowledge categories supported by the sixth node, so the sixth node can determine which child node corresponds to which professional knowledge category, and then send the training data in the training dataset corresponding to the professional knowledge category to the corresponding child node.
[0490] In this implementation, the fifth node sends information to the sixth node indicating the training dataset for the first node. This helps determine the training data to be provided to the first node for training, thereby improving the accuracy of the first node in determining the reward prediction result.
[0491] In one possible implementation of the invention, the fourth information may also indicate the format of the data used to train the first node, so that the sixth node prepares training data for the first node to ensure that the training data meets the format requirements.
[0492] The format here refers to the format of the training data that the sixth node wants to send to the first node. The format can be communicated to the sixth node to prepare the corresponding training data. Furthermore, this format indication can also be carried in other information exchanged between the fifth and sixth nodes; this is not limited here. The format can also be fixed; in this case, it is not necessary to indicate the format to the sixth node.
[0493] In this implementation, the format of the data used to train the first node is sent from the fifth node to the sixth node. This facilitates sending training data to the first node for training purposes, thereby improving the accuracy of the first node in determining the reward prediction result.
[0494] S905: The sixth node sends training data to the first node, and the first node receives training data from the sixth node, wherein the training data corresponds to the feature information of the first node.
[0495] Depending on the type of the first node, the sixth node can send training data in different ways.
[0496] In the case where the first node is of the first type and includes an integrated node associated with multiple professional knowledge categories, the sixth node determines the first training dataset based on the fourth information and sends the training data of the first training dataset to the first node based on the third information, so that the first node performs training based on the training data of the first training dataset.
[0497] In the case where the first node is of the second type and includes multiple child nodes, each of the multiple child nodes is associated with a professional knowledge category. The sixth node determines the second training dataset based on the third information and sends the training data of the second training dataset to each of the multiple child nodes based on the fourth information, so that the first node performs training based on the training data of the second training dataset.
[0498] The first and second training datasets here can both be included in the human feedback dataset controlled by the sixth node. The first training dataset includes multiple training data sets, each consisting of a first data portion and a second data portion, where the second data portion indicates the expertise category of the first data portion; the second training dataset includes multiple sets of training data sets, each set of training data sets being associated with a predefined expertise category among multiple expertise categories.
[0499] As referenced above Figure 7B In the case of I-ERP, each piece of training data in the first training dataset includes an observation segment and an additional professional knowledge segment, which indicates the feature information of the observation segment. The observation segment here corresponds to the first data segment described above, and the additional professional knowledge segment corresponds to the second data segment described above.
[0500] Therefore, the training dataset is determined by the sixth node based on the type of the first node and the third information, and sent to the first node based on the fourth information, so that the first node can perform training based on the training data in the training dataset. In other words, the sixth node prepares different types of training data for the first node based on its type. After the training data is determined, the sixth node and the first node can exchange training data through at least one path indicated by the third information. Thus, this implementation method helps improve the accuracy of the first node in determining the reward prediction result.
[0501] S906: The first node performs training based on the training data from the sixth node.
[0502] The steps S901 to S905 described above are used to prepare for training the first node. In step S906, the first node performs training based on the training data from the sixth node.
[0503] Figure 9B One possible implementation of the above process is illustrated. Specifically, a human feedback training data request is sent from the RLCF controller to the controller of the human feedback dataset. In response, the controller of the human feedback dataset analyzes the human feedback training data request and returns supported professional knowledge category information. The supported professional knowledge category information may include one or more IDs of professional knowledge categories supported by the controller of the human feedback dataset, and the amount of training data for each required professional knowledge category that the human feedback dataset can support. After receiving the supported professional knowledge category information, the RLCF controller (e.g., the NET4AI controller) determines the number of candidate nodes for implementing the ERP, the type of ERP (ERP or I-ERP in the RLCF system α described above), hyperparameters (e.g., DNN structure, training method, detailed algorithm), and initial parameters of the one or more ERPs to be trained (e.g., initial weights and biases on neurons and links). For multiple ERPs in the RLCF system α, the associated professional knowledge category of each ERP (which needs to be one of the required professional knowledge categories) should also be determined by the RLCF controller based on the supported professional knowledge category information. Then, the RLCF controller configures one or more network elements to implement each ERP using the determined hyperparameters and initial parameters, and configures one or more connections or one or more APIs between the human feedback dataset and the one or more network elements used to implement each ERP. Then, the RLCF controller notifies one or more network elements used to implement each ERP of the human feedback dataset regarding routing information, as well as the training data requirements for each ERP. The training data requirements for each ERP may include: (1) the ID of the associated professional knowledge category of the ERP (for I-ERP, this information is one or more IDs of all one or more professional knowledge categories that can be used to train the I-ERP); (2) the format of the training data samples.
[0504] After the ERP training configuration is completed, each ERP can be trained independently according to the following steps: (1) The controller of the human feedback dataset prepares training data for each ERP to ensure that the training data request is received in step 1.5, and sends the prepared training data to the target ERP according to the routing information received in step 1.5; (2) Given that the prepared training data has been received in step 2.1, each ERP can independently perform its own training process.
[0505] Figure 9B Step 1.1 corresponds to Figure 9A Step S901, Figure 9B Step 1.2 corresponds to Figure 9A Step S902, Figure 9B Step 1.3 corresponds to Figure 9AStep S903, Figure 9B Step 1.4 corresponds to Figure 9A Step S904, Figure 9B The training data feed path information shown in step 1.5 corresponds to Figure 9A In step S904, the third message is sent. Figure 9B Step 2.1 corresponds to Figure 9A Step S905, Figure 9B Step 2.2 corresponds to Figure 9A Step S906.
[0506] It's important to note that the training of the first node can be completed offline. This means that the trained first node can be reused by different RLCF tasks before the RLCF system executes. Furthermore, the first node can also update itself after the RLCF system executes. For example, when executing an RLCF task initiated by an RL client, the RL client can provide feedback data, which can be added to the human feedback dataset controlled by the sixth node, thereby updating the first node's training dataset. The first node can then be trained again to improve its performance.
[0507] In this data processing method, the first node and the sixth node are configured based on a first response indicating the ability of the sixth node to support the training of the first node, and the training of the first node is performed based on training data corresponding to the feature information of the first node. The first node trained with training data associated with different professional knowledge categories can provide more accurate reward predictions, that is, improve the accuracy of the first node in determining the reward prediction results.
[0508] In the aforementioned ERP training process, through the interaction between the fifth and sixth nodes, the fifth node can be configured based on the capabilities of the sixth node, thereby preparing both the first and sixth nodes for training. This, in turn, enables the trained first node to determine more accurate reward prediction results. Furthermore, the training data for the first node, corresponding to its feature information, is received from the first node and used to perform training, which helps improve the accuracy of the first node's reward prediction results.
[0509] As mentioned above, after training the first node, the RLCF controller (the fifth node) can configure the corresponding network elements to prepare for the subsequent execution of the entire RLCF system. The following section combines... Figure 10A and Figure 10B Describe the configuration process.
[0510] Figure 10A This is a flowchart of a second data processing method according to one or more embodiments of the present invention. Figure 10B It corresponds to Figure 10AThe diagram illustrates one possible implementation of the process, where the main entities are represented as RL clients, RL algorithms, ECs, ERPs, and RLCF controllers, with different names used to describe the information exchanged between the different entities. However, the principles shown in both diagrams are similar. This method can be implemented by a third node (RL client), a second node (RL algorithm), a fourth node (EC), a first node (one or more ERPs), and a fifth node (RLCF controller) of the RLCF system. Figure 10A As shown, the method may include the following steps.
[0511] S1001: The third node sends a fourth indication to the fifth node, wherein the fourth indication indicates the type of feedback data provided by the third node regarding the first observation.
[0512] Specifically, the fourth indication sent by the third node to the fifth node indicates the type of feedback data provided by the third node to the fifth node regarding the first observation. Here, the first observation is the state of the environment obtained in response to the operation indication of the second node. During the execution of the RLCF system, the second node can output an operation to the environment, and the state of the environment can then change in response to that operation. Therefore, the third node provides the first / fourth node with an observation indicating the state of the environment (the first observation). Simultaneously, the third node also provides the fourth node with feedback data regarding the first observation, enabling the first node to determine a first prediction (reward prediction) for the operation of the second node. Thus, the feedback data provided by the third node is used to determine the final reward prediction for the operation of the second node, allowing the third node to inform the fifth node of the type of feedback data during configuration so that the fifth node can configure the relevant nodes of that type.
[0513] In one possible implementation of the invention, the feedback data includes a first feedback type and a second feedback type. The first feedback type indicates the third node's evaluation of the observation, and the second feedback type indicates the third node's characteristic information. Specifically, the first feedback type can be reward-based feedback, where the value of the feedback data can be interpreted as a reward determined by the customer based on the observation. The second feedback type can be non-reward-based feedback, where the value of the feedback data may not be interpreted as a reward, but may represent one or more categories of expertise of the third node.
[0514] In one possible implementation of the invention, for example, before sending the fourth instruction, the third node may send a third request to the fifth node, wherein the third request indicates descriptive information corresponding to the target task. Specifically, the third node sends a third request to the fifth node indicating descriptive information corresponding to the target task to notify the fifth node of the descriptive information corresponding to the target task, wherein the descriptive information corresponding to the target task may include a problem description or specification of the target task. Accordingly, the fifth node receives the third request sent by the third node. At this time, sending the third request can be regarded as initiating the target task. In one possible implementation, the third request may also indicate the performance requirements for performing the target task. For example, the performance requirements may include a convergence time threshold, a minimum prediction accuracy requirement, etc.
[0515] In this implementation, sending a third request from the third node to the fifth node, which indicates the description information corresponding to the target task, is beneficial to the execution of the target task.
[0516] S1002: The fifth node obtains the fourth instruction from the third node.
[0517] In one possible implementation of the invention, the third node can autonomously send the type of feedback data without being triggered by a second request from the fifth node, thus allowing the fifth node to directly obtain the fourth instruction.
[0518] In one possible implementation of the invention, the fifth node sends a second request to the third node to request the type of feedback data, in order to query the type of feedback data of the first observation provided by the third node, thereby configuring the system to execute the target task subsequently initiated by the third node, and thus enabling the system to be well configured to execute the target task.
[0519] In one possible implementation of the invention, the second request may also be used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node; the fourth instruction may also indicate at least one of the items requested by the fifth node.
[0520] In this implementation, the fifth node sends a second request to the third node requesting at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, so as to configure the system for executing the target task initiated by the third node based on these items in a more flexible manner, thereby enabling the system to be well configured for executing the target task.
[0521] S1003: The fifth node is configured based on the fourth instruction to execute the target task initiated by the third node.
[0522] Specifically, after receiving the type of feedback data from the third node, the fifth node can be configured based on the type of feedback data provided by the third node regarding the first observation to execute the target task initiated by the third node. For example, the configuration here may include the selection of functional nodes, the operation of corresponding functional nodes, and the connections between related functional nodes. All the content required for subsequent execution of the RLCF system can be configured in this step.
[0523] In one possible implementation of the invention, when the third node only provides the type of feedback data to the fifth node, the fifth node can determine the relevant functional nodes, their operations, and the connections between them, without considering specific requirements related to the target task. For example, historical data can be used to determine which nodes have participated in processing the same type of feedback data.
[0524] In one possible implementation of the invention, when the third node sends a third request indicating descriptive information corresponding to the target task in addition to the fourth instruction, the fifth node can then be configured based on the type of feedback data and the descriptive information indicated by the third node.
[0525] For example, this configuration can be performed as follows:
[0526] The fifth node determines the first and second nodes based on the description information corresponding to the target task in order to execute the target task;
[0527] The fifth node, based on the type of feedback data and the first node, determines the operations to be performed by the first, second, and fourth nodes to execute the target task, wherein the fourth node is used to provide predictions for the observations provided by the third node;
[0528] The fifth node configures network elements and connections to implement the operations performed by the first, second, and fourth nodes, based on the operations performed by the first, second, and fourth nodes.
[0529] The fifth node notifies the third node of the configuration to execute the target task based on the network elements and connections used to implement the operations performed by the first, second, and fourth nodes.
[0530] Specifically, the fifth node, based on the problem description or specification of the target task in the description information, determines the first and second nodes to execute the target task. For example, the type of the first node (I-ERP or multiple ERPs) and the specific algorithm implemented by the second node. Then, based on the type of feedback data and the first node, the fifth node determines the RLCF execution flow for executing the target task. Specifically, the operations performed by the first, second, and fourth nodes execute the target task. The operations performed by the first, second, and fourth nodes are then used by the fifth node to configure the network elements and connections for implementing the operations performed by the first, second, and fourth nodes, and based on this, notify the third node of the configuration to execute the target task. The node operations here include not only the execution logic associated with the node but also the specific algorithm implemented by the node and the relevant parameters used in the algorithm. For example, configuring the second node's operations may also include configuring the hyperparameters of the algorithm implemented by the second node, because the hyperparameters are related to the description information; configuring the fourth node's operations may also include setting a preset classification algorithm in the fourth node. Furthermore, the configuration operations here may also include configuring the scheduling start time of the applied RLCF execution process.
[0531] Furthermore, the RLCF execution process can be implemented in multiple ways (described in detail below), so the fifth node can be configured with relevant nodes based on a specific process. Additionally, the fourth node itself can provide predictions for observations provided by the third node, and can also help provide such predictions; therefore, the operations performed by the fourth node and the specific algorithms used by the fourth node (e.g., the classification algorithm used by the fourth node and the parameters of the classification algorithm) can also be configured in this step.
[0532] In this implementation, by determining the first and second nodes to execute the target task, determining the operations to be performed by the first, second, and fourth nodes, configuring the network elements and connections to implement the operations performed by the first, second, and fourth nodes, and notifying the third node of the configuration to execute the target task, the system for executing the target task can be well configured.
[0533] In one possible implementation of the present invention, configuring network elements and connections for implementing operations performed by the first node, the second node, and the fourth node includes:
[0534] Configure network elements to implement operations performed by the first node, second node, and fourth node;
[0535] Configure the connection between the first node and the network element that implements the second node, the connection between the first node and the network element that implements the fourth node, and the connection between the first node and the third node;
[0536] Configure the connection between the fourth node and the network element that implements the first node, the connection between the fourth node and the network element that implements the second node, and the connection between the fourth node and the third node;
[0537] Configure the connection between the second node and the network element that implements the first node, the connection between the second node and the network element that implements the fourth node, and the connection between the second node and the third node.
[0538] The network element configuration determines the network elements used to implement the first, second, and fourth nodes, while the connection configuration prepares the network elements for operation. It should be noted that the connection configurations shown above cover all possible connections between the first, second, third, and fourth nodes. However, some connections can be configured based on the interactions between these nodes; this is not a limitation here. Furthermore, although these nodes are described using different serial numbers, some of these nodes can be integrated into a single node; in this case, a single network element is sufficient to implement these integrated nodes.
[0539] In one possible implementation of the present invention, determining the first node to perform the target task includes:
[0540] The fifth node determines at least one candidate node that has performed a historical task as the first node based on the description information corresponding to the target task, wherein the similarity between the historical task and the target task is above a preset threshold.
[0541] Specifically, historical data from past tasks is used to select the first node. Since some candidate nodes may have already participated in similar tasks, these candidate nodes that have performed historical tasks and whose similarity to the target task is above a preset threshold can be determined as the first node by the fifth node. The preset threshold here represents the similarity between different tasks and can be set according to actual circumstances; it is not limited here.
[0542] In this implementation, since the first node is determined to be at least one candidate node that has executed a historical task, and the similarity between the historical task and the target task is above a preset threshold, reusing the first node makes configuration easier.
[0543] In one possible implementation of the present invention, determining the first node to perform the target task includes:
[0544] The fifth node determines the type of the first node based on the description information corresponding to the target task. The type of the first node includes a first type and a second type. The first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of which is associated with one feature information among the multiple feature information.
[0545] When the first node is determined to be of the second type, the fifth node determines all available candidate nodes as child nodes of the first node.
[0546] Specifically, the description of the type of the first node can be found in the above embodiment. Here, the determination of the type can consider multiple factors, such as the potential processing capacity required for the description information, the availability of candidate nodes, etc. When selecting I-ERP, all available candidate nodes can be selected as child nodes of the first node.
[0547] In this implementation, by determining the type of the first node based on the description information corresponding to the target task, the system for executing the target task can be well configured.
[0548] After configuring the connection and network elements, the fifth node can notify the third node of the configuration to execute the target task, so that the third node can interact with the relevant nodes during the subsequent RLCF execution process.
[0549] In one possible implementation of the invention, the fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates the network elements for implementing the operations performed by the first node and the connection between the third node and the first node, thereby the third node receives the third configuration notification from the fifth node. This can be applied to situations where the third node only interacts with the first node.
[0550] In this implementation, the fifth node sends a third configuration notification to the third node. This third configuration notification indicates the network elements used to implement the operations performed by the first or fourth node, as well as the connection between the third node and the first or fourth node, and enables the third node to be well configured for the system to perform the target task.
[0551] In one possible implementation of the invention, the fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates network elements for implementing operations performed by the fourth node and the connection between the third node and the fourth node, thereby the third node receives the third configuration notification from the fifth node. This can be applied to situations where the third node only interacts with the fourth node.
[0552] In addition, the third configuration notification also instructs on the execution logic related to the third node. Upon receiving the third configuration notification, the third node performs configuration based on the third configuration notification.
[0553] It should be noted that in some cases, the third node may not only interact with the first or fourth node, but may also need to know the location of the second node, since the second node was selected by the fifth node. In this case, the third configuration notification can also indicate the network elements used to implement the operations performed by the second node, as well as the connection between the third node and the second node.
[0554] In addition to the third configuration notification, the fifth node can also notify other nodes (including the first, second, and third nodes) of the corresponding operations.
[0555] In one possible implementation of the present invention, the fifth node sends a first configuration notification, a second configuration notification, and a fourth configuration notification to the first node, the second node, and the fourth node, respectively. The first, second, and fourth configuration notifications respectively indicate the operations to be performed by the first, second, and fourth nodes for executing the target task. The operations performed by the first, second, and fourth nodes are determined by the fifth node based on description information corresponding to the target task and the type of feedback data. After receiving the first, second, and fourth configuration notifications, the first, second, and fourth nodes perform configurations based on the first, second, and fourth configuration notifications, respectively.
[0556] It should be noted that, as mentioned above, some nodes can be clustered into one node. Therefore, in this case, one configuration notification is sufficient for the clustered node. For example, if the first node and the fourth node are clustered into one node, then the first configuration notification and the fourth configuration notification can be clustered into one configuration notification.
[0557] Figure 10B The diagram illustrates one possible implementation of the above process. Specifically, in step 1, the RL client sends an RLCF problem-solving request (corresponding to the third request mentioned above) to the RLCF controller to trigger the execution of the RLCF task (target task). The RLCF problem-solving request may include information about the RLCF problem (or a general AI problem) to be solved, such as a problem description or specification, or it may also include performance requirements for the execution of the RLCF task, such as a convergence time threshold, minimum prediction accuracy requirements, etc.
[0558] Upon receiving the RLCF problem resolution request, in steps 2 and 3, the RLCF controller sends a customer feedback information request (corresponding to the second request mentioned above) to the RL customer to query customer feedback information (corresponding to the fourth instruction mentioned above). After receiving the customer feedback information request, the RL customer sends customer feedback information to the RLCF controller. The customer feedback information includes all customer feedback information queried in the customer feedback information request.
[0559] In step 4, the RLCF controller then analyzes the received customer feedback information to determine the type of customer feedback involved in the RLCF task (corresponding to the feedback data described above). The type of customer feedback can be reward-based or non-reward-based feedback as described above. Given the determined type of customer feedback, the RLCF controller resolves the information from the RLCF problem received in step 1 and determines the following:
[0560] The hyperparameters of the RL algorithm used by the RLCF task to solve the requested RLCF problem, the RLCF system applied to the RLCF task (i.e., RLCF system α or RLCF system β), and one or more trained ERPs participating in the RLCF task.
[0561] One or more trained ERPs participating in an RLCF task can be selected based on historical data (e.g., one or more trained ERPs from one or more past RLCF tasks that participated in solving the same or similar RLCF problems); or if RLCF system α is applied, all available ERPs are selected by default.
[0562] The RLCF execution process for an application used to perform RLCF tasks. The application's RLCF execution process should be determined based on the application's RLCF system and the type of customer feedback. The planned start time for the application's RLCF execution process should also be determined in this step.
[0563] The classification algorithm (and its related parameters) executed by the EC. The classification algorithm (and its related parameters) executed by the EC should be determined based on the type of customer feedback and the RLCF execution process of the application.
[0564] Then, in steps 5 through 8, the RLCF controller configures one or more network entities that implement one or more selected ERPs. These one or more network entities are configured through: one or more interfaces / connections to one or more network entities implementing the RL algorithm; one or more interfaces / connections to one or more network entities implementing the EC; one or more interfaces / connections to the RL client; and the ERP-related execution logic and steps defined during the RLCF execution of the application. The RLCF controller selects and configures one or more network entities to implement the EC. These one or more network entities are configured through: the determined classification algorithm and its associated parameters; one or more interfaces / connections to one or more network entities implementing one or more selected ERPs; one or more interfaces / connections to one or more network entities implementing the RL algorithm; one or more interfaces / connections to the RL client; and the EC-related execution logic and steps defined during the RLCF execution of the application. The RLCF controller selects and configures one or more network entities to implement the RL algorithm. The one or more network entities are configured through the following components: hyperparameters of the RL algorithm; one or more interfaces / connections with one or more network entities implementing one or more selected ERPs; one or more interfaces / connections with one or more network entities implementing ECs; one or more interfaces / connections with the RL client; and the execution logic and steps related to the RL algorithm defined during the RLCF execution of the application. The RLCF controller confirms the configured RLCF task information with the RL client. The RLCF task information includes: information on one or more interfaces / connections between the RL client and one or more network entities implementing the RL algorithm, one or more selected ERPs, and ECs; and the execution logic and steps related to the RL client defined during the RLCF execution of the application.
[0565] During the configuration process to execute target tasks initiated by a third node, feedback data provided by the third node can be used to help configure relevant nodes in the system used to execute the target tasks, thereby adapting the system to the third node and improving the accuracy of executing RLCF tasks initiated by the third node.
[0566] During the RLCF configuration process described above, since the configuration is performed based on a fourth instruction indicating the type of feedback data about the first observation provided by the third node to execute the target task initiated by the third node, the feedback data provided by the third node can be used to help configure the system for executing the target task.
[0567] As mentioned above, after training the first node and configuring the nodes, the entire RLCF system can be executed. The following section combines... Figures 11 to 15B Describe the execution process.
[0568] Figure 11 This is a flowchart of a third data processing method according to one or more embodiments of the present invention. The method can be implemented by a first node (one or more ERPs), a second node (RL algorithm), a third node (RL client), and a fourth node (EC), referred to in the description as the first node, second node, third node, and fourth node, respectively. Figure 11 As shown, the method may include the following steps.
[0569] S1101: The first node acquires the first observation, wherein the first observation is in response to the operation of the second node indicating the state of the environment.
[0570] Here, the first observation is the state of the environment obtained in response to the operation instruction of the second node. During the execution of the RLCF system, the second node can output an operation to the environment, and the state of the environment can then change in response to that operation. Therefore, the third node provides the first node with an observation indicating the state of the environment (the first observation). Simultaneously, the third node also provides feedback data about the first observation to the fourth node, enabling the first node to determine a first prediction (reward prediction) for the operation performed by the second node.
[0571] In one possible implementation of the present invention, a first node is associated with multiple feature informations, and the first feature information is included in the multiple feature informations. As described in the above embodiments, the first node can be of a first type or a second type. A node of the first type includes an integrated node associated with multiple feature informations, and a node of the second type includes multiple child nodes, each of which is associated with one feature information among the multiple feature informations. That is, the node of the first type is the I-ERP in the above-described RLCF system β, while the multiple child nodes of the node of the second type are the multiple ERPs in the above-described RLCF system α. Therefore, when the first node is of the first type, the first node as an integrated node can be associated with multiple feature informations; when the first node is of the second type, each child node of the first node can be associated with one feature information.
[0572] In this implementation, the first node includes multiple child nodes. The parameters of each child node can be optimized / trained to be associated with one of the multiple feature information. Thus, each child node can provide a reward prediction result based on the feature information, and all child nodes can provide a more accurate reward prediction result.
[0573] Furthermore, feature information is used to characterize the features of the first node and can be a professional knowledge category. The following explanation uses a professional knowledge category as an example, but it should be understood that this approach applies equally when the feature information is other types of information.
[0574] In one possible implementation of the present invention, the first node is trained using a first training dataset, which includes multiple training data sets. Each training data set includes a first data portion and a second data portion, with the second data portion indicating the feature information of the first data portion.
[0575] Specifically, the first training dataset may be included in (or under the control of) the sixth node, and includes multiple training data sets, each including an observation segment and additional professional knowledge segments indicating feature information of the observation segments; the second training dataset includes multiple sets of training data, each set of training data being associated with predefined feature information from multiple feature information sets. Furthermore, the sixth node can be used to train the first node according to steps S901 to S906 of the data processing method shown in Figure 9.
[0576] In this implementation, since the first node is trained using training data, which includes a first data portion and a second data portion containing feature information indicating the first data portion, the first node can provide a more accurate reward prediction result based on the first feature information indicated by the first indication and feedback data about the first observation.
[0577] In one possible implementation of the present invention, the first node is trained using a second training dataset, which includes multiple sets of training data, each set of training data corresponding to predefined feature information in multiple feature information.
[0578] Specifically, the second training dataset may be included in the sixth node, and includes multiple sets of training data, each set corresponding to predefined feature information from multiple feature information. Furthermore, the sixth node can be used to train the first node according to steps S901 to S906 of the data processing method shown in Figure 9.
[0579] In this implementation, since the first node is trained using training data corresponding to predefined feature information from multiple feature information, the reward prediction results provided by the first node will be more accurate.
[0580] To diversify the way the first node receives the first observation, in one possible implementation of the invention, the third node directly sends the first observation to each of the multiple child nodes of the first node of the first type or the first node of the second type, the first observation being in response to the state of the environment indicated by the operation of the second node.
[0581] In another possible implementation of the invention, the third node sends a first observation to the fourth node, the first observation being in response to the state of the environment indicated by the operation of the second node, and upon receiving the first observation, the fourth node broadcasts the first observation to each of the multiple child nodes of the first node of the first type or the first node of the second type. In this implementation, the first observation is received by the fourth node from the third node and broadcast to each of the multiple child nodes of the first node. Since the third node no longer needs to send the first observation to each of these child nodes, the processing load on the third node is reduced, and the first node is able to determine a more accurate reward prediction result based on the first observation.
[0582] S1102: The third node sends feedback data about the first observation to the fourth node, and the fourth node receives the feedback data from the third node.
[0583] Specifically, the third node sends feedback data about the first observation to the fourth node. The feedback data can be of two types: a first feedback type and a second feedback type. The first feedback type indicates the third node's evaluation of the observation, while the second feedback type indicates characteristic information of the third node. Specifically, the first feedback type can be reward-based feedback, where the value of the feedback data can be interpreted as a reward determined by the client based on the observation. The second feedback type can be non-reward-based feedback, where the value of the feedback data cannot be interpreted as a reward, but can represent one or more categories of expertise of the third node.
[0584] After receiving feedback data, the fourth node can calculate the final prediction itself, i.e., the final reward for the operation of the second node, or it can help the first node calculate the final prediction. This final prediction is related to the first feature information of the third node in the environment. Therefore, generally, when the first node is of type one, the fourth node can determine the first feature information of the third node based on the feedback data, and the fourth node can provide this information to the first node so that the first node can determine the final prediction. When the first node is of type two, the fourth node can also provide the first feature information of the third node to the first node so that the first node can determine the final prediction. When the first node is of type two, or when the fourth node can determine the final prediction itself, the final prediction can be sent from either the first or fourth node to the second node. These different implementations will become clear through the following description.
[0585] It should be noted that in the current RL step, the third node may not send feedback data to the fourth node. In this case, the fourth node can use historical feedback data (which may correspond to historical observations) for subsequent operations.
[0586] S1103: The first node obtains a first prediction of the operation for the second node based on the first observation.
[0587] Upon acquiring the first observation, the first node can determine the first prediction. This first prediction can be either the final prediction mentioned above, or a third prediction, which is used by the fourth node to generate a fourth prediction as the final prediction.
[0588] S1104: The first node or the fourth node sends the final prediction to the second node, and the second node receives the final prediction of the operation performed by the second node on the environment from the first node or the fourth node.
[0589] S1105: The second node is updated based on the final prediction.
[0590] After receiving the final prediction, the second node can be updated (e.g., updating parameters related to the algorithm implemented by the second node) to obtain a better reward in the next round of execution.
[0591] Generally, the node that determines the final prediction and the node that sends the final prediction to the second node can be the same or different; they can be the first node or the fourth node, depending on the type of the first node and the specific execution logic. To illustrate this process more clearly, the following will explain different scenarios for generating the final prediction, especially the interaction between the first and fourth nodes, with the help of different accompanying diagrams.
[0592] Scenario 1: The first node is of type 1
[0593] At this point, the first node is I-ERP, such as Figure 12A As shown, determining the final prediction may include the following steps:
[0594] S1201: The first node obtains the first observation from the third node.
[0595] For a detailed description, please refer to the relevant content in step S1101.
[0596] S1202: The fourth node receives feedback data from the third node.
[0597] For a detailed description, please refer to the relevant content in step S1102.
[0598] S1203: The fourth node determines the first feature information of the third node based on the feedback data.
[0599] Specifically, the fourth node can use a preset classification algorithm configured in the above configuration process to determine the first feature information. The configuration of this preset classification algorithm in the fourth node can be described in detail in step S1003. Determining the first feature information helps the first node output a prediction more suitable for the third node.
[0600] In this implementation, the first feature information of the third node is determined based on the feedback data from the third node regarding the first observation. The feedback data provided by the third node can be used to help identify the first feature information of the third node.
[0601] In one possible implementation of the invention, the classification algorithm used by the EC to determine the customer's expertise may require one or more rounds to converge. Before the classification algorithm converges, the determined first feature information can be represented as a weighted average of multiple expertise categories, where each expertise category should be an expertise category associated with the first node.
[0602] S1204: The fourth node sends a first instruction to the first node, wherein the first instruction indicates the first characteristic information of the third node.
[0603] Since the first indication indicates the first feature information and is determined by the fourth node based on the feedback data about the first observation provided by the third node, the feedback data provided by the third node can be used to help identify the first feature information of the third node.
[0604] S1205: The first node determines the first prediction based on the first observation and the first indication from the fourth node.
[0605] Specifically, as one of the multiple feature information associated with the first node, the first feature information can also be the professional knowledge category of the third node.
[0606] As described above, when training the first node, since the first node (I-ERP) is trained using a first training dataset that includes multiple training data, each training data includes a first data part (observation fragment part) and a second data part (additional expertise fragment), and the second data part indicates the feature information of the first data part, the input data consists of the first observation and the first feature information indicated by the first indication, and the first node can therefore obtain the first prediction based on the first observation and the first feature information.
[0607] S1206: The first node sends the first prediction to the second node, and the second node receives the first prediction.
[0608] S1207: The second node is updated based on the first prediction.
[0609] For a detailed description, please refer to the relevant content in steps S1104 and S1105.
[0610] Steps S1201 to S1207 described above can be executed iteratively in each RL step. Here, an RL step refers to an RL training step within an RL training round.
[0611] Figure 12B It corresponds to Figure 12A One possible implementation of the process is shown, where the main entities are represented as RL clients, RL algorithms, EC, and I-ERP, and the information exchanged between the different entities is described using different names. However, the principles shown in the two figures are similar.
[0612] Specifically, in Figure 12B In step 1, the RL algorithm interacts with the environment to complete the traditional RL execution steps without updating the RL algorithm's parameters / one or more models. Then, in step 2, the RL client obtains the current observation from the environment (corresponding to the first observation mentioned above) and sends it to the I-ERP. Then, in step 3, the RL client determines client feedback (corresponding to the feedback data mentioned above) based on the current observation (and in some embodiments, one or more past observations), and then sends the client feedback to the EC. In some embodiments, the RL client may choose not to send client feedback in the current RL step.
[0613] Then, in step 4, the EC runs a classification algorithm based on the customer feedback received in step 3 to determine the expertise of the RL customer (corresponding to the first feature information mentioned above). The classification algorithm used by the EC to determine the customer's expertise may require one or more rounds to converge. Before the classification algorithm converges, the determined customer expertise can be represented as a weighted average of multiple expertise categories, where each expertise category should be an associated expertise category of the ERP.
[0614] After determining the expertise of the RL customer, in step 5, the EC sends the determined customer expertise customer category indicator (corresponding to the first indicator mentioned above) to the I-ERP.
[0615] Then, in step 6, I-ERP concatenates the observations received in step 2 and the customer expertise indicators received in step 5 to form input data, and uses the input data to perform reward prediction to generate the final predicted reward (corresponding to the final prediction mentioned above). Then, in step 7, I-ERP sends the calculated final predicted reward to the RL algorithm, and in step 8, the RL algorithm updates the parameters / one or more models based on the received final predicted reward.
[0616] Scenario 2: The first node is of type 2
[0617] In this case, the first node is of the second type and therefore includes multiple child nodes, each associated with a feature (a category of expertise).
[0618] Generally, the first or fourth node can determine the final prediction, and the feedback data used to determine the final prediction can be of either a first feedback type or a second feedback type. The first feedback type of the feedback data indicates the third node's evaluation of the observation, while the second feedback type indicates the third node's feature information. Therefore, the determination of the final prediction can be described below for these cases.
[0619] Figure 13A and Figure 13B The process of the first node determining the final prediction (referred to as the first prediction in this embodiment) is shown, wherein the feedback data is a second feedback type (non-reward-based feedback).
[0620] like Figure 13A As shown, determining the final prediction may include the following steps:
[0621] S1301: The first node obtains the first observation from the third node.
[0622] For a detailed description, please refer to the relevant content in step S1101.
[0623] S1302: The fourth node receives feedback data from the third node.
[0624] For a detailed description, please refer to the relevant content in step S1102.
[0625] S1303: The fourth node determines the second instruction based on the feedback data.
[0626] The second indication is used to indicate the first feature information of the third node. As one of multiple feature information associated with the first node, the first feature information can also be the professional knowledge category of the third node.
[0627] In one possible implementation of the invention, the fourth node determines the professional knowledge category of the third node by running a classification algorithm based on feedback data. In another possible implementation, the classification algorithm used by the fourth node to determine one or more professional knowledge categories of the third node may require one or more rounds to converge.
[0628] In one possible implementation, the second indicator may indicate only the first feature information of the third node. In another possible implementation, the second indicator may indicate both the first feature information of the third node and the weights corresponding to its child nodes. These weights can be determined by the fourth node by comparing the first feature information of the third node with the feature information associated with each child node. For example, if the first feature information of the third node and the feature information associated with its child nodes are the same, the weight is 1; otherwise, the weight is 0.
[0629] S1304: The fourth node sends a second instruction to the first node, and the first node receives the second instruction.
[0630] Specifically, since the first node includes multiple child nodes, the fourth node can send a second indication to each of the multiple child nodes. If the second indication indicates the first feature information of the third node, all child nodes can receive the same second indication; and if the second indication also indicates the weight corresponding to the child node, each child node can receive its corresponding second indication, because the indication carries its weight.
[0631] S1305: The first node obtains a first prediction of the operation for the second node based on the first observation.
[0632] Specifically, for each of the multiple child nodes, the child node can determine a second prediction for the operation against the second node based on a first observation, and determine the weight corresponding to the second prediction based on a second indication from the fourth node. Thus, the first child node can determine the first prediction as a weighted sum of the second predictions from the multiple child nodes, thereby enabling the first child node to provide a more accurate reward prediction. Here, the second prediction is the initial prediction determined by each child node and needs to be weighted.
[0633] As described in the above embodiments, since the child nodes of the first node have been trained based on training data from a second training dataset associated with predefined feature information, the child node is able to provide a second prediction for the first observation, and the second prediction of each child node is related to its feature information.
[0634] It should be noted that when the multiple child nodes of the first node execute the steps of the above data processing method, they can be implemented as sub-modules of the first node.
[0635] In one possible implementation of the invention, for each of a plurality of child nodes, a second instruction indicates the first feature information of a third node. Upon receiving the second instruction, the child node determines a weight corresponding to a second prediction based on the second instruction and the feature information associated with the child node. For example, each child node may compare its associated feature information (e.g., expertise category) with the first feature information of the third node; if they are the same, the weight is 1; if they are different, the weight is 0.
[0636] In this implementation, for each of the multiple child nodes, a second indication indicating the first feature information of the third node is sent to the multiple child nodes to determine the weight corresponding to the second prediction. Therefore, the determined weight can be used to calibrate the initial second prediction, thereby enabling the first child node to provide a more accurate reward prediction result based on the calibrated second prediction.
[0637] After determining the weights corresponding to the second prediction based on the second indication from the fourth node, the first child node among the multiple child nodes obtains a weighted sum of the second predictions of the multiple child nodes as the first prediction, based on the second prediction of each of the multiple child nodes and the weights corresponding to the second prediction. Here, the first child node can be a predefined anchor node or an anchor node selected by the fourth node from the multiple child nodes based on the similarity between the feature information associated with each child node and the first feature information of the third node.
[0638] In this implementation, the first child node, which serves as the anchor node, undertakes more interaction tasks than other child nodes. This first child node can be predefined or selected by the fourth node, thereby simplifying the operation of the data processing method.
[0639] In one possible implementation, each child node can determine its own weighted prediction and then send its weighted sum to the first child node, enabling the first node to determine the entire weighted sum as the first prediction (final prediction). For example, a child node can multiply its second prediction by its weights to obtain its weighted prediction. In another possible implementation, each child node can send its second prediction and corresponding weights to the first child node, allowing the first child node to determine the first prediction as the weighted sum of the second predictions of all child nodes based on the second predictions of each of the child nodes and the corresponding weights. In yet another possible implementation, some child nodes can calculate their own weighted predictions, while others do not. The first child node can then determine the weighted sum based on what it receives from the other child nodes. The method of obtaining the weighted sum is not limited here.
[0640] In one possible implementation of the invention, for each of the plurality of child nodes, the second indication further indicates the weight corresponding to the child node. In this case, the determination of the weight is completed by the fourth node as described above.
[0641] It should be noted that each of the multiple child nodes determines the second prediction based on the second indication, which may be received by the first node in this instance or previously received; the present invention does not limit this.
[0642] S1306: The first node sends the first prediction to the second node, and the second node receives the first prediction.
[0643] Specifically, since the first child node obtained the first prediction in step S1305, it will send the first prediction to the second node.
[0644] Therefore, the first prediction sent by the first node to the second node is used to update the second node, making the parameters / one or more models of the second node updated according to the first prediction more accurate.
[0645] S1307: The second node is updated based on the first prediction.
[0646] For a detailed description, please refer to the relevant content in steps S1104 and S1105.
[0647] Steps S1301 to S1307 described above can be executed iteratively within each RL step. Here, an RL step refers to one RL training step in an RL training round. Furthermore, it should be noted that step S1303 and the determination of the second prediction for each child node can be executed in parallel or sequentially; this is not limited here.
[0648] Figure 13B It corresponds to Figure 13A The diagram illustrates one possible implementation of the process, where the main entities are represented as RL clients, RL algorithms, ECs, and ERPs, and the information exchanged between different entities is described using different names. However, the principles shown in both diagrams are similar.
[0649] Specifically, in Figure 13B In step 1, the RL algorithm interacts with the environment to complete the traditional RL execution steps without updating the RL algorithm's parameters / one or more models. Then, in step 2, the RL client obtains the current observation (corresponding to the first observation mentioned above) from the environment and sends it to the ERP. Then, in step 3, the RL client determines client feedback (corresponding to the feedback data mentioned above) based on the current observation (and in some embodiments, one or more past observations), and then sends the client feedback to the EC. In some embodiments, the RL client may choose not to send client feedback in the current RL step.
[0650] Then, in step 4, the EC interacts with the ERPs to predict the final predicted reward (corresponding to the first prediction mentioned above). Each ERP independently calculates its predicted reward based on the current observations received in step 2 (corresponding to the third prediction mentioned above). The EC runs a classification algorithm (e.g., KNN) to determine the expertise of RL customers (corresponding to the first feature information mentioned above) based on customer feedback received in step 3. In some embodiments, the classification algorithm used by the EC to determine customer expertise may require one or more rounds to converge. Before the classification algorithm converges, the determined customer expertise can be represented as a weighted average of multiple expertise categories, where each expertise category should be an associated expertise category of the ERP. The EC sends customer expertise composition information to each ERP (step 4.3). The customer expertise composition information includes the weights of the associated expertise categories of each ERP. For a converged classification algorithm, the weights sent to ERPs with the same associated expertise category as the customer expertise category are 1, and the weights sent to one or more other ERPs are 0. Each ERP then synchronizes the received customer expertise composition information with each other (e.g., sends the received customer expertise composition information to the anchor ERP), which calculates the final predicted reward as a weighted average of the predicted rewards of all ERPs. During the calculation process, the weight applied to each ERP predicted reward is equal to the corresponding weight received by the ERP in the customer expertise composition information. In a certain RL step where customer expertise composition information was not received (i.e., step 4.2 was not executed because no new customer feedback was received in the RL step), one or more weights received in past RL steps can be applied to calculate the final predicted reward.
[0651] Figure 14A and Figure 14B The process is illustrated whereby the fourth node determines the final prediction (represented as the first prediction in this embodiment) and the second node sends the final prediction to the first node, wherein the feedback data is of the first feedback type (reward-based feedback).
[0652] like Figure 14A As shown, determining the final prediction may include the following steps:
[0653] S1401: The first node obtains the first observation from the third node.
[0654] For a detailed description, please refer to the relevant content in step S1101.
[0655] S1402: The fourth node receives feedback data from the third node.
[0656] For a detailed description, please refer to the relevant content in step S1102.
[0657] S1403: The first node determines a third prediction of the operation for the second node based on the first observation.
[0658] Specifically, since the first node includes multiple child nodes, each child node can determine the third prediction as the initial prediction. The process of a child node determining the third prediction is the same as that described in step S1305 for determining the second prediction, and will not be repeated here for clarity. These two predictions are represented by different names simply because they are used in different processes.
[0659] S1404: The first node sends the third prediction to the fourth node, and the fourth node receives the third prediction.
[0660] S1405: The fourth node determines the first prediction for the operation of the second node based on the third prediction and the feedback data from the third node.
[0661] Specifically, the fourth node can determine the final prediction of the operation for the second node based on the third prediction and feedback data from the third node regarding the first observation, wherein the final prediction corresponds to the first feature information of the third node in the environment.
[0662] In one or more embodiments of the present invention, determining the final prediction based on the third prediction and feedback data from the third node regarding the first observation includes:
[0663] The fourth node determines the similarity between the feedback data and the third prediction from each of the multiple child nodes;
[0664] The fourth node determines the final prediction from each of the multiple child nodes based on the third prediction and the similarity between the feedback data and the third prediction.
[0665] In this implementation, the final prediction is determined from each of the multiple child nodes based on the third prediction and the similarity between the feedback data and the third prediction, which helps to provide more accurate reward prediction results.
[0666] For example, if at least one third prediction has a similarity greater than or equal to a preset threshold, the fourth node determines the third prediction with the highest similarity as the final prediction. If no third prediction has a similarity greater than or equal to the preset threshold, the fourth node determines the weight of the third prediction corresponding to each of the multiple child nodes based on the similarity between the feedback data and the third prediction, and calculates a weighted sum of the third predictions of the multiple child nodes as the final prediction based on the third predictions of each of the multiple child nodes and the corresponding weights. If the fourth node does not receive feedback data, it uses historical feedback data to determine the final prediction.
[0667] In this implementation, the final prediction can be determined in multiple ways based on the third prediction, thereby providing a more accurate reward prediction result.
[0668] In the current situation, the final prediction determined by the fourth node is called the first prediction.
[0669] S1406: The fourth node sends the first prediction to the first node, and the first node receives the first prediction from the fourth node.
[0670] In one possible implementation, the fourth node can select a first child node from among the first node's multiple child nodes and send a third indication indicating the first prediction to the first child node. The first child node can be selected based on the similarity of the third predictions of the multiple child nodes. For example, the first child node serving as the anchor node could be the child node with the highest similarity between the third prediction and the final prediction. In another possible implementation, the first child node can also be predefined. Furthermore, since all child nodes know their respective third predictions, if the third prediction with the highest similarity is selected as the final prediction, and the corresponding child node is selected as the first child node, the fourth node can issue a third indication only to the first child node, instead of sending the first prediction to the first child node.
[0671] In this implementation, each of the multiple child nodes determines a third prediction and sends it to a fourth node. This allows the fourth node to determine a third instruction based on the third prediction and feedback data, instructing the first prediction (which would be a more accurate reward prediction result for the second node's operation). The fourth node then sends this third instruction to the first child node, enabling the first child node to provide a more accurate reward prediction result. Furthermore, this calculation is performed by the fourth node, reducing the processing load on the first node. Since the first prediction can be sent from the first child node to the second node, the fourth node does not need to interact directly with the second node, thus simplifying its operation.
[0672] S1407: The first node sends the first prediction to the second node, and the second node receives the first prediction.
[0673] Specifically, since the first child node obtained the first prediction in step S1406, it will send the first prediction to the second node.
[0674] S1408: The second node is updated based on the first prediction.
[0675] For a detailed description, please refer to the relevant content in steps S1104 and S1105.
[0676] It should be noted that the fourth node can also send the first prediction directly to the second node.
[0677] Steps S1401 to S1408 described above can be executed iteratively in each RL step. Here, an RL step refers to an RL training step within an RL training round.
[0678] Figure 14B It corresponds to Figure 14A The diagram illustrates one possible implementation of the process, where the main entities are represented as RL clients, RL algorithms, ECs, and ERPs, and the information exchanged between different entities is described using different names. However, the principles shown in both diagrams are similar.
[0679] Specifically, in Figure 14B In step 1, the RL algorithm interacts with the environment to complete the traditional RL execution steps without updating the RL algorithm's parameters / one or more models. Then, in step 2, the RL client obtains the current observation (corresponding to the first observation mentioned above) from the environment and sends it to the ERP. Then, in step 3, the RL client determines client feedback (corresponding to the feedback data mentioned above) based on the current observation (and in some embodiments, one or more past observations), and then sends the client feedback to the EC. In some embodiments, the RL client may choose not to send client feedback in the current RL step.
[0680] Then, in step 4, each ERP independently calculates its predicted ERP reward (corresponding to the third prediction described above) based on the current observation received in step 2. Each ERP then sends its predicted ERP reward to the EC. After receiving predicted ERP rewards from all ERPs, the EC calculates the final predicted reward (corresponding to the first prediction described above) based on the received predicted ERP rewards and the customer feedback received in step 3. In some embodiments, the EC compares the similarity level between each predicted ERP reward (received in 4.2) and the customer feedback (past customer feedback values can be used if no customer feedback is received in this RL step). In this step, the predicted ERP reward with the highest similarity to the customer feedback is selected as the final predicted reward. In some embodiments, if one or more similarity levels between the customer feedback and all or all of the predicted ERP rewards are below a predetermined threshold, the final predicted reward can be determined as a weighted average of the multiple predicted ERP rewards received in step 4.2. The weight assigned to each predicted ERP reward can be determined by the similarity level between the customer feedback value and the predicted ERP reward value.
[0681] Next, the EC sends the final predicted reward information to one of these ERPs (i.e., the anchor ERP). The final predicted reward information may include the final predicted reward value. The anchor ERP can be predetermined during the RLCF configuration process, or it can be the ERP whose predicted reward is selected as the final predicted reward. In the latter case, the final predicted reward information can only contain one indicator. The anchor ERP that receives the final predicted reward information in step 4.4 can forward the received final predicted reward (or its own generated ERP predicted reward) to the RL algorithm.
[0682] Figure 15A and Figure 15B The process is illustrated in which the fourth node determines the final prediction (referred to as the fourth prediction in this embodiment) and sends the final prediction to the second node, wherein the feedback data is of either a first feedback type or a second feedback type (reward-based / non-reward-based feedback).
[0683] like Figure 15A As shown, determining the final prediction may include the following steps:
[0684] S1501: The first node obtains the first observation from the third node.
[0685] For a detailed description, please refer to the relevant content in step S1101.
[0686] S1502: The fourth node receives feedback data from the third node.
[0687] For a detailed description, please refer to the relevant content in step S1102.
[0688] S1503: The first node determines a third prediction of the operation for the second node based on the first observation.
[0689] For a detailed description, please refer to the relevant content in step S1403.
[0690] S1504: The first node sends the third prediction to the fourth node, and the fourth node receives the third prediction.
[0691] S1505: The fourth node determines the fourth prediction for the operation of the second node based on the third prediction and feedback data from the third node.
[0692] When the feedback data is of the first feedback type (reward-based feedback), the fourth node first determines the first feature information of the third node based on the feedback data. For example, the fourth node determines the professional knowledge category of the third node by running a classification algorithm based on the feedback data. In one or more embodiments of the present invention, the classification algorithm used by the fourth node to determine one or more professional knowledge categories of the third node may require one or more rounds to converge.
[0693] Then, for each of the multiple child nodes, the fourth node determines the weight of the third prediction corresponding to the child node based on the first feature information of the third node and the feature information associated with the child node. Based on the third prediction of each child node and the weight corresponding to the third prediction, the fourth node calculates a weighted sum of the third predictions of the multiple child nodes as the final prediction, which is the fourth prediction. The weights here can be determined by the fourth node by comparing the first feature information of the third node with the feature information associated with each child node. For example, if the first feature information of the third node and the feature information associated with the child node are the same, the weight is 1; otherwise, the weight is 0.
[0694] When the feedback data is of the second feedback type (non-reward-based feedback), determining the fourth prediction can be described in detail in step S1405, referring to the determination of the first prediction. Both predictions are final predictions; the difference lies in... Figure 14A In the process shown, the first prediction, which is the final prediction, is sent from the fourth node to the second node, but in Figure 15A In the process shown, the fourth prediction, which is the final prediction, is sent from the fourth node to the second node.
[0695] S1506: The fourth node sends the fourth prediction to the second node, and the second node receives the fourth prediction.
[0696] S1507: The second node is updated based on the fourth prediction.
[0697] In this implementation, each of the multiple child nodes determines a third prediction and sends it to a fourth node. This allows the fourth node to determine a fourth prediction (which would be a more accurate reward prediction for the operation of the second node) based on the third prediction and feedback data about the first observation provided by the third node. This enables the fourth node to provide a more accurate reward prediction. Furthermore, since this calculation is performed by the fourth node, the processing load on the first node is reduced, and the first prediction can be sent directly from the fourth node to the second node, thus reducing signaling overhead.
[0698] For a detailed description, please refer to the relevant content in steps S1104 and S1105.
[0699] Figure 15B It corresponds to Figure 15A The diagram illustrates one possible implementation of the process, where the main entities are represented as RL clients, RL algorithms, ECs, and ERPs, and the information exchanged between different entities is described using different names. However, the principles shown in both diagrams are similar.
[0700] Specifically, in Figure 15B In step 1, as shown, the RL algorithm interacts with the environment to complete the traditional RL execution steps without updating the RL algorithm's parameters / one or more models. Then, in step 2, the RL client obtains the current observations from the environment (corresponding to the first observations mentioned above).
[0701] Then, in step 3, the EC interacts with the ERPs to predict the final predicted reward (corresponding to the fourth prediction mentioned above). The specific interactions in this step include: the EC broadcasting the observations received in the current RL step to all ERPs. Each ERP independently calculates its predicted reward based on the current observations received in step 3.1. Each ERP sends its predicted reward to the EC (corresponding to the third prediction mentioned above). After receiving one or more predicted rewards from all or more ERPs, the EC calculates the final predicted reward based on the one or more predicted rewards received in step 3.3 and the customer feedback received in step 2. If reward-based feedback is applied, the EC... Figure 14B The same method described in step 4.3 is used to calculate the final predicted reward. If non-reward-based feedback is applied, the EC is calculated via... Figure 13B The same method described in step 4.2 is used to calculate the final predicted reward. The EC then sends the calculated final predicted reward to the RL algorithm.
[0702] By obtaining the final prediction for updating the second node related to the first feature information of the third node in the environment, the first or fourth node can provide more accurate reward prediction results, thereby enabling the second node to generate an operation that matches the first feature information of the third node.
[0703] It should be noted that some of these nodes can also be integrated into a single node, or implemented as different functional modules on the same network element. In this case, the information exchange between these functional modules can be considered as internal information exchange. The information exchanged between these functional modules and other nodes can also be integrated accordingly. For example, if the second node and the first node are implemented in the same network element, the information received or sent by the second node and the first node will be considered as being sent to or from the same network element. For clarity, this will not be elaborated further here.
[0704] Embodiments of this invention provide a method for performing reinforcement learning using customer feedback, and provide multiple ERPs or integrated ERPs (I-ERPs) to represent preferences / behaviors of different expertise categories. Furthermore, based on customer feedback, the EC can dynamically categorize customers into a expertise category and adjust the final reward prediction based on the customer's expertise (via the ERP or I-ERP). Additionally, one or more trained ERPs (or I-ERPs) and ECs can be reused by different RLCF tasks.
[0705] Furthermore, embodiments of the present invention provide new NFs and signaling to support RLCF, as well as signaling between ERP (I-ERP), EC, and network elements (PSF) implementing RL algorithms / models. For example, ERP or I-ERP can be implemented as DP function (processing service function (PSF)) or CP function (task control function (TCF)), and EC can be implemented as CP function (TCF) or DP function (PSF).
[0706] According to the technical solution provided by this invention, an ERP pre-trained with data from customer expertise categories can provide more accurate reward predictions compared to a general reward predictor. Furthermore, the final predicted reward can be dynamically determined based on customer feedback, achieving faster convergence compared to retraining a general reward predictor using a large amount of customer feedback.
[0707] During the RLCF execution process described above, by obtaining the first prediction related to the first feature information of the third node in the environment, which is used to update the second node, the first node is able to provide a more accurate reward prediction result, thereby enabling the second node to generate an operation that matches the first feature information of the third node.
[0708] Figure 16 This is a schematic diagram of a data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an RLCF controller. Figure 16 As shown, the data processing device 1600 may include:
[0709] The sending module 1601 is used to send a first request to the sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information;
[0710] The receiving module 1602 is configured to receive a first response from the sixth node, wherein the first response is determined by the sixth node based on the first request and instructs the sixth node to support the ability to train the first node;
[0711] Processing module 1603 is used to configure the first node and the sixth node to perform training based on the first response.
[0712] In one possible implementation of the present invention, the processing module 1603 is used for:
[0713] Based on the first response, first configuration information and second configuration information of the first node are determined, wherein the first configuration information indicates the type of the first node and the second configuration information indicates the parameters used to train the first node;
[0714] Configure the first node based on the first configuration information and the second configuration information;
[0715] The sending module 1601 is used for:
[0716] Notify the sixth node of the configuration of the first node.
[0717] In one possible implementation of the present invention, the processing module 1603 is used for:
[0718] The type of the first node is determined based on the first response, wherein the type of the first node includes a first type and a second type. The first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of which is associated with one feature information among the multiple feature information.
[0719] The sending module 1601 is used for:
[0720] Notify the sixth node of the type of the first node.
[0721] In one possible implementation of the invention, the first response includes an indication of feature information supported by the sixth node, and an indication of the sixth node including the data required to train the first node.
[0722] Processing module 1603 is used for:
[0723] Based on the indication of the feature information supported by the sixth node, determine the number of candidate nodes to implement the first node;
[0724] Based on the indication from the sixth node, including the data required to train the first node, the second configuration information is determined.
[0725] In one possible implementation of the present invention, the second configuration information includes parameters of a first type and parameters of a second type, wherein the parameters of the first type remain unchanged during the training of the first node, and the parameters of the second type are updated based on the training data provided by the sixth node during the training of the first node.
[0726] In one possible implementation of the present invention, the parameters of the first type include at least one of the structure of the neural network for training the first node and the method for training the first node;
[0727] The second type of parameters includes at least one of the initial weights and biases of the neural network.
[0728] In one possible implementation of the present invention, the processing module 1603 is used for:
[0729] Based on the first configuration information and the second configuration information, network elements for implementing the first node and the connection between the sixth node and the first node are configured.
[0730] In one possible implementation of the present invention, the sending module 1601 is used for:
[0731] Send a third message to the sixth node, wherein the third message indicates at least one path between the first node and the sixth node.
[0732] In one possible implementation of the present invention, the sending module 1601 is used for:
[0733] Send the fourth message to the sixth node, where the fourth message indicates the training dataset of the first node.
[0734] In one possible implementation of the invention, the fourth information also includes the format of the data used to train the first node.
[0735] In one possible implementation of the invention, the first information includes at least one indication of at least one feature information, and the second information includes a minimum amount of data associated with each of the at least one feature information.
[0736] This data processing device can be applied to the above. Figure 9A and Figure 9B The fifth node described in the method embodiment shown, and / or may be the one described above. Figure 9A and Figure 9B The fifth node described in the illustrated method embodiment. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 9A and Figure 9B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0737] Figure 17 This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be a human feedback dataset. Figure 17 As shown, the data processing device 1700 may include:
[0738] The receiving module 1701 is used to receive a first request from the fifth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information;
[0739] The determination module 1702 is used to determine a first response based on a first request, wherein the first response indicates that the sixth node supports the ability to train the first node;
[0740] The sending module 1703 is used to send the first response to the fifth node.
[0741] In one possible implementation of the invention, the first response includes an indication of feature information supported by the sixth node, and an indication of the sixth node including the data required to train the first node.
[0742] Module 1702 is used for:
[0743] Indication of the feature information supported by the sixth node based on the first information;
[0744] Based on the second information, an indication is determined for the sixth node, including the data required to train the first node.
[0745] In one possible implementation of the present invention, the receiving module 1701 is further configured to:
[0746] Receive third information from the fifth node, wherein the third information indicates at least one path between the first node and the sixth node.
[0747] In one possible implementation of the present invention, the receiving module 1701 is further configured to:
[0748] Receive a training request from the first node, wherein the training request indicates the feature information of the first node;
[0749] The determination module 1702 is also used for:
[0750] Based on the type of the first node and the training request, determine the training data to be used for the first node;
[0751] The sending module 1703 is also used for:
[0752] Training data is sent to the first node based on the third information, so that the first node can perform training based on the training data.
[0753] In one possible implementation of the present invention, the receiving module 1701 is further configured to:
[0754] Receive the fourth information from the fifth node, where the fourth information indicates the training dataset of the first node.
[0755] In one possible implementation of the present invention, the determining module 1702 is further configured to:
[0756] When the first node is determined to be of the first type and includes an integrated node associated with multiple feature information, the first training dataset is determined based on the third information.
[0757] The sending module 1703 is also used for:
[0758] The training data of the first training dataset is sent to the first node based on the fourth information, so that the first node can perform training based on the training data of the first training dataset.
[0759] The determination module 1702 is also used for:
[0760] When it is determined that the first node is of the second type and includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information, the second training dataset is determined based on the third information.
[0761] The sending module 1703 is also used for:
[0762] The training data of the second training dataset is sent to each of the multiple child nodes based on the fourth information, so that the first node can perform training based on the training data of the second training dataset.
[0763] The first training dataset includes multiple training data sets, each of which includes a first data part and a second data part, with the second data part indicating the feature information of the first data part; the second training dataset includes multiple sets of training data, each set of training data being associated with predefined feature information from multiple feature information sets.
[0764] In one possible implementation of the invention, the type of the first node is predefined or notified by the sixth node.
[0765] This data processing device can be applied to the above. Figure 9A and Figure 9B The sixth node described in the method embodiment shown, and / or may be the one described above. Figure 9A and Figure 9B The sixth node described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 9A and Figure 9B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0766] Figure 18 This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be one or more ERP systems shown in FIG. 7. Figure 8The I-ERP shown is an example. Figure 18 As shown, the data processing device 1800 may include:
[0767] The receiving module 1801 is used to receive training data from the sixth node, wherein the training data corresponds to the feature information of the first node;
[0768] Training module 1802 is used to perform training based on training data from the sixth node.
[0769] In one possible implementation of the present invention, the data processing apparatus 1800 further includes:
[0770] The sending module 1803 is used to send a training request to the sixth node, wherein the training request indicates the feature information of the first node.
[0771] In one possible implementation of the present invention, the first node is of a first type and includes an integrated node associated with multiple feature information;
[0772] Receiver module 1801 is used for:
[0773] The sixth node receives training data from the first training dataset, wherein the first training dataset includes multiple training data, each training data includes a first data part and a second data part, and the second data part indicates the feature information of the first data part.
[0774] In one possible implementation of the present invention, the first node is of the second type and includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information;
[0775] Receiver module 1801 is used for:
[0776] The sixth node receives training data from the second training dataset, which includes multiple sets of training data, each set of training data being associated with predefined feature information from multiple feature information.
[0777] This data processing device can be applied to the above. Figure 9A and Figure 9B The first node described in the method embodiment shown, and / or may be the one described above. Figure 9A and Figure 9B The first node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 9A and Figure 9B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0778] Figure 19This is a schematic diagram of another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an RLCF controller. Figure 19 As shown, the data processing device 1900 may include:
[0779] Acquisition module 1901 is used to acquire a fourth indication from a third node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node;
[0780] Configuration module 1902 is used to configure based on the fourth instruction to execute the target task initiated by the third node.
[0781] In one possible implementation of the present invention, the data processing apparatus 1900 further includes:
[0782] The first sending module 1903 is used to send a second request to the third node, wherein the second request is used to request the type of feedback data;
[0783] The acquisition module 1901 is used to receive the fourth instruction from the third node.
[0784] In one possible implementation of the invention, the second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein...
[0785] The fourth instruction also instructs at least one of the items requested by the fifth node.
[0786] In one possible implementation of the present invention, the acquisition module 1901 is used for:
[0787] Receive a third request from the third node, wherein the third request indicates description information corresponding to the target task.
[0788] In one possible implementation of the present invention, the configuration module 1902 is used for:
[0789] Based on the description information corresponding to the target task, the first node and the second node are determined to execute the target task;
[0790] Based on the type of feedback data and the first node, the operations to be performed by the first node, the second node, and the fourth node are determined to execute the target task, wherein the fourth node is used to provide predictions for the observations provided by the third node;
[0791] Based on the operations performed by the first node, the second node, and the fourth node, configure the network elements and connections used to implement the operations performed by the first node, the second node, and the fourth node;
[0792] The data processing device 1900 also includes:
[0793] The second sending module 1904 is used to: notify the third node of the configuration to execute the target task based on the network elements and connections used to implement the operations performed by the first node, the second node and the fourth node.
[0794] In one possible implementation of the present invention, the configuration module 1902 is used for:
[0795] Based on the description information corresponding to the target task, at least one candidate node that has performed a historical task is identified as the first node, wherein the similarity between the historical task and the target task is above a preset threshold.
[0796] In one possible implementation of the present invention, the configuration module 1902 is used for:
[0797] Based on the description information corresponding to the target task, the type of the first node is determined. The type of the first node includes a first type and a second type. The first type of node includes an integrated node associated with multiple feature information. The second type of node includes multiple child nodes, and each of the multiple child nodes is associated with one feature information among the multiple feature information.
[0798] When the first node is determined to be of the second type, all available candidate nodes are determined to be child nodes of the first node.
[0799] In one possible implementation of the present invention, the second transmitting module 1904 is used for:
[0800] Send a third configuration notification to the third node, wherein the third configuration notification indicates the network element used to implement the operation performed by the first node and the connection between the third node and the first node; or
[0801] A third configuration notification is sent to the third node, wherein the third configuration notification indicates the network element used to implement the operation performed by the fourth node and the connection between the third node and the fourth node.
[0802] In one possible implementation of the present invention, the second transmitting module 1904 is further configured to:
[0803] Send a first configuration notification to the first node, wherein the first configuration notification indicates the operation to be performed by the first node;
[0804] Send a second configuration notification to the second node, wherein the second configuration notification indicates the operation to be performed by the second node;
[0805] Send a fourth configuration notification to the fourth node, wherein the fourth configuration notification indicates the operation to be performed by the fourth node.
[0806] In one possible implementation of the present invention, the type of feedback data includes a first feedback type and a second feedback type. The first feedback type of the feedback data indicates the evaluation of the observation by the third node, and the second feedback type of the feedback data indicates the feature information of the third node.
[0807] This data processing device can be applied to the above. Figure 10A and Figure 10B The fifth node described in the method embodiment shown, and / or may be the one described above. Figure 10A and Figure 10B The fifth node described in the illustrated method embodiment. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 10A and Figure 10B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0808] Figure 20 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be one or more ERP systems shown in FIG. 7. Figure 8 The I-ERP shown is an example. Figure 20 As shown, the data processing device 2000 may include:
[0809] The receiving module 2001 is configured to receive a first configuration notification from the fifth node, wherein the first configuration notification indicates an operation performed by the first node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node.
[0810] Configuration module 2002 is used for configuration based on the first configuration notification.
[0811] This data processing device can be applied to the above. Figure 10A and Figure 10B The first node described in the method embodiment shown, and / or may be the one described above. Figure 10A and Figure 10B The first node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 10A and Figure 10B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0812] Figure 21 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an RL algorithm. For example... Figure 21 As shown, the data processing device 2100 may include:
[0813] The receiving module 2101 is configured to receive a second configuration notification from the fifth node, wherein the second configuration notification indicates an operation performed by the second node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node.
[0814] Configuration module 2102 is used for configuration based on the second configuration notification.
[0815] This data processing device can be applied to the above. Figure 10A and Figure 10B The second node described in the method embodiment shown, and / or may be the one described above. Figure 10A and Figure 10B The second node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 10A and Figure 10B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0816] Figure 22 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. This data processing apparatus may be an RL client. For example... Figure 22 As shown, the data processing device 2200 may include:
[0817] The sending module 2201 is used to send a fourth indication to the fifth node so that the fifth node can be configured based on the fourth indication to execute the target task initiated by the third node, wherein the fourth indication indicates the type of feedback data provided by the third node.
[0818] In one possible implementation of the present invention, the data processing apparatus 2200 further includes:
[0819] The first receiving module 2202 is used to receive a second request from the fifth node, wherein the second request is used to request the type of feedback data.
[0820] In one possible implementation of the invention, the second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein...
[0821] The fourth instruction also instructs at least one of the items requested by the fifth node.
[0822] In one possible implementation of the present invention, the sending module 2201 is further configured to:
[0823] Send a third request to the fifth node, wherein the third request indicates description information corresponding to the target task.
[0824] In one possible implementation of the present invention, the data processing apparatus further includes:
[0825] The second receiving module 2203 is configured to receive a third configuration notification from the fifth node, wherein the third configuration notification indicates the network element for implementing the operation performed by the first node and the connection between the third node and the first node; or the third configuration notification indicates the network element for implementing the operation performed by the fourth node and the connection between the third node and the fourth node.
[0826] Configuration module 2204 is used for configuration based on a third configuration notification.
[0827] This data processing device can be applied to the above. Figure 10A and Figure 10B The third node described in the method embodiments shown, and / or may be the one described above. Figure 10A and Figure 10B The third node described in the illustrated method embodiment. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 10A and Figure 10B The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0828] Figure 23 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an EC (Electronic Processing Unit). Figure 23 As shown, the data processing device 2300 may include:
[0829] The receiving module 2301 is configured to receive a fourth configuration notification from the fifth node, wherein the fourth configuration notification indicates an operation to be performed by the fourth node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node.
[0830] Configuration module 2302 is used for configuration based on the fourth configuration notification.
[0831] This data processing device can be applied to the above. Figure 10A and Figure 10B The fourth node described in the method embodiments shown, and / or may be the one described above. Figure 10A and Figure 10B The fourth node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 10A and Figure 10BThe embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0832] Figure 24 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be one or more ERP systems shown in FIG. 7. Figure 8 The I-ERP shown is an example. Figure 24 As shown, the data processing device 2400 may include:
[0833] The acquisition module 2401 is used to acquire a first observation, wherein the first observation responds to the state of the environment indicated by the operation of the second node;
[0834] The processing module 2402 is used to obtain a first prediction of the operation for the second node based on the first observation, wherein the first prediction is related to the first feature information of the third node in the environment, and the first prediction is used to update the second node.
[0835] In one possible implementation of the present invention, the first node is associated with multiple feature information, and the first feature information is included in the multiple feature information.
[0836] In one possible implementation of the present invention, the processing module 2402 is used for:
[0837] A first prediction is determined based on a first observation and a first indication from a fourth node, wherein the first indication indicates first feature information of a third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
[0838] In one possible implementation of the present invention, the first node is trained using a first training dataset, which includes multiple training data sets. Each training data set includes a first data portion and a second data portion, with the second data portion indicating the feature information of the first data portion.
[0839] In one possible implementation of the present invention, the first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
[0840] In one possible implementation of the present invention, the processing module 2402 is used for:
[0841] For each of the multiple child nodes, a second prediction for the operation of the second node is determined based on a first observation, and a weight corresponding to the second prediction is determined based on a second indication from a fourth node, wherein the second indication indicates first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
[0842] For the first child node among multiple child nodes, the weighted sum of the second predictions of the multiple child nodes is obtained as the first prediction based on the second prediction of each child node among the multiple child nodes and the weight corresponding to the second prediction.
[0843] In one possible implementation of the present invention, the processing module 2402 is used for:
[0844] For each of the multiple child nodes, a weight corresponding to the second prediction is determined based on the second indication and the feature information associated with the child node.
[0845] In one possible implementation of the invention, for each of the plurality of child nodes, the second indication further indicates the weight corresponding to the child node.
[0846] In one possible implementation of the present invention, the processing module 2402 is further configured to:
[0847] For each of the multiple child nodes, a third prediction is made based on the first observation to determine the operation for the second node;
[0848] The data processing device 2400 also includes:
[0849] The first sending module 2403 is configured to: send a third prediction to a fourth node for each of a plurality of child nodes, so that the fourth node determines a third indication indicating the first prediction based on the third prediction and feedback data about the first observation provided by the third node;
[0850] Module 2401 is used for:
[0851] For the first child node among multiple child nodes, a third indication is received, wherein the third indication indicates the first prediction sent by the fourth node.
[0852] In one possible implementation of the invention, the first child node is predefined or selected by the fourth node.
[0853] In one possible implementation of the invention, the first prediction includes a third prediction determined by a plurality of child nodes.
[0854] Processing module 2402 is used for:
[0855] For each of the multiple child nodes, a third prediction is made based on the first observation to determine the operation for the second node;
[0856] The data processing device 2400 also includes:
[0857] The second sending module 2404 is configured to: send a third prediction to a fourth node for each of the plurality of child nodes, so that the fourth node determines a fourth prediction for the operation of the second node based on the third prediction and feedback data about the first observation provided by the third node, wherein the fourth prediction corresponds to the first feature information of the third node in the environment.
[0858] In one possible implementation of the present invention, the first node is trained using a second training dataset, which includes multiple sets of training data, each set of training data corresponding to predefined feature information in multiple feature information.
[0859] In one possible implementation of the present invention, the data processing apparatus 2400 further includes:
[0860] The third sending module 2405 is used to send the first prediction to the second node to update the second node.
[0861] In one possible implementation of the present invention, the acquisition module 2401 is used for:
[0862] Receive the first observation from the third node; or
[0863] The first observation received from the fourth node's broadcast.
[0864] This data processing device can be applied to the above. Figure 11 The first node described in the method embodiment shown, and / or may be the one described above. Figure 11 The first node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 11 The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0865] Figure 25 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an EC (Electronic Processing Unit). Figure 25 As shown, the data processing device 2500 may include:
[0866] The receiving module 2501 is used to receive feedback data about the first observation from the third node, wherein the first observation is in response to the state of the operation indication environment of the second node.
[0867] In one possible implementation of the present invention, the first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
[0868] The receiving module 2501 is also used for:
[0869] The fourth node receives the third prediction from each of its multiple child nodes;
[0870] The data processing device 2500 may further include:
[0871] The first processing module 2502 is configured to: determine a final prediction of the operation for the second node based on the third prediction and feedback data from the third node regarding the first observation, wherein the final prediction corresponds to the first feature information of the third node in the environment.
[0872] In one possible implementation of the present invention, the first processing module 2502 is used for:
[0873] Determine the similarity between the feedback data and the third prediction for each of the multiple child nodes;
[0874] The final prediction is determined from each of the multiple child nodes based on the third prediction and the similarity between the feedback data and the third prediction.
[0875] In one possible implementation of the present invention, the first processing module 2502 is used for:
[0876] If there is at least one third prediction with a similarity greater than or equal to a preset threshold, the third prediction with the highest similarity will be determined as the final prediction.
[0877] In the absence of a third prediction with a similarity greater than or equal to a preset threshold, the weight of the third prediction corresponding to each of the multiple child nodes is determined based on the similarity between the feedback data and the third prediction. Based on the third prediction of each of the multiple child nodes and the weight corresponding to the third prediction, the weighted sum of the third predictions of the multiple child nodes is calculated as the final prediction.
[0878] In one possible implementation of the present invention, the final prediction is the first prediction;
[0879] The first processing module 2502 is also used for:
[0880] Based on the similarity of the third predictions of multiple child nodes, the first child node is selected from multiple child nodes;
[0881] The data processing device 2500 may further include:
[0882] The first sending module 2503 is used to send a third indication indicating the first prediction to the first child node.
[0883] In one possible implementation of the second aspect, the final prediction is the fourth prediction;
[0884] The data processing device 2500 may further include:
[0885] The second sending module 2504 is used to send the fourth prediction to the second node.
[0886] In one possible implementation of the present invention, the data processing apparatus 2500 further includes:
[0887] The second processing module 2505 is used to determine the first feature information of the third node based on the feedback data.
[0888] In one possible implementation of the present invention, the data processing apparatus further includes:
[0889] The third sending module 2506 is used to send a first indication to the first node, wherein the first indication indicates the first characteristic information of the third node.
[0890] In one possible implementation of the present invention, the first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
[0891] The data processing device 2500 may further include:
[0892] The fourth sending module 2507 is used to send a second indication to each of the multiple child nodes, wherein the second indication indicates the first characteristic information of the third node; or
[0893] The second processing module 2505 is used to: determine the weight corresponding to the feature information associated with the child node for each of the plurality of child nodes; the data processing device 2500 may further include a fifth sending module 2508, used to: send a second instruction to the child node, wherein the second instruction indicates the first feature information of the third node and the weight corresponding to the child node.
[0894] In one possible implementation of the present invention, the first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
[0895] Receiver module 2501 is used for:
[0896] Receive a third prediction from each of the multiple child nodes;
[0897] The second processing module 2505 is used for:
[0898] For each of the multiple child nodes, the weight of the third prediction corresponding to the child node is determined based on the first feature information of the third node and the feature information associated with the child node.
[0899] Based on the third prediction of each of the multiple child nodes and the weight corresponding to the third prediction, the weighted sum of the third predictions of the multiple child nodes is calculated as the final prediction, where the final prediction is the fourth prediction.
[0900] The data processing device 2500 also includes a sixth sending module 2509, used to send a fourth prediction to the second node.
[0901] In one possible implementation of the present invention, the second processing module 2505 is used for:
[0902] Based on the feedback data and the preset classification algorithm, the first feature information of the third node is determined.
[0903] In one possible implementation of the present invention, the receiving module 2501 is further configured to:
[0904] Receive the first observation from the third node;
[0905] The data processing device 2500 also includes:
[0906] Broadcast module 2510 is used to broadcast the first observation to each of multiple child nodes.
[0907] This data processing device can be applied to the above. Figure 11 The fourth node described in the method embodiments shown, and / or may be the one described above. Figure 11 The fourth node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 11 The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0908] Figure 26 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. This data processing apparatus may be an RL client. For example... Figure 26 As shown, the data processing device 2600 may include:
[0909] The sending module 2601 is configured to: send feedback data about the first observation to the fourth node so that the fourth node determines the first feature information of the third node or the final prediction of the operation performed by the second node on the environment, wherein the first observation indicates the state of the environment in response to the operation of the second node; the final prediction corresponds to the first feature information, and the final prediction is either the first prediction determined by the first node or the fourth prediction determined by the fourth node.
[0910] In one possible implementation of the present invention, the sending module 2601 is used for:
[0911] Send the first observation to the first or fourth node.
[0912] This data processing device can be applied to the above. Figure 11 The third node described in the method embodiments shown, and / or may be the one described above. Figure 11 The third node described in the illustrated method embodiment. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 11 The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0913] Figure 27 This is a schematic diagram of yet another data processing apparatus according to one or more embodiments of the present invention. The data processing apparatus may be an RL algorithm. For example... Figure 27 As shown, the data processing device 2700 may include:
[0914] The receiving module 2701 is configured to receive a final prediction of the operation performed by the second node on the environment, wherein the final prediction corresponds to the first feature information of the third node in the environment; the final prediction is based on a first observation and feedback data about the first observation, the first observation responding to the operation of the second node indicating the state of the environment, and the final prediction is either a first prediction determined by the first node or a fourth prediction determined by the fourth node.
[0915] Update module 2702 is used to update based on the final prediction.
[0916] This data processing device can be applied to the above. Figure 11 The second node described in the method embodiment shown, and / or may be the one described above. Figure 11 The second node is described in the method embodiment shown. Those skilled in the art should understand that, in conjunction with the relevant descriptions of the above modules in the embodiments of the present invention, it can be understood as... Figure 11 The embodiments of the present invention shown herein are related to the description of the data processing apparatus method.
[0917] One embodiment of the present invention provides a first node including processing circuitry for performing any of the above-described data processing methods. It should be understood that the first node can perform the steps executed by the first node in the above method embodiments, which will not be elaborated further here.
[0918] One embodiment of the present invention provides a second node including processing circuitry for performing any of the above-described data processing methods. It should be understood that the second node can perform the steps executed by the second node in the above method embodiments, which will not be elaborated further here.
[0919] One embodiment of the present invention provides a third node, including processing circuitry for performing any of the above-described data processing methods. It should be understood that the third node can perform the steps executed by the third node in the above method embodiments, which will not be elaborated further here.
[0920] One embodiment of the present invention provides a fourth node, including processing circuitry for performing any of the above-described data processing methods. It should be understood that the fourth node can perform the steps executed by the fourth node in the above method embodiments, which will not be elaborated further here.
[0921] One embodiment of the present invention provides a fifth node, including processing circuitry for performing any of the above-described data processing methods. It should be understood that the fifth node can perform the steps executed by the fifth node in the above method embodiments, which will not be elaborated further here.
[0922] One embodiment of the present invention provides a sixth node, including processing circuitry for performing any of the above-described data processing methods. It should be understood that the sixth node can perform the steps executed by the sixth node in the above method embodiments, which will not be elaborated further here.
[0923] One embodiment of the present invention provides a data processing system, including a first node, a fifth node, and a sixth node. The first node is used to execute the steps performed by the first node in any of the above-described data processing methods, the fifth node is used to execute the steps performed by the fifth node in any of the above-described data processing methods, and the sixth node is used to execute the steps performed by the sixth node in any of the above-described data processing methods.
[0924] One embodiment of the present invention provides a data processing system, including a first node, a second node, a third node, a fourth node, and a fifth node. The first node is used to execute steps performed by the first node in any of the aforementioned data processing methods; the second node is used to execute steps performed by the second node in any of the aforementioned data processing methods; the third node is used to execute steps performed by the third node in any of the aforementioned data processing methods; the fourth node is used to execute steps performed by the fourth node in any of the aforementioned data processing methods; and the fifth node is used to execute steps performed by the fifth node in any of the aforementioned data processing methods.
[0925] One embodiment of the present invention provides a data processing system, including a first node, a second node, a third node, and a fourth node. The first node is used to execute the steps performed by the first node in any of the above-described data processing methods; the second node is used to execute the steps performed by the second node in any of the above-described data processing methods; the third node is used to execute the steps performed by the third node in any of the above-described data processing methods; and the fourth node is used to execute the steps performed by the fourth node in any of the above-described data processing methods.
[0926] One embodiment of the present invention provides a computer-readable medium storing computer-executable instructions that, when executed by a processor, cause the processor to perform any of the above-described data processing methods.
[0927] One embodiment of the present invention provides a computer program product including computer-executable instructions, which, when executed by a processor, cause the processor to perform any of the above-described data processing methods.
[0928] Although the present invention describes methods and processes by steps performed in a certain order, one or more steps in the methods and processes may be omitted or modified as appropriate. Where appropriate, one or more steps may be performed in an order other than that described.
[0929] It should be noted that the expression "at least one of A or B" used in this document is interchangeable with the expression "A and / or B". It refers to a list from which you can choose either A or B, or A and B. Similarly, the expression "at least one of A, B, or C" used in this document is interchangeable with "A and / or B and / or C" or "A, B, and / or C". It refers to a list from which you can choose: A or B or C, or A and B, or A and C, or B and C, or all of A, B, and C. The same principle applies to longer lists with the same format.
[0930] Although the invention has been described at least partially in terms of method, those skilled in the art will understand that the invention is also directed to various components for performing at least some aspects and features of the method, whether by hardware components, software, or any combination thereof. Accordingly, the technical solutions of the invention can be embodied in the form of a software product. Suitable software products can be stored in pre-recorded storage devices or other similar non-volatile or non-transitory computer-readable media, including DVDs, CD-ROMs, USB flash drives, removable hard drives, or other storage media. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, server, or network device) to perform examples of the methods disclosed herein. Machine-executable instructions can be in the form of sequences of code, configuration information, or other data that, when executed, cause a machine (e.g., a processor or other processing device) to perform the steps in the methods provided in the examples of the invention.
[0931] The invention may be embodied in other specific forms without departing from the subject matter of the claims. The exemplary embodiments described are merely illustrative in all respects and not restrictive. Features selected from one or more of the foregoing embodiments may be combined to create alternative embodiments not explicitly described, and features suitable for such combinations will be understood within the scope of the invention.
[0932] All values and sub-ranges within the scope of this disclosure are also disclosed. Furthermore, while the systems, devices, and processes disclosed and illustrated herein may include a specific number of elements / components, these systems, devices, and components may be modified to include more or fewer of such elements / components. For example, although any disclosed element / component may be a single quantity, embodiments disclosed herein may be modified to include multiple such elements / components. The subject matter described herein is intended to cover and include all suitable technical changes.
[0933] Although embodiments have been described above with reference to the accompanying drawings, those skilled in the art will understand that variations and modifications can be made without departing from the scope defined by the appended claims.
Claims
1. A data processing method, comprising: The first node acquires a first observation, wherein the first observation is in response to the operation of the second node indicating the state of the environment; The first node obtains a first prediction of the operation for the second node based on the first observation, wherein the first prediction is related to a first feature information of a third node in the environment, and the first prediction is used to update the second node.
2. The method according to claim 1, characterized in that, The first node is associated with multiple feature information, and the first feature information is included in the multiple feature information.
3. The method according to claim 2, characterized in that, The first prediction obtained by the first node based on the first observation for the operation against the second node includes: The first node determines the first prediction based on the first observation and a first indication from the fourth node, wherein the first indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
4. The method according to claim 3, characterized in that, The first node is trained using a first training dataset, which includes multiple training data sets. Each training data set includes a first data part and a second data part, where the second data part indicates the feature information of the first data part.
5. The method according to claim 2, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
6. The method according to claim 5, characterized in that, The first prediction obtained by the first node based on the first observation for the operation against the second node includes: For each of the plurality of child nodes, the child node determines a second prediction of the operation for the second node based on the first observation, and the child node determines a weight corresponding to the second prediction based on a second indication from a fourth node, wherein the second indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node. The first child node among the plurality of child nodes obtains a weighted sum of the second predictions of the plurality of child nodes as the first prediction, based on the second prediction of each of the plurality of child nodes and the weight corresponding to the second prediction.
7. The method according to claim 6, characterized in that, The child node determines the weight corresponding to the second prediction by including: The child node determines the weight corresponding to the second prediction based on the second indication and the feature information associated with the child node.
8. The method according to claim 6, characterized in that, For each of the plurality of child nodes, the second indication further indicates the weight corresponding to the child node.
9. The method according to claim 5, further comprising: For each of the plurality of child nodes, the child node determines a third prediction for the operation of the second node based on the first observation, and the child node sends the third prediction to a fourth node so that the fourth node determines a third indication indicative of the first prediction based on the third prediction and feedback data about the first observation provided by the third node. The first prediction obtained by the first node based on the first observation for the operation against the second node includes: The first child node among the plurality of child nodes receives the third indication, wherein the third indication indicates the first prediction sent by the fourth node.
10. The method according to any one of claims 6 to 9, characterized in that, The first child node is either predefined or selected by the fourth node.
11. The method according to claim 5, characterized in that, The first prediction includes a third prediction determined by the plurality of child nodes respectively; The first prediction obtained by the first node based on the first observation for the operation against the second node includes: For each of the plurality of child nodes, the child node determines a third prediction of the operation for the second node based on the first observation; The method further includes: For each of the plurality of child nodes, each of the plurality of child nodes sends the third prediction to the fourth node, so that the fourth node determines a fourth prediction for the operation of the second node based on the third prediction and feedback data provided by the third node regarding the first observation, wherein the fourth prediction corresponds to the first feature information of the third node in the environment.
12. The method according to any one of claims 5 to 11, characterized in that, The first node is trained using a second training dataset, which includes multiple sets of training data, each set of training data corresponding to predefined feature information among the multiple feature information.
13. The method according to any one of claims 1 to 10, characterized in that, Also includes: The first node sends the first prediction to the second node to implement the update of the second node.
14. The method according to any one of claims 1 to 13, characterized in that, The first node acquires the first observation including: The first node receives the first observation from the third node; or The first node receives the first observation broadcast by the fourth node.
15. A data processing method, comprising: The fourth node receives feedback data from the third node regarding the first observation, wherein the first observation is in response to the state of the environment indicated by the operation of the second node.
16. The method according to claim 15, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The method further includes: The fourth node receives a third prediction from each of the plurality of child nodes; The fourth node determines a final prediction of the operation for the second node based on the third prediction and the feedback data from the third node regarding the first observation, wherein the final prediction corresponds to the first feature information of the third node in the environment.
17. The method according to claim 16, characterized in that, The fourth node determines the final prediction by including: The fourth node determines the similarity between the feedback data and the third prediction from each of the plurality of child nodes; The fourth node determines the final prediction from each of the plurality of child nodes based on the third prediction and the similarity between the feedback data and the third prediction.
18. The method according to claim 17, characterized in that, The fourth node determines the final prediction from each of the plurality of child nodes based on the third prediction and the similarity between the feedback data and the third prediction, including: If there is at least one third prediction with a similarity greater than or equal to a preset threshold, the fourth node will determine the third prediction with the highest similarity as the final prediction. In the absence of a third prediction with a similarity greater than or equal to the preset threshold, the fourth node determines the weight of the third prediction corresponding to each of the plurality of child nodes based on the similarity between the feedback data and the third prediction, and the fourth node calculates the weighted sum of the third predictions of the plurality of child nodes as the final prediction based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction.
19. The method according to any one of claims 16 to 18, characterized in that, The final prediction is the first prediction; The method further includes: The fourth node selects a first child node from the plurality of child nodes based on the similarity of the third prediction of the plurality of child nodes; The fourth node sends a third indication to the first child node, indicating the first prediction.
20. The method according to any one of claims 16 to 18, characterized in that, The final prediction is the fourth prediction; The method further includes: The fourth node sends the fourth prediction to the second node.
21. The method of claim 15, further comprising: The fourth node determines the first feature information of the third node based on the feedback data.
22. The method of claim 21, further comprising: The fourth node sends a first indication to the first node, wherein the first indication indicates the first feature information of the third node.
23. The method according to claim 21, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The method further includes: The fourth node sends a second indication to each of the plurality of child nodes, wherein the second indication indicates the first feature information of the third node; or For each of the plurality of child nodes, the fourth node determines a weight corresponding to the feature information associated with the child node, and the fourth node sends a second indication to the child node, wherein the second indication indicates the first feature information of the third node and the weight corresponding to the child node.
24. The method according to claim 21, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The method further includes: The fourth node receives a third prediction from each of the plurality of child nodes; For each of the plurality of child nodes, the fourth node determines the weight of the third prediction corresponding to the child node based on the first feature information of the third node and the feature information associated with the child node; The fourth node calculates a weighted sum of the third predictions of the plurality of child nodes as the final prediction based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction, wherein the final prediction is the fourth prediction; The fourth node sends the fourth prediction to the second node.
25. The method according to any one of claims 21 to 24, characterized in that, The fourth node determines the first feature information of the third node based on the feedback data, including: The fourth node determines the first feature information of the third node based on the feedback data and the preset classification algorithm.
26. The method according to any one of claims 15 to 25, characterized in that, Also includes: The fourth node receives the first observation from the third node; The fourth node broadcasts the first observation to each of the plurality of child nodes.
27. A data processing method, comprising: The third node sends feedback data about the first observation to the fourth node, so that the fourth node determines the first feature information of the third node or the final prediction of the operation performed by the second node on the environment, wherein the first observation is in response to the operation of the second node indicating the state of the environment; the final prediction corresponds to the first feature information, and the final prediction is either the first prediction determined by the first node or the fourth prediction determined by the fourth node.
28. The method of claim 27, further comprising: The third node sends the first observation to the first node or the fourth node.
29. A data processing method, comprising: The second node receives a final prediction of the operation performed by the second node on the environment, wherein the final prediction corresponds to first feature information of a third node in the environment; the final prediction is based on a first observation and feedback data about the first observation, the first observation being in response to the operation of the second node indicating the state of the environment, and the final prediction is either a first prediction determined by the first node or a fourth prediction determined by the fourth node. The second node is updated based on the final prediction.
30. A data processing method, comprising: The fifth node sends a first request to the sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information; The fifth node receives a first response from the sixth node, wherein the first response is determined by the sixth node based on the first request and instructs the sixth node to support the ability to train the first node; The fifth node configures the first node and the sixth node to perform the training based on the first response.
31. The method according to claim 30, characterized in that, The fifth node, based on the first response, configures the first node and the sixth node to perform the training, including: The fifth node determines the first configuration information and the second configuration information of the first node based on the first response, wherein the first configuration information indicates the type of the first node, and the second configuration information indicates the parameters used to train the first node. The fifth node configures the first node based on the first configuration information and the second configuration information; The fifth node notifies the sixth node of the configuration of the first node.
32. The method according to claim 31, characterized in that, The fifth node determines the first configuration information of the first node, including: The fifth node determines the type of the first node based on the first response, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information; The fifth node notifies the sixth node of the configuration of the first node, including: The fifth node informs the sixth node of the type of the first node.
33. The method according to claim 32, characterized in that, The first response includes an indication of feature information supported by the sixth node, and an indication from the sixth node of the data required to train the first node; The fifth node further determines the first configuration information of the first node by including: The fifth node determines the number of candidate nodes to implement the first node based on the indication of the feature information supported by the sixth node; The fifth node determines the second configuration information of the first node, including: The fifth node determines the second configuration information based on the indication from the sixth node, which includes the data required to train the first node.
34. The method according to any one of claims 31 to 33, characterized in that, The second configuration information includes parameters of a first type and parameters of a second type, wherein the parameters of the first type remain unchanged during the training of the first node, and the parameters of the second type are updated based on the training data provided by the sixth node during the training of the first node.
35. The method according to claim 34, characterized in that, The parameters of the first type include at least one of the structure of the neural network used to train the first node and the method used to train the first node; The parameters of the second type include at least one of the initial weights and biases of the neural network.
36. The method according to any one of claims 31 to 35, characterized in that, The fifth node configures the first node based on the first configuration information and the second configuration information, including: The fifth node configures network elements for implementing the first node and the connection between the sixth node and the first node based on the first configuration information and the second configuration information.
37. The method according to any one of claims 31 to 36, characterized in that, The fifth node notifies the sixth node of the configuration of the first node, including: The fifth node sends a third message to the sixth node, wherein the third message indicates at least one path between the first node and the sixth node.
38. The method according to claim 37, characterized in that, The fifth node's notification of the first node's configuration to the sixth node also includes: The fifth node sends a fourth message to the sixth node, wherein the fourth message indicates the training dataset of the first node.
39. The method according to claim 38, characterized in that, The fourth piece of information also indicates the format of the data used to train the first node.
40. The method according to any one of claims 30 to 39, characterized in that, The first information includes at least one indication of at least one feature information, and the second information includes the minimum amount of data associated with each of the at least one feature information.
41. A data processing method, comprising: The sixth node receives a first request from the fifth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information; The sixth node determines a first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node; The sixth node sends the first response to the fifth node.
42. The method according to claim 41, characterized in that, The first response includes an indication of feature information supported by the sixth node, and an indication from the sixth node of the data required to train the first node; The sixth node determines the first response based on the first request, including: The sixth node determines the indication of the feature information supported by the sixth node based on the first information; The sixth node determines, based on the second information, the indication that includes the data required to train the first node.
43. The method according to claim 41 or 42, further comprising: The sixth node receives third information from the fifth node, wherein the third information indicates at least one path between the first node and the sixth node.
44. The method of claim 43, further comprising: The sixth node receives a training request from the first node, wherein the training request indicates the feature information of the first node; The sixth node determines the training data for the first node based on the type of the first node and the training request. The sixth node sends the training data to the first node based on the third information, so that the first node performs the training based on the training data.
45. The method of claim 43, further comprising: The sixth node receives fourth information from the fifth node, wherein the fourth information indicates the training dataset of the first node.
46. The method of claim 45, further comprising: When the first node is determined to be of the first type and includes an integrated node associated with multiple feature information, the sixth node determines the first training dataset based on the fourth information, and the sixth node sends the training data of the first training dataset to the first node based on the third information, so that the first node performs the training based on the training data of the first training dataset. When it is determined that the first node is of the second type and includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information, the sixth node determines the second training dataset based on the fourth information, and the sixth node sends the training data of the second training dataset to each of the multiple child nodes based on the third information, so that the first node performs the training based on the training data of the second training dataset. The first training dataset includes multiple training data sets, each training data set including a first data part and a second data part, wherein the second data part indicates the feature information of the first data part; The second training dataset includes multiple sets of training data, each set of training data being associated with predefined feature information from the multiple feature information.
47. The method according to claim 46, characterized in that, The type of the first node is either predefined or notified by the sixth node.
48. A data processing method, comprising: The first node receives training data from the sixth node, wherein the training data corresponds to the feature information of the first node; The first node performs training based on the training data from the sixth node.
49. The method of claim 48, further comprising: The first node sends a training request to the sixth node, wherein the training request indicates the feature information of the first node.
50. The method according to claim 48, characterized in that, The first node is of the first type and includes an integrated node associated with multiple feature information; The first node receives the training data from the sixth node, including: The first node receives training data from the sixth node in the first training dataset, wherein the first training dataset includes multiple training data, each training data includes a first data part and a second data part, and the second data part indicates the feature information of the first data part.
51. The method according to claim 48, characterized in that, The first node is of the second type and includes multiple child nodes, each of the multiple child nodes being associated with one of the multiple feature information; The first node receives the training data from the sixth node, including: The first node receives training data from the sixth node in the second training dataset, wherein the second training dataset includes multiple sets of training data, and each set of training data is associated with predefined feature information in the multiple feature information.
52. A data processing method, comprising: The fifth node obtains a fourth indication from the third node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node; The fifth node is configured based on the fourth instruction to execute the target task initiated by the third node.
53. The method according to claim 52, characterized in that, The fifth node obtains the fourth instruction from the third node including: The fifth node sends a second request to the third node, wherein the second request is used to request the type of the feedback data; The fifth node receives the fourth instruction from the third node.
54. The method according to claim 53, characterized in that, The second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein, The fourth instruction also indicates at least one of the items requested by the fifth node.
55. The method according to any one of claims 52 to 54, further comprising: The fifth node receives a third request from the third node, wherein the third request indicates description information corresponding to the target task.
56. The method according to claim 55, characterized in that, The fifth node performs the configuration based on the fourth instruction to execute the target task initiated by the third node, including: The fifth node determines the first node and the second node to execute the target task based on the description information corresponding to the target task; The fifth node, based on the type of the feedback data and the first node, determines the operations to be performed by the first node, the second node, and the fourth node to execute the target task, wherein the fourth node is used to provide predictions for the observations provided by the third node; The fifth node configures network elements and connections for implementing the operations performed by the first node, the second node, and the fourth node, based on the operations performed by the first node, the second node, and the fourth node. The fifth node, based on the network elements and the connections used to implement the operations performed by the first node, the second node, and the fourth node, notifies the third node of the configuration to perform the target task.
57. The method according to claim 56, characterized in that, The fifth node determines the first node to execute the target task based on the description information corresponding to the target task, including: The fifth node determines at least one candidate node that has performed a historical task as the first node based on the description information corresponding to the target task, wherein the similarity between the historical task and the target task is above a preset threshold.
58. The method according to claim 56, characterized in that, The fifth node determines the first node to execute the target task based on the description information corresponding to the target task, including: The fifth node determines the type of the first node based on the description information corresponding to the target task, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information; When the first node is determined to be of the second type, the fifth node determines all available candidate nodes as child nodes of the first node.
59. The method according to any one of claims 56 to 58, characterized in that, The fifth node, based on the network elements and connections used to implement the operations performed by the first, second, and fourth nodes, notifies the third node of the configuration to perform the target task, including: The fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the first node and the connection between the third node and the first node; or The fifth node sends a third configuration notification to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the fourth node and the connection between the third node and the fourth node.
60. The method of claim 59, further comprising: The fifth node sends a first configuration notification to the first node, wherein the first configuration notification indicates the operation to be performed by the first node; The fifth node sends a second configuration notification to the second node, wherein the second configuration notification indicates the operation to be performed by the second node; The fifth node sends a fourth configuration notification to the fourth node, wherein the fourth configuration notification indicates the operation to be performed by the fourth node.
61. The method according to any one of claims 52 to 60, characterized in that, The feedback data includes a first feedback type and a second feedback type. The first feedback type indicates the third node's evaluation of the observation, and the second feedback type indicates the third node's feature information.
62. A data processing method, comprising: The first node receives a first configuration notification from the fifth node, wherein the first configuration notification indicates an operation to be performed by the first node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The first node is configured based on the first configuration notification.
63. A data processing method, comprising: The second node receives a second configuration notification from the fifth node, wherein the second configuration notification indicates an operation to be performed by the second node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The second node is configured based on the second configuration notification.
64. A data processing method, comprising: The third node sends a fourth instruction to the fifth node so that the fifth node can be configured based on the fourth instruction to execute the target task initiated by the third node, wherein the fourth instruction indicates the type of feedback data provided by the third node.
65. The method of claim 64, further comprising: The third node receives a second request from the fifth node, wherein the second request is used to request the type of the feedback data.
66. The method according to claim 64 or 65, characterized in that, The second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein, The fourth instruction also indicates at least one of the items requested by the fifth node.
67. The method according to any one of claims 64 to 66, further comprising: The third node sends a third request to the fifth node, wherein the third request indicates description information corresponding to the target task.
68. The method according to any one of claims 64 to 67, further comprising: The third node receives a third configuration notification from the fifth node, wherein the third configuration notification indicates a network element for implementing an operation performed by the first node and a connection between the third node and the first node; or the third configuration notification indicates a network element for implementing an operation performed by the fourth node and a connection between the third node and the fourth node; The third node is configured based on the third configuration notification.
69. A data processing method, comprising: The fourth node receives a fourth configuration notification from the fifth node, wherein the fourth configuration notification indicates an operation to be performed by the fourth node to execute a target task initiated by the third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The fourth node is configured based on the fourth configuration notification.
70. A data processing method, comprising: The first node acquires a first observation, wherein the first observation is in response to the operation of the second node indicating the state of the environment; The third node sends feedback data about the first observation to the fourth node; The first node obtains a first prediction for the operation of the second node based on the first observation, wherein the first prediction is related to the first feature information of the third node in the environment, the first prediction is a final prediction or a third prediction, the third prediction is generated by the fourth node as the final prediction, the final prediction corresponds to the first feature information, and the final prediction is based on the first observation and the feedback data. The fourth node receives the feedback data from the third node; The second node receives the final prediction of the operation performed by the second node on the environment from the first node or the fourth node; The second node is updated based on the final prediction.
71. A data processing method, comprising: The fifth node sends a first request to the sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information; The sixth node receives the first request from the fifth node; The sixth node determines a first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node; The sixth node sends the first response to the fifth node; The fifth node receives the first response from the sixth node; The fifth node configures the first node and the sixth node to perform the training based on the first response. The sixth node sends training data to the first node, wherein the training data corresponds to the feature information of the first node; The first node receives the training data from the sixth node; The first node performs the training based on the training data from the sixth node.
72. A data processing method, comprising: The third node sends a fourth indication to the fifth node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node; The fifth node obtains the fourth instruction from the third node; The fifth node is configured based on the fourth instruction to execute the target task initiated by the third node.
73. A data processing system, comprising: The first node is used to execute the data processing method according to any one of claims 1 to 14; The second node is used to execute the data processing method according to claim 29; The third node is used to execute the data processing method according to claim 27 or 28; The fourth node is used to perform the data processing method according to any one of claims 15 to 26.
74. A data processing system, comprising: The first node is used to execute the data processing method according to any one of claims 48 to 51; The fifth node is used to execute the data processing method according to any one of claims 30 to 40; The sixth node is used to perform the data processing method according to any one of claims 41 to 47.
75. A data processing system, comprising: The first node is used to execute the data processing method according to claim 62; The second node is used to execute the data processing method according to claim 63; The third node is used to execute the data processing method according to any one of claims 64 to 68. The fourth node is used to execute the data processing method according to claim 69; The fifth node is used to execute the data processing method according to any one of claims 52 to 61.
76. A data processing apparatus, comprising: The acquisition module is used to acquire a first observation, wherein the first observation is in response to the operation indication environment state of the second node; A processing module is configured to obtain a first prediction of the operation for the second node based on the first observation, wherein the first prediction is related to first feature information of a third node in the environment, and the first prediction is used to update the second node.
77. The data processing apparatus according to claim 76, characterized in that, The first node is associated with multiple feature information, and the first feature information is included in the multiple feature information.
78. The data processing apparatus according to claim 77, characterized in that, The processing module is used for: The first prediction is determined based on the first observation and a first indication from the fourth node, wherein the first indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node.
79. The data processing apparatus according to claim 78, characterized in that, The first node is trained using a first training dataset, which includes multiple training data sets. Each training data set includes a first data part and a second data part, where the second data part indicates the feature information of the first data part.
80. The data processing apparatus according to claim 77, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information.
81. The data processing apparatus according to claim 80, characterized in that, The processing module is used for: For each of the plurality of child nodes, a second prediction for the operation of the second node is determined based on the first observation, and a weight corresponding to the second prediction is determined based on a second indication from a fourth node, wherein the second indication indicates the first feature information of the third node and is determined by the fourth node based on feedback data about the first observation provided by the third node. For the first child node among the plurality of child nodes, based on the second prediction of each child node among the plurality of child nodes and the weight corresponding to the second prediction, the weighted sum of the second predictions of the plurality of child nodes is obtained as the first prediction.
82. The data processing apparatus according to claim 81, characterized in that, The processing module is used for: For each of the plurality of child nodes, the weight corresponding to the second prediction is determined based on the second indication and the feature information associated with the child node.
83. The data processing apparatus according to claim 81, characterized in that, For each of the plurality of child nodes, the second indication further indicates the weight corresponding to the child node.
84. The data processing apparatus according to claim 80, characterized in that, The processing module is also used for: For each of the plurality of child nodes, a third prediction for the operation on the second node is determined based on the first observation; The data processing device further includes: The first sending module is configured to: send the third prediction to the fourth node for each of the plurality of child nodes, so that the fourth node determines a third indication indicating the first prediction based on the third prediction and feedback data about the first observation provided by the third node; The acquisition module is used for: For the first child node among the plurality of child nodes, the third indication is received, wherein the third indication indicates the first prediction sent by the fourth node.
85. The data processing apparatus according to any one of claims 81 to 84, characterized in that, The first child node is either predefined or selected by the fourth node.
86. The data processing apparatus according to claim 80, characterized in that, The first prediction includes a third prediction determined by the plurality of child nodes respectively; The processing module is used for: For each of the plurality of child nodes, a third prediction for the operation on the second node is determined based on the first observation; The data processing device further includes: The second sending module is configured to: send the third prediction to the fourth node for each of the plurality of child nodes, so that the fourth node determines a fourth prediction for the operation of the second node based on the third prediction and feedback data about the first observation provided by the third node, wherein the fourth prediction corresponds to the first feature information of the third node in the environment.
87. The data processing apparatus according to any one of claims 80 to 86, characterized in that, The first node is trained using a second training dataset, which includes multiple sets of training data, each set of training data corresponding to predefined feature information among the multiple feature information.
88. The data processing apparatus according to any one of claims 76 to 85, characterized in that, The data processing device further includes: The third sending module is configured to: send the first prediction to the second node to implement the update of the second node.
89. The data processing apparatus according to any one of claims 76 to 88, characterized in that, The acquisition module is used for: Receive the first observation from the third node; or Receive the first observation broadcast by the fourth node.
90. A data processing apparatus, comprising: A receiving module is configured to receive feedback data from a third node regarding a first observation, wherein the first observation is in response to the state of the environment indicated by the operation of a second node.
91. The data processing apparatus according to claim 90, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The receiving module is also used for: The fourth node receives a third prediction from each of the plurality of child nodes; The data processing device further includes: A first processing module is configured to: determine a final prediction of the operation for the second node based on the third prediction and the feedback data from the third node regarding the first observation, wherein the final prediction corresponds to the first feature information of the third node in the environment.
92. The data processing apparatus according to claim 91, characterized in that, The first processing module is used for: Determine the similarity between the feedback data and the third prediction from each of the plurality of child nodes; The final prediction is determined from each of the plurality of child nodes based on the third prediction and the similarity between the feedback data and the third prediction.
93. The data processing apparatus according to claim 92, characterized in that, The first processing module is used for: If there is at least one third prediction with a similarity greater than or equal to a preset threshold, the third prediction with the highest similarity will be determined as the final prediction. In the absence of a third prediction with a similarity greater than or equal to the preset threshold, the weight of the third prediction corresponding to each of the plurality of child nodes is determined based on the similarity between the feedback data and the third prediction, and the weighted sum of the third predictions of the plurality of child nodes is calculated as the final prediction based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction.
94. The data processing apparatus according to any one of claims 91 to 93, characterized in that, The final prediction is the first prediction; The first processing module is also used for: Based on the similarity of the third prediction of the plurality of child nodes, a first child node is selected from the plurality of child nodes; The data processing device further includes: The first sending module is used to send a third indication to the first child node, indicating the first prediction.
95. The data processing apparatus according to any one of claims 91 to 93, characterized in that, The final prediction is the fourth prediction; The data processing device further includes: The second sending module is used to send the fourth prediction to the second node.
96. The data processing apparatus according to claim 90, characterized in that, The data processing device further includes: The second processing module is used to determine the first feature information of the third node based on the feedback data.
97. The data processing apparatus according to claim 96, characterized in that, The data processing device further includes: The third sending module is used to send a first indication to the first node, wherein the first indication indicates the first feature information of the third node.
98. The data processing apparatus according to claim 96, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The data processing device further includes: The fourth sending module is configured to send a second indication to each of the plurality of child nodes, wherein the second indication indicates the first feature information of the third node; or The second processing module is configured to: determine a weight corresponding to feature information associated with each of the plurality of child nodes; the data processing device further includes a fifth sending module, configured to: send a second indication to the child node, wherein the second indication indicates the first feature information of the third node and the weight corresponding to the child node.
99. The data processing apparatus according to claim 96, characterized in that, The first node includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information; The receiving module is used for: Receive a third prediction from each of the plurality of child nodes; The second processing module is used for: For each of the plurality of child nodes, based on the first feature information of the third node and the feature information associated with the child node, the weight of the third prediction corresponding to the child node is determined; Based on the third prediction of each of the plurality of child nodes and the weight corresponding to the third prediction, a weighted sum of the third predictions of the plurality of child nodes is calculated as the final prediction, wherein the final prediction is a fourth prediction; The data processing device further includes a sixth sending module, used to send the fourth prediction to the second node.
100. The data processing apparatus according to any one of claims 96 to 99, characterized in that, The second processing module is used for: Based on the feedback data and the preset classification algorithm, the first feature information of the third node is determined.
101. The data processing apparatus according to any one of claims 90 to 100, characterized in that, The receiving module is also used for: Receive the first observation from the third node; The data processing device further includes: A broadcast module is used to broadcast the first observation to each of the plurality of child nodes.
102. A data processing apparatus, comprising: A sending module is configured to: send feedback data about a first observation to a fourth node, so that the fourth node determines a first feature information of the third node or a final prediction of an operation performed by the second node on the environment, wherein the first observation responds to the operation of the second node indicating the state of the environment; the final prediction corresponds to the first feature information, and the final prediction is a first prediction determined by the first node or a fourth prediction determined by the fourth node.
103. The data processing apparatus according to claim 102, characterized in that, The sending module is also used for: Send the first observation to the first node or the fourth node.
104. A data processing apparatus, comprising: A receiving module is configured to receive a final prediction of an operation performed by the second node on the environment, wherein the final prediction corresponds to first feature information of a third node in the environment; the final prediction is based on a first observation and feedback data about the first observation, the first observation being in response to the operation of the second node indicating the state of the environment, and the final prediction is either a first prediction determined by the first node or a fourth prediction determined by the fourth node. An update module is used to update the prediction based on the final prediction.
105. A data processing apparatus, comprising: A sending module is configured to send a first request to a sixth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information; A receiving module is configured to receive a first response from the sixth node, wherein the first response is determined by the sixth node based on the first request and indicates that the sixth node supports the ability to train the first node; A processing module is configured, based on the first response, to configure the first node and the sixth node to perform the training.
106. The data processing apparatus according to claim 105, characterized in that, The processing module is used for: Based on the first response, first configuration information and second configuration information of the first node are determined, wherein the first configuration information indicates the type of the first node, and the second configuration information indicates the parameters used to train the first node; Configure the first node based on the first configuration information and the second configuration information; The sending module is used for: The configuration of the first node is notified to the sixth node.
107. The data processing apparatus according to claim 106, characterized in that, The processing module is used for: The type of the first node is determined based on the first response, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information; The sending module is used for: The sixth node is notified of the type of the first node.
108. The data processing apparatus according to claim 107, characterized in that, The first response includes an indication of feature information supported by the sixth node, and an indication from the sixth node of the data required to train the first node; The processing module is used for: Based on the indication of the feature information supported by the sixth node, the number of candidate nodes for implementing the first node is determined; The second configuration information is determined based on the indication from the sixth node, which includes the data required to train the first node.
109. The data processing apparatus according to any one of claims 106 to 108, characterized in that, The second configuration information includes parameters of a first type and parameters of a second type, wherein the parameters of the first type remain unchanged during the training of the first node, and the parameters of the second type are updated based on the training data provided by the sixth node during the training of the first node.
110. The data processing apparatus according to claim 109, characterized in that, The parameters of the first type include at least one of the structure of the neural network used to train the first node and the method used to train the first node; The parameters of the second type include at least one of the initial weights and biases of the neural network.
111. The data processing apparatus according to any one of claims 106 to 110, characterized in that, The processing module is used for: Based on the first configuration information and the second configuration information, network elements for implementing the first node and the connection between the sixth node and the first node are configured.
112. The data processing apparatus according to any one of claims 106 to 111, characterized in that, The sending module is used for: Send a third message to the sixth node, wherein the third message indicates at least one path between the first node and the sixth node.
113. The data processing apparatus according to claim 112, characterized in that, The sending module is used for: Send a fourth message to the sixth node, wherein the fourth message indicates the training dataset of the first node.
114. The data processing apparatus according to claim 113, characterized in that, The fourth piece of information also includes the format of the data used to train the first node.
115. The data processing apparatus according to any one of claims 105 to 114, characterized in that, The first information includes at least one indication of at least one feature information, and the second information includes the minimum amount of data associated with each of the at least one feature information.
116. A data processing apparatus, comprising: A receiving module is configured to receive a first request from a fifth node, wherein the first request includes first information and second information, the first information indicating feature information for training the first node, and the second information indicating data attributes associated with the feature information; A determining module is configured to determine a first response based on the first request, wherein the first response indicates that the sixth node supports the ability to train the first node; The sending module is used to send the first response to the fifth node.
117. The data processing apparatus according to claim 116, characterized in that, The first response includes an indication of feature information supported by the sixth node, and an indication from the sixth node of the data required to train the first node; The determining module is used for: The indication of the feature information supported by the sixth node is determined based on the first information; The indication for the sixth node, including the data required to train the first node, is determined based on the second information.
118. The data processing apparatus according to claim 116 or 117, characterized in that, The receiving module is also used for: The third information is received from the fifth node, wherein the third information indicates at least one path between the first node and the sixth node.
119. The data processing apparatus according to claim 118, characterized in that, The receiving module is also used for: Receive a training request from the first node, wherein the training request indicates feature information of the first node; The determining module is also used for: Based on the type of the first node and the training request, determine the training data for the first node; The sending module is also used for: The training data is sent to the first node based on the third information, so that the first node performs the training based on the training data.
120. The data processing apparatus according to claim 118, characterized in that, The receiving module is also used for: The fourth information is received from the fifth node, wherein the fourth information indicates the training dataset of the first node.
121. The data processing apparatus according to claim 120, characterized in that, The determining module is also used for: When the first node is determined to be of the first type and includes an integrated node associated with multiple feature information, a first training dataset is determined based on the third information; the sending module is further configured to: send the training data of the first training dataset to the first node based on the fourth information, so that the first node performs the training based on the training data of the first training dataset; The determining module is further configured to: when determining that the first node is of the second type and includes multiple child nodes, and each of the multiple child nodes is associated with one of the multiple feature information, determine a second training dataset based on the third information; the sending module is further configured to: send the training data of the second training dataset to each of the multiple child nodes based on the fourth information, so that the first node performs the training based on the training data of the second training dataset; The first training dataset includes multiple training data sets, each training data set including a first data part and a second data part, wherein the second data part indicates the feature information of the first data part; The second training dataset includes multiple sets of training data, each set of training data being associated with predefined feature information from the multiple feature information.
122. The data processing apparatus according to claim 121, characterized in that, The type of the first node is either predefined or notified by the sixth node.
123. A data processing apparatus, comprising: A receiving module is configured to receive training data from the sixth node, wherein the training data corresponds to the feature information of the first node; A training module is used to perform the training based on the training data from the sixth node.
124. The data processing apparatus according to claim 123, characterized in that, The data processing device further includes: A sending module is used to send a training request to the sixth node, wherein the training request indicates the feature information of the first node.
125. The data processing apparatus according to claim 123, characterized in that, The first node is of the first type and includes an integrated node associated with multiple feature information; The receiving module is used for: The sixth node receives training data from the first training dataset, wherein the first training dataset includes multiple training data sets, each training data set including a first data part and a second data part, and the second data part indicates the feature information of the first data part.
126. The data processing apparatus according to claim 123, characterized in that, The first node is of the second type and includes multiple child nodes, each of the multiple child nodes being associated with one of the multiple feature information; The receiving module is used for: The sixth node receives training data from the second training dataset, wherein the second training dataset includes multiple sets of training data, and each set of training data is associated with predefined feature information among the multiple feature information.
127. A data processing apparatus, comprising: An acquisition module is used to acquire a fourth indication from a third node, wherein the fourth indication indicates the type of feedback data about the first observation provided by the third node; The configuration module is used to configure based on the fourth instruction to execute the target task initiated by the third node.
128. The data processing apparatus according to claim 127, characterized in that, The data processing device further includes: A first sending module is configured to send a second request to the third node, wherein the second request is used to request the type of the feedback data; The acquisition module is used to receive the fourth instruction from the third node.
129. The data processing apparatus according to claim 128, characterized in that, The second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein, The fourth instruction also indicates at least one of the items requested by the fifth node.
130. The data processing apparatus according to any one of claims 127 to 129, characterized in that, The acquisition module is also used for: A third request is received from the third node, wherein the third request indicates description information corresponding to the target task.
131. The data processing apparatus according to claim 130, characterized in that, The configuration module is used for: Based on the description information corresponding to the target task, a first node and a second node are determined to execute the target task; Based on the type of the feedback data and the first node, the operations to be performed by the first node, the second node, and the fourth node are determined to execute the target task, wherein the fourth node is used to provide a prediction for the observations provided by the third node; Based on the operations performed by the first node, the second node, and the fourth node, configure network elements and connections for implementing the operations performed by the first node, the second node, and the fourth node; The data processing device further includes: The second sending module is configured to: notify the third node of the configuration for executing the target task based on the network element and the connection used to implement the operation performed by the first node, the second node, and the fourth node.
132. The data processing apparatus according to claim 131, characterized in that, The configuration module is used for: Based on the description information corresponding to the target task, at least one candidate node that has performed a historical task is determined as the first node, wherein the similarity between the historical task and the target task is above a preset threshold.
133. The data processing apparatus according to claim 131, characterized in that, The configuration module is used for: Based on the description information corresponding to the target task, the type of the first node is determined, wherein the type of the first node includes a first type and a second type, the first type of node includes an integrated node associated with multiple feature information, and the second type of node includes multiple child nodes, each of the multiple child nodes being associated with one feature information among the multiple feature information; When the first node is determined to be of the second type, all available candidate nodes are determined to be the child nodes of the first node.
134. The data processing apparatus according to any one of claims 131 to 133, characterized in that, The second sending module is used for: Send a third configuration notification to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the first node and the connection between the third node and the first node; or A third configuration notification is sent to the third node, wherein the third configuration notification indicates the network element for implementing the operation performed by the fourth node and the connection between the third node and the fourth node.
135. The data processing apparatus according to claim 134, characterized in that, The second sending module is also used for: Send a first configuration notification to the first node, wherein the first configuration notification indicates the operation to be performed by the first node; Send a second configuration notification to the second node, wherein the second configuration notification indicates the operation to be performed by the second node; A fourth configuration notification is sent to the fourth node, wherein the fourth configuration notification indicates the operation to be performed by the fourth node.
136. The data processing apparatus according to any one of claims 127 to 135, characterized in that, The feedback data includes a first feedback type and a second feedback type. The first feedback type indicates the third node's evaluation of the observation, and the second feedback type indicates the third node's feature information.
137. A data processing apparatus, comprising: A receiving module is configured to receive a first configuration notification from a fifth node, wherein the first configuration notification indicates an operation performed by the first node to execute a target task initiated by a third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The configuration module is used to perform configuration based on the first configuration notification.
138. A data processing apparatus, comprising: A receiving module is configured to receive a second configuration notification from a fifth node, wherein the second configuration notification indicates an operation performed by the second node to execute a target task initiated by a third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The configuration module is used to perform configuration based on the second configuration notification.
139. A data processing apparatus, comprising: A sending module is configured to send a fourth indication to a fifth node so that the fifth node can be configured based on the fourth indication to execute a target task initiated by the third node, wherein the fourth indication indicates the type of feedback data provided by the third node.
140. The data processing apparatus according to claim 139, characterized in that, The data processing device further includes: A first receiving module is configured to receive a second request from the fifth node, wherein the second request is used to request the type of the feedback data.
141. The data processing apparatus according to claim 139 or 140, characterized in that, The second request is used to request at least one of the following: the format of the feedback data, the frequency at which the third node provides the feedback data, or the interface required by the third node, wherein, The fourth instruction also indicates at least one of the items requested by the fifth node.
142. The data processing apparatus according to any one of claims 139 to 141, characterized in that, The sending module is also used for: A third request is sent to the fifth node, wherein the third request indicates description information corresponding to the target task.
143. The data processing apparatus according to any one of claims 139 to 142, characterized in that, The data processing device further includes: The second receiving module is configured to receive a third configuration notification from the fifth node, wherein the third configuration notification indicates a network element for implementing an operation performed by the first node and a connection between the third node and the first node; or the third configuration notification indicates a network element for implementing an operation performed by the fourth node and a connection between the third node and the fourth node; The configuration module is used to perform configuration based on the third configuration notification.
144. A data processing apparatus, comprising: A receiving module is configured to receive a fourth configuration notification from a fifth node, wherein the fourth configuration notification indicates an operation to be performed by the fourth node to execute a target task initiated by a third node, wherein the operation is determined by the fifth node based on description information corresponding to the target task and the type of feedback data provided by the third node; The configuration module is used to perform configuration based on the fourth configuration notification.
145. A first node comprising a processing circuit for performing a data processing method according to any one of claims 1 to 14, 48 to 51 and 62.
146. A second node comprising processing circuitry for performing the data processing method according to claim 29 or 63.
147. A third node comprising processing circuitry for performing the data processing method according to any one of claims 27, 28, and 64 to 68.
148. A fourth node comprising processing circuitry for performing the data processing method according to any one of claims 15 to 26 and 69.
149. A fifth node comprising processing circuitry for performing the data processing method according to any one of claims 30 to 40 and 52 to 61.
150. A sixth node comprising processing circuitry for performing the data processing method according to any one of claims 41 to 47.
151. A computer-readable medium storing computer-executable instructions, which, when executed by a processor, cause the processor to perform a data processing method according to any one of claims 1 to 72.
152. A computer program product comprising computer-executable instructions, which, when executed by a processor, cause the processor to perform the data processing method according to any one of claims 1 to 72.