Robot control model training method and device
By evaluating and adjusting the training frequency and data priorities of each node in federated learning, the problems of data inhomogeneity and quality are solved, and the training efficiency and model performance are improved, and the risk of overfitting is avoided.
Patent Information
- Application Number
- CN202510137017.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
In federated learning, due to the uneven amount of data uploaded by various enterprises or the uneven data quality, the training efficiency of the robot control model is reduced, the model performance is reduced, and the risk of overfitting is increased.
By evaluating the privacy and scarcity of the training data reported by each node on the aggregated server side, dynamically adjusting the training frequency and data priority of each node, thereby controlling the contribution and participation frequency of each node in federated learning training.
It effectively ensures the fairness of the federated learning training process, improves the data consistency of the global model, and avoids the decline in model performance, training efficiency and increased risk of overfitting.
Smart Images

Figure CN119973987A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method and device for a robot control model. Background Art
[0002] At present, when enterprises in the industrial Internet system control various robots in the production line based on artificial intelligence models, they are limited by limited data samples and cannot effectively train the robot control models. This problem is particularly serious for small and medium-sized enterprises. With the development of federated learning technology, small and medium-sized enterprises will jointly upload their training data to the federated learning aggregation server for training, which not only solves the problem of insufficient training samples, but also ensures the privacy of each enterprise's data.
[0003] However, with the operation of the federated learning training mode of the above robot control model, due to the uneven amount of data uploaded by each enterprise or the uneven quality of the uploaded data, for example, some enterprises have long been lazy in uploading their own data and only enjoy the training results issued by the federated learning aggregation server, or some enterprises upload low-quality data such as daily operation logs and maintenance records of robots every time, which is of little help to the data of federated learning training. This common problem of uneven distribution of training data in federated learning is gradually exposed, which may lead to a variety of adverse consequences, such as reduced model performance, reduced training efficiency, and increased risk of overfitting. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a training method, device, electronic device and storage medium for a robot control model.
[0005] In a first aspect, an embodiment of the present invention provides a method for training a robot control model, the method comprising:
[0006] Confirming that the first training frequency identifier carried by the target node meets the preset conditions;
[0007] Determine a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority;
[0008] Update the second private data priority corresponding to the current training round according to the robot training data reported by each node;
[0009] The model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier are sent to the target node.
[0010] Optionally, the confirming that the first training frequency identifier carried by the target node meets a preset condition specifically includes:
[0011] The current time is compared with the time in the first training frequency identifier. If the time in the first training frequency identifier is later, the preset condition is met.
[0012] Optionally, updating the second private data priority corresponding to the current training round according to the robot training data reported by each node specifically includes:
[0013] Extract the data type from the robot training data collected in the current training round;
[0014] Calculate the aggregation degree of the robot training data based on the data type and group the training data;
[0015] Different second private data priorities are assigned to different data types in each training data group according to the aggregation degree of each training data group.
[0016] Optionally, determining the second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority specifically includes:
[0017] Determining the privacy level of each training data sample in the training data, and then calculating the privacy weight of the training data;
[0018] Divide the robot training data into a private data sample set and a non-private data sample set, and calculate a first private data priority of the robot training data according to the private data sample set;
[0019] A second training frequency identifier corresponding to the target node is obtained by calculation according to the privacy weight and the first privacy data priority.
[0020] Optionally, the private data sample set includes the robot's position data, motion trajectory data, working environment data and task data.
[0021] Optionally, the non-private data sample set includes basic control parameters, device status data, public environment data and non-specific task data.
[0022] Optionally, the privacy weight of the robot training data is determined according to the amount and data type of the training data.
[0023] In a second aspect, an embodiment of the present invention provides a training device for a robot control model, the device comprising:
[0024] A training admission judgment module is used to confirm that the first training frequency identifier carried by the target node meets the preset conditions;
[0025] A training contribution determination module, configured to determine a second training frequency identifier corresponding to the target node according to a privacy weight of the robot training data reported by the target node and a first private data priority;
[0026] A data priority calculation module, used to update the second private data priority corresponding to the current training round according to the robot training data reported by each node;
[0027] A parameter delivery module is used to deliver the model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier to the target node.
[0028] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0029] one or more processors;
[0030] A memory for storing one or more programs;
[0031] When the one or more programs are executed by the one or more processors, the one or more processors execute the method as described in the first aspect.
[0032] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having executable instructions stored thereon, wherein when the executable instructions are executed by a processor, the processor executes the method as described in the first aspect.
[0033] In the training method, device, electronic device and storage medium of the robot control model provided by the embodiments of the present invention, the aggregation server will evaluate the contribution of the training data reported by each participating node to the global model training from the two dimensions of privacy and scarcity of training data in each round of federated learning training, and control the frequency of each node participating in the federated learning training according to the contribution, and provide guidance for each node to collect and report data in the next round, that is, effectively ensure the fairness of the federated learning training process to all participants, and improve the data consistency of the global model being trained, thereby avoiding the occurrence of situations such as degradation of model performance, reduced training efficiency, and increased risk of overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solution of the embodiments of the present application, the drawings required for use in the embodiments of the present application are briefly introduced below.
[0035] Figure 1 A schematic diagram of a flow chart of a training method for a robot control model provided by an embodiment of the present invention;
[0036] Figure 2 A flow chart of a training frequency calculation method provided by an embodiment of the present invention;
[0037] Figure 3 A schematic diagram of a process for determining the priority of private data provided by an embodiment of the present invention;
[0038] Figure 4 A schematic diagram of the structure of a training device for a robot control model provided by an embodiment of the present invention;
[0039] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0041] Similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0042] At present, when enterprises in the industrial Internet system control various robots in the production line based on artificial intelligence models, they are limited by limited data samples and cannot effectively train the robot control models. This problem is particularly serious for small and medium-sized enterprises. With the development of federated learning technology, small and medium-sized enterprises will jointly upload their training data to the federated learning aggregation server for training, which not only solves the problem of insufficient training samples, but also ensures the privacy of each enterprise's data.
[0043] However, with the operation of the federated learning training mode of the above robot control model, due to the uneven amount of data uploaded by each enterprise or the uneven quality of the uploaded data, for example, some enterprises have long been lazy in uploading their own data and only enjoy the training results issued by the federated learning aggregation server, or some enterprises upload low-quality data such as daily operation logs and maintenance records of robots every time, which is of little help to the data of federated learning training. This common problem of uneven distribution of training data in federated learning is gradually exposed, which may lead to a variety of adverse consequences, such as reduced model performance, reduced training efficiency, and increased risk of overfitting.
[0044] Based on this, an embodiment of the present invention provides a training method for a robot control model. Figure 1 A schematic flow chart of a training method for a robot control model provided in an embodiment of the present invention is shown.
[0045] Step S110: confirming that the first training frequency identifier carried by the target node meets a preset condition.
[0046] The execution entities of each step in the embodiment of the present invention are all aggregation servers in the federated learning platform. The aggregation server is mainly responsible for collecting, aggregating and distributing model updates in federated learning. Specifically, it allows each participating node (in the embodiment of the present invention, each industrial Internet company or other registered entity) to regularly share model updates trained locally on its dataset while retaining its dataset locally. These updates are then aggregated by the aggregation server to generate a global model. The aggregation server in the federated learning platform has the characteristics of protecting data privacy, improving training efficiency, supporting scalability and fault tolerance, etc.
[0047] In an embodiment of the present invention, the aggregation server performs federated learning training according to preset rounds. In a round of training, the aggregation server performs centralized training based on the training data reported by each node, and then sends the updated global model to each node, which is considered to be a round of training completed. Specifically in this step, since different nodes have different data contributions to federated learning training, the aggregation server will assign different training frequencies to different nodes. The data contribution in the embodiment of the present invention refers to the quantity and quality of robot training data participating in federated learning provided by the node. When the data contribution of a node is too low, the training frequency assigned to it by the aggregation server will also be relatively low, which means that the node cannot participate in the next few rounds of federated learning training and cannot obtain the parameter update results of the global model.
[0048] In the embodiment of the present invention, the training frequency assigned to the node is calculated by its data contribution, which is embodied in the form of a training frequency identifier. The first training frequency identifier in this step is determined and notified by the aggregation server when the target node last participated in the federated learning training round. That is to say, in this step, before each target node participates in the current round of federated learning training, the aggregation server will check the first training frequency identifier carried by the target node, and the target node will be allowed to participate in this training only if it meets the conditions.
[0049] Specifically, confirming that the first training frequency identifier carried by the target node meets the preset conditions can be achieved in the following way: comparing the current moment with the moment in the first training frequency identifier, if the moment in the first training frequency identifier is later, it meets the preset access conditions. That is, when the target node wants to participate in federated learning training, the current moment needs to have reached the target moment set for it by the aggregation server. This step can prevent low-contribution nodes that do not meet the access conditions from participating in the training rounds of federated learning in advance.
[0050] Step S120: determining a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority.
[0051] Specifically, as attached Figure 2 As shown, the training frequency calculation method in step S120 can be implemented through the following steps S121 to S123.
[0052] Step S121, determining the privacy level of each training data sample in the training data, and then calculating the privacy weight of the training data;
[0053] Step S122, dividing the robot training data into a private data sample set and a non-private data sample set, and calculating a first private data priority of the robot training data according to the private data sample set;
[0054] Step S123: Calculate a second training frequency identifier corresponding to the target node according to the privacy weight and the first privacy data priority.
[0055] Specifically, in the context of federated learning in the embodiments of the present invention, robot control parameter data can be divided into private data and non-private data according to the possibility of privacy leakage, the importance and scarcity of the data, etc. For the data of robot control parameters, the following types of data are generally regarded as sensitive data: location data, i.e. the specific geographical location or spatial coordinates of the robot's work; motion trajectory data, i.e. the path, speed and acceleration of the robot's movement and other motion parameters; working environment data, i.e. the environmental information of the robot, such as temperature, humidity, light, etc.; task-related data, i.e. the specific task type, task parameters and task results performed by the robot. The following types of data are generally regarded as non-private data: basic control parameters, such as basic control instructions such as the robot's movement speed, acceleration, steering angle, etc.; device status data, such as the robot's power status, network connection status, sensor status, etc.; public environment data, such as environmental parameters such as temperature and humidity in public areas; non-specific task data, such as the robot's daily operation log and maintenance records.
[0056] It is worth noting that the above-mentioned private data and non-private data are only listed for illustration. The specific data included in the actual application is determined according to the corresponding scenario of the robot application.
[0057] Furthermore, in an embodiment of the present invention, the training data also needs to be distinguished by different privacy levels according to the requirements of the model training in the federated learning training. For example, position data and motion trajectory data have the highest privacy level, which can be set to 1.0; non-specific task data has the lowest privacy level, which can be set to 0.1. The above is only an example. In actual applications, how the aggregation server sets the privacy level according to the type of training data is determined according to the corresponding scenario of the robot application. For a specific model, once the pre-setting is completed before the federated learning training, the privacy level of each relevant training data type will no longer change.
[0058] In an embodiment of the present invention, the training data uploaded by the target node in a federated learning training round includes multiple training data samples, each of which has a corresponding privacy level. Combining the privacy levels of these training data samples, the privacy weight of the entire training data uploaded by the target node in the training round can be determined. The privacy weight of the robot training data is determined based on the amount and data type of the training data, and can be specifically calculated according to the following formula:
[0059]
[0060] Where W represents the privacy weight of the entire training data, n represents the number of training data samples, and w i represents the privacy weight of the i-th data sample, s i Represents some importance indicator of the i-th data sample, such as sample size, information amount, etc.
[0061] In addition, for the entire training data, in addition to distinguishing the privacy level of each sample, it is also necessary to determine the priority of private data. The priority of private data in the embodiment of the present invention is set to evaluate the scarcity of training data. The priority of different training data is dynamically determined by the aggregation server based on the training data collected from all nodes. Generally speaking, non-private data is easier for the aggregation server to obtain and is generally uniformly set to the lowest priority. Only private data has different priorities.
[0062] Therefore, in the embodiment of the present invention, after the robot receives the robot training data uploaded by the target node, it first divides the data into a private data sample set and a non-private data sample set, and extracts the private data sample set. Therefore, it is necessary to calculate the first private data priority of the robot training data based on the private data sample set. The first private data priority in the embodiment of the present invention is used to evaluate the priority of the robot training data uploaded by the target node in this round of federated learning training. The specific calculation method can consider the following formula:
[0063]
[0064] Where P represents the first private data priority result, n represents the number of samples in the private data sample set, and p i represents the priority data of the i-th private data sample, t i represents the weight of the i-th private data sample, which can be determined based on factors such as the importance of the sample, the amount of information, and the sample size. Among them, it can be understood that the specific priority p of each sample i The value is calculated and determined by the aggregation server during the previous round of federated learning training. Accordingly, during this round of federated learning training, the aggregation server will calculate and determine the priority values of training samples of different data types in the next round.
[0065] After determining the privacy weight and the first private data priority of the robot training data reported by the target node, the embodiment of the present invention calculates the second training frequency identifier corresponding to the target node based on the privacy weight and the first private data priority, that is, determines the time when the target node can participate in the federated learning training next time based on the above two aspects of the target node's contribution to this round of federated learning training.
[0066] Specifically, in order to determine the second training frequency identifier of the target node according to the contribution of the privacy weight and the privacy data priority, and make the second training frequency identifier smaller when the privacy weight and the privacy data priority are larger, and the time for the next federated learning training is shorter, the following formula can be used:
[0067] f=f max -(α·w+β·p)
[0068] Where w represents the privacy weight, p represents the first private data priority, f represents the second training frequency identifier, α and β are weight coefficients for adjusting the privacy weight and the private data priority, which can be adjusted according to specific needs. max is the maximum possible value of the second training frequency identifier. Among them, α and β can be used to balance the influence of privacy weight and private data priority. For example, if we think that privacy weight is more important, we can increase the value of α; if we think that private data priority is more important, we can increase the value of β. In order to ensure that the value of f is always within a reasonable range (for example, non-negative value), some boundary conditions can be added. For example:
[0069] f=max(0,f max -(α·w+β·p))
[0070] In addition, if the values of privacy weight and privacy data priority are quite different, they can be normalized to make their contributions to the final result more balanced.norm and p norm are the normalized privacy weight and the first private data priority, respectively. Then the formula can be further modified as follows:
[0071] f=max(0,f max -(α·w norm +β·p norm ))
[0072] Normalization can be achieved as follows:
[0073]
[0074] Among them, w min and w max is the minimum and maximum value of the privacy weight, p min and p max is the minimum and maximum value of the first private data priority. In this way, the formula can dynamically adjust the second training frequency identifier according to the changes in the privacy weight and the private data priority, while ensuring the rationality and effectiveness of the result. In summary, the second training frequency identifier calculated here objectively evaluates the contribution of the robot training data uploaded by the target node this time through the two dimensions of privacy and scarcity, and provides a basis for the reasonable allocation of federated learning resources to the target node during the next training.
[0075] Step S130: Update the second private data priority corresponding to the current training round according to the robot training data reported by each node.
[0076] As mentioned above, the private data priority in the embodiment of the present invention is an indicator designed by the aggregation server to evaluate the scarcity of training data. The private data priority is not only used at the aggregation server to evaluate the scarcity of training data uploaded by each node, but is also dynamically updated by the aggregation server in each round based on the training data collected from all nodes, and is sent to each node at the end of each round of training to guide each node to improve the quality of the robot training data uploaded next time.
[0077] Among them, the first private data priority is calculated and determined by the aggregation server during the previous round of federated learning training. Correspondingly, during this round of federated learning training, the aggregation server will calculate and determine the second private data priority of training samples of different data types in the next round.
[0078] Specifically, Figure 3 As shown, the method for determining the priority of private data in step S130 can be specifically implemented according to steps S131 to S133.
[0079] Step S131, extracting the data type in the robot training data collected in the current training round.
[0080] Step S132, calculating the aggregation degree of the robot training data based on the data type and grouping the training data.
[0081] Step S133: assigning different second private data priorities to different data types in each training data group according to the aggregation degree of each training data group.
[0082] Specifically, for the robot training data collected from each node in the current round, the aggregation server first needs to clean the collected data, remove duplicate, invalid or abnormal data, and then standardize the cleaned data to ensure that data from different sources are comparable. Then the training data can be grouped according to the type of robot training data mentioned in the full text. After the grouping is completed, the aggregation server calculates the aggregation degree of each group of data, which can be achieved by calculating indicators such as similarity, consistency or correlation of the data. Each group of data is sorted according to the calculated aggregation degree, and the data group with high aggregation degree is ranked first. Data with high aggregation degree indicates that they have high similarity or consistency, indicating that this type of data has a relatively rich data source, which is reflected in the embodiment of the present invention as low scarcity. This type of data will be marked with a lower second private data priority value. On the contrary, the training data type in the group with low aggregation degree has high data scarcity and will be marked with a higher second private data priority value. The specific assigned value can be set according to the actual needs of the user.
[0083] As described in this step, the aggregation server will dynamically update the priority values of different private data in each round of federated learning training, that is, the second private data priority in this step, as a reflection of the scarcity of current training data. In specific implementation, the second private data priority can be directly calculated using the robot training data of the current round; or after calculating the scarcity of the robot training data of the current round, the calculation result can be updated to the priority value of the historical total training data in the form of weight adjustment for numerical update.
[0084] Step S140: Send the model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier to the target node.
[0085] After step S130 is completed, the aggregation server can perform the current round of federated learning training. The embodiment of the present invention does not limit the specific model type adopted by the robot control model. Exemplarily, a federated deep learning model can be used to train the robot control model. This type of model usually uses a deep neural network as a function approximator, and is collaboratively trained between multiple robots or data holders through federated learning. For example, a convolutional neural network is suitable for processing image data, such as target detection and recognition in robot vision tasks, and can be collaboratively trained between multiple robots through federated learning to improve the accuracy and generalization ability of the model. Another example is a recurrent neural network (RNN) or a long short-term memory network (LSTM), which is suitable for processing sequence data, such as robot motion trajectory prediction and control. The federated learning method can enable multiple robots to share information in the sequence data, thereby improving the prediction and control performance of the model. The specific model type used should be determined according to the application scenario of the robot in the industrial Internet, and determined by negotiation among the participants in the federated learning.
[0086] The trained model parameters will be sent to the nodes participating in this round of federated learning training as normal. In addition, unlike the general federated learning process, the embodiment of the present invention will also send the second private data priority and the second training frequency identifier to the target node.
[0087] Among them, the second private data priority is the scarcity index of the current training data of different types of robots calculated by the aggregation server. This index is the same for each node and is used to guide each node on how to increase its contribution when reporting training data next time, so as to avoid the situation where it cannot participate in federated learning training at a high frequency. The second training frequency identifier is an indicator that directly evaluates the current contribution of the target node to the federated learning training. This indicator is different for each node and is used to indicate the round in which the target node can participate in the next federated learning training. The target node can refer to the second training frequency identifier to implement targeted data type training data collection within a specific time. It can be understood that each participating node in the federated learning training collects and reports training data according to the second private data priority and the second training frequency identifier received by each node, which are given to the target node. The two indicators, namely, effectively ensure the fairness of the federated learning training process to all participants, and improve the data consistency of the global model trained, thereby avoiding the occurrence of situations such as model performance degradation, reduced training efficiency, and increased overfitting risk.
[0088] In the training method of the robot control model provided by the embodiment of the present invention, the aggregation server will evaluate the contribution of the training data reported by each participating node to the global model training from the two dimensions of privacy and scarcity of training data in each round of federated learning training, and control the frequency of each node participating in the federated learning training according to the contribution, and provide guidance for each node to collect and report data in the next round, thereby effectively ensuring the fairness of the federated learning training process to all participants and improving the data consistency of the global model being trained, thereby avoiding the occurrence of situations such as degradation of model performance, reduced training efficiency, and increased risk of overfitting.
[0089] Based on any of the above embodiments, Figure 4 The following is a schematic diagram showing the structure of a training device for a robot control model provided by an embodiment of the present invention, and the specific contents are as follows:
[0090] The training admission judgment module 410 is used to confirm that the first training frequency identifier carried by the target node meets the preset conditions;
[0091] A training contribution determination module 420, configured to determine a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority;
[0092] A data priority calculation module 430, configured to update the second private data priority corresponding to the current training round according to the robot training data reported by each node;
[0093] The parameter sending module 440 is used to send the model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier to the target node.
[0094] In the training device of the robot control model provided by the embodiment of the present invention, the aggregation server will evaluate the contribution of the training data reported by each participating node to the global model training from the two dimensions of privacy and scarcity of training data in each round of federated learning training, and control the frequency of each node participating in the federated learning training according to the contribution, and provide guidance for each node to collect and report data in the next round, thereby effectively ensuring the fairness of the federated learning training process to all participants and improving the data consistency of the global model being trained, thereby avoiding the occurrence of situations such as degradation of model performance, reduced training efficiency, and increased risk of overfitting.
[0095] Based on any of the above embodiments, Figure 4The physical structure diagram of the electronic device provided by the embodiment of the present invention is shown, and the electronic device may include: a processor (processor) 510, a communication interface (CommunicationsInterface) 520, a memory (memory) 530 and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the following method:
[0096] Confirming that the first training frequency identifier carried by the target node meets the preset conditions;
[0097] Determine a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority;
[0098] Update the second private data priority corresponding to the current training round according to the robot training data reported by each node;
[0099] The model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier are sent to the target node.
[0100] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the embodiment of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0101] On the other hand, an embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in each of the above embodiments is implemented, for example, including:
[0102] Confirming that the first training frequency identifier carried by the target node meets the preset conditions;
[0103] Determine a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority;
[0104] Update the second private data priority corresponding to the current training round according to the robot training data reported by each node;
[0105] The model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier are sent to the target node.
[0106] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a robot control model, characterized in that: The method comprises: Confirming that the first training frequency identifier carried by the target node meets the preset conditions; Determine a second training frequency identifier corresponding to the target node according to the privacy weight of the robot training data reported by the target node and the first private data priority; Update the second private data priority corresponding to the current training round according to the robot training data reported by each node; The model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier are sent to the target node.
2. The training method of the robot control model according to claim 1, characterized in that: The step of confirming that the first training frequency identifier carried by the target node meets a preset condition specifically includes: The current time is compared with the time in the first training frequency identifier. If the time in the first training frequency identifier is later, the preset condition is met.
3. The training method of the robot control model according to claim 1, characterized in that: The updating of the second private data priority corresponding to the current training round according to the robot training data reported by each node specifically includes: Extract the data type from the robot training data collected in the current training round; Calculate the aggregation degree of the robot training data based on the data type and group the training data; Different second private data priorities are assigned to different data types in each training data group according to the aggregation degree of each training data group.
4. The training method of the robot control model according to claim 1, characterized in that: The determining, according to the privacy weight of the robot training data reported by the target node and the first private data priority, a second training frequency identifier corresponding to the target node specifically includes: Determining the privacy level of each training data sample in the training data, and then calculating the privacy weight of the training data; Divide the robot training data into a private data sample set and a non-private data sample set, and calculate a first private data priority of the robot training data according to the private data sample set; A second training frequency identifier corresponding to the target node is obtained by calculation according to the privacy weight and the first privacy data priority.
5. The training method of the robot control model according to claim 4, characterized in that: The private data sample set includes the robot's position data, motion trajectory data, working environment data and task data.
6. The training method of the robot control model according to claim 4, characterized in that: The non-private data sample set includes basic control parameters, device status data, public environment data and non-specific task data.
7. The training method of the robot control model according to claim 4, characterized in that: The privacy weight of the robot training data is determined according to the amount and data type of the training data.
8. A training device for a robot control model, characterized in that: The device comprises: A training admission judgment module is used to confirm that the first training frequency identifier carried by the target node meets the preset conditions; A training contribution determination module, configured to determine a second training frequency identifier corresponding to the target node according to a privacy weight of the robot training data reported by the target node and a first private data priority; A data priority calculation module, used to update the second private data priority corresponding to the current training round according to the robot training data reported by each node; A parameter delivery module is used to deliver the model parameters of the current round of federated learning training, the second private data priority and the second training frequency identifier to the target node.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors execute the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, characterized in that: When the executable instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Information processing method and device, equipment, storage medium and computer program product
CN112906864A
Federal learning model training method for large-scale industrial chain privacy calculation
CN114169412A
Fair privacy calculation method based on federated node contribution
CN116306910A
Federal machine learning method and device, storage medium and processor
CN117521783A
Federal learning method and device, electronic equipment and computer readable storage medium
CN118095475A