A model training method, apparatus and system, as well as a prediction method and apparatus
Through the distributed model training method, the coordinated work of the first device and the second device is used to solve the data diversity and security of network element equipment, and efficient model training and prediction are realized, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN201910744658.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2039-08-13
AI Technical Summary
In traditional network intelligent systems, network element equipment data are diverse and inconsistent, centralized data collection is difficult, high bandwidth occupies, and there is a risk of user privacy leakage, resulting in difficulty in model training and security management.
Using the distributed model training method, the second device participating in the training is determined through the first device, the training job and parameters are sent, the second device conducts local training and uploads incremental parameters. The first device summarizes the results, avoids data transmission, and utilizes the parallel computing power of multiple devices.
It reduces bandwidth usage, reduces the difficulty of data security management, realizes the scalability and efficiency of model training, and is suitable for a variety of application scenarios.
Smart Images

Figure CN112396070B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to, but are not limited to, the fields of artificial intelligence and telecom network, and particularly to a model training method, apparatus and system, and a prediction method and apparatus. Background Art
[0002] In the traditional network intelligent system cooperation mode, the operation support system (OSS) collects data from each network element, centrally processes data and trains models on the OSS, and then distributes the trained models to each network element for inference. However, this method has several problems. First, the data generated by network elements are diverse, and there is no unified data for network element devices of different equipment manufacturers, and some data are not opened by equipment manufacturers, which causes difficulties in centralized data collection and processing. Second, the amount of data generated by network elements is huge, and collecting data will occupy a large amount of backhaul bandwidth, increasing the network deployment cost. Finally, centralized data collection may bring the risk of user privacy leakage to third parties, increasing the difficulty of data security management. That is to say, in some application scenarios (for example, when equipment manufacturers do not open data), model training cannot be carried out, collecting the data required for model training will occupy a large bandwidth, and the difficulty of data security management is relatively high. Summary of the Invention
[0003] Embodiments of the present invention provide a model training method, apparatus and system, and a prediction method and apparatus, which can be applicable to model training in various application scenarios, reduce the occupied bandwidth, and reduce the difficulty of data security management.
[0004] Embodiments of the present invention provide a model training method, including:
[0005] The first device determines the second device participating in the model training according to the description of the data required for model training in the model training job, and sends part or all of the model training job to the second device;
[0006] The first device executes the first model training code in the model training job to implement that for the j-th training step, when the model training is not completed, the first device sends the model parameters corresponding to the j-th training step to the second device; receives the increment of the model parameters corresponding to the j-th training step uploaded by the second device, and calculates the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device.
[0007] Embodiments of the present invention provide a model training method, including:
[0008] The second device receives some or all of the model training jobs sent by the first device;
[0009] For the j-th training step, the second device receives the model parameters corresponding to the j-th training step sent by the first device, performs model training based on the model parameters corresponding to the j-th training step and some or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step to the first device.
[0010] An embodiment of the present invention provides a model training device, including a processor and a computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are executed by the processor, any of the above model training methods is implemented.
[0011] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above model training methods are implemented.
[0012] An embodiment of the present invention provides a model training device, including:
[0013] A model training job distribution module, configured to determine a second device participating in model training according to a description of data required for model training in a model training job, and send some or all of the model training job to the second device;
[0014] A first model training module, configured to execute the first model training code in the model training job to implement that for the j-th training step, when the model training is not completed, the first device sends the model parameters corresponding to the j-th training step to the second device; receives the model parameter increment corresponding to the j-th training step uploaded by the second device, and calculates the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device.
[0015] An embodiment of the present invention provides a model training device, including:
[0016] A model training job receiving module, configured to receive some or all of the model training jobs sent by the first device;
[0017] A second model training module, configured to for the j-th training step, receive the model parameters corresponding to the j-th training step sent by the first device, perform model training based on the model parameters corresponding to the j-th training step and some or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step, and upload the model parameter increment corresponding to the j-th training step to the first device.
[0018] An embodiment of the present invention provides a model training system, including:
[0019] A first device, configured to determine a second device participating in model training according to a description of data required for model training in a model training job, send part or all of the model training job to the second device; execute a first model training code in the model training job to implement that for the j-th training step, when the model training is not completed, send model parameters corresponding to the j-th training step to the second device; receive a model parameter increment corresponding to the j-th training step uploaded by the second device, and calculate model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device;
[0020] A second device, configured to receive part or all of the model training job sent by the first device; for the j-th training step, receive model parameters corresponding to the j-th training step sent by the first device, perform model training according to the model parameters corresponding to the j-th training step and part or all of the model training job to obtain a model parameter increment corresponding to the j-th training step, and upload the model parameter increment corresponding to the j-th training step to the first device.
[0021] An embodiment of the present invention provides a prediction method, including:
[0022] Obtain data required for prediction, and extract key features from the data required for prediction;
[0023] Input the key features into the model corresponding to the trained model parameters in any of the above model training methods, and output a predicted value.
[0024] An embodiment of the present invention provides a prediction device, including a processor and a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are executed by the processor, any of the above prediction methods is implemented.
[0025] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above prediction methods are implemented.
[0026] An embodiment of the present invention provides a prediction device, including:
[0027] A data acquisition module, configured to acquire data required for prediction;
[0028] A key feature extraction module, configured to extract key features from the data required for prediction;
[0029] A prediction module, configured to input the key features into any of the above model training methods, and output a prediction value using the model corresponding to the trained model parameters.
[0030] Embodiments of the present invention include: A first device determines a second device participating in model training according to the description of the data required for model training in a model training job, and sends part or all of the model training job to the second device; executes the first model training code in the model training job to implement that for the j-th training step, when the model training is not completed, the first device sends the model parameters corresponding to the j-th training step to the second device; receives the model parameter increment corresponding to the j-th training step uploaded by the second device, and calculates the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device. After the embodiments of the present invention send the model training job to the second device for distributed model training, the model training results of the second device are summarized in the first device, avoiding the transmission of training data between the first device and the second device, being applicable to model training in various application scenarios (for example, when device manufacturers do not open data), reducing the occupied bandwidth, reducing the difficulty of data security management, and at the same time, making full use of the parallel computing capabilities of multiple second devices to achieve the scalability of the model training system.
[0031] Other features and advantages of the embodiments of the present invention will be described in the following specification, and partly become obvious from the specification, or are understood by implementing the embodiments of the present invention. The objectives and other advantages of the embodiments of the present invention can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings are used to provide a further understanding of the technical solutions of the embodiments of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solutions of the embodiments of the present invention, and do not constitute a limitation to the technical solutions of the embodiments of the present invention.
[0033] Figure 1 It is a flowchart of a model training method proposed in an embodiment of the present invention;
[0034] Figure 2 It is a flowchart of a model training method proposed in another embodiment of the present invention;
[0035] Figure 3 It is a schematic diagram of the installation and deployment of CTE and DTE provided by an embodiment of the present invention;
[0036] Figure 4 It is an internal architecture diagram of CTE and DTE provided by an embodiment of the present invention;
[0037] Figure 5 The architecture diagram of the model training systems for Example 1 and Example 2 of the embodiments of the present invention;
[0038] Figure 6 The architecture diagram of the model training system for Example 3 of the embodiments of the present invention;
[0039] Figure 7 The schematic diagram of the structural composition of the model training device proposed in another embodiment of the present invention;
[0040] Figure 8 The schematic diagram of the structural composition of the model training device proposed in another embodiment of the present invention;
[0041] Figure 9 The schematic diagram of the structural composition of the model training system proposed in another embodiment of the present invention. Detailed implementation manners
[0042] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other arbitrarily.
[0043] The steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0044] See Figure 1 , an embodiment of the present invention proposes a model training method, including:
[0045] Step 100, the first device determines the second device participating in the model training according to the description of the data required for the model training in the model training job, and sends part or all of the model training job to the second device.
[0046] In the embodiments of the present invention, the model training job includes: the description of the data required for the model training, the data processing code, the first model training code, and the second model training code.
[0047] The model in the embodiments of the present invention can be any kind of model, such as an artificial intelligence (AI) model, a deep learning model, a machine learning model, etc.
[0048] In the embodiments of the present invention, the model training job can be pre-set in the first device, or can be deployed by the user on the first device, that is, the user inputs the model training job, and the first device receives the model training job input by the user.
[0049] In an embodiment of the present invention, the model training job can be carried and sent in a start model training message (such as a deploy_training_job message).
[0050] In an embodiment of the present invention, for example, when the description of the data required for model training includes a list of cell IDs, the second devices participating in the model training are the second devices corresponding to all base stations covered by all cells in the list of cell IDs; when the description of the data required for model training includes a list of device IDs (such as Packet Transport Network (PTN) device IDs), the second devices participating in the model training are the second devices corresponding to all devices in the list of device IDs.
[0051] Step 101: The first device executes the first model training code in the model training job to implement, for the j-th training step, when the model training is not finished, the first device sends the model parameters corresponding to the j-th training step to the second device; receives the increment of the model parameters corresponding to the j-th training step uploaded by the second device, and calculates the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device.
[0052] In an embodiment of the present invention, the first device can send a start training step message (such as a start_training_step message) to the second device. The start training step message carries the model parameters corresponding to the j-th training step, and can also carry the number of the j-th training step. The number of the j-th training step is not necessary.
[0053] In an exemplary instance, the model parameters corresponding to the j-th training step can be represented in the form of a vector W, that is, the model parameter vector W corresponding to the j-th training step. Of course, it can also not be represented in the form of a vector. The embodiment of the present invention does not limit the specific representation form of the model parameters.
[0054] Similarly, the increment of the model parameters corresponding to the j-th training step can also be represented in the form of a vector. Of course, it can also not be represented in the form of a vector. The embodiment of the present invention does not limit the specific representation form of the increment of the model parameters.
[0055] In an embodiment of the present invention, the first device can receive a training step end message (such as a training_step_finished message) sent by the i-th second device. The training step end message carries the increment of the model parameters corresponding to the j-th training step. The training step end message also carries any one or more of the following: the number of training data used in the j-th training step, the loss function value corresponding to the j-th training step, and other content specified to be uploaded in the model training code.
[0056] In an embodiment of the present invention, calculating the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device includes:
[0057] Calculating the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device;
[0058] Calculating the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step.
[0059] In an embodiment of the present invention, various methods can be used to calculate the global model parameter increment corresponding to the j-th training step. In an exemplary example, before calculating the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device, the method further includes:
[0060] The first device receives the number of training data used in the j-th training step uploaded by the second device;
[0061] Calculating the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device includes:
[0062] According to the formula Calculating the global model parameter increment corresponding to the j-th training step;
[0063] Where Is the global model parameter increment corresponding to the j-th training step, #D ij Is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in model training, ΔW ij Is the model parameter increment corresponding to the j-th training step uploaded by the i-th second device.
[0064] In an embodiment of the present invention, various methods can be used to calculate the model parameters corresponding to the (j + 1)-th training step. In an exemplary example, calculating the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step includes:
[0065] According to the formula Calculating the model parameters corresponding to the (j + 1)-th training step;
[0066] Where W j+1 Is the model parameter corresponding to the (j + 1)-th training step, W jis the model parameter corresponding to the j-th training step, and α is the learning rate, which is a constant. is the global model parameter increment corresponding to the j-th training step.
[0067] In another embodiment of the present invention, when the model training ends, the method further includes any one or more of the following:
[0068] The first device sends a stop model training message (such as a delete_training_job message) to the second device;
[0069] The first device does not send the model parameter corresponding to the j-th training step to the second device.
[0070] In the implementation of the present invention, various methods can be used to determine whether the model training ends, and these methods can be implemented in the first model training code in the model training job.
[0071] In an exemplary instance, determining whether the model training ends includes any one or more of the following:
[0072] When j is greater than or equal to the maximum number of training steps, it is determined that the model training ends;
[0073] When j is less than the maximum number of training steps, it is determined that the model training has not ended.
[0074] The maximum number of training steps can be specified by the user when deploying the model training job.
[0075] In another exemplary instance, determining whether the model training ends includes any one or more of the following:
[0076] When the difference between the averages of the loss functions corresponding to any two adjacent training steps from the (j - m + 1)-th training step to the j-th training step is less than or equal to a preset threshold, it is determined that the model training ends;
[0077] When there are at least two adjacent training steps from the (j - m + 1)-th training step to the j-th training step whose difference between the averages of the loss functions is greater than the preset threshold, it is determined that the model training has not ended.
[0078] That is, when the average value of the loss function has not changed significantly in consecutive m training steps, it is considered that the model training has been completed.
[0079] In an exemplary instance, according to the formula calculate the average value of the loss function corresponding to the j-th training step;
[0080] where is the average value of the loss function corresponding to the j-th training step, is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in model training, L ij is the loss function value corresponding to the j-th training step uploaded by the i-th device.
[0081] See Figure 2 , another embodiment of the present invention proposes a model training method, including:
[0082] Step 200, the second device receives part or all of the model training jobs sent by the first device.
[0083] In the embodiments of the present invention, the second device may receive a start model training message (such as a deploy_training_job message) sent by the first device, and obtain part or all of the model training jobs from the model training message.
[0084] In the embodiments of the present invention, the model training job includes: a description of the data required for model training, data processing code, first model training code, and second model training code.
[0085] In the embodiments of the present invention, after the second device receives part or all of the model training jobs sent by the first device, it may start the model training job.
[0086] Step 201, for the j-th training step, the second device receives the model parameters corresponding to the j-th training step sent by the first device, performs model training based on the model parameters corresponding to the j-th training step and part or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step to the first device.
[0087] In the embodiments of the present invention, the second device may receive a start training step message (such as a start_training_step message) sent by the first device. The start training step message carries the model parameters corresponding to the j-th training step, and may also carry the number of the j-th training step. The number of the j-th training step is not necessary.
[0088] In an exemplary instance, the model parameters corresponding to the j-th training step may be represented in the form of a vector W, that is, the model parameter vector W corresponding to the j-th training step. Of course, it may also not be represented in the form of a vector. The embodiments of the present invention do not limit the specific representation form of the model parameters.
[0089] Similarly, the model parameter increment corresponding to the j-th training step can also be represented in the form of a vector. Of course, it can also be represented in a form other than a vector, and the embodiment of the present invention does not limit the specific representation form of the model parameter increment.
[0090] In an exemplary embodiment, performing model training according to the model parameters corresponding to the j-th training step and part or all of the model training job to obtain the model parameter increment corresponding to the j-th training step includes:
[0091] Execute the data processing code to obtain the training data corresponding to the j-th training step from the network element corresponding to itself according to the description of the data required for the model training, process the training data corresponding to the j-th training step to obtain the training sample corresponding to the j-th training step, and execute the second model training code to perform model training according to the model parameters corresponding to the j-th training step and the training sample corresponding to the j-th training step to obtain the model parameter increment corresponding to the j-th training step.
[0092] In another exemplary embodiment, after the second device receives part or all of the model training job sent by the first device, the method further includes:
[0093] Executing the data processing code to obtain training data from the network element corresponding to itself according to the description of the data required for the model training, and processing the training data to obtain training samples;
[0094] The step of performing model training according to the model parameters corresponding to the j-th training step and part or all of the model training job to obtain the model parameter increment corresponding to the j-th training step includes:
[0095] The second model training code is executed to perform model training according to the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step.
[0096] That is to say, the training data corresponding to different training steps may be the same or different, that is, the training data corresponding to different training steps may be obtained at one time after receiving part or all of the model training job, or different training data may be obtained in real time in each training step. The embodiments of the present invention are not limited to this.
[0097] In an embodiment of the present invention, the network element corresponding to the second device may be one or multiple. When the second device is installed as a single board inside the network element device, there is only one network element corresponding to the second device, which is the network element where the single board is located; when the second device is independently deployed outside the network element in the form of a separate device, the second device may be connected to one or multiple network element devices. In this case, the network elements corresponding to the second device are the network element devices of the same level as the second device, which may be one or multiple.
[0098] In an embodiment of the present invention, the second device may send a training step end message (such as a training_step_finished message) to the first device, and the training step end message carries the model parameter increment corresponding to the j-th training step.
[0099] In another embodiment of the present invention, the method further includes: the second device receives a stop model training message (such as a delete_training_job message) sent by the first device. After receiving the stop model training message, the second device ends this process and no longer performs model training.
[0100] In another embodiment of the present invention, the method further includes any one or more of the following:
[0101] The second device uploads the number of training data used in the j-th training step to the first device;
[0102] The second device uploads the loss function value corresponding to the j-th training step to the first device.
[0103] In an embodiment of the present invention, the training step end message further carries any one or more of the following: the number of training data used in the j-th training step, the loss function value corresponding to the j-th training step, and other content specified to be uploaded in the model training code.
[0104] After the embodiment of the present invention distributes the model training job to the second device for distributed model training, it then aggregates the model training results of the second device in the first device, avoiding the transmission of training data between the first device and the second device, being applicable to model training in various application scenarios (for example, when the device manufacturer does not open the data), reducing the occupied bandwidth, reducing the difficulty of data security management, and at the same time, making full use of the parallel computing capabilities of multiple second devices to achieve the scalability of the model training system.
[0105] In an embodiment of the present invention, such as Figure 3As shown in the figure, the software and hardware devices of the training engine (TE, Training Engine) can be installed and deployed on the OSS and network elements respectively, and the OSS and multiple network elements can be formed into a data parallel distributed model training system to jointly complete model training.
[0106] Among them, the OSS is a system in the communication network, responsible for the operation and management of the communication network. It includes subsystems at two levels. The element management system (EMS, Element Management System) is responsible for the management at the network element level, and the network management system (NMS, Network Management System) is responsible for the management at the network level across multiple network elements. The network element devices in the communication network usually include two subsystems. The operation, administration and maintenance (OAM, Operation Administration and Maintenance) subsystem is responsible for the operation, administration and maintenance of this network element, and the protocol stack (Protocal Stack) subsystem is responsible for implementing the protocol stack function of the network element.
[0107] Then, the first device can be the training engine deployed on the OSS, that is, the centric training engine (CTE, Centric Training Engine), which is mainly responsible for model training job management, distribution of model training jobs, synchronization of training steps, summary and update of model parameters, etc.; the second device can be the training engine deployed on the network element, that is, the distributed training engine (DTE, Distributed Training Engine), which is mainly responsible for training the model distributed by the CTE using the local data of the network element and uploading the model parameter increments generated in each training step to the CTE.
[0108] The DTE can be installed as a single board inside the network element device, or can be independently deployed outside the network element in the form of a separate device and connected to one or more network element devices. To accelerate the process of model training, the DTE can include dedicated computing acceleration hardware required for accelerating model training, such as a graphics processing unit (GPU, Graphics Processing Unit), digital signal processing (DSP, Digital Signal Processing), field programmable gate array (FPGA, Field-Programmable Gate Array) or application specific integrated circuit (ASIC, Application specific Integrated Circuit), as Figure 4 shown.
[0109] For example, as Figure 4 shown, the CTE includes three components: a job manager, a DTE controller, and a training coordinator.
[0110] Job Manager: Responsible for the lifecycle management of multiple simulation training job instances, allowing the CTE to execute multiple distributed model training jobs in parallel.
[0111] DTE Controller: Responsible for implementing the interaction between the CTE and the DTE during the model training process, including the selection of DTEs participating in distributed model training, the distribution of model training jobs to each DTE, and the communication for each training step, such as the distribution of the CTE's model parameters and the collection of the model parameter increments of each DTE.
[0112] Training Coordinator: Executes the first model training code to control the training steps, calculate the global model parameter increments, update the global model parameters, and determine whether the model training is completed, etc. The CTE hardware can use general-purpose servers.
[0113] DTE: Includes three components: a data collector, a data processor, and a model trainer.
[0114] Data Collector: Responsible for parsing the description of the data required for model training in the model training job, and obtaining the corresponding original training data from the OAM subsystem of the network element according to the description of the data required for model training.
[0115] Data Processor: Provides a runtime library for data processing algorithms, executes the data processing code in the model training job, and processes the original training data into training samples required for model training.
[0116] Model Trainer: Provides a runtime library for model training algorithms such as machine learning and deep learning, executes the second model training code, and uses the training samples to train the model to obtain model parameter increments.
[0117] In the embodiments of the present invention, the network element can be any network element, such as a base station, etc.
[0118] The following lists several examples to illustrate the implementation process of the above model training method. The examples listed are only for convenience of explanation and are not used to limit the protection scope of the embodiments of the present invention.
[0119] Example 1
[0120] This example illustrates the model training method of the embodiments of the present invention by taking the training of a radio access network (RAN) coverage prediction model as an example.
[0121] Such as Figure 5As shown in the figure, CTE is deployed in the OSS of the radio access network (RAN), and DTE is deployed in 2 / 3 / 4 / 5G base stations. Since a large number of 2 / 3 / 4G base stations already exist in the existing network, in order to avoid modifying the existing hardware, DTE is deployed in an external manner; DTE is deployed in an internal manner in 5G base stations.
[0122] Step A: The user deploys the model training job to be performed on the CTE through the OSS.
[0123] The model training job mainly includes:
[0124] 1) The data required for model training is described in YAML language, including a list of cell identity identifiers (ID, Identifier), as well as the cell configuration data, antenna configuration data, and measurement report data corresponding to each cell;
[0125] 2) The data processing code can be written in Python language. The data processor of DTE can execute the data processing code to complete data processing; the main function of the data processing code is to extract the key features corresponding to each cell (as shown in Table 1) from the configuration data and measurement report data corresponding to each cell, and generate training samples.
[0126]
[0127]
[0128] Table 1
[0129] 3) The first model training code and the second model training code can be written using the Python-based SDK provided by DTE. The training coordinator of CTE can execute the first model training code to complete the update of model parameters. The model trainer of DTE can execute the second model training code to complete model training, and use dedicated hardware to accelerate the training calculation process. The main function of the second model training code is to establish a deep neural network model (MLP, Multi-Layer Preceptron): its input is the features shown in Table 1, and the output is the predicted value of the reference signal receiving power (RSRP, Reference Signal Receiving Power) of the wireless coverage field strength; the mean square error (MSE, Mean Square Error) is used as the objective function; and the hyperparameters of the model training network (such as the maximum number of training steps, training end strategy, etc.);
[0130] Step B: The CTE determines the DTEs participating in model training based on the list of cell IDs included in the model training job, that is, all DTEs in all base stations covered by all cells in the cell ID list, and sends the data required for model training, data processing code, and second model training code in the model training job to the DTEs;
[0131] Step C: The DTE executes the data processing code to obtain training data (i.e., the cell ID list, and the cell configuration data, antenna configuration data, and measurement report data corresponding to each cell) from the base station corresponding to itself according to the description of the data required for the model training, and processes the training data to obtain training samples;
[0132] Step D: The CTE executes the first model training code to implement that for the j-th training step, the CTE sends the model parameters corresponding to the j-th training step to the DTE; the DTE receives the model parameters corresponding to the j-th training step, executes the second model training code to implement model training based on the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step and the loss function value corresponding to the j-th training step to the CTE; the CTE calculates the average value of the loss function according to the formula When the average value of the loss function does not continue to decrease in 20 consecutive training steps, the training is completed. The CTE uploads the model corresponding to the trained model parameters to the specified location of the OSS.
[0133] Step E: The OSS system obtains the trained RAN coverage prediction model, and this RAN coverage prediction model can be used for subsequent prediction of the coverage of the wireless network.
[0134] Example 2
[0135] This example illustrates the model training method of the embodiment of the present invention by taking the training of the RAN cell traffic prediction model as an example.
[0136] As Figure 5 shown, the CTE is deployed in the OSS of the radio access network (RAN), and the DTE is deployed in the 2 / 3 / 4 / 5G base stations. Since a large number of 2 / 3 / 4G base stations already exist in the existing network, in order to avoid transforming the existing hardware, the DTE is deployed in an external manner; the DTE is deployed in an internal manner for the 5G base stations.
[0137] Step A: The user deploys the model training job to be performed on the CTE through the OSS system.
[0138] The model training job mainly includes:
[0139] 1) The data required for model training is described in YAML language, including a list of cell identity identifiers (ID, Identifier), the key performance indicator (KPI, Key Performance Indicator) data corresponding to each cell, and the data collection period;
[0140] 2) The data processing code can be written in Python language. The data processor of DTE can execute the data processing code to complete data processing; the main function of the data processing code is to extract the key features corresponding to each cell (as shown in Table 2) from the cell KPI data corresponding to each cell, and generate training samples.
[0141]
[0142]
[0143] Table 2
[0144] 3) The first model training code and the second model training code can be written using the Python-based SDK provided by DTE. The training coordinator of CTE can execute the first model training code to complete the update of model parameters. The model trainer of DTE can execute the second model training code to complete model training and use dedicated hardware to accelerate the training calculation process. The main function of the second model training code is to establish a deep neural network model (Conv3DNet + LSTM): its input is the features shown in Table 2, and the output is the predicted number of call detail reports (CDR, Call Detail Report) for cell user access requests; MSE is used as the objective function; and the hyperparameters for training the network (such as the maximum number of training steps, training end strategy, etc.);
[0145] Step B: CTE determines the DTEs participating in model training according to the list of cell IDs included in the model training job, that is, all DTEs in all base stations covered by all cells in the cell ID list, and sends the data, data processing code, and second model training code required for model training in the model training job to the DTEs;
[0146] Step C: The DTE executes the data processing code to obtain training data from the base station corresponding to itself according to the description of the data required for model training, and processes the training data to obtain training samples;
[0147] Step D: The CTE executes the first model training code so that for the j-th training step, the CTE sends the model parameters corresponding to the j-th training step to the DTE. After receiving the model parameters corresponding to the j-th training step, the DTE executes the second model training code to perform model training based on the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step and the loss function value corresponding to the j-th training step to the CTE. The CTE calculates the average value of the loss function according to the formula When the average value of the loss function does not continue to decrease in 20 consecutive training steps, the training is completed. The CTE uploads the model corresponding to the trained model parameters to the specified location of the OSS.
[0148] Step E: The OSS system obtains the trained RAN cell traffic prediction model, which can be used to predict the cell traffic (such as predicting the cell traffic volume) in the future.
[0149] Example 3
[0150] This example illustrates the model training method of the embodiments of the present invention by taking the training of the optical module fault prediction model of the wired bearer network (BN) as an example.
[0151] As Figure 6 shown, the CTE is deployed in the OSS of the wired bearer network (BN), and the DTE is deployed in the PTN device in a built-in manner.
[0152] Step A: The user deploys the model training job to be performed on the CTE through the OSS.
[0153] The model training job mainly includes:
[0154] 1) The data required for training is described in the YAML language, including the list of PTN device IDs, and the optical module monitoring data, optical module alarm data, and data collection period corresponding to each PTN device;
[0155] 2) The data processing code can be written in the Python language, and the data processor of the DTE can execute the data processing code to complete data processing; the main function of the data processing code is to extract the key features corresponding to each PTN device (as shown in Table 3) from the optical module monitoring data and optical module alarm data corresponding to each PTN device, and generate training samples.
[0156] Feature Description Datetime Collection Time Pn Manufacturer Number Sn Optical Module Serial Number Txpower Transmission Power Biascurrent Bias Current Temperature Operating Temperature Voltage Voltage Fault Whether Fault Alarms
[0157] Table 3
[0158] 3) The first model training code and the second model training code can be written using the SDK based on the Python language provided by DTE. The training coordinator of CTE can execute the first model training code to complete the update of model parameters. The model trainer of DTE can execute the second model training code to complete model training and use dedicated hardware to accelerate the training calculation process. The main function of the second model training code is to establish a logistic regression model: its input is the features shown in Table 3, and the output is whether the optical module fails (0 - no failure, 1 - failure); cross-entropy is used as the objective function; and the hyperparameters of the training network (such as the maximum number of training steps, training end strategy, etc.);
[0159] Step B: CTE determines the DTEs participating in model training according to the list of PTN device IDs included in the model training job, that is, all DTEs in all PTN devices deployed in the list of PTN device IDs, and sends the data, data processing code, and second model training code required for model training in the model training job to the DTEs;
[0160] Step C: The DTE executes the data processing code to obtain training data from the PTN device corresponding to itself according to the description of the data required for model training, and processes the training data to obtain training samples;
[0161] Step D: CTE executes the first model training code to implement that for the j-th training step, CTE sends the model parameters corresponding to the j-th training step to the DTE; the DTE receives the model parameters corresponding to the j-th training step, executes the model training code to implement model training based on the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step and the loss function value corresponding to the j-th training step to CTE; CTE calculates the average value of the loss function according to the formula When the average value of the loss function does not continue to decrease in 20 consecutive training steps, the training is completed. CTE uploads the model corresponding to the trained model parameters to the specified location in OSS.
[0162] Step E: The OSS system obtains the trained optical module fault prediction model, and this optical module fault prediction model can be used to predict whether the optical module fails in the future.
[0163] Another embodiment of the present invention proposes a model training device, including a processor and a computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are executed by the processor, the above-mentioned any model training method is implemented.
[0164] Another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above model training methods are implemented.
[0165] See Figure 7 , another embodiment of the present invention provides a model training device (such as the above first device), including:
[0166] A model training job distribution module 701, configured to determine a second device participating in model training according to the description of the data required for model training in the model training job, and send part or all of the model training job to the second device;
[0167] A first model training module 702, configured to execute the first model training code in the model training job to implement that for the j-th training step, when the model training is not completed, the first device sends the model parameters corresponding to the j-th training step to the second device; receive the model parameter increment corresponding to the j-th training step uploaded by the second device, and calculate the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device.
[0168] In the embodiment of the present invention, the first model training module 702 is further configured to: when the model training is completed, execute any one or more of the following:
[0169] Send a stop model training message to the second device;
[0170] Do not send the model parameters corresponding to the j-th training step to the second device.
[0171] In the embodiment of the present invention, the model training job distribution module 701 is further configured to: receive a model training job.
[0172] In the embodiment of the present invention, the model training job includes: a description of the data required for model training, data processing code, first model training code, and second model training code.
[0173] In the embodiment of the present invention, the first model training module 702 is specifically configured to calculate the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device in the following manner:
[0174] Calculate the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device;
[0175] Calculate the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step.
[0176] In the embodiment of the present invention, the first model training module 702 is further configured to:
[0177] Receive the number of training data used in the j-th training step uploaded by the second device;
[0178] The first model training module 702 is specifically configured to calculate the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device in the following manner:
[0179] According to the formula Calculate the global model parameter increment corresponding to the j-th training step;
[0180] Where Is the global model parameter increment corresponding to the j-th training step, #D ij Is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in model training, ΔW ij Is the model parameter increment corresponding to the j-th training step uploaded by the i-th second device.
[0181] In the embodiment of the present invention, the first model training module 702 is specifically configured to calculate the model parameter corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step in the following manner:
[0182] According to the formula Calculate the model parameter corresponding to the (j + 1)-th training step;
[0183] Where W j+1 Is the model parameter corresponding to the (j + 1)-th training step, W j Is the model parameter corresponding to the j-th training step, α is the learning rate, Is the global model parameter increment corresponding to the j-th training step.
[0184] In the embodiment of the present invention, the first model training module 702 is specifically configured to determine whether the model training is completed by using any one or more of the following methods:
[0185] When j is greater than or equal to the maximum number of training steps, it is determined that the model training is completed;
[0186] When j is less than the maximum number of training steps, it is determined that the model training is not completed.
[0187] In the embodiment of the present invention, the first model training module 702 is specifically configured to determine whether the model training is completed by using any one or more of the following methods:
[0188] When the difference between the averages of the loss functions corresponding to any two adjacent training steps from the (j−m + 1)-th training step to the j-th training step is less than or equal to a preset threshold, it is determined that the model training is completed;
[0189] When there are at least two adjacent training steps from the (j−m + 1)-th training step to the j-th training step for which the difference between the averages of the loss functions is greater than the preset threshold, it is determined that the model training is not completed.
[0190] In an embodiment of the present invention, the first model training module 702 is further configured to:
[0191] According to the formula calculate the average value of the loss function corresponding to the j-th training step;
[0192] wherein, is the average value of the loss function corresponding to the j-th training step, #D ij is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in the model training, and L ij is the value of the loss function corresponding to the j-th training step uploaded by the i-th device.
[0193] The specific implementation process of the above model training device is the same as that of the model training method in the foregoing embodiment, and will not be elaborated here.
[0194] See Figure 8 , another embodiment of the present invention provides a model training device (such as the above-mentioned second device), including:
[0195] A model training job receiving module 801, configured to receive some or all of the model training jobs sent by the first device;
[0196] A second model training module 802, configured to, for the j-th training step, receive the model parameters corresponding to the j-th training step sent by the first device, perform model training based on the model parameters corresponding to the j-th training step and some or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step, and upload the model parameter increment corresponding to the j-th training step to the first device.
[0197] In an embodiment of the present invention, the second model training module 802 is further configured to:
[0198] Receive a stop model training message sent by the first device.
[0199] In an embodiment of the present invention, the model training job includes: a description of the data required for model training, data processing code, and second model training code;
[0200] The second model training module 802 is specifically configured to implement model training using the following method to obtain a model parameter increment corresponding to the j-th training step based on the model parameters corresponding to the j-th training step and part or all of the model training job:
[0201] Execute the data processing code to obtain the training data corresponding to the j-th training step from the network element corresponding to itself according to the description of the data required for model training, process the training data corresponding to the j-th training step to obtain the training samples corresponding to the j-th training step, and execute the second model training code to implement model training based on the model parameters corresponding to the j-th training step and the training samples corresponding to the j-th training step to obtain the model parameter increment corresponding to the j-th training step.
[0202] In an embodiment of the present invention, the second model training module 802 is further configured to:
[0203] Execute the data processing code to obtain training data from the network element corresponding to itself according to the description of the data required for model training, process the training data to obtain training samples; execute the second model training code to implement model training based on the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step.
[0204] In an embodiment of the present invention, the second model training module 802 is further configured to execute any one or more of the following:
[0205] Upload the number of training data used in the j-th training step to the first device;
[0206] Upload the loss function value corresponding to the j-th training step to the first device.
[0207] The specific implementation process of the above model training device is the same as that of the model training method in the foregoing embodiment, and will not be described in detail here.
[0208] See Figure 9 , another embodiment of the present invention proposes a model training system, including:
[0209] The first device 901 is configured to determine a second device participating in model training according to a description of data required for model training in a model training job, send a part or all of the model training job to the second device; execute a first model training code in the model training job to implement, for the j-th training step, when the model training is not finished, send the model parameters corresponding to the j-th training step to the second device; receive the increment of the model parameters corresponding to the j-th training step uploaded by the second device, and calculate the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device.
[0210] The second device 902 is configured to receive a part or all of the model training job sent by the first device; for the j-th training step, receive the model parameters corresponding to the j-th training step sent by the first device, perform model training based on the model parameters corresponding to the j-th training step and the part or all of the model training job to obtain the increment of the model parameters corresponding to the j-th training step, and upload the increment of the model parameters corresponding to the j-th training step to the first device.
[0211] In an embodiment of the present invention, the first device 901 is further configured to: when the model training is finished, execute any one or more of the following:
[0212] Send a stop model training message to the second device;
[0213] Do not send the model parameters corresponding to the j-th training step to the second device;
[0214] The second device 902 is further configured to: receive the stop model training message sent by the first device.
[0215] In an embodiment of the present invention, the first device 901 is further configured to: receive a model training job.
[0216] In an embodiment of the present invention, the model training job includes: a description of data required for model training, data processing code, first model training code, and second model training code.
[0217] In an embodiment of the present invention, the first device 901 is specifically configured to calculate the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device in the following manner:
[0218] Calculate the global model parameter increment corresponding to the j-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device;
[0219] Calculate the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step.
[0220] In an embodiment of the present invention, the first device 901 is further configured to:
[0221] Receive the number of training data used in the j-th training step uploaded by the second device;
[0222] The second device 902 is further configured to perform any one or more of the following:
[0223] Upload the number of training data used in the j-th training step to the first device;
[0224] Upload the loss function value corresponding to the j-th training step to the first device;
[0225] The first device 901 is specifically configured to calculate the global model parameter increment corresponding to the j-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device in the following manner:
[0226] According to the formula Calculate the global model parameter increment corresponding to the j-th training step;
[0227] Where Is the global model parameter increment corresponding to the j-th training step, #D ij Is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in model training, ΔW ij Is the model parameter increment corresponding to the j-th training step uploaded by the i-th second device.
[0228] In an embodiment of the present invention, the first device 901 is specifically configured to calculate the model parameter corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step in the following manner:
[0229] According to the formula Calculate the model parameter corresponding to the (j + 1)-th training step;
[0230] Where W j+1 Is the model parameter corresponding to the (j + 1)-th training step, W j Is the model parameter corresponding to the j-th training step, α is the learning rate, Is the global model parameter increment corresponding to the j-th training step.
[0231] In an embodiment of the present invention, the first device 901 is specifically configured to determine whether model training is completed in any one or more of the following manners:
[0232] When j is greater than or equal to the maximum number of training steps, it is determined that the model training is completed;
[0233] When j is less than the maximum number of training steps, it is determined that the model training is not completed.
[0234] In the embodiment of the present invention, the first device 901 is specifically configured to determine whether the model training is completed by using any one or more of the following methods:
[0235] When the difference between the averages of the loss functions corresponding to any two adjacent training steps from the (j - m + 1)-th training step to the j-th training step is less than or equal to a preset threshold, it is determined that the model training is completed;
[0236] When there are at least two adjacent training steps in the (j - m + 1)-th training step to the j-th training step whose difference between the averages of the loss functions is greater than the preset threshold, it is determined that the model training is not completed.
[0237] In the embodiment of the present invention, the first device 901 is further configured to:
[0238] According to the formula Calculate the average value of the loss function corresponding to the j-th training step;
[0239] Wherein, Is the average value of the loss function corresponding to the j-th training step, #D ij Is the number of training data corresponding to the j-th training step uploaded by the i-th second device, N is the number of second devices participating in the model training, and L ij Is the loss function value corresponding to the j-th training step uploaded by the i-th device.
[0240] In the embodiment of the present invention, the second device 902 is specifically configured to implement obtaining the model parameter increment corresponding to the j-th training step by using the following method according to the model parameter corresponding to the j-th training step and part or all of the model training jobs:
[0241] Execute the data processing code to obtain the training data corresponding to the j-th training step from the network element corresponding to itself according to the description of the data required for the model training, process the training data corresponding to the j-th training step to obtain the training samples corresponding to the j-th training step, and execute the second model training code to implement model training according to the model parameter corresponding to the j-th training step and the training samples corresponding to the j-th training step to obtain the model parameter increment corresponding to the j-th training step.
[0242] In the embodiment of the present invention, the second device 902 is further configured to:
[0243] Execute the data processing code to obtain training data from the network element corresponding to itself according to the description of the data required for the model training, and process the training data to obtain training samples; execute the second model training code to perform model training according to the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step.
[0244] The specific implementation process of the above model training system is the same as that of the model training method in the foregoing embodiment, and will not be elaborated here.
[0245] Another embodiment of the present invention provides a prediction method, including:
[0246] Obtain the data required for prediction, and extract key features from the data required for prediction;
[0247] Input the key features into the model corresponding to the trained model parameters in any of the above model training methods, and output a predicted value.
[0248] Another embodiment of the present invention provides a prediction device, including a processor and a computer-readable storage medium, and instructions are stored in the computer-readable storage medium, and when the instructions are executed by the processor, any of the above prediction methods is implemented.
[0249] Another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above prediction methods are implemented.
[0250] Another embodiment of the present invention provides a prediction device, including:
[0251] A data acquisition module, configured to acquire the data required for prediction;
[0252] A key feature extraction module, configured to extract key features from the data required for prediction;
[0253] A prediction module, configured to input the key features into the model corresponding to the trained model parameters in any of the above model training methods, and output a predicted value.
[0254] The following lists several examples to illustrate the implementation process of the above prediction method. The examples listed are only for convenience of description and are not used to limit the protection scope of the embodiments of the present invention.
[0255] Example 4
[0256] This example illustrates the prediction method of the embodiments of the present invention by taking the prediction of the wireless network coverage based on the RAN coverage prediction model trained in Example 1 as an example. The method includes:
[0257] Step A, the OSS system obtains the data required for predicting the wireless network coverage, including: a list of cell IDs, and the cell configuration data, antenna configuration data, and measurement report data corresponding to each cell;
[0258] Step B, the OSS system extracts the key features of each cell from the cell configuration data, antenna configuration data, and measurement report data corresponding to each cell (as shown in Table 1);
[0259] Step C, the OSS system inputs the key features corresponding to each cell into the RAN coverage prediction model trained in Example 1, and outputs the predicted value of the radio coverage field strength RSRP of each cell;
[0260] Step D, the OSS system displays the predicted value of the radio coverage field strength RSRP of each cell to the user.
[0261] Example 5
[0262] This example illustrates the prediction method of the embodiments of the present invention by taking the prediction of the cell traffic (such as call volume) based on the RAN cell traffic prediction model trained in Example 2 as an example. The method includes:
[0263] Step A, the OSS system obtains the data required for predicting the cell traffic, including: a list of cell IDs, and the cell KPI data corresponding to each cell in the most recent two weeks;
[0264] Step B, the OSS system extracts the key features corresponding to each cell from the cell KPI data corresponding to each cell (as shown in Table 2);
[0265] Step C, the OSS system inputs the key features corresponding to each cell into the RAN cell traffic prediction model, and outputs the predicted value of the traffic of each cell;
[0266] Step D, the OSS system displays the predicted value of the traffic of each cell to the user.
[0267] Example 6
[0268] This example illustrates the prediction method of the embodiments of the present invention by taking the prediction of the optical module failure based on the wired BN optical module failure prediction model trained in Example 3 as an example. The method includes:
[0269] Step A: The OSS system obtains the data required for optical module fault prediction, including: a list of PTN device IDs, and the optical module monitoring data, optical module alarm data, and data collection time periods corresponding to each PTN device.
[0270] Step B: The OSS system extracts the key features corresponding to each PTN device from the optical module monitoring data, optical module alarm data, and data collection time periods corresponding to each PTN device (as shown in Table 3).
[0271] Step C: The OSS system inputs the key features corresponding to each PTN device into the optical module fault prediction model and outputs the predicted values indicating whether the optical modules corresponding to each PTN device have failed.
[0272] Step D: The OSS system displays the predicted values indicating whether the optical modules corresponding to each PTN device have failed to the user.
[0273] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some components or all components can be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0274] Although the embodiments disclosed in the embodiments of the present invention are as above, the content described is only an embodiment adopted for the convenience of understanding the embodiments of the present invention, and is not intended to limit the embodiments of the present invention. Any person skilled in the art within the field to which the embodiments of the present invention pertain may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the embodiments of the present invention. However, the scope of patent protection for the embodiments of the present invention shall still be subject to the scope defined by the appended claims.
Claims
1. A model training method, comprising: The first device determines a second device participating in model training according to the description of the data required for model training in the model training job, and sends part or all of the model training job to the second device; The first device executes the first model training code in the model training job to implement that for the j-th training step, when the model training is not finished, the first device sends the model parameters corresponding to the j-th training step to the second device; receives the increment of the model parameters corresponding to the j-th training step uploaded by the second device, and calculates the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device; Wherein, the first device is a central training engine deployed on the OSS; The second device is a distributed training engine deployed on a network element managed by the OSS.
2. The method according to claim 1, characterized in that, When the model training ends, the method further includes any one or more of the following: The first device sends a stop model training message to the second device; The first device does not send the model parameters corresponding to the j-th training step to the second device.
3. The method according to claim 1, wherein Before the first device determines a second device participating in model training according to the description of the data required for model training in the model training job, the method further includes: the first device receives the model training job.
4. The method according to any one of claims 1 to 3, characterized in that Wherein, The model training job includes: a description of the data required for model training, data processing code, first model training code, and second model training code.
5. The method according to any one of claims 1 to 3, characterized in that Wherein, Calculating the model parameters corresponding to the (j + 1)-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device includes: Calculating the global model parameter increment corresponding to the j-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device; Calculating the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step.
6. The method according to claim 5, wherein Before calculating the global model parameter increment corresponding to the j-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device, the method further includes: The first device receives the number of training data used in the j-th training step uploaded by the second device; Calculating the global model parameter increment corresponding to the j-th training step according to the increment of the model parameters corresponding to the j-th training step uploaded by the second device includes: According to the formula calculate the global model parameter increment corresponding to the j-th training step; Among them, is the global model parameter increment corresponding to the j-th training step, , is the number of training data corresponding to the j-th training step uploaded by the i-th second device, and N is the number of second devices participating in model training. is the model parameter increment corresponding to the j-th training step uploaded by the i-th second device.
7. The method according to claim 5, characterized in that, Wherein, Calculating the model parameters corresponding to the (j + 1)-th training step according to the global model parameter increment corresponding to the j-th training step includes: According to the formula calculate the model parameters corresponding to the (j + 1)-th training step; wherein, is the model parameter corresponding to the (j + 1)-th training step, is the model parameter corresponding to the j-th training step, is the learning rate, is the global model parameter increment corresponding to the j-th training step.
8. The method according to any one of claims 1 to 3, characterized in that Wherein, Determining whether the model training ends includes any one or more of the following: When j is greater than or equal to the maximum number of training steps, it is determined that the model training ends; When j is less than the maximum number of training steps, it is determined that the model training is not finished.
9. The method according to any one of claims 1 to 3, characterized in that Wherein, Determining whether the model training ends includes any one or more of the following: When the difference between the averages of the loss functions corresponding to any two adjacent training steps from the (j - m + 1)-th training step to the j-th training step is less than or equal to a preset threshold, it is determined that the model training ends; When the difference between the averages of the loss functions corresponding to at least two adjacent training steps from the (j - m + 1)-th training step to the j-th training step is greater than a preset threshold, it is determined that the model training has not ended.
10. The method according to claim 9, wherein wherein, According to the formula calculate the average value of the loss function corresponding to the j-th training step; wherein, is the average value of the loss function corresponding to the j-th training step, , is the number of training data corresponding to the j-th training step uploaded by the i-th second device, and N is the number of second devices participating in model training, is the loss function value corresponding to the j-th training step uploaded by the i-th device.
11. A model training method, comprising: The second device receives part or all of the model training jobs sent by the first device; For the j-th training step, the second device receives the model parameters corresponding to the j-th training step sent by the first device, performs model training according to the model parameters corresponding to the j-th training step and part or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step, and uploads the model parameter increment corresponding to the j-th training step to the first device; wherein, the first device is a central training engine deployed on the OSS; the second device is a distributed training engine deployed on the network element managed by the OSS.
12. The method according to claim 11, wherein The method further includes: The second device receives the stop model training message sent by the first device.
13. The method according to claim 11 or 12, characterized in that, wherein, the model training job includes: a description of the data required for model training, data processing code, first model training code, and second model training code; The performing model training according to the model parameters corresponding to the j-th training step and part or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step includes: executing the data processing code to obtain the training data corresponding to the j-th training step from the network element corresponding to itself according to the description of the data required for model training, processing the training data corresponding to the j-th training step to obtain the training samples corresponding to the j-th training step, and executing the second model training code to perform model training according to the model parameters corresponding to the j-th training step and the training samples corresponding to the j-th training step to obtain the model parameter increment corresponding to the j-th training step.
14. The method according to claim 11 or 12, characterized in that wherein, the model training job includes: a description of the data required for model training, data processing code, first model training code, and second model training code; after the second device receives the model training job sent by the first device, the method further includes: executing the data processing code to obtain training data from the network element corresponding to itself according to the description of the data required for model training, and processing the training data to obtain training samples; The performing model training according to the model parameters corresponding to the j-th training step and part or all of the model training jobs to obtain the model parameter increment corresponding to the j-th training step includes: executing the second model training code to perform model training according to the model parameters corresponding to the j-th training step and the training samples to obtain the model parameter increment corresponding to the j-th training step.
15. The method according to claim 13, wherein The method further includes any one or more of the following: The second device uploads the number of training data used in the j-th training step to the first device; The second device uploads the loss function value corresponding to the j-th training step to the first device.
16. A model training device, comprising a processor and a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and characterized in that, When the instruction is executed by the processor, the model training method according to any one of claims 1 to 15 is implemented.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the model training method according to any one of claims 1 to 15.
18. A model training device, comprising: A model training job distribution module, configured to determine a second device participating in model training according to a description of data required for model training in a model training job, and send part or all of the model training job to the second device; A first model training module, configured to execute the first model training code in the model training job to implement, for the j-th training step, when the model training is not completed, the first device sends the model parameters corresponding to the j-th training step to the second device; receive the model parameter increment corresponding to the j-th training step uploaded by the second device, and calculate the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device; Wherein, the first device is a central training engine deployed on the OSS; The second device is a distributed training engine deployed on a network element managed by the OSS.
19. The device according to claim 18, characterized in that, The first model training module is further configured to: When the model training is completed, execute any one or more of the following: The first device sends a stop model training message to the second device; The first device does not send the model parameters corresponding to the j-th training step to the second device.
20. A model training device, comprising: A model training job receiving module, configured to receive part or all of the model training job sent by the first device; A second model training module, configured to, for the j-th training step, the second device receives the model parameters corresponding to the j-th training step sent by the first device, perform model training based on the model parameters corresponding to the j-th training step and part or all of the model training job to obtain the model parameter increment corresponding to the j-th training step, and upload the model parameter increment corresponding to the j-th training step to the first device; Wherein, the first device is a central training engine deployed on the OSS; The second device is a distributed training engine deployed on a network element managed by the OSS.
21. A model training system, comprising: A first device, configured to determine a second device participating in model training according to a description of data required for model training in a model training job, and send part or all of the model training job to the second device; execute the first model training code in the model training job to implement, for the j-th training step, when the model training is not completed, send the model parameters corresponding to the j-th training step to the second device; receive the model parameter increment corresponding to the j-th training step uploaded by the second device, and calculate the model parameters corresponding to the (j + 1)-th training step according to the model parameter increment corresponding to the j-th training step uploaded by the second device; A second device, configured to receive part or all of a model training job sent by a first device; for the j-th training step, receive the model parameters corresponding to the j-th training step sent by the first device, perform model training based on the model parameters corresponding to the j-th training step and part or all of the model training job to obtain the model parameter increment corresponding to the j-th training step, and upload the model parameter increment corresponding to the j-th training step to the first device; wherein, the first device is a central training engine deployed on the OSS; the second device is a distributed training engine deployed on a network element managed by the OSS.
22. A prediction method, comprising: obtaining data required for prediction, and extracting key features from the data required for prediction; inputting the key features into the model training method according to any one of claims 1 to 15, and outputting a prediction value in the model corresponding to the trained model parameters.
23. A prediction device, comprising a processor and a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, characterized in that, When the instruction is executed by the processor, the prediction method according to claim 22 is implemented.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the prediction method according to claim 22 are implemented.
25. A prediction device, comprising: a data acquisition module, configured to acquire data required for prediction; a key feature extraction module, configured to extract key features from the data required for prediction; a prediction module, configured to input the key features into the model training method according to any one of claims 1 to 15, and output a prediction value in the model corresponding to the trained model parameters.
Citation Information
Patent Citations
Method for establishing data analysis model
CN111190487A
Multi-task federated learning method and system for large-scale federated learning tasks
CN117608814A