Fault prediction model generation method, electronic equipment and storage medium

By implementing federated learning in enterprise network service scenarios, sending model training instructions to base station devices and aggregating local model parameters to generate a global fault prediction model, it solves the problems of high failure prediction cost and complex model training in enterprise network service scenarios, and achieves low-cost and efficient fault prediction.

CN120151222APending Publication Date: 2025-06-13ZTE CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311714215.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Enterprise-oriented network service scenarios have challenges in network failure prediction and lightweight implementation. The prior art is difficult to achieve effective failure prediction without increasing hardware resources and storage costs.

Method used

By implementing federated learning between the network management server and the base station equipment, model training instructions are sent to all access base station equipment, local model parameters of the target base station equipment are received and aggregated, and a global fault prediction model is generated. This method does not need to store a large amount of data from base station devices on the network management server, but only requires a small amount of data aggregation capability to complete global model training.

Benefits of technology

It realizes low-cost and lightweight fault prediction model generation, which reduces the computing pressure of network management servers and the communication overhead of base station equipment, and improves the generalization ability and accuracy of fault prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151222A_ABST
    Figure CN120151222A_ABST
Patent Text Reader

Abstract

The invention provides a fault prediction model generation method, electronic equipment and a storage medium, and the method comprises the steps: sending a model training instruction to all access base station equipment, the model training instruction carrying a fault prediction initial model; receiving first local model parameters of the at least two target base station devices based on the model training instruction, wherein the first local model parameters are model parameters trained for a fault prediction initial model based on local data of the target base station; and performing aggregation processing on the first local model parameters of the at least two target base station devices to obtain first global model parameters, and sending the first global model parameters to the at least two target base station devices, so that the at least two target base station devices generate a global fault prediction model based on the first global model parameters. According to the method, the fault prediction model can be generated with low cost and light weight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and particularly to a method for generating a fault prediction model, an electronic device, and a storage medium. Background Art

[0002] The network service scenario for enterprises (To-Business, ToB) is extremely sensitive to network faults in business, and network fault prediction needs to be carried out in a timely manner to ensure network stability and reliability. Summary of the Invention

[0003] This application provides a method for generating a fault prediction model, an electronic device, and a storage medium, which are used to realize the generation of a fault prediction model with low cost and light weight.

[0004] An embodiment of this application provides a method for generating a fault prediction model, which is applied to a network management server. The method includes: sending a model training instruction to all access base station devices, where the model training instruction carries an initial fault prediction model; receiving first local model parameters of at least two target base station devices based on the model training instruction, where the first local model parameters are model parameters obtained by training the initial fault prediction model based on local data of the target base station; performing an aggregation process on the first local model parameters of the at least two target base station devices to obtain first global model parameters, and sending the first global model parameters to the at least two target base station devices, so that the at least two target base station devices respectively generate a global fault prediction model based on the first global model parameters.

[0005] An embodiment of this application provides a method for generating a fault prediction model, which is applied to a base station device. The method includes: receiving a model training instruction from a network management server, where the model training instruction carries an initial fault prediction model; in response to the model training instruction, performing model training on the initial fault prediction model based on local data to obtain first local model parameters; sending the first local model parameters to the network management server, where the network management server is used to perform an aggregation process on the first local model parameters of at least two of the base station devices received based on the model training instruction to obtain first global model parameters; receiving the first global model parameters from the network management server, and generating a corresponding global fault prediction model based on the first global model parameters.

[0006] An embodiment of this application provides an electronic device, including: one or more processors; a memory, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the methods in the embodiments of this application.

[0007] An embodiment of the present application provides a storage medium storing a computer program, which when executed by a processor implements any of the methods in the embodiments of the present application.

[0008] According to the method for generating a fault prediction model in an embodiment of the present application, the network management server can send the initial fault prediction model to all access base station devices through a model training instruction, and receive the first local model parameters of at least two target base station devices based on the model training instruction. Then, aggregate the first local model parameters of the at least two target base station devices to obtain the first global model parameter, and send the first global model parameter to the at least two target base station devices, so that the at least two base station devices can respectively generate a global fault prediction model based on the first global model parameter; according to this method, the network management server does not need to store the relevant data of the base station devices required for training, thus saving storage costs, and only needs a small amount of data aggregation ability to complete global model training and obtain global model parameters, which is beneficial to realizing the low cost and light weight of the network management server, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0009] More descriptions about the above embodiments and other aspects of the present application and their implementation manners are provided in the drawings description, the specific implementation manner and the claims. Description of the Drawings

[0010] Figure 1 The application scenario diagram showing the fault prediction method and device provided by the embodiment of the present application.

[0011] Figure 2 The schematic diagram showing the federated learning process of the network management server and the base station devices it manages provided by the embodiment of the present application.

[0012] Figure 3 The flowchart of a method for generating a fault prediction model provided by the embodiment of the present application.

[0013] Figure 4 The flowchart of a method for generating a fault prediction model provided by the embodiment of the present application.

[0014] Figure 5 The detailed flowchart showing the method for generating a fault prediction model provided by the exemplary embodiment of the present application.

[0015] Figure 6 The schematic diagram showing the model training process of multiple categories provided by the exemplary embodiment of the present application.

[0016] Figure 7 The block diagram of a device for generating a fault prediction model provided by the embodiment of the present application.

[0017] Figure 8A block diagram of a fault prediction model generation device provided by an embodiment of the present application.

[0018] Figure 9 It is a structural diagram showing an exemplary hardware architecture of an electronic device capable of implementing the methods and devices according to the embodiments of the present invention. Detailed implementation manners

[0019] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined arbitrarily with each other.

[0020] In related scenarios, a fault prediction model for network fault prediction can be obtained through pre-training.

[0021] In the related art, the fault prediction model can use the ToB network management server as an edge node for local training, and the trained model parameters are uploaded to the cloud separate server aggregation point to form a global model. This method is suitable for scenarios with sufficient hardware resources and no low-cost constraints, and it is usually difficult for enterprises facing ToB services to bear; in the related art, the fault prediction model can also perform fault prediction based on federated learning. This method is usually strongly bound to device service data and requires specific data of professional devices for specific optimization, and it is not applicable to the fault prediction of base station communication devices.

[0022] Federated learning (Federated machine learning / Federated Learning, FL) is a distributed machine learning technology or machine learning framework. The goal of federated learning is to enable all participating parties to jointly model on the basis of ensuring data privacy, security, legality, and compliance, so as to carry out efficient model learning and improve the effect of artificial intelligence (AI) models.

[0023] Based on this, it is necessary to provide a fault prediction method applicable to the ToB scenario, enabling the ToB network management to achieve the fault prediction of base station communication devices in a lighter and lower-cost manner.

[0024] Figure 1 Shows an application scenario diagram of the fault prediction method and device provided by the embodiments of the present application.

[0025] As Figure 1As shown in the figure, the application scenarios of the embodiments of the present application may include a network management server 101, multiple base station devices 102, and a network 103. The network management server 101 is located at the network management level, and the base station devices 102 are located at the base station level. A connection is established between the network management server 101 and each base station device 102 through a communication link provided by the network 103. The communication link may include links of various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0026] In some ToB scenarios, such as industrial parks, intelligent industrial bases, etc., the network management server 101 may be a server that provides various services by a ToB network management server, and the network management server 101 may be a layer 2 network management device.

[0027] It should be noted that the fault prediction methods and devices provided by the embodiments of the present application include the methods and devices applied to the network management server 101, and the fault prediction methods and devices applied to each base station device 102. The network management server 101 may also be a server or a server cluster different from the network management server 101 and capable of communicating with the network management server 101 and / or the base station devices 102. Correspondingly, the fault prediction device provided by the embodiments of the present application may also be set in a server or a server cluster different from the network management server 101 and capable of communicating with the network management server 101 and / or the base station devices 102.

[0028] It should be understood that the fault prediction method of the embodiments of the present application may be implemented by a processor calling computer-readable program instructions stored in a memory. Figure 1 The numbers of the network management server and the base station devices in [description] are merely illustrative. According to actual needs, any number of network management servers and base station devices may be provided.

[0029] Figure 2 The figure shows a schematic diagram of the federated learning process of the network management server provided by the embodiments of the present application and the base station devices it manages. Figure 2 Same as Figure 1 the same or equivalent structures in [description] are denoted by the same reference numerals.

[0030] In Figure 2 In [description], based on the distributed computing power of multiple base station devices 102, each base station device 102 locally trains a model through a training data set, and each base station device 102 uploads the model parameters of its respective trained model, such as the gradient parameters ω1, ω2, ω3,..., ωn of n models, where n is an integer greater than or equal to 1, and each gradient parameter is the model parameter of a model trained by a base station device 102, to the network management server 101; the network management server 101 aggregates the received gradient parameters of the models to obtain all models, and the model parameters (ω of the global model *(t)) is broadcast to multiple base station devices 102, and each base station device 102 obtains a global model according to the model parameters of the global model and uses the global model to perform fault prediction locally at the base station device.

[0031] In the embodiments of the present application, model training is performed locally at the base station device through the distributed computing power of the base station device 102, which is beneficial to greatly reducing the computing pressure on the network management server 101; moreover, since the base station device 102 uploads the model parameters after local training at the base station device, it is beneficial to reduce the communication overhead and pressure from the base station device 102 to the network management server 101; the network management server 101 does not need to store a large amount of relevant data of the base station device 101 required for training, and only needs a small amount of data aggregation ability to complete the global model training and obtain the model parameters of the global model, which is beneficial to realizing the low cost and lightweight of the network management server 101, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0032] In some embodiments, there can be multiple gradient parameters corresponding to the models trained by each target base station device. Through the Transmission Control Protocol (TCP) channel, the gradient parameters are transmitted to the network management server. The network management server can perform aggregation calculation and analysis on the model parameters uploaded by different base station devices, rather than analyzing and training the actual acquisition data of the base stations, which can ensure lightweight data storage and low Input / Output (I / O) occupancy. The model parameters obtained through a certain aggregation algorithm are then broadcast to all base stations, and the model parameters generated during the storage process will be stored in the network management for easy iterative update of the model.

[0033] The fault prediction model generation method in the embodiments of the present application involves collaborative interactions at multiple levels. Specifically, it mainly involves the interaction between the base station level and the network management level. It should be understood that the multiple base station devices 102 in the embodiments of the present application can be multiple base station devices with certain computing power and storage capabilities, such as 5th Generation Mobile Communication Technology (5G) base station devices, and the network management server 101 can be a network management server for ToB and can take over 5G base station devices.

[0034] In a first aspect, the embodiments of the present application provide a fault prediction model generation method, which can be applied to a network management server.

[0035] Figure 3 This is a flowchart of a fault prediction model generation method provided by the embodiments of the present application. Refer to Figure 3 , the method may include the following steps.

[0036] S310. Send a model training instruction to all base station devices connected, where the model training instruction carries an initial fault prediction model.

[0037] In this step, the number of all base station devices connected is T, where T is an integer greater than or equal to 2. The target base station device can be all base station devices within the jurisdiction of the network management server.

[0038] S320. Receive first local model parameters of at least two target base station devices based on the model training instruction. The first local model parameters are model parameters obtained by training the initial fault prediction model based on local data of the target base stations.

[0039] In this step, the number of at least two target base station devices is T1, where T1 is an integer greater than or equal to 1 and less than or equal to T. In some scenarios, for any base station device connected, during the process of model training according to the received model training instruction, if there are situations such as going out of service, being offline (disconnected from the network management), and / or station outage (service interruption caused by power failure), the connection between the base station device and the network management server will be disconnected, resulting in the base station device stopping model training and being unable to report local model parameters. At this time, the target base station devices will not include this base station device. In other scenarios, during the process of model training on the local side of any base station device, if the computing power resources of the base station device are insufficient to support continued model training, it will also cause the base station device to stop model training and be unable to report local model parameters (since the model training is not completed, the trained model parameters cannot be obtained). Based on the situations described in the above scenarios, the phenomenon of T1 being less than T will occur.

[0040] S330. Aggregate the first local model parameters of at least two target base station devices to obtain first global model parameters, and send the first global model parameters to at least two target base station devices so that the at least two target base station devices can respectively generate a global fault prediction model based on the first global model parameters.

[0041] In this step, the model parameters of the global model can be aggregated through a certain algorithm. As an example, an aggregation algorithm can be used to aggregate the first local model parameters of multiple (at least two) target base station devices received to form global model parameters, which are called the first global model parameters.

[0042] The aggregation algorithm can be, for example, the Stochastic Controlled Averaging for Federated Learning (SCAFFOLD) algorithm, or it can also be an algorithm for weighted aggregation or averaging. For example, the weighted sum of multiple first local model parameters is calculated to obtain the first global model parameter; or, the average value of multiple first local model parameters is calculated and then used as the first global model parameter.

[0043] In this step, the network management server can send the first global model parameter to at least two target base station devices by means of broadcasting; or it can also adopt the method of sending one by one to send the first global model parameter to each target base station device respectively. The specific sending method can be customized according to actual needs, and the embodiments of the present application do not make specific limitations.

[0044] According to the fault prediction method of the embodiments of the present application, the network management server can send the fault prediction initial model to all access base station devices through a model training instruction, and receive the first local model parameters of at least two target base station devices based on the model training instruction, and then perform aggregation processing on the first local model parameters of at least two target base station devices to obtain the first global model parameter, and send the first global model parameter to at least two target base station devices, so that at least two base station devices can respectively generate a global fault prediction model based on the first global model parameter; according to this method, the network management server does not need to store the relevant data of the base station devices required for training, thus saving storage costs, and only needs a small amount of data aggregation ability to complete the global model training and obtain the global model parameter, which is beneficial to realizing the low cost and light weight of the network management server, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0045] In some embodiments, before step S310, the method further includes the following steps:

[0046] S11, determine the computing power of each target base station device according to the computing power resources of each target base station device obtained in advance.

[0047] In this step, the computing power resources include but are not limited to at least one of the following resource items: the CPU, memory, and storage bandwidth of the target base station device. Among them, the storage bandwidth refers to the amount of information accessed by the memory per unit time, and is also called the number of bits or bytes read / written by the memory per unit time.

[0048] In some embodiments, the computing power of each target base station device can be obtained by conversion according to its own computing power resources. As an implementation, the computing power parameter values of at least one computing power resource of each target base station device can be input into a pre-trained computing power measurement model to obtain the computing power (computing ability) evaluation value of the corresponding target base station device output by the computing power measurement model. As an example, the computing power parameter values of the computing power resources can include: CPU frequency, number of CPU cores, storage speed of the memory, storage capacity of the memory, storage bandwidth value, etc.

[0049] In some embodiments, each target base station device can be classified according to the computing power evaluation value. As an example, the computing power levels include: high-level computing power, medium-level computing power, and low-level computing power. The computing ability of the corresponding target base station device is characterized by the computing power level.

[0050] Exemplarily, the target base station devices whose computing power evaluation values are greater than or equal to the first evaluation value threshold are used as base station devices with high-level computing power; the target base station devices whose computing power evaluation values are less than the first evaluation value threshold and greater than or equal to the second evaluation value threshold are used as base station devices with medium-level computing power; the target base station devices whose computing power evaluation values are less than the second evaluation value threshold are used as base station devices with low-level computing power; where the first evaluation threshold is greater than the second evaluation threshold.

[0051] S12. Send a data collection instruction to each target base station device. The data collection instruction includes the number of categories of local data determined according to the computing ability of each target base station device, and is used to instruct each target base station device to collect the corresponding number of categories of local data.

[0052] In this step, first, the network management server can pre-configure the service guarantee level target based on the Service Level Agreement (SLA), and configure the threshold of the service quality guarantee target parameter corresponding to the service guarantee level target, which is simply called the guarantee target set based on SLA. This guarantee target can be characterized by the corresponding 5G Quality Of Service Identifier (5QI) parameter. As an example, if the SLA is not configured, the default 5QI parameter value can be used to characterize the service quality guarantee target parameter.

[0053] Then, the key data recording function or key data collection function of the target base station device can be enabled on the network management server, which is abbreviated as "Black Box". Among them, the black box is a technology that records data in different dimensions according to the actual occurrence time of abnormal points and analyzes the cause of problems in real time based on the recorded data. It is used to record and collect key data when the base station device has network service anomalies and equipment failures. That is to say, the black box can be used to record and store key data at the terminal device (Users Experience, UE) level when the target base station device fails to meet the guarantee objectives set by the SLA or a device failure alarm occurs. Exemplarily, the key data includes at least one of the following: measurement index data, performance index data, and log data.

[0054] Considering that the content collected by the black box and the storage process consume system resources (CPU, memory, storage bandwidth, etc.) and the stored content is relatively large, in order to conveniently control the impact of the black box on services after startup, hierarchical control can be introduced to determine the content collected by the black box according to the level, so as to adjust the consumption of system resources and reduce the impact on services.

[0055] As an example, the levels of the black box corresponding to low-level computing power include the first-level black box; the levels of the black box corresponding to medium-level computing power include the first-level black box and the second-level black box; the levels of the black box corresponding to high-level computing power include the first-level black box, the second-level black box, and the third-level black box.

[0056] Among them, for the first-level black box: the data collected through the black box technology (abbreviated as black box content) includes measurement index data, which supports carrier / cell level / UE level measurement index statistics, the collection time is at the second level granularity, and the data volume is the smallest; for the second-level black box, the black box content includes performance data, the collection time is at the second level granularity, and the data volume is medium; for the third-level black box, the black box content includes log data, the collection time is at the second level granularity, and the data volume is the largest. That is to say, the higher the computing power of the base station device, the more types and quantities of key data that can be collected.

[0057] That is to say, if the computing power of the target base station device corresponds to low-level computing power, the number of categories of local data that the target base station device needs to collect is 1, specifically including: measurement index data; if the computing power of the target base station device corresponds to medium-level computing power, the number of categories of local data that the target base station device needs to collect is 2, specifically including: measurement index data and performance data; if the computing power of the target base station device corresponds to high-level computing power, the number of categories of local data that the target base station device needs to collect is 3, specifically including: measurement index data, performance data, and log data.

[0058] Further, when the number of categories of local data that the target base station device needs to collect is 1, only the key measurement index data can be collected. When the number of categories of local data that the target base station device needs to collect is 2, only the key measurement index data and the key performance data can be collected. When the number of categories of local data that the target base station device needs to collect is 3, only the key measurement index data, the key performance data, and the key log data can be collected. Among them, the key measurement index data, the key performance data, and the key log data are the data in the measurement index data, the performance data, and the log data that are determined to play an important role in data analysis and problem location, respectively.

[0059] Through the above steps S11 - S12, the network management server can determine the computing power of each target base station device according to the computing power resources of each target base station device, so as to control each target base station device to turn on the key data collection function of the base station black box, and instruct each target base station device to collect the local data of the corresponding number of categories as the model training data.

[0060] In the embodiment of the present application, the network management server can set the data to be collected by the target base station device using the black box. The setting content includes the specific data types and collection content to be collected, and dynamically controls the data set of the local training of the base station based on the policies and algorithms of the base station service characteristics.

[0061] In some embodiments, before step S320, the method further includes the following steps:

[0062] S21, obtain at least one traffic evaluation index of each target base station device.

[0063] In this step, the traffic evaluation index is used to evaluate the size of the traffic carried by the target base station device. As an example, the traffic evaluation index includes but is not limited to at least one of the following: the proportion of service connections (denoted as c), the service SLA guarantee index (denoted as p), and the consumption rate of the service level objective (SLO) error budget (denoted as v).

[0064] As an example, the way to obtain the traffic evaluation index can be: each target base station device regularly reports the index value of the traffic evaluation index, so as to realize the real-time observation of the traffic of each target base station device by the network management service.

[0065] S22, determine the corresponding number of training rounds for each target base station device based on at least one traffic evaluation index of each target base station device and the training round comparison table.

[0066] Specifically, the training round comparison table is used to indicate the correspondence between the traffic volume evaluation metrics and the training rounds. By querying the training round comparison table based on the traffic volume evaluation metrics, the training rounds corresponding to at least one traffic volume evaluation metric of each target base station device can be obtained.

[0067] In this step, the correspondence can be a preset mapping relationship between at least one traffic volume evaluation metric and the training rounds. Based on this mapping relationship and at least one traffic volume evaluation metric, the training rounds of each target base station device can be determined.

[0068] S23. Send the corresponding training rounds to at least two target base station devices. The training rounds are used to represent the number of times of model training for the initial fault prediction model.

[0069] In this step, when the training round R of a certain target base station device is equal to 0, it indicates that the traffic volume of this target base station device is heavy, and it can be considered that this target base station device does not have enough resources for model training. Therefore, it can not participate in the current model training and parameter reporting, and the network management server can stop collecting the model parameters of this base station. Only after the selected base stations complete the training will they upload their respective model parameters.

[0070] Exemplarily, the network management server can send the corresponding training rounds to the target base station devices with training rounds greater than zero through a separate instruction each time.

[0071] Exemplarily, before step S310, the training rounds corresponding to at least one traffic volume evaluation metric of each target base station device can be calculated through the above steps S21 - S22, and a corresponding model training instruction can be generated for each target base station device with a training round greater than zero; the model training instruction carries: the training round of this target base station device and the initial fault prediction model; then, in step S310, each generated model training instruction (which not only carries the initial fault prediction model but also the training round of the corresponding target base station device) can be sent to the corresponding target base station device. After step S310, the training rounds corresponding to at least one traffic volume evaluation metric of each target base station device can be recalculated through the above steps S21 - S22 again, so as to be used to re - determine each target base station device with a training round greater than zero, and through a separate instruction, send the corresponding training rounds to the re - determined target base station devices with training rounds greater than zero.

[0072] In this embodiment, the network management server can combine the policies of the base station services, observe the traffic volume carried on each target base station device in real time, and regularly determine the number of training rounds for each target base station device according to the traffic volume evaluation metrics reported by each target base station device. The base stations with a training round number equal to zero may not participate in the current model training and parameter reporting, thereby reducing the communication frequency of interacting model parameters between the target base station devices and the network management server, and achieving the effect of dynamically adjusting the number of model training participants according to the resource limitations of the target base station devices.

[0073] In the embodiment of the present application, considering that in federated learning, before the system reaches the target accuracy, many rounds of communication are required between the edge devices (base stations) and the federated learning server (network management server), that is, the base stations need to upload the trained model of each round to the network management server. Since each model may contain millions of parameters, if communication is performed frequently, plus waiting for the responses of all participants, it will cause a large communication overhead and become a bottleneck in model training. Therefore, how to reduce the communication cost of federated learning has become a crucial issue. For this, the business characteristics of the base stations can be combined, and in the process, through the method of the above embodiment, the communication frequency of interacting model parameters between the target base station devices and the network management server can be reduced, and the number of participants can be dynamically adjusted according to the resource limitations of the target base station devices.

[0074] In some embodiments, the training round comparison table is used to represent the corresponding relationship between the traffic volume evaluation value and the training round; the above step S22 may specifically include the following steps.

[0075] S31, perform weighted fusion on multiple traffic volume evaluation metrics of each target base station device to obtain the traffic volume evaluation value corresponding to each target base station device.

[0076] Specifically, the corresponding relationship between the traffic volume evaluation metric and the training round can be represented by the following expression (1):

[0077] R(Rounds) = N - N*(k1*c + k2*p + k3*v) (1)

[0078] In the above expression (1), R(Rounds) is the number of training rounds, and the value of R is the floor value of the value of N - N*(k1*c + k2*p + k3*v); k1, k2, and k3 are preset weight values, N is a hyperparameter, and the values of c, p, and v are all greater than or equal to 0 and less than or equal to 1. The meanings of c, p, and v are as described in the above steps and will not be elaborated here. As an example, k1 is 0.5, k2 is 0.3, and k3 can be 0.2. The embodiment of the present application does not make specific limitations.

[0079] It should be understood that the value of N, and the values of k1, k2, and k3 can all be customarily set and adjusted according to the training effect in the actual application scenario, and the embodiments of the present application do not make specific limitations.

[0080] S32. According to the traffic volume evaluation value of each target base station device and the corresponding relationship between the traffic volume evaluation value and the training round, determine the training round corresponding to each target base station device, where the training round is inversely proportional to the corresponding traffic volume evaluation value.

[0081] Exemplarily, if the traffic volume evaluation value of the target base station device is higher, it can be understood that the business of this base station device is busier, the resource utilization rate is higher, the quality of the training dataset is higher at this time, and the number of training rounds can be less; on the contrary, it means that the quality of the dataset is average, and the number of training rounds can be more.

[0082] In this embodiment, the network management server can, every once in a while, calculate the training round of each target base station device through weighted calculation of the traffic volume evaluations collected during this period, so as to update the target base station devices participating in the model training this time, which is conducive to dynamically updating the communication frequency of the interaction model parameters between the target base station device and the network management server according to the actual business situation of the target base station device, and dynamically adjusting the number of model training participants (target base station devices with a training round greater than zero).

[0083] In some embodiments, another way to reduce the communication frequency of the interaction model parameters between the target base station device and the network management server can be: after the target base station device performs multiple rounds of training for a predetermined number of training rounds, report the model parameters to the network management server.

[0084] In some embodiments, after aggregating the first local model parameters of at least two target base station devices in step S330 to obtain the first global model parameter, the method may further include the following steps: in the case that the global fault prediction model does not reach the training end condition, perform at least one global model update until the updated global fault prediction model reaches the training end condition. In each global model update, perform the following operations: determine the target base station devices participating in this global model update; send a model update instruction to the determined target base station devices; receive the second local model parameters of at least two of the determined target base station devices based on the model update instruction; aggregate the second local model parameters of at least two of the determined target base station devices to obtain a second global model parameter, and send the second global model parameter to at least two of the determined target base station devices, so that at least two of the determined target base station devices respectively generate a new global fault prediction model based on the second global model parameter.

[0085] Exemplarily, the training end condition may be at least one of the following: the global fault prediction model is a converged model, and the model accuracy has reached the accuracy requirement. For example, the model accuracy is greater than or equal to a predetermined accuracy threshold, and the accuracy threshold can be set according to the actual federated learning scenario and specific federated learning requirements, and no limitation is imposed thereon.

[0086] In this embodiment, through the interaction of model parameters between the network management server and the target base station device for multiple times, until the training end condition is met, the global model update can be stopped.

[0087] In some embodiments, sending a model update instruction to the determined target base station device may include: in the model update instruction, carrying the training round of the current global model update, so that at least two determined target base station devices obtain second local model parameters after performing model training for the training round of the current global model update respectively.

[0088] Specifically, the second local model parameter is the local model parameter obtained by the target base station device participating in the current global model update after performing model training for the training round of the current global model update.

[0089] Exemplarily, the training round of each target base station device among multiple target base station devices can be calculated according to the above steps S21 and S22, and the target base station devices with a training round greater than zero are used as the target base station devices participating in the current global model update.

[0090] In this embodiment, the training round of the current global model update is carried in the model update instruction. The model update instruction is used to instruct the target base station device participating in the current global model update to report the second local model parameter after performing model training for the corresponding training round, that is, the local model parameter obtained after performing model training for the training round of the current global model update.

[0091] In the embodiment of the present application, the network management server can use federated learning to perform multi-base station node distributed local learning training at the base station layer, train and update their respective local models with the diverse local key data sets of multiple base station device nodes, and then upload the local model parameters (such as gradient parameters) to the network management level (network management server) for weighted aggregation to form global model parameters, and then broadcast and distribute the model parameters to all base station device nodes at the base station layer. The base station device uses the global model to predict the faults of the device.

[0092] In some embodiments, the network management server can receive the training data uploaded by multiple target base station devices, and perform sample learning based on all the received training data and then perform fault prediction.

[0093] Among them, the data volume of the training data uploaded by each base station device is less than or equal to the first data volume threshold, and the sum of the data volumes of all the training data is less than or equal to the first total data volume threshold. That is to say, the network management server can collect the data of all base stations to the network management server, perform few-shot learning, and then perform fault prediction. Moreover, it is required that the data volume cannot be too large. Therefore, the data collected on the base station side must be key, high-quality, and easy-to-extract feature data, such as data collected based on black box technology or more accurate UE-level end-user portrait data. After these data are uploaded to the network management server, they can be stored uniformly. Exemplarily, each base station device can upload a data volume greater than or equal to a predetermined number of days (for example, data collected within 10-15 days). The sum of the data volumes uploaded by multiple base station devices is less than the first total data volume threshold, and the first total data volume threshold is the data volume that the computing power of the network management server can bear. Moreover, the data upload time occupies bandwidth, and the occupation of this bandwidth can be other time periods outside the predetermined peak period to avoid the business peak period.

[0094] According to the fault prediction model generation method of the embodiments of the present application, the network management server can send the initial fault prediction model to all access base station devices through a model training instruction, and receive the first local model parameters of at least two target base station devices based on the model training instruction. Then, aggregate the first local model parameters of at least two target base station devices to obtain the first global model parameters, and send the first global model parameters to at least two target base station devices, so that at least two base station devices can generate a global fault prediction model based on the first global model parameters respectively; according to this method, the network management server does not need to store the relevant data of the base station devices required for training, thus saving storage costs, and only needs a small amount of data aggregation ability to complete the global model training and obtain the global model parameters, which is beneficial to realizing the low cost and light weight of the network management server, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0095] In a second aspect, the embodiments of the present application provide a fault prediction model generation method, which can be applied to base station devices.

[0096] Figure 4 It is a flowchart of a fault prediction model generation method provided by the embodiments of the present application. Refer to Figure 4 and the method may include the following steps.

[0097] S410, receive a model training instruction from the network management server, where the model training instruction carries an initial fault prediction model.

[0098] S420, in response to the model training instruction, perform model training on the initial fault prediction model based on local data to obtain first local model parameters.

[0099] S430. Send the first local model parameter to the network management server, which is used to aggregate the first local model parameters of at least two base station devices received based on the model training instruction to obtain the first global model parameter.

[0100] S440. Receive the first global model parameter from the network management server and generate a corresponding global fault prediction model based on the first global model parameter.

[0101] According to the fault prediction model generation method of the embodiments of the present application, the base station device can receive the model training instruction from the network management server, obtain the initial fault prediction model in the model training instruction, perform model training on the initial fault prediction model based on local data at the local of the base station device, and send the trained first local model parameter to the network management server, and receive the first global model parameter obtained by aggregating the first local model parameters of at least two base station devices by the network management server, so as to generate a corresponding global fault prediction model based on the first global model parameter. According to this method, the distributed computing power of different base station devices (network devices at the edge nodes) can be used for feature extraction and model training, and then the trained local model parameters are reported, so that the network management server does not need to store the relevant data of the base station devices required for training, saving storage costs, and only a small amount of data aggregation ability is required to complete the global model training and obtain the model parameters of the global model, which is beneficial to the low cost and lightweight of the network management server, and the generalization ability of the trained fault prediction model is stronger and the accuracy is higher.

[0102] In some embodiments, before performing model training on the initial fault prediction model based on local data in step S420, the method further includes the following steps.

[0103] S41. Receive the data collection instruction from the network management server; S42. In response to the data collection instruction, start collecting local data, and the number of categories of the local data matches the computing power of the base station device itself.

[0104] In this step, the base station device can, in response to the data collection instruction, start the data collection function of the local data at the local of the base station device. The number of categories of the local data matches the computing power of the base station device itself, and generate the training data of the local fault prediction model according to the collected local data, so as to perform model training on the initial fault prediction model according to the training data.

[0105] As an example, if the computing power of the base station device itself corresponds to low-level computing power, the number of categories of local data that the target base station device needs to collect is 1, specifically including: measurement index data; if the computing power of the target base station device corresponds to medium-level computing power, the number of categories of local data that the target base station device needs to collect is 2, specifically including: measurement index data and performance data; if the computing power of the target base station device corresponds to high-level computing power, the number of categories of local data that the target base station device needs to collect is 3, specifically including: measurement index data, performance data, and log data; for details, refer to the description of the above embodiments, which will not be elaborated here.

[0106] Further, when the number of categories of local data that the target base station device needs to collect is 1, only the key measurement index data can be collected; when the number of categories of local data that the target base station device needs to collect is 2, only the key measurement index data and key performance data can be collected; when the number of categories of local data that the target base station device needs to collect is 3, only the key measurement index data, key performance data, and key log data can be collected. Among them, the key measurement index data, key performance data, and key log data are the data in the measurement index data, performance data, and log data that are determined to play an important role in data analysis and problem location, respectively.

[0107] As an example, taking the prediction of quality degradation faults caused by external interference as an example, the key measurement indexes (which can also be abbreviated as measurement indexes) include but are not limited to at least one of the following: data such as longitude and latitude GNSS, tracking area TA, serving cell RSRP value, serving neighbor cell RSRP, serving cell SINR, etc.; the key performance indexes (which can also be abbreviated as performance indexes) include KPIs in different dimensions, such as the number of RBs occupied by the UE, uplink and downlink PRB utilization rates, single-user Bler index, cell slice uplink and downlink packet loss rates, etc.; the key log data (which can also be abbreviated as log data) includes base station operation period data, operation logs, security logs, audit logs, etc.

[0108] In some embodiments, the priority of the key measurement index data is higher than that of the key performance data, and the priority of the key performance data is higher than that of the key log data. Therefore, when the computing power of the base station device itself is low, the relevant data with higher priority is preferentially collected.

[0109] In this embodiment, the base station side can use the black box technology to hierarchically record the measurement indexes, performance indexes, log data, etc. when the fault occurs as the training data set. The black box technology can perform real-time data analysis and tagging, which can improve the accuracy of the training model and the convergence speed of the global model.

[0110] In some embodiments, the model structure of the initial fault prediction model includes: a convolutional neural network model and a long short-term neural network model connected in series; when the number of categories of local data is greater than 1, the training data includes multiple types of training data. In this embodiment, the step of training the initial fault prediction model based on local data in S420 to obtain the first local model parameters may specifically include: for the local data of each category, using the convolutional neural network model to extract features to obtain the first features corresponding to the local data of each category; using the long short-term neural network to process the first features of the local data of each category respectively to obtain the second features corresponding to the local data of each category; fusing the second features of the local data of each category to obtain fused features; and performing model training based on the fused features to obtain the first local model parameters.

[0111] Exemplarily, for the training data of the i-th type, using the convolutional neural network model to extract features to obtain the first features of the training data of the i-th type, where i is greater than or equal to 1 and less than or equal to the number of categories; using the long short-term neural network to process the first features to obtain the second features of the training data of the i-th type; fusing the second features of the training data of each type to obtain fused features; and performing model training based on the fused features to obtain the local fault prediction model.

[0112] In this embodiment, the Convolutional Neural Network (CNN) is used to extract data features from the spatial dimension, and the Long Short Term Memory networks (LSTM) is used to obtain correlations from the time dimension using time-series data; the base station device can adopt the combined method of cascading CNN + LSTM and overall parallelism of multiple types of data sources for local training, to make more accurate and efficient predictions of faults from both spatial and time dimensions, and can flexibly allocate three different types of training data with different priorities based on the dynamic scheduling algorithm of the base station service characteristics (traffic volume).

[0113] In some embodiments, step S420 may specifically include: S51, in response to a model training instruction, determining the number of training rounds, where the number of training rounds comes from the network management server and is determined by the network management server based on at least one traffic volume evaluation index of the base station device and the training round comparison table for the base station device; S52, after performing model training for the number of training rounds on the initial fault prediction model, obtaining the first local model parameters.

[0114] In this embodiment, if the traffic evaluation index values of different base station devices are different, the number of training rounds between different base station devices is different. If the number of training rounds is zero, the number of base stations participating in each training may also be different, so that only the base station devices with non-zero rounds perform model training for the corresponding number of training rounds. Since the traffic that a base station device can carry is proportional to the amount of resources included in the base station, in this embodiment, it is possible to dynamically adjust the data volume of model training participants according to the limitations of base station resources.

[0115] In some embodiments, after step S430, the method further includes: S61, receiving a model update instruction from a network management server, where the model update instruction is an instruction generated by the network management server when the global fault prediction model does not reach the training end condition; S62, in response to the model update instruction, continuing to perform model training on the initial fault prediction model based on local data to obtain second local model parameters; S63, sending the second local model parameters to the network management server, where the network management server is used to perform aggregation processing on the second local model parameters of at least two base station devices to obtain second global model parameters; S64, receiving the second global model parameters generated by the network management server; S65, generating a new global fault prediction model according to the second global model parameters.

[0116] In this embodiment, the base station device can perform model training for the corresponding number of training rounds according to the number of training rounds of the current global model update carried in the model update instruction, and then report the second local model parameters, that is, report the local model parameters obtained after model training for the number of training rounds of the current global model update, so as to realize the model update of the global model.

[0117] In the embodiments of the present application, multiple base station devices in the base station layer can be used as distributed edge computing nodes, and a deep learning algorithm with a model structure in which CNN and LSTM are connected in series is used to train a training data set to generate a model. The training data set is data of abnormal time points captured by the "black box" technology. The data is divided into 3 different levels, including measurement indicators, performance indicators, and log data. After the training data set is decoded, its data format is converted and cleaned, and then data analysis and model training are performed. The network management server can control which type (priority level) or which levels of data are used for local training according to the model training effect.

[0118] In the embodiments of the present application, fault prediction can be processed by targeting base station - type communication devices and meeting the requirements of enterprises for low network management costs and lightweight, so as to solve the problem that the ToB scenario is extremely sensitive to business network faults. By utilizing the distributed computing power of network devices at different edge nodes and using the fine - tuned data (the data at different levels collected based on the black - box technology in the above - mentioned embodiments) processed specially on the base station network devices for feature extraction and model training, the ToB network management only needs a small amount of data aggregation ability to complete global model training. The generalization ability of fault prediction is stronger, the accuracy is higher, and it is more lightweight and low - cost.

[0119] According to the method for generating a fault prediction model in the embodiments of the present application, the base station device can receive a model training instruction from the network management server, obtain the initial fault prediction model in the model training instruction, perform model training on the initial fault prediction model based on local data locally on the base station device, and send the trained first local model parameters to the network management server. Then, receive the first global model parameters obtained by the network management server through aggregating the first local model parameters of at least two base station devices, and a corresponding global fault prediction model can be generated based on the first global model parameters. According to this method, the distributed computing power of different base station devices (network devices at the edge nodes) can be utilized for feature extraction and model training, and then the trained local model parameters are reported, so that the network management server does not need to store the relevant data of the base station devices required for training, saving storage costs, and only needs a small amount of data aggregation ability to complete global model training and obtain the model parameters of the global model, which is conducive to realizing the low - cost and lightweight of the network management server, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0120] To better understand the present application, the following will describe the processing flow of the method for generating a fault prediction model in the exemplary embodiments of the present application through Figure 5 and Figure 6 , and describe the processing flow of the method for generating a fault prediction model in the exemplary embodiments of the present application.

[0121] Figure 5 Figure 14 shows the detailed flowchart of the method for generating a fault prediction model provided by the exemplary embodiments of the present application. As Figure 5 shown, the method includes the following steps.

[0122] S501, the network management server enables access to all base stations.

[0123] In this step, as a base station access condition, the base station is enabled and incorporated into the management scope of the network management server. Usually, before performing fault prediction based on federated learning, the network management server needs to enable the base station and complete the construction of the wireless communication network through a series of projects.

[0124] S502, the network management server turns on the base station black - box data recording switch.

[0125] In this step, the network management server can turn on the data recording switch that allows the target base station device to perform data recording, so as to set the local data to be collected by each target base station device based on the black box.

[0126] The black box level can be set to the first-level black box, the second-level black box, and the third-level black box. For the types of collected data corresponding to different levels of black boxes, reference can be made to the description of the above embodiments, which will not be elaborated here.

[0127] As an example, the specific strategies between the black box level setting, the types and contents of data collected by the base station, and the corresponding deep learning training methods include: if the network management server sets the black box level to the first level, the base station only records and collects the measurement index data at the abnormal point moment, and the local training of the base station is based on the cascading of CNN and LSTM, and uses the measurement index for training and fault prediction; secondly, if the network management server sets the black box level to the second level, the base station will record and collect the data of the measurement index and the performance index at the abnormal point moment. The local training of the base station is to use these two parts of data to perform parallel training based on the relevant algorithms of the model structure in which CNN and LSTM are connected in series, and finally output the result after aligning the data through the Merge operation; and if the network management server sets the black box level to the third level, the base station will record and collect the data of the measurement index, the performance index, and the log data at the abnormal point moment. The local training of the base station is to use these three parts of data to perform parallel training based on the cascading algorithm of CNN and LSTM, and finally output the result after aligning the data through the Merge operation.

[0128] In the embodiment of the present application, the black box can be triggered based on an anomaly or a condition, records the key data in the time period near the abnormal time point, and will automatically analyze some abnormal reasons in real time, automatically label the collected data, improve the data quality of the collected data, and can save the time for manually labeling some data, which has a great improvement effect on the accuracy of the training model for fault prediction.

[0129] S503, the network management server starts the federated learning process.

[0130] S504, as shown in "whether the number of base stations is not less than 2", the network management server determines whether the number of target base station devices is not less than 2.

[0131] In this step, the network management server determines whether the number of accessed target base station devices is greater than 2. If so, continue to execute S505; if not, end the process. In this step, if the number of target base station devices is less than 2, the process can be ended and wait for the next trigger; otherwise, enter the process of iteratively updating the model of federated learning.

[0132] It should be understood that in this step, the lower limit value of the number of target base station devices can also be other values, such as 3 or 5, which can be specifically customized according to the actual situation, and the embodiments of the present application do not make specific limitations.

[0133] S505. The network management server performs initial model distribution, and each base station downloads the initial model to its own base station device.

[0134] In this step, each target base station device can download the initial fault prediction model for fault prediction from the network management server. After the download is completed, a completion notification is sent to the ToB network management. Each base station starts local model training, and the trained model needs to be locally stored. This function reports the model parameters of the model after training.

[0135] S506. As shown in "Whether each base station has received the initial model", the network management server determines whether each base station has received the initial fault prediction model; if so, step S507 is executed, and if not, step S505 is executed.

[0136] In this step, if a corresponding message returned by the target base station device in response to receiving the initial fault prediction model is received, it is determined that the target base station device has received the initial fault prediction model.

[0137] S507. Each base station uses local black box data to train the local model.

[0138] In this step, a model structure with CNN and LSTM in series can be used to extract the local features and long-term dependence rules of time series data from the spatial and temporal dimensions. Specifically, the local features of time series data can be extracted by using CNN first, and then the extracted feature sequence is input into LSTM for classification (the classification result, that is, the prediction result includes: normal or faulty).

[0139] Specifically, the time series data can be segmented into windows of a fixed length, and then CNN is used to extract the local features of each window. Then, these feature sequences are input into LSTM, and LSTM will learn the long-term dependence relationship of the time series data, and output the final classification result through the fully-connected layer and the activation function (Softmax).

[0140] In the embodiments of the present application, the training data set can be from the normal operation data of the base station (normal metrics and data), the "black box" data recorded at the abnormal points, and can also include topology (Topo) data (such as the metrics and data of resource nodes in the Topo path). The "black box" data specifically includes measurement metric data, performance metric data, log data, etc. The intelligent engine platform server of the base station device can perform real-time analysis on the cause of the anomaly at the abnormal moment based on the collected data, and perform operations such as format conversion, decoding, data cleaning, and tagging on the "black box" data.

[0141] In some embodiments, the network management server can issue which level of data in the black box is used for training according to certain policies (such as the CPU usage rate of the target base station device, the degree of model convergence, etc.). The base station will select the corresponding level of data for training according to the current black box level. During the training process, a model structure with a series connection of CNN and LSTM can be used to train with one type of data and multiple types of data, and then their training results can be fused together by means such as merging or splicing. For example, the outputs of different levels can be connected together to form a larger feature vector, and then it can be input into the fully connected layer or the output layer for the final prediction task.

[0142] In some embodiments, a model structure with a series connection of CNN and LSTM can be used for a single level of data type. Model training includes: after inputting the key measurement data of the base station device, the CNN unit extracts the corresponding features. Specifically, in the CNN unit, the convolutional layer is used to expand the depth, the pooling layer is used to reduce the number of parameters by dimensionality reduction, and the fully connected layer converts the features into a one-dimensional vector, thus completing the feature extraction work of the CNN unit. The LSTM layer learns the change law of the key measurement data of the base station based on the features extracted by the CNN to predict the probability of future faults. Finally, the output layer outputs the prediction result.

[0143] As an example, if the black box setting is level one, and the data collected through the black box technology is the key measurement metrics, at this time, only the model training process based on CNN + LSTM can be enabled. Specifically, CNN can be used to extract the features of the key measurement metrics, the LSTM layer learns the change law features of the key measurement metrics of the base station based on the features extracted by the CNN, and the output result of the final LSTM unit, after being flattened by the Flatten layer to pull the data into a one-dimensional vector, passes through multiple fully connected layers and then passes through the activation layer (Softmax activation function) for prediction to obtain the prediction result during the model training process.

[0144] In the embodiments of the present application, a model structure in which CNN and LSTM are connected in series is proposed, as well as a method for parallel training of multiple types of data sources. It is necessary to consider how to merge and splice the results output by each type of data through the model structure in which CNN and LSTM are connected in series. The embodiments of the present application can propose a custom fusion method based on specific tasks according to the data type and data priority of the data collected by the base station device.

[0145] As an example, if the black box setting is level two, turn on CNN+LSTM, use parallel training of measurement metrics and performance metrics. The output results of the two LSTM units are flattened into one-dimensional vectors through Flatten, and weighted addition with specific weights is used for merging and splicing.

[0146] In this example, based on the importance and priority of different types of data in the base station device, it can be set that the weight assigned to the training output result of the measurement metric type data (represented by R2_MR_CNN_LSTM) is 0.7, and the weight assigned to the training output result of the performance metric type data (represented by R2_NI_CNN_LSTM) is 0.3. Therefore, the merging strategy of the Merge layer can be expressed as: 0.7*R2_MR_CNN_LSTM + 0.3*R2_NI_CNN_LSTM. The fused one-dimensional vector is input into the multi-layer Fully-connected layer and then passed through the Softmax activation function for prediction to obtain the prediction result during the model training process.

[0147] It should be understood that the weights of R2_NI_CNN_LSTM and R2_MR_CNN_LSTM can be set according to actual needs, as long as the sum of the two weights is 1.

[0148] Figure 6 The figure shows a schematic diagram of the model training process for multiple categories in an exemplary embodiment of the present application. In Figure 6 it, each CNN can be called a Convolutional Neural Network (CNN) unit 610, each LSTM can be called a Long Short-Term Neural Network (LSTM) unit 620. Each CNN unit includes: at least one hierarchical module and a Flatten layer. Each hierarchical module includes at least: a Convolution Layer and a Pooling layer.

[0149] In some embodiments, each hierarchical module may further include the following hierarchical levels (represented by the "..." module after the pooling layer) for example: an Activation Layer, a Pooling Layer, and a FullyConnected Layer.

[0150] Reference Figure 6 Figure 6

[0151] Exemplarily, based on the importance and priority of these three types of data in the base station device, we set the weight assigned to the training output result of the measurement metric type data (denoted as R3_MR_CNN_LSTM) to be 0.5, the weight assigned to the training output result of the performance metric type data (denoted as R3_NI_CNN_LSTM) to be 0.3, and the weight assigned to the training output result of the log type data (denoted as R3_LOG_CNN_LSTM) to be 0.2. So the final Merge strategy is 0.5 * R3_MR_CNN_LSTM + 0.3 * R3_NI_CNN_LSTM + 0.2 * R3_LOG_CNN_LSTM.

[0152] It should be understood that the weights of R3_MR_CNN_LSTM, R3_NI_CNN_LSTM, and R3_LOG_CNN_LSTM can be customized according to actual needs, as long as the sum of the weights of the three is 1.

[0153] In this embodiment, model training is carried out using three types of data and based on a serial CNN and LSTM model structure. The parallel output results are weighted and fused to output a one-dimensional vector. After integrating the features, aligning the data, and adjusting the data compatibility of the three types of data, the final result is output through a multi-layer Fully-connected layer followed by a Softmax activation.

[0154] S508, each base station encrypts and uploads the trained model parameters to the network management server for model gradient aggregation.

[0155] In this step, the network management server receives the encrypted model parameters uploaded by each target base station device, decrypts the model parameter ciphertext according to a pre-determined decryption method to obtain the decrypted model parameters, and then performs gradient aggregation on the decrypted multiple model parameters.

[0156] In the embodiments of the present application, after the model training of all participating base station devices is completed, the gradient parameters of the local model are encrypted and uploaded to the network management server (or called the network management center server). Due to the optimization strategy proposed above, there will be a phenomenon that the number of training rounds of each base station is different and the number of base stations participating in training each time is also uncertain, which will cause a certain degree of data heterogeneity. Therefore, the network management server uses the SCAFFOLD aggregation algorithm to correct and aggregate the model gradients of the participating base stations to form a global model gradient parameter. SCAFFOLD can handle the data of base stations that do not fully participate, such as disconnection from the network management caused by base station outages and the number of base stations participating in training is not fixed, etc., that is, it allows only some base stations to participate in training and a certain degree of data heterogeneity, and uses the pre-trained model to make up for the impact of missing data. In addition, the SAFFOLD aggregation algorithm has a good inhibitory effect on the "client drift" phenomenon and has better robustness and adaptability when processing the data of only some participating base stations, making the global model converge faster.

[0157] S509, as shown in "Whether all base stations have transmitted the model gradient parameters", the network management server determines whether all base stations have transmitted the model gradient parameters. If so, step S510 is executed; if not, step S508 is continued.

[0158] S510, the network management server broadcasts the aggregated global model gradient parameters to each base station device.

[0159] In this step, the network management server determines the global model according to the aggregated global model gradient parameters, and each base station device can update the model according to the global model gradient to obtain the current global model for fault prediction.

[0160] The network management server can broadcast the generated global model gradient parameters to all base station devices managed by the network management server. After each base station device updates its local model, it generates the current global model for fault prediction according to the received global model gradient parameters.

[0161] S511, as shown in "Whether the global model converges", the network management server determines whether the global model converges. If so, step S512 is executed; if not, step S507 is executed.

[0162] End the model training, and each base station device performs model prediction.

[0163] In this step, the network management server can broadcast an instruction to end the model training to multiple target base station devices. Each base station device stops the model training according to the received instruction, and uses the current global model for fault prediction as the global model for fault prediction to perform fault prediction locally at the base station device.

[0164] In this step, the network management server can iteratively execute the above steps S504 - S506 until the actual effect of the global model meets certain conditions (for example, the model accuracy rate is greater than or equal to the preset accuracy threshold) or the model converges, and then this stage of federated learning can be stopped. Each base station uses the optimal model trained currently for fault prediction.

[0165] Through the above steps S501 - S512, a lightweight and low - cost scheme for base station fault prediction based on federated learning is proposed. Model training can be directly performed at the edge nodes (base station devices) where data is produced, without the need to transmit massive amounts of data to the upper - layer network management server for model training. Thus, a large amount of transmission bandwidth and storage media can be saved, and even some cloud resources can be saved, meeting the requirements of most ToB enterprises for high reliability and low cost of network services. In the embodiments of the present application, data near the abnormal points collected in real - time on the base station side can be used to generate time - series data and index - related data locally after data encoding and data format conversion, and then through data cleaning, data analysis, and model training. Finally, the model parameters trained by each base station are uploaded to the ToB network management for aggregation through a certain algorithm to generate a global model.

[0166] In an actual application scenario, the fault prediction model generation method in the embodiments of the present application can be implemented as a fault prediction method based on federated learning. Aiming at the operation and maintenance capabilities of networks and devices with requirements for low - latency and high - reliability services in ToB private networks, the distributed computing power of multiple edge - node network devices (base stations) is used for local model training. The data set on which the model training depends is specially designed, greatly improving the accuracy and training efficiency of the model. And a dynamic scheduling algorithm based on base - station services and data characteristics is proposed, which can intelligently allocate the behaviors of base stations. The trained model parameters are uploaded to the ToB network management for aggregation to generate a global model, and then the global model parameters are sent down to the edge nodes locally for fault prediction of network devices. This will greatly reduce the computing pressure on the ToB network management and the communication overhead and pressure from the base stations to the ToB network management. In addition, since the data uploaded by network devices is only the model parameters after local training, it avoids the ToB network management from storing a large amount of relevant data of network devices, achieving the lightweight of the ToB network management.

[0167] The fault prediction model generation method in the embodiments of the present application can, based on federated learning, adopt a dynamic adjustment strategy based on service availability to schedule the distributed computing power of edge nodes such as base stations, and propose a coordination algorithm based on base - station services and data characteristics, which is beneficial to solving the communication overhead and bottleneck between the ToB network management and base stations, and is also beneficial to optimizing the aggregation efficiency problem of the central server (network management server) in federated learning.

[0168] It can be understood that, without violating the principle logic, the above-mentioned method embodiments mentioned in this application can be combined with each other to form combined embodiments. Due to space limitations, this application will not elaborate further. Those skilled in the art can understand that in the above-mentioned method of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0169] In addition, this application also provides a fault prediction model generation device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any one of the fault prediction model generation methods provided by this application. The corresponding technical solutions and descriptions can be referred to the corresponding records in the method part and will not be elaborated here.

[0170] In a third aspect, an embodiment of this application provides a fault prediction model generation device, which is applied to a network management server.

[0171] Figure 7 is a block diagram of a fault prediction model generation device provided by an embodiment of this application. Refer to Figure 7 , an embodiment of this application provides a fault prediction model generation device, and the fault prediction model generation device 700 may include the following modules.

[0172] A first sending module 710, configured to send a model training instruction to all access base station devices, where the model training instruction carries an initial fault prediction model;

[0173] A first receiving module 720, configured to receive first local model parameters of at least two target base station devices based on the model training instruction, where the first local model parameters are model parameters obtained by training the initial fault prediction model based on local data of the target base station;

[0174] An aggregation module 730, configured to perform aggregation processing on the first local model parameters of at least two target base station devices to obtain first global model parameters;

[0175] The first sending module 710 is further configured to send the first global model parameters to at least two target base station devices, so that at least two target base station devices respectively generate a global fault prediction model based on the first global model parameters.

[0176] In some embodiments, the fault prediction model generation device 700 further includes: a data recording activation module, configured to determine the computing power of each target base station device according to the computing power resources of each target base station device obtained in advance before sending the model training instruction to all access base station devices; send a data collection instruction to each target base station device, where the data collection instruction includes the number of categories of local data determined according to the computing power of each target base station device, and is used to instruct each target base station device to collect local data corresponding to the number of categories.

[0177] In some embodiments, the fault prediction model generation device 700 further includes: a training round indication module, configured to obtain at least one traffic volume evaluation index of each target base station device before receiving the first local model parameters of at least two target base station devices based on a model training instruction; determine the corresponding training round of each target base station device based on at least one traffic volume evaluation index of each target base station device and a training round look-up table; and send the corresponding training round to at least two target base station devices, where the training round is used to represent the number of times of model training for the initial fault prediction model.

[0178] In some embodiments, the training round look-up table is used to characterize the corresponding relationship between the traffic volume evaluation value and the training round; when determining the corresponding training round of each target base station device based on at least one traffic volume evaluation index of each target base station device and the training round look-up table, the training round indication module is specifically configured to: perform weighted fusion on multiple traffic volume evaluation indexes of each target base station device to obtain a traffic volume evaluation value corresponding to each target base station device; and determine the corresponding training round of each target base station device according to the traffic volume evaluation value of each target base station device and the corresponding relationship between the traffic volume evaluation value and the training round, where the training round is inversely proportional to the corresponding traffic volume evaluation value.

[0179] In some embodiments, the fault prediction model generation device 700 further includes: a model update module, configured to perform at least one global model update until the updated global fault prediction model reaches a training end condition after aggregating the first local model parameters of at least two target base station devices to obtain a first global model parameter and when the global fault prediction model does not reach the training end condition, and perform the following operations in each global model update: determine the target base station devices participating in the current global model update; send a model update instruction to the determined target base station devices; receive the second local model parameters of at least two determined target base station devices based on the model update instruction; aggregate the second local model parameters of at least two determined target base station devices to obtain a second global model parameter, and send the second global model parameter to at least two determined target base station devices, so that at least two determined target base station devices respectively generate a new global fault prediction model based on the second global model parameter.

[0180] In some embodiments, when sending the model update instruction to the determined target base station devices, the first sending module 710 is further configured to: carry the training round of the current global model update in the model update instruction, so that at least two determined target base station devices obtain the second local model parameters after performing model training for the training round of the current global model update respectively.

[0181] According to the fault prediction model generation device of the embodiments of the present application, the network management server can send the initial fault prediction model to all access base station devices through a model training instruction, and receive the first local model parameters of at least two target base station devices based on the model training instruction. Then, the network management server performs an aggregation process on the first local model parameters of the at least two target base station devices to obtain the first global model parameters, and sends the first global model parameters to the at least two target base station devices, so that the at least two base station devices can respectively generate a global fault prediction model based on the first global model parameters; according to this method, the network management server does not need to store the relevant data of the base station devices required for training, thus saving storage costs, and only requires a small amount of data aggregation ability to complete the global model training and obtain the global model parameters, which is beneficial to realizing the low cost and lightweight of the network management server, and the generalization ability of the trained fault prediction model is also stronger and the accuracy is higher.

[0182] Fourthly, the embodiments of the present application provide a fault prediction model generation device, which is applied to a base station device.

[0183] Figure 8 It is a block diagram of a fault prediction model generation device provided by the embodiments of the present application. Refer to Figure 8 In this regard, the embodiments of the present application provide a fault prediction model generation device, and the fault prediction model generation device 800 may include the following modules.

[0184] The second receiving module 810 is configured to receive a model training instruction from the network management server, and the model training instruction carries an initial fault prediction model;

[0185] The training module 820 is configured to, in response to the model training instruction, perform model training on the initial fault prediction model based on local data to obtain the first local model parameters;

[0186] The second sending module 830 is configured to send the first local model parameters to the network management server, and the network management server is configured to perform an aggregation process on the first local model parameters of at least two base station devices received based on the model training instruction to obtain the first global model parameters;

[0187] The second receiving module 810 is further configured to receive the first global model parameters from the network management server and generate a corresponding global fault prediction model based on the first global model parameters.

[0188] In some embodiments, the fault prediction model generation device 800 further includes: a recording and starting module, configured to receive a data collection instruction from the network management server before performing model training on the initial fault prediction model based on local data; in response to the data collection instruction, start collecting local data, and the number of categories of the local data matches the computing power of the base station device itself.

[0189] In some embodiments, the model structure of the initial fault prediction model includes: a convolutional neural network model and a long short-term neural network model connected in series; when the number of categories of local data is greater than 1, the training data includes multiple types of training data; when the training module 820 is used to perform model training on the initial fault prediction model based on local data to obtain the first local model parameters, it is specifically used for: for the local data of each category, respectively use the convolutional neural network model to extract features to obtain the first features corresponding to the local data of each category; use the long short-term neural network to process the first features of the local data of each category respectively to obtain the second features corresponding to the local data of each category; fuse the second features of the local data of each category to obtain fused features; perform model training according to the fused features to obtain the first local model parameters.

[0190] In some embodiments, the training module 820 is specifically used for: in response to a model training instruction, determine the number of training rounds, where the number of training rounds comes from the network management server and is the number of training rounds corresponding to the base station device determined by the network management server based on at least one traffic evaluation index of the base station device and a training rounds comparison table; after performing model training for the number of training rounds on the initial fault prediction model, obtain the first local model parameters.

[0191] In some embodiments, the fault prediction model generation device 800 further includes: a training and updating module, configured to, after sending the first local model parameters to the network management server, receive a model update instruction from the network management server, where the model update instruction is an instruction generated by the network management server when the global fault prediction model has not reached the training end condition; in response to the model update instruction, continue to perform model training on the initial fault prediction model based on local data to obtain second local model parameters; send the second local model parameters to the network management server, where the network management server is used to perform an aggregation process on the second local model parameters of at least two base station devices to obtain second global model parameters; receive the second global model parameters generated by the network management server; generate a new global fault prediction model according to the second global model parameters.

[0192] For the fault prediction model generation device according to the embodiments of the present application, the base station device can receive the model training instruction from the network management server, obtain the initial fault prediction model in the model training instruction, perform model training on the initial fault prediction model based on local data locally at the base station device, and send the obtained first local model parameters to the network management server. After receiving the first global model parameters obtained by aggregating the first local model parameters of at least two base station devices by the network management server, the corresponding global fault prediction model can be generated based on the first global model parameters. According to this method, the distributed computing power of different base station devices (network devices at the edge nodes) can be utilized for feature extraction and model training, and then the obtained local model parameters are reported, so that the network management server does not need to store the relevant data of the base station devices required for training, saving storage costs, and only a small amount of data aggregation ability is required to complete the global model training and obtain the model parameters of the global model, which is conducive to realizing the low cost and lightweight of the network management server, and the generalization ability of the trained fault prediction model is stronger and the accuracy is higher.

[0193] It should be clear that the present invention is not limited to the specific configurations and processes described and illustrated in the above embodiments. For the convenience and conciseness of description, the detailed descriptions of known methods are omitted here, and the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0194] Each module in the above-mentioned fault prediction model generation device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0195] In the embodiments of the present application, the low-cost requirements of the ToB scenario for server resources can be solved or addressed. When doing intelligent applications, such as fault prediction, for the solutions in the related technologies that consume a large amount of computing resources, storage resources, and bandwidth resources, a lightweight fault prediction model generation method can be used to complete complex and heavy training services by utilizing the capabilities of distributed multi-node edge computing base stations. In some scenarios, when performing model training locally at the base station, the dataset can come from specially designed "black box" data. The black box encapsulates the key data when an anomaly occurs and can analyze the cause in real time and label it, making it more convenient to obtain high-quality data; the black box data can be graded to enable the network management server to flexibly control which levels of data are used for training. Based on the model structure of the cascade of CNN and LSTM, a scheme of overall parallelism of multiple data types is used in combination with the black box level switch, and dynamic scheduling is performed according to a certain resource availability situation.

[0196] In the embodiments of the present application, the ToB industry pays more attention to data security and privacy. Traditional solutions generally transmit user-related data to a server center outside the campus for large-scale model training, which may lead to the risk of leaking user privacy. This solution can better protect the privacy and security of user data in scenarios where the network management server is not in the campus.

[0197] In addition, the invention of this application also proposes an optimization algorithm based on base station business characteristics and collected data for the communication overhead problem of the most critical central server (network management server) and edge computing node (base station equipment) in federated learning, ensuring that this communication overhead is reduced to a certain extent, and can meet the timeliness and high reliability requirements of ToB enterprise intelligent business production. The fault prediction model can take effect locally on the base station equipment in a timely manner without the need for distribution through the network management, saving the distribution process and avoiding the risk of network failures during the distribution process.

[0198] Based on the description of the above embodiments, the embodiments of the present application can be applied to the management of communication equipment fault prediction in ToB enterprises with requirements such as low cost, high timeliness, and high data security, helping enterprises to complete forward-looking complex businesses in a lightweight, efficient, and high-quality manner.

[0199] In a fifth aspect, an embodiment of the present application also provides an electronic device.

[0200] Reference Figure 9 The electronic device includes: at least one processor 901; at least one memory 902, and one or more I / O interfaces 903; wherein the memory 902 stores one or more computer programs that can be executed by at least one processor 901, and the one or more computer programs are executed by at least one processor 901 so that the at least one processor 901 can perform the above method.

[0201] Among them, the processor is a device with data processing capabilities, including but not limited to the central processing unit (CPU); the memory is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, and can realize information exchange between the memory and the processor, including but not limited to the data bus (Bus), etc.

[0202] Embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor / processing core, the above-mentioned method is implemented. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0203] Those of ordinary skill in the art can understand that all or some of the steps, systems, and functional modules / units in the devices disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations.

[0204] In a hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation.

[0205] Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH), or other disk memories; compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical disc memories; magnetic cartridges, tapes, magnetic disk storage, or other magnetic memories; and any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0206] The present application has disclosed exemplary embodiments, and although specific terms are employed, they are used only and should be interpreted only as general illustrative meanings and not for the purpose of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly specified, features, characteristics, and / or elements described in connection with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in connection with other embodiments. Accordingly, those skilled in the art will understand that various forms and details of changes may be made without departing from the scope of the present application as set forth by the appended claims.

Claims

1. A method for generating a fault prediction model, wherein, applied to a network management server, the method includes: Sending a model training instruction to all access base station devices, where the model training instruction carries an initial fault prediction model; Receiving first local model parameters of at least two target base station devices based on the model training instruction, where the first local model parameters are model parameters obtained by training the initial fault prediction model based on local data of the target base stations; Performing an aggregation process on the first local model parameters of the at least two target base station devices to obtain first global model parameters, and sending the first global model parameters to the at least two target base station devices, so that the at least two target base station devices respectively generate a global fault prediction model based on the first global model parameters.

2. The method according to claim 1, wherein, Before sending the model training instruction to all access base station devices, the method further includes: Determining the computing power of each target base station device according to the computing power resources of each target base station device obtained in advance; Sending a data collection instruction to each target base station device, where the data collection instruction includes the number of categories of local data determined according to the computing power of each target base station device, and is used to instruct each target base station device to collect local data corresponding to the number of categories.

3. The method according to claim 1, wherein, Before receiving the first local model parameters of at least two target base station devices based on the model training instruction, it further includes: Obtaining at least one traffic volume evaluation index of each target base station device; Determining the corresponding number of training rounds for each target base station device based on the at least one traffic volume evaluation index of each target base station device and a training round comparison table; Sending the corresponding number of training rounds to the at least two target base station devices, where the number of training rounds is used to represent the number of times of model training for the initial fault prediction model.

4. The method according to claim 3, wherein, The training round comparison table is used to characterize the corresponding relationship between the traffic volume evaluation value and the number of training rounds; The determining the corresponding number of training rounds for each target base station device based on the at least one traffic volume evaluation index of each target base station device and the training round comparison table includes: Performing weighted fusion on multiple traffic volume evaluation indexes of each target base station device to obtain a traffic volume evaluation value corresponding to each target base station device; Determining the corresponding number of training rounds for each target base station device according to the traffic volume evaluation value of each target base station device and the corresponding relationship between the traffic volume evaluation value and the number of training rounds, where the number of training rounds is inversely proportional to the corresponding traffic volume evaluation value.

5. The method according to claim 1, wherein, After performing an aggregation process on the first local model parameters of the at least two target base station devices to obtain first global model parameters, it further includes: In the case that the global fault prediction model does not reach the training end condition, performing at least one global model update until the updated global fault prediction model reaches the training end condition, and performing the following operations in each global model update: Determine the target base station devices participating in the current global model update; Send a model update instruction to the determined target base station devices; Receive the second local model parameters of at least two of the determined target base station devices based on the model update instruction; Aggregate the second local model parameters of the at least two determined target base station devices to obtain second global model parameters, and send the second global model parameters to the at least two determined target base station devices, so that the at least two determined target base station devices respectively generate new global fault prediction models based on the second global model parameters.

6. The method according to claim 5, wherein, The sending the model update instruction to the determined target base station devices further includes: Carry the number of training rounds of the current global model update in the model update instruction, so that the at least two determined target base station devices obtain the second local model parameters after respectively performing model training for the number of training rounds of the current global model update.

7. A method for generating a fault prediction model, wherein, Applied to a base station device, the method includes: Receive a model training instruction from a network management server, and the model training instruction carries an initial fault prediction model; In response to the model training instruction, perform model training on the initial fault prediction model based on local data to obtain first local model parameters; Send the first local model parameters to the network management server, and the network management server is used to aggregate the first local model parameters of at least two of the base station devices received based on the model training instruction to obtain first global model parameters; Receive the first global model parameters from the network management server and generate a corresponding global fault prediction model based on the first global model parameters.

8. The method according to claim 7, wherein, Before performing model training on the initial fault prediction model based on local data, it further includes: Receive a data collection instruction from the network management server; In response to the data collection instruction, start collecting local data, and the number of categories of the local data matches the computing power of the base station device itself.

9. The method according to claim 8, wherein, The model structure of the initial fault prediction model includes: a cascaded convolutional neural network model and a long short-term neural network model; when the number of categories of the local data is greater than 1, the performing model training on the initial fault prediction model based on local data to obtain first local model parameters includes: For each category of local data, respectively use a convolutional neural network model to extract features to obtain first features corresponding to each category of local data; Use a long short-term neural network to process the first features of each category of local data respectively to obtain second features corresponding to each category of local data; Fuse the second features of each category of local data to obtain a fused feature; Perform model training according to the fused feature to obtain first local model parameters.

10. The method according to claim 9, wherein, In response to the model training instruction, performing model training on the initial fault prediction model based on local data to obtain first local model parameters, including: In response to the model training instruction, determining a training round, where the training round is from the network management server and is determined by the network management server based on at least one traffic evaluation index of the base station device and a training round comparison table for the base station device; After performing model training for the training round on the initial fault prediction model, obtaining first local model parameters.

11. The method according to claim 7, wherein, after sending the first local model parameters to the network management server, the method further includes: receiving a model update instruction from the network management server, where the model update instruction is an instruction generated by the network management server when the global fault prediction model does not reach the training end condition; in response to the model update instruction, continuing to perform model training on the initial fault prediction model based on the local data to obtain second local model parameters; sending the second local model parameters to the network management server, where the network management server is configured to perform an aggregation process on the second local model parameters of at least two base station devices to obtain second global model parameters; receiving the second global model parameters generated by the network management server; generating a new global fault prediction model according to the second global model parameters.

12. An electronic device, including: at least one processor; a memory storing at least one program, where when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-6 or any one of claims 7-11.

13. A storage medium, wherein, the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-6 or any one of claims 7-11.

Citation Information

Cited By

  • A model determination method, device, storage medium, and program product

    CN122634187A