Model training method, apparatus and system
By using federated learning, the central node collaborates with multiple first nodes to train the AI model, solving the problem of low recognition rate under data privacy protection and achieving efficient business content recognition and improved recognition capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-08-23
- Publication Date
- 2026-04-21
AI Technical Summary
Under the premise of meeting data privacy protection requirements, how to improve the recognition rate of artificial intelligence models, especially when there is little local sample data at the node, makes it difficult for existing technologies to achieve efficient business content recognition.
By employing a federated learning approach, model training information is sent from a central node to multiple first nodes. The AI model is trained using local data and model training configuration information, and the model parameter update information is aggregated into second model parameter update information to update the AI models of each node, thereby achieving distributed training and improved recognition capabilities.
While ensuring data privacy protection, the recognition rate of AI models has been improved, the recognition capabilities of nodes have been synchronized and the training computing power has been shared, adapting to the different data of nodes and rapidly improving the recognition capabilities.
Smart Images

Figure CN115718868B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a model training method, apparatus and system. Background Technology
[0002] With the rapid development of mobile internet in recent years, new applications such as augmented reality (AR) / virtual reality (VR) and 4K high-definition video have emerged, leading to explosive growth in mobile data services. What operators urgently need is differentiated billing based on service content.
[0003] To achieve differentiated billing based on service content, it is necessary to identify the service content. Currently, service awareness (SA) technology is used to identify service content. SA technology, based on packet header analysis, can deeply analyze the characteristics of layers 4 to 7 protocols carried in data packets, and is an application layer information-based detection and control technology.
[0004] SA (Self-Service) can achieve intelligence and automation based on artificial intelligence (AI) models. However, SA based on AI models requires a large number of samples to be trained before recognition. If the local sample data of each node (or node) is transferred to each other, it does not meet the requirements of data privacy protection; if the number of samples is too small, it will result in a low recognition rate for business content.
[0005] Therefore, how to improve the recognition rate of AI models while meeting the requirements of data privacy protection has become an urgent problem to be solved. Summary of the Invention
[0006] This application provides a model training method, apparatus, and system to maximize the recognition rate of AI models while meeting data privacy protection requirements.
[0007] In a first aspect, embodiments of this application provide a model training method, which can be executed by a central node. The method includes: the central node sending model training information to at least two first nodes, the model training information including an artificial intelligence (AI) model and model training configuration information, the AI model being used to identify the category to which a data stream belongs; the central node receiving at least two first model parameter update messages from the at least two first nodes, the first model parameter update messages being model parameter update messages after training the AI model based on local data of the first node corresponding to the first model parameter update message and the model training configuration information; and the central node sending second model parameter update information to one of the at least two first nodes, the second model parameter update information being obtained based on the at least two first model parameter update messages, the second model parameter update information being used to update the model parameters of the AI model of the first node.
[0008] The above method can meet the data privacy protection requirements because the first model parameter update information is obtained by training the first node using local data and model training configuration information respectively. Moreover, since the second model parameter update information is obtained based on at least two first model parameter update information from at least two first nodes, the recognition rate of the AI model updated with the second model parameter update information can be higher, thereby maximizing the recognition rate of the AI model.
[0009] In one possible design, the category to which the data stream belongs includes at least one of the following: the application to which the data stream belongs; the type or protocol to which the service content of the data stream belongs; or, the message characteristic rules of the data stream.
[0010] In the above design, the AI model can be used to identify the application to which the data stream belongs. Since identifying the application of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the application to which the data stream belongs. Alternatively, in the above solution, the AI model can be used to identify the type or protocol of the business content of the data stream. Again, since identifying the type or protocol of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the type or protocol to which the data stream belongs. Or, in the above solution, the AI model can be used to identify the feature rules of the data stream. Since identifying the message feature rules of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the message feature rules of the data stream.
[0011] In one possible design, after the central node receives at least two first model parameter update messages from the at least two first nodes, the design further includes: the central node sending the second model parameter update messages and the AI model to the second node; the second model parameter update messages are used to update the model parameters of the AI model of the second node.
[0012] The above design updates the model parameters of the AI model of the second node by using the second model parameter update information obtained from the update information of the first model parameters after the AI model is trained using local data and model training configuration information from at least two first nodes. This allows the AI model of the second node to combine the update information of the first model parameters from at least two first nodes, thereby maximizing the recognition rate of the AI model of the second node.
[0013] In one possible design, the model training configuration information further includes: a training result accuracy threshold; the training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
[0014] The above design allows the central node to control the accuracy of the training results of the first node's AI model by setting a training result accuracy threshold in the model training configuration information. When the accuracy of the AI model trained by the first node reaches the training result accuracy threshold, the model training can be stopped.
[0015] Secondly, embodiments of this application provide a model training method, which can be executed by a first node. The method includes: the first node receiving a model training message, the model training message including an AI model and model training configuration information, the AI model being used to identify the category to which a data stream belongs; the first node sending first model parameter update information, the first model parameter update information being model parameter update information after training the AI model based on the first node's local data and the model training configuration information; and the first node receiving second model parameter update information, the second model parameter update information being obtained based on at least two first model parameter update messages from at least two first nodes, the second model parameter update information being used to update the model parameters of the AI model of the first node.
[0016] The above method can meet the requirements of data privacy protection because the first node trains the AI model based on local data and model training configuration information; and because the second model parameter update information is obtained based on the first model parameter update information of at least two first nodes, the recognition rate of the AI model updated with the second model parameter update information can be higher, thereby maximizing the recognition rate of the AI model.
[0017] In one possible design, the category to which the data stream belongs includes at least one of the following: the application to which the data stream belongs; the type or protocol to which the service content of the data stream belongs; or, the message characteristic rules of the data stream.
[0018] In the above design, the AI model can be used to identify the application to which the data stream belongs. Since identifying the application of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the application to which the data stream belongs. Alternatively, in the above solution, the AI model can be used to identify the type or protocol of the business content of the data stream. Again, since identifying the type or protocol of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the type or protocol to which the data stream belongs. Or, in the above solution, the AI model can be used to identify the feature rules of the data stream. Since identifying the message feature rules of the data stream using an AI model eliminates the need for human intervention, it also improves the ability to identify the message feature rules of the data stream.
[0019] In one possible design, the method further includes: the first node determining that the first node trains the AI model updated with the second model parameter update information based on local data.
[0020] In the above design, the first node uses local data to train the AI model updated with the second model parameter update information, which can make the final trained AI model have a higher recognition rate for local data.
[0021] In one possible design, the method is executed by an application (APP) deployed on a cloud platform or edge computing platform.
[0022] The above design allows the APP deployed on a cloud platform or edge computing platform to execute the method of the first node, which can decouple the APP from the cloud platform or edge computing platform and minimize the modification to the existing cloud platform or edge computing platform.
[0023] In one possible design, after receiving the model training message and before sending the first model parameter update information, the method further includes: receiving the first model parameter update information from the first node through the server module of the APP; sending the first model parameter update information includes: sending the first model parameter update information through the client module of the APP.
[0024] The above design sets up a server module and a client module in the APP. The server module enables communication with the first node, and the client module enables communication with the outside world, which can give full play to the APP's information transmission function.
[0025] In one possible design, the first model parameter update information is obtained after the model parameters of the trained AI model have been successfully verified based on the model training configuration information.
[0026] The above design verifies the model parameters of the trained AI model before sending the first model parameter update information, which can ensure that the model parameters of the AI model are consistent before and after training, and can minimize the impact on training effect due to inconsistencies in the model parameters of the AI model before and after training.
[0027] In one possible design, the model training configuration information further includes: a training result accuracy threshold; the training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
[0028] The above design allows the central node to control the accuracy of the training results of the first node's AI model by setting a training result accuracy threshold in the model training configuration information. When the accuracy of the AI model trained by the first node reaches the training result accuracy threshold, the model training can be stopped.
[0029] In one possible design, after receiving the second model parameter update information, the method further includes: the first node acquiring the identification result and message feature rules for identifying the data stream based on the AI model updated with the second model parameter update information; and updating the service-aware SA feature library according to the message feature rules.
[0030] The above design improves the recognition rate of the SA feature library by updating the message feature rules identified by the AI model obtained from model training to the SA feature library.
[0031] Thirdly, embodiments of this application provide a model training apparatus, which includes modules for executing the first aspect or any possible design of the first aspect. Alternatively, the model training apparatus includes modules for executing the second aspect or any possible design of the second aspect.
[0032] Fourthly, embodiments of this application provide a model training apparatus, which includes a processor and a memory. The memory stores computer execution instructions, and when the processor runs, the processor executes the computer execution instructions in the memory to perform operational steps of any possible design method of any of the first to second aspects using the hardware resources in the controller.
[0033] Fifthly, embodiments of this application provide a model training system, including the model training apparatus provided in the third or fourth aspect above.
[0034] Sixthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described above.
[0035] In a seventh aspect, based on the same inventive concept as in the first aspect, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described above.
[0036] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0037] Figure 1 A schematic diagram of a federated learning architecture provided for an embodiment of this application;
[0038] Figure 2 A schematic diagram illustrating an application scenario provided in an embodiment of this application;
[0039] Figure 3 A schematic flowchart illustrating a model training method provided in an embodiment of this application;
[0040] Figure 4 A flowchart illustrating another model training method provided in an embodiment of this application;
[0041] Figure 5 This is a schematic diagram of the architecture of a model training system provided in an embodiment of this application;
[0042] Figure 6 This is a schematic diagram showing the distribution of sample data for each first node in scenario one.
[0043] Figure 7 This is a schematic diagram showing the recognition accuracy of each first-node model and the federated model in scenario one.
[0044] Figure 8 This is a schematic diagram showing the distribution of sample data for each first node in scenario two.
[0045] Figure 9 This is a schematic diagram showing the recognition accuracy of each first-node model and the federated model in scenario two.
[0046] Figure 10 This is a schematic diagram showing the recall rate of small sample applications in each first node in scenario two and the recall rate after federated learning.
[0047] Figure 11 This is a schematic diagram showing the distribution of sample data for each first node in scenario three.
[0048] Figure 12 This is a schematic diagram showing the recognition accuracy of each first-node model and the federated model in scenario three.
[0049] Figure 13 This is a schematic diagram showing the recall rate after applying federated learning to each first node in scenario three when there are no samples.
[0050] Figure 14 This is a schematic diagram illustrating the recognition accuracy after fine-tuning and initializing the federated model and retraining in Scenario 4.
[0051] Figure 15 This diagram illustrates the recognition accuracy of the retrained model after fine-tuning and initializing the federated model in Scenario 4 for the new application.
[0052] Figure 16 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0053] Figure 17 This is a schematic diagram of another model training device provided in an embodiment of this application;
[0054] Figure 18 This is a schematic diagram of another model training device provided in an embodiment of this application. Detailed Implementation
[0055] AI-based SA (Search Engine Optimization) technology can achieve intelligent and automated identification. For example, the main process of traffic identification based on AI algorithms includes: a training phase, which involves capturing traffic, preprocessing it, and then using it as input to the AI model; and a testing phase, which involves feeding the preprocessed traffic into the AI model for classification, and taking the traffic category with the highest probability in the classifier results as the final traffic prediction result.
[0056] However, the SA technology based on AI algorithms has the following drawbacks:
[0057] First, a large amount of labeled data is needed for model training. The more complex the content to be identified and the wider the scope of recognition, the larger the amount of data required. The model's generalization ability is also related to the amount of training data. Furthermore, the training data must be protected for privacy, limited to local use, and cannot be shared externally. For example, if the traffic of the application that a node needs to identify is too small (small sample size), the node may either lack the ability to identify this application or have poor generalization ability and low recognition accuracy.
[0058] Secondly, AI model training and inference consume a significant amount of processor performance. For example, if a single node has high performance requirements and the training time is too long, it will affect the rapid iteration of the model and cause the recognition ability to not be updated quickly.
[0059] Secondly, app traffic exhibits regional distribution characteristics; even the same app can have different traffic patterns in different regions. For example, when traffic patterns change rapidly, the ability of node A's AI model to identify a new app cannot be passed on to node B. Furthermore, due to privacy concerns, node data cannot be exported for model training, and data collected through methods such as dial-up testing is insufficient for training requirements.
[0060] In view of this, embodiments of this application provide a model training method, apparatus, and system for maximizing the recognition rate of AI models while meeting data privacy protection requirements.
[0061] Before providing a detailed explanation of the embodiments of this application, the system architecture involved in the embodiments of this application will be introduced first.
[0062] Figure 1 This is a schematic diagram of a federated learning architecture provided for an embodiment of this application. For ease of understanding, let's first combine... Figure 1 The scenarios and processes of federated learning are illustrated by example.
[0063] Federated learning is an encrypted distributed machine learning technique that refers to the collaborative construction of an AI model by participating parties without sharing local data. Its core principle is: each participant trains the AI model locally, then encrypts and uploads only the updated parts of the model to the coordinating node, where they are aggregated and integrated with the updated parts from other participants to form a federated learning model. This federated learning model is then distributed to all participants from the cloud. Through repeated local training and integration, a better AI model is ultimately obtained.
[0064] See Figure 1 In a federated learning scenario, there may be a coordinating node and multiple participating nodes. The coordinating node acts as the coordinator in the federated learning process and can be deployed in the cloud. The participating nodes are the participants in the federated learning process and also the owners of the dataset. For ease of understanding and distinction, in this embodiment, the coordinating node is referred to as the central node (e.g., in...). Figure 1 Marked as 110), the participating node is referred to as the first node (e.g., in...). Figure 1 (marked as 120 and 121).
[0065] The central node 110 and the first nodes 120 and 121 can be any node that supports data transmission (such as a network node). For example, the central node can be a server, or a parameter server, or an aggregation server. The first nodes can be clients, such as mobile terminals or personal computers.
[0066] The central node 110 can be used to maintain the federated learning model. First nodes 120 and 121 can obtain the federated learning model from the central node 110 and train it locally using their local training datasets to obtain a local model. After training the local model, the first nodes 120 and 121 can send the local model to the central node 110 so that the central node 110 can update or optimize the federated learning model. This process is repeated for multiple iterations until the federated learning model converges or reaches a preset iteration stopping condition (e.g., reaching the maximum number of iterations or the longest training time).
[0067] The embodiments of this application can be applied to various scenarios such as service packages based on SA technology, service management based on SA technology, or auxiliary operations based on SA technology.
[0068] For example, for data traffic packages with specific business targets proposed by operators, the solution provided in the embodiments of this application can be used to train an AI model, and then the trained AI model can be used to identify different applications for subsequent implementation of different billing and control strategies.
[0069] For example, service management based on SA technology can include bandwidth control, congestion control, or service guarantees after identifying service content. For instance, taking congestion control after identifying service content as an example, some countries or regions do not allow the use of certain types of software, such as Voice over Internet Protocol (VoIP). There are many types of VoIP applications, and their versions or protocols are updated frequently. Many applications are also encrypted, thus requiring SA technology to support the detection and control of VoIP software.
[0070] For example, with the continuous enrichment of content on the internet, operators urgently need to analyze the content transmitted on the network. By analyzing the transmission traffic, operators can better formulate business operation and maintenance strategies. This necessitates SA (Service Assist) technology to identify the traffic of different applications.
[0071] The solutions provided in the embodiments of this application are illustrated below through specific application scenarios.
[0072] Based on the above, the embodiments of this application can be applied to Figure 2 In the scenario shown. For example... Figure 2 As shown, the central node 110 can be deployed in the cloud, and the first nodes 120 and 121 can be deployed on cloud platforms respectively. Alternatively, the first nodes 120 and 121 can also be deployed on edge computing platforms respectively.
[0073] Central node 110 can download the AI model to first nodes 120 and 121. Then, first nodes 120 and 121 can train the model using local data and upload the local model parameter updates (also known as the first model parameter update information) to central node 110. Central node 110 will aggregate the local model parameter updates received from first nodes 120 and 121 to obtain shared federated model parameter updates (also known as the second model parameter updates), and then send them to first nodes 120 and 121. First nodes 120 and 121 will update their local AI models before training according to the shared federated model parameter updates to obtain the final federated model.
[0074] like Figure 2 As shown, data streams from the Internet pass through nodes 120 and 121, reaching users A and B respectively. Nodes 120 and 121 can perform SA (Application Classification) on the data streams flowing through their local nodes using a trained federated model. For example, they can identify the application to which the data stream belongs. For instance, a data stream might be identified as belonging to different applications; data stream 1 could be identified as application 1, application 2, application 3, application 4, application 5, and application 6. The accuracy for application 1 is 99%, for application 2 it is 80%, for application 3 it is 78%, for application 4 it is 72%, for application 5 it is 68%, and for application 6 it is 40%. Generally, after identification, the data stream is identified as belonging to application 1. When outputting the classification results, the names and specific accuracy scores of the applications ranked second to fifth in terms of identification accuracy can also be output as classification results.
[0075] In this scenario, AI model parameter updates can be synchronized while ensuring data privacy protection requirements, enabling normalized training across different nodes. Distributed training of AI models can also be achieved, with training computational power distributed across different nodes to avoid performance bottlenecks. Furthermore, differentiated or personalized data can be processed for different nodes, enabling rapid enhancement of recognition capabilities across them. Additionally, in cases where data samples are scarce or missing, federated learning can be used to acquire the recognition capabilities of other nodes.
[0076] Based on the above, in another possible scenario, the central node 110 can be deployed on NAIE, and the first nodes 120 and 121 can be deployed on edge computing platforms, respectively. For example, the first nodes 120 and 121 can be deployed in the local data center of operator A and the local data center of operator B, respectively. In this scenario, the recognition capabilities of each data center can be aggregated and integrated to improve the overall recognition capability while ensuring that the original data does not leave the local area. In addition, the training computing power can be distributed to enable rapid iteration of recognition capabilities for key applications and key time periods.
[0077] Based on the above, this application provides a model training method, which can be performed by... Figure 1 or Figure 2 The central node 110 and the first node 120 or 121 in the process are executed. For example... Figure 3 As shown, the method includes:
[0078] S301, the central node sends model training messages to at least two first nodes; correspondingly, the first node among the at least two first nodes receives the model training messages;
[0079] The model training message includes an artificial intelligence (AI) model and model training configuration information. The AI model is used to identify the category to which the data stream belongs.
[0080] It should be noted that, Figure 3 Only one first node is shown in the figure. This is only an illustrative example and does not limit the embodiments of this application.
[0081] The central node can be deployed in the cloud, for example, on an artificial intelligence engine.
[0082] The first node can be deployed on a cloud platform or an edge computing platform. For example, it can be deployed on a clouded multiple service engine (CloudMSE) or a multi-access edge computing (MEC).
[0083] In one possible implementation, the central node or first node can be implemented using a container service, or through one or more virtual machines (VMs), or by one or more processors, or by one or more computers.
[0084] In addition, the model training configuration information also includes: training result accuracy threshold;
[0085] The training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
[0086] In one possible implementation, the AI model can employ a neural network model. For example, the neural network model can be a convolutional neural network (CNN) model, a recurrent neural network model, etc. CNN is a deep learning network structure widely used in image recognition; it is a feedforward neural network where artificial neurons respond to surrounding units. As an example, this application embodiment can use a CNN model. Because the convolutional layers of a CNN model are multi-layered, multiple convolution calculations are performed on the original data. The data processing of the CNN model is more complex and can extract more and more complex traffic features, which is helpful for SA identification. Furthermore, the CNN model has strong generalization ability and is not sensitive to the position of traffic features in the message, requiring no special pre-processing of the data stream to be identified, thus exhibiting strong adaptability to different network environments. The following description will use the CNN model as an example.
[0087] For example, the AI model uses the CNN model. See Table 1. The model training configuration information may include the required parameters in Table 1, or it may include both the required and optional parameters in Table 1, or it may include some of the required parameters and some of the optional parameters. For example, one of parameters 4 and 5 can be included in the model training configuration information.
[0088] Table 1:
[0089]
[0090]
[0091]
[0092] Among them, parameter 27 was dynamically generated during training using parameter 25, and parameter 28 was dynamically generated during training using parameter 26.
[0093] The AI model supports a certain number of protocols, and different protocols can correspond to different applications or different types of applications. The AI model can be configured to recognize a range of categories as needed. For example, the range of protocols to be recognized by the AI model can be flexibly configured according to business needs. For ease of description, the configured protocol information is referred to as the selected protocols. For example, the selected protocols may include one or more of the following information: protocol level, number of protocols, protocol name, protocol number, etc. The selected protocols may also include AI instance number and AI instance name. The protocol level, protocol name, and protocol number can be pre-configured and do not need to be changed later. Of course, the protocol level, protocol name, protocol number, etc., can also be changed according to needs, and this embodiment does not limit this. The selected protocol ID list corresponding to parameter 3 refers to the list of identifier IDs of each protocol included in the selected protocols.
[0094] For example, the AI model and its training configuration information can be sent to the first node in the form of a configuration file. This configuration file could include the following three files: 1. a .caffemodel file, representing the initial AI model; 2. a .proto file, containing the definition information of the AI model's parameters, such as those in the table above; 3. a *netproto.txt file, containing the parameter values of the AI model's parameters, such as the parameter values in the table above. The *netproto.txt file, matched with the AI model, can be used for loading and validating the AI model.
[0095] For example, the category to which the data stream belongs may include at least one of the following: the application to which the data stream belongs; the type or protocol to which the service content of the data stream belongs; or the message characteristic rules of the data stream, etc.
[0096] For example, based on business needs, a data stream can be identified as belonging to an application. For instance, data stream A can be identified as belonging to WeChat, data stream B can be identified as belonging to YouTube, data stream C can be identified as belonging to iQiyi, and so on.
[0097] For example, based on business needs, the business content of a data stream can be identified as belonging to different types. For instance, the business content of data stream D can be identified as video, specifically as a WeChat video; the business content of data stream E can be identified as IP telephony; the business content of data stream F can be identified as image, and so on.
[0098] For example, based on business requirements, the business content of a data stream can be identified as belonging to a specific protocol. Different protocols can correspond to different applications or different types of applications. For instance, the business content of data stream G can be identified as belonging to the BT (Bit Torrent) protocol, the business content of data stream H can be identified as belonging to the MSN (Microsoft Network) protocol, and the business content of data stream F can be identified as belonging to the SMTP (Simple Mail Transfer Protocol), and so on.
[0099] For example, depending on business needs, AI models can also be used to extract message feature rules from data streams. For instance, for an unknown data stream, message feature rules can be extracted during the process of identifying the unknown data stream; that is, message feature rules are extracted from the identification process within the unknown data stream. In this embodiment, after obtaining the feature rules, they can be updated in the SA feature library.
[0100] For example, the central node can compress the model training information before sending it, and the first node can decompress the received model training information to obtain the AI model and model training configuration information; or the central node can encrypt the model training information before sending it, and the first node can decrypt the received model training information to obtain the AI model and model training configuration information; or the central node can compress and encrypt the model training information before sending it, and the first node can decompress and decrypt the received model training information to obtain the AI model and model training configuration information.
[0101] In one possible implementation, the central node, acting as the coordinator, can select at least two first nodes from a pool of first nodes to participate in federated learning and send model training information to these two first nodes. For example, the first nodes can pre-register with the central node, which then selects at least two first nodes from the registered pool to participate in federated learning. Alternatively, the central node can randomly select at least two first nodes or select them according to preset rules. For instance, the central node can select at least two first nodes that store business data that meets the training requirements from a pool of first nodes based on the data distribution information of the multiple first nodes, and these selected nodes will participate in the model training process with the AI model.
[0102] For example, after successful registration, the first node can also send heartbeat messages to the central node periodically or in real time; the heartbeat message may include the status information of the first node.
[0103] Additionally, the central node can send the AI model and its training configuration information to at least two first nodes in the form of training tasks. For example, the central node can send training task information, including the AI model and its training configuration information, to at least two first nodes. For instance, before sending the training task information to the at least two first nodes, the central node can first send a training task notification message to the first nodes, which may include a task identifier ID. Then, the first node receiving the training task notification message can send a training task query message to the central node, which may include a task ID. After receiving the training task query message, the central node then sends the training task information to the first nodes.
[0104] Before sending model training information to at least two first nodes, the central node can perform initialization settings, including: selecting the first nodes to participate in federated learning; setting the convergence algorithm; setting parameters such as the number of iterations for model training, as shown in Table 1; creating a federated learning instance and selecting SA as the instance type, for example, selecting the instance for AISA traffic identification; initializing an AI model, such as initializing an AISA traffic identification model, or injecting pre-trained model parameters and weights into an AI model. Initializing an AI model means injecting initial model parameters and weights into the AI model.
[0105] S302, the first node of the at least two first nodes trains the AI model based on local data and model training configuration information;
[0106] For example, local data can be obtained by the first node performing feature recognition on the collected data stream through the SA engine, obtaining the classification result of the data stream, and then labeling the data stream according to the classification result.
[0107] The first node can load the AI model based on the received model training configuration information. For example, it can load the AI model according to the parameters in Table 1, such as parameters 1, 3-4, 10-14, and 16-30 in Table 1. Then, it can train the AI model based on local data and the model training configuration information, such as parameters 3, 8, 15, 31, and 32 in Table 1. For example, the process of training the AI model may include: dividing the local data into K batches according to a predetermined size B (batch size); and training the AI model K times based on the K batches. The value of B is the parameter value of parameter 31 in Table 1. When the parameter value of the number of training iterations in parameter 8 in Table 1 is 1, the above training process only needs to be executed once; when the parameter value of the number of training iterations in parameter 8 is 2, the above training process needs to be executed twice, and so on.
[0108] In addition, the model training configuration information also includes: a training result accuracy threshold. This threshold indicates the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information. For example, parameter 15 in Table 1 is a training result parameter, which can include the training result accuracy threshold. When the accuracy of the AI model trained by the first node reaches this threshold, training can stop. The central node can pre-set the training result accuracy threshold and then include it in the model training configuration information to send to the first node, thus controlling the model training of the first node. Training result accuracy refers to the proportion of successfully identified samples out of the total identified samples. For example, if the total number of samples is 100, and 90 samples are identified, and 80 of those 90 samples are successfully identified, then the accuracy is 80 / 90.
[0109] For example, taking the training of a CNN using supervised learning as an example, the model training process mainly includes: the collection and labeling stage of the training dataset, the training stage, and the validation stage.
[0110] In the training dataset collection and labeling phase, this embodiment of the application can use the SA engine to obtain the training set for training the CNN model, eliminating the need for manual intervention. This not only improves efficiency but also reduces human resources. Deep learning training relies on large labeled datasets, which not only need labeling but also need to be updated regularly to ensure the learning of new features. For example, based on the successful recognition results of the SA engine, the messages to be trained can be labeled to form a training dataset, which can also be updated regularly.
[0111] CNN training can employ the BP algorithm, also known as Back Progagation or Error Back Progagation. The basic idea of the BP algorithm is forward propagation and backward propagation. Forward propagation involves the input sample being fed into the input layer, processed layer by layer through each hidden layer, and then propagated to the output layer. If the actual output of the output layer does not match the expected output, backward propagation of the error begins. Back propagation involves propagating the output back to the input layer layer by layer in some form, distributing the error among all units in each layer, thus obtaining the error of each unit. This error is used as the basis for adjusting the weights of individual units. Network learning is completed during the weight modification process. When the error reaches the expected value, network learning ends. These two stages are iterated repeatedly until the network's response to the input reaches the predetermined target range.
[0112] For example, a CNN may include an input layer, hidden layers, and an output layer. The hidden layers may include multiple layers, which is not limited in this application. The input training samples are fed into the model to be trained through the input layer, processed layer by layer by the hidden layers, and then passed to the output layer. If the actual output of the output layer does not match the expected output, backpropagation of the error begins. Backpropagation involves propagating the output back to the input layer layer by layer in some form, distributing the error among each layer to obtain the error of each layer. This error serves as the basis for adjusting the weights of each layer. The training process involves modifying the weights multiple times to finally obtain the CNN model. The training process ends when the error reaches the expected value.
[0113] This application embodiment can employ cross-validation to train the model. The training dataset can be divided into two parts: one part is used to train the model (as training samples), and the other part is used to verify the accuracy of the network model (verification samples). After a CNN model is trained, the verification samples are used to verify whether the trained model can accurately identify the data stream and provide the recognition accuracy. When the recognition accuracy reaches a set threshold, the CNN model can be determined to be suitable for subsequent feature recognition. If the recognition accuracy does not reach the set threshold, training can continue until the recognition accuracy reaches the set threshold.
[0114] Furthermore, after completing model training and validation to obtain the AI model, the AI model obtained from training and validation can be used to perform feature recognition on the Unknown data stream to obtain classification results.
[0115] For example, taking supervised learning to train a CNN as an example, the model training process can also include an inference and recognition stage. For instance, in the inference and recognition process, the input is an unlabeled dataset (messages), i.e., the dataset to be recognized, and the output is the label classification of the dataset (messages), i.e., the recognition result. For example, the dataset to be recognized can be divided into two parts: one part consists of data messages whose recognition results have been labeled by the SA engine, which can be used to verify the functionality between the SA engine and the AI model, proving the accuracy of the AI model's recognition; the other part consists of data messages whose recognition results have not been labeled by the SA engine, which can be used to find messages that demonstrate the differences in capabilities between the SA engine and the AI model.
[0116] S303, the first node of the at least two first nodes determines the first model parameter update information;
[0117] The first model parameter update information is the model parameter update information after AI model training based on the local data of the first node and the model training configuration information.
[0118] For example, the first node can locally back up the received AI model common model as old_model, and then train it to obtain the new_model. The parameter update information of the first model is calculated as grad_n = new_model - old_model.
[0119] For example, the process of calculating the first model parameter update information includes: according to the model structure, expanding the model parameters of the trained AI model into a one-dimensional array Y, expanding the model parameters of the untrained AI model into a one-dimensional array X, and using the one-dimensional data Z obtained by subtracting X from Y as the first model parameter update. For example, assuming the data X is (1, 2, 3, 4, 5, 6) and the array Y is (2, 3, 4, 5, 6, 7), then the data Z = YX, which is (1, 1, 1, 1, 1, 1).
[0120] S304, the first node of the at least two first nodes sends first model parameter update information; correspondingly, the central node receives at least two first model parameter update information from the at least two first nodes;
[0121] The first model parameter update information is the model parameter update information after training the AI model based on the local data of the first node and the model training configuration information.
[0122] For example, the first model parameter update information is sent after the model parameters of the trained AI model have been successfully verified based on the model training configuration information.
[0123] Alternatively, the first node can compress the first model parameter update information and send it to the central node, which can then decompress it to obtain the first model parameter update information; or the first node can encrypt the first model parameter update information and send it to the central node, which can then decrypt it to obtain the first model parameter update information; or the first node can compress and encrypt the first model parameter update information and send it to the central node, which can then decompress and decrypt it to obtain the first model parameter update information.
[0124] For example, the first model parameter update information is sent after the model parameters of the trained AI model have been successfully verified based on the model training configuration information. That is, after the first node calculates the first model parameter update information, it needs to verify whether the structure of the trained AI model is consistent with that of the AI model before training. If the structure is consistent, the first model parameter update information is sent to the central node; if the structure is inconsistent, an error is reported. This ensures that the structure of the AI model before and after training is completely consistent, avoiding any impact on the training effect.
[0125] For example, the consistency of the AI model's structure before and after training can be verified based on the model training configuration information. For instance, some or all of the parameters 3 to 5, 10 to 13, and 16 to 30 in Table 1 can be used as verification parameters. That is, comparing whether these parameters of the AI model before and after training are consistent. When these parameters are completely consistent, it is determined that the structure of the AI model before and after training is consistent; when these parameters are not completely consistent, it is determined that the structure of the AI model before and after training is inconsistent.
[0126] In addition, the first node can also send other parameter values to the central node simultaneously, such as training time, training data volume, training result accuracy, and training result recall. Training result recall refers to the proportion of identified samples out of the total samples. For example, if the total samples are 100 and 90 samples are identified, the recall rate is 90 / 100.
[0127] The first node can send the first model parameter update information to the central node as the result of the task execution. For example, the first node sends the execution result information of the training task to the central node; this execution result information includes: task ID, task execution success, first model parameter update information, and may also include training time, training data volume, training result accuracy, training result recall, etc. As another example, when training fails, the execution result information may include: task ID, task execution failure, reason for failure, etc. This execution result information can be sent after compression and / or encryption.
[0128] S305, the central node uses a preset aggregation algorithm to aggregate at least two first model parameter update information to obtain second model parameter update information;
[0129] For example, the preset aggregation algorithm can be an averaging algorithm, a weighted averaging algorithm, a Federated Averaging algorithm, or a stochastic gradient descent (SVRG) algorithm, etc. The following section uses the weighted averaging algorithm as an example to introduce the aggregation calculation:
[0130] For example, suppose there are K distributed first nodes, and the dataset of each first node is P. k That is (x i y i ), i∈P k Sample size n k =|P k The total sample size of each first node is
[0131]
[0132]
[0133] Assume that the central node will share the model parameters ω at time t. t The data is distributed to each first node, and the model update at the first node uses gradient descent:
[0134]
[0135] The update process of the central node model can be achieved through model aggregation:
[0136]
[0137]
[0138] Alternatively, the model update amount Δω of each first node can be used. k To aggregate model:
[0139]
[0140]
[0141] Subsequent central nodes use ω t+1 The updated model is then distributed to each first node, enabling the aggregation and enhancement of the local model through the federated learning mechanism to achieve business objectives.
[0142] In the above formulas (1), (2), (3), (4), (5), (6), and (7), i represents the sample number, ranging from 1 to n, n represents the total number of samples, lowercase k represents the number of the first node, and K represents the number of the first nodes, ranging from 1 to uppercase K. k Let ω represent the number of samples at node k, t represent time t, t+1 represent the next time t, and ω represent the model parameter update. t ω represents the model parameter update at time t. t+1 This represents the model parameter update at time t+1. f represents the model parameter update of the first node k at time t+1. i (ω) represents the model parameter update corresponding to sample i, F k (ω) represents the model parameter update for the first node k, F k (ω t f(ω) represents the model parameter update of the first node k at time t, and f(ω) represents the model parameter update after aggregation. t ) represents the model parameter update after convergence at time t, α represents the preset coefficients, and gk represents
[0143] S306, the central node sends the second model parameter update information to the first node among the at least two first nodes; correspondingly, the first node among the at least two first nodes receives the second model parameter update information.
[0144] The second model parameter update information is obtained based on the at least two first model parameter update information, and the second model parameter update information is used to update the model parameters of the AI model of the first node.
[0145] For example, the method steps performed by the first node can be executed by an application (APP) deployed on a cloud platform or edge computing platform.
[0146] In one possible implementation, after receiving the model training message and before sending the first model parameter update information, the method further includes: receiving the first model parameter update information from the first node through the server module of the APP; sending the first model parameter update information includes: sending the first model parameter update information through the client module of the APP.
[0147] In another embodiment of this application, after the first node receives the second model parameter update information, the method may further include: the first node obtaining the identification result and message feature rules for identifying the data stream based on the AI model updated with the second model parameter update information; and the first node updating the service-aware SA feature library according to the message feature rules.
[0148] Alternatively, the central node can compress the second model parameter update information before sending it, and the first node can decompress it to obtain the second model parameter update information; or the central node can encrypt the second model parameter update information before sending it, and the first node can decrypt it to obtain the second model parameter update information; or the central node can compress and encrypt the second model parameter update information before sending it, and the first node can decompress and decrypt it to obtain the second model parameter update information.
[0149] S307, the first node of the at least two first nodes updates the AI model before this training according to the second model parameter update information.
[0150] For example, suppose the second model parameter update information is a one-dimensional array W, where W is (4, 5, 6, 7, 8, 9), and the model parameters of the AI model before training are expanded into a one-dimensional array X, which is (1, 2, 3, 4, 5, 6). Then the calculated one-dimensional array V = W + X should be (5, 6, 7, 8, 9, 10). Then, the model parameters of the AI model are restored according to this array, and the AI model before training is updated based on the restored model parameters.
[0151] In one possible implementation, the above steps S302 to S307 can be performed iteratively until the model convergence condition is met, ending federated learning and obtaining the final AI model.
[0152] In addition, during the above process, the first node can send heartbeat messages to the central node periodically or in real time. The heartbeat messages can include the status information of the training task, such as received, running, completed, failed, etc.
[0153] Furthermore, after the federated learning process is complete, if the first node needs to go offline, it can send a registration request message to the central node. The central node will then register the first node and send a registration success response message to the first node, at which point the first node will successfully go offline.
[0154] In another embodiment of this application, based on the above, the method may further include:
[0155] S308, the first node trains the updated AI model based on local data and model training configuration information.
[0156] In this way, the first node uses local data to train the AI model after federated learning. This not only improves the recognition rate of the AI model through federated learning, but also makes the trained AI model more adaptable to local data, further improving the recognition rate and meeting local business needs.
[0157] In another embodiment of this application, based on the above, the method may further include:
[0158] S309, The first node identifies the data stream based on the updated AI model, including the identification results and message feature rules;
[0159] S310, the first node updates the Service Awareness (SA) feature library according to the message feature rules.
[0160] Among these features, characteristic rules can be used to achieve rapid SA (Service Provider) identification. Compared with AI models, using an SA feature library is faster. In this embodiment, the acquired message characteristic rules are quickly added to the SA feature library, ensuring that the SA has the ability to quickly identify the traffic that meets the added characteristic rules, without needing to undergo further AI model identification. Furthermore, extracting message characteristic rules during the data flow identification process using an AI model meets the requirements of automated product operation and maintenance, eliminating the need for manual upgrades to the SA feature library, thus improving efficiency and reducing costs.
[0161] In another embodiment of this application, based on the above, the method may further include:
[0162] S311, the central node sends the second model parameter update information and the AI model to the second node;
[0163] The second model parameter update information is used to update the model parameters of the AI model of the second node.
[0164] S312, the second node updates the AI model based on the second model parameter update information.
[0165] In this context, the second node is the node that did not participate in federated learning. Therefore, the federated model obtained through federated learning can be directly applied to non-federated scenarios, meaning the AI model obtained through federated learning can be applied to the second node that did not participate in federated learning, thereby improving the recognition rate of the second node's AI model.
[0166] For example, if the second node requires it, the model carrying these parameters can be directly exported to the non-federated node (i.e., the second node). By reading these parameters, the non-federated node can fully understand the AI model structure, and the AI model can run automatically. These parameters can include parameters 4 to 5, 14, and 16 to 33 in Table 1.
[0167] For example, the second node can also directly modify the federated model before use, facilitating the use of different models on different nodes. For instance, if the business performance on some nodes (i.e., the second node) is not as expected after using the federated model, the model parameters of the federated model can be modified. Alternatively, the federated model can be taken offline and run on non-federated nodes by modifying the model parameters to achieve better business performance. The model parameters that can be modified here can include parameters 4 to 5, 14, and 16 to 33 in Table 1.
[0168] The technical solution provided in this application satisfies data privacy protection requirements because the first model parameter update information is obtained by training the first node using local data and model training configuration information. Furthermore, since the second model parameter update information is obtained based on at least two first model parameter update information from at least two first nodes, the recognition rate of the AI model updated using the second model parameter update information is higher, thereby maximizing the recognition rate of the AI model. Additionally, this application embodiment can also distribute training computing power among nodes, avoiding problems such as excessively long training times and high performance consumption for a single node.
[0169] In another embodiment of this application, based on the above, such as Figure 4 As shown, prior to S301, the method may further include:
[0170] S401, the first node sends a registration request message to the central node; correspondingly, the central node receives the registration request message.
[0171] It should be noted that there can be multiple first nodes participating in model training. Figure 3 Only one first node is shown in the figure. This is only an illustrative example and does not limit the embodiments of this application.
[0172] For example, the registration request message may include the name, identifier, etc. of the first node, or it may also include information such as the amount of local data of the first node.
[0173] S402 After successful registration, the central node sends a registration success response message to the first node; correspondingly, the first node receives the registration success response message.
[0174] When registration fails, the central node can send a registration failure response message to the first node; correspondingly, the first node receives the registration failure response message.
[0175] After successful registration, the first node can continuously send heartbeat messages to the central node.
[0176] S403, the central node performs initialization settings;
[0177] For example, the initialization settings can be performed by the management unit of the central node. The initialization settings may include:
[0178] 1. Select the first node to participate in this round of federated learning;
[0179] In one possible implementation, at least two first nodes participating in federated learning can be selected randomly or according to preset rules from multiple registered first nodes.
[0180] 2. Configure the aggregation algorithm;
[0181] 3. Configure model training settings;
[0182] For example, parameters such as the number of iterations for model training can be set, such as setting some or all of the parameters in Table 1;
[0183] 4. Create a federated learning instance and select SA as the instance type;
[0184] For example, select this instance for AISA traffic identification.
[0185] 5. Initialize an AI model.
[0186] For example, initializing an AISA traffic identification model, or injecting pre-trained model parameters and weights into an AI model. Initializing an AI model means injecting initial model parameters and weights into that AI model.
[0187] S404, the central node sends a training task notification message to at least two selected first nodes; the first node of the at least two first nodes receives the training task notification message accordingly.
[0188] The training task notification message may include a task ID, which is used to notify the first node of the training task.
[0189] S405, the first node of the at least two first nodes sends a training task query message to the central node;
[0190] The training task query message may include a task ID, which is used to query the central node for training tasks.
[0191] As an example, in S201, the model training message can be carried in the training task message and sent to the first node. For example, the central node sends training task information to the at least two first nodes; correspondingly, the first node among the at least two first nodes receives the training task information; wherein, the training task information includes the initial AI model and model training configuration information. The AI model and model training configuration information have been described in the previous embodiment and will not be repeated here.
[0192] In addition, the central node can compress and / or encrypt the training task information before sending it to the first node. After receiving the information, the first node decompresses and / or decrypts it to obtain the AI model and model training configuration information.
[0193] Furthermore, during the training process, the first node can continuously send heartbeat messages to the central node to inform the central node of the status of the training task, such as receiving the task, the task running, the task completed, and the task failing.
[0194] For example, in S204, the first model parameter update information can be carried in the task execution result and sent to the central node. For instance, the first of the at least two first nodes sends the task execution result to the central node; the corresponding central node receives at least two task execution results sent by the at least two first nodes; wherein the task execution result includes the first model parameter update information.
[0195] Furthermore, the task execution results may also include one or more of the following: training time, training data volume, training result accuracy, or training result recall, etc.
[0196] The first node can compress and / or encrypt the task execution result before sending it to the central node, and the central node can decompress and / or decrypt it to obtain the task execution result.
[0197] In one possible implementation, another round of training can be performed, that is, jump to step 305 to train the updated AI model again until the model converges and the federated learning ends.
[0198] The technical solution provided by this invention uses a federated learning mechanism to transmit the updated model parameters of each local first node after local training between different nodes. This allows a first node to enhance its local recognition ability even if the required APP traffic is too low or the recognition effect is poor, by transmitting the recognition capabilities of other first nodes.
[0199] Furthermore, when a first node needs to identify a large number of apps and local training alone cannot complete the identification in a short time, federated learning of other first nodes can rapidly expand the identification capability, improve the training efficiency of single-node models, and enable rapid iteration. Additionally, the identification capability of new apps can be quickly transferred between different first nodes. For example, if sites A and B participate in federated learning to train an AI model, the identification capability of site A can be transferred to site B, allowing site B to have the same identification capability as site A even if a particular app has low traffic (small sample size) or no traffic.
[0200] Furthermore, federated learning involves transmitting model parameter updates between different nodes, which are unrelated to the original data, thus fully satisfying privacy requirements. For example, since site A's local data does not leave site A, site B can improve its recognition capabilities without relying on source data from site A.
[0201] Furthermore, federated learning can also distribute training computational power, enabling rapid iteration of recognition capabilities for key applications and time periods. For example, site B can share the computational power of site A, resulting in improved performance compared to site A training alone.
[0202] exist Figure 1 Based on this, embodiments of this application provide an architecture for a model training system, such as... Figure 5 As shown, a federated learning server FLS1101 and a convergence unit 1102 are set at the central node 110, and a federated learning client FLC and a training unit are set at each first node. For example, FLC1201 and training unit 1202 are set at the first node 120, and FLC1211 and training unit 1212 are set at the first node 121.
[0203] The FLS1101 communicates with each FLC1201 and 1211 via wired or wireless connections.
[0204] like Figure 5 As shown, the aggregation unit 1102 can be any unit that supports data aggregation, and can be co-located with FLS1101 within the central node 110 to work together with FLS1101 to achieve federated learning. The training units 1201 and 1212 can be any units that support AI model training, and can be co-located with FLC1201 and 1211 respectively within the first node to work together with FLS1101 to achieve federated learning.
[0205] For example, FLS1101 and aggregation unit 1101 can be implemented using different container services, or they can be implemented using one or more virtual machines (VMs), or they can be implemented using one or more processors, or they can be implemented using one or more computers.
[0206] For example, FLC1201, 1211 and training units 1201, 1212 can be implemented using different container services, or they can be implemented using one or more virtual machines (VMs), or they can be implemented using one or more processors, or they can be implemented using one or more computers.
[0207] As an example, training units 1201 and 1212 may be artificial intelligence service awareness (AISA) deployed in the first nodes 120 and 121 respectively. AISA may also be called artificial intelligence recognition function, or it may be named by other names. For the sake of description in this embodiment, it is referred to as AISA.
[0208] For example, AISA can be used to classify collected data streams based on the SA feature library to obtain classification results. The SA feature library can be located within AISA or outside of AISA, connected via an interface. AISA can also include an SA engine. The SA engine is used to perform feature recognition on the collected data streams based on the SA feature library. AISA can use the SA engine to perform feature recognition on the collected data streams, obtain classification results for the data streams, and then label the data streams according to the classification results; then, the labeled data streams are used as the training dataset for the AI model, and AISA trains the AI model based on the training dataset.
[0209] For example, nodes 120 and 121 can also deploy SA recognition engines, such as the SA@AI engine. The SA recognition engine can submit a model training request to AISA as an application that needs to perform SA recognition, configure the selected protocol, and then AISA will collect data, train the model, extract rules, output the recognition results and rules, and update the SA feature library.
[0210] For example, AI models can be deployed on a cloud platform. The cloud platform can register with the Federated Learning Server (FLS) on a central node in the cloud and communicate with each other via status messages. After the AI model collects data and trains the model, it outputs recognition results. The cloud platform can forward interactive data, such as the recognition results from the AI model, to the FLS. The FLS accepts the model parameters uploaded by the cloud platform, performs model aggregation and fusion, and then distributes the shared model back to the cloud platform. In this way, the aggregation and fusion of AI models enhances the model's capabilities.
[0211] As an example, FLC1201 and 1211 can receive data from local nodes and forward it to the central node 110. AISA can also use FLC1201 and 1211 to synchronize status messages with the central node FLS1101, export, upload, and download AI model parameter updates, and upload and download training time, data volume, and recognition results (recall, precision), etc. At the central node 110, the federated learning server FLS1101 is responsible for receiving data from FLC1201 and 1211, detecting the status of the first nodes 120 and 121, distributing training tasks, receiving model parameter updates uploaded by each distributed first node 120 and 121, aggregating and merging model parameter updates, and distributing the aggregated model parameter updates, etc.
[0212] As another example, server and client modules can be configured in FLC1201 and 1211, client modules in training units 1201 and 1212, and a server module in FLS1101. The server modules in FLC1201 and 1211 are configured to connect to the client modules in training units 1202 and 1212, respectively, handling data transmission between FLC1201 and 1211 and their respective training units. The client modules in FLC1201 and 1211 are configured to connect to the server module in FLS1101, handling data transmission between FLC1201 / 1211 and FLS1101.
[0213] For example Figure 5 As shown, server module 11011 can be set in FLS1101, client module 12011 and server module 12012 can be set in FLC1201, client module 12021 can be set in training unit 1202, client module 12111 and server module 12112 can be set in FLC1211, and client module 12121 can be set in training unit 1212.
[0214] For example, client modules 12011, 12021, 12111, and 12121 can be HTTP / HTTPS clients, and server modules 11011, 12012, and 12112 can be HTTP / HTTPS servers.
[0215] In one example, FLC1201 and 1211 can be deployed as applications (APPs) on the first nodes 120 and 121, respectively. The first nodes 120 and 121 can be nodes of an edge computing platform or a cloud platform. For example, FLC1201 and 1211 can be deployed as applications (APPs) on an edge computing platform. Alternatively, FLC1201 and 1211 can be directly deployed as an application on the edge computing platform or cloud platform, and then FLC1201 and 1211 can act as proxies for the connection between the first nodes 120 and 121 and FLS110.
[0216] In another example, FLC1201 and 1211 can be deployed as statically linked libraries on the first nodes 120 and 121, respectively. For instance, FLC1201 and 1211 can be integrated into a virtual machine (VM) deploying AISA. The VM can then configure interfaces for external connections for FLC1201 and 1211. Furthermore, the IP address and parameters for connecting FLC1201 and 1211 to FLS1101 can be configured through the VM's login portal. These parameters include the name and identifier of FLC1201 and 1211, and the username and password used for registration with FLS1101.
[0217] based on Figure 5 The architecture shown is for Figure 3 The method shown allows FLS1101 to perform the operations performed by the central node 110, and FLC1201 and 1211 to perform the operations performed by the first nodes 120 and 121, respectively.
[0218] based on Figure 5 The architecture shown is for Figure 4 The method shown allows FLS1101 and aggregation unit 1102 to cooperate in executing the operations performed by central node 110. FLS1101 is responsible for receiving and sending data, while aggregation unit 1102 is responsible for aggregating at least two first model parameter update information using a preset aggregation algorithm to obtain second model parameter update information. FLC1201 and 1211, along with training units 1202 and 1212, can respectively execute the operations performed by first nodes 120 and 121. Training units 1202 and 1212 are responsible for training the AI model, while other operations are handled by FLC1201 and 1211.
[0219] The following examples illustrate the effects achieved by the solutions provided in the embodiments of this application through specific application scenarios.
[0220] Scenario 1: Federated Learning Experiment with Few Sample Nodes
[0221] In this scenario, the sample sizes of the three local nodes (i.e., the first node, hereinafter referred to as local) are 10%, 30%, and 60% of the total sample size, respectively. The total sample set is randomly distributed among the local nodes proportionally. The sample data distribution of each local node in this scenario is as follows: Figure 6 As shown, since local1 is only allocated 10% of the total data, it becomes a minority node. A federated learning experiment was conducted to verify whether minority nodes can improve the model's recognition ability after federated learning.
[0222] like Figure 7 As shown, the recognition accuracies of the models trained on the three local nodes based on local data are 76.0%, 90.9%, and 95.6%, respectively. When trained using the full dataset (i.e., the total sample set), the recognition accuracy of the model reaches 97.1%. When performing federated learning on the three local nodes, the model parameter fusion strategy performs parameter fusion once per epoch, resulting in a federated model with a recognition accuracy of 95.7%. This demonstrates that the recognition accuracy of the model trained locally is lower due to the smaller number of training samples in local1, while the recognition accuracy of the model trained locally in local3 is higher due to the larger number of samples. Through federated learning, each local node obtains a federated model with higher recognition accuracy, and the recognition accuracy is improved to varying degrees compared to local training. In particular, the recognition accuracy of the model trained on the small sample node, local1, is significantly improved through federated learning.
[0223] Scenario 2: Federated Learning Experiments for Few-Sample Applications
[0224] In Scenario 2, the number of samples is similar across different locales, but the distribution of application samples differs. Each locale has some applications with small sample sizes, but these small applications have a sufficiently large number of samples in other locales. The sample data distribution across different locales in Scenario 2 is as follows: Figure 8 As shown, a federated learning experiment was conducted to verify whether the small sample applications of each local application can achieve good recognition capabilities after federated learning.
[0225] like Figure 9As shown, the recognition accuracies of the models trained on the three local datasets were 83.6%, 81.9%, and 82.8%, respectively. When trained on the full dataset, the recognition accuracy reached 97.1%. After federated learning, the recognition accuracy of the federated model reached 95.6%. This demonstrates that because each local dataset contains applications with small sample sizes, the local training model's recognition accuracy for these small sample applications is not high. However, after federated learning, the recognition accuracy and recall of the federated model for these small sample applications are significantly improved. For example... Figure 10 The diagram shows scenario two: the recall rates of small sample applications in each local model and the recall rate after federated learning. The recall rates of small sample applications in each local model are low, while the recall rate of the federated model is significantly improved. For example, application A has a recall rate of only 46.9% in local model 1, but a recall rate of 95.6% in the federated model; application B has a recall rate of only 13.8% in local model 1, but a recall rate of 94.4% in the federated model; and application C has a recall rate of only 49.9% in local model 2. For example, application D has a recall rate of only 7% in the local model, but a recall rate of 98.5% when using the federated model; application D has a recall rate of only 42.3% in the local model, but a recall rate of 97.6% when using the federated model; application E has a recall rate of only 32.1% in the local model, but a recall rate of 96.4% when using the federated model; application F has a recall rate of only 59.1% in the local model, but a recall rate of 97.3% when using the federated model.
[0226] Scenario 3: Federated Learning Experiment to Expand the Number of Recognition Applications
[0227] In Scenario 3, the number of samples is similar across different locales, but the sample distribution varies across applications. Some applications have no samples in one or two locales. The sample data distribution across different locales in Scenario 3 is as follows: Figure 11 As shown, a federated learning experiment was conducted to verify whether each local instance can expand the number of applications it can identify after federated learning.
[0228] like Figure 12As shown, the recognition accuracies of the models trained on the three local datasets are 87%, 87.5%, and 86.9%, respectively. When trained on the full dataset, the recognition accuracy reaches 98.4%. After federated learning, the recognition accuracy of the federated model reaches 97.9%. This demonstrates that when some applications lack training samples in each local dataset, the locally trained models lack the ability to recognize these applications, resulting in low overall recognition accuracy. Through federated learning, the resulting federated model achieves higher recognition accuracy and recall for applications without local samples, and the overall recognition accuracy of the federated model is significantly improved. For example... Figure 13 Scenario 3: Recall after applying federated learning when there are no samples in each locale. When some applications in each locale lack training samples, the recall is 0, but the recall of the federated model is significantly improved. For example, APP1 has a recall of 0 in locale 1, but the recall increases to 98.1% with the federated model. Similarly, APP2 has a recall of 0 in locale 2, but the recall increases to 97.0% with the federated model. APP3 has a recall of 0 in locale 3, but the recall increases to 98.5% with the federated model. APP4 has a recall of 0 in both locale 2 and locale 3, but the recall increases to 96.0% with the federated model. APP5 has a recall of 0 in both locale 1 and locale 3, but the recall increases to 99.3% with the federated model. APP6 has a recall of 0 in both locale 1 and locale 2, but the recall increases to 99.9% with the federated model.
[0229] Scenario 4: Fine-tuning Experiment of Federated Learning Based on Federated Model
[0230] After federated learning is completed, a federated model is obtained. When new application categories are generated in a certain locale, the recognition rate can be improved by either fine-tuning the model using federated learning or re-initializing and then re-training it. The effects of these two methods on federated learning are then tested. Fine-tuning (Federated-finefune) refers to training a pre-trained model for several more rounds until it converges; in this scenario, it refers to training the federated model for several more rounds until it converges. Re-initialization (Federated-init) refers to initializing the federated model.
[0231] For example Figure 14As shown, based on federated learning, the federated model, after fine-tuning (i.e., after 500 rounds of federated training), achieves a recognition accuracy of 95.7%. Furthermore, after 2000 rounds of federated training on the initialized federated model, the recognition accuracy also reaches 95.7%. This demonstrates that fine-tuning of the federated model based on federated learning can achieve the same model accuracy as retraining from initialization with fewer iterations.
[0232] For example Figure 15 As shown, for the new application APP_T, the recognition accuracy of the model after fine-tuning the federated model is 99.4%, and the recognition accuracy of the model after federated training on the initialized federated model is 99.4%. For the new application APP_D, the recognition accuracy of the model after fine-tuning the federated model is 99.8%, and the recognition accuracy of the model after federated training on the initialized federated model is 99.8%. For the new application APP_Y, the recognition accuracy of the model after fine-tuning the federated model is 99.3%, and the recognition accuracy of the model after federated training on the initialized federated model is 99.2%. For the new application APP_M, the recognition accuracy of the model after fine-tuning the federated model is 99.7%, and the recognition accuracy of the model after federated training on the initialized federated model is 99.5%. Therefore, it can be seen that the model trained by fine-tuning the federated model based on federated learning can achieve the same or even higher recognition accuracy for new applications as the model after initial retraining.
[0233] The technical solutions provided in this application can enable AI-based SA (Service Assist) technology to train AI models capable of high-precision and high-performance recognition in networks, even when training samples are scarce or computing power is insufficient, and privacy protection is required. For example, traditional AI-based SA requires a large number of data samples for model training, but the means of reasonably collecting data samples are very limited, and if the requirements are not met, the overall performance will be affected. The solutions provided in this application can improve the overall recognition capability of AI models even when training samples are scarce. As another example, traditional AI-based traffic identification generally has a long computation time, which can easily lead to significant performance consumption on a single node. The solutions provided in this application can achieve distributed model training, avoiding the performance bottleneck of a single node. Furthermore, with the increasing demand for network data traffic privacy protection, the solutions provided in this application can achieve distributed computing of AI technology in networks, protecting the security of original data and providing secure and reliable intelligent traffic identification services.
[0234] Based on the above, Figure 16 The structure of a model training apparatus provided in an embodiment of this application is shown. For example... Figure 16 As shown, this model training device can be deployed in Figure 1 or Figure 2 The central node shown, or deployed in Figure 5 FLS in the middle.
[0235] The model training device includes a first transmitting module 1601 and a receiving module 1602.
[0236] The first sending module 1601 is used to send model training messages to at least two first nodes and send second model parameter update information to the first node among the at least two first nodes. The model training messages include an artificial intelligence (AI) model and model training configuration information. The AI model is used to identify the category to which the data stream belongs. The second model parameter update information is obtained based on the at least two first model parameter update information and is used to update the model parameters of the AI model of the first node.
[0237] The receiving module 1602 is used to receive at least two first model parameter update messages from the at least two first nodes. The first model parameter update messages are model parameter update messages after training the AI model based on the local data of the first node corresponding to the first model parameter update messages and the model training configuration information.
[0238] For example, the category to which the data stream belongs includes at least one of the following: the application to which the data stream belongs; the type or protocol to which the service content of the data stream belongs; and the message characteristic rules of the data stream.
[0239] In one possible implementation, the apparatus further includes: a second sending module 1603, configured to send the second model parameter update information and the AI model to the second node; the second model parameter update information is used to update the model parameters of the AI model of the second node.
[0240] In addition, the model training configuration information may also include: a training result accuracy threshold; the training result accuracy threshold is used to indicate the training result accuracy of the AI model trained by the first node based on local data and the model training configuration information.
[0241] Based on the above, Figure 17 The structure of a model training apparatus provided in an embodiment of this application is shown. For example... Figure 17 As shown, this model training device can be deployed in Figure 1 or Figure 2 The first node shown, or deployed in Figure 5 FLC in the middle.
[0242] The model training device includes a receiving module 1701 and a transmitting module 1702.
[0243] The receiving module 1701 is used to receive model training messages and second model parameter update information. The model training messages include an AI model and model training configuration information. The AI model is used to identify the category to which the data stream belongs. The second model parameter update information is obtained based on at least two first model parameter update information of at least two first nodes. The second model parameter update information is used to update the model parameters of the AI model of the first node.
[0244] The sending module 1702 is used to send first model parameter update information, which is the model parameter update information after training the AI model based on the local data of the first node and the model training configuration information.
[0245] For example, the category to which the data stream belongs includes at least one of the following: the application to which the data stream belongs; the type or protocol to which the service content of the data stream belongs; and the message characteristic rules of the data stream.
[0246] For example, the apparatus may further include: a training module 1703, configured to determine that the first node trains the AI model updated with the second model parameter update information based on local data.
[0247] For example, the receiving module and the sending module can be modules of an application (APP) deployed on a cloud platform or edge computing platform.
[0248] For example, the receiving module 1701 and the sending module 1702 are client modules of the APP; the APP also includes a server module; the server module is used to send the first model parameter update information to the first node.
[0249] In one possible implementation, the first model parameter update information is sent after the model parameters of the trained AI model have been successfully verified based on the model training configuration information.
[0250] In addition, the model training configuration information also includes: a training result accuracy threshold; the training result accuracy threshold is used to indicate the training result accuracy of the AI model trained by the first node based on local data and the model training configuration information.
[0251] In one example, the device may further include:
[0252] The acquisition module 1704 is used to acquire the identification results and message feature rules for identifying data streams based on the AI model updated using the second model parameter update information.
[0253] The update module 1705 is used to update the Service Awareness (SA) feature library according to the message feature rules.
[0254] The module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into a single processor, exist as separate physical units, or be integrated into a single unit. The integrated units described above can be implemented in hardware or as software functional units.
[0255] Based on the above, Figure 18 The structure of a model training apparatus provided in an embodiment of this application is shown. For example... Figure 18 As shown, the model training device 1800 can be deployed in Figure 1 or Figure 2 The central node shown, or deployed in Figure 5 The FLS in the model. Alternatively, the model training device 1800 can be deployed in... Figure 1 or Figure 2 The first node shown, or deployed in Figure 5 FLC in the middle.
[0256] The model training device 1800 may include a communication interface 1810 and a processor 1820. Optionally, the model training device 1800 may also include a memory 1830. The memory 1830 may be located inside or outside the model training device. The functions implemented by the model training device in the above embodiments can all be implemented by the processor 1820. The processor 1820 receives data streams through the communication interface 1810 and uses them to implement the model training method described in any of the above embodiments. During implementation, each step of the processing flow can be completed by the integrated logic circuits in the hardware of the processor 1820 or by software instructions to complete the model training method described in any of the above embodiments. For simplicity, further details are omitted here. The program code executed by the processor 1820 to implement the model training method described in any of the above embodiments can be stored in the memory 1830. The memory 1830 and the processor 1820 are coupled.
[0257] The processors involved in the embodiments of this application can be general-purpose processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0258] The coupling in the embodiments of this application is an indirect coupling or communication connection between devices, modules, or modules, which can be electrical, mechanical, or other forms, and is used for information interaction between devices, modules, or modules.
[0259] The processor may work in conjunction with memory. Memory can be non-volatile memory, such as hard disk drives (HDDs) or solid-state drives (SSDs), or it can be volatile memory, such as random-access memory (RAM). Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited to this.
[0260] This application does not limit the specific connection medium between the communication interface, processor, and memory. For example, the memory, processor, and communication interface can be connected via a bus. The bus can be categorized as an address bus, data bus, control bus, etc.
[0261] Based on the above embodiments, this application also provides a computer storage medium storing a software program. When read and executed by one or more processors, the software program can implement the model training method provided in any of the above embodiments. The computer storage medium may include various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory, random access memory, magnetic disk, or optical disk.
[0262] Based on the above embodiments, this application also provides a computer program product containing instructions, which, when run on a computer, causes the computer to execute the model training method provided in any of the above embodiments.
[0263] Based on the above embodiments, this application also provides a chip, which includes a processor for implementing the model training methods provided in any one or more of the above embodiments. Optionally, the chip further includes a memory for storing necessary program instructions and data executed by the processor. This chip can be composed of individual chips or may include chips and other discrete devices.
[0264] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0265] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0266] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0267] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0268] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
Claims
1. A model training method, characterized in that, The method includes: Send a model training message to at least two first nodes. The model training message includes an artificial intelligence (AI) model and model training configuration information. The AI model is used to identify the category to which the data stream belongs. Receive at least two first model parameter update messages from the at least two first nodes, wherein the first model parameter update messages are model parameter update messages after training the AI model based on the local data of the first node corresponding to the first model parameter update messages and the model training configuration information; Send second model parameter update information to the first node among the at least two first nodes. The second model parameter update information is obtained based on the at least two first model parameter update information. The second model parameter update information is used to update the model parameters of the AI model of the first node. The model training configuration information includes a protocol identifier list, the name of the model loader, and information about the model structure. The protocol identifier list, the name of the model loader, and the information about the model structure are used for loading the AI model and for the first node to verify the AI model trained based on the local data of the first node. The category to which the data stream belongs includes at least one of the following: The application to which the data stream belongs; the type or protocol to which the business content of the data stream belongs; the message characteristic rules of the data stream.
2. The method as described in claim 1, characterized in that, After receiving at least two first model parameter update information from the at least two first nodes, the method further includes: The second model parameter update information and the AI model are sent to the second node; the second model parameter update information is used to update the model parameters of the AI model of the second node.
3. The method as described in claim 1 or 2, characterized in that, The model training configuration information also includes: a training result accuracy threshold; The training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
4. A model training method, characterized in that, The method includes: Receive a model training message, which includes an AI model and model training configuration information, wherein the AI model is used to identify the category to which the data stream belongs; Send first model parameter update information, which is the model parameter update information after training the AI model based on the local data of the first node and the model training configuration information; Receive second model parameter update information, which is obtained based on at least two first model parameter update information of at least two first nodes, and the second model parameter update information is used to update the model parameters of the AI model of the first node; The model training configuration information includes a protocol identifier list, the name of the model loader, and information about the model structure. The protocol identifier list, the name of the model loader, and the information about the model structure are used for loading the AI model and for the first node to verify the AI model trained based on the local data of the first node. The category to which the data stream belongs includes at least one of the following: The application to which the data stream belongs; the type or protocol to which the business content of the data stream belongs; the message characteristic rules of the data stream.
5. The method as described in claim 4, characterized in that, The method further includes: It is determined that the first node trains the AI model updated with the second model parameter update information based on local data.
6. The method as described in claim 4 or 5, characterized in that, The method is executed by an application (APP) deployed on a cloud platform or edge computing platform.
7. The method as described in claim 6, characterized in that, After receiving the model training message and before sending the first model parameter update information, the method further includes: The APP's server module receives the first model parameter update information from the first node; The sending of the first model parameter update information includes: The update information of the first model parameters is sent through the client module of the APP.
8. The method according to any one of claims 4 to 7, characterized in that, The first model parameter update information is sent after the model parameters of the trained AI model have been successfully verified based on the model training configuration information.
9. The method according to any one of claims 4 to 8, characterized in that, The model training configuration information also includes: a training result accuracy threshold; The training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
10. The method according to any one of claims 4 to 9, characterized in that, After receiving the second model parameter update information, the method further includes: Obtain the identification results and message feature rules for the data stream based on the AI model updated using the second model parameter update information; Update the Service Awareness (SA) feature library according to the message feature rules.
11. A model training device, characterized in that, The device includes: The first sending module is used to send model training messages to at least two first nodes and send second model parameter update information to the first node among the at least two first nodes. The model training messages include an artificial intelligence (AI) model and model training configuration information. The AI model is used to identify the category to which the data stream belongs. The second model parameter update information is obtained based on the at least two first model parameter update information and is used to update the model parameters of the AI model of the first node. The receiving module is configured to receive at least two first model parameter update messages from the at least two first nodes. The first model parameter update messages are model parameter update messages after training the AI model based on the local data of the first node corresponding to the first model parameter update messages and the model training configuration information. The model training configuration information includes a protocol identifier list, the name of the model loader, and information about the model structure. The protocol identifier list, the name of the model loader, and the information about the model structure are used for loading the AI model and for the first node to verify the AI model trained based on the local data of the first node. The category to which the data stream belongs includes at least one of the following: The application to which the data stream belongs; the type or protocol to which the business content of the data stream belongs; the message characteristic rules of the data stream.
12. The apparatus as claimed in claim 11, characterized in that, The device further includes: The second sending module is used to send the second model parameter update information and the AI model to the second node; the second model parameter update information is used to update the model parameters of the AI model of the second node.
13. The apparatus as claimed in claim 11 or 12, characterized in that, The model training configuration information also includes: a training result accuracy threshold; The training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
14. A model training device, characterized in that, The device includes: A receiving module is used to receive model training messages and second model parameter update information. The model training messages include an AI model and model training configuration information. The AI model is used to identify the category to which the data stream belongs. The second model parameter update information is obtained based on at least two first model parameter update information of at least two first nodes. The second model parameter update information is used to update the model parameters of the AI model of the first node. The sending module is used to send first model parameter update information, which is the model parameter update information after training the AI model based on the local data of the first node and the model training configuration information. The model training configuration information includes a protocol identifier list, the name of the model loader, and information about the model structure. The protocol identifier list, the name of the model loader, and the information about the model structure are used for loading the AI model and for the first node to verify the AI model trained based on the local data of the first node. The category to which the data stream belongs includes at least one of the following: The application to which the data stream belongs; the type or protocol to which the business content of the data stream belongs; the message characteristic rules of the data stream.
15. The apparatus as claimed in claim 14, characterized in that, The device further includes: The training module is used to determine whether the first node trains the AI model updated with the second model parameter update information based on local data.
16. The apparatus as claimed in claim 14 or 15, characterized in that, The receiving module and the sending module are modules of an application (APP) deployed on a cloud platform or edge computing platform.
17. The apparatus as claimed in claim 16, characterized in that, The receiving module and the sending module are client modules of the APP; The app also includes a server-side module; The server module is used to send the first model parameter update information to the first node.
18. The apparatus as claimed in any one of claims 14 to 17, characterized in that, The first model parameter update information is sent after the model parameters of the trained AI model have been successfully verified based on the model training configuration information.
19. The apparatus as claimed in any one of claims 14 to 18, characterized in that, The model training configuration information also includes: a training result accuracy threshold; The training result accuracy threshold is used to indicate the accuracy of the training result when the first node trains the AI model based on local data and the model training configuration information.
20. The apparatus as claimed in any one of claims 14 to 19, characterized in that, The device further includes: The acquisition module is used to acquire the recognition results and message feature rules for identifying the data stream based on the AI model updated using the second model parameter update information. The update module is used to update the Service Awareness (SA) feature library according to the message feature rules.
21. A model training device, characterized in that, Including the processor and memory, of which: The memory is used to store program code; The processor is configured to read and execute program code stored in the memory to implement the method as described in any one of claims 1 to 10.
22. A model training system, characterized in that, Includes the apparatus as claimed in any one of claims 11 to 13, and includes the apparatus as claimed in any one of claims 14 to 20.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a software program that, when read and executed by one or more processors, is used to implement the method described in any one of 1 to 10.
Citation Information
Patent Citations
Model training method and device
CN111612153A
Data stream classification method, device and system
CN112861894A