Model management method and system
By cleaning the data set before the AI model update and deleting the subset of data that is not used for update, the problem of pseudo-data impact during the AI model update process is solved, and the model's robustness and prediction performance are improved.
Patent Information
- Application Number
- CN202311796301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
During the update process, existing AI models may be affected by pseudo-data generated by cyber attack devices, resulting in reduced model robustness.
By cleaning the data set before the AI model is updated, deleting data subsets whose statistical distribution features are different from the training data set and are not used for updates, thereby improving the robustness of the model.
It effectively reduces the probability of pseudo-data in the dataset used for model updates, and improves the robustness and prediction performance of the model.
Smart Images

Figure CN120196943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a model management method and system. Background Art
[0002] To improve the intelligence and automation levels of networks, technologies such as artificial intelligence (AI) and machine learning (ML) are applied in more and more fields. For example, model building, deployment, training, and application can be performed based on an AI system.
[0003] Currently, to improve the robustness of an AI model, it is necessary to manage the life cycle of the model. For example, over time, the prediction accuracy of an AI model may decrease, that is, the prediction error becomes larger. Therefore, to solve this problem, it is usually necessary to regularly retrain the AI model with new data, that is, update the AI model, to maintain the prediction performance of the AI model. However, the dataset used to update the AI model may include data generated by network attack devices, which affects the robustness of the AI model. Summary of the Invention
[0004] Embodiments of this application provide a model management method and system for enhancing the robustness of a model.
[0005] In a first aspect, a model management method is provided. The method includes: determining a first algorithm according to a first probability; where the first probability is the probability that a first data subset exists in a first dataset, the statistical distribution characteristics of the first data subset are different from those of the training dataset used to train a first model, and the first data subset is used to update the first model; the first algorithm is used to delete a second data subset from the first dataset to obtain a second dataset, the statistical distribution characteristics of the second data subset are different from those of the training dataset, the second data subset is not used to update the first model, and the second dataset is used to update the first model.
[0006] In the embodiments of this application, before using the first dataset to update the first model, the first dataset can be cleaned, which helps reduce the probability that the dataset used for model update contains false data (such as data generated by network attack devices) and enhances the robustness of the model. In addition, the algorithm for cleaning the first dataset determined according to the probability that a first data subset with statistical distribution characteristics different from those of the training dataset for training the model exists in the first dataset can improve the accuracy of cleaning the second data and further enhance the robustness of the model.
[0007] The above method can be implemented by the model management functional entity or the model training functional entity. For example, if the above method is implemented by the model management functional entity, the specific implementation method is as follows:
[0008] In a possible implementation manner, the model management functional entity may also send a first indication message to the model training functional entity or receive a first indication message from the model training functional entity. The first indication message is used to indicate the first algorithm. The first algorithm may be determined by the model management functional entity or by the model training functional entity. If the first algorithm is determined by the model management functional entity, the model management functional entity may send the first indication message to the model training functional entity to instruct the model training functional entity to delete a second data subset in the first dataset based on the first algorithm. If the first algorithm is determined by the model training functional entity, the model training functional entity may send the first indication message to the model management functional entity to inform the model management functional entity of the first algorithm used when deleting the second data subset in the first dataset.
[0009] In a possible implementation manner, before sending the first indication message to the model training functional entity, the model management functional entity may also receive a second indication message from the model training functional entity. The second indication message is used to indicate the first probability. Since the first algorithm is determined based on the first probability, if the first algorithm is determined by the model management functional entity, the model management functional entity also needs to receive the first probability from the model training functional entity.
[0010] In a possible implementation manner, the model management functional entity may also determine the first algorithm based on the first probability. The first algorithm determined based on the first probability can improve the accuracy of cleaning the second data and further improve the robustness of the model.
[0011] In a possible implementation, the model management functional entity may also send third indication information to the model training functional entity, or receive third indication information from the model training functional entity. The third indication information is used to indicate a second algorithm, and the second algorithm is used to determine the first probability. The second algorithm may be determined by the model management functional entity or by the model training functional entity. If the second algorithm is determined by the model management functional entity, the model management functional entity may send the third indication information to the model training functional entity to indicate that the model training functional entity determines the first probability of the existence of the first data subset in the first data set based on the second algorithm. If the first algorithm is determined by the model training functional entity, the model training functional entity may send the third indication information to the model management functional entity to inform the model management functional entity of the first algorithm used when it determines the first probability. Among them, if the second algorithm is determined by the model training functional entity, the third indication information and the foregoing second indication information may be the same information or different information.
[0012] In a possible implementation, the model management functional entity may also send fourth indication information to the model inference functional entity, or receive fourth indication information from the model inference functional entity. The fourth indication information is used to indicate a third algorithm, and the third algorithm is used to delete the third data subset in the third data set to obtain the first data set. The third data set is collected by the model inference functional entity. The statistical distribution characteristics of the third data subset are different from those of the training data set, and the third data subset is not used to update the first model. By cleaning the third data set collected by the model inference functional entity, the amount of data sent by the model inference functional entity to the model training functional entity can be reduced, which helps to reduce the transmission overhead between entities and improve the data transmission speed.
[0013] In a possible implementation, the model management functional entity may also receive first information from the model training functional entity. The first information includes the second data subset and / or the value of the first parameter of the data in the second data subset. The value of the first parameter is determined based on the first algorithm, and the first parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model. The model training functional entity sends the deleted second data subset and / or the value of the first parameter of the data in the second data subset to the model management functional entity, so that the model management functional entity can learn the statistical distribution characteristics of the second data subset and optimize the first algorithm.
[0014] In a possible implementation, the model management functional entity may further optimize the first algorithm based on the second data subset and / or the value of the first parameter of the data in the second data subset. The first algorithm optimized based on the second data subset and / or the value of the first parameter of the data in the second data subset can improve the accuracy of data cleaning, thereby improving the robustness of the updated model.
[0015] In a possible implementation, the first parameter includes one or more of the following: the change amount of the loss value, which is used to indicate the influence degree of the data on the loss function of the first model; the second probability, which is used to indicate the probability that the data is abnormal data; or, the difference degree of the statistical distribution characteristics, which is used to indicate the difference degree of the statistical distribution characteristics of the data in the second data subset and the statistical distribution characteristics of the data in the training data subset. The first parameter can be any parameter that can measure the influence degree of the data on the model prediction accuracy. The parameters shown in the embodiments of the present application are only examples. In other embodiments, the first parameter may further include more or fewer parameters, and the embodiments of the present application do not limit the types of parameters included in the first parameter.
[0016] In a possible implementation, the first information further includes one or more of the following: the value of the second parameter, which is used to indicate the inference accuracy of the updated first model; the identifier of the first model; the version number of the updated first model; the update time of the first model; the data volume included in the third data set used to update the first model; or, the duration required for the updated first model to perform inference. The model training functional entity may send the relevant information of the model to the model management functional entity, so that the model management functional entity can adjust the strategy for managing the model based on the relevant information of the model.
[0017] In a possible implementation, the model management functional entity may further receive second information from the model inference functional entity, and the second information includes the value of the third parameter of the third data subset and / or the data in the third data subset. The model inference functional entity sends the value of the third parameter of the deleted third data subset and / or the data in the third data subset to the model management functional entity, so that the model management functional entity can learn the statistical distribution characteristics of the third data subset and optimize the third algorithm.
[0018] In a possible implementation, the model management functional entity may further optimize a third algorithm based on a third data subset and / or the value of a third parameter of the data in the third data subset, where the value of the third parameter is determined based on the third algorithm, the third algorithm is used to delete the third data subset, the third parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model, the statistical distribution characteristics of the third data subset are different from those of the training data set, and the third data subset is not used to update the first model. The third algorithm optimized based on the third data subset and / or the value of the third parameter of the data in the third data subset can improve the accuracy of data cleaning, thereby improving the robustness of the updated model.
[0019] In a possible implementation, the third parameter includes one or more of the following: the change amount of the loss value, which is used to indicate the degree of influence of the data on the loss function of the first model; the second probability, which is used to indicate the probability that the data is abnormal data; or, the degree of difference in statistical distribution characteristics, which is used to indicate the degree of difference in the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set. The third parameter can be any parameter that can measure the degree of influence of the data on the model prediction accuracy. The parameters shown in the embodiments of this application are only examples. In other embodiments, the third parameter may further include more or fewer parameters. The embodiments of this application do not limit the types of parameters included in the third parameter. It should be understood that the third parameter may be the same as the foregoing first parameter or different from the foregoing first parameter. The embodiments of this application do not limit the relationship between the first parameter and the third parameter.
[0020] For another example, if the above method is implemented by the model training functional entity, the specific implementation is as follows:
[0021] In a possible implementation, the model training functional entity may further send a first indication message to the model management functional entity or receive a first indication message from the model management functional entity, where the first indication message is used to indicate the first algorithm.
[0022] In a possible implementation, the model training functional entity may further determine the first probability based on a second algorithm.
[0023] In a possible implementation, the model training functional entity may further send a second indication message to the model management functional entity, where the second indication message is used to indicate the first probability.
[0024] In a possible implementation, the model training functional entity may also send third indication information to the model management functional entity, or receive third indication information from the model management functional entity, where the third indication information is used to indicate the second algorithm.
[0025] In a possible implementation, the model training functional entity may also receive the first data set from the model inference functional entity.
[0026] In a possible implementation, the model training functional entity may also optimize the first algorithm based on the second data subset and / or the value of the first parameter of the data in the second data subset, where the first parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model.
[0027] In a possible implementation, the model training functional entity may also send first information to the model management functional entity, where the first information includes the second data subset and / or the value of the first parameter of the data in the second data subset.
[0028] In a possible implementation, the first parameter includes one or more of the following: the change amount of the loss value, where the change amount of the loss value is used to indicate the degree of influence of the data on the loss function of the first model; the second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, the degree of difference in statistical distribution characteristics, where the degree of difference in statistical distribution characteristics is used to indicate the degree of difference in the statistical distribution characteristics of the data in the second data subset and the statistical distribution characteristics of the data in the training data set.
[0029] In a possible implementation, the first information further includes one or more of the following: the value of the second parameter, where the second parameter is used to indicate the inference accuracy of the updated first model; the identifier of the first model; the version number of the updated first model; the update time of the first model; the data volume included in the third data set used to update the first model; or, the duration required for the inference of the updated first model.
[0030] In a possible implementation, the model training functional entity may also send third information to the model inference functional entity, where the third information includes one or more of the following: the identifier of the first model; the version number of the updated first model; the update time of the first model; or, the storage address of the updated first model. The model training functional entity may send the model parameter information of the model to the model inference functional entity, so that the model inference functional entity can update the model parameters of the first model based on this third information.
[0031] Second aspect, a model management method is provided. This method can be implemented by a model inference functional entity. The method includes: deleting a third data subset in a third data set based on a third algorithm to obtain a first data set. The third data set is collected by the model inference functional entity. The statistical distribution characteristics of the third data subset are different from those of the training data set used to train the first model, and the third data subset is not used to update the first model; sending the first data set to a model training functional entity.
[0032] In a possible implementation manner, the model inference functional entity may also send fourth indication information to a model management functional entity, or receive fourth indication information from the model management functional entity. The fourth indication information is used to indicate the third algorithm.
[0033] In a possible implementation manner, deleting the third data subset in the third data set based on the third algorithm includes: determining the value of a third parameter of the data in the third data set based on the third algorithm. The third parameter is used to indicate the influence degree of the data on the accuracy of the first model inference; deleting the third data subset, where the value of the third parameter of the data in the third data subset is greater than or equal to a first threshold.
[0034] In a possible implementation manner, the model inference functional entity may also optimize the third algorithm based on the third data subset and / or the value of the third parameter of the third data subset.
[0035] In a possible implementation manner, the model inference functional entity may also send second information to the model management functional entity. The second information includes the third data subset and / or the value of the third parameter of the data in the third data subset.
[0036] In a possible implementation manner, the third parameter includes one or more of the following: the change amount of the loss value, which is used to indicate the influence degree of the data on the loss function of the first model; the second probability, which is used to indicate the probability that the data is abnormal data; or, the difference degree of the statistical distribution characteristics, which is used to indicate the difference degree between the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set.
[0037] In a possible implementation manner, the model inference functional entity may also receive third information from the model training functional entity. The third information includes one or more of the following: the identifier of the first model; the version number of the updated first model; the update time of the first model; or, the storage address of the updated first model.
[0038] In a possible implementation manner, the model inference functional entity may further update the model parameters of the first model based on the third information.
[0039] In a third aspect, a model management system is provided, including a model management functional entity and a model training functional entity for executing any implementation method in the first aspect, and a model inference functional entity for executing any implementation method in the second aspect.
[0040] In a fourth aspect, an embodiment of the present application provides a communication device, and the device has a function of implementing any method implemented by the model management functional entity in the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0041] In a fifth aspect, an embodiment of the present application provides a communication device, and the device has a function of implementing any method implemented by the model training functional entity in the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0042] In a sixth aspect, an embodiment of the present application provides a communication device, and the device has a function of implementing any method implemented by the model training functional entity in the second aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0043] In a seventh aspect, an embodiment of the present application provides a communication device. The communication device includes a communication interface and a processor. Optionally, it further includes a memory. Wherein, the memory is used to store a computer program, and the processor is coupled to the memory and the communication interface. When the processor reads the computer program or instruction, the communication device executes the method executed by the model management functional entity in the first aspect, or the communication device executes the method executed by the model training functional entity in the first aspect, or the communication device executes the method executed by the model inference functional entity in the second aspect.
[0044] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium is used to store a computer program. When the computer program runs on a computer, the computer is caused to execute the method provided in the first aspect or the second aspect.
[0045] In a ninth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program runs on a computer, the computer is caused to execute the method described in the first aspect or the second aspect.
[0046] In a tenth aspect, a chip system is provided, including a processor and an interface. The processor is configured to call and run instructions from the interface, so that the chip system implements the method described in the first aspect or the second aspect above.
[0047] For the beneficial effects of the second aspect to the tenth aspect, refer to the beneficial effects of the first aspect, and will not be repeated here. Description of the Drawings
[0048] Figure 1 Process of beam management for a wireless AI model;
[0049] Figure 2A and Figure 2B Schematic diagrams of two network architectures provided by embodiments of the present application;
[0050] Figure 2C Structural diagram of a model management system provided by embodiments of the present application;
[0051] Figure 3 Schematic diagram of a scenario for dataset collection;
[0052] Figure 4 and Figure 5 Flowcharts of two model management methods provided by embodiments of the present application;
[0053] Figure 6 Structural diagram of a communication device provided by embodiments of the present application;
[0054] Figure 7 Structural diagram of another communication device provided by embodiments of the present application. Detailed Embodiments
[0055] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. The embodiments of the present application can be applied to various mobile communication systems, such as: new radio (NR) systems, long term evolution (LTE) systems, advanced long term evolution (LTE-A) systems, future communication systems, and other communication systems, which are not limited here. For ease of understanding, some terms in the embodiments of the present application are explained below. For the sake of clarity of the embodiments of the present invention, the following is a unified introduction to some content and concepts related to the embodiments of the present invention.
[0056] (1) The terminal device in the embodiments of this application is a device with wireless transceiver functions. It can be a fixed device, a mobile device, a handheld device (such as a mobile phone), a wearable device, a vehicle-mounted device, or a wireless device (such as a communication module, a modem, or a chip system, etc.) built into the above devices. This terminal device is used to connect people, things, machines, etc., and can be widely used in various scenarios, such as including but not limited to the following scenarios: cellular communication, device-to-device (D2D) communication, vehicle to everything (V2X) communication, machine-to-machine / machine-type communications (M2M / MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearables, smart transportation, smart city, drones, robots, etc. This terminal device can sometimes be referred to as a user equipment (UE), a terminal, an access station, a UE station, a remote station, a wireless communication device, or a user device, etc.
[0057] (2) A network management system (NMS), as a management service consumer, is responsible for the operation, management, and maintenance of the network. The network management system is a traditional cross-domain management system or can also be a collection of partial management functions for cross-domain management.
[0058] (3) An element management system (EMS), as a management service provider, is used to manage network elements. The element management system can be a traditional domain management system or a single-domain management system, or can also be a collection of partial management functions for single-domain management. The network management system and the element management system can be collectively referred to as the 3rd generation partnership project (3GPP) management system, or can also be collectively referred to as the operations administration and maintenance (OAM) module.
[0059] (4) Network element (NE), which is an entity that provides network services. A network element is a management service provider, including core network elements, radio access network elements, transmission network elements, etc. For example, core network elements may include, but are not limited to, access and mobility management function (AMF) entities, session management function (SMF) entities, policy control function (PCF) entities, network data analysis function (NWDAF) entities, network repository function (NRF), gateways, etc.
[0060] Radio access network elements may include, but are not limited to: base stations, evolved Node Bs in Long Term Evolution (LTE) systems or Long Term Evolution-Advanced (LTE-A) systems, which may be abbreviated as eNBs or e-NodeBs, transmission reception points (TRPs), next-generation Node Bs (gNBs) in 5th generation (5G) mobile communication systems, next-generation Node Bs in 6th generation (6G) mobile communication systems, base stations in future mobile communication systems, access nodes in Wi-Fi systems, etc. Additionally, it may also be access network devices in an Open Radio Access Network (ORAN) system. Optionally, the network device may also be a module or unit that performs some functions of a base station. For example, the network device may be a Central Unit (CU), a Distributed Unit (DU), a CU-Control Plane (CP), a CU-User Plane (UP), or a Radio Unit (RU), etc. Among them, the CU can perform the functions of the radio resource control protocol and the Packet Data Convergence Protocol (PDCP) of the base station, and can also perform the function of the Service Data Adaptation Protocol (SDAP); the DU can perform the functions of the radio link control layer and the Medium Access Control (MAC) layer of the base station, and can also perform some or all of the functions of the physical layer. In different systems, the CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art can understand their meanings. For example, in the ORAN system, the CU may also be called O-CU, the DU may also be called Open (O)-DU, the CU-CP may also be called O-CU-CP, the CU-UP may also be called O-CUP-UP, and the RU may also be called O-RU.
[0061] (5) The terms "system" and "network" in the embodiments of the present application may be used interchangeably. "At least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. The following "at least one" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0062] Unless otherwise stated, the ordinal numbers such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority, or importance of multiple objects.
[0063] In addition, the terms "including" and "having" in the embodiments of the present application, the claims, and the drawings are not exclusive. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules and may further include unlisted steps or modules.
[0064] Currently, AI / ML technologies can be applied in multiple fields to solve problems that cannot be solved by traditional algorithms. For example, a wireless AI model (which can also be simply referred to as a model in the following embodiments) can be used for feedback of channel state information (CSI), positioning, beam management, energy-saving management, load balancing, and mobility optimization, etc. For the sake of easy understanding, first, in combination with Figure 1 An exemplary description of the implementation process of the wireless AI model will be given.
[0065] Please refer to Figure 1 , for the process of beam management of the wireless AI model.
[0066] S101: The terminal device scans all beams to obtain data for model training or updating.
[0067] Among them, the data for training or updating may include, for example, the reference signal receiving power (RSRP) corresponding to all beams and the best beam information (such as the best beam identification ID).
[0068] S102: The terminal device sends the obtained data to the network device. Correspondingly, the network device receives the data from the terminal device.
[0069] S103: The network device trains or updates a model based on the data.
[0070] S104: The terminal device sparsely samples all beams to obtain synchronization signal block (SSB) beams.
[0071] S105: The terminal device scans the SSB beams to obtain measurement results of the SSB beams.
[0072] S106: The terminal device sends the measurement results of the SSB beams to the network device. Correspondingly, the network device receives the measurement results from the terminal device.
[0073] Wherein, the measurement results are, for example, the reference signal receiving power (RSRP) corresponding to the SSB beams and the best beam information (such as the best beam identification ID).
[0074] S107: The model in the network device predicts its corresponding K best beams based on the measurement results (such as RSRP) of the SSB beams.
[0075] Wherein, the K best beams are the K beams with the greatest possibility of being the best beams among the SSB beams.
[0076] S108: The network device sends the beam IDs of the K best beams to the terminal device. Correspondingly, the terminal device receives the beam IDs of the K best beams from the network device.
[0077] S109: The terminal device determines the channel state information reference signal (CSI-RS) corresponding to the K best beams.
[0078] S110: The terminal device sends the beam ID of the best CSI-RS beam to the network device.
[0079] S111: The network device sends data to the terminal device through the best CSI-RS beam.
[0080] Please refer to Figure 2A , which is a schematic diagram of a network architecture provided by an embodiment of this application. Figure 2AIn this case, the network architecture includes a management service (MnS) producer and an MnS consumer. Among them, the management service producer is used to provide network management services externally, and the MnS consumer can call the management service to provide network management services externally.
[0081] Please refer to Figure 2B , which is a schematic diagram of another network architecture provided by the embodiments of this application. Figure 2B In this case, the network architecture includes network elements, a network element management system, and a network management system.
[0082] Please refer to Figure 2C , which is the model management system architecture provided by the embodiments of this application. Figure 2C In this case, the model management system architecture includes: a model management function (MMF) entity, a model training function (MTF) entity, and a model inference function (MIF) entity. The MMF entity, the MTF entity, and the MIF entity can communicate with each other pairwise. Among them, the MMF entity can be used to implement the management of the model life cycle, can trigger the MTF entity to perform model training or update, and can trigger the MIF entity to perform inference. The MTF entity can be used to train and update the model, and output the model file after the model training or update is completed. The MIF entity can be used for model inference, that is, inputting data into the model to obtain the corresponding predicted output result. Optionally, the MIF entity can also collect the output results of the model to obtain a data set for updating the model. Optionally, other entities can also be included in the model management system architecture. For example, a model evaluation function (MEF) entity can also be included to evaluate the performance of the model. The embodiments of this application do not limit the types of entities included in the model management system architecture.
[0083] Optionally, taking model update as an example, the model management system can be connected to a model update trigger module (not shown in the figure), and the model update trigger module can send model management information to the model management system.
[0084] For example, the model update trigger module may send information such as the identifier of the model to be updated (e.g., Model 1), the type and requirements of model management, etc. to the model management functional entity. Among them, the identifier of the model is the identification code of the model, which is used to distinguish different models. The type and requirements of model management refer to the management measures and execution standards taken to ensure the credibility and robustness of the model during the model management process. Such management measures may include, for example, the verification of data integrity, the monitoring method of model performance (e.g., continuous monitoring), and ensuring the robustness of the model update process. For example, the data used for model training is an 8*4 RSRP matrix. To verify the integrity of the data, for example, it is to verify whether there is missing data in the RSRP matrix. If there is missing data, it is determined that the data is incomplete; or, it can also be verified whether the data has corresponding labels. Taking Model 1 used for beam management as an example, the label may be, for example, the best beam ID. If there is no label, it is determined that the data is incomplete. The model management functional entity can determine the model management strategy based on this model management information and model management constraint conditions. Among them, the model management constraint conditions are, for example, the built-in constraint conditions of the model management functional entity, which are used to constrain the accuracy of model prediction and data overhead. The data overhead refers to, for example, the proportion of the data that is excluded from the collected data. The model management strategy may include, for example, the algorithm used in the model update process.
[0085] The model update trigger module may also send information such as the model identifier of Model 1 and the model update strategy to the model training functional entity. Among them, the model update strategy can be used to indicate the way of model update and the conditions for ending the model update. The way of model update may be, for example, parameter fine-tuning update, or it can also be incremental update. The conditions for ending the model update may include, for example, that the accuracy of the updated model prediction reaches a threshold (e.g., Threshold 1), and / or the model update duration exceeds the preset update duration.
[0086] The model update module may also send information such as the identifier of Model 1 and the dataset collection strategy to the model inference functional entity. Among them, the dataset collection strategy can be used to indicate the type of data to be collected. The data type includes, for example, data such as RSRP and channel impulse response (CIR). Optionally, the dataset collection strategy can also indicate how to select the collected data to ensure that the collected data is the latest, most relevant, and most diverse. The model inference functional entity can collect data based on this dataset collection strategy to obtain a dataset for model update (e.g., the second dataset described below).
[0087] Among them, the model management functional entity can be Figure 2A the management service provider shown, or it can also be Figure 2AThe management service consumer, model training functional entity, and model inference functional entity shown are Figure 2A the management service provider shown. Therefore, the following deployment scenarios may exist for this model management system architecture:
[0088] Scenario 1: The model management functional entity is Figure 2A the management service consumer shown, and the model training functional entity and model inference functional entity are Figure 2A the management service providers shown. The model management functional entity is deployed in Figure 2B the network management system shown, and the model training functional entity and model inference functional entity are deployed in Figure 2B the network element management system shown.
[0089] Scenario 2: The model management functional entity is Figure 2A the management service consumer shown, and the model training functional entity and model inference functional entity are Figure 2A the management service providers shown. The model management functional entity is deployed in Figure 2B the network management system shown, the model training functional entity is deployed in Figure 2B the network element management system shown, and the model inference functional entity is deployed in Figure 2B the network element shown.
[0090] Scenario 3: The model management functional entity is Figure 2A the management service consumer shown, and the model training functional entity and model inference functional entity are Figure 2A the management service providers shown. The model management functional entity is deployed in Figure 2B the network management system shown, and the model training functional entity and model inference functional entity are deployed in Figure 2B the network element shown.
[0091] Scenario 4: The model management functional entity, model training functional entity, and model inference functional entity are Figure 2A the management service providers shown. The model management functional entity, model training functional entity, and model inference functional entity are all deployed in Figure 2B the network element management system shown.
[0092] Scenario 5: The model management functional entity, model training functional entity, and model inference functional entity are Figure 2A the management service providers shown. The model management functional entity and model training functional entity are deployed in Figure 2B the network element management system shown, and the model inference functional entity is deployed in Figure 2B the network element shown.
[0093] Scenario 6: The model management functional entity, model training functional entity, and model inference functional entity are Figure 2AThe management service provider shown. The model management functional entity is deployed in Figure 2B The network management system shown, and the model training functional entity and the model inference functional entity are deployed in Figure 2B The network element shown.
[0094] The above model management system can be used for model training and model updating. When performing model training or model updating, the model management functional entity can instruct the model training functional entity to perform model training or updating based on the data collected by the model inference functional entity, so as to generate a model or update the model (such as updating model parameters). Among them, the data collected by the model inference functional entity is generally the data actually generated by the terminal device. Therefore, a terminal device used for network attack (which can also be called a network attack device) can perform a network attack by reporting false data, resulting in a relatively low performance of the trained or updated model. For example, please refer to Figure 3 , the network attack device can report forged data (such as false RSRP) to the network device (such as gNB), so that the data set used for updating the model collected by the model inference functional entity includes the forged data and the normal data (such as normal RSRP) reported by the terminal device that does not perform network attack (which can also be called a non-network attack device, that is, a normal device). Therefore, after the model training functional entity updates the model based on the data collected by the model inference functional entity, a wrong prediction model may be obtained, such as a wrong classification boundary, ultimately resulting in a decline in the prediction performance of the updated model.
[0095] In view of this, the embodiments of the present application provide a model management method, which is used to clean the data set for updating the model before model updating, reduce the probability of false data existing in the data for model updating, and improve the robustness of the model.
[0096] The following will introduce the method provided by the embodiments of the present application in detail in combination with the above Figure 1 、 Figures 2A to 2C introduction.
[0097] Please refer to Figure 4 and Figure 5 , which is a model management method provided by the embodiments of the present application. This method can be implemented by the model management system shown in Figure 2C . The embodiments of the present application take the model management system for model updating as an example. Figure 4 and Figure 5 , all optional steps in the embodiments of the present application are represented by dotted lines.
[0098] S1: Determine a first algorithm based on a first probability.
[0099] Over time, the data distribution of the input model may change, resulting in an increase in the model prediction error, that is, the model drift phenomenon occurs. Therefore, it is necessary to use the collected dataset to update the model to maintain the prediction performance of the model. As mentioned above, there may be pseudo data generated by network attack devices in the dataset used to update the model (such as the second data subset and the third data subset described below), resulting in low robustness of the model. Currently, in order to maintain the prediction performance of the model, before updating the model, some data in the collected dataset can be deleted. For example, the pseudo data (or also called poisoned data) in the collected dataset can be deleted based on the sample detection method of the loss threshold. Among them, the sample detection method based on the loss threshold can be, for example, out-of-distribution detector (OOD Detection), or it can also be methods such as mahalanobis distance-based score, etc.
[0100] However, the core idea of deleting the pseudo data in the collected dataset (such as dataset 1) based on the sample detection method of the loss threshold is: it is considered that dataset 1 only contains normal data and pseudo data. Among them, the statistical distribution characteristics of the normal data are the same as those of the training dataset used during model training, and the statistical distribution characteristics of the pseudo data are quite different from those of the training dataset. Therefore, the pseudo data in dataset 1 can be deleted based on the principle that the loss performance differences of normal data and pseudo data on the same model are relatively large. That is, when deleting the pseudo data in dataset 1 based on the sample detection method of the loss threshold, all data with statistical distribution characteristics different from those of training data 1 in this dataset 1 may be deleted. However, there may still be data in this dataset 1 with statistical distribution characteristics different from those of the training dataset, but generated due to the reason of model prediction accuracy during model application, that is, there is data in dataset 1 that is not generated by network attack devices and has statistical distribution characteristics different from those of the training dataset (such as also called drift data). If this part of the data is deleted, the prediction performance of the model updated based on this dataset 1 is still low.
[0101] Therefore, when it is determined that model update is required, for example, when the first model needs to be updated, an algorithm (e.g., the first algorithm) for deleting spurious data (e.g., the second data subset) from the first data set (e.g., the first data set) can be determined based on the first probability that there is drifting data (e.g., the first data subset) in the data set. Optionally, the determination of the first algorithm based on the first probability can be implemented by a model management functional entity in the model management system, or can also be implemented by a model training functional entity. If the determination of the first algorithm based on the first probability is implemented by the model management functional entity, then before executing S1, the following S401 - S408 can also be executed. Please refer to Figure 4 。
[0102] S401: The model inference functional entity sends the first data set to the model training functional entity. Correspondingly, the model training functional entity receives the first data set from the model inference functional entity. Delete the third data subset from the third data set based on the third algorithm.
[0103] Among them, the first data set can be the original data set (e.g., the third data set) collected by the model inference functional entity, or can also be the data set obtained after the model inference functional entity deletes some data from the third data set. Therefore, optionally, S402 can also be executed before executing S401.
[0104] S402: The model inference functional entity deletes the third data subset from the third data set based on the third algorithm.
[0105] The data in the third data subset is, for example, the data generated by the aforementioned network attack device. Before the model inference functional entity sends the first data set to the model training functional entity, it may delete the spurious data (i.e., the third data subset) in the third data set based on a third algorithm, and obtain the first data set. The third algorithm is, for example, the aforementioned sample detection method based on a loss threshold. For example, the model inference functional entity may determine the value of a third parameter corresponding to each data in the third data set based on the third algorithm, and delete the data in the third data set whose value of the third parameter is greater than or equal to a first threshold, so as to obtain the first data set. The set formed by the data whose value of the third parameter is greater than or equal to the first threshold is the aforementioned third data subset. Among them, the third parameter is used to indicate the degree of influence of the data on the accuracy of the first model inference. The third parameter may include, for example, one or more of the following: the change amount of the loss value, which is used to indicate the degree of influence of the data on the loss function of the first model; the second probability, which is used to indicate the probability that the data is abnormal data; or, the degree of difference in statistical distribution characteristics, which is used to indicate the degree of difference in the statistical distribution characteristics between the third data subset and the statistical distribution characteristics of the training data set.
[0106] Among them, if the change amount of the loss value corresponding to the data is greater than or equal to the first threshold, or the probability that the data is abnormal data is greater than or equal to the first threshold, or the degree of difference in the statistical distribution characteristics of the data relative to the statistical distribution characteristics of the training data set is greater than or equal to the first threshold, it indicates that the statistical distribution characteristics of the data are quite different from the statistical distribution characteristics of the training data set. Therefore, it can be considered that the data is spurious data. Therefore, the data can be deleted from the third data set. The set formed by the data deleted from the third data set is the third data subset.
[0107] Optionally, the model inference functional entity may also determine that the data in the third data set whose value of the third parameter is less than or equal to a second threshold is the aforementioned normal data, that is, the data with the same statistical distribution characteristics as the statistical distribution characteristics of the training data set, and determine that the data in the third data set whose value of the third parameter is greater than the second threshold and less than the first threshold is suspicious data. The suspicious data is, for example, data for which it is impossible to confirm whether the data is spurious data, drift data, or normal data based on the third algorithm. Therefore, the model inference functional entity may also add a suspicious identifier to the suspicious data in the third data set, so that some data in the first data set obtained after deleting the third data subset in the third data set carry the suspicious identifier. Among them, the suspicious data may include one or more of the following types of data: the aforementioned drift data, spurious data, or normal data.
[0108] The above-mentioned third algorithm can be determined by, for example, the model inference functional entity, or can also be indicated by the model management functional entity. If the third algorithm is determined by the model management functional entity, before executing S402, S403 can also be executed, and the model management functional entity sends fourth indication information to the model inference functional entity. Correspondingly, the model inference functional entity receives the fourth indication information from the model management functional entity. The fourth indication information is used to instruct the model inference functional entity to delete the third data subset in the third dataset based on the third algorithm. If the third algorithm is determined by the model inference functional entity, the model inference functional entity can also execute S404 after determining the third algorithm, and the model inference functional entity sends fourth indication information to the model management functional entity. Correspondingly, the model management functional entity receives the fourth indication information from the model inference functional entity. The fourth indication information is used to inform the model management functional entity that the first dataset it sends to the model training functional entity is the dataset obtained after deleting some data in the to-be-collected third dataset, and the algorithm used to delete this part of the data is the third algorithm. It should be understood that
[0109] Optionally, the fourth indication information can directly indicate the third algorithm or can also indicate the identifier of the third algorithm. For example, the protocol predefines multiple algorithms, and each of the multiple algorithms has a unique identifier. The identifier can be, for example, the number of the algorithm. Therefore, the fourth indication information can directly indicate the number. When the model management functional entity or the model inference functional entity receives the fourth indication information, it can determine the algorithm to be used based on the identifier indicated by the fourth indication information. Optionally, the identifier of the first model can also be included in the fourth indication information.
[0110] Optionally, the model inference functional entity can also learn the statistical distribution characteristics of the third data subset based on the third data subset and / or the value of the third parameter of the data in the third data subset to optimize the third algorithm. Or, the model inference functional entity can also send the third data subset and / or the value of the third parameter of the data in the third data subset to the model management functional entity, so that the model management functional entity learns the statistical distribution characteristics of the third data subset based on the third data subset and / or the value of the third parameter of the data in the third data subset to optimize the third algorithm. For example, the model inference functional entity can send second information to the model management functional entity, and the second information includes the third data subset and / or the value of the third parameter of the data in the third data subset. Optionally, if the third algorithm is determined by the model inference functional entity, the second information and the fourth indication information can be the same information or can also be different information, which is not limited in the embodiments of the present application.
[0111] Optionally, the above first threshold and second threshold may be pre-configured. For example, corresponding first threshold and second threshold are pre-configured for different algorithms; alternatively, the first threshold and the second threshold may also be indicated by the fourth indication information. Therefore, optionally, the fourth indication information may further include the first threshold and the second threshold.
[0112] S405: The model training functional entity determines the first probability based on the second algorithm.
[0113] The second algorithm may be, for example, a maximum mean discrepancy (MMD) algorithm and an algorithm combining Kolmogorov-Smirnov (KS) test and Bonferroni correction. The model training functional entity may determine the probability that a first data subset exists in the first data set based on the second algorithm.
[0114] For example, if the second algorithm is the MMD algorithm, the model training functional entity may obtain the aforementioned training data set, and calculate the statistical parameters of the first data set and the training data set respectively, such as mean embeddings, and calculate the unbiased estimates of MMD and MMD squared based on the statistical parameters of the first data set and the mean embeddings of the training data set using the Hilbert space, so as to obtain the probability that a first data subset exists in the first data set.
[0115] If the second algorithm is an algorithm combining KS test and Bonferroni correction, the model training functional entity may obtain the aforementioned training data set, and calculate the statistical parameters of the first data set and the training data set respectively, such as empirical cumulative density functions, and determine the KS test parameter based on the empirical cumulative density functions of the first data set and the training data set, such as the largest difference, so as to obtain the probability that a first data subset exists in the first data set.
[0116] Optionally, if the first dataset includes data carrying the aforementioned suspicious identifier, the model training functional entity may also transform the statistical distribution characteristics of the data carrying the suspicious identifier based on the statistical distribution characteristics of the training dataset. For example, the model training functional entity may transform the data carrying the suspicious identifier in the first dataset into a first distribution on the latent space through a selective batch normalization (SBN) layer, and the statistical distribution characteristics of the first distribution are the same as those of the aforementioned training dataset. The model training functional entity may determine the probability of the existence of the first data subset in the first dataset after the distribution transformation based on the second algorithm.
[0117] The second algorithm for the model training functional entity to determine the first probability may be determined by the model training functional entity, or may also be indicated by the model management functional entity. If the second algorithm is indicated by the model management functional entity, then before executing S405, S406 may also be executed, and the model management functional entity sends the third indication information to the model training functional entity. Correspondingly, the model training functional entity receives the third indication information from the model management functional entity, and the third indication information is used to instruct the model training functional entity to determine the first probability of the existence of the first data subset in the first dataset based on the second algorithm. If the second algorithm is determined by the model training functional entity, then after executing S405, S407 may also be executed, and the model training functional entity sends the third indication information to the model management functional entity. Correspondingly, the model management functional entity receives the third indication information from the model training functional entity, and the third indication information is used to inform the model management functional entity that the algorithm used for determining the first probability of the existence of the first data subset in the first dataset is the second algorithm.
[0118] Optionally, the third indication information may directly indicate the second algorithm or may indicate the identifier of the second algorithm. Optionally, the identifier of the first model may also be included in the third indication information.
[0119] S408: The model training functional entity sends the second indication information to the model management functional entity, and the second indication information is used to indicate the first probability. Correspondingly, the model management functional entity receives the second indication information from the model training functional entity.
[0120] When the model management functional entity receives the second indication information from the model training functional entity, it can obtain the first probability that the first data subset exists in the first dataset based on the second indication information, and execute the foregoing S1, that is, determine the first algorithm based on the first probability. For example, if the first probability is greater than a threshold (for example, threshold 2), it indicates that the probability that the first data subset exists in the first dataset is relatively high. Since the statistical distribution characteristics of the first data subset and the second data subset are different from the statistical distribution characteristics of the foregoing training dataset, when deleting the second data subset, it is also necessary to detect the first data subset in the first dataset to avoid deleting the first data subset. Therefore, the model management functional entity can determine that the algorithm for detecting drift data is the first algorithm. If the first probability is less than or equal to threshold 2, it indicates that the probability that the first data subset exists in the first dataset is relatively low, that is, the first dataset may only include the foregoing pseudo data and normal data. Therefore, when deleting the second data subset, it is not necessary to detect the first data subset in the first dataset. Therefore, the model management functional entity can determine that the algorithm not used for detecting drift data is the first algorithm, and the first algorithm can be, for example, the foregoing OOD Detection. Optionally, the first algorithm and the foregoing third algorithm can be the same algorithm or different algorithms. If the first algorithm and the foregoing third algorithm are the same algorithm, the relevant threshold corresponding to the first algorithm is different from the relevant threshold corresponding to the third algorithm. For example, the third algorithm corresponds to a first threshold and a second threshold, while the first algorithm corresponds to a threshold (for example, the third threshold described below). Optionally, the second indication information and the third indication information in S407 can be the same indication information or different indication information. If the second indication information and the third indication information in S407 are not the same indication information, then S407 and S408 can be executed simultaneously, or S407 can be executed before S408, or S407 can be executed after S408. The embodiments of the present application do not limit the execution order of S407 and S408.
[0121] Optionally, the model training functional entity can also determine the uncertainty quantification information of the first probability. For example, the model training functional entity can determine the uncertainty quantification information of the first probability through parameters such as Bayesian posterior probability, and send the uncertainty quantification information of the first probability to the model management functional entity. The uncertainty quantification information is used to quantify the accuracy of the first probability. Among them, the uncertainty quantification information of the first probability and the first probability can be sent through the same information or through different information. The embodiments of the present application do not limit the sending method of the first probability and the uncertainty quantification information of the first probability.
[0122] The model management functional entity can determine the first algorithm based on the first probability and the uncertainty quantification information of the first probability. For example, the first algorithms determined by the model management functional entity include the following situations.
[0123] Situation 1: The first probability is 80%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 20%, indicating that the accuracy of the first probability is low, that is, the possibility of the existence of the first data subset in the first dataset is low, and there is no need to detect the first data subset in the first dataset. The model management functional entity can determine the algorithm for not detecting drift data as the first algorithm.
[0124] Situation 2: The first probability is 20%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 20%, indicating that the accuracy of the first probability is low, that is, the possibility of the existence of the first data subset in the first dataset is high, and it is necessary to detect the first data subset in the first dataset. The model management functional entity can determine the algorithm for detecting drift data as the first algorithm.
[0125] Situation 3: The first probability is 80%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 80%, indicating that the accuracy of the first probability is high, that is, the possibility of the existence of the first data subset in the first dataset is high, and it is necessary to detect the first data subset in the first dataset. The model management functional entity can determine the algorithm for detecting drift data as the first algorithm.
[0126] Optionally, the algorithm for detecting drift data determined based on Situation 3 and the algorithm for detecting drift data determined based on Situation 2 can be the same algorithm or different algorithms. For example, the first algorithm determined in Situation 2 is an algorithm with a more complex detection process, while the first algorithm determined in this Situation 3 is an algorithm with a simpler detection process.
[0127] Situation 4: The first probability is 20%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 80%, indicating that the accuracy of the first probability is high, that is, the possibility of the existence of the first data subset in the first dataset is low, and there is no need to detect the first data subset in the first dataset. The model management functional entity can determine the algorithm for not detecting drift data as the first algorithm.
[0128] Optionally, the algorithm for not detecting drift data determined based on Situation 4 and the algorithm for not detecting drift data determined in Situation 1 can be the same algorithm or different algorithms. For example, the first algorithm determined in Situation 1 is an algorithm with a more complex detection process, while the first algorithm determined in this Situation 4 is an algorithm with a simpler detection process.
[0129] After the model management functional entity executes S1, it can send the first algorithm to the model training functional entity. Therefore, optionally, S409 to S412 can also be executed.
[0130] S409: The model management functional entity sends first indication information to the model training functional entity. Correspondingly, the model training functional entity receives the first indication information from the model management functional entity.
[0131] The first indication information is used to indicate the first algorithm. For example, the first indication information can directly indicate the first algorithm, or can indicate the identifier of the first algorithm. Optionally, the identifier of the first model can also be included in the first indication information.
[0132] S410: The model training functional entity deletes the second data subset in the first data set based on the first algorithm to obtain a second data set.
[0133] When receiving the first indication information, the model training functional entity can clean the second data subset in the first data set based on the first algorithm. Optionally, the model training functional entity can determine the value of the first parameter corresponding to each data in the first data set based on the first algorithm, and delete the data in the first data set whose value of the first parameter is greater than or equal to the third threshold to obtain a second data set. The set of data whose value of the first parameter is greater than or equal to the third threshold forms the second data subset. Among them, the first parameter is used to indicate the influence degree of the data on the inference accuracy of the first model. The first parameter can include, for example, one or more of the following: the change amount of the loss value, and the change amount of the loss value is used to indicate the influence degree of the data on the loss function of the first model. The change amount of the loss value is, for example, the change amount of the loss value of the loss function of the first model caused by the data (for example, the data in the second data subset) relative to the loss value of the loss function when the first model is trained; the second probability, and the second probability is used to indicate the probability that the data is abnormal data; or, the difference degree of the statistical distribution characteristics, and the difference degree of the statistical distribution characteristics is used to indicate the difference degree between the statistical distribution characteristics of the second data subset and the statistical distribution characteristics of the training data set.
[0134] Among them, if the change amount of the loss value corresponding to the data is greater than or equal to the third threshold, or the probability that the data is abnormal data is greater than or equal to the third threshold, or the difference degree of the statistical distribution characteristics of the data relative to the statistical distribution characteristics of the training data set is greater than or equal to the third threshold, it indicates that the statistical distribution characteristics of the data are quite different from the statistical distribution characteristics of the training data set. Therefore, it can be considered that the data is pseudo data. Therefore, the data can be deleted from the first data set, and the set of data deleted from the first data set forms the second data subset.
[0135] As described above, there may be data carrying suspicious identifiers in the first dataset. Therefore, when the model training functional entity deletes the second data subset from the first dataset based on the first algorithm, it can also determine whether there is data carrying the suspicious identifier in the first dataset. If there is data carrying the suspicious identifier in the first dataset, the model training functional entity can determine the value of the first parameter corresponding to the data carrying the suspicious identifier in the first dataset based on the first algorithm, and delete the data whose value of the first parameter is greater than or equal to the third threshold from the first dataset to obtain the second dataset. In this way, for the data in the first dataset that does not carry the suspicious identifier, the value of the corresponding first parameter does not need to be calculated, which can reduce the computational burden and improve the model update efficiency.
[0136] Optionally, the model training functional entity can also learn the statistical distribution characteristics of the second data subset based on the second data subset and / or the value of the first parameter of the data in the second data subset to optimize the first algorithm. Alternatively, the model training functional entity can also send the second data subset and / or the value of the first parameter of the data in the second data subset to the model management functional entity, so that the model management functional entity can learn the statistical distribution characteristics of the second data subset based on the second data subset and / or the value of the first parameter of the data in the second data subset to optimize the first algorithm. For example, the model training functional entity can send the first information to the model management functional entity, and the first information includes the second data subset and / or the value of the first parameter of the data in the second data subset.
[0137] Optionally, the above-mentioned third threshold can be pre-configured; or, the third threshold can also be indicated by the first indication information. Therefore, optionally, the first indication information can also include the third threshold.
[0138] S411: The model training functional entity updates the first model based on the second dataset.
[0139] After obtaining the second dataset, the model training functional entity can update the first model based on the second dataset, and end the update of the first model when the accuracy of the updated first model prediction reaches the aforementioned threshold 1, or the update duration of the first model reaches the preset update duration, to obtain the updated first model.
[0140] Optionally, the aforementioned model training functional entity may send the first information to the model management functional entity before or after executing S411. If the model training functional entity sends the first information to the model management functional entity after executing S411, the first information may further include one or more of the following: the value of a second parameter for indicating the accuracy of the updated first model inference; the identifier of the first model; the version number of the updated first model; the update time of the first model; the data volume included in the third data set used to update the first model; or, the duration required for the updated first model inference.
[0141] If the model training functional entity sends the first information to the model management functional entity before executing S411, then after executing S411, the model training functional entity may further send the fourth information to the model management functional entity, and the fourth information includes one or more of the following: the value of a second parameter for indicating the accuracy of the updated first model inference; the identifier of the first model; the version number of the updated first model; the update time of the first model; the data volume included in the third data set used to update the first model; or, the duration required for the updated first model inference.
[0142] S412: The model training functional entity sends the third information to the model inference functional entity. Correspondingly, the model inference functional entity receives the third information from the model training functional entity.
[0143] The third information includes one or more of the following: the identifier of the first model, the storage address of the updated first model, the update time of the first model, or, the version number of the updated first model.
[0144] After receiving the third information, the model inference functional entity may update the model parameters of the first model according to the third information. For example, the model inference functional entity may obtain the model parameters of the first model based on the storage address of the updated first model in the third information, and update the model parameters of the first model stored locally based on the identifier of the first model in the third information, such as modifying the model parameters of the first model stored locally, and storing the version number of the updated first model.
[0145] The above first information, second information, third information, and fourth information may further include other information. For example, the first information, third information, and fourth information may further include information such as the robustness of the updated first model. The embodiments of the present application do not limit the information included in the first information, second information, third information, and fourth information.
[0146] If in the foregoing S1, it is determined based on the first probability that the first algorithm is implemented by the model training functional entity, then before executing S1, the following S501 to S508 may also be executed. Please refer to Figure 5 。
[0147] S501: The model inference functional entity sends the first data set to the model training functional entity. Correspondingly, the model training functional entity receives the first data set from the model inference functional entity. Delete the third data subset in the third data set based on the third algorithm.
[0148] S502: The model management functional entity deletes the third data subset in the third data set based on the third algorithm.
[0149] S503: The model management functional entity sends the fourth indication information to the model inference functional entity. Correspondingly, the model inference functional entity receives the fourth indication information from the model management functional entity.
[0150] S504: The model inference functional entity sends the fourth indication information to the model management functional entity. Correspondingly, the model management functional entity receives the fourth indication information from the model inference functional entity.
[0151] S505: The model training functional entity determines the first probability based on the second algorithm.
[0152] S506: The model management functional entity sends the third indication information to the model training functional entity. Correspondingly, the model training functional entity receives the third indication information from the model management functional entity.
[0153] S507: The model training functional entity sends the third indication information to the model management functional entity. Correspondingly, the model management functional entity receives the third indication information from the model training functional entity.
[0154] For the relevant descriptions of the foregoing S501 to S507, please refer to the relevant descriptions of the corresponding steps of the foregoing S401 to S407, which will not be elaborated here.
[0155] S508: The model training functional entity sends the second indication information to the model management functional entity, and the second indication information is used to indicate the first probability. Correspondingly, the model management functional entity receives the second indication information from the model training functional entity.
[0156] In the embodiments of the present application, since it is determined based on the first probability that the first algorithm is implemented by the model training functional entity, the model training functional entity sending the second indication information to the model management functional entity can enable the model management functional entity to optimize the foregoing third algorithm based on the first probability.
[0157] After determining the first probability based on the second algorithm, the model training functional entity may execute the aforementioned S1, that is, determine the first algorithm based on the first probability. For the relevant description of the model training functional entity determining the second algorithm based on the first probability, reference may be made to the relevant description of the model management functional entity determining the second algorithm based on the first probability in S408, and for the relevant description of the relationship between the second indication information and the third indication information in S507, reference may be made to the relevant description of the relationship between the second indication information and the third indication information in S407 in S408, which will not be elaborated here.
[0158] Optionally, the model training functional entity may also determine the uncertainty quantification information of the first probability. For example, the model training functional entity may determine the uncertainty quantification information of the first probability through parameters such as Bayesian posterior probability. The uncertainty quantification information is used to quantify the accuracy of the first probability. Among them, the uncertainty quantification information of the first probability may be sent through the same information as the first probability, or may be sent through different information. The embodiments of the present application do not limit the sending manner of the first probability and the uncertainty quantification information of the first probability.
[0159] For the relevant description of the first algorithm determined by the model training functional entity based on the first probability and the uncertainty quantification information of the first probability, reference may be made to the relevant description of the first algorithm determined by the model management functional entity based on the first probability and the uncertainty quantification information of the first probability in S408, such as cases 1 to 4 in S408.
[0160] S509: The model training functional entity sends the first indication information to the model management functional entity. Correspondingly, the model management functional entity receives the first indication information from the model training functional entity.
[0161] For the relevant description of the first indication information, reference may be made to the relevant description of the first indication information in S409 above, which will not be elaborated here. In the embodiments of the present application, the first indication information is used to inform the model management functional entity that the algorithm for deleting the second data subset in the first data set is the first algorithm. Optionally, the first indication information in the embodiments of the present application and the aforementioned second indication information and / or third indication information may be the same indication information or different indication information. For example, the first indication information, the second indication information, and the third indication information are the same indication information; or the first indication information and the second indication information are the same indication information, and are different from the third indication information; or the first indication information and the third indication information are the same indication information, and are different from the second indication information; or the first indication information, the second indication information, and the third indication information are different indication information. It can be understood that if the first indication information and the second indication information are the same indication information and are different from the third indication information; or the first indication information and the third indication information are the same indication information and are different from the second indication information, it indicates that the second indication information and the third indication information are different indication information. If the first indication information, the second indication information, and the third indication information are the same indication information, then the second indication information and the third indication information may or may not be the same indication information.
[0162] S510: The model training functional entity deletes the second data subset in the first data set based on the first algorithm to obtain the second data set.
[0163] S511: The model training functional entity updates the first model based on the second data set.
[0164] S512: The model training functional entity sends the third information to the model inference functional entity. Correspondingly, the model inference functional entity receives the third information from the model training functional entity.
[0165] For the relevant description of S510 - S512 above, reference may be made to the relevant description of S410 - S412 above, which will not be elaborated here.
[0166] In an embodiment of the present application, taking the example that the model management functional entity, the model training functional entity, and the model inference functional entity are deployed on different devices, for example, taking the aforementioned scenario 2 as an example. When at least two of the model management functional entity, the model training functional entity, and the model inference functional entity are deployed on the same device (such as the aforementioned scenarios 1, 3 to 6), the at least two functional entities can be regarded as one functional entity, or the at least two functional entities can obtain each other's relevant processing results. Taking scenario 3 as an example, if the model training functional entity and the model inference functional entity are deployed in the same network element, the model training functional entity can obtain the first data set from the model inference functional entity without sending it through signaling.
[0167] In the above technical solution, considering that there may be drift data in the collected data set, before the model training functional entity cleans the first data set, the algorithm for cleaning the first data set determined by the probability that there is a first data subset in the first data set with statistical distribution characteristics different from those of the training data set of the training model can improve the accuracy of cleaning the second data and further improve the robustness of the model.
[0168] Figure 6 The structural schematic diagram of a communication device provided by an embodiment of the present application is given. The communication device 600 may be Figure 4 or Figure 5 the model management functional entity described in the embodiment shown, for implementing the method corresponding to the model management functional entity in the above method embodiment; or the communication device 600 may be Figure 4 or Figure 5 the model training functional entity described in the embodiment shown, for implementing the method corresponding to the model training functional entity in the above method embodiment; or the communication device 600 may be Figure 4 or Figure 5 the model inference functional entity described in the embodiment shown, for implementing the method corresponding to the model inference functional entity in the above method embodiment.
[0169] The communication device 600 includes at least one processor 601. The processor 601 can be used for internal processing of the device to implement certain control processing functions. Optionally, the processor 601 includes instructions. Optionally, the processor 601 can store data. Optionally, different processors can be independent devices, can be located at different physical locations, and can be located on different integrated circuits. Optionally, different processors can be integrated into one or more processors, for example, integrated on one or more integrated circuits.
[0170] Optionally, the communication device 600 includes one or more memories 603 for storing instructions. Optionally, data may also be stored in the memory 603. The processor and the memory may be provided separately or integrated together.
[0171] Optionally, the communication device 600 includes a communication line 602 and at least one communication interface 604. Among them, since the memory 603, the communication line 602, and the communication interface 604 are all optional, they are Figure 6 represented by dashed lines in the figure.
[0172] Optionally, the communication device 600 may further include a transceiver and / or an antenna. Among them, the transceiver may be used to send information to other devices or receive information from other devices. The transceiver may be referred to as a transceiver, a transceiver circuit, an input / output interface, etc., and is used to implement the transceiver function of the communication device 600 through the antenna. Optionally, the transceiver includes a transmitter and a receiver. Exemplarily, the transmitter may be used to generate a radio frequency signal from a baseband signal, and the receiver may be used to convert a radio frequency signal into a baseband signal.
[0173] The processor 601 may include a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application solution.
[0174] The communication line 602 may include a path for transmitting information between the above components.
[0175] The communication interface 604 uses any device of the transceiver type for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), wired access networks, etc.
[0176] The memory 603 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 603 can exist independently and be connected to the processor 601 through the communication line 602. Alternatively, the memory 603 can also be integrated with the processor 601.
[0177] Among them, the memory 603 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 601 to execute. The processor 601 is used to execute the computer execution instructions stored in the memory 603, so as to implement Figure 4 or Figure 5 the steps performed by the model management functional entity, the model training functional entity, or the model inference functional entity described in the embodiments shown.
[0178] Optionally, the computer execution instructions in the embodiments of this application can also be referred to as application code, and this application does not make specific limitations on this.
[0179] In a specific implementation, as an embodiment, the processor 601 can include one or more CPUs, such as Figure 6 CPU0 and CPU1 in
[0180] In a specific implementation, as an embodiment, the communication device 600 can include multiple processors, such as Figure 6 the processor 601 and the processor 605 in
[0181] When Figure 6When the device shown is a chip, such as a chip of a model management functional entity, a model training functional entity, or a model inference functional entity, the chip includes a processor 601 (which may also include a processor 605), a communication line 602, and a communication interface 604. Optionally, it may include a memory 603. Specifically, the communication interface 604 may be an input interface, a pin, a circuit, etc. The memory 603 may be a register, a cache, etc. The processor 601 and the processor 605 may be a general-purpose CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of a program for the semantic communication method in any of the above embodiments.
[0182] Embodiments of the present application may divide functional modules for the device according to the above method examples. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. For example, in the case of dividing each functional module corresponding to each function, Figure 7 A schematic diagram of a device is shown. The device 700 may be the model management functional entity, the model training functional entity, or the model inference functional entity involved in the above method embodiments, or a chip in the model management functional entity, the model training functional entity, or the model inference functional entity. The device 700 includes a sending unit 701, a processing unit 702, and a receiving unit 703.
[0183] It should be understood that the device 700 may be used to implement the steps executed by the model management functional entity, the model training functional entity, or the model inference functional entity in the model management method of the embodiments of the present application. Related features may refer to any one of the above Figure 4 or Figure 5 shown embodiments, which will not be elaborated here.
[0184] Optionally, Figure 7 the functions / implementation processes of the sending unit 701, the receiving unit 703, and the processing unit 702 in Figure 6 may be implemented by the processor 601 in Figure 7 calling computer-executable instructions stored in the memory 603. Or, Figure 6 the function / implementation process of the processing unit 702 in Figure 7 may be implemented by the processor 601 in Figure 6 calling computer-executable instructions stored in the memory 603, and
[0185] Optionally, when the device 700 is a chip or a circuit, the functions / implementation processes of the sending unit 701 and the receiving unit 703 can also be implemented through pins or circuits, etc.
[0186] The present application also provides a computer-readable storage medium storing a computer program or instructions. When the computer program or instructions are run, the methods executed by the model management functional entity, the model training functional entity, or the model inference functional entity in the foregoing method embodiments are implemented. In this way, the functions described in the above embodiments can be implemented in the form of software functional units and sold or used as independent products. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0187] The present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer is caused to execute the methods executed by the model management functional entity, the model training functional entity, or the model inference functional entity in any of the foregoing method embodiments.
[0188] The embodiments of the present application also provide a processing device, including a processor and an interface; the processor is used to execute the methods executed by the model management functional entity, the model training functional entity, or the model inference functional entity involved in any of the foregoing method embodiments.
[0189] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0190] In the embodiments of the present application, the various illustrative logical units and circuits described can be implemented or operated to perform the described functions by a design of a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0191] The steps of the methods or algorithms described in the embodiments of this application can be directly embedded in hardware, software units executed by a processor, or a combination of both. The software units can be stored in a RAM, flash memory, ROM, erasable programmable read-only memory (EPROM), EEPROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC, and the ASIC can be provided in a network device. Optionally, the processor and the storage medium can also be provided in different components of the network device.
[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in multiple blocks or blocks.
[0193] The content in the various embodiments of this application can be referred to each other. Without special instructions and logical conflicts, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0194] It can be understood that in the embodiments of this application, the network device can execute some or all of the steps in the embodiments of this application. These steps or operations are only examples. In the embodiments of this application, other operations or various deformations of the operations can also be executed. In addition, the various steps can be executed in different orders presented in the embodiments of this application, and it is possible not to execute all the operations in the embodiments of this application.
Claims
1. A model management method, characterized in that, The method includes: Determining a first algorithm according to a first probability; wherein, The first probability is the probability that a first data subset exists in a first dataset, and the statistical distribution characteristics of the first data subset are different from those of the training dataset used to train the first model, and the first data subset is used to update the first model; The first algorithm is used to delete a second data subset from the first dataset to obtain a second dataset. The statistical distribution characteristics of the second data subset are different from those of the training dataset, and the second data subset is not used to update the first model, and the second dataset is used to update the first model.
2. The method according to claim 1, wherein The method further includes: Sending first indication information, where the first indication information is used to indicate the first algorithm.
3. The method according to claim 1 or 2, characterized in that, The first probability is determined by a model training functional entity based on a second algorithm.
4. The method according to claim 3, characterized in that, The method further includes: Sending second indication information to a model management functional entity, where the second indication information is used to indicate the first probability.
5. The method according to claim 3 or 4, characterized in that, The method further includes: Sending third indication information, where the third indication information is used to indicate the second algorithm.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Receiving the first dataset from a model inference functional entity.
7. The method according to any one of claims 1 to 3 and 5, characterized in that The method further includes: Sending fourth indication information to the model inference functional entity, or receiving fourth indication information from the model inference functional entity. The fourth indication information is used to indicate a third algorithm, and the third algorithm is used to delete a third data subset from a third dataset to obtain the first dataset. The third dataset is collected by the model inference functional entity, and the statistical distribution characteristics of the third data subset are different from those of the training dataset, and the third data subset is not used to update the first model.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Optimizing the first algorithm based on the second data subset and / or the value of a first parameter of the data in the second data subset. The value of the first parameter is determined based on the first algorithm, and the first parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model.
9. The method according to claim 8, wherein The method further includes: Sending first information to a model management functional entity, where the first information includes the second data subset and / or the value of the first parameter of the data in the second data subset.
10. The method according to claim 8 or 9, characterized in that, The first parameter includes one or more of the following: The change amount of the loss value, where the change amount of the loss value is used to indicate the degree of influence of the data on the loss function of the first model; The second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics, where the degree of difference in statistical distribution characteristics is used to indicate the degree of difference in the statistical distribution characteristics of the data in the second data subset and the data in the training dataset.
11. The method according to claim 9 or 10, characterized in that, The first information further includes one or more of the following: The value of a second parameter, where the second parameter is used to indicate the inference accuracy of the updated first model; The identifier of the first model; The version number of the updated first model; The update time of the first model; The data volume included in the third dataset used to update the first model; or, The duration required for the inference of the updated first model.
12. The method according to any one of claims 1 to 3, 5, and 7, characterized in that The method further includes: Optimizing a third algorithm based on a third data subset and / or a value of a third parameter of data in the third data subset, where the value of the third parameter is determined based on a third algorithm for deleting the third data subset, the third parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model, the statistical distribution characteristics of the third data subset are different from those of the training data set, and the third data subset is not used to update the first model.
13. The method according to claim 12, wherein The method further includes: Receiving second information from a model inference functional entity, where the second information includes the third data subset and / or the value of the third parameter of the data in the third data subset.
14. The method according to claim 12 or 13, characterized in that The third parameter includes one or more of the following: The change amount of the loss value, where the change amount of the loss value is used to indicate the degree of influence of the data on the loss function of the first model; The second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics, where the degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the third data subset and those of the data in the training data set.
15. The method according to any one of claims 1 to 6, 8 to 11, characterized in that, The method further includes: Sending third information to a model inference functional entity, where the third information includes one or more of the following: The identifier of the first model; The version number of the updated first model; The update time of the first model; or, The storage address of the updated first model.
16. A model management method, characterized in that, The method includes: Deleting a third data subset in a third data set based on a third algorithm to obtain a first data set, where the third data set is collected by a model inference functional entity, the statistical distribution characteristics of the third data subset are different from those of the training data set used to train the first model, and the third data subset is not used to update the first model; Sending the first data set to a model training functional entity.
17. The method according to claim 16, wherein The method further includes: Sending fourth indication information to a model management functional entity, or receiving fourth indication information from the model management functional entity, where the fourth indication information is used to indicate the third algorithm.
18. The method according to claim 16 or 17, characterized in that, Deleting the third data subset in the third data set based on the third algorithm includes: Determining the value of the third parameter of the data in the third data set based on the third algorithm, where the third parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model; Deleting the third data subset, where the value of the third parameter of the data in the third data subset is greater than or equal to a first threshold.
19. The method according to any one of claims 16 to 18, characterized in that, The method further includes: Optimizing the third algorithm based on the third data subset and / or the value of the third parameter of the third data subset.
20. The method according to claim 18 or 19, characterized in that, The method further includes: Sending second information to a model management functional entity, where the second information includes the third data subset and / or the value of the third parameter of the data in the third data subset.
21. The method according to any one of claims 18 to 20, characterized in that The third parameter includes one or more of the following: The change amount of the loss value, where the change amount of the loss value is used to indicate the degree of influence of the data on the loss function of the first model; The second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics, where the degree of difference in statistical distribution characteristics is used to indicate the degree of difference in the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data subset.
22. The method according to any one of claims 16 to 21, characterized in that, The method further includes: Receiving third information from a model training functional entity, where the third information includes one or more of the following: The identifier of the first model; The version number of the updated first model; The update time of the first model; or, The storage address of the updated first model.
23. The method according to claim 22, wherein The method further includes: Updating the model parameters of the first model based on the third information.
24. A model management system, characterized in that, Including a model management functional entity and a model training functional entity for executing the method according to any one of claims 1 to 15, and a model inference functional entity for executing the method according to any one of claims 16 to 23.
25. A communication device, characterized in that, Including a processor and a memory, where the memory is coupled to the processor, and the processor is configured to call computer instructions in the memory to execute the method according to any one of claims 1 to 15, or execute the method according to any one of claims 16 to 23.
26. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are called by the computer, they are used to execute the method according to any one of claims 1 to 15, or execute the method according to any one of claims 16 to 23.
27. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 15, or causes the computer to execute the method according to any one of claims 16 to 23.
28. A computer program, characterized in that, Including program code, and when the computer runs the program code, the program code executes the method according to any one of claims 1 to 15, or the program code executes the method according to any one of claims 16 to 23.
29. A chip, characterized in that, The chip is coupled to the memory and is configured to read and execute program instructions stored in the memory to implement the method according to any one of claims 1 to 15, or implement the method according to any one of claims 16 to 23.