Model management method and system
By cleaning the data set before the AI model is updated and deleting a large data subset with different statistical distribution characteristics from the training data set, the problem of possible pseudo-data when updating the AI model is solved, and the robustness and prediction performance of the model are improved.
Patent Information
- Application Number
- PCT/CN2024/133995
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-26
AI Technical Summary
When updating AI models, the prior art may contain pseudo-data generated by cyber attack devices, affecting the robustness of the model.
By cleaning the data set before the model is updated, the first probability is determined to determine the algorithm for cleaning, deleting a large data subset with different statistical distribution characteristics from the training data set, and improving the accuracy of data cleaning.
Reduces the probability of pseudo-data in the dataset used for model updates, and improves the robustness and prediction performance of the model.
Smart Images

Figure CN2024133995_26062025_PF_FP_ABST
Abstract
Description
Model management method and system
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 22, 2023, with application number 202311796301.5 and application name “A Model Management Method and System”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of communication technology, and in particular to a model management method and system. Background Art
[0004] To improve the intelligence and automation of networks, technologies such as artificial intelligence (AI) and machine learning (ML) are being applied in an increasing number of fields. For example, AI systems can be used to build, deploy, train, and apply models.
[0005] Currently, to improve the robustness of AI models, it's necessary to manage the model's lifecycle. For example, over time, the accuracy of AI model predictions may decrease, meaning the prediction error increases. To address this issue, it's often necessary to regularly retrain the AI model with new data—that is, update the AI model—to maintain its predictive performance. However, the dataset used to update the AI model may include data generated by cyberattack devices, compromising the AI model's robustness. Summary of the Invention
[0006] Embodiments of the present application provide a model management method and system for improving the robustness of a model.
[0007] In a first aspect, a model management method is provided, the method comprising: determining a first algorithm based on a first probability; wherein the first probability is a probability that a first data subset exists in a first data set, the statistical distribution characteristics of the first data subset are different from the statistical distribution characteristics of a training data set used to train a first model, and the first data subset is used to update the first model; the first algorithm is used to delete a second data subset in the first data set to obtain a second data set, the statistical distribution characteristics of the second data subset are different from the statistical distribution characteristics of the training data set, the second data subset is not used to update the first model, and the second data set is used to update the first model.
[0008] In an embodiment of the present application, before using the first dataset to update the first model, the first dataset can be cleaned, which helps reduce the probability of false data (such as data generated by a network attack device) in the dataset used for model update and improves the robustness of the model. In addition, the algorithm used to clean the first dataset, which is determined by the probability of the first data subset in the first dataset having statistical distribution characteristics different from the training dataset of the training model, can improve the accuracy of cleaning the second data and further improve the robustness of the model.
[0009] The above method can be implemented by a model management functional entity or a model training functional entity. For example, if the above method is implemented by a model management functional entity, the specific implementation method is as follows:
[0010] In a possible implementation, the model management function entity may further send first indication information to the model training function entity, or receive first indication information from the model training function entity, where the first indication information is used to indicate the first algorithm. The first algorithm may be determined by the model management function entity or by the model training function entity. If the first algorithm is determined by the model management function entity, the model management function entity may send the first indication information to the model training function entity to instruct the model training function entity to delete the second data subset in the first data set based on the first algorithm; if the first algorithm is determined by the model training function entity, the model training function entity may send the first indication information to the model management function entity to inform the model management function entity of the first algorithm used when deleting the second data subset in the first data set.
[0011] In one possible implementation, before sending the first indication information to the model training function entity, the model management function entity may further receive second indication information from the model training function entity, where the second indication information is used to indicate the first probability. Since the first algorithm is determined based on the first probability, if the first algorithm is determined by the model management function entity, the model management function entity also needs to receive the first probability from the model training function entity.
[0012] In a possible implementation, the model management function entity may further determine the first algorithm based on the first probability. The first algorithm determined based on the first probability may improve the accuracy of cleaning the second data and further improve the robustness of the model.
[0013] In a possible implementation, the model management function entity may further send third indication information to the model training function entity, or receive third indication information from the model training function entity, wherein the third indication information is used to indicate a second algorithm, and the second algorithm is used to determine the first probability. The second algorithm may be determined by the model management function entity or by the model training function entity. If the second algorithm is determined by the model management function entity, the model management function entity may send the third indication information to the model training function entity to instruct the model training function entity to determine the first probability that the first data subset exists in the first data set based on the second algorithm; if the first algorithm is determined by the model training function entity, the model training function entity may send the third indication information to the model management function entity to inform the model management function entity of the first algorithm used when determining the first probability. If the second algorithm is determined by the model training function entity, the third indication information and the aforementioned second indication information may be the same information or different information.
[0014] In a possible embodiment, the model management functional entity may also send a fourth indication message to the model reasoning functional entity, or receive a fourth indication message from the model reasoning functional entity, wherein the fourth indication message is used to indicate a third algorithm, wherein the third algorithm is used to delete a third data subset in a third data set to obtain the first data set, wherein the third data set is collected by the model reasoning functional entity, and the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set, and the third data subset is not used to update the first model. By having the model reasoning functional entity clean the third data set it collects, the amount of data sent by the model reasoning functional entity to the model training functional entity can be reduced, which helps to reduce the transmission overhead between entities and improve the data transmission speed.
[0015] In one possible implementation, the model management functional entity may also receive first information from the model training functional entity, where the first information includes a value of a first parameter of the second data subset and / or the data in the second data subset, the value of the first parameter being determined based on the first algorithm, and the first parameter being used to indicate the degree of influence of the data on the inference accuracy of the first model. The model training functional entity sends the value of the first parameter of the deleted second data subset and / or the data in the second data subset to the model management functional entity, so that the model management functional entity can learn the statistical distribution characteristics of the second data subset and optimize the first algorithm.
[0016] In one possible implementation, the model management function entity may further optimize the first algorithm based on the second data subset and / or the value of the first parameter of the data in the second data subset. The first algorithm optimized based on the second data subset and / or the value of the first parameter of the data in the second data subset may improve the accuracy of data cleaning, thereby improving the robustness of the updated model.
[0017] In one possible embodiment, the first parameter includes one or more of the following: a change in the loss value, the change in the loss value is used to indicate the degree of influence of the data on the loss function of the first model; a second probability, the second probability is used to indicate the probability that the data is abnormal data; or a degree of difference in statistical distribution characteristics, the degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the second data subset and the statistical distribution characteristics of the data in the training data set. The first parameter can be any parameter that can measure the degree of influence of data on the prediction accuracy of the model. The parameters shown in the embodiment of the present application are only examples. In other embodiments, the first parameter can also include more or fewer parameters. The embodiment of the present application does not limit the types of parameters included in the first parameter.
[0018] In one possible implementation, the first information further includes one or more of the following: a value of a second parameter, the second parameter being used to indicate the accuracy of the updated first model reasoning; an identifier of the first model; a version number of the updated first model; an update time of the first model; the amount of data included in a third data set used to update the first model; or the duration required for reasoning of the updated first model. The model training function entity may send relevant information about the model to the model management function entity, so that the model management function entity may adjust a policy for managing the model based on the relevant information about the model.
[0019] In one possible implementation, the model management function entity may further receive second information from the model reasoning function entity, where the second information includes the value of a third parameter of the third data subset and / or the data in the third data subset. The model reasoning function entity sends the deleted third data subset and / or the value of the third parameter of the data in the third data subset to the model management function entity, so that the model management function entity can learn the statistical distribution characteristics of the third data subset and optimize the third algorithm.
[0020] In one possible embodiment, the model management functional entity may also optimize a third algorithm based on a third data subset and / or the value of a third parameter of the data in the third data subset, where the value of the third parameter is determined based on the third algorithm, the third algorithm is used to delete the third data subset, the third parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model, the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set, and the third data subset is not used to update the first model. The third algorithm optimized based on the third data subset and / or the value of the third parameter of the data in the third data subset can improve the accuracy of data cleaning, thereby improving the robustness of the updated model.
[0021] In one possible embodiment, the third parameter includes one or more of the following: a change in the loss value, the change in the loss value is used to indicate the degree of influence of the data on the loss function of the first model; a second probability, the second probability is used to indicate the probability that the data is abnormal data; or, a degree of difference in statistical distribution characteristics, the degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set. The third parameter can be any parameter that can measure the degree of influence of data on the accuracy of model prediction. The parameters shown in the embodiment of the present application are only examples. In other embodiments, the third parameter can also include more or fewer parameters. The embodiment of the present application does not limit the types of parameters included in the third parameter. It should be understood that the third parameter can be the same as the aforementioned first parameter, or it can be different from the aforementioned first parameter. The embodiment of the present application does not limit the relationship between the first parameter and the third parameter.
[0022] For another example, if the above method is implemented by a model training functional entity, the specific implementation method is as follows:
[0023] In a possible implementation, the model training functional entity may also send first indication information to the model management functional entity, or receive first indication information from the model management functional entity, where the first indication information is used to indicate the first algorithm.
[0024] In a possible implementation, the model training functional entity may also determine the first probability based on a second algorithm.
[0025] In a possible implementation, the model training functional entity may further send second indication information to the model management functional entity, where the second indication information is used to indicate the first probability.
[0026] In a possible implementation, the model training functional entity may further send third indication information to the model management functional entity, or receive third indication information from the model management functional entity, where the third indication information is used to indicate the second algorithm.
[0027] In a possible implementation, the model training functional entity may also receive the first data set from the model reasoning functional entity.
[0028] In one possible embodiment, the model training functional entity can also optimize the first algorithm based on the value of a first parameter of the second data subset and / or the data in the second data subset, where the first parameter is used to indicate the degree of influence of the data on the reasoning accuracy of the first model.
[0029] In a possible implementation, the model training functional entity may further send first information to the model management functional entity, where the first information includes the value of a first parameter of the second data subset and / or the data in the second data subset.
[0030] In one possible embodiment, the first parameter includes one or more of the following: a change in the loss value, which is used to indicate the degree of influence of the data on the loss function of the first model; a second probability, which is used to indicate the probability that the data is abnormal data; or a degree of difference in statistical distribution characteristics, which is used to indicate the degree of difference in the statistical distribution characteristics of the data in the second data subset and the statistical distribution characteristics of the data in the training data set.
[0031] In one possible embodiment, the first information also includes one or more of the following: the value of a second parameter, where the second parameter is used to indicate the accuracy of the reasoning of the updated first model; the identifier of the first model; the version number of the updated first model; the update time of the first model; the amount of data included in the third data set used to update the first model; or the time required for reasoning of the updated first model.
[0032] In one possible implementation, the model training function entity may further send third information to the model reasoning function entity, where the third information includes one or more of the following: an identifier of the first model; a version number of the updated first model; an update time of the first model; or a storage address of the updated first model. The model training function entity may send model parameter information of the model to the model reasoning function entity, so that the model reasoning function entity may update the model parameters of the first model based on the third information.
[0033] In a second aspect, a model management method is provided, which can be implemented by a model reasoning functional entity. The method includes: deleting a third data subset in a third data set based on a third algorithm to obtain a first data set, wherein the third data set is collected by the model reasoning functional entity, the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set used to train the first model, and the third data subset is not used to update the first model; and sending the first data set to the model training functional entity.
[0034] In a possible implementation, the model reasoning function entity may further send fourth indication information to the model management function entity, or receive fourth indication information from the model management function entity, where the fourth indication information is used to indicate the third algorithm.
[0035] In one possible implementation, deleting a third data subset in a third data set based on a third algorithm includes: determining a value of a third parameter of the data in the third data set based on the third algorithm, the third parameter being used to indicate a degree of influence of the data on the accuracy of inference of the first model; and deleting the third data subset, wherein the value of the third parameter of the data in the third data subset is greater than or equal to a first threshold.
[0036] In a possible implementation, the model reasoning functional entity may further optimize the third algorithm based on the third data subset and / or the value of a third parameter of the third data subset.
[0037] In a possible implementation, the model reasoning functional entity may further send second information to the model management functional entity, where the second information includes the third data subset and / or a value of a third parameter of the data in the third data subset.
[0038] In one possible embodiment, the third parameter includes one or more of the following: a change in the loss value, which is used to indicate the degree of influence of the data on the loss function of the first model; a second probability, which is used to indicate the probability that the data is abnormal data; or a degree of difference in statistical distribution characteristics, which is used to indicate the degree of difference between the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set.
[0039] In a possible embodiment, the model reasoning functional entity may also receive third information from the model training functional entity, wherein the third information includes one or more of the following: an identifier of the first model; a version number of the updated first model; an update time of the first model; or a storage address of the updated first model.
[0040] In a possible implementation, the model reasoning function entity may also update the model parameters of the first model based on the third information.
[0041] In a third aspect, a model management system is provided, comprising a model management functional entity and a model training functional entity for executing any implementation method in the first aspect, and a model reasoning functional entity for executing any implementation method in the second aspect.
[0042] In a fourth aspect, embodiments of the present application provide a communication device that implements any of the methods implemented by the model management functional entity in the first aspect. This functionality can be implemented via hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functionality.
[0043] In a fifth aspect, an embodiment of the present application provides a communication device that has the function of implementing any method implemented by the model training functional entity in the first aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0044] In a sixth aspect, an embodiment of the present application provides a communication device having the function of implementing any method implemented by the model training functional entity in the second aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0045] In a seventh aspect, an embodiment of the present application provides a communication device. The communication device includes a communication interface and a processor, and optionally, a memory. The memory is used to store a computer program, and the processor is coupled to the memory and the communication interface. When the processor reads the computer program or instructions, the communication device executes the method performed by the model management functional entity in the first aspect, or the method performed by the model training functional entity in the first aspect, or the method performed by the model inference functional entity in the second aspect.
[0046] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is run on a computer, the computer executes the method provided in the first or second aspect above.
[0047] In a ninth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed on a computer, enables the computer to execute the method described in the first or second aspect above.
[0048] In the tenth aspect, a chip system is provided, comprising a processor and an interface, wherein the processor is used to call and execute instructions from the interface so that the chip system implements the method described in the first or second aspect above.
[0049] For the beneficial effects of the second to tenth aspects mentioned above, refer to the beneficial effects of the first aspect and will not be repeated. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 shows the process of beam management performed by the wireless AI model;
[0051] 2A and 2B are schematic diagrams of two network architectures provided in embodiments of the present application;
[0052] FIG2C is a structural diagram of a model management system provided in an embodiment of the present application;
[0053] FIG3 is a schematic diagram of a scenario for collecting a data set;
[0054] Figures 4 and 5 are flowcharts of two model management methods provided in embodiments of the present application;
[0055] FIG6 is a structural diagram of a communication device provided in an embodiment of the present application;
[0056] FIG7 is a structural diagram of another communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. The embodiments of the present application can be applied to various mobile communication systems, such as new radio (NR) systems, long term evolution (LTE) systems, advanced long term evolution (LTE-A) systems, future communication systems and other communication systems, without limitation. For ease of understanding, some of the terms in the embodiments of the present application are explained below. In order to make the embodiments of the present application clearer, some of the contents and concepts related to the embodiments of the present application are uniformly introduced here.
[0058] (1) The terminal device in the embodiment of the present application is a device with wireless transceiver function, which can be a fixed device, a mobile device, a handheld device (such as a mobile phone), a wearable device, a vehicle-mounted device, or a wireless device built into the above device (such as a communication module, a modem, or a chip system, etc.). The terminal device is used to connect people, objects, machines, etc., and can be widely used in various scenarios, such as but not limited to the following scenarios: cellular communication, device-to-device communication (D2D), vehicle to everything (V2X), machine-to-machine / machine-type communication (M2M / MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drone, robot and other scenarios. The terminal device may sometimes be referred to as user equipment (UE), terminal, access station, UE station, remote station, wireless communication device, or user equipment, etc.
[0059] (2) Network management system (NMS) is a management service consumer responsible for network operation, management, and maintenance. The network management system is a traditional cross-domain management system, or it can be a partial set of management functions for cross-domain management.
[0060] (3) The element management system (EMS) is a management service provider used to manage network elements. The EMS can be a traditional domain management system or a single-domain management system, or it can be a partial set of management functions for single-domain management. The network management system and the EMS can be collectively referred to as the 3rd Generation Partnership Project (3GPP) management system, or the operations administration and maintenance (OAM) module.
[0061] (4) Network element (NE): An entity that provides network services. A network element is a management service provider and includes core network elements, radio access network elements, transport network elements, etc. For example, core network elements may include but are not limited to access and mobility management function (AMF) entity, session management function (SMF) entity, policy control function (PCF) entity, network data analysis function (NWDAF) entity, network repository function (NRF), gateway, etc.
[0062] A radio access network element may include, but is not limited to, a base station, an evolved Node B in a long term evolution (LTE) system or an advanced long term evolution (LTE-A), which may be referred to as an eNB or e-NodeB, a transmission reception point (TRP), a next generation NodeB (gNB) in a fifth generation (5G) mobile communication system, a next generation base station in a sixth generation (6G) mobile communication system, a base station in a future mobile communication system, or an access node in a Wi-Fi system, and may also be an access network device in an open access network (ORAN) system. Optionally, a network device may also be a module or unit that performs some of the functions of a base station, for example, a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU). Among them, the CU can complete the functions of the radio resource control protocol and packet data convergence protocol (PDCP) of the base station, and can also complete the function of the service data adaptation protocol (SDAP); the DU can complete the functions of the radio link control layer and medium access control (MAC) layer of the base station, and can also complete the functions of part or all of the physical layer. In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in the ORAN system, CU can also be called O-CU, DU can also be called open (open, O)-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CUP-UP, and RU can also be called O-RU.
[0063] (5) The terms "system" and "network" in the embodiments of the present application can be used interchangeably. "At least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. The following at least one item or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0064] Unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects.
[0065] In addition, the terms "including" and "having" in the embodiments, claims, and drawings of this application are not exclusive. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules and may also include steps or modules that are not listed.
[0066] Currently, AI / ML technologies can be applied in multiple fields to solve problems that traditional algorithms cannot. For example, the wireless AI model (hereinafter referred to as the model) can be used to feedback channel state information (CSI), positioning, beam management, energy saving management, load balancing, and mobility optimization. For ease of understanding, the implementation process of the wireless AI model is first illustrated with reference to Figure 1.
[0067] Please refer to Figure 1 for the beam management process for the wireless AI model.
[0068] S101: The terminal device scans the full beam to obtain data for model training or updating.
[0069] The data used for training or updating may include, for example, reference signal receiving power (RSRP) corresponding to all beams and best beam information (for example, best beam identification ID).
[0070] S102: The terminal device sends the acquired data to the network device. Correspondingly, the network device receives the data from the terminal device.
[0071] S103: The network device trains or updates a model based on the data.
[0072] S104: The terminal device performs sparse sampling on the full beam to obtain a synchronization signal block (SSB) beam.
[0073] S105: The terminal device scans the SSB beam to obtain a measurement result of the SSB beam.
[0074] S106: The terminal device sends the measurement result of the SSB beam to the network device. Correspondingly, the network device receives the measurement result from the terminal device.
[0075] The measurement result is, for example, the reference signal receiving power (RSRP) corresponding to the SSB beam and the best beam information (for example, the best beam identifier ID).
[0076] S107: The model in the network device predicts the corresponding K best beams based on the measurement results (such as RSRP) of the SSB beam.
[0077] The K best beams are the K beams with the greatest probability of being the best beams among the SSB beams.
[0078] S108: The network device sends the beam IDs of the K best beams to the terminal device. Correspondingly, the terminal device receives the beam IDs of the K best beams from the network device.
[0079] S109: The terminal device determines the channel state information reference signal (CSI-RS) corresponding to the K best beams.
[0080] S110: The terminal device sends the beam ID of the best CSI-RS beam to the network device.
[0081] S111: The network device sends data to the terminal device through the optimal CSI-RS beam.
[0082] Please refer to Figure 2A, which is a schematic diagram of a network architecture provided in an embodiment of the present application. In Figure 2A, the network architecture includes a management service (MnS) provider (producer) and a management service consumer (MnS consumer). The management service provider is used to provide network management services to the outside world, and the management service consumer can call the management service to provide network management services to the outside world.
[0083] Please refer to Figure 2B, which is a schematic diagram of another network architecture provided by an embodiment of the present application. In Figure 2B, the network architecture includes network elements, a network element management system, and a network management system.
[0084] Please refer to Figure 2C, which shows the model management system architecture provided in an embodiment of the present application. In Figure 2C, the model management system architecture includes: a model management function (MMF) entity, a model training function (MTF) entity, and a model inference function (MIF) entity. The model management function entity, the model training function entity, and the model inference function entity can communicate with each other. Among them, the model management function entity can be used to implement the management of the model life cycle, can trigger the model training function entity to perform model training or update, and can trigger the model inference function entity to perform inference. The model training function entity can be used to train and update the model and output the model file after the model training or update is completed. The model inference function entity can be used for model inference, that is, inputting data into the model to obtain the corresponding predicted output result. Optionally, the model inference function entity can also collect the output results of the model to obtain a data set for updating the model. Optionally, the model management system architecture can also include other entities, for example, it can also include a model evaluation function (MEF) entity for evaluating the performance of the model. The embodiment of the present application does not limit the types of entities included in the model management system architecture.
[0085] Optionally, taking model update as an example, the model management system may be connected to a model update trigger module (not shown in the figure), and the model update trigger module may send model management information to the model management system.
[0086] For example, the model update trigger module can send the identifier of the model to be updated (e.g., Model 1) and information such as the model management type and requirements to the model management functional entity. The model identifier is the model's identification code, used to distinguish different models. The model management type and requirements refer to the management measures and implementation standards taken to ensure the model's credibility and robustness during the model management process. Such management measures may include, for example, data integrity verification, monitoring methods for model performance (e.g., continuous monitoring), and ensuring the robustness of the model update process. For example, if the data used for model training is an 8*4 RSRP matrix, verifying the data integrity may involve, for example, verifying whether data is missing from the RSRP matrix. If data is missing, the data is determined to be incomplete. Alternatively, the data may be verified to have a corresponding label. For example, taking Model 1 for beam management as an example, the label may be, for example, the optimal beam ID. If there is no label, the data is determined to be incomplete. The model management functional entity may determine the model management strategy based on the model management information and model management constraints. Model management constraints, for example, are constraints built into the model management functional entity. These constraints are used to constrain the accuracy of model predictions and data overhead, such as the proportion of excluded data to collected data. Model management strategies, for example, may include algorithms used during model updates.
[0087] The model update trigger module may also send information such as the model identifier and model update strategy of the model 1 to the model training functional entity. The model update strategy may be used to indicate the method of model update and the conditions for terminating the model update. The method of model update may be, for example, a parameter fine-tuning update or an incremental update. The conditions for terminating the model update may include, for example, that the accuracy of the updated model prediction reaches a threshold (for example, threshold 1) and / or that the model update duration exceeds a preset update duration.
[0088] The model update module can also send information such as the identifier of model 1 and the data set collection strategy to the model reasoning functional entity. Among them, the data set collection strategy can be used to indicate the type of data collected, which data type includes, for example, RSRP, channel impulse response (CIR) and other data. Optionally, the data set collection strategy can also indicate how to select the collected data to ensure that the collected data is the latest, most relevant and most diverse data. The model reasoning functional entity can collect data based on the data set collection strategy to obtain a data set for model updating (such as the second data set described below).
[0089] The model management functional entity can be the management service provider shown in Figure 2A, or it can also be the management service consumer shown in Figure 2A. The model training functional entity and the model inference functional entity are the management service providers shown in Figure 2A. Therefore, the model management system architecture can have the following deployment scenarios:
[0090] Scenario 1: The model management functional entity is the management service consumer, as shown in Figure 2A. The model training functional entity and the model inference functional entity are the management service providers, as shown in Figure 2A. The model management functional entity is deployed in the network management system, as shown in Figure 2B. The model training functional entity and the model inference functional entity are deployed in the network element management system, as shown in Figure 2B.
[0091] Scenario 2: The model management functional entity is the management service consumer shown in Figure 2A, and the model training functional entity and model inference functional entity are the management service providers shown in Figure 2A. The model management functional entity is deployed in the network management system shown in Figure 2B, the model training functional entity is deployed in the network element management system shown in Figure 2B, and the model inference functional entity is deployed in the network element shown in Figure 2B.
[0092] Scenario 3: The model management functional entity is the management service consumer shown in Figure 2A, and the model training functional entity and model inference functional entity are the management service providers shown in Figure 2A. The model management functional entity is deployed in the network management system shown in Figure 2B, and the model training functional entity and model inference functional entity are deployed in the network element shown in Figure 2B.
[0093] Scenario 4: The model management functional entity, model training functional entity, and model reasoning functional entity are the management service providers shown in Figure 2A. The model management functional entity, model training functional entity, and model reasoning functional entity are all deployed in the network element management system shown in Figure 2B.
[0094] Scenario 5: The model management functional entity, model training functional entity, and model reasoning functional entity are the management service providers shown in Figure 2A. The model management functional entity and model training functional entity are deployed in the network element management system shown in Figure 2B, and the model reasoning functional entity is deployed in the network element shown in Figure 2B.
[0095] Scenario 6: The model management functional entity, model training functional entity, and model reasoning functional entity are the management service providers shown in Figure 2A. The model management functional entity is deployed in the network management system shown in Figure 2B, and the model training functional entity and model reasoning functional entity are deployed in the network element shown in Figure 2B.
[0096] The above-mentioned model management system can be used for model training and model updating. During model training or model updating, the model management functional entity can instruct the model training functional entity to perform model training or updating based on data collected by the model reasoning functional entity to generate a model or update the model (e.g., update model parameters). The data collected by the model reasoning functional entity is generally data actually generated by the terminal device. Therefore, a terminal device used to conduct a network attack (also referred to as a network attack device) can conduct a network attack by reporting false data, resulting in lower performance of the trained or updated model. For example, referring to FIG3 , a network attack device can report forged data (e.g., a false RSRP) to a network device (e.g., a gNB), so that the data set collected by the model reasoning functional entity for updating the model includes the forged data and normal data (e.g., normal RSRP) reported by terminal devices that do not conduct network attacks (also referred to as non-network attack devices, i.e., normal devices). Therefore, after the model training functional entity performs a model update based on the data collected by the model reasoning functional entity, it may obtain an incorrect prediction model, such as an incorrect classification boundary, which ultimately leads to a decrease in the prediction performance of the updated model.
[0097] In view of this, an embodiment of the present application provides a model management method for cleaning the data set used to update the model before the model is updated, reducing the probability of pseudo data existing in the data used for model updating, and improving the robustness of the model.
[0098] The method provided in the embodiment of the present application is described in detail below in conjunction with the introduction of the above-mentioned Figures 1 and 2A to 2C.
[0099] Please refer to Figures 4 and 5 for a model management method provided in an embodiment of the present application. This method can be implemented by the model management system shown in Figure 2C. This embodiment of the present application uses the model management system for model update as an example. In Figures 4 and 5, all optional steps in the embodiment of the present application are represented by dotted lines.
[0100] S1: Determine a first algorithm based on a first probability.
[0101] As time goes by, the data distribution of the input model may change, resulting in an increase in the model prediction error, that is, the model drift phenomenon occurs. Therefore, it is necessary to use the collected data set to update the model to maintain the prediction performance of the model. As mentioned above, the data set used to update the model may contain pseudo data generated by the network attack device (such as the second data subset and the third data subset described below), resulting in low robustness of the model. At present, in order to maintain the prediction performance of the model, before updating the model, part of the data in the collected data set can be deleted. For example, the pseudo data (or poisonous data) in the collected data set can be deleted based on a sample detection method based on a loss threshold. Among them, the sample detection method based on the loss threshold can be, for example, an out-of-distribution detector (OOD Detection), or it can also be a method such as a Mahalanobis distance-based score.
[0102] However, the core concept behind the loss threshold-based sample detection method for deleting pseudo data from a collected dataset (e.g., dataset 1) is that dataset 1 is assumed to contain only normal data and pseudo data. The statistical distribution characteristics of the normal data are identical to those of the training dataset used during model training, while the statistical distribution characteristics of the pseudo data differ significantly from those of the training dataset. Therefore, pseudo data in dataset 1 can be deleted based on the principle that the loss performance of normal data and pseudo data on the same model differs significantly. That is, when the loss threshold-based sample detection method deletes pseudo data from dataset 1, it may delete all data in dataset 1 whose statistical distribution characteristics differ from those of the training dataset. However, dataset 1 may also contain data whose statistical distribution characteristics differ from those of the training dataset but were generated during model application due to model prediction accuracy issues. This means that dataset 1 may contain data whose statistical distribution characteristics differ from those of the training dataset (e.g., drift data) generated by non-network attack devices. If this data is deleted, the prediction performance of the updated model based on dataset 1 will still be low.
[0103] Therefore, when it is determined that a model update is required, for example, when the first model needs to be updated, an algorithm (for example, a first algorithm) for deleting pseudo data (for example, a second data subset) in the first data set can be determined based on a first probability that drift data (for example, a first data subset) exists in the data set (for example, the first data set). Optionally, the determination of the first algorithm based on the first probability can be implemented by a model management functional entity in the model management system, or it can also be implemented by a model training functional entity. If the determination of the first algorithm based on the first probability is implemented by the model management functional entity, then before executing S1, the following S401 to S408 can also be executed, please refer to Figure 4.
[0104] S401: The model inference function entity sends a first data set to the model training function entity. Correspondingly, the model training function entity receives the first data set from the model inference function entity and deletes a third data subset in the third data set based on a third algorithm.
[0105] The first data set may be an original data set (e.g., a third data set) collected by the model reasoning functional entity, or may be a data set obtained by the model reasoning functional entity after deleting some data from the third data set. Therefore, optionally, S402 may be performed before S401.
[0106] S402: The model reasoning functional entity deletes a third data subset in the third data set based on a third algorithm.
[0107] The data in the third data subset is, for example, the data generated by the aforementioned network attack device. Before sending the first data set to the model training function entity, the model reasoning function entity may delete the pseudo data (i.e., the third data subset) in the third data set based on a third algorithm to obtain the first data set. The third algorithm is, for example, the aforementioned sample detection method based on the loss threshold. For example, the model reasoning function entity may determine the value of the third parameter corresponding to each data in the third data set based on the third algorithm, delete the data whose value of the third parameter is greater than or equal to the first threshold from the third data set, and obtain the first data set. The set formed by the data whose value of the third parameter is greater than or equal to the first threshold is the aforementioned third data subset. Among them, the third parameter is used to indicate the degree of influence of the data on the accuracy of the first model reasoning. The third parameter may, for example, include one or more of the following: the change in the loss value, and the change in the loss value is used to indicate the degree of influence of the data on the loss function (loss) of the first model. The change in the loss value is, for example, the change in the loss value of the loss function of the first model caused by the data (for example, the data in the third data subset) relative to the loss value of the loss function when the first model is trained; the second probability, the second probability is used to indicate the probability that the data is abnormal data; or the degree of difference in statistical distribution characteristics, the degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the third data subset and the statistical distribution characteristics of the training data set.
[0108] Among them, if the change in the loss value corresponding to the data is greater than or equal to the first threshold, or the probability that the data is abnormal data is greater than or equal to the first threshold, or the degree of difference between the statistical distribution characteristics of the data and the statistical distribution characteristics of the training data set is greater than or equal to the first threshold, it indicates that the statistical distribution characteristics of the data are significantly different from the statistical distribution characteristics of the training data set, and therefore the data can be considered to be pseudo data, and therefore the data can be deleted from the third data set, and the set formed by the data deleted from the third data set is the third data subset.
[0109] Optionally, the model reasoning functional entity may also determine that the data whose value of the third parameter in the third data set is less than or equal to the second threshold is the aforementioned normal data, that is, the data whose statistical distribution characteristics are the same as the statistical distribution characteristics of the training data set, and determine that the data whose value of the third parameter in the third data set is greater than the second threshold and less than the first threshold is suspicious data, and the suspicious data is, for example, based on the third algorithm, it is impossible to confirm whether the data is pseudo data, drift data or normal data. Therefore, the model reasoning functional entity may also add a suspicious identifier to the suspicious data in the third data set, so that part of the data in the first data set obtained after deleting the third data subset in the third data set carries the suspicious identifier. The suspicious data may include one or more of the following data: the aforementioned drift data, pseudo data or normal data.
[0110] The above-mentioned third algorithm can be determined by the model reasoning functional entity, or it can be instructed by the model management functional entity. If the third algorithm is determined by the model management functional entity, then before executing S402, S403 can also be executed, and the model management functional entity sends a fourth indication information to the model reasoning functional entity. Accordingly, the model reasoning functional entity receives the fourth indication information from the model management functional entity. The fourth indication information is used to instruct the model reasoning functional entity to delete the third data subset in the third data set based on the third algorithm. If the third algorithm is determined by the model reasoning functional entity, then after determining the third algorithm, the model reasoning functional entity can also execute S404, and the model reasoning functional entity sends a fourth indication information to the model management functional entity. Accordingly, the model management functional entity receives the fourth indication information from the model reasoning functional entity. The fourth indication information is used to inform the model management functional entity that the first data set sent to the model training functional entity is the data set obtained after deleting part of the data in the collected third data set, and the algorithm used to delete the part of the data is the third algorithm.
[0111] Optionally, the fourth indication information may directly indicate the third algorithm or may indicate an identifier of the third algorithm. For example, the protocol predefines multiple algorithms, each of which has a unique identifier, such as the algorithm number. Therefore, the fourth indication information may directly indicate the number. Upon receiving the fourth indication information, the model management function entity or the model reasoning function entity may determine the algorithm to be used based on the identifier indicated by the fourth indication information. Optionally, the fourth indication information may also include the identifier of the first model.
[0112] Optionally, the model reasoning functional entity may also learn the statistical distribution characteristics of the third data subset based on the value of the third parameter of the data in the third data subset and / or the third data subset to optimize the third algorithm. Alternatively, the model reasoning functional entity may also send the value of the third parameter of the third data subset and / or the data in the third data subset to the model management functional entity, so that the model management functional entity learns the statistical distribution characteristics of the third data subset based on the value of the third parameter of the data in the third data subset and / or the third data subset to optimize the third algorithm. For example, the model reasoning functional entity may send second information to the model management functional entity, and the second information includes the value of the third parameter of the third data subset and / or the data in the third data subset. Optionally, if the third algorithm is determined by the model reasoning functional entity, the second information and the fourth indication information may be the same information or different information, which is not limited in the embodiments of the present application.
[0113] Optionally, the first threshold and the second threshold may be preconfigured, for example, corresponding first and second thresholds may be preconfigured for different algorithms; or, the first and second thresholds may be indicated by the fourth indication information. Therefore, the fourth indication information may also optionally include the first and second thresholds.
[0114] S405: The model training functional entity determines the first probability based on the second algorithm.
[0115] The second algorithm may be, for example, a maximum mean discrepancy (MMD) algorithm and a Kolmogorov-Smirnov (KS) test combined with a Bonferroni correction. The model training function may determine the probability that the first data subset exists in the first data set based on the second algorithm.
[0116] For example, if the second algorithm is an MMD algorithm, the model training functional entity can obtain the aforementioned training data set, and calculate the statistical parameters of the first data set and the training data set respectively, such as mean embeddings, and use the Hilbert space to calculate the unbiased estimates of MMD and MMD square based on the statistical parameters of the first data set and the mean embeddings of the training data set to obtain the probability that the first data subset exists in the first data set.
[0117] If the second algorithm is an algorithm combining the KS test and the Bonferroni correction, the model training function entity can obtain the aforementioned training data set and calculate the statistical parameters of the first data set and the training data set respectively, such as the empirical cumulative density functions, and determine the KS test parameters based on the cumulative density functions of the first data set and the training data set, such as the largest difference, to obtain the probability that the first data subset exists in the first data set.
[0118] Optionally, if the first data set includes data carrying the aforementioned suspicious identifier, the model training functional entity may further convert the statistical distribution characteristics of the data carrying the suspicious identifier based on the statistical distribution characteristics of the training data set. For example, the model training functional entity may convert the data carrying the suspicious identifier in the first data set into a first distribution in a latent space through a selective batch normalization (SBN) layer, and the statistical distribution characteristics of the first distribution are the same as the statistical distribution characteristics of the aforementioned training data set. The model training functional entity may determine the probability of the first data subset existing in the first data set after the converted distribution based on the second algorithm.
[0119] The second algorithm used by the model training functional entity to determine the first probability may, for example, be determined by the model training functional entity, or may be indicated by the model management functional entity. If the second algorithm is indicated by the model management functional entity, then before executing S405, S406 may be further executed, and the model management functional entity sends a third indication message to the model training functional entity. Accordingly, the model training functional entity receives the third indication message from the model management functional entity, and the third indication message is used to instruct the model training functional entity to determine the first probability that the first data subset exists in the first data set based on the second algorithm. If the second algorithm is determined by the model training functional entity, then after executing S405, S407 may be further executed, and the model training functional entity sends a third indication message to the model management functional entity. Accordingly, the model management functional entity receives the third indication message from the model training functional entity, and the third indication message is used to inform the model management functional entity that the algorithm used to determine the first probability that the first data subset exists in the first data set is the second algorithm.
[0120] Optionally, the third indication information may directly indicate the second algorithm, or may indicate an identifier of the second algorithm. Optionally, the third indication information may further include an identifier of the first model.
[0121] S408: The model training function entity sends second indication information to the model management function entity, where the second indication information is used to indicate the first probability. Correspondingly, the model management function entity receives the second indication information from the model training function entity.
[0122] Upon receiving the second indication information from the model training functional entity, the model management functional entity may obtain a first probability that the first data subset exists in the first data set based on the second indication information, and execute the aforementioned S1, i.e., determine the first algorithm based on the first probability. For example, if the first probability is greater than a threshold value (e.g., threshold value 2), it indicates that the probability of the first data subset existing in the first data set is high. Since the statistical distribution characteristics of the first data subset and the statistical distribution characteristics of the second data subset are different from the statistical distribution characteristics of the aforementioned training data set, when deleting the second data subset, it is necessary to detect the first data subset in the first data set to prevent the first data subset from being deleted. Therefore, the model management functional entity may determine that the algorithm used to detect drift data is the first algorithm. If the first probability is less than or equal to threshold value 2, it indicates that the probability of the first data subset existing in the first data set is low, i.e., the first data set may only include the aforementioned pseudo data and normal data. Therefore, when deleting the second data subset, it is not necessary to detect the first data subset in the first data set. Therefore, the model management functional entity may determine that the algorithm not used to detect drift data is the first algorithm, which may be, for example, the aforementioned OOD Detection. Optionally, the first algorithm and the aforementioned third algorithm may be the same algorithm or different algorithms. If the first algorithm and the aforementioned third algorithm are the same algorithm, the relevant threshold corresponding to the first algorithm is different from the relevant threshold corresponding to the third algorithm. For example, the third algorithm corresponds to the first threshold and the second threshold, while the first algorithm corresponds to one threshold (for example, the third threshold described below). Optionally, the second indication information and the third indication information in S407 may be the same indication information or different indication information. If the second indication information and the third indication information in S407 are not the same indication information, S407 may be executed simultaneously with S408, or S407 may be executed before S408, or S407 may be executed after S408. The embodiment of the present application does not limit the execution order of S407 and S408.
[0123] Optionally, the model training functional entity may also determine uncertainty quantification information of the first probability. For example, the model training functional entity may determine the uncertainty quantification information of the first probability through parameters such as the Bayesian posterior probability, and send the uncertainty quantification information of the first probability to the model management functional entity. The uncertainty quantification information is used to quantify the accuracy of the first probability. The uncertainty quantification information of the first probability may be sent through the same information as the first probability, or may be sent through different information. The embodiment of the present application does not limit the manner in which the first probability and the uncertainty quantification information of the first probability are sent.
[0124] The model management function entity may determine the first algorithm based on the first probability and the uncertainty quantification information of the first probability. For example, the first algorithm determined by the model management function entity includes the following situations.
[0125] Case 1: The first probability is 80%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 20%, indicating that the accuracy of the first probability is low, that is, the possibility of the existence of the first data subset in the first data set is low, and there is no need to detect the first data subset in the first data set. The model management functional entity can determine that the algorithm that does not detect drift data is the first algorithm.
[0126] Case 2: The first probability is 20%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 20%, indicating that the accuracy of the first probability is low, that is, the possibility of the existence of the first data subset in the first data set is high, and the first data subset in the first data set needs to be detected. The model management functional entity can determine that the algorithm used to detect drift data is the first algorithm.
[0127] Case 3: The first probability is 80%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 80%, indicating that the accuracy of the first probability is relatively high, that is, the possibility of the existence of the first data subset in the first data set is relatively high, and the first data subset in the first data set needs to be detected. The model management functional entity can determine that the algorithm used to detect drift data is the first algorithm.
[0128] Optionally, the algorithm for detecting drift data determined based on scenario 3 and the algorithm for detecting drift data determined based on scenario 2 may be the same algorithm or different algorithms. For example, the first algorithm determined in scenario 2 is an algorithm with a more complex detection process, while the first algorithm determined in scenario 3 is an algorithm with a simpler detection process.
[0129] Case 4: The first probability is 20%, and the uncertainty quantification information corresponding to the first probability indicates that the accuracy of the first probability is 80%, indicating that the accuracy of the first probability is relatively high, that is, the possibility of the existence of the first data subset in the first data set is low, and there is no need to detect the first data subset in the first data set. The model management functional entity can determine that the algorithm that does not detect drift data is the first algorithm.
[0130] Optionally, the algorithm for not detecting drift data determined based on case 4 may be the same algorithm as or different from the algorithm for not detecting drift data determined based on case 1. For example, the first algorithm determined in case 1 is an algorithm with a relatively complex detection process, while the first algorithm determined in case 4 is an algorithm with a relatively simple detection process.
[0131] After executing S1, the model management functional entity may send the first algorithm to the model training functional entity, and therefore optionally, may also execute S409 to S412.
[0132] S409: The model management function entity sends first indication information to the model training function entity. Correspondingly, the model training function entity receives the first indication information from the model management function entity.
[0133] The first indication information is used to indicate the first algorithm. For example, the first indication information may directly indicate the first algorithm or may indicate an identifier of the first algorithm. Optionally, the first indication information may further include an identifier of the first model.
[0134] S410: The model training functional entity deletes a second data subset in the first data set based on the first algorithm to obtain a second data set.
[0135] Upon receiving the first indication information, the model training functional entity may cleanse the second data subset in the first data set based on the first algorithm. Optionally, the model training functional entity may determine the value of the first parameter corresponding to each data in the first data set based on the first algorithm, and delete the data whose value of the first parameter is greater than or equal to the third threshold from the first data set to obtain a second data set, wherein the set formed by the data whose value of the first parameter is greater than or equal to the third threshold is the second data subset. The first parameter is used to indicate the degree of influence of the data on the accuracy of the first model inference. The first parameter may, for example, include one or more of the following: a change in the loss value, which is used to indicate the degree of influence of the data on the loss function (loss) of the first model. The change in the loss value is, for example, the change in the loss value of the loss function of the first model caused by the data (for example, the data in the second data subset) relative to the loss value of the loss function when the first model is trained; a second probability, which is used to indicate the probability that the data is abnormal data; or a degree of difference in statistical distribution characteristics, which is used to indicate the degree of difference between the statistical distribution characteristics of the second data subset and the statistical distribution characteristics of the training data set.
[0136] Among them, if the change in the loss value corresponding to the data is greater than or equal to the third threshold, or the probability that the data is abnormal data is greater than or equal to the third threshold, or the degree of difference between the statistical distribution characteristics of the data and the statistical distribution characteristics of the training data set is greater than or equal to the third threshold, it indicates that the statistical distribution characteristics of the data are significantly different from the statistical distribution characteristics of the training data set, and therefore the data can be considered to be pseudo data, and therefore the data can be deleted from the first data set, and the set formed by the data deleted from the first data set is the second data subset.
[0137] As previously mentioned, the first data set may contain data carrying a suspicious identifier. Therefore, when the model training functional entity deletes the second data subset from the first data set based on the first algorithm, it can also determine whether there is data carrying the suspicious identifier in the first data set. If there is data carrying the suspicious identifier in the first data set, the model training functional entity can determine the value of the first parameter corresponding to the data carrying the suspicious identifier in the first data set based on the first algorithm, and delete the data whose value of the first parameter is greater than or equal to the third threshold from the first data set to obtain the second data set. In this way, for data in the first data set that does not carry a suspicious identifier, the value of the corresponding first parameter can be omitted, which can reduce the computational burden and improve the efficiency of model updates.
[0138] Optionally, the model training functional entity may also learn the statistical distribution characteristics of the second data subset based on the value of the first parameter of the second data subset and / or the data in the second data subset to optimize the first algorithm. Alternatively, the model training functional entity may also send the value of the first parameter of the second data subset and / or the data in the second data subset to the model management functional entity, so that the model management functional entity learns the statistical distribution characteristics of the second data subset based on the value of the first parameter of the second data subset and / or the data in the second data subset to optimize the first algorithm. For example, the model training functional entity may send first information to the model management functional entity, and the first information includes the value of the first parameter of the second data subset and / or the data in the second data subset.
[0139] Optionally, the third threshold may be pre-configured; or, the third threshold may be indicated by the first indication information. Therefore, the first indication information may also optionally include the third threshold.
[0140] S411: The model training functional entity updates the first model based on the second data set.
[0141] After obtaining the second data set, the model training functional entity can update the first model based on the second data set, and end the updating of the first model when the prediction accuracy of the updated first model reaches the aforementioned threshold 1, or the update time of the first model reaches the preset update time, thereby obtaining the updated first model.
[0142] Optionally, the aforementioned sending of the first information by the model training functional entity to the model management functional entity may be performed before or after executing S411. If the sending of the first information by the model training functional entity to the model management functional entity is performed after executing S411, the first information may further include one or more of the following: the value of a second parameter, the second parameter being used to indicate the accuracy of reasoning of the updated first model; the identifier of the first model; the version number of the updated first model; the update time of the first model; the amount of data included in the third data set used to update the first model; or the duration required for reasoning of the updated first model.
[0143] If the model training functional entity sends the first information to the model management functional entity before executing S411, then after executing S411, the model training functional entity may also send fourth information to the model management functional entity, and the fourth information includes one or more of the following: the value of the second parameter, the second parameter is used to indicate the accuracy of the updated first model reasoning; the identifier of the first model; the version number of the updated first model; the update time of the first model; the amount of data included in the third data set used to update the first model; or the time required for the updated first model reasoning.
[0144] S412: The model training functional entity sends the third information to the model reasoning functional entity. Correspondingly, the model reasoning functional entity receives the third information from the model training functional entity.
[0145] The third information includes one or more of the following: an identifier of the first model, a storage address of the updated first model, an update time of the first model, or a version number of the updated first model.
[0146] After receiving the third information, the model reasoning function entity may update the model parameters of the first model according to the third information. For example, the model reasoning function entity may obtain the model parameters of the first model based on the storage address of the updated first model in the third information, and update the model parameters of the locally stored first model based on the identifier of the first model in the third information, such as modifying the model parameters of the locally stored first model and storing the version number of the updated first model.
[0147] The above-mentioned first information, second information, third information and fourth information may also include other information. For example, the first information, third information and fourth information may also include information such as the robustness of the updated first model. The embodiment of the present application does not limit the information included in the first information, second information, third information and fourth information.
[0148] If in the aforementioned S1, it is determined based on the first probability that the first algorithm is implemented by the model training functional entity, then before executing S1, the following S501 to S508 may also be executed, please refer to Figure 5.
[0149] S501: The model inference function entity sends a first data set to the model training function entity. Correspondingly, the model training function entity receives the first data set from the model inference function entity and deletes a third data subset in the third data set based on a third algorithm.
[0150] S502: The model management function entity deletes a third data subset in the third data set based on a third algorithm.
[0151] S503: The model management function entity sends fourth indication information to the model reasoning function entity. Correspondingly, the model reasoning function entity receives the fourth indication information from the model management function entity.
[0152] S504: The model reasoning function entity sends fourth indication information to the model management function entity. Correspondingly, the model management function entity receives the fourth indication information from the model reasoning function entity.
[0153] S505: The model training functional entity determines the first probability based on the second algorithm.
[0154] S506: The model management function entity sends third indication information to the model training function entity. Correspondingly, the model training function entity receives the third indication information from the model management function entity.
[0155] S507: The model training function entity sends third indication information to the model management function entity. Correspondingly, the model management function entity receives the third indication information from the model training function entity.
[0156] For the description of S501 to S507 above, reference may be made to the description of the corresponding steps S401 to S407 above, which will not be repeated here.
[0157] S508: The model training function entity sends second indication information to the model management function entity, where the second indication information is used to indicate the first probability. Correspondingly, the model management function entity receives the second indication information from the model training function entity.
[0158] In an embodiment of the present application, since the first algorithm is determined to be implemented by the model training functional entity based on the first probability, the model training functional entity sends the second indication information to the model management functional entity, which enables the model management functional entity to optimize the aforementioned third algorithm based on the first probability.
[0159] After determining the first probability based on the second algorithm, the model training functional entity may execute the aforementioned S1, i.e., determine the first algorithm based on the first probability. For the relevant description of the model training functional entity determining the second algorithm based on the first probability, reference may be made to the relevant description of the model management functional entity determining the second algorithm based on the first probability in S408, and for the relevant description of the relationship between the second indication information and the third indication information in S507, reference may be made to the relevant description of the relationship between the second indication information in S408 and the third indication information in S407, which will not be repeated here.
[0160] Optionally, the model training functional entity may further determine uncertainty quantification information of the first probability. For example, the model training functional entity may determine the uncertainty quantification information of the first probability using parameters such as the Bayesian posterior probability, and the uncertainty quantification information is used to quantify the accuracy of the first probability. The uncertainty quantification information of the first probability may be sent through the same information as the first probability, or may be sent through different information. The embodiments of the present application do not limit the manner in which the first probability and the uncertainty quantification information of the first probability are sent.
[0161] The relevant description of the first algorithm determined by the model training functional entity based on the first probability and the uncertainty quantification information of the first probability can refer to the relevant description of the first algorithm determined by the model management functional entity based on the first probability and the uncertainty quantification information of the first probability in S408, such as cases 1 to 4 in S408.
[0162] S509: The model training function entity sends first indication information to the model management function entity. Correspondingly, the model management function entity receives the first indication information from the model training function entity.
[0163] For the description of the first indication information, please refer to the description of the first indication information in S409 above, and will not be repeated here. In the embodiment of the present application, the first indication information is used to inform the model management function entity that the algorithm used to delete the second data subset in the first data set is the first algorithm. Optionally, the first indication information in the embodiment of the present application and the second indication information and / or the third indication information above may be the same indication information or different indication information. For example, the first indication information, the second indication information, and the third indication information are the same indication information; or the first indication information and the second indication information are the same indication information and the third indication information are different indication information; or the first indication information and the third indication information are the same indication information and the second indication information are different indication information; or the first indication information and the third indication information are different indication information. It can be understood that if the first indication information and the second indication information are the same indication information and the third indication information are different indication information; or the first indication information and the third indication information are the same indication information and the second indication information are different indication information, it means that the second indication information and the third indication information are different indication information. If the first indication information is the same as the second indication information and the third indication information, then the second indication information and the third indication information may be the same as the indication information, or may not be the same as the indication information.
[0164] S510: The model training functional entity deletes a second data subset in the first data set based on the first algorithm to obtain a second data set.
[0165] S511: The model training functional entity updates the first model based on the second data set.
[0166] S512: The model training functional entity sends the third information to the model reasoning functional entity. Correspondingly, the model reasoning functional entity receives the third information from the model training functional entity.
[0167] The related descriptions of S510 to S512 mentioned above can refer to the related descriptions of S410 to S412 mentioned above, which will not be repeated here.
[0168] In the embodiment of the present application, the model management functional entity, the model training functional entity, and the model reasoning functional entity are deployed on different devices as an example, such as the aforementioned scenario 2. When at least two functional entities among the model management functional entity, the model training functional entity, and the model reasoning functional entity are deployed on the same device (such as the aforementioned scenarios 1, 3 to 6), the at least two functional entities can be regarded as one functional entity, or the at least two functional entities can obtain relevant processing results from each other. Taking scenario 3 as an example, if the model training functional entity and the model reasoning functional entity are deployed on the same network element, the model training functional entity can obtain the first data set from the model reasoning functional entity without sending it through signaling.
[0169] In the above technical solution, considering that there may be drift data in the collected data set, before the model training functional entity cleans the first data set, it can use the probability of the presence of a first data subset in the first data set that is different from the statistical distribution characteristics of the training data set of the training model to determine the algorithm for cleaning the first data set. This can improve the accuracy of cleaning the second data and further improve the robustness of the model.
[0170] Figure 6 shows a schematic diagram of the structure of a communication device provided in an embodiment of the present application. The communication device 600 may be the model management functional entity described in the embodiment shown in Figure 4 or Figure 5, used to implement the method corresponding to the model management functional entity in the above method embodiment; or the communication device 600 may be the model training functional entity described in the embodiment shown in Figure 4 or Figure 5, used to implement the method corresponding to the model training functional entity in the above method embodiment; or the communication device 600 may be the model reasoning functional entity described in the embodiment shown in Figure 4 or Figure 5, used to implement the method corresponding to the model reasoning functional entity in the above method embodiment.
[0171] The communication device 600 includes at least one processor 601. Processor 601 can be used for internal processing of the device, implementing certain control processing functions. Optionally, processor 601 includes instructions. Optionally, processor 601 can store data. Optionally, different processors can be independent devices, located in different physical locations, or on different integrated circuits. Optionally, different processors can be integrated into one or more processors, for example, on one or more integrated circuits.
[0172] Optionally, the communication device 600 includes one or more memories 603 for storing instructions. Optionally, data may also be stored in the memories 603. The processor and memory may be provided separately or integrated together.
[0173] Optionally, the communication device 600 includes a communication line 602 and at least one communication interface 604. Since the memory 603, the communication line 602 and the communication interface 604 are all optional, they are indicated by dotted lines in FIG6 .
[0174] Optionally, the communication device 600 may further include a transceiver and / or an antenna. The transceiver may be used to send information to or receive information from other devices. The transceiver may be referred to as a transceiver, a transceiver circuit, an input / output interface, etc., and is used to implement the transceiver function of the communication device 600 via an antenna. Optionally, the transceiver includes a transmitter and a receiver. For example, the transmitter may be used to generate a radio frequency signal from a baseband signal, and the receiver may be used to convert the radio frequency signal into a baseband signal.
[0175] The processor 601 may include a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.
[0176] Communication link 602 may include a pathway for transmitting information between the aforementioned components.
[0177] The communication interface 604 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), wired access network, etc.
[0178] The memory 603 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 603 may exist independently and be connected to the processor 601 via the communication line 602. Alternatively, the memory 603 may be integrated with the processor 601.
[0179] The memory 603 is used to store computer-executable instructions for executing the solution of the present application, and the execution is controlled by the processor 601. The processor 601 is used to execute the computer-executable instructions stored in the memory 603, thereby implementing the steps performed by the model management functional entity, the model training functional entity, or the model inference functional entity in the embodiments shown in Figures 4 or 5.
[0180] Optionally, the computer-executable instructions in the embodiments of the present application may also be referred to as application code, which is not specifically limited in the embodiments of the present application.
[0181] In a specific implementation, as an embodiment, the processor 601 may include one or more CPUs, such as CPU0 and CPU1 in FIG6 .
[0182] In a specific implementation, as an embodiment, the communication device 600 may include multiple processors, such as processor 601 and processor 605 in FIG6 . Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0183] When the device shown in Figure 6 is a chip, for example, a chip of a model management functional entity, a model training functional entity, or a model reasoning functional entity, the chip includes a processor 601 (and may also include a processor 605), a communication line 602, and a communication interface 604. Optionally, it may include a memory 603. Specifically, the communication interface 604 may be an input interface, a pin, or a circuit. The memory 603 may be a register, a cache, or the like. The processor 601 and the processor 605 may be a general-purpose CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the semantic communication method of any of the above embodiments.
[0184] The embodiment of the present application can divide the functional modules of the device according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. For example, in the case of dividing each functional module according to each function, Figure 7 shows a schematic diagram of a device. The device 700 can be the model management functional entity, model training functional entity or model reasoning functional entity involved in the above-mentioned various method embodiments, or a chip in the model management functional entity, model training functional entity or model reasoning functional entity. The device 700 includes a sending unit 701, a processing unit 702 and a receiving unit 703.
[0185] It should be understood that the device 700 can be used to implement the steps performed by the model management functional entity, the model training functional entity or the model reasoning functional entity in the model management method of the embodiment of the present application. The relevant features can refer to any of the embodiments shown in Figures 4 or 5 above, and will not be repeated here.
[0186] Optionally, the functions / implementation processes of the sending unit 701, the receiving unit 703, and the processing unit 702 in FIG7 may be implemented by the processor 601 in FIG6 calling computer-executable instructions stored in the memory 603. Alternatively, the functions / implementation processes of the processing unit 702 in FIG7 may be implemented by the processor 601 in FIG6 calling computer-executable instructions stored in the memory 603, and the functions / implementation processes of the sending unit 701 and the receiving unit 703 in FIG7 may be implemented by the communication interface 604 in FIG6.
[0187] Optionally, when the device 700 is a chip or a circuit, the functions / implementation processes of the sending unit 701 and the receiving unit 703 can also be implemented through pins or circuits.
[0188] The present application also provides a computer-readable storage medium, which stores a computer program or instruction. When the computer program or instruction is executed, the method performed by the model management functional entity, the model training functional entity or the model reasoning functional entity in the aforementioned method embodiment is implemented. In this way, the functions described in the above embodiments can be implemented in the form of software functional units and sold or used as independent products. Based on this understanding, the technical solution of the present application can be essentially or in other words, the part that contributes or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk.
[0189] The present application also provides a computer program product, which includes: computer program code, which, when running on a computer, enables the computer to execute the method performed by the model management functional entity, the model training functional entity or the model reasoning functional entity in any of the aforementioned method embodiments.
[0190] An embodiment of the present application also provides a processing device, including a processor and an interface; the processor is used to execute the method performed by the model management functional entity, model training functional entity or model reasoning functional entity involved in any of the above method embodiments.
[0191] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0192] The various illustrative logic units and circuits described in the embodiments of the present application can be implemented or operated by a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.
[0193] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, software units executed by a processor, or a combination of the two. The software units can be stored in RAM, flash memory, ROM, erasable programmable read-only memory (EPROM), EEPROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Alternatively, the storage medium can also be integrated into the processor. The processor and storage medium can be provided in an ASIC, which can be provided in a network device. Alternatively, the processor and storage medium can also be provided in different components in the network device.
[0194] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0195] The contents of the various embodiments of this application can refer to each other. If there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0196] It is understood that in the embodiments of the present application, the network device may perform some or all of the steps in the embodiments of the present application. These steps or operations are merely examples. In the embodiments of the present application, other operations or variations of various operations may also be performed. In addition, the various steps may be performed in a different order than those presented in the embodiments of the present application, and it is possible that not all operations in the embodiments of the present application need to be performed.
Claims
1. A model management method, characterized in that: The method comprises: Determine a first algorithm according to a first probability; wherein, The first probability is a probability that a first data subset exists in the first data set, a statistical distribution feature of the first data subset is different from a statistical distribution feature of a training data set used to train the first model, and the first data subset is used to update the first model; The first algorithm is used to delete a second data subset in the first data set to obtain a second data set, the statistical distribution characteristics of the second data subset are different from the statistical distribution characteristics of the training data set, the second data subset is not used to update the first model, and the second data set is used to update the first model.
2. The method according to claim 1, characterized in that The method further comprises: Send first indication information, where the first indication information is used to indicate the first algorithm.
3. The method according to claim 1 or 2, characterized in that The first probability is determined by the model training functional entity based on the second algorithm.
4. The method according to claim 3, characterized in that The method further comprises: Sending second indication information to the model management function entity, where the second indication information is used to indicate the first probability.
5. The method according to claim 3 or 4, characterized in that The method further comprises: Send third indication information, where the third indication information is used to indicate the second algorithm.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: The first data set is received from a model reasoning functional entity.
7. The method according to any one of claims 1 to 3 and 5, characterized in that: The method further comprises: Send fourth indication information to the model reasoning functional entity, or receive fourth indication information from the model reasoning functional entity, the fourth indication information is used to indicate a third algorithm, the third algorithm is used to delete a third data subset in a third data set to obtain the first data set, the third data set is collected by the model reasoning functional entity, the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set, and the third data subset is not used to update the first model.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: The first algorithm is optimized based on the value of a first parameter of the second data subset and / or the data in the second data subset, wherein the value of the first parameter is determined based on the first algorithm, and the first parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model.
9. The method according to claim 8, characterized in that The method further comprises: Sending first information to a model management function entity, where the first information includes the second data subset and / or a value of a first parameter of data in the second data subset.
10. The method according to claim 8 or 9, characterized in that The first parameter includes one or more of the following: A change in the loss value, where the change in the loss value is used to indicate the degree of influence of the data on the loss function of the first model; A second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the second data subset and the statistical distribution characteristics of the data in the training data set.
11. The method according to claim 9 or 10, characterized in that The first information also includes one or more of the following: a value of a second parameter, the second parameter being used to indicate the accuracy of the updated first model inference; the identification of the first model; The version number of the first model after the update; The update time of the first model; the amount of data included in the third data set used to update the first model; or, The time required for inference of the updated first model.
12. The method according to any one of claims 1 to 3, 5 and 7, characterized in that: The method further comprises: A third algorithm is optimized based on a third data subset and / or a value of a third parameter of the data in the third data subset, the value of the third parameter is determined based on the third algorithm, the third algorithm is used to delete the third data subset, the third parameter is used to indicate the degree of influence of the data on the inference accuracy of the first model, the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set, and the third data subset is not used to update the first model.
13. The method according to claim 12, characterized in that The method further comprises: Receive second information from the model reasoning function entity, where the second information includes the third data subset and / or a value of a third parameter of the data in the third data subset.
14. The method according to claim 12 or 13, characterized in that The third parameter includes one or more of the following: A change in the loss value, where the change in the loss value is used to indicate the degree of influence of the data on the loss function of the first model; A second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set.
15. The method according to any one of claims 1 to 6 and 8 to 11, characterized in that: The method further comprises: Sending third information to the model reasoning function entity, wherein the third information includes one or more of the following: the identification of the first model; The version number of the first model after the update; The update time of the first model; or, The storage address of the updated first model.
16. A model management method, characterized in that: The method comprises: Deleting a third data subset in the third data set based on a third algorithm to obtain a first data set, wherein the third data set is collected by the model reasoning functional entity, the statistical distribution characteristics of the third data subset are different from the statistical distribution characteristics of the training data set used to train the first model, and the third data subset is not used to update the first model; The first data set is sent to a model training functional entity.
17. The method according to claim 16, characterized in that The method further comprises: Send fourth indication information to the model management function entity, or receive fourth indication information from the model management function entity, where the fourth indication information is used to indicate the third algorithm.
18. The method according to claim 16 or 17, characterized in that Deleting a third data subset in the third data set based on a third algorithm includes: Determining a value of a third parameter of the data in the third data set based on the third algorithm, the third parameter being used to indicate the degree of influence of the data on the inference accuracy of the first model; The third data subset is deleted, wherein the value of the third parameter of the data in the third data subset is greater than or equal to the first threshold.
19. The method according to any one of claims 16 to 18, characterized in that: The method further comprises: The third algorithm is optimized based on the third data subset and / or a value of a third parameter of the third data subset.
20. The method according to claim 18 or 19, characterized in that The method further comprises: Second information is sent to the model management function entity, where the second information includes the third data subset and / or a value of a third parameter of the data in the third data subset.
21. The method according to any one of claims 18 to 20, characterized in that: The third parameter includes one or more of the following: A change in the loss value, where the change in the loss value is used to indicate the degree of influence of the data on the loss function of the first model; A second probability, where the second probability is used to indicate the probability that the data is abnormal data; or, The degree of difference in statistical distribution characteristics is used to indicate the degree of difference between the statistical distribution characteristics of the data in the third data subset and the statistical distribution characteristics of the data in the training data set.
22. The method according to any one of claims 16 to 21, characterized in that: The method further comprises: Receive third information from the model training functional entity, where the third information includes one or more of the following: the identification of the first model; The version number of the first model after the update; The update time of the first model; or, The storage address of the updated first model.
23. The method of claim 22, wherein: The method further comprises: Model parameters of the first model are updated based on the third information.
24. A model management system, characterized in that: It comprises a model management functional entity and a model training functional entity for executing the method as described in any one of claims 1 to 15, and a model reasoning functional entity for executing the method as described in any one of claims 16 to 23.
25. A communication device, characterized in that: The method comprises a processor and a memory, wherein the memory is coupled to the processor, and the processor is used to call computer instructions in the memory to execute the method according to any one of claims 1 to 15, or to execute the method according to any one of claims 16 to 23.
26. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called by the computer, the computer-executable instructions are used to execute the method according to any one of claims 1 to 15, or to execute the method according to any one of claims 16 to 23.
27. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 15, or the computer is caused to execute the method according to any one of claims 16 to 23.
28. A computer program, characterized in that The method comprises a program code, and when the computer runs the program code, the program code executes the method according to any one of claims 1 to 15, or the program code executes the method according to any one of claims 16 to 23.
29. A chip, characterized in that: The chip is coupled to the memory and is used to read and execute program instructions stored in the memory to implement the method according to any one of claims 1 to 15, or to implement the method according to any one of claims 16 to 23.
Citation Information
Patent Citations
Promotion data processing method, model training method, system and storage medium
CN113570398A
Model training method, related device and storage medium
CN115392405A
Rail transit monitoring data cleaning method and system based on Bayesian reasoning
CN115700494A
Methods and apparatus to estimate cardinality through ordered statistics
US20230120709A1