Communication of information for optimizing data samples for model training

By querying and determining the validity information of data samples in AI/ML model training, and optimizing data samples, the problems of high model training time and resource cost in the prior art are solved, and more efficient model training is achieved.

CN120153379APending Publication Date: 2025-06-13ALCATEL LUCENT SHANGHAI BELL CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101628.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In artificial intelligence (AI)/machine learning (ML) model training, it is difficult for the existing technology to effectively optimize data samples, resulting in increased model training time and resource costs, and the cost of data acquisition and computing resources is huge.

Method used

By communicating between the first device and the second device, validity information of the data samples for model training is queried and determined, indicating the usefulness of the data samples or data features for model training performance, and sending an effectiveness report to optimize the data samples.

Benefits of technology

By optimizing data samples, the amount of data that needs to be collected, sent, stored and preprocessed is reduced, the performance and efficiency of model training is improved, and resource consumption is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153379A_ABST
    Figure CN120153379A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to equipment, a method, a device and computer readable storage medium ML abstract behavior management. A method includes receiving, at a first device, a first request from a second device, the first request for querying validity information for a plurality of data samples for model training at the first device; determining validity information for at least one of the plurality of data samples for model training, the validity information indicating usefulness of the at least one data sample or the at least one data feature for performance of the model training; and transmitting a first report indicating the validity information to the second device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various exemplary embodiments of the present disclosure generally relate to the field of telecommunications, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for communicating information for optimizing data samples for model training. Background Art

[0002] Basically, for an artificial intelligence (AI) / machine learning (ML) application workflow, data sample collection is an original operation. In terms of model training accuracy and training time, data sample quality and data sample volume may have a significant impact on the execution of AI / ML model training.

[0003] Generally speaking, in order to train an ML model, high-quality and a large number of training data samples are necessary. That is to say, in order to obtain a model with high accuracy, the general practice is to collect as many data samples as possible, and then feed the collected data samples for preprocessing and model training. However, this increases the time and resource costs for model training. Therefore, it is necessary to classify and optimize data samples so that the performance of model training can be improved accordingly. Summary of the Invention

[0004] In a first aspect of the present disclosure, a first device is provided. The first device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the first device to at least perform: receiving a first request from a second device, the first request being for querying validity information of a plurality of data samples for model training at the first device; determining validity information of at least one data sample among the plurality of data samples for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training; and sending a first report indicating the validity information to the second device.

[0005] In a second aspect of the present disclosure, a second device is provided. The second device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the second device to at least perform: sending a first request to the first device, the first request being for querying validity information of a plurality of data samples for model training at the first device; and receiving a first report from the first device, the first report indicating validity information of at least one data sample among the plurality of data samples used for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training, the validity information being determined by the first device during model training by using the plurality of data samples.

[0006] In a third aspect of the present disclosure, a first device is provided. The first device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the first device to at least perform: receiving a second request from a second device, the second request being for information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity value indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of model training; determining the information at least partially based on the second request; and sending a second report to the second device, the second report indicating the information associated with the first set of a plurality of data samples.

[0007] In a fourth aspect of the present disclosure, a second device is provided. The second device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the second device to at least perform: sending a second request to a first device, the second request being for information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity value indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of model training; and receiving a second report from the first device, the second report indicating the information associated with the first set of a plurality of data samples.

[0008] In a fifth aspect of the present disclosure, a method is provided. The method includes: receiving, at a first device, a first request from a second device, the first request being for querying validity information of a plurality of data samples for model training at the first device; determining validity information of at least one data sample among the plurality of data samples for model training, the validity information indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of model training; and sending a first report indicating the validity information to the second device.

[0009] In a sixth aspect of the present disclosure, a method is provided. The method includes: sending, at a second device, a first request to a first device, the first request querying validity information of a plurality of data samples for model training at the first device; and receiving, from the second device, a first report indicating validity information of at least one data sample among the plurality of data samples used for model training, the validity information indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of model training, the validity information being determined by the first device during model training by using the plurality of data samples.

[0010] In a seventh aspect of the present disclosure, a method is provided. The method includes: receiving, at a first device, a second request from a second device, the second request being for information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity value indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of model training, and determining the information at least partially based on the second request; and sending, to the second device, a second report indicating the information associated with the first set of a plurality of data samples.

[0011] In an eighth aspect of the present disclosure, a method is provided. The method includes: sending, at a second device, a second request to a first device, the second request being for information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity value indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of model training; and receiving, from the first device, a second report indicating the information associated with the first set of a plurality of data samples.

[0012] In a ninth aspect of the present disclosure, a first device is provided. The first device includes: means for receiving a first request from a second device, the first request being for querying validity information of a plurality of data samples for model training at the first device; means for determining validity information of at least one data sample among the plurality of data samples for model training, the validity information indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of model training; and means for sending a first report indicating the validity information to the second device.

[0013] In a tenth aspect of the present disclosure, a second device is provided. The second device includes: means for sending a first request to a first device, the first request being for querying validity information of a plurality of data samples for model training at the first device; and means for receiving a first report from the second device, the first report indicating validity information of at least one data sample among the plurality of data samples used for model training, the validity information indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of model training, the validity information being determined by the first device during model training by using the plurality of data samples.

[0014] In an eleventh aspect of the present disclosure, a first device is provided. The first device includes: a component for receiving a second request from a second device, the second request being directed to information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training; a component for determining the information at least partially based on the second request; and a component for sending a second report to a second device, the second report indicating the information associated with the first set of a plurality of data samples.

[0015] In a twelfth aspect of the present disclosure, a second device is provided. The second device includes: a component for sending a second request to a first device, the second request being directed to information associated with a first set of a plurality of data samples for model training, a first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training; and a component for receiving a second report from the first device, the second report indicating the information associated with the first set of a plurality of data samples.

[0016] In a thirteenth aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions stored thereon for causing a device to at least execute the method according to the fifth aspect.

[0017] In a fourteenth aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions stored thereon for causing a device to at least execute the method according to the sixth aspect.

[0018] In a fifteenth aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions stored thereon for causing a device to at least execute the method according to the seventh aspect.

[0019] In a sixteenth aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium includes instructions stored thereon for causing a device to at least execute the method according to the eighth aspect.

[0020] It should be understood that the summary section is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Some example embodiments will now be described with reference to the accompanying drawings, in which:

[0022] Figures 1A to 1C An example communication environment in which example embodiments of the present disclosure may be implemented;

[0023] Figure 2A and Figure 2B An example signaling diagram showing a process for optimizing data samples for model training according to some example embodiments of the present disclosure;

[0024] Figure 3 Another example signaling diagram showing a process for optimizing data samples for model training according to some example embodiments of the present disclosure;

[0025] Figure 4 Yet another example signaling diagram showing a process for optimizing data samples for model training according to some example embodiments of the present disclosure;

[0026] Figure 5 Yet another example signaling diagram showing a process for optimizing data samples for model training according to some example embodiments of the present disclosure;

[0027] Figure 6 Yet another example signaling diagram showing a process for optimizing data samples for model training according to some example embodiments of the present disclosure;

[0028] Figure 7 A flowchart of a method implemented at a first device according to some example embodiments of the present disclosure;

[0029] Figure 8 A flowchart of a method implemented at a second device according to some example embodiments of the present disclosure;

[0030] Figure 9 A flowchart of a method implemented at a first device according to some example embodiments of the present disclosure;

[0031] Figure 10 A flowchart of a method implemented at a second device according to some example embodiments of the present disclosure;

[0032] Figure 11 A simplified block diagram of a device suitable for implementing example embodiments of the present disclosure; and

[0033] Figure 12 A block diagram of an example computer-readable medium according to some example embodiments of the present disclosure.

[0034] Throughout the drawings, like or similar reference numerals denote like or similar elements. Detailed Description

[0035] The principles of the present disclosure will now be described with reference to some example implementations. It should be understood that the description of these embodiments is for illustrative purposes only and helps those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. The embodiments described herein can be implemented in various ways other than those described below.

[0036] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0037] References in this disclosure to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but not necessarily every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to combine such feature, structure, or characteristic with other embodiments, whether or not explicitly described.

[0038] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0039] As used herein, "at least one of the following: <list of two or more elements>" and "at least one of <list of two or more elements>" and similar phrases, where the list of two or more elements is joined by "and" or "or", mean at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.

[0040] As used herein, unless explicitly stated otherwise, performing the step "in response to A" does not indicate that the step is performed immediately after A occurs, and may include one or more intermediate steps.

[0041] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It will be further understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including", when used herein, specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0042] The term "circuitry" as used in this application can refer to one or more or all of the following: (a) only hardware circuit implementations (such as implementations in only analog and / or digital circuits) and (b) combinations of hardware circuits and software, such as, where applicable: (i) combinations of (one or more) analog and / or digital hardware circuits and software / firmware and (ii) any part of (one or more) hardware processors (including (one or more) digital signal processors) with software and (one or more) memories, which work together to enable a device (such as a mobile phone or a server) to perform various functions) and (c) (one or more) hardware circuits and / or (one or more) processors, such as (one or more) microprocessors or a part of (one or more) microprocessors, which require software (such as firmware) to operate, but may not have the software when not required to operate.

[0043] The definition of circuitry applies to all uses of the term in this application, including any claims. As a further example, as used in this application, the term circuitry also encompasses implementations of only hardware circuits or processors (or one or more processors) or a part of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also encompasses, for example (and if applicable to a particular claim element), a baseband integrated circuit or a processor integrated circuit for a mobile device or a server, a cellular network device or a similar integrated circuit in other computing or network devices.

[0044] As used herein, the term "communication network" refers to a network that follows any suitable communication standard, such as New Radio (NR), Long-Term Evolution (LTE), LTE-Advanced (LTE-A), Wideband Code Division Multiple Access (WCDMA), High-Speed Packet Access (HSPA), NarrowBand Internet of Things (NB-IoT), etc. In addition, the communication between the terminal device and the network device in the communication network can be performed according to any suitable generation communication protocol, including but not limited to the first generation (1G), second generation (2G), 2.5G, 2.75G, third generation (3G), fourth generation (4G), 4.5G, fifth generation (5G) communication protocol and / or any other protocol currently known or to be developed in the future. Embodiments of the present disclosure can be applied to various communication systems. Given the rapid development of communication, of course, there will be future types of communication technologies and systems that can implement the present disclosure. The scope of the present disclosure should not be limited to the above systems.

[0045] It should be noted that the title of any section / subsection provided herein is not intended to be restrictive. Embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments disclosed in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different section / subsections in any way.

[0046] As briefly mentioned above, in terms of model training accuracy and training time, the quality and quantity of data samples may have a significant impact on the execution of AI / ML model training.

[0047] In addition, generally speaking, in order to train an ML model, high-quality and a large number of training data samples are necessary. That is, in order to obtain a model with high accuracy, the general practice is to collect as many data samples as possible, and then feed the collected data samples for preprocessing and model training. However, model training is usually performed without considering the possible different contributions of different parts of the input data samples to the accuracy of the trained model. As a result, the costs of both data acquisition and computing resources are huge, because even if the data samples are unnecessary, the data samples still need to be calculated by the ML model.

[0048] In some example embodiments, the ML model training function may optimize training data samples during or before model training. For example, by simply resampling the training data to use only a portion of the collected training data. However, such embodiments are not suitable for mobile networks because in mobile networks, the amount of data samples is huge, resources are constrained, but the accuracy of the model is strict. Without insight into the validity of the training data samples, AI / ML model training must utilize the complete data set, which will consume more resources, such as complete data collection, data transmission or storage, data preparation for model training, etc. As a result, there is a significant increase in the network load of data collection at all network elements (including radio nodes, core network functions, etc.) and additional time for transmission (along with additional power consumption) for training the ML model. In addition, in the case where training data is collected from terminal devices (i.e., UEs) or IoT devices, more battery and computing power will be required to be allocated for redundant data collection. Another problem is that since all data samples are treated equally importantly, this may affect the efficiency and accuracy of the model training process and results. This is especially useful for conventional model training / retraining use cases.

[0049] In some example embodiments, it is proposed that based on some internal understanding of the data samples (such as lightweight gradient boosting machine (GMB)), the model can use only a portion of the available data samples. In a specific example embodiment, the number of training data samples can be reduced using the feature of gradient-based one-sided sampling (GOSS). As a specific example embodiment, it is noted that the gradient for each data instance in the gradient boosting decision tree (GBDT) provides useful information for data sampling because the instances associated with small gradients have small training errors, indicating that the samples have been well trained.

[0050] Although those data samples with small gradients may be discarded operationally. However, this changes the data distribution and may harm the accuracy of the learned model. Therefore, GOSS can be used to remedy this defect. Thus, lightweight GBM can be much faster than other gradient-based boosting algorithms without negative impacts on model performance. In view of this, algorithms such as lightweight GBM can be used for wireless communication.

[0051] In short, by selecting high-quality training data samples (which contribute more to model training), the number of training data samples can be reduced, and the expected training accuracy can be achieved in a shorter training time. At the same time, data collection, data storage, and data preprocessing can be minimized, thus significantly improving training efficiency. In view of this, it is necessary to label the training data samples with tags, where the tags can indicate whether the data samples are useful or useless for model training, or indicate the contribution value (for example, a weight between 0 and 1). In addition, such tag information is expected to be known to the device that consumes the tag information. Specifically, in a wireless communication network, it is desirable to report the contribution information of different training data samples based on insights into how different parts of the data contribute differently to model accuracy.

[0052] Some example embodiments of the present disclosure provide a solution for optimizing data samples. Specifically, a first device may determine validity information of at least one data sample among a plurality of data samples for model training, where the validity information indicates the usefulness of the at least one data sample or at least one data feature for the performance of model training. In addition, the first device may send a first report indicating the validity information to a second device. In this way, the second device.

[0053] Consuming the validity information can distinguish data samples with higher contribution values from other data samples.

[0054] In addition, as mentioned above, it is generally necessary to collect data samples from all network elements and domains of 5GS, including 5G NR gNB, core network functions, and data collection from terminal devices, or any other device that generates data for analysis. Therefore, it is expected that unnecessary data sample collection can be avoided. In some example embodiments of the present disclosure, validity information can be obtained. However, the validity information is not sufficient to determine how to collect data samples and which data to collect to optimize the number of data samples used for training and the associated computing resources spent on model training.

[0055] Specifically, as mentioned above, the validity information can indicate information such as whether a specific data sample is useful for model training and information about the degree to which the data sample is useful for model training. Although such information can provide valuable feedback to the management service (MnS) training consumer about the effectiveness of the training dataset during the training process, and can even point out the most valuable data samples in the given dataset for which training has been performed, this "one-time", simple classification of data instances is of little significance for improving future (re)training processes without further processing the available data to obtain more easily understandable insights.

[0056] In other words, a single / independent observation regarding whether a certain sample or feature contributes to the model gradient at a given timestamp cannot provide an understanding of whether using such a sample or feature in subsequent training / retraining will contribute to the accuracy of the same model, nor can it provide an understanding of whether it will contribute to the training efficiency of other models related to the same use case. If the data sample collection process is performed without considering the validity information, the network load and energy consumption for data collection at all network elements will increase significantly, the AI / ML model training will be based on a data set including unnecessary data, consuming more resources (such as computing, energy and taking more time), and since all data samples are treated equally importantly, the efficiency and accuracy of the model training process will decrease. Therefore, analyzing the patterns of the most effective training data samples generated will be very helpful for improving data sample collection in order to optimize the quantity and amount of data samples to be used for model training.

[0057] In other words, further analysis of the data related to the importance of data samples during training is needed.

[0058] Some exemplary embodiments of the present disclosure provide another solution for optimizing data samples. Specifically, a first device may determine information associated with a first set of multiple data samples used for model training, where the first set of multiple data samples may have higher validity values. In this way, data sample collection can be focused on high-quality data samples.

[0059] In this way, the number of data samples that need to be collected, sent, stored and preprocessed is reduced, thereby further improving the model training performance.

[0060] It should be clear that the specific terms for data types / attributes are for illustrative purposes only and do not imply any limitation. In some other example embodiments, the terms for data types / attributes may be changed to any suitable terms. The present disclosure is not limited in this regard.

[0061] Example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Example environment and working principle

[0062] Figure 1A An example communication environment 100 is shown in which example embodiments of the present disclosure can be implemented. In the communication environment 100, there is a first device 110 and a second device 120. Optionally, the functions of model training 116 and LM entity 115 may also be included in the communication environment 100.

[0063] In some example embodiments, validity information can be generated and consumed, as Figure 1B shown, Figure 1BAn example communication environment 150 is shown. In this case, the second device 120 can be a management service (MnS) consumer / model training consumer, such as an AI / ML MnS consumer, and the first device 110 can be an MnS producer, such as an AI / ML MnS producer / model training producer.

[0064] In some example embodiments, information (referred to as analysis information for discussion purposes) associated with a first set of multiple data samples having a relatively high validity value can be generated and consumed. In this case, the first device 110 can be an MnS producer and the second device 120 can be an MnS consumer. The first device 110 is a producer for training data analysis, and the second device 120 is a consumer for training data analysis. The first device 110 can be a management data analysis (MDA) producer, and the second device 120 can be an MDA consumer, or the first device 110 can be an AIMLT producer, and the second device 120 can be an AIMLT consumer.

[0065] It should be understood that the above examples of the first device 110 and the second device 120 are for illustrative purposes only and do not imply any limitation. In the present disclosure, for example embodiments related to validity information, the first device 110 can be a device that provides validity information, and the second device 120 can be a device that receives validity information. Additionally, for example embodiments related to analysis information, the first device 110 can be a device that provides analysis information, and the second device 120 can be a device that receives analysis information. That is, the specific device types of the first device 110 and the second device 120 can change according to the specific application scenario and the specific network structure. The present disclosure is not limited in this regard.

[0066] Communication in the communication environment 100 can be implemented according to any suitable communication protocol(s), including but not limited to cellular communication protocols of the first generation (1G), second generation (2G), third generation (3G), fourth generation (4G), fifth generation (5G), sixth generation (6G), etc., wireless local area network communication protocols such as Institute of Electrical and Electronics Engineers (IEEE) 802.11, and / or any other protocol known currently or to be developed in the future. In addition, the communication can utilize any suitable wireless communication technology, including but not limited to: Code Division Multiple Access (CDMA), Frequency Division Multiple Access (FDMA), Time Division Multiple Access (TDMA), Frequency Division Duplexing (FDD), Time Division Duplexing (TDD), Multiple-Input Multiple-Output (MIMO), Orthogonal Frequency Division Multiple Access (OFDM), Discrete Fourier Transform Spread OFDM (DFT-s-OFDM), and / or any other technology known currently or to be developed in the future. Example general process for ML abstract behavior management

[0067] Hereinafter, for illustrative purposes, some example embodiments are described with the first device 110 operating as an MnS producer and the second device 120 operating as an MnS consumer. However, in some other example embodiments, the operations described in connection with the first device 110 can be implemented at a device other than an MnS consumer, and the operations described in connection with the second device 120 can be implemented at a device other than an MnS producer.

[0068] Figure 2A An example signaling diagram of a process 200 for optimizing data samples for model training according to some example embodiments of the present disclosure is shown.

[0069] In operation, the first device 110 receives 210 a first request from the second device 120, the first request for querying validity information of a plurality of data samples for model training at the first device. The first device 110 determines 260 the validity information of at least one data sample among the plurality of data samples for model training, the validity information indicating the usefulness of the at least one data sample or at least one data feature for the performance of model training. Then, the first device 110 sends 290 a first report indicating the validity information to the second device 120.

[0070] In some example embodiments, the function of sending the first report, the parameter(s) for obtaining the validity information, and the parameter(s) for generating the first report can be configured, such as by defining and configuring data types and attribute parameters. Specifically, in some example embodiments, at least one parameter is configured for the first device 110. The at least one parameter includes at least one of the following: a first indication for enabling or disabling the function of sending the first report, the first report indicating the validity information, or at least one attribute parameter for determining the validity information.

[0071] In some example embodiments, the first device 110 receives 210 a configuration indicating at least one parameter from the second device.

[0072] In some example embodiments, the validity information is specific to at least one of the following: type of inference, training data source, data characteristics for model training, or set of data characteristics for model training.

[0073] In some example embodiments, the first report includes at least one of the following: at least one corresponding validity value of at least one data sample, or storage information of the first report.

[0074] In some example embodiments, the first request for querying validity information of a plurality of data samples used for model training at the first device may include a second indication instructing the first device 110 to provide the validity information.

[0075] In some example embodiments, the first report includes a third indication indicating one of the following: the validity information requested by the second device 120 has been successfully determined; the validity information requested by the second device 120 has been partially successfully determined; the validity information requested by the second device 120 has not been determined; or determining the validity information requested by the second device is not supported.

[0076] In some example embodiments, the first device 110 is a producer for model training, and the second device 120 is a consumer for model training.

[0077] In this way, a data validity tagging process is proposed, where the process enables the training function to be configured to and perform the process to identify and provide a report on data samples useful for ML model training and their degree of usefulness for ML model training.

[0078] In some example embodiments, the ML model training function enabling data validity tagging is capable of generating information on how each training data instance or feature contributes to model training. Such information can be part of a training report or can be transmitted as a special data validity report.

[0079] For better understanding, some specific example embodiments will still be referred to Figure 2A Some specific example embodiments are discussed. Additionally, the AI ML Training Consumer (AIMLT) is described as an example of the second device 120, and the AIMLT Producer is described as an example of the first device 110. It should be understood that these specific device types can be changed to other device types. The present disclosure is not limited in this regard.

[0080] In operation, an AIMLT consumer can optionally configure 210 the training data validity feedback to be enabled at the AIMLT producer. In some example embodiments, the AIMLT consumer initiates model training when the training data validity feedback is enabled.

[0081] In some example embodiments, the AIMLT consumer requests 220 model training. Optionally, the AIMLT consumer can send a first request 220 to the AIMLT producer.

[0082] In some example embodiments, the AIMLT producer sends 230 a response to the model training request to indicate that the request has been accepted but is not yet complete. The AIMLT producer starts model training using the complete training data set.

[0083] In some example embodiments, the AIMLT producer prepares 240 a complete set of training data samples for model training.

[0084] In some example embodiments, the AIMLT producer starts 250 model training using the complete training data set.

[0085] In some example embodiments, the AIMLT producer collects 260 the results of the effectiveness of the contribution of each training data instance to model training.

[0086] In some example embodiments, after AIML model training using the original data samples, the original data samples can be divided into two parts: One part of the data instances is important for model training and is called model training important data instances (MT-IDI). The other part makes no contribution or negligible contribution to model accuracy and is called model training redundant data instances (MT-RDI).

[0087] In some example embodiments, the AIMLT producer marks 270 the original complete data samples as MT-RDI and MT-IDI.

[0088] In some example embodiments, the AIMLT producer optionally obtains 270 the results of the usefulness / importance of each data feature.

[0089] In some example embodiments, the AIMLT producer generates 280 a training data validity report.

[0090] In some example embodiments, the AIMLT producer notifies 290 the AIMLT consumer of the AIML training report (such as the training data validity report).

[0091] Still referring toFigure 2A Let's discuss the workflow for generating reports on the contribution or effectiveness of training data to training data.

[0092] In some example embodiments, MT-RDIs can be safely discarded from model training because they do not contain (or only contain very little) additional information for model training accuracy.

[0093] In some example embodiments, the training producer can have two data sets after training, namely MT-IDI and MT-RDI. In other words, mark the entire training data sample as a useful data sample or a useless data sample.

[0094] In some example embodiments, for MT-IDI, the training producer can generate importance information of data features. In some example embodiments, machine learning algorithms (e.g., random forests, different types of decision tree-based models) can be used to generate importance information of data features.

[0095] In addition, in order to result in the reports described below, the Information Object Class (IOC) needs to be integrated in the NRM.

[0096] In some example embodiments, the AI / ML training capability is provided by the AIMLT MnS producer to one or more consumers.

[0097] In some example embodiments, AI / ML training can be triggered by requests from one or more AIMLT MnS consumers. To trigger AI / ML training, the AIMLT MnS consumer requests the AIMLT MnS producer to train an AI / ML model or an AI / ML-enabled function. In the AI / ML training request, the consumer should specify the inference type, which indicates the function or purpose of the AI / ML entity, such as CoverageProblemAnalysis (coverage problem analysis).

[0098] In some example embodiments, the AIMLT MnS producer can perform the training according to the specified inference type. The AIMLT MnS consumer can provide one or more data sources containing training data, which are considered as input candidates for training. To obtain effective training results, the consumer can also specify their requirements for model performance (e.g., accuracy, etc.) in the training request.

[0099] In some example embodiments, the AIMLT MnS producer provides a response to the consumer indicating whether the request is accepted.

[0100] In some example embodiments, if a request is accepted, the AIMLT MnS producer decides when to start AI / ML training considering the (multiple) requests from the (multiple) consumers. Once the training is decided, the AIMLT producer can select the training data considering the candidate training data provided by the consumers. Since the training data directly affects the algorithms and performance of the trained AI / ML entities, the AIMLT MnS producer can inspect the training data provided by the consumers and decide not to select any of them, select some of them, or select all of them. Additionally, the AIMLT MnS producer can select some other available training data. The AIMLT producer also uses the selected training data to train the AI / ML entities and provides the training results (including the location of the trained AI / ML entities, etc.) to the (multiple) AIMLT MnS consumers.

[0101] In some example embodiments, the training report can include information about the validity of the data in the model training per request from the consumers.

[0102] In some example embodiments, a new attribute for providing feedback on the validity of the training data to an existing IOC AIML training request is added as shown in Table 1 below. For example, the new attribute can be called trainingDataEffectivenessFeedback. Table 1 Examples of Attributes

[0103] In some example embodiments, the attribute of trainingDataEffectivenessFeedback will be a flag to allow the consumers to request whether to generate the training data validity for model training. It has a default value of FALSE, indicating that the consumers have not requested the training data validity feedback. When the value has been set to TRUE, the AIMLT MnS producer will address the requests related to the training data validity. In addition to the explicit requests from the consumers, the producer can decide to provide the trainingDataEffectivenessReport in the training report.

[0104] In some example embodiments, new data types and some attributes are added to the IOC AIMLTraingReport. The relevant descriptions are shown in Tables 2 and 3 below. Table 2 Examples of Attributes Table 3 Examples of Attribute Constraints

[0105] In some example embodiments, the attribute of trainingDataEffectivenessIndication will be a flag indicating the on or off state of the training data effectiveness report.

[0106] In some example embodiments, the data type represents a report of training data for model training effectiveness. When the training data effectiveness report is enabled (enabled by the consumer or initiated by the producer), at the end of model training, the report should be included as part of the MOI of the AIMLTrainingReport.

[0107] In some example embodiments, the attribute of trainingDataEffectivenessFeedbackResult represents the overall status of a report on a training data instance effectiveness feedback request, which has the following values: SUCCESSFUL_FEEDBACK indicates that the training data effectiveness feedback request was successfully executed and a report was generated. UNSUCCESSFUL_FEEDBACK indicates that the training data effectiveness feedback request was not successfully executed and no report was generated. PARTIAL_SUCCESSFUL_FEEDBACK indicates that the training data effectiveness feedback request was partially successfully executed (e.g., feature effectiveness was successful but data instance effectiveness was not, or vice versa, or only partial data instance effectiveness reporting was successful), and a report was generated. NOTSUPPORT indicates that the training data effectiveness feedback request is not supported.

[0108] In some example embodiments, the attribute of trainingDataInstanceEffectivenessInfo represents the overall effectiveness information of the data used for model training.

[0109] In some example embodiments, the attribute of trainingDataSourceEffectivenessInfo represents the effectiveness information of the data used for model training for each training data source (e.g., provided by the producer or provided by the consumer).

[0110] In some example embodiments, the attribute of trainingDataFeatureEffectivenessInfo represents the overall effectiveness information of the training features used for model training.

[0111] In some example embodiments, the attribute of trainingDataInstanceEffectivenessType represents the type of effectiveness being modeled, and it is a list of enumerations: BINARY_TYPE, indicating whether the training data instance is useful or not, and is indicated only by on and off. WEIGHT_TYPE, indicating that the contribution is measured by a weight or a real value from 0 to 1. 0 indicates no contribution at all, and the larger the value, the more the data instance contributes to model training.

[0112] In some example embodiments, when the file retrieval NRM fragment is supported by the MnS producer, the attribute of "linkToReportFiles" should be supported.

[0113] The following Tables 5 and 6 are further examples of attributes and attribute constraints. Examples of Attributes in Table 5 Examples of Attribute Constraints in Table 5

[0114] In some example embodiments, a training data effectiveness report (i.e., the first report) is introduced.

[0115] During AI / ML model training, training data effectiveness information can be provided for the following aspects: 1) whether a specific sample / data instance is useful or not for training, and / or the degree to which such a sample / data instance is useful for training; 2) whether a training feature of the training data (e.g., one measurement type among all types of KPIs / measurements in the input training data) is useful or not for training, and / or the degree of contribution of the training feature to model training.

[0116] Although such information can provide valuable feedback to the MnS training consumer regarding the effectiveness of the training data set during the training process, and can even point out the most valuable data instances (and / or training features) in the given data set for which training has been performed, such a "one-time", simple classification of data instances is of little significance for improving future (re)training processes if the available data is not further processed to obtain more understandable insights.

[0117] Example embodiments related to effectiveness information can be used for one of multiple use cases.

[0118] For AI / ML model training, a large number of data instances may not add value. For example, if only a portion contributes to the actual model training, the other portion will be discarded by some well-designed algorithms, such as LightGBM, which includes the GOSS feature. This feature will only use the top N data instances with high gradients for model training and discard all the remaining data instances associated with small gradients, which have negligible or no contribution to model training. Therefore, LightGBM is much faster than other gradient-based boosting algorithms without negatively impacting model performance.

[0119] Obviously, this selection of high-quality training data samples (which contribute more to model training) can reduce the amount of training data while still achieving the expected training accuracy with minimal training time and minimizing data acquisition, data storage, and data preprocessing, thus greatly improving the overall training efficiency.

[0120] Therefore, it is necessary to provide means for AI / ML model training producers to report the effectiveness of the data used in training, or the effectiveness of a particular sample / data instance, or the effectiveness of training features (specific input data, e.g., one measurement type among all types of KPIs / measurements), or both. Using the insights in such reports, consumers and producers can take measures to optimize model training, e.g., to ensure minimizing storage and processing costs by discarding less effective raw data without compromising the ability to obtain sufficient historical information to fully train the AI / ML application. In other words, the management system needs to generate the effectiveness of the data used in training, e.g., based on the actual contribution to model training.

[0121] In some example embodiments, the 3GPP management system should have the ability to allow authorized consumers to configure the AI / ML training function to report the effectiveness of the data used for model training.

[0122] In some example embodiments, the 3GPP management system should have the ability to report the effectiveness of the data used in training, including: Indicating or marking which data instances are useful or not useful; Indicating which data features are useful or not useful, or negatively contributing to model training.

[0123] In some example embodiments, the 3GPP management system should have the ability to allow authorized consumers to activate the AI / ML training function / entity to mark the effectiveness of the data used for training.

[0124] In addition, for better understanding, reference may be made to Figure 2B , which shows an example signaling diagram of a process 295 for optimizing data samples for model training according to some example embodiments of the present disclosure.

[0125] The above process has been mainly discussed with respect to example embodiments related to validity information. As described above, validity information can be used as an input parameter for determining information that can be used to optimize a set of data samples (i.e., information associated with a first set of multiple data samples having a higher validity value).

[0126] First, reference is made to Figure 3 , which shows an example signaling diagram of a process 300 for optimizing data samples for model training according to some example embodiments of the present disclosure. For the purpose of discussion, process 300 will be discussed with reference to Figures 1A to 1C , for example, by using a first device 110 and a second device 120.

[0127] In operation, the first device 110 receives 310 a second request from the second device 120, the second request being for information associated with a first set of multiple data samples for model training, wherein a first set of multiple validity values of the first set of multiple data samples is higher than a second set of multiple validity values of a second set of multiple data samples for model training, and wherein the validity value indicates the degree of usefulness of at least one data sample or at least one data feature for the performance of model training.

[0128] The first device 110 determines 320 the information at least in part based on the second request, and then sends 330 a second report to the second device 120, the second report indicating the information associated with the first set of multiple data samples.

[0129] In some example embodiments, the first device 110 determines 322 the information by itself. Alternatively, the first device 110 determines 322 the information by requesting information from another device. Specifically, in some example embodiments, the first device 110 generates a third request for the information based on the second request, and sends 324 the third request for the information to a third device, the third device being a management data analysis (MDA) producer. Then, the first device 110 receives the information from the third device.

[0130] In some example embodiments, the second request includes at least one parameter for determining the first set of multiple data samples, the at least one parameter including at least one of the following: a validity threshold, version information of the model, or a specific use case of one or more models.

[0131] In some example embodiments, the first device 110 determines information by determining a first set of a plurality of data samples based on at least one parameter and determining information based on context information of the first set of a plurality of data samples.

[0132] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training.

[0133] Alternatively, in some example embodiments, the first device is a producer for management data analytics (MDA), and the second device is a consumer for MDA.

[0134] In some example embodiments, the information includes at least one of the following: a combination of data features corresponding to the first set of a plurality of data samples, identification information of at least one network object corresponding to the first set of a plurality of data samples, geographical location information of at least one network object, network area information corresponding to the first set of a plurality of data samples, or acquisition time information corresponding to the first set of a plurality of data samples.

[0135] It can be seen that the operations of the first device 120 and the third device during the process including actions 324, 325, and 326 are similar to the operations of the second device 120 and the first device 110 with respect to the process including actions 310, 320, and 330. Therefore, the description of the actions 310, 320, and 330 can be applied to actions 324, 325, and 326.

[0136] Now refer to Figure 4 , which shows an example signaling diagram of a process 400 for optimizing data samples for model training according to some example embodiments of the present disclosure. For the purpose of discussion, the process 400 will be discussed with reference to Figures 1A to 1C , for example, by using the first device 110 and the second device 120.

[0137] In operation, the second device 120 may send 410 a first request to the first device 110, where the first request may request the first device 110 to provide validity information. The first device 110 may send 420 a response indicating that the first request is accepted. Then, the first device 110 determines 430 the validity information. Next, a first report indicating the validity information may be provided to the second device 120.

[0138] Based on the first report, the second device 120 can send a second request 450 to the third device, and the second request is for information associated with a first set of multiple data samples used for model training. The third device can determine 460 the relevant information and send the information to the second device 120. It can be seen that during the process including actions 460 and 470, the operations of the second device 120 and the third device are similar to the operations of the second device 120 and the first device 110 with respect to the process including actions 320 and 330. Therefore, the descriptions of actions 320 and 330 can be applied to actions 460 and 470.

[0139] According to some embodiments of the present disclosure, a method is introduced, which is used to analyze data to understand which data is most useful for ML model training / retraining.

[0140] In some example embodiments, the results of such analysis may include insights into combinations of different data features, data collected on certain network objects, data collected during certain time periods, or any combination thereof that contributes most to model training. Such insights are derived from information related to the importance or effectiveness of different training data samples or features, and the sample or feature comes from the output of an ML model training function that can label its data samples and features with their effectiveness.

[0141] In some example embodiments, the effective training data pattern analysis consumer may request the producer to identify the pattern of the data used for training that contributes most to the training process (e.g., input data that causes significant gradient changes).

[0142] In some example embodiments, the consumer can specify in his request whether the most effective training data should be derived for a specific version of the ML model, all versions of the ML model, or one of all ML models related to a use case / problem.

[0143] In some example embodiments, the effective training data pattern analysis producer derives and provides the pattern of the data used for training that contributes most to the training process, that is, the effective training data pattern. Such patterns may include the following information (or any combination of the following information): A set of data features that, when used simultaneously (combined) as input to model training, have significant effectiveness for the training process. This can be expressed, for example, by a list of DNs (distinguished names) of the most important performance metrics or KPIs for model training. A list of DNs (distinguished names) of network objects from which the most effective data features are collected, A description of the area from which the most effective data features are collected, which can be expressed, for example, by a cell list (E-UTRAN-CGI or NG-RAN CGI), a tracking area list (identified by the TAC - tracking area code). Information about the geographical location (latitude and longitude) of the network object from which the most effective data features are collected, or information about a larger geographical area (Multiple) time windows during which the most effective data features have been collected.

[0144] In some example embodiments, an effective training data pattern is derived by learning the association between the importance of data instances during ML model training and the context (e.g., time or geographical location) in which a given data instance is collected.

[0145] In some example embodiments, such a pattern can be used as a benchmark for improving data collection to be used for model training. The EffectiveTrainingDataPatternAnalytics management service can be built into the MnS training producer. On the other hand, this capability can be part of the MDAS producer in the management domain, and the NWDAF in the core domain, or the RAN intelligence in the RAN domain.

[0146] In some example embodiments, the input to the analysis is a tag indicating the importance of a specific data sample or feature (such as validity information), and similar to the following 1, where the highlighted columns and rows provide the results of training data classification in terms of validity (true / false) and validity level / weight (indicating the degree of contribution of the data instance or each feature to the gradient). This classification provides the validity / importance tag of the data instances used during the training of an AI / MLEntity with a given aIMLEntity ID (or optionally together with all past trained mLEntityVersions) or all AIMLEntities of a given inference type. The validity / importance tag provides information on how useful a specific data instance is in AI / MLEntity training. Table 7 Example tags for the validity of each sample used for data training ID Timestamp Cell_name ERAB_Fail_Resource Avg_Lat_All_5QI DL_PRB_Util_Percent Validity / Level 1 2019-04-19T15:00:00Z CELL90FDD3 0 16 8.1 T / 0.8 2 2019-04-19T15:15:00Z CELL90FDD3 0 9 5.4 T / 0.9 3 2019-04-19T15:30:00Z CELL90FDD3 0 16 9.2 T / 0.7 4 2019-04-19T15:45:00Z CELL90FDD3 0 9 8.5 T / 0.8 5 2019-04-19T16:00:00Z CELL90FDD3 0 7 10 T / 0.66 6 2019-04-19T16:15:00Z CELL90FDD3 0 20 15.9 T / 0.77 7 2019-04-19T16:30:00Z CELL90FDD3 0 23 8.1 T / 0.69 8 2019-04-19T16:45:00Z CELL90FDD3 0 39 10.9 T / 0.88 9 2019-04-19T15:00:00Z CELL92FDD4 0 106 53.3 F / 0.0001 10 2019-04-19T15:15:00Z CELL92FDD4 0 77 46.5 F / 0.002 11 2019-04-19T15:30:00Z CELL92FDD4 0 75 63.9 F / 0.003 12 2019-04-19T15:45:00Z CELL92FDD4 0 86 52.2 F / 0.005 13 2019-04-19T16:00:00Z CELL92FDD4 0 92 57.4 F / 0.01 14 2019-04-19T16:15:00Z CELL92FDD4 0 150 92.1 F / 0.009 15 2019-04-19T16:30:00Z CELL92FDD4 0 149 81.6 F / 0.0034 16 2019-04-19T16:45:00Z CELL92FDD4 0 66 38.4 F / 0.005 n / a Validity / Level T / 0.8 F / 0.0001 T / 0.61 T / 0.8 n / a

[0147] In Table 7 above, the first row is the header of the data, and each column is called a data feature; each row except the header row is called a data instance or data sample. The header row and the ID column are not used for model training. If there is no other unique key to identify the data instance, the ID is used to uniquely identify the data instance.

[0148] In some example embodiments, the output of the analysis is an effectiveTrainingDataPattern, which can be represented as a general effectiveTrainingDataPattern (i.e., data type) (e.g., used as the output from an MDA producer), having the following attributes, each of which is a separate data type: effectiveDataFeatures: A combination of data features that have the greatest impact on ML training, effectiveNetworkObjects: A list of network objects from which effective data is collected, effectiveLocation: A description of the geographical location or area from which effective data is collected, effectiveTime: The moment or period during which effective data is collected, effectiveDataHybrid: Any combination of the above attribute data types (e.g., effectiveNetworkObjects and effectiveTime indicate the network objects and time from which effective data is collected).

[0149] In some example embodiments, this pattern of most effective training can represent the characteristics of an AIML entity, and thus for each AIML entity, the corresponding description of the pattern of the most effective training data can be part of the AIML entity. Tables 8 and 9 below show examples of newly defined attributes and attribute constraints. Table 8 Attribute Examples Table 9 Examples of Attribute Constraints

[0150] In some example embodiments, an effectiveTrainingDataPattern can be derived and associated with a specific tuple of mLEntityId, inerenceType, and mLEntityVersion. However, the effectiveTrainingDataPattern can be derived for a certain mLEntityId that is valid for all versions of the AIMLEntity. Additionally, the MnS producer can derive the effectiveTrainingDataPattern for a certain inference type, which should be a good indicator of which data is most effective for training an ML model related to some specific use cases indicated by the inference type.

[0151] Examples of the end-to-end flow are described below. In some example embodiments, the MnS consumer (e.g., an operator) may be interested in obtaining insights into which data (e.g., combinations of features collected on which network objects and when) contributes most to the effectiveness of the training process. The MnS consumer can issue a new request proposed in the present invention, the AI / ML training analysis request, by specifying at least one of the following inputs: an aIMLEntity ID (optionally together with the mLEntityVersion) or the inerenceType.

[0152] Optionally, the MnS consumer can indicate a threshold above which data instances are considered valid in model training and the consumer requests an analysis based on this threshold. Note that in this step, we describe a general request for performing an analysis on training data, however, such a request can be an extension of a training request. Additionally, as indicated above, the ability to perform an analysis on the effectiveness of training data can be built into the MnS training producer, the MDAS producer in the management domain, the NWDAF in the core domain, or the RAN intelligence in the RAN domain. Section 6.4 provides more details on these different options.

[0153] In some example embodiments, the MnS producer relies on the tagging of training data. The MnS producer only considers the most important data during MLEntity training and analyzes its context data, e.g., according to the network object and / or geographical location from which a specific data instance was collected, the moment when a specific data instance was collected, etc.

[0154] In the case where the MnS consumer provides a specific threshold, based on this specific threshold, the data should be considered valid or not, and the MnS producer will use this as a benchmark for performing the requested analysis. That is, the MnS producer can only consider data samples above the desired validity threshold and analyze which features contribute the most, especially for which combination of different features, the most important impact on the gradient has been observed. Once the most important combination of data features is identified, the MnS producer can identify at which moments the input data samples have the greatest impact for this specific feature combination. Among such moments, the most influential time periods can be obtained, such as peak hours on weekdays (7 am to 9 am and 4 pm to 6 pm).

[0155] Similarly, the MnS producer can identify each network object from which such the most influential data instances are collected. By further analyzing and aggregating such network objects, the MnS producer can identify the geographical area, such as an urban area, from which the most influential data instances have been collected.

[0156] In some example embodiments, the MnS producer provides the analysis result to the MnS consumer, and the analysis result includes a description of the pattern that the most important data instances have. The pattern description may include the following attributes: The combination of data features that have the greatest impact on training, which can be expressed by, for example, the DN (distinguished name) of the performance metric or KPI that is most important for model training. For example, the combination from Table 7: Avg_Lat_All_5QI and DL_PRB_Util_Percent can be indicated. The DN (distinguished name) of the object on which the data is selected and collected. For example, from Table 7: CELL90FDD3. A description of the area from which the most effective data is collected, which can be expressed by, for example, a cell list (E-UTRAN-CGI or NG-RAN CGI), a tracking area list (identified by TAC - tracking area code). Information on the geographical location (latitude and longitude) of the object from which the most effective data is collected or information on a larger geographical area, The (multiple) time windows during which the most effective data is collected. For example, from Table 7: 2019-04-19T15:00 to 16:00, or Any combination of the above attributes.

[0157] In some example embodiments, such a pattern can be provided for a specific version of the AIML entity, for all versions of the AIML entity, or for all AIML entities related to certain use cases / issues (e.g., "coverage and capacity optimization").

[0158] Figure 5 and Figure 6 FIG. 6 shows an example signaling diagram for processes 500 and 600 for optimizing data samples for model training according to some example embodiments of the present disclosure.

[0159] In some example embodiments, the ability to perform an analysis on the training data validity can be built into the MnS training producer, the MDA service (MDAS) producer in the management domain, the network data analytics function (NWDAF) in the core domain, or the RAN intelligence in the RAN domain. Figure 6 FIG. 11 shows an option where such an ability is part of the MnS training producer and the MDAS producer ( Figure 6 the MDA_producer in the MDAS_producer)). The request to perform an analysis on the training data validity is an extension of the AIMLT training (AIMLTraining) request.

[0160] In Figure 6 a specific embodiment, the ability to request an analysis of the training data validity is part of the AIMLT Producer or the MDA_producer.

[0161] In some example embodiments, the consumer initiates model training using the enabled training data validity feedback including the analysis. The producer sends a response to the model training request indicating that the request is accepted but not yet completed. If the analysis ability is available at the AIMLT producer, it performs an analysis on the validity of the data used for model training (step 2). Otherwise, the AIMLT producer can request the MDA_producer to perform such an analysis. The MDA_producer generates an analysis according to what is requested by the AIMLT producer. The MDA_producer notifies the AIMLT producer of the resulting analysis on the training data validity. Finally, the consumer is notified of the requested analysis on the training data validity.

[0162] Alternatively, in some other example embodiments, the consumer can directly request the training data validity from the MDA producer. Action 10 is an interaction between the consumer and the AIMLT producer to perform training and provide a flag indicating the training data validity. The consumer requests the MDA_producer for a validity analysis on the validity of the data used for model training. The MDA producer generates a validity report by performing an analysis. Then, the MDA_producer notifies the consumer of the resulting analysis on the training data validity.

[0163] In some example embodiments, during ML model training, information can be provided as to whether a particular sample is useful or not useful for training, and as to the degree to which the sample is useful for training. Although such information can provide valuable feedback to the MnS training consumer regarding the effectiveness of the training dataset during the training process, and can even point out the most valuable data instances in the given dataset on which training has been performed, such a "one-time", simple classification of data instances is of little significance for improving future (re)-training processes if the available data is not further processed to obtain more easily understandable insights.

[0164] In some example embodiments, a single / independent observation as to whether a particular sample or feature at a given timestamp contributes to the model gradient does not provide an understanding as to whether using such a sample or feature will contribute to the model accuracy of the same model in further training / re-training, or whether it will contribute to the training efficiency of other models related to the same use case.

[0165] To have such an understanding, further analysis of the data related to the importance of data instances during training is needed. Patterns of the most effective training data generated from the analysis will be very helpful in improving data collection in order to optimize the quantity and quality of the data to be used for training.

[0166] In some example embodiments, the 3GPP management system should have the ability to allow an authorized consumer to request an analysis of the data used for model training regarding the effectiveness of such data during model training.

[0167] In some example embodiments, the 3GPP management system should have the ability to generate and provide to the consumer a pattern of highly effective data for a particular version of a model, all versions of a single model, or all models related to a particular use case for training the model. Example method

[0168] Figure 7 A flowchart of an example method 700 implemented at a first device in accordance with some example embodiments of the present disclosure is shown. For purposes of discussion, method 700 will be described from the perspective of the first device 110 in Figure 1A the following.

[0169] At block 710, the first device receives a first request from a second device, the first request querying validity information for a plurality of data samples for model training at the first device.

[0170] At block 720, a first device determines validity information for at least one data sample among a plurality of data samples for model training, the validity information indicating the usefulness of the at least one data sample or at least one data feature for the performance of model training.

[0171] At block 730, the first device sends a first report indicating the validity information to a second device.

[0172] In some example embodiments, the method further includes: configuring at least one parameter for the first device, the at least one parameter including at least one of the following: a first indication for enabling or disabling the function of sending the first report, the first report indicating the validity information, or at least one attribute parameter for determining the validity information.

[0173] In some example embodiments, the method further includes: receiving, from the second device, a configuration indicating the at least one parameter.

[0174] In some example embodiments, the validity information is specific to at least one of the following: an inference type, a training data source, a data feature for model training, or a set of data features for model training.

[0175] In some example embodiments, model training at the first device is triggered by receiving a first request.

[0176] In some example embodiments, the first request for triggering model training includes a second indication instructing the first device to provide the validity information.

[0177] In some example embodiments, the first report includes a third indication indicating one of the following: the validity information requested by the second device has been successfully determined; the validity information requested by the second device has been partially successfully determined; the validity information requested by the second device has not been determined; or determining the validity information requested by the second device is not supported.

[0178] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training.

[0179] Figure 8 A flowchart of an example method 800 implemented at a second device according to some example embodiments of the present disclosure is shown. For purposes of discussion, method 800 will be described from Figure 1A the perspective of the second device 120 in

[0180] At block 810, the second device sends a first request to the first device, the first request querying the validity information of a plurality of data samples for model training at the first device.

[0181] At block 820, a second device receives a first report from a first device, the first report indicating validity information for at least one data sample among a plurality of data samples used for model training, the validity information indicating the usefulness of the at least one data sample or at least one data feature for the performance of model training, the validity information being determined by the first device during model training by using the plurality of data samples.

[0182] In some example embodiments, the method further includes: sending a configuration indicating at least one parameter, the at least one parameter including at least one of the following: a first indication for enabling or disabling a function of sending the first report, the first report indicating validity information, or at least one attribute parameter for determining the validity information.

[0183] In some example embodiments, the first report includes at least one of the following: a validity value of the corresponding at least one data sample, or storage information of the first report.

[0184] In some example embodiments, the validity information is specific to at least one of the following: an inference type, a training data source, a data feature for model training, or a set of data features for model training.

[0185] In some example embodiments, model training at the first device is triggered by receiving a first request.

[0186] In some example embodiments, the first report includes a third indication of at least one of the following: the validity information requested by the second device has been successfully determined, the validity information requested by the second device has been partially successfully determined, the validity information requested by the second device has not been determined, or determining the validity information requested by the second device is not supported.

[0187] In some example embodiments, the method further includes: configuring the second device to enable the second device to send a first request including a second indication.

[0188] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training.

[0189] Figure 9 A flowchart of an example method 900 implemented at a third device according to some example embodiments of the present disclosure is shown. For purposes of discussion, method 900 will be described from the perspective of the first device 1a0 in Figure 1A which.

[0190] At block 910, a first device receives a second request from a second device, the second request being for information associated with a first set of a plurality of data samples for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity values indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of model training.

[0191] At block 920, the first device determines the information at least in part based on the second request.

[0192] At block 930, the first device sends a second report to the second device, the second report indicating the information associated with the first set of a plurality of data samples.

[0193] In some example embodiments, determining the information includes: generating, based on the second request, a third request for the information; sending the third request for the information to a third device, the third device being a producer for management data analysis (MDA); and receiving the information from the third device.

[0194] In some example embodiments, the second request includes at least one parameter for determining the first set of a plurality of data samples, the at least one parameter including at least one of the following: a validity threshold, version information of a model, or a specific use case of one or more models.

[0195] In some example embodiments, determining the information includes: determining the first set of a plurality of data samples based on the at least one parameter, and determining the information based on context information of the first set of a plurality of data samples.

[0196] In some example embodiments, the information includes at least one of the following: a combination of data features corresponding to the first set of a plurality of data samples, identification information of at least one network object corresponding to the first set of a plurality of data samples, geographical location information of at least one network object, network region information corresponding to the first set of a plurality of data samples, acquisition time information corresponding to the first set of a plurality of data samples.

[0197] In some example embodiments, the first device is a producer for model training and the second device is a consumer for model training, the first device is a producer for training data analysis and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA) and the second device is a consumer for MDA.

[0198] Figure 10 A flowchart of an example method 1000 implemented at a fourth device according to some example embodiments of the present disclosure is shown. For purposes of discussion, method 1000 will be described from the perspective of the second device 120 in FIG. 1.

[0199] At block 1010, a first device sends a second request to a second device, the second request being for information associated with a first set of multiple data samples for model training, the first set of multiple validity values of the first set of multiple data samples being higher than the second set of multiple validity values of a second set of multiple data samples for model training, the validity values indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of model training.

[0200] At block 1020, the first device receives a second report from the second device, the second report indicating information associated with the first set of multiple data samples.

[0201] In some example embodiments, the second request includes at least one parameter for determining the first set of multiple data samples, the at least one parameter including at least one of the following: a validity threshold, version information of one or more models, or a specific use case of one or more models.

[0202] In some example embodiments, the information includes at least one of the following: a combination of data features corresponding to the first set of multiple data samples, identification information of at least one network object corresponding to the first set of multiple data samples, geographical location information of at least one network object, network area information corresponding to the first set of multiple data samples, or acquisition time information corresponding to the first set of multiple data samples.

[0203] In some example embodiments, the first device is a producer for model training and the second device is a consumer for model training, the first device is a producer for training data analysis and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA) and the second device is a consumer for MDA. Example device, equipment and medium

[0204] In some example embodiments, a first apparatus (e.g., Figure 1A the first device 110 in Figure 1A ) capable of performing any of the methods 700 may include components for performing the corresponding operations of the method 700. The components may be implemented in any suitable form. For example, the components may be implemented in circuitry or software modules. The first apparatus may be implemented as or included in

[0205] In some example embodiments, the first device includes: components for receiving a first request from a second device, the first request for querying validity information of a plurality of data samples for model training at the first device; components for determining validity information of at least one data sample among the plurality of data samples for model training, the validity information indicating the usefulness of the at least one data sample or at least one data feature for the performance of model training; and components for sending a first report indicating the validity information to the second device.

[0206] In some example embodiments, it further includes: components for configuring at least one parameter for the first device, the at least one parameter including at least one of the following: a first indication for enabling or disabling the function of sending the first report, the first report indicating the validity information, or at least one attribute parameter for determining the validity information.

[0207] In some example embodiments, it further includes: components for receiving, from the second device, a configuration indicating the at least one parameter.

[0208] In some example embodiments, the validity information is specific to at least one of the following: the inference type, the training data source, the data features for model training, or the set of data features for model training.

[0209] In some example embodiments, the first report includes at least one of the following: at least one corresponding validity value of at least one data sample, or storage information of the first report.

[0210] In some example embodiments, model training at the first device is triggered by receiving the first request.

[0211] In some example embodiments, the first report includes a third indication indicating one of the following: the validity information requested by the second device has been successfully determined; the validity information requested by the second device has been partially successfully determined; the validity information requested by the second device has not been determined; or determining the validity information requested by the second device is not supported.

[0212] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training.

[0213] In some example embodiments, the first device further includes components for performing other operations in some example embodiments of method 700 or the first device 110. In some example embodiments, the components include: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the first device to execute.

[0214] In some example embodiments, the second device capable of performing any of the methods 800 (e.g.,Figure 1A The second device 120) in can include components for performing the corresponding operations of method 800. The components can be implemented in any suitable form. For example, the components can be implemented in a circuit system or a software module. The second device can be implemented as or included in Figure 1A the second device 120 in.

[0215] In some example embodiments, the second device includes: a component for sending a first request to the first device to trigger model training at the first device; and a component for receiving a first report from the second device, the first report indicating validity information of at least one data sample among a plurality of data samples used for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training, and the validity information being determined by the first device during model training by using the plurality of data samples.

[0216] In some example embodiments, it further includes: a component for sending a configuration indicating at least one parameter, the at least one parameter including at least one of the following: a first indication for enabling or disabling the function of sending the first report, the first report indicating validity information, or at least one attribute parameter for determining the validity information.

[0217] In some example embodiments, the first report includes at least one of the following: a validity value of the corresponding at least one data sample, or storage information of the first report.

[0218] In some example embodiments, the validity information is specific to at least one of the following: an inference type, a training data source, a data feature for model training, or a data feature set for model training.

[0219] In some example embodiments, the first request for triggering model training includes a second indication for instructing the first device to provide validity information.

[0220] In some example embodiments, the first report includes a third indication of at least one of the following: the validity information requested by the second device has been successfully determined, the validity information requested by the second device has been partially successfully determined, the validity information requested by the second device has not been determined, or determining the validity information requested by the second device is not supported.

[0221] In some example embodiments, it further includes: a component for configuring the second device to enable the second device to send the first request including the second indication.

[0222] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training.

[0223] In some example embodiments, the second device further includes components for performing method 800 or other operations in some example embodiments of the second device 120. In some example embodiments, the components include: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the second device to perform.

[0224] In some example embodiments, a third device (e.g., Figure 1A the first device 110 in Figure 1A ) that can perform any of the methods 900 may include components for performing the corresponding operations of method 900. The components may be implemented in any suitable form. For example, the components may be implemented in a circuit system or a software module. The first device may be implemented as or included in

[0225] In some example embodiments, the first device includes: components for receiving a second request from the second device, the second request being for information associated with a first set of a plurality of data samples for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than the second set of a plurality of validity values of a second set of a plurality of data samples for model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of model training; components for determining the information at least in part based on the second request; and components for sending a second report to the second device, the second report indicating the information associated with the first set of a plurality of data samples.

[0226] In some example embodiments, the components for determining the information include: components for generating a third request for the information based on the second request; components for sending the third request for the information to a third device, the third device being a producer for managing data analysis (MDA); and components for receiving the information from the third device.

[0227] In some example embodiments, the second request includes at least one parameter for determining the first set of a plurality of data samples, the at least one parameter including at least one of the following: a validity threshold, version information of the model, or a specific use case of one or more models.

[0228] In some example embodiments, the components for determining the information include: components for determining the first set of a plurality of data samples based on the at least one parameter, and components for determining the information based on the context information of the first set of a plurality of data samples.

[0229] In some example embodiments, the information includes at least one of the following: a combination of data characteristics corresponding to a first set of a plurality of data samples, identification information of at least one network object corresponding to the first set of a plurality of data samples, geographical location information of at least one network object, network area information corresponding to the first set of a plurality of data samples, or acquisition time information corresponding to the first set of a plurality of data samples.

[0230] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training, the first device is a producer for training data analysis, and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA), and the second device is a consumer for MDA.

[0231] In some example embodiments, the first device further includes components for performing method 900 or other operations in some example embodiments of the first device 110. In some example embodiments, the components include: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the execution of the third device.

[0232] In some example embodiments, the second device (e.g., Figure 1A the second device 120 in) capable of performing any of the methods 1000 may include components for performing the corresponding operations of method 1000. The components may be implemented in any suitable form. For example, the components may be implemented in a circuit system or a software module. The fourth device may be implemented as or included in Figure 1A the second device 120 in.

[0233] In some example embodiments, the second device includes: a component for sending a second request to the first device, the second request being for information associated with a first set of a plurality of data samples for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than the second set of a plurality of validity values of the second set of a plurality of data samples for model training, and the validity information indicating the usefulness of at least one data sample or at least one data characteristic for the performance of model training; and a component for receiving a second report from the first device, the second report indicating the information associated with the first set of a plurality of data samples.

[0234] In some example embodiments, the second request includes at least one parameter for determining the first set of a plurality of data samples, the at least one parameter including at least one of the following: a validity threshold, version information of one or more models, or a specific use case of one or more models.

[0235] In some example embodiments, the information includes at least one of the following: a combination of data features corresponding to a first set of multiple data samples, identification information of at least one network object corresponding to the first set of multiple data samples, geographical location information of at least one network object, network area information corresponding to the first set of multiple data samples, and acquisition time information corresponding to the first set of multiple data samples.

[0236] In some example embodiments, the first device is a producer for model training, and the second device is a consumer for model training. The first device is a producer for training data analysis, and the second device is a consumer for training data analysis. Or the first device is a producer for management data analysis (MDA), and the second device is a consumer for MDA.

[0237] In some example embodiments, the fourth device further includes components for performing method 1000 or other operations in some example embodiments of the second device 120. In some example embodiments, the components include: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the fourth device to execute.

[0238] Figure 11 is a simplified block diagram of a device 1100 suitable for implementing example embodiments of the present disclosure. The device 1100 may be provided to implement a communication device, such as Figure 1A the first device 110 or the second device 120 shown in. As shown, the device 1100 includes one or more processors 1110, one or more memories 1120 coupled to the processors 1110, and one or more communication modules 1140 coupled to the processors 1110.

[0239] The communication module 1140 is used for two-way communication. The communication module 1140 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interface may represent any interface necessary for communicating with other network elements. In some example embodiments, the communication module 1140 may include at least one antenna.

[0240] As a non-limiting example, the processor 1110 may be of any type suitable for a local technology network and may include one or more of the following: a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture. The device 1100 may have multiple processors, such as an application-specific integrated circuit chip that is clocked subordinate to a synchronous main processor over time.

[0241] Memory 1120 may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, read-only memory (ROM) 1124, electrically programmable read-only memory (EPROM), flash memory, hard disk, optical disc (CD), digital video disc (DVD), optical disc, laser disc, and other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, random access memory (RAM) 1122 and other volatile memories that do not persist during power loss.

[0242] Computer program 1130 includes computer-executable instructions executed by an associated processor 1110. The instructions of program 1130 may include instructions for performing the operations / actions of some example embodiments of the present disclosure. Program 1130 may be stored in a memory, such as ROM 1124. Processor 1110 may execute any suitable actions and processes by loading program 1130 into RAM 1122.

[0243] Example embodiments of the present disclosure may be implemented by means of program 1130 such that device 1100 may perform any process of the present disclosure as discussed with reference to Figures 2A to 10 Example embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.

[0244] In some example embodiments, program 1130 may be tangibly embodied in a computer-readable medium, which may be included in device 1100 (such as in memory 1120) or in other storage devices accessible by device 1100. Device 1100 may load program 1130 from the computer-readable medium into RAM 1122 for execution. In some example embodiments, the computer-readable medium may include any type of non-transitory storage medium, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc. As used herein, the term "non-transitory" is a limitation of the medium itself (i.e., tangible, rather than a signal), rather than a limitation on data storage persistence (e.g., RAM versus ROM).

[0245] Figure 12 An example of a computer-readable medium 1200 is shown that may be in the form of a CD, DVD, or other optical storage disc. Program 1130 is stored on computer-readable medium 1200.

[0246] Generally, the various embodiments of the present disclosure may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, a microprocessor, or other computing devices. Although the various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.

[0247] Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a computer-readable medium (such as a non-transitory computer-readable medium). The computer program product includes computer-executable instructions, such as those included in program modules executed in a device on a target physical or virtual processor, to perform any of the methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules may be combined or split as needed among program modules. The machine-executable instructions for program modules may be executed within local or distributed devices. In a distributed device, program modules may be located in both local and remote storage media.

[0248] The program code for performing the methods of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0249] In the context of the present disclosure, the computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc.

[0250] A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium will include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0251] Moreover, although the operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some instances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment unless expressly stated otherwise. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments unless expressly stated otherwise.

[0252] Although the present disclosure has been described in language specific to structural features and / or methodological acts, it is to be understood that the disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A first device, comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the first device to at least perform: receive a first request from a second device, the first request for querying validity information of a plurality of data samples for model training at the first device; determine validity information of at least one data sample among the plurality of data samples for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training; and send a first report indicating the validity information to the second device.

2. The first device according to claim 1, wherein the first device is further caused to perform: configure at least one parameter for the first device, the at least one parameter including at least one of the following: a first indication for enabling or disabling the function of sending the first report indicating the validity information, or at least one attribute parameter for determining the validity information.

3. The first device according to claim 2, wherein the first device is further caused to perform: receive a configuration indicating the at least one parameter from a second device.

4. The first device according to any one of claims 1 to 3, wherein the validity information is specific to at least one of the following: inference type, training data source, data features for the model training, or a set of data features for the model training.

5. The first device according to any one of claims 1 to 4, wherein the first report includes at least one of the following: at least one corresponding validity value of the at least one data sample, or storage information of the first report.

6. The first device according to any one of claims 1 to 5, wherein the model training at the first device is triggered by receiving the first request.

7. The first device according to claim 1, wherein the first report includes a third indication indicating one of the following: the validity information requested by the second device has been successfully determined; the validity information requested by the second device has been partially successfully determined; the validity information requested by the second device has not been determined; or determining the validity information requested by the second device is not supported.

8. The first device according to any one of claims 1 to 7, wherein the first device is a producer for model training and the second device is a consumer for model training.

9. A second device, comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the second device to at least perform: send a first request to a first device, the first request for querying validity information of a plurality of data samples for model training at the first device; and Receive a first report from the second device, the first report indicating validity information of at least one data sample among a plurality of data samples used for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training, and the validity information being determined by the first device during the execution of the model training by using the plurality of data samples.

10. The second device according to claim 9, wherein the second device is further caused to perform: Send an indication of the configuration of the at least one parameter, the at least one parameter including at least one of the following: A first indication for enabling or disabling the function of sending the first report, the first report indicating the validity information, or At least one attribute parameter for determining the validity information.

11. The second device according to claim 9 or 10, wherein the first report includes at least one of the following: The validity value of the corresponding at least one data sample, or The storage information of the first report.

12. The second device according to any one of claims 9 to 11, wherein the validity information is specific to at least one of the following: Inference type, Training data source, Data features for the model training, or A set of data features for the model training.

13. The second device according to any one of claims 9 to 12, wherein the model training at the first device is triggered by receiving the first request.

14. The second device according to claim 9, wherein the first report includes a third indication of at least one of the following: The validity information requested by the second device has been successfully determined, The validity information requested by the second device has been partially successfully determined, The validity information requested by the second device has not been determined, or Determining the validity information requested by the second device is not supported.

15. The second device according to claim 13, wherein the second device is further caused to perform: Configure the second device to enable the second device to send the first request including a second indication.

16. The second device according to any one of claims 9 to 14, wherein the first device is a producer for model training, and the second device is a consumer for model training.

17. A first device, comprising: At least one processor; and At least one memory storing instructions which, when executed by the at least one processor, cause the first device to at least perform: Receive a second request from a second device, the second request being for information associated with a first set of a plurality of data samples used for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than the second set of a plurality of validity values of a second set of a plurality of data samples used for the model training, the validity values indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of the model training; Determine the information at least partially based on the second request; and Send a second report to the second device, the second report indicating the information associated with the first set of multiple data samples.

18. The second device according to claim 17, wherein determining the information comprises: generating a third request for the information based on the second request; sending the third request for the information to a third device, the third device being a producer for management data analysis (MDA); and receiving the information from the third device.

19. The second device according to claim 18, wherein the second request includes at least one parameter for determining the first set of multiple data samples, the at least one parameter including at least one of the following: a validity threshold, version information of a model, or specific use cases of one or more models.

20. The second device according to claim 19, wherein determining the information comprises: determining the first set of multiple data samples based on the at least one parameter, and determining the information based on context information of the first set of multiple data samples.

21. The second device according to claim 17, wherein the information includes at least one of the following: a combination of data characteristics corresponding to the first set of multiple data samples, identification information of at least one network object corresponding to the first set of multiple data samples, geographical location information of the at least one network object, network area information corresponding to the first set of multiple data samples, or acquisition time information corresponding to the first set of multiple data samples.

22. The first device according to any one of claims 17 to 21, wherein, the first device is a producer for model training and the second device is a consumer for model training, the first device is a producer for training data analysis and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA) and the second device is a consumer for MDA.

23. A second device, comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the second device to at least perform: sending a second request to a first device, the second request being for information associated with a first set of multiple data samples for model training, the first set of multiple validity values of the first set of multiple data samples being higher than the second set of multiple validity values of a second set of multiple data samples for the model training, the validity values indicating the degree of usefulness of at least one data sample or at least one data feature for the performance of the model training; and receiving a second report from the first device, the second report indicating the information associated with the first set of multiple data samples.

24. The second device according to claim 23, wherein the second request includes at least one parameter for determining the first set of multiple data samples, the at least one parameter including at least one of the following: a validity threshold, version information of one or more models, or specific use cases of one or more models.

25. The second device according to claim 23, wherein the information includes at least one of the following: A combination of data characteristics corresponding to the first set of multiple data samples, Identification information of at least one network object corresponding to the first set of multiple data samples, Geographical location information of the at least one network object, Network area information corresponding to the first set of multiple data samples, or Collection time information corresponding to the first set of multiple data samples.

26. The second device according to any one of claims 23 to 25, wherein, the first device is a producer for model training and the second device is a consumer for model training, the first device is a producer for training data analysis and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA) and the second device is a consumer for MDA.

27. A method, comprising: Receiving, at a first device, a first request from a second device, the first request for querying validity information of a plurality of data samples for model training at the first device; Determining validity information of at least one data sample among the plurality of data samples for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training; and Sending a first report indicating the validity information to the second device.

28. The method according to claim 27, further comprising: Configuring at least one parameter for the first device, the at least one parameter including at least one of the following: A first indication for enabling or disabling the function of sending the first report indicating the validity information, or At least one attribute parameter for determining the validity information.

29. The method according to claim 28, further comprising: Receiving a configuration indicating the at least one parameter from the second device.

30. The method according to any one of claims 27 to 29, wherein the validity information is specific to at least one of the following: Inference type, Training data source, Data features for the model training, or A set of data features for the model training.

31. The method according to any one of claims 27 to 30, wherein the first report includes at least one of the following: At least one corresponding validity value of the at least one data sample, or Storage information of the first report.

32. The method according to any one of claims 27 to 31, wherein the model training at the first device is triggered by receiving the first request.

33. The method according to claim 27, wherein the first report includes a third indication indicating one of the following: The validity information requested by the second device has been successfully determined; The validity information requested by the second device has been partially successfully determined; The validity information requested by the second device has not been determined; or Determining the validity information requested by the second device is not supported.

34. The method according to any one of claims 27 to 33, wherein the first device is a producer for model training, and the second device is a consumer for model training.

35. A method, comprising: sending, at a second device, a first request to a first device, the first request querying validity information of a plurality of data samples for model training at the first device; and receiving, at the second device, a first report indicating the validity information of at least one data sample among the plurality of data samples used for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training, the validity information being determined by the first device during performing the model training by using the plurality of data samples.

36. The method according to claim 35, further comprising: sending a configuration indicating the at least one parameter, the at least one parameter including at least one of the following: a first indication for enabling or disabling a function of sending the first report indicating the validity information, or at least one attribute parameter for determining the validity information.

37. The method according to claim 35 or 36, wherein the first report includes at least one of the following: a validity value of the corresponding at least one data sample, or storage information of the first report.

38. The method according to any one of claims 35 to 37, wherein the validity information is specific to at least one of the following: inference type, training data source, data features for the model training, or a set of data features for model training.

39. The method according to any one of claims 35 to 38, wherein the model training at the first device is triggered by receiving the first request.

40. The method according to claim 35, wherein the first report includes a third indication of at least one of the following: the validity information requested by the second device has been successfully determined, the validity information requested by the second device has been partially successfully determined, the validity information requested by the second device has not been determined, or determining the validity information requested by the second device is not supported.

41. The method according to claim 39, further comprising: configuring the second device to enable the second device to send the first request including a second indication.

42. The method according to any one of claims 35 to 40, wherein the first device is a producer for model training, and the second device is a consumer for model training.

43. A method, comprising: Receive, at a first device, a second request from a second device, the second request being for information associated with a first set of a plurality of data samples for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for the model training, the validity values indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of the model training; Determine the information at least in part based on the second request; And Send a second report to the second device, the second report indicating the information associated with the first set of a plurality of data samples.

44. The method according to claim 43, wherein determining the information Comprises: Based on the second request, generate a third request for the information; Send the third request for the information to a third device, the third device being a producer for managed data analytics (MDA); And Receive the information from the third device.

45. The method according to claim 44, wherein the second request includes at least one parameter for determining the first set of a plurality of data samples, the at least one parameter including at least one of the following: A validity threshold, Version information of the model, or A specific use case of one or more models.

46. The method according to claim 45, wherein determining the information Comprises: Determine the first set of a plurality of data samples based on the at least one parameter, and Determine the information based on context information of the first set of a plurality of data samples.

47. The method according to claim 43, wherein the information includes at least one of the following: A combination of data features corresponding to the first set of a plurality of data samples, Identification information of at least one network object corresponding to the first set of a plurality of data samples, Geographical location information of the at least one network object, Network area information corresponding to the first set of a plurality of data samples, or Collection time information corresponding to the first set of a plurality of data samples.

48. The method according to any one of claims 43 to 47, Wherein, The first device is a producer for model training, and the second device is a consumer for model training; The first device is a producer for training data analytics, and the second device is a consumer for training data analytics, or The first device is a producer for managed data analytics (MDA), and the second device is a consumer for MDA.

49. A method, Comprises: At a second device, send a second request to a first device, the second request being for information associated with a first set of a plurality of data samples for model training, the first set of a plurality of validity values of the first set of a plurality of data samples being higher than a second set of a plurality of validity values of a second set of a plurality of data samples for the model training, the validity values indicating a degree of usefulness of at least one data sample or at least one data feature for the performance of the model training; And Receive a second report from the first device, the second report indicating the information associated with the first set of a plurality of data samples.

50. The method according to claim 49, wherein the second request includes determining at least one parameter of the first set of multiple data samples, and the at least one parameter includes at least one of the following: a validity threshold, version information of one or more models, or specific use cases of one or more models.

51. The method according to claim 49, wherein the information includes at least one of the following: a combination of data characteristics corresponding to the first set of multiple data samples, identification information of at least one network object corresponding to the first set of multiple data samples, geographical location information of the at least one network object, network area information corresponding to the first set of multiple data samples, or acquisition time information corresponding to the first set of multiple data samples.

52. The method according to any one of claims 49 to 51, wherein the first device is a producer for model training and the second device is a consumer for model training; the first device is a producer for training data analysis and the second device is a consumer for training data analysis, or the first device is a producer for management data analysis (MDA) and the second device is a consumer for MDA.

53. A first apparatus, comprising: means for receiving a first request from a second device, the first request for querying validity information of multiple data samples for model training at the first device; means for determining validity information of at least one data sample among the multiple data samples for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training; and means for sending a first report indicating the validity information to the second device.

54. A second apparatus, comprising: means for sending a first request to a first device, the first request for querying validity information of multiple data samples for model training at the first device; and means for receiving a first report from the second device, the first report indicating validity information of at least one data sample among the multiple data samples used for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training, the validity information being determined by the first device during performing the model training by using the multiple data samples.

55. A first apparatus, comprising: means for receiving a second request from a second device, the second request being directed to information associated with a first set of multiple data samples for model training, the first set of multiple validity values of the first set of multiple data samples being higher than the second set of multiple validity values of a second set of multiple data samples for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training; means for determining the information at least partially based on the second request; and A component for sending a second report to the second device, the second report indicating the information associated with the first set of multiple data samples.

56. A second device, comprising: A component for sending a second request to a first device, the second request being for information associated with a first set of multiple data samples for model training, the first set of multiple validity values of the first set of multiple data samples being higher than the second set of multiple validity values of a second set of multiple data samples for the model training, the validity information indicating the usefulness of at least one data sample or at least one data feature for the performance of the model training; and A component for receiving a second report from the first device, the second report indicating the information associated with the first set of multiple data samples.

57. A computer-readable medium comprising instructions stored thereon for causing a device to perform at least the method according to any one of claims 27 to 34 or the method according to any one of claims 35 to 42 or the method according to any one of claims 43 to 48 or the method according to any one of claims 49 to 52.