Reuse of data for training machine learning models

By implementing devices for identifying and reusing existing stored data in the device, the problem of limited storage capacity of the device is solved, the storage demand for new user data is reduced, and data utilization is improved.

CN120011800APending Publication Date: 2025-05-16NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411609175.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-11-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

As more and more software applications collect and store user data for training machine learning models, storage requirements for devices increase, especially on devices with limited storage capacity, such as smartphones, resulting in storage difficulties.

Method used

By implementing a device in the device, the device includes means for receiving new user data requests, means for identifying existing stored data suitable for training machine learning models based on the ontology, and means for providing access to identified existing stored data, thereby reducing storage requirements for new user data.

Benefits of technology

By reusing existing stored data, the solution reduces the storage demand for new user data and reduces the requirements for device storage capacity, especially for devices with limited storage capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011800A_ABST
    Figure CN120011800A_ABST
Patent Text Reader

Abstract

Example embodiments may relate to systems, methods, and / or computer programs for reusing data for training machine learning models. In an example, an apparatus includes means to receive collected new user data for training a machine learning model associated with an application. The apparatus may also include means for identifying existing stored data suitable for training the machine learning model based on the ontology. The apparatus may also include a component to provide access to the identified existing stored data in response to the identification data being suitable for training the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments may relate to systems, methods, and / or computer programs for reusing data for training machine learning models. Background Art

[0002] Software applications are increasingly using machine learning models. These models may need to be fine-tuned for specific users to improve performance. For example, medical-related applications may need to obtain user data to establish a baseline for the user, or an artificial intelligence (AI) assistant may collect user voice data to adapt the user's voice and speaking style to improve speech recognition performance for the user. As many applications collect and store user data for training their respective machine learning models, this increases the burden of storage requirements on devices. The increased storage requirements may cause difficulties for devices with limited storage capacity, such as smartphones. Summary of the invention

[0003] The scope of protection considered for various embodiments of the present invention is defined by the independent claims. Embodiments and features described in this specification that do not fall within the scope of the independent claims, if any, are to be interpreted as examples that aid in understanding the various embodiments of the present invention.

[0004] According to a first aspect, an apparatus is described, comprising: a component for receiving a request to collect new user data, the new user data being used to train a machine learning model associated with an application; a component for identifying, based on an ontology, existing stored data suitable for training the machine learning model; and a component for providing access to the identified existing stored data in response to identifying that the data is suitable for training the machine learning model.

[0005] In some examples, the request may also include a modality of the new user data. In some examples, the means for identifying existing stored data may also include: means for identifying the existing stored data based on the modality.

[0006] In some examples, the request may also include data indicating one or more tags for training the machine learning model. In some examples, the component for identifying existing stored data may also include: a component for determining one or more items related to the one or more tags based on the ontology; and a component for identifying existing stored data having metadata, the metadata including at least one of the one or more items related.

[0007] The apparatus may also include means for processing the identified existing stored data to enhance the suitability of the existing stored data for training the machine learning model.In some examples, the means for providing access to the identified existing stored data may include means for providing access to a processed version of the existing stored data.

[0008] In some examples, means for processing the identified existing stored data may include means for applying signal processing to the identified existing stored data.

[0009] The apparatus may also include: a component for generating tags for the identified existing stored data for training a machine learning model associated with the application.

[0010] In some examples, the component for generating the tag may include: a component for generating one or more hidden tasks from the identified existing stored data; and a component for labeling the identified existing stored data based on the one or more hidden tasks.

[0011] In some examples, a component for generating one or more hidden tasks from identified existing stored data may include a component for generating one or more hidden tasks from identified existing stored data based on an optimized labeling function and multiple machine learning models configured to perform the hidden tasks, wherein the optimization is based on a consistency score between the outputs of the multiple machine learning models when performing the hidden tasks.

[0012] In some examples, multiple machine learning models have the same architecture but different starting parameter values.

[0013] In some examples, the hidden task is a random classification task.

[0014] In some examples, means for labeling the identified existing stored data based on the one or more hidden tasks may include means for labeling the existing stored data based on an active learning model.

[0015] In some examples, the components for marking identified existing stored data based on one or more hidden tasks may include at least one of the following: components for providing a subset of existing stored data for one or more hidden tasks to a user for manual marking; components for receiving manual markings of a subset of existing stored data from a user; or components for automatically marking the remaining existing stored data based on the received manual markings.

[0016] The apparatus may also include means for modifying metadata of the identified existing stored data to indicate reusability of the data for training a machine learning model.

[0017] In some examples, components for modifying metadata of identified existing stored data to indicate reusability of the data for training a machine learning model may include components for modifying the metadata to include labels determined for the data.

[0018] The apparatus may also include components for training a machine learning model based on the identified existing stored data.

[0019] In some examples, the ontology may have been generated based on co-occurrences of tags in data items of one or more data sets.In some examples, the apparatus may also include a component for generating the ontology.

[0020] According to a second aspect, a method is described, comprising: receiving, by a device, a request from an application to collect new user data, the new user data being used to train a machine learning model associated with the application; identifying, by the device based on an ontology, existing stored data suitable for training the machine learning model; and providing, by the device, access to the identified existing stored data in response to the request of the application.

[0021] In some examples, the request may further include a modality of the user data. In some examples, identifying the existing stored data may further include: identifying the existing stored data based on the modality.

[0022] In some examples, the request may also include: data indicating one or more tags for training the machine learning model. In some examples, identifying existing stored data may also include: determining one or more items related to the one or more tags based on the ontology; and identifying existing stored data having metadata including at least one of the related one or more items.

[0023] The method may also include processing the identified existing stored data to enhance the suitability of the existing stored data for training the machine learning model.In some examples, providing access to the identified existing stored data may include providing access to a processed version of the existing stored data.

[0024] In some examples, processing the identified existing stored data may include applying signal processing to the identified existing stored data.

[0025] The method may also include generating tags for the identified existing stored data for use in training a machine learning model associated with the application.

[0026] In some examples, generating the tags may include: generating one or more hidden tasks from the identified existing stored data; and labeling the identified existing stored data based on the one or more hidden tasks.

[0027] In some examples, generating one or more hidden tasks from identified existing stored data can include generating one or more hidden tasks from identified existing stored data based on an optimization labeling function and multiple machine learning models configured to perform the hidden tasks, wherein the optimization is based on a consistency score between outputs of the multiple machine learning models when performing the hidden tasks.

[0028] In some examples, multiple machine learning models have the same architecture but different starting parameter values.

[0029] In some examples, the hidden task is a random classification task.

[0030] In some examples, labeling the identified existing stored data based on the one or more hidden tasks may include labeling the existing stored data based on an active learning model.

[0031] In some examples, marking identified existing stored data based on one or more hidden tasks can include at least one of: providing a subset of existing stored data for one or more hidden tasks to a user for manual marking; receiving manual markings of a subset of existing stored data from a user; or automatically marking the remaining existing stored data based on the received manual markings.

[0032] The method may also include modifying metadata of the identified existing stored data to indicate reusability of the data for training the machine learning model.

[0033] In some examples, modifying metadata of identified existing stored data to indicate reusability of the data for training a machine learning model can include modifying the metadata to include tags determined for the data.

[0034] The method may also include training a machine learning model based on the identified existing stored data.

[0035] In some examples, the ontology can be generated based on co-occurrence of tags in data items of one or more data sets.In some examples, the method can also include generating the ontology.

[0036] According to a third aspect, there is provided a computer program product comprising a set of instructions which, when executed on an apparatus, are configured to cause the apparatus to perform a method as defined in any preceding method.

[0037] According to a fourth aspect, there is provided a (non-transitory) computer-readable medium comprising program instructions which, when executed by a device, cause the device to perform at least the following operations: the device receives a request from an application to collect new user data, the new user data being used to train a machine learning model associated with the application; the device identifies, based on an ontology, existing stored data suitable for training the machine learning model; and, in response to the request of the application, the device provides access to the identified existing stored data.

[0038] The program instructions of the third aspect may also perform operations according to any of the aforementioned method definitions of the second aspect.

[0039] According to a fifth aspect, there is provided a device comprising: one or more processors; and at least one memory storing instructions which, when executed by the one or more processors, cause the device to at least receive a request from an application to collect new user data, the new user data being used to train a machine learning model associated with the application; identify existing stored data suitable for training the machine learning model based on an ontology; and provide access to the identified existing stored data in response to a request from the application.

[0040] The computer program code of the fifth aspect may also perform operations according to any of the aforementioned method definitions of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Example embodiments will now be described by way of non-limiting examples with reference to the accompanying drawings, in which:

[0042] Figure 1 is a block diagram of an apparatus that may be useful in understanding example embodiments;

[0043] Figure 2 is a flowchart illustrating operations according to one or more example embodiments;

[0044] Figure 3 is a block diagram illustrating operation according to one or more example embodiments;

[0045] Figure 4 is a schematic diagram of a portion of an ontology that may be useful in understanding an example embodiment;

[0046] Figure 5 is a flowchart illustrating operations according to one or more example embodiments;

[0047] Figure 6 is a flowchart illustrating operations according to one or more example embodiments;

[0048] Figure 7 is a flowchart illustrating operations according to one or more example embodiments;

[0049] Figure 8 is a flowchart illustrating operations according to one or more example embodiments;

[0050] Fig. 9 is a block diagram of a federated learning system that may be useful in understanding example embodiments;

[0051] Fig.10 is a block diagram of an apparatus that may be configured according to one or more example embodiments;

[0052] Fig.11 is an illustration of an example medium that may store instructions that, when executed by an apparatus, may cause the apparatus to perform one or more example embodiments. DETAILED DESCRIPTION

[0053] Example embodiments relate to an apparatus, method, and computer program involving reusing existing stored data for training a machine learning (ML) model.

[0054] The term "label" as used herein may refer to any form of descriptor or label used to describe a particular data item. The term "label" is not intended to imply a limitation to supervised learning, nor is it limited to labels used in the context of supervised learning.

[0055] The term "request" as used herein may include a single data transfer event or multiple data transfer events. A request may include all data transfers required for a device to perform the requested operation, and such transfers may occur in the form of multiple transfers over a period of time.

[0056] Software applications are increasingly using machine learning models. These models may need to be fine-tuned for specific users to improve performance. For example, medical-related applications may need to obtain user data to establish a baseline for the user, or an AI (artificial intelligence) assistant may collect user voice data to adapt the user's voice and speaking style to improve the voice recognition performance for the user. As many applications collect and store user data for training their respective machine learning models, data previously collected for one application may be reused to train the machine learning model of another application. Example embodiments provide components for identifying existing stored data suitable for training machine learning models to avoid the need to collect and store more user data. In this way, storage requirements are reduced. This is particularly useful for devices with limited storage capacity, such as smartphones and other portable devices.

[0057] Figure 1is a block diagram illustrating an apparatus, in this case device 100, which may be helpful in understanding the example embodiments. Device 100 includes one or more storage media 101 and one or more sensors that can be used to collect user data. For example, device 100 includes: a microphone 102, a camera 103, a GNSS (Global Navigation Satellite System) sensor, such as a GPS (Global Positioning System) sensor 104, and / or an accelerometer 105. It is understood that the device may include other sensors, or certain sensors may be omitted as needed by those skilled in the art. For example, other potential sensors include at least one of the following sensors: a temperature sensor, a humidity sensor, a proximity sensor, a heart rate sensor, a blood pressure sensor, a motion sensor, an inertial measurement sensor, or a position sensor, etc.

[0058] The device 100 also includes an operating system 106, which can be configured to control access to the device's hardware resources, such as one or more storage media 101 and one or more sensors. The device 100 may also include one or more existing applications. For example, the device 100 includes two existing applications 107a and 107b, but it will be appreciated that the device may include a greater or lesser number of applications. The existing applications 107a and 107b each include machine learning models 108a and 108b. Both existing applications 107a and 107b have previously collected data, or have access to previously collected data, for training their respective machine learning models. Previously collected data 111 is stored on the storage medium 101 of the device 100.

[0059] Device 100 also includes a "new" application 109. For example, new application 109 can be installed by a user of device 100. New application 109 also includes a machine learning model 110 that requires training on user data. New application 109 is configured to request collection of new user data for training its machine learning model 110. For example, the request can include a request to access one or more sensors of device 100. The request can be sent to and processed by operating system 106.

[0060] The device 100 is configured to receive a request from a new application 109. However, rather than immediately approving the request, the device 100 is configured to determine whether any existing stored data 111 previously collected by existing applications 107a, 107b is suitable for training a machine learning model 110 associated with the new application 109. In this regard, the device 100 is configured to identify, for example, existing stored data 111 suitable for training a machine learning model 110 based on an ontology 112. This process will be described in more detail below. The existing stored data 111 may also include existing data stored on the device 100 that was not acquired for the specific purpose of training a machine learning model. In addition, the existing stored data 111 need not be collected for the same precise task performed by the machine learning model 110 of the new application 109.

[0061] If suitable existing data is identified, the device 100 is configured to provide access to the identified existing data in response to the request of the new application 109. The new application 109 can then use the identified existing data to continue training the machine learning model 110. Alternatively, if no suitable existing data is identified, the device 100 can be configured to allow the request of the new application 109 to collect new user data, and can allow access to one or more sensors to do so.

[0062] However, the user may opt-out of any data collection process and search of stored data, or alternatively, active permission may be sought from the user before any such activity is performed by the device.

[0063] Example applications may include applications involving healthcare, such as cough analysis to predict potential diseases or breathing analysis to determine any anomalies from captured audio, visual, and / or other sensor signals. Other example applications include speech recognition, for example, recognizing voice commands to control / operate a device. Another example application may include facial recognition, for example, to prove that a user has allowed access to a device. Embodiments may also be applied in industrial environments. For example, applications may include detecting machine failures and environmental hazard monitoring from audio, visual, and / or other sensor inputs. More generally, example applications may involve any classification or detection task, such as classifying input received from one or more sensors of device 100 into one or more appropriate categories, or locating a specific entity of interest within an input signal.

[0064] Device 100 may be any suitable form of computing device. For example, device 100 may be a wireless communication device, a smartphone, a desktop computer, a laptop computer, a tablet computer, a smart watch, a smart ring, a digital assistant, an AR (augmented reality) headset, a VR (virtual reality) headset, a television, an over-the-top (OTT) device, a vehicle, or some form of Internet of Things (IoT) device, or any combination thereof.

[0065] Although Figure 1 Not shown, the device 100 includes one or more processors, such as a central processing unit (CPU), one or more memories, such as main memory or random access memory (RAM), and optionally one or more additional hardware units, such as one or more floating point units (FPUs), graphics processing units (GPUs), tensor processing units (TPUs), and digital signal processors (DSPs).

[0066] Figure 2 is a flow diagram of operations 200 that may be performed, for example, by an apparatus, according to one or more example embodiments. For example, the operations may be performed by Figure 1 The operation may be performed by the device 100 of the present invention. The operation may be a processing operation performed by hardware, software, firmware, or a combination thereof. The order shown does not necessarily represent the order of processing. The device may include or at least include one or more components for performing the operation, wherein the component may include one or more processors, controllers, or circuits, which, when executing computer-readable instructions, can cause it to perform the operation.

[0067] The first operation 201 includes receiving a request from an application to collect new user data for training a machine learning model associated with the application. For example, the application may generate a request when it is newly installed on a device.

[0068] The second operation 202 includes identifying existing stored data suitable for training the machine learning model based on the ontology, for example, in the device 100. As discussed above, the existing stored data may have been previously collected by one or more existing applications for the purpose of training the machine learning model associated with these one or more existing applications. The existing machine learning models do not need to perform the same tasks as the machine learning model associated with the requesting application. Additionally, or alternatively, the existing stored data may include existing stored data that was not acquired for the specific purpose of training one or more machine learning models. The identification of suitable existing data will be described in more detail below.

[0069] The third operation 203 includes: in response to the application's request, providing access to the identified existing stored data. This may be an alternative to allowing the application to collect new user data. In this way, the existing stored data may be used to train the machine learning model instead of collecting and storing new user data, which may reduce the amount of storage available to the device and other applications.

[0070] Figure 3 is a block diagram of an example pipeline or process 300 for processing existing stored data for training a machine learning model. According to one or more example embodiments, the operations of the pipeline 300 may be performed, for example, by a device. For example, the operations may be performed by Figure 1 The operation may be performed by the device 100 of the present invention. The operation may be a processing operation performed by hardware, software, firmware, or a combination thereof. The order shown does not necessarily represent the order of processing. The device may include or at least include one or more components for performing the operation, wherein the component may include one or more processors, circuits, or controllers, which, when executing computer-readable instructions, can cause the operation to be performed.

[0071] The operations of pipeline 300 can be performed in response to receiving a request to collect data (such as new data) for training a machine learning model as discussed above. Pipeline 300 includes a first data identification stage 301. In this stage, as discussed in further detail below, stored data that is potentially suitable for training a machine learning model is identified based on an ontology.

[0072] The second stage 302 of the pipeline 300 includes data enhancement. In this stage, the identified stored data can be processed to enhance the suitability of the data for training the machine learning model. Example operations in this second stage will be described in more detail below.

[0073] The third stage 303 of the pipeline 300 includes data labeling. In this stage, the identified stored data can be labeled according to the requirements of the machine learning model. Example operations in this third stage will be described in more detail below.

[0074] It is understood that the data augmentation and data labeling stages may be optional and may be omitted depending on the nature of the machine learning model. For example, an unsupervised learning model may not require data labeling, or if the data is collected for the same or similar purpose, augmentation may not be required.

[0075] Example operations for identifying existing stored data suitable for training a machine learning model will now be described in more detail. As discussed above, the existing stored data is identified based on an ontology. Generally, an ontology is a graph showing the relationships between concepts and entities. An ontology can be generated or received using a publicly available dataset, such as "AudioSet: An Ontology and Manually Labeled Dataset for Audio Events" by Gemmeke, Jort F. et al. In 2017 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), IEEE 2017, pp. 776-780. Generally, the AudioSet dataset includes labeled sound events for Youtube videos.

[0076] Figure 4A portion of an ontology 400 generated from an audio collection dataset is shown. The ontology 400 includes a collection of nodes, each node representing a concept or action, such as music, speech, crying, coughing, etc. (hereinafter referred to as "items"). The ontology 400 also includes a collection of edges connecting certain nodes in the graph to represent the relationship between the concepts / actions / items represented by the nodes. Each edge is also associated with a probability indicating the likelihood of the node connection. The probability can be generated based on the co-occurrence of the labels used for each data item in the dataset. The probability can provide an indication of the likelihood that a data item with a particular node label is related to a task with a connecting node label.

[0077] Figure 5 is a flow diagram illustrating example operations for identifying existing stored data suitable for training a machine learning model based on an ontology.

[0078] The first operation 501 may include obtaining one or more tags associated with training the machine learning model. Such tags may be included in a request to collect new user data through an associated application. The one or more tags may also be provided in a separate transmission, but for purposes of this disclosure, the separate transmissions will be considered part of one related request. Additionally, or alternatively, any metadata relating to the application itself, such as a title, description, or associated keywords, may also be used.

[0079] The second operation 502 may include using the ontology to determine one or more items related to the one or more tags. This may include searching the ontology for each of the one or more tags and providing items represented by connecting nodes. These items may be items with direct connections to the tag nodes, or may be items with indirect connections with a threshold maximum degree. When the ontology includes a probability indicating the likelihood that two nodes are related, a threshold may be used to determine whether to provide the item as a related item.

[0080] The third operation 503 can include identifying existing stored data with metadata, and the metadata includes at least one item in the relevant one or more items, or at least one tag in the one or more tags. The metadata can include the file name, and / or the file descriptor, and / or any other appropriate metadata. In this regard, the previously collected data can be stored in the file descriptor field of the data file with appropriate tags to promote reusability and help search for related data.

[0081] Additionally or alternatively to using tags, existing stored data can be identified based on the modality of the data required for training the machine learning model, for example, the type of data or the type of data content. For example, the modality can be audio data, or the modality can be a more specific form, such as voice data or music data. In another example, the modality can be image data or video data or audio-visual data. In another example, the modality can be location data or other sensor data. The required modality can be included in the request of the application. Therefore, in some embodiments, existing stored data is also identified based on modality, and additional filters are provided for identifying existing data suitable for training the machine learning model. In some embodiments, data with the required modality can be first identified before processing using the ontology. Alternatively, the modality can be used to filter the data identified using the ontology.

[0082] review Figure 3 , the data enhancement stage 302 will now be described in more detail. As discussed above, the identified existing stored data can be processed to enhance the suitability of the data for training the machine learning model. In some embodiments, the processing includes applying signal processing to the identified existing stored data. Typically, devices such as smartphones have hardware processing units, such as digital signal processors, that are capable of efficiently performing signal processing operations. By using signal processing techniques, such hardware can be exploited to increase efficiency and reduce latency.

[0083] An example of specific enhancement will now be described using time series data. In one example, if the desired data is speech data, signal processing can be applied to the identified audio data to enhance components associated with human speech in the audio signal and / or weaken components associated with background noise or other sounds. In another example, if the application is a medical application for analyzing breathing, signal processing can be used to enhance frequencies associated with breathing. In another example, the identified existing stored data may include metadata indicating the presence of irrelevant sounds. Signal processing methods can be used to filter or weaken such sounds. Signal processing methods can also be applied to any other time series data, such as sensor data. Example signal processing operations that can be performed include Fourier transforms, wavelet analysis, change point detection, filtering, statistical enhancement, and the like.

[0084] For image data, example operations may also include cropping and scaling or any other appropriate image transformation. In one example, the purpose of the machine learning model may be to recognize a user's face. A general face detection model may be used to identify potential faces in existing image data, and data cropping may be performed to remove non-facial elements in the image. More generally, an object detector may be used to detect potential objects in an image, and cropping may be performed to remove the background.

[0085] Although some specific examples and modalities have been discussed above, it will be appreciated that the embodiments are not limited to these examples and modalities.

[0086] Data enhancement operations may depend on the specific application environment and the specific content of the data. Specific enhancement operations may be predetermined and may be associated with certain items in the ontology or associated with certain applications. The list of enhancement operations may be updated from time to time to cover new applications. For example, updates may be provided through operating system updates. Additionally or alternatively, the application may also specify the enhancement operations to be performed, for example in the application's request.

[0087] In response to an application's request to collect user data, access may be provided to a processed / enhanced version of the identified existing stored data.

[0088] review Figure 3 , the data labeling phase 303 will now be described in more detail. As discussed above, at this phase, the identified stored data may be labeled according to the requirements of the machine learning model, which may be included in the application's request. Figure 6 is a flow chart 600 illustrating example operations at this stage.

[0089] The first operation 601 may include a hidden task identification from the identified existing stored data. Generally, a task may be viewed as some allocation or partitioning of data according to one or more categories. The identified existing stored data may be storage metadata including one or more labels. These labels provide known tasks or data partitions. The hidden task may be a data partition where the differences and labels may not be initially known. Thus, the goal of the hidden task identification may be to determine one or more partitions of the identified existing stored data. Figure 7 Provides additional details about hidden task identifiers.

[0090] The second operation 602 may include labeling the identified existing stored data based on the one or more hidden tasks identified from the first operation 601. As discussed, the one or more hidden tasks may provide one or more partitions of the identified existing stored data. Such data partitions may be used to initialize a process for labeling existing data used to train a machine learning model associated with the new application. For example, labeling may be performed based on an active learning model. Figure 8 Provide additional details.

[0091] Reference now Figure 7 , an example operation 700 for hiding task identification will be described. As an example only, the hiding task is considered to be binary classification of data into a first category and a second category. However, it should be understood that other tasks deemed appropriate by those skilled in the art may be used.

[0092] The first operation 701 may include labeling the data using a labeling function to generate an initial labeling of the data. That is, the function may be used to assign each data item to the first category or the second category. The labeling function may include an adjustable parameter, and the value of the adjustable parameter may be randomly initialized. Therefore, the initial labeling of the data may be a random assignment, and the hidden task may be a random classification task.

[0093] The second operation 702 may include initializing multiple machine learning models configured to perform the hidden task. For example, if the hidden task is a binary classification, the machine learning model is configured to provide an output for classifying the input into a first class or a second class. Multiple machine learning models may have the same architecture, but with different starting parameter values. The starting parameter values ​​may be determined randomly. For example, the architecture of a machine learning model in a machine learning model associated with an existing application may be used. In one embodiment, two instances of the machine learning model may be created, but with different parameter initializations.

[0094] The third operation 703 may include optimizing the labeling function and multiple machine learning models based on a consistency score between the outputs of the multiple machine learning models when performing the hiding task. For example, the consistency score may be based on the number of identical outputs produced by multiple neural networks. When the consistency score is low, the label assignments produced by the labeling function are likely not to represent any identifiable differences in the data and may still be random in nature. When the consistency score is high, the labeling function is likely to have produced data partitions corresponding to certain identifiable differences in the data. Therefore, by optimizing based on the consistency score, data partitions corresponding to potential identifiable differences in the data can be generated. The process can be repeated to identify additional potential data partitions / hiding tasks, and these can be used to initialize the process for labeling existing data for training machine learning models associated with new applications.

[0095] Optimization based on consistency scores can be performed using any appropriate optimization method. For example, when multiple machine learning models are neural networks, stochastic gradient descent and back propagation can be used. The optimization can alternate between updating adjustable parameters of the labeling function to optimize the data partitioning and updating adjustable parameters of multiple machine learning models to perform hidden tasks / classify the data. The consistency score between the two machine learning models can be calculated based on the difference or distance calculated between the output of the first model and the output of the second model. This can be combined with any appropriate supervised loss function, such as mean squared error or cross entropy error, to obtain an overall loss function for optimization. The consistency score can be calculated using a retained validation set or test set.

[0096] Reference now Figure 8, an example operation 800 for marking identified existing stored data based on one or more hidden tasks will now be described. The first operation 801 may include providing a subset of existing stored data for one or more hidden tasks to a user for manual labeling. For example, when the task is a binary classification, a subset of data for the first class and a subset of data for the second class may be selected and provided to the user for manual labeling. The subset of data may be selected according to an active learning model, which may determine which data items are most beneficial to receive labels for. In some embodiments, the active learning model may include a graph having nodes representing data items and their labels, and weighted edges representing similarity measures between nodes. It may be understood by those skilled in the art that other forms of active learning models may be used as long as they are deemed appropriate. The data partition provided by the hidden task may be used to initialize the active learning model. Depending on the circumstances, there may be multiple active learning models that appropriately correspond to each item in the hidden task / data partition. An uncertainty measure of whether a data item is correctly labeled may be used to determine which data items are included in a subset for presentation to the user. For example, the data item with the lowest confidence may be provided to the user for manual labeling, or the data item that generates the most information gain may be provided to the user for manual labeling.

[0097] The second operation 802 may include receiving manual labels for a subset of the data provided in the first operation 801. The third operation 803 may include automatically labeling the remaining identified existing stored data based on the received manual labels. For example, the received manual labels may be used to update the active learning model. The active learning model may use a graph diffusion process to propagate the labels through the graph to update the labels of each non-manually labeled data item. For example, the active learning model may be repeated. Figure 8 The process continues until all labels have a threshold confidence, or until a fixed number of iterations have been performed to provide the final output labeling. In one example, manually labeling approximately 20% of the data in each category is sufficient to achieve accurate labeling for the remaining data using active learning.

[0098] In some embodiments, the selected data may be presented to the user at a specific appropriate moment. These may be determined based on an empirical sampling method to avoid overwhelming the user and to promote user cooperation with the labeling process. For example, when the user is not engaged in physical activity, and / or when the user is not engaged in intensive tasks on the device. The presentation may also be separated in time. The user may be presented with an option to label the data based on the labels required for training the machine learning model associated with the new application. When the user cannot identify any data partitioning provided by the hidden task corresponding to the labels required for training the machine learning model, it may be considered that there is no suitable storage data for training the machine learning model. In these cases, permission may be granted to the new application to collect new user data. When the user is able to identify the data corresponding to the required labels, the remaining existing data may be automatically labeled as described above and provided in response to the application's request. In some embodiments, the metadata of the identified existing storage data is modified to include a label, for example, the file descriptor field of the data item may be modified to include the label of the data item. In addition, this promotes the reusability of the data and facilitates subsequent searches for data for other applications that have been installed.

[0099] Reference now Fig. 9 , a block diagram of an example federated learning system is shown. Any of the above example embodiments can be used in a federated learning environment. In general, federated learning is a method for training machine learning models in a distributed manner. Machine learning models can be trained using data stored locally at multiple participating devices without sending the data from the devices. Therefore, federated learning provides anonymity and security for participating devices.

[0100] System 900 may include a server 901 connected to a plurality of client devices 903a...n via a network 902. Client devices 903a...n may be Figure 1 1. The server 901 may include any suitable form of computing device or system capable of communicating with multiple client devices 903a…n. The network 902 may include any form of data network, such as a wired or wireless network, including but not limited to a radio access network (RAN) that complies with 3G, 4G, 5G (generation) protocols or any future RAN protocol. Alternatively, the network 103 may include a Wi-Fi (Wireless Fidelity) network (IEEE 802.11), or a similar network, other forms of local area networks or wide area networks (LAN or WAN), or a short-range wireless network utilizing protocols such as Bluetooth or similar.

[0101] Server 901 may be configured to centrally maintain machine learning models. Server 901 may be configured to send a copy of the machine learning model to a client device when requested. Each client device may be configured to train the machine learning model using local data. The client device may be configured to send updates to the machine learning model based on training using local data. Server 901 may be configured to apply updates received from each client device to the centrally maintained machine learning model. Server 901 may be configured to send an updated machine learning model in response to any additional request from the client device. Alternatively, server 901 may be configured to send a copy of the machine learning model accordingly without a request from the client device.

[0102] In some cases, the client devices 903a...n may need to collect data before being able to begin the process of training a machine learning model. Figures 1 to 8 The described method allows a client device to use existing stored data on the device to identify and train machine learning. In some embodiments, this can replace the collection of new data, thereby reducing the necessary storage requirements, or this can enable the training process to begin earlier without having to wait for the data collection process to complete.

[0103] Fig.10 An apparatus according to some example embodiments is shown, which may include a device 100. The apparatus may be configured to perform the operations described herein, such as the operations described with reference to any of the disclosed processes above. The apparatus includes at least one processor 1001 and at least one memory 1002 directly or closely connected to the processor. The memory 1002 includes at least one random access memory (RAM) 1002b and at least one read-only memory (ROM) 1002a. Computer program code (software) 1003 is stored in the ROM 1002a. The apparatus may be connected to a transmitter 1004 (TX) and a receiver 1005 (RX). Optionally, the apparatus may be connected to a user interface 1006 (UI) for instructing the apparatus and / or for outputting data. At least one processor 1001, at least one memory 1002 and computer program code 1003 are arranged such that the apparatus performs at least a method according to any of the previous processes, such as involving Figure 2 and Figure 3 and Figures 5 to 8 Alternatively or additionally, the apparatus comprises one or more circuits to operate any of the processes disclosed above or any part thereof, as described in reference Figure 2 and Figure 3 and Figures 5 to 8 The features involved are described in the flowchart.

[0104] As used in this application, the term "circuitry" may refer to one or more or all of the following:

[0105] (a) pure hardware circuit implementation (such as implementation with analog and / or digital circuit systems only) and

[0106] (b) a combination of hardware circuitry and software such as (where applicable):

[0107] (ii) a combination of analog and / or (multiple) digital hardware circuits and software / firmware and

[0108] (ii) any portion of hardware processor(s) working in conjunction with software (including digital signal processor(s), software and memory(s) to enable a device, such as a mobile phone or server, to perform various functions) and hardware circuit(s) and / or processor(s), such as microprocessor(s) or portion(s) of microprocessor(s), which requires software (e.g. firmware) for operation, but where software is not required for operation, the software may not be present.

[0109] This definition of circuitry applies to all uses of the term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" also covers hardware circuitry alone, or a processor (or multiple processors), or a portion of a hardware circuitry or processor, and its (or therein) accompanying software and / or firmware implementation. For example and if applicable to a particular claim element, the term "circuitry" also covers a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device.

[0110] Fig.11 A non-transient medium 1100 according to some embodiments is shown. The non-transient medium 1100 is a computer readable storage medium. It can be, for example, a CD (Compact Disc), a DVD (Digital Video Disc), a USB (Universal Serial Bus) memory stick, a Blu-ray Disc, etc. The non-transient medium 1100 stores computer program code, causing the apparatus to perform methods of any of the aforementioned processes, for example, as disclosed in relation to the flow charts and the features involved therein.

[0111] The names of network elements, protocols and methods are based on current standards. In other versions or other technologies, the names of these network elements and / or protocols and / or methods may be different, as long as they provide the corresponding functions. For example, embodiments may be deployed in 2G / 3G / 4G / 5G networks and subsequent generations of 3GPP (3rd Generation Partnership Project), and may also be deployed in non-3GPP radio networks, such as Wi-Fi.

[0112] The memory may be volatile or non-volatile. It may be, for example, RAM, SRAM (static random access memory), flash memory, FPGA (field programmable gate array), RAM block, DVD, CD, USB stick, and Blu-ray disc.

[0113] Unless otherwise stated or otherwise clear from the context, a statement that two entities are different means that they perform different functions. This does not necessarily mean that they are based on different hardware. That is, each entity described in this specification can be based on different hardware, or some or all entities can be based on the same hardware. This does not necessarily mean that they are based on different software. That is, each entity described in this specification can be based on different software, or some or all entities can be based on the same software. Each entity described in this specification can be embodied in the cloud.

[0114] Implementation of any of the blocks, devices, systems, techniques, or methods described above includes implementation as hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof, as non-limiting examples. Some embodiments may be implemented in the cloud.

[0115] It should be understood that what has been described above is what is presently considered to be the preferred embodiment. However, it should be noted that the description of the preferred embodiment is given by way of example only and that various modifications may be made without departing from the scope defined by the appended claims.

[0116] Applicants hereby independently disclose each individual feature described herein and any combination of two or more such features, as long as such feature or combination can be performed based on this specification as a whole and according to the common general knowledge of those skilled in the art, regardless of whether such feature or feature combination solves any problem disclosed herein, and does not limit the scope of the claims. Applicants point out that the disclosed aspects / examples may consist of any such single feature or feature combination. In view of the above description, it will be apparent to those skilled in the art that various modifications may be made within the scope of the present disclosure.

[0117] As used herein, “at least one of: ” and “at least one of ” and similar words, where a list of two or more elements is connected by “and” or “or”, mean at least any one element, or at least any two or more elements, or at least all elements.

[0118] Although the basic novel features applied to the examples thereof have been shown and described and pointed out, it will be understood that various omissions and substitutions and changes in the form and details of the described apparatus and methods may be made by those skilled in the art without departing from the scope of the present disclosure. For example, it is expressly intended that all combinations of those elements and / or method steps that perform substantially the same function in substantially the same manner to achieve the same results are within the scope of the present disclosure. Furthermore, it will be recognized that the structures and / or elements and / or method steps shown and / or described in connection with any disclosed form or example may be incorporated into any other disclosed or described or suggested form or example as a general matter of design choice. Furthermore, in the claims, parts-plus-function clauses are intended to cover the structures described herein as performing the functions described, and include not only structural equivalents but also equivalent structures.

Claims

1. A device for machine learning, comprising: means for receiving a request to collect new user data for use in training a machine learning model associated with an application; means for identifying, based on the ontology, existing stored data suitable for training the machine learning model; as well as Means for providing access to the identified existing stored data in response to identifying that the data is suitable for training the machine learning model.

2. The apparatus of claim 1 , wherein the request further comprises a modality for the new user data, and wherein the means for identifying the existing stored data further comprises: Means for identifying said existing stored data based on said modality.

3. The apparatus according to claim 1 or 2, wherein the request further comprises: data indicating one or more labels for training the machine learning model, and wherein the means for identifying the existing stored data further comprises: means for determining, based on the ontology, one or more terms related to the one or more tags; and Means for identifying said existing stored data having metadata including at least one of said one or more items that are related.

4. The device according to any one of the preceding claims, further comprising: means for processing the identified existing stored data to enhance suitability of the existing stored data for training the machine learning model; and Wherein the means for providing access to the identified existing stored data comprises: means for providing access to a processed version of the existing stored data.

5. The apparatus of claim 4, wherein the means for processing the identified existing stored data comprises: Means for applying signal processing to the identified existing stored data.

6. The device according to any one of the preceding claims, further comprising: Means for generating tags for the identified existing stored data for use in training the machine learning model associated with the application.

7. The apparatus of claim 6, wherein the means for generating a label comprises: means for generating one or more hidden tasks from the identified existing stored data; as well as Means for marking the identified existing stored data based on the one or more hidden tasks.

8. The apparatus of claim 7, wherein the means for generating the one or more hidden tasks from the identified existing stored data comprises: A component for generating the one or more hidden tasks from the identified existing stored data based on optimizing a labeling function and a plurality of machine learning models configured to perform the hidden tasks, wherein the optimization is based on a consistency score between outputs of the plurality of machine learning models when performing the hidden tasks.

9. The apparatus of claim 8, wherein the plurality of machine learning models have the same architecture but different starting parameter values.

10. The apparatus according to any one of claims 7 to 9, wherein the hidden task is a random classification task.

11. The apparatus according to any one of claims 7 to 10, wherein the means for marking the identified existing stored data based on the one or more hidden tasks comprises: Means for labeling the existing stored data based on an active learning model.

12. The apparatus according to any one of claims 7 to 11, wherein the means for marking the identified existing stored data based on the one or more hidden tasks comprises at least one of the following: means for providing a user with a subset of said existing stored data for said one or more hidden tasks for manual labeling; means for receiving from a user manual markings of said subset of said existing stored data; or Means for automatically marking remaining said existing stored data based on said received manual markings.

13. The apparatus according to any one of the preceding claims, further comprising: Means for modifying the metadata of the identified existing stored data to indicate reusability of the data for training a machine learning model.

14. A method of machine learning, comprising: receiving, by a device, from an application, a request to collect new user data for use in training a machine learning model associated with the application; identifying, by the device based on the ontology, existing stored data suitable for training the machine learning model; as well as In response to the request by the application, access to the identified existing stored data is provided by the device.

15. A computer readable medium comprising program instructions, which, when executed by a device, cause the device to perform at least the following operations: receiving, by the device, from an application, a request to collect new user data, the new user data being used to train a machine learning model associated with the application; identifying, by the device based on the ontology, existing stored data suitable for training the machine learning model; and In response to the request by the application, access to the identified existing stored data is provided by the device.