Accurate and efficient inference in multi-device environment

By extracting features in the common feature space in a multi-device environment and training task-specific models, the problem of inconsistent inference performance of machine learning models in a multi-device environment is solved, and more efficient and accurate inference effect is achieved.

CN119948495APending Publication Date: 2025-05-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066593.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-07-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In a multi-device environment, a single machine learning model has difficulty accurately processing data captured by different devices, resulting in inconsistent inference performance across devices and deployment scenarios.

Method used

Features in the common feature space are extracted from multiple types of client devices by using client device-specific feature extractors and training task-specific models based on these features, deploying them on each client device.

Benefits of technology

This method reduces the computational cost of training machine learning models, improves inference accuracy in multi-device environments, and ensures stable performance in different devices and deployment scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948495A_ABST
    Figure CN119948495A_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for training and using a machine learning model in a multi-device network environment. An example computer-implemented method for network communication performed by a host device includes: extracting, using a client device-specific feature extractor, a feature set from a data set associated with a client device, wherein the feature set includes a subset of features in a common feature space; training a task-specific model based on the extracted set of features and one or more other sets of features associated with other client devices, wherein the set of features associated with the other client devices includes one or more subsets of features in the common feature space; and deploying a respective version of the task-specific model to each respective client device of the plurality of client devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application serial number 17 / 935,067, filed on September 23, 2022, which is hereby incorporated by reference into this application.

[0003] introduction

[0004] Various aspects of the present disclosure relate to machine learning models for use in a multi-device environment.

[0005] In many computing environments, many different types of devices may perform similar tasks. For example, devices such as smart phones, wearable devices, Internet of Things (IoT) devices, desktop computers, laptop computers, tablet computers, connected home devices, etc. may include components that allow voice commands to be captured and processed to trigger the execution of one or more actions on these devices or relative to other devices in the network. However, these devices may have different input components that capture data differently, and may have different processing capabilities that affect how these devices can process the captured data. For example, a smart phone or tablet computer with multiple processors may be able to use one or more machine learning models to process input data faster and more accurately (e.g., using a larger number of quantization bins, larger data types, etc.) compared to a wearable device or IoT device with a less capable processor. In addition, even across the same type of device, different models of devices may have data capture components that capture data at different quality levels. For example, a high-end device may be able to capture CD-quality audio content (e.g., using 16 bits and a 44.1kHz sampling rate) utilizing a microphone that can capture an audible frequency range (e.g., frequencies between 20 Hz and 20 kHz), while a low-end device may capture lower quality audio content (e.g., using fewer bits and / or a lower sampling rate) utilizing a microphone that captures a smaller frequency range (e.g., frequencies between 80 Hz and 255 Hz, corresponding to the range between the low and high ends of human speech audio).

[0006] In addition, different devices may be deployed in different scenarios. For example, mobile phones may be used in various environments with different environmental noise (e.g., wind noise) characteristics, while devices in motor vehicles may be used in environments with relatively consistent environmental noise (e.g., wind noise, road noise, etc.) characteristics. In another example, devices such as Internet-enabled smart devices in a home may operate in an environment with little noise or sporadic background noise at different times. Because different devices may be deployed in different scenarios, a single machine learning model may not be able to accurately process the captured data and trigger the execution of appropriate actions based on the results of processing the captured data on these devices.

[0007] Therefore, there is a need for techniques for accurately performing inference using machine learning models. Summary of the invention

[0008] Certain aspects provide a computer-implemented method for network communication by a host device. The method generally includes: extracting a feature set from a data set associated with a client device using a client device-specific feature extractor, wherein the feature set includes a feature subset in a common feature space; training a task-specific model based on the extracted feature set and one or more other feature sets associated with other client devices, wherein the feature sets associated with the other client devices include one or more feature subsets in the common feature space; and deploying a respective version of the task-specific model to each respective client device in a plurality of client devices.

[0009] Other aspects provide a computer-implemented method for network communication by a client device. The method generally includes: sending a data set associated with the client device to a host device; receiving a version of a task-specific model trained based on a feature set extracted from at least the sent data set; receiving input for processing; generating an inference based on the received input and the received version of the task-specific model; and performing one or more actions based on the inference.

[0010] Other aspects provide: a processing system configured to perform the aforementioned methods and those described herein; a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the aforementioned methods and those described herein; a computer program product embodied on a computer-readable storage medium, the computer program product comprising code for performing the aforementioned methods and those further described herein; and a processing system comprising components for performing the aforementioned methods and those further described herein.

[0011] The following description and associated drawings set forth in detail certain illustrative features of the one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings depict certain aspects of the various features of the disclosure and therefore are not to be considered limiting of the scope of the disclosure.

[0013] Figure 1 An example network environment is depicted in which multi-device inference is performed.

[0014] Figure 2 An example multi-device environment is depicted in accordance with aspects of the present disclosure, in which a machine learning model is trained and used to generate inferences in a multi-device network based on features in a common feature space extracted from data captured by various client devices.

[0015] Figure 3 Depicted is a pipeline for multi-device inference using machine learning models and features in a common feature space extracted from data from different client devices in a multi-device network in accordance with aspects of the present disclosure.

[0016] Figure 4 Depicted are examples of scaling a machine learning model for client devices in a multi-device network in accordance with aspects of the present disclosure.

[0017] Figure 5A An example of device registration in a multi-device network according to aspects of the present disclosure is depicted.

[0018] Figure 5B Depicted are example device-to-device mappings in a machine learning model for performing inference in a multi-device network in accordance with aspects of the present disclosure.

[0019] Fig. 6A and Figure 6B Depicted are example interactions between a host device and two client devices for training a machine learning model to generate inferences based on data from different devices encoded into a common feature space, in accordance with aspects of the present disclosure.

[0020] Figure 7 Depicted are example operations that may be performed for training a machine learning model for multiple client devices in a network in accordance with aspects of the present disclosure.

[0021] Figure 8 Depicted are example operations that may be performed for generating inferences from input data using a machine learning model trained for multiple client devices in accordance with aspects of the present disclosure.

[0022] Fig. 9Depicted are example implementations of a processing system on which a machine learning model is trained for multiple client devices in a network in accordance with aspects of the present disclosure.

[0023] Fig.10 Depicted are example implementations of a processing system on which a machine learning model trained for multiple client devices is used to generate inferences from input data in accordance with aspects of the present disclosure.

[0024] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements and features of one aspect may be beneficially incorporated in other aspects without further recitation. DETAILED DESCRIPTION

[0025] Aspects of the present disclosure provide apparatus, methods, processing systems, and computer-readable media for training and using machine learning models to perform inference on devices with different capabilities in a networked environment.

[0026] In many networked environments, such as environments where multiple devices communicate via wireless communication networks, many devices with different capabilities may be deployed. Because these devices have different capabilities, training and deploying machine learning models to process inputs collected by these devices may be a computationally complex and resource-intensive process. For example, different machine learning models may be trained for different devices, even if these models may ultimately be trained to perform the same task. However, because there may be an intractable number of combinations of deployment scenarios and device types (e.g., high-end devices capable of high-resolution data capture, low-end devices capable of low-resolution data capture, and devices with capabilities between the capabilities of high-end devices and low-end devices), it may not be possible to generate a machine learning model for each of these types of devices and deployment scenarios that will accurately process input data captured by devices deployed in a networked environment.

[0027] However, because multiple devices can exist in any given environment, multiple devices can capture data from the same input source. For example, assume that a smartphone and a connected home device are located in the same room. These two devices can record the same input from the same source (e.g., a user speaking a voice command within a networked environment to instantiate a specific operation). Because these devices deployed in the same environment can capture input data from the same source, these devices can collaborate to accurately process the input data captured in that environment.

[0028] Various aspects of the present disclosure provide techniques and apparatus for training and using machine learning models to accurately process input data in a multi-device environment. As discussed in further detail herein, a host device can be trained using data from multiple types of client devices, and can use a client device-specific feature extractor that extracts device-independent features to extract features from data received from these client devices. In this way, a machine learning model can be trained to extract features in the same feature space from input data with different characteristics, and inferences generated by the machine learning model can be generated based on features in a common space rather than features in a unique feature space for each different type of device and the environment in which these devices are deployed. Because the machine learning model described herein is trained based on data in a common space, the computational cost of training a machine learning model to perform inference operations on various devices or using inputs from various devices can be reduced. In addition, the accuracy of these inferences can be improved because inferences can be generated using a machine learning model trained on a larger data set in a common space rather than multiple smaller data sets from different spaces.

[0029] Inference by example in a multi-device network environment with multiple devices

[0030] Figure 1 An example network environment 100 is depicted in which inferences are performed using machine learning models and data captured by various devices in the network environment.

[0031] As illustrated, in the network environment 100, one or more client devices 102 (but for simplicity, only one client device 102 is illustrated) and one or more host devices in the cloud environment 104 can use machine learning models to participate in training and inference operations. Typically, the host device in the cloud environment 104 may have more computing power than the client device 102 in the network environment 100; for example, the host device in the cloud environment 104 may include a significantly larger number of processing units (e.g., central processing units (CPUs) or other general-purpose processors, graphics processing units (GPUs), digital signal processors (DSPs), and / or other special-purpose processors) that are more powerful than the processing units installed at the client device 102 (e.g., can support a greater number of operations per second). Therefore, in many cases, the host device in the cloud environment 104 can train a common model that can be used across devices (but with varying accuracy, depending on the capabilities of any given device and the environment in which these devices are deployed), and deploy the model to the host device or one or more client devices 102 as illustrated. In some cases, user registration and refinement of the machine learning model to allow the machine learning model to be used by a particular user of the client device 102 may be performed on the client device 102, as user registration and refinement may be operations that are less computationally complex and therefore use fewer computing resources than initial training of the machine learning model from a training data set. In some aspects, the client devices 102 in the network environment 100 may not be fixed and may be selected based on various parameters, such as the capabilities of the client devices 102, the power states of the client devices 102, and the like.

[0032] After the machine learning models are deployed (e.g., to one or both of the client device 102 or the host device in the cloud environment 104), input data to be processed using these machine learning models can be captured at the client device 102. For example, Figure 1 As illustrated, the client device 102 may capture a recording of a user's voice command. The recording (e.g., raw audio data) or feature data extracted from the raw audio data (e.g., Mel-frequency cepstral coefficients (MFCCs)) may be processed using a model at either or both of the client device 102 or the host device in the cloud environment 104 to detect keywords in the recording and trigger execution of one or more operations at the client device based on the detected keywords.

[0033] As discussed, the client devices 102 in the network environment 100 may have different capabilities and may be deployed in different environments. Therefore, a common model trained by a host device in the cloud environment 104 may not be executed (e.g., may have poor inference performance, or have different inference performance on different devices). These performance issues may be caused in part by the fact that each client device 102 generates data with different characteristics due to the capabilities of these devices and the environments in which these devices are deployed.

[0034] In order to improve the performance of machine learning models deployed in a multi-device environment, various aspects of the present disclosure provide techniques and apparatus that allow features in a common feature space to be extracted from data captured by different devices in different environments. When extracting data into a common feature space and using the data in the common feature space to train machine learning models and generate inferences for captured input data, various aspects of the present disclosure can allow accurate inferences to be made on different client devices without the computational expense of inferring training multiple models (e.g., for different devices and deployment scenarios).

[0035] Figure 2 An example multi-device environment 200 is depicted in which a machine learning model is trained and used to generate inferences in a multi-device network based on features in a common feature space extracted from data captured by various client devices, in accordance with aspects of the present disclosure.

[0036] As illustrated, environment 200 includes a host device 202 and one or more client devices 204, each of which is communicatively connected to host device 202 via connection 206. Host device 202 is generally configured to train a machine learning model based on features in a common space extracted from input data from client device 204 using a device-specific feature extractor, as discussed in further detail below. In some aspects, host device 202 may also process input data (e.g., voice recordings) to generate inferences using a trained machine learning model, and may retrain a machine learning model and / or a device-specific feature extractor based on inferences generated from input data received from client device 204.

[0037] In some aspects, client devices 204 may represent a subset of devices that belong to a trusted ecosystem. A trusted ecosystem may include, for example, devices produced by a specific manufacturer or having components produced by a specific manufacturer, devices associated with a specific user, and the like. Typically, client devices 204 that belong to a trusted ecosystem may participate in the training and use of machine learning models as discussed herein, and other client devices that do not belong to a trusted ecosystem may not participate in the training and use of machine learning models as discussed herein. In some aspects, client devices 204 that are trusted devices may be registered with host device 202 (e.g., via an authentication process that authenticates these devices and verifies that these devices are part of the trusted ecosystem). In this case, client devices that are not registered with host device 202 may not participate in the training and use of machine learning models as discussed herein, but client devices 204 that are registered with host device 202 may participate in the training and use of machine learning models as discussed herein.

[0038] In general, connection 206 may be a bidirectional communication connection between host device 202 and client device 204. For example, as illustrated, a communication link may exist between host device 202 and each respective client device 204 in environment 200, such that client device 204A communicates with host device 202 via connection 206A, client device 204B communicates with host device 202 via communication link 206B, and client device 204C communicates with host device 202 via communication link 206C. In one example, as illustrated, an uplink channel between client device 204 and host device 202 may carry features in a common feature space extracted from input captured at client device 204 for registration and processing using a machine learning model deployed at host device 202. A downlink channel between host device 202 and client device 204 may carry a model to be deployed locally at client device 204 (e.g., for detecting keywords in a captured voice sample that trigger a defined function). The downlink channel may also carry information generated by a machine learning model deployed at the host device 202 for a given input, such as authentication information, information identifying a command to be executed at the client device 204, and the like.

[0039] As illustrated, there may be many different types of devices in environment 200. For example, environment 200 may include client device 204A (illustrated as a smart phone), client device 204B (illustrated as an edge device), client device 204C (illustrated as a vehicle), and other devices (not illustrated). Because each of these devices has different capabilities and capture components, and because the user (or other data source) may not be located equidistantly from each of these devices, client devices 204A, 204B, and 204C may capture different data for the same input. For example, data captured at a client device 204 that is closer to a user (or other data source) may be louder or otherwise have greater fidelity than data captured at a client device 204 that is farther away from the user (or other data source). Additionally, the fidelity of the captured data may also vary based on the capabilities of the capture components at each of these client devices; for example, a lower cost device such as an infrastructure device (e.g., client device 204B) may have less capable capture components than a higher cost device such as client device 204A (smartphone) or 204C (vehicle) and, therefore, may not be able to generate as accurate inferences for a given input as either of client devices 204A, 204C.

[0040] In order to compensate for or at least adjust for differences in inputs captured by multiple client devices 204 in a multi-device environment 200, features used as inputs into a machine learning model may be extracted from the inputs captured by these client devices, and these features may be located in the same feature space. By extracting features from inputs captured by different sources into the same feature space, a common machine learning model can be used to accurately generate inferences for such data. In addition, as discussed in further detail below, correlations between features extracted by different devices in the environment 200 and the resulting inferences generated by these devices can be used to refine the machine learning models deployed at the host device 202 and the client device 204, which can further improve the inference accuracy of the inferences generated by each of the client devices in the environment 200.

[0041] Figure 3 Depicted is an example pipeline 300 for multi-device inference using machine learning models and features in a common feature space extracted from data from different client devices in accordance with aspects of the present disclosure.

[0042] The pipeline 300 generally allows a client device 304 (and its user) to provide information to the host device 202 to train and use a machine learning model based on features in a common feature space. As illustrated, the pipeline 300 includes a plurality of device-specific feature extractors 302, each of which may be associated with a specific client device 204 in a multi-device network environment. Typically, the device-specific feature extractors 302 are trained to extract features located in a common feature space from input data associated with (e.g., captured by) the client devices 204 associated with the device-specific feature extractors 302 (e.g., from the corresponding users of these client devices). Because the device-specific feature extractors 302 extract features in a common feature space from the input, each of the feature extractors 302A, 302B, and 302C can therefore be trained to generate device-independent and, in some cases, environment-independent data. Thus, a machine learning model can be trained at host device 202 using data in the same feature space rather than different data in various spaces, which can reduce the likelihood that the model will overfit to a specific scenario (e.g., a specific device from which the majority of data is received, at the expense of inference accuracy for data captured by other devices in the environment) or underfit to a wide variety of scenarios (e.g., generating the same or similar inferences for data captured by different devices in different environments when such similarity in the inference outputs is not guaranteed).

[0043] For example, in an environment where the devices 204 perform inferences based on audio data captured by the devices 204, the device-specific feature extractor 302 allows the audio data captured by these client devices 204 to be mapped to similar or identical features regardless of the quality of the audio data captured by these devices 204. In this way, the device-specific feature extractor 302 effectively normalizes the captured audio data across these devices 204. By normalizing the input data and mapping the input data to the same feature space, the machine learning models discussed herein can generate accurate inferences for any device 204 in a multi-device network environment.

[0044] As illustrated, the features in the common feature space extracted by the device-specific feature extractor 302 may be input into a user identifier / semantic decryptor 306, which may, for example, implement one or more machine learning models to authenticate a user of the client device 204 and / or decrypt an input voice recording to identify one or more functions to be invoked and executed at the host device 202 or the client device 204. At block 308, the machine learning model implemented in the user identifier / semantic decryptor 306 may be trained to minimize or at least reduce a task-specific loss, which may be defined with respect to a specific task to be performed on one or both of the host device 202 or the client device 204. The task-specific loss minimized (or at least reduced) at block 308 may be different for different applications of the machine learning model deployed at the host device 202. For example, a first task-specific loss may be defined for user voice authentication, while a second task-specific loss may be defined for triggering various functions to be performed at the host device 202 and / or the client device 204.

[0045] The feature extractors 302 may be trained using the feature registration loss calculated at block 310 to extract a set of features in a common feature space from inputs generated by the client devices 204 associated with each feature extractor 302. In general, the feature registration loss calculated at block 310 may allow the feature extractors 302 to compensate for or at least adjust for specific properties of each client device 204 and the environments in which those client devices 204 are deployed.

[0046] Figure 4 An example of scaling a machine learning model for a client device (e.g., client device 204A, 204B, or 204C) in a multi-device network is depicted in accordance with aspects of the present disclosure.

[0047] Typically, to train a machine learning model, each client device 204 provides both capability information and local data captured at the client device 204 to the host device 202. The capability information may include, for example, information about the amount of memory at the client device 204 (e.g., total memory, free memory after loading the operating system and other associated software components, memory available at different accelerators (if any) installed on the client device, etc.), processing power (the number of operations per second supported by the processor or other data that may represent the processing power of the client device, such as the processor ID, the number of processing cores, etc.), power utilization attributes, etc. This capability information may be used to scale the machine learning model so that different machine learning models may be deployed to different devices depending on the capabilities of those devices. More generally, devices with greater computing power may receive models that allow data to be processed with higher fidelity (e.g., bit rate, number of quantization intervals, etc.) than devices with less computing power. In another example, because a device connected to a larger battery or main power source can draw more power from the power source than a device connected to a smaller battery and not connected to main power, the device connected to the larger battery or main power source may be configured with a model that allows data to be processed with a higher fidelity (e.g., bit rate, number of quantization intervals, etc.) than a model deployed to the device connected to the smaller battery and not connected to main power.

[0048] To train the machine learning model, the host device 202 therefore receives local data D 1 , D 2 , ..., D N , capability information C of each client device 204 1 , C 2 , ..., C N , and the expected performance Acc of the model deployed to each client device 204 1 , Acc 2 , ..., ACC N The resulting model It can be expressed by the following expression:

[0049]

[0050] The resulting model The model may be trained without considering the capability information and expected performance defined for each client device 204 in the environment. In order to customize the model deployed to each client device 204 to consider the capability information and expected performance of each client device 204, the host device 202 may scale (or prune) the trained model. To fit the capability information C of each client device 1 , C 2 , ..., CN and the expected performance Acc of the model deployed to each client device 1 , Acc 2 , ..., Acc N .

[0051] Typically, a scaled (or pruned) model Or the model of the i-th client device among a total of n client devices in a multi-device environment can be represented by the following expression:

[0052]

[0053] In some aspects, the model of the i-th client device The expected performance may not be achieved i (Given the ability C i In this case, the expected performance of the i-th client device can be considered a limiting factor for the scaled model generated for each client device. In order to generate a model that allows execution according to the expected performance Acc, the capabilities can be changed at the host device 202 to achieve the expected performance of the model trained for the i-th client device 204. The resulting scaled model can then be deployed from the host device 202 to the appropriate client device 204, where the updated capability information is mapped to the specific client device 204.

[0054] Various techniques may be used to prune the model for each client device 204. For example, the pruned model may be a sub-model, or a trained model. For example, if the trained model is implemented as a decision tree, the pruned model of the client device 204 may be better than the trained model Shallow model, so that the depth of the pruned model is less than the trained model On the other hand, in the trained model In the case of being implemented as a neural network, the pruned model of the client device 204 may include fewer neurons, fewer layers, or be otherwise smaller than the trained model. Model.

[0055] In yet other examples, host device 202 can generate a scaled model for client device 204 by changing the quantization interval size of the model. Multiple quantization intervals or categories into which data can be classified can be used for training. A larger number of quantization intervals can use a larger number of bits to represent the classification generated for the input data, which can use a larger amount of computing resources when performing inference on the input data. To allow the use of fewer computing resources to perform inference, the scaled model can use a smaller number of larger quantization intervals into which data can be classified. By reducing the number of quantization intervals and correspondingly increasing the size of each quantization interval, fewer computing resources can be used to perform inference because fewer bits can be used to represent the quantization interval, at the expense of inference accuracy.

[0056] In some environments, a user may use multiple devices to trigger various operations on one or more client devices in the environment. Because registration can be a time-consuming process, and because client devices in the environment may be configured with machine learning models trained to generate inferences based on features extracted into a common feature space, some aspects of the present disclosure allow for the use of device-to-device mapping functions to facilitate user registration on client devices in the environment. Typically, these device-to-device mapping functions are used, such as Figure 5A As illustrated, a user may register on a first client device 204A, and registration information (eg, a registration vector) may be communicated to other client devices in the environment (eg, a second client device 204C).

[0057] In order to allow a user to register on multiple client devices 204 using a registration process on a single device, properties or characteristics of each client device 204 may be used to transform a registration vector from a vector appropriate for the first client device to a vector appropriate for the second client device via a mapping function between the first client device and the second client device. The mapping function may be implemented, for example, via a non-parametric model, an autoencoder model, a generative adversarial model (GAN), or other model that allows for transformation from data generated by a machine learning model on a first client device to data to be generated by a machine learning model on a second client device.

[0058] Figure 5B Depicted are example device-to-device mappings in a machine learning model for performing inference in a multi-device network in accordance with aspects of the present disclosure.

[0059] As illustrated, input x may be received at client devices 204A and 204C. Client device 204A may generate an impulse response F 1 (x) is used for input x, and the client device 204C may generate an impulse response F 2 (x) for input x. To use the registration vector generated by client device 204A at client device 204C for input x, x may be mapped using the mapping function M 1→2 (x) is transformed into In one example, in the mapping function M 1→2 When (x) is implemented as a GAN, the impulse response F 1 (x) and F 2 (x) may be recorded in parallel. The impulse response of each of the client devices 204A and 204C typically embeds various characteristics of the client devices 204A and 204C, such as device capabilities at each device, quality of capture components at each device, etc. The impulse response may be used to identify the F 1 (x) and F 2 (x) to train the neural network to generate two mapping functions. The first mapping function can map the impulse response from the first client device to the impulse response of the second client device (e.g., according to the expression M 1→2 (x)), and the second mapping function may map the impulse response from the second client device to the first client device (e.g., according to the expression M 2→1 (x)).

[0060] In some aspects, devices in a multi-device network environment may teach each other via knowledge distillation and continuous learning. Fig. 6A and Figure 6B Depicted are example interactions between a host device (e.g., host device 202) and two client devices (both of client devices 204A, 204B, or 204C) for training a machine learning model to generate inferences based on data from the different devices encoded into a common feature space in accordance with aspects of the present disclosure.

[0061] like Fig. 6A As illustrated, devices in a multi-device environment may capture the same input with different signal-to-noise ratios (or other quality metrics). In this example, client device 204A may be closer to the source of the data input than client device 204B, and therefore, the quality of the captured input at client device 204A may be greater than the quality of the captured input at client device 204B. Therefore, the machine learning model deployed at client device 204A may generate predictions with a higher confidence level than the machine learning model deployed at client device 204B. In order to transfer knowledge from a stronger device (e.g., client device 204A) to a weaker device (e.g., client device 204B), host device 202 may use the captured input and the inferences generated by client devices 204A and 204B to generate updates to the global machine learning model, and distribute the scaled model to client devices 204 in the multi-device network.

[0062] like Figure 6B As illustrated, the host device 202 thus receives a pair of input x and inference y from the first client device 204A, designated as (x 1,y 1 ), and receives a pair of input x and inference y from the second client device 204B, designated as (x 2 ,y 2 ). When teaching a peer client device in a multi-device environment, the host device 202 may calculate the host tag y for the input x according to the following formula:

[0063]

[0064] where c i Representing and inferring y i The confidence level associated with the inference generated by the i-th client device. Confidence level c i The following properties may be observed:

[0065]

[0066] and

[0067] c i ∝modelSize

[0068] Thus, the confidence level may be proportional to the inverse of the distance between the input data source and the i-th client device, and may be proportional to the size of the model. Thus, devices closer to the input source may be used to teach devices farther away, because inferences generated by devices closer to the input source (and therefore with higher confidence levels) may have knowledge that may be used as teaching information for weaker devices, such as intermediate features, output softmax distributions, etc. Similarly, devices with larger models may have better recognition results, and therefore may also be used to teach other devices with smaller and therefore weaker models.

[0069] After scaling the updated model for each client device according to the capability information and target performance information associated with each client device 204 in the multi-device environment, the scaled updated model may be distributed to the client devices 204 for use in performing subsequent inferences.

[0070] Example operations for training a machine learning model for multiple client devices in a network

[0071] Figure 7 An example of a method 700 for training a machine learning model for multiple client devices in a network according to aspects of the present disclosure is shown. In some examples, the method 700 may be performed by a host device (e.g., Figures 2 to 6B exemplified host device 202).

[0072] As illustrated, method 700 begins at block 705, where a feature set is extracted from a data set associated with a client device (e.g., client device 204A, 204B, or 204C) using a client device-specific feature extractor, wherein the feature set includes a subset of features in a common feature space. In some cases, the operation of this block refers to the process described in reference to FIG. Fig. 9 The described circuit for extracting and / or code for extracting may be or may be executed by the circuit and / or the code.

[0073] Method 700 then proceeds to block 710, where a task-specific model is trained based on the extracted feature set and one or more other feature sets associated with other client devices, wherein the feature sets associated with the other client devices include one or more feature subsets in the common feature space. In some cases, the operation of this block refers to the example of reference Fig. 9 The described circuits for training and / or codes for training may be executed by the circuits and / or codes.

[0074] The method 700 then proceeds to block 715, where a respective version of the task-specific model is deployed to each respective client device in the plurality of client devices. In some cases, the operation of this block refers to the operation described in reference to Fig. 9 The described circuits for deployment and / or codes for deployment, or may be executed by the circuits and / or the codes.

[0075] In some aspects, the client device-specific feature extractors are associated with a class of devices having common capabilities.

[0076] In some aspects, the method 700 further includes: training the client device specific feature extractor to extract the same set of features from inputs having the same labels generated by multiple client devices. In some cases, the operation of this block refers to the Fig. 9 The described circuits for training and / or codes for training may be executed by the circuits and / or codes.

[0077] In some aspects, the method 700 further includes: for each respective client device, generating a respective version of the task-specific model based on the capability information and the target performance associated with the respective client device. In some cases, the operation of this block refers to the example of reference Fig. 9 In some aspects, generating a corresponding version of the task-specific model includes: pruning (or scaling) the model based on the capability information and the target performance, so that the pruned model includes a portion of the task-specific model (e.g., as described above with respect to Figure 4As discussed, a pruned (or scaled) model typically includes various techniques that cause the scaled model to be smaller than the task-specific model, such as a model having a depth that is less than the depth of the task-specific model, a model having fewer convolutional layers or neurons than the task-specific model, etc. More generally, a pruned model may allow inference to be performed using fewer computing resources than the amount of computing resources used to perform inference using the task-specific model. In some aspects, generating a corresponding version of the task-specific model includes: adjusting a quantization interval size of the corresponding version of the task-specific model so that the quantization interval size of the corresponding version of the task-specific model is associated with an interval that is larger than a corresponding quantization interval size of the task-specific model (e.g., as described above with respect to Figure 4 discussed).

[0078] In some aspects, method 700 further includes: receiving an indication that the performance of the corresponding version of the task-specific model of the corresponding client device fails to meet the target performance of the corresponding client device; generating a revised corresponding version of the task-specific model based on the received indication; and deploying the revised corresponding version of the task-specific model to the corresponding client device. In some cases, the operation of this block refers to as described in reference Fig. 9 The circuit for receiving and / or the code for receiving described herein may be executed by the circuit and / or the code. In some aspects, the revised corresponding version of the task-specific model includes a pruned version of the corresponding version of the task-specific model. In some aspects, the revised corresponding version of the task-specific model includes a version of the task-specific model having quantization intervals associated with an interval size that is larger than the interval size associated with the corresponding version of the task-specific model.

[0079] In some aspects, the method 700 further includes: generating a mapping between the first type of client device and the second type of client device. In some cases, the operation of this block refers to the same as in reference Fig. 9 Circuitry for generating and / or code for generating described herein, or executable by the circuitry and / or code.In some aspects, the mapping includes a function that transforms an input associated with a first type of client device into an input associated with a second type of client device.

[0080] In some aspects, method 700 further includes: training a machine learning model to map input from a first type of client device to input from a second type of client device based on device attributes, samples generated by the first type of client device, and samples generated by the second type of client device. In some cases, the operation of this block refers to as described in reference Fig. 9 The described circuits for training and / or codes for training may be executed by the circuits and / or codes.

[0081] In some aspects, the method 700 further includes: training a machine learning model to map input from a first type of client device or a second type of client device to input from any client device in the plurality of client devices based on device attributes, samples generated by the first type of client device, and samples generated by the second type of client device. In some cases, the operation of this block refers to as described in reference Fig. 9 The described circuits for training and / or codes for training may be executed by the circuits and / or codes.

[0082] In some aspects, method 700 further includes: computing a common label based on data from the client device and data from another client device; updating the task-specific model based on a first pairing between the data from the client device and the common label and a second pairing between the data from the another client device and the common label; and deploying the updated task-specific model to the client device and the another client device. In some cases, the operations of this block refer to the example of reference 700. Fig. 9 The described circuits for computing, circuits for updating, circuits for deploying and / or codes for computing, codes for updating and codes for deploying may be or may be executed by these circuits and / or codes.

[0083] In some aspects, the data from the client device includes a label generated for an input of the client device and a confidence level associated with the label.

[0084] In some aspects, the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with multiple client devices, or a signal-to-noise ratio (SNR) associated with an input to the client device.

[0085] In some aspects, a plurality of client devices comprises the client device.

[0086] In some aspects, the plurality of client devices may include client devices that are members of a trusted ecosystem. These client devices may register as part of the trusted ecosystem prior to participating in the training of the machine learning model.

[0087] Example operations for generating inferences using a machine learning model trained on multiple client devices

[0088] Figure 8 An example of a method 800 for generating inferences from input data using a machine learning model trained for multiple client devices in accordance with aspects of the present disclosure is shown. In some examples, the method 800 may be performed by a client device, such as Figures 2 to 6B An exemplary client device 204 (eg, 204A, 204B, or 204C) is shown.

[0089] As illustrated, method 800 begins at block 805, where a data set associated with a client device is sent to a host device (e.g., host device 202). In some cases, the operation of this block refers to the operation described in reference to Fig.10 The described circuits for transmitting and / or codes for transmitting, or executable by the circuits and / or the codes.

[0090] Method 800 then proceeds to block 810, where a version of a task-specific model trained based on at least a feature set extracted from the transmitted data set is received. In some cases, the operation of this block refers to the example of reference Fig.10 The described circuits for receiving and / or codes for receiving may be or may be executed by the circuits and / or codes.

[0091] Method 800 then proceeds to block 815, where input is received for processing. In some cases, the operation of this block refers to the Fig.10 The described circuits for receiving and / or codes for receiving may be or may be executed by the circuits and / or codes.

[0092] The method 800 then proceeds to block 820, where inferences are generated based on the received input and the received version of the task-specific model. In some cases, the operation of this block refers to the example of reference Fig.10 The described circuit for generating and / or code for generating, or can be executed by the circuit and / or the code.

[0093] Method 800 then proceeds to block 825, where one or more actions are performed based on the inference. In some cases, the operation of this block refers to the Fig.10 Circuits for performing and / or code for performing are described, or can be performed by such circuits and / or such code.

[0094] In some aspects, the version of the mission-specific model is based on capability information and target performance associated with the client device.

[0095] In some aspects, the version of the task-specific model includes a pruned version of the task-specific model, such that the pruned model includes a portion of the task-specific model.

[0096] In some aspects, the version of the task-specific model includes a version of the task-specific model having quantization intervals associated with an interval size that is larger than an interval size associated with a corresponding version of the task-specific model.

[0097] In some aspects, method 800 further includes: sending to the host device an indication that the performance of the version of the task-specific model of the client device fails to meet the target performance of the client device. In some cases, the operation of this block refers to as described in reference Fig.10 The described circuits for transmitting and / or codes for transmitting, or executable by the circuits and / or the codes.

[0098] In some aspects, the method 800 further includes: receiving a revised version of the task-specific model from the host device based on the indication. In some cases, the operation of this block refers to the following example. Fig.10 The described circuits for receiving and / or codes for receiving may be or may be executed by the circuits and / or codes.

[0099] In some aspects, the revised version of the task-specific model includes a pruned version of the task-specific model.

[0100] In some aspects, the revised version of the task-specific model includes another version of the task-specific model having quantization intervals associated with an interval size that is larger than the interval size associated with the revised version of the task-specific model.

[0101] In some aspects, method 800 further includes: receiving an updated task-specific model, wherein the updated task-specific model is based on: a common tag based on data from the client device and data from another client device; and a first pairing between the data from the client device and the common tag and a second pairing between the data from the another client device and the common tag. In some cases, the operation of this block refers to as described in reference Fig.10 The described circuits for receiving and / or codes for receiving may be or may be executed by the circuits and / or codes.

[0102] In some aspects, the data from the client device includes a label generated for an input of the client device and a confidence level associated with the label.

[0103] In some aspects, the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with multiple client devices, or a signal-to-noise ratio (SNR) associated with an input to the client device.

[0104] Example process for training and generating inferences using a machine learning model trained on multiple client devices Management System

[0105] Fig. 9 An example processing system 900 is depicted for training a machine learning model for multiple client devices in a network, such as described herein, for example, with respect to Figure 7 As described.

[0106] The processing system 900 includes a central processing unit (CPU) 902, which may be a multi-core CPU in some examples. Instructions executed at the CPU 902 may be loaded, for example, from a program memory associated with the CPU 902, or may be loaded from the memory 924.

[0107] The processing system 900 also includes additional processing components customized for specific functions, such as a graphics processing unit (GPU) 904 , a digital signal processor (DSP) 906 , a neural processing unit (NPU) 908 , a multimedia processing unit 910 , and a wireless connectivity component 912 .

[0108] An NPU, such as 908, is generally a dedicated circuit configured to implement control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. An NPU is sometimes alternatively referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.

[0109] NPUs such as 908 are configured to accelerate the execution of common machine learning tasks such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs may be instantiated on a single chip such as a system on a chip (SoC), while in other examples, multiple NPUs may be part of a dedicated neural network accelerator.

[0110] The NPU can be optimized for training or inference, or in some cases configured to balance performance between the two. For NPUs that can perform both training and inference, the two tasks can generally still be performed independently.

[0111] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation that involves inputting an existing data set (usually labeled or tagged), iterating on the data set, and then adjusting model parameters (such as weights and biases) to improve model performance. Typically, optimization based on error predictions involves passing back through the layers of the model and determining the gradient to reduce the prediction error.

[0112] NPUs designed to accelerate inference are generally configured to operate on complete models. Thus, such NPUs can be configured to input a new piece of data and quickly process the new piece through an already trained model to generate a model output (e.g., inference).

[0113] In one specific implementation, the NPU 908 is part of one or more of the CPU 902 , the GPU 904 , and / or the DSP 906 .

[0114] In some examples, wireless connectivity component 912 may include, for example, subcomponents for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., 4G LTE), fifth generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. Wireless connectivity component 912 is further connected to one or more antennas 914.

[0115] The processing system 900 may also include one or more sensor processing units 916 associated with any manner of sensors, one or more image signal processors (ISPs) 918 associated with any manner of image sensors, and / or a navigation processor 920, which may include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.

[0116] The processing system 900 may also include one or more input and / or output devices 922, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, and the like.

[0117] In some examples, one or more of the processors of processing system 900 may be based on the ARM or RISC-V instruction set.

[0118] The processing system 900 also includes a memory 924, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 924 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 900.

[0119] Specifically, in this example, the memory 924 includes a feature extraction component 924A, a task-specific model training component 924B, and a model deployment component 924C. The depicted components, as well as other components not depicted, may be configured to perform various aspects of the methods described herein.

[0120] Generally, the processing system 900 and / or its components may be configured to perform the methods described herein.

[0121] It is worth noting that in other aspects, aspects of the processing system 900 can be omitted, such as where the processing system 900 is a server computer, etc. For example, in other aspects, the multimedia processing unit 910, the wireless connectivity component 912, the sensor processing unit 916, the ISP 918, and / or the navigation processor 920 can be omitted. In addition, aspects of the processing system 900 can be distributed, such as training a model and using the model to generate inferences, such as user-verified predictions.

[0122] Fig.10 An example processing system 1000 is depicted for generating inferences using a machine learning model trained for multiple client devices in a network, such as described herein, for example, with respect to Figure 8 As described.

[0123] Processing system 1000 includes a central processing unit (CPU) 1002, which in some examples may be a multi-core CPU. Instructions executed at CPU 1002 may be loaded, for example, from a program memory associated with CPU 1002, or may be loaded from memory 1024.

[0124] The processing system 1000 also includes additional processing components customized for specific functions, such as a graphics processing unit (GPU) 1004 , a digital signal processor (DSP) 1006 , a neural processing unit (NPU) 1008 , a multimedia processing unit 1010 , and a wireless connectivity component 1012 .

[0125] An NPU, such as 1008, is generally a dedicated circuit configured to implement control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. An NPU is sometimes alternatively referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.

[0126] NPUs such as 1008 are configured to accelerate the execution of common machine learning tasks such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs may be instantiated on a single chip such as a system on a chip (SoC), while in other examples, multiple NPUs may be part of a dedicated neural network accelerator.

[0127] The NPU can be optimized for training or inference, or in some cases configured to balance performance between the two. For NPUs that can perform both training and inference, the two tasks can generally still be performed independently.

[0128] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation that involves inputting an existing data set (usually labeled or tagged), iterating on the data set, and then adjusting model parameters (such as weights and biases) to improve model performance. Typically, optimization based on error predictions involves passing back through the layers of the model and determining the gradient to reduce the prediction error.

[0129] NPUs designed to accelerate inference are generally configured to operate on complete models. Thus, such NPUs can be configured to input a new piece of data and quickly process the new piece through an already trained model to generate a model output (e.g., inference).

[0130] In one specific implementation, the NPU 1008 is part of one or more of the CPU 1002 , the GPU 1004 , and / or the DSP 1006 .

[0131] In some examples, wireless connectivity component 1012 may include, for example, subcomponents for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., 4G LTE), fifth generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. Wireless connectivity component 1012 is further connected to one or more antennas 1014.

[0132] The processing system 1000 may also include one or more sensor processing units 1016 associated with any manner of sensors, one or more image signal processors (ISPs) 1018 associated with any manner of image sensors, and / or a navigation processor 1020, which may include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.

[0133] The processing system 1000 may also include one or more input and / or output devices 1022 , such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, and the like.

[0134] In some examples, one or more of the processors of processing system 1000 may be based on the ARM or RISC-V instruction set.

[0135] The processing system 1000 also includes a memory 1024, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1024 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 1000.

[0136] Specifically, in this example, memory 1024 includes a data sending component 1024A, a data receiving component 1024B, an inference generating component 1024C, and an action taking component 1024D. The depicted components, as well as other components not depicted, may be configured to perform various aspects of the methods described herein.

[0137] Generally, the processing system 1000 and / or its components may be configured to perform the methods described herein.

[0138] It is worth noting that in other aspects, aspects of the processing system 1000 can be omitted, such as where the processing system 1000 is a server computer, etc. For example, in other aspects, the multimedia processing unit 1010, the wireless connectivity component 1012, the sensor processing unit 1016, the ISP 1018, and / or the navigation processor 1020 can be omitted. In addition, aspects of the processing system 1000 can be distributed, such as training a model and using the model to generate inferences, such as user-verified predictions.

[0139] Sample Clauses

[0140] Specific implementation details of various aspects of the disclosure are described in the following numbered clauses.

[0141] Item 1: A computer-implemented method for network communication by a host device, the method comprising: extracting a feature set from a data set associated with a client device using a client device-specific feature extractor, wherein the feature set includes a subset of features in a common feature space; training a task-specific model based on the extracted feature set and one or more other feature sets associated with other client devices, wherein the feature sets associated with the other client devices include one or more subsets of features in the common feature space; and deploying a corresponding version of the task-specific model to each corresponding client device among a plurality of client devices.

[0142] Clause 2: The method of clause 1, wherein the client device-specific feature extractor is associated with a class of devices having common capabilities.

[0143] Clause 3: The method of clause 1 or 2, further comprising: training the client device-specific feature extractor to extract the same set of features from inputs having the same labels generated by the multiple client devices.

[0144] Clause 4: The method of any one of clauses 1 to 3, further comprising: for each respective client device, generating the respective version of the task-specific model based on capability information and target performance associated with the respective client device.

[0145] Clause 5: The method of clause 4, wherein generating the corresponding version of the task-specific model comprises: pruning the model based on the capability information and the target performance, so that the pruned model includes a portion of the task-specific model.

[0146] Clause 6: A method according to clause 4 or 5, wherein generating the corresponding version of the task-specific model includes: adjusting the quantization interval size of the corresponding version of the task-specific model so that the quantization interval size of the corresponding version of the task-specific model is associated with an interval larger than the corresponding quantization interval size of the task-specific model.

[0147] Clause 7: According to the method described in any one of clauses 4 to 6, the method also includes: receiving an indication that the performance of the corresponding version of the task-specific model of the corresponding client device fails to meet the target performance of the corresponding client device; generating a revised corresponding version of the task-specific model based on the received indication; and deploying the revised corresponding version of the task-specific model to the corresponding client device.

[0148] Clause 8: The method of clause 7, wherein the revised respective version of the task-specific model comprises a pruned version of the respective version of the task-specific model.

[0149] Clause 9: A method according to clause 7 or 8, wherein the revised corresponding version of the task-specific model includes a version of the task-specific model having quantization intervals associated with an interval size that is larger than the interval size associated with the corresponding version of the task-specific model.

[0150] Clause 10: The method of any one of clauses 1 to 9, further comprising: generating a mapping between client devices of the first type and client devices of the second type.

[0151] Clause 11: The method of clause 10, wherein the mapping comprises a function that transforms input associated with the first type of client device into input associated with the second type of client device.

[0152] Clause 12: The method according to clause 10 or 11, further comprising: training a machine learning model to map input from the first type of client device to input from the second type of client device based on device attributes, samples generated by the first type of client device, and samples generated by the second type of client device.

[0153] Clause 13: A method according to any one of clauses 10 to 12, further comprising: training a machine learning model to map input from the first type of client device or the second type of client device to input from any client device among the multiple client devices based on device attributes, samples generated by the first type of client device, and samples generated by the second type of client device.

[0154] Clause 14: A method according to any one of clauses 1 to 13, the method further comprising: calculating a common label based on data from the client device and data from another client device; updating the task-specific model based on a first pairing between the data from the client device and the common label and a second pairing between the data from the other client device and the common label; and deploying the updated task-specific model to the client device and the other client device.

[0155] Clause 15: The method of clause 14, wherein the data from the client device comprises a label generated for the input of the client device and a confidence level associated with the label.

[0156] Clause 16: A method according to clause 15, wherein the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with the multiple client devices, or a signal-to-noise ratio (SNR) associated with the input to the client device.

[0157] Clause 17: The method of any one of clauses 1 to 16, wherein the plurality of client devices includes the client device.

[0158] Clause 18: The method of any one of clauses 1 to 17, further comprising: registering the plurality of client devices into a trusted ecosystem, wherein only devices in the trusted ecosystem are allowed to participate in training the task-specific model.

[0159] Clause 19: A computer-implemented method for network communications by a client device, the method comprising: sending a data set associated with the client device to a host device; receiving a version of a task-specific model trained based on a feature set extracted from at least the sent data set; receiving input for processing; generating inferences based on the received input and the received version of the task-specific model; and performing one or more actions based on the inferences.

[0160] Clause 20: The method of clause 19, wherein the version of the task-specific model is based on capability information and a target performance associated with the client device.

[0161] Clause 21: A method according to clause 19 or 20, wherein the version of the task-specific model comprises a pruned version of the task-specific model, such that the pruned model comprises a portion of the task-specific model.

[0162] Clause 22: The method of clause 19 or 20, wherein the version of the task-specific model comprises a version of the task-specific model having quantization intervals associated with an interval size that is larger than an interval size associated with the task-specific model.

[0163] Clause 23: A method according to any one of clauses 19 to 22, the method further comprising: sending an indication to the host device that the performance of the version of the task-specific model of the client device fails to meet the target performance of the client device; and receiving a revised version of the task-specific model from the host device based on the indication.

[0164] Clause 24: The method of clause 23, wherein the revised version of the task-specific model comprises a pruned version of the task-specific model.

[0165] Clause 25: A method according to clause 23 or 24, wherein the revised version of the task-specific model comprises another version of the task-specific model having quantization intervals associated with an interval size that is larger than the interval size associated with the version of the task-specific model.

[0166] Clause 26: The method according to clause 23 or 24 further includes receiving an updated task-specific model, wherein the updated task-specific model is based on: a common tag based on the data from the client device and the data from another client device; and a first pairing between the data from the client device and the common tag and a second pairing between the data from the other client device and the common tag.

[0167] Clause 27: The method of clause 26, wherein the data from a client device comprises a label generated for input to the client device and a confidence level associated with the label.

[0168] Clause 28: A method according to clause 26 or 27, wherein the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with the multiple client devices, or a signal-to-noise ratio (SNR) associated with the input to the client device.

[0169] Clause 29: A processing system, the processing system comprising: a memory having executable instructions stored thereon; and a processor configured to execute the executable instructions so as to cause the processing system to perform the method according to any one of clauses 1 to 28.

[0170] Clause 30: A processing system comprising: means for performing the method of any one of clauses 1 to 27.

[0171] Clause 31: A non-transitory computer readable medium having stored thereon instructions which, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of clauses 1 to 28.

[0172] Clause 32: A computer program product embodied on a computer readable storage medium, the computer readable storage medium comprising: code for performing the method according to any one of clauses 1 to 28.

[0173] Additional considerations

[0174] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limited to the scope, applicability or aspects set forth in the claims. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, without departing from the scope of the present disclosure, the functions and arrangements of the elements discussed may be changed. Various examples may omit, replace or add various processes or components as appropriate. For example, the described methods may be performed in a sequence different from that described, and various steps may be added, omitted or combined. In addition, the features described for some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functionality or structures and functionality that are supplementary or alternative to the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of the present invention.

[0175] As used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0176] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items (including single members). For example, "at least one of a, b, or c" is intended to cover a, b, c, ab, ac, bc, and abc, as well as any combination with multiple of the same elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).

[0177] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, a database, or another data structure), ascertaining, etc. In addition, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. In addition, "determining" may include resolving, selecting, choosing, establishing, etc.

[0178] The method disclosed herein includes one or more steps or actions for implementing the method. The steps and / or actions of the method can be interchangeable with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. In addition, the various operations of the method described above can be performed by any appropriate component capable of performing the corresponding function. The component may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs) or processors. Typically, in the case of operations illustrated in the accompanying drawings, those operations may have corresponding corresponding components with similar numbers plus functional components.

[0179] The following claims are not intended to be limited to the various aspects shown herein, but should be given the full scope consistent with the language of the claims. Within the claims, unless specifically stated otherwise, the reference to an element in the singular form is not intended to mean "one and only one", but "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. Any claim element should not be interpreted according to the provisions of 35 U.S.C. § 112 (f), unless the phrase "parts for..." is used to explicitly record the element, or in the case of a method claim, the phrase "step for..." is used to record the element. All structural and functional equivalents of the elements of the various aspects described throughout the present disclosure that are known or will be known later to a person of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly recorded in the claims.

Claims

1. A computer-implemented method for network communication by a host device, the method comprising: extracting a feature set from a data set associated with the client device using a client device specific feature extractor, wherein the feature set comprises a subset of features in a common feature space; training a task-specific model based on the extracted feature set and one or more other feature sets associated with other client devices, wherein the other feature sets associated with the other client devices include one or more feature subsets in the common feature space; as well as A respective version of the task-specific model is deployed to each respective client device of a plurality of client devices.

2. The method of claim 1, wherein the client device specific feature extractor is associated with a class of devices having common capabilities.

3. The method according to claim 1, further comprising: The client device specific feature extractor is trained to extract a same set of features from inputs having the same labels generated by the plurality of client devices.

4. The method according to claim 1, further comprising: For each respective client device, the respective version of the task-specific model is generated based on capability information and a target performance associated with the respective client device.

5. The method of claim 4, wherein generating the corresponding version of the task-specific model comprises: The model is pruned based on the capability information and the target performance such that the pruned model includes a portion of the task-specific model.

6. The method of claim 4, wherein generating the corresponding version of the task-specific model comprises: A quantization interval size of the corresponding version of the task specific model is adjusted such that the quantization interval size of the corresponding version of the task specific model is associated with an interval that is larger than a corresponding quantization interval size of the task specific model.

7. The method according to claim 4, further comprising: receiving an indication that performance of the respective version of the task-specific model for the respective client device fails to meet the target performance for the respective client device; generating a revised corresponding version of the task-specific model based on the received indication; as well as The revised respective versions of the task-specific models are deployed to the respective client devices. 8 . The method of claim 7 , wherein the revised respective version of the task-specific model comprises a pruned version of the respective version of the task-specific model.

9. The method of claim 7, wherein the revised corresponding version of the task-specific model comprises a version of the task-specific model having quantization intervals associated with an interval size that is larger than an interval size associated with the corresponding version of the task-specific model.

10. The method according to claim 1, further comprising: A mapping between client devices of the first type and client devices of the second type is generated.

11. The method of claim 10, wherein the mapping comprises a function that transforms input associated with the first type of client device into input associated with the second type of client device.

12. The method according to claim 10, further comprising: A machine learning model is trained to map input from the first type of client devices to input from the second type of client devices based on device attributes, samples generated by the first type of client devices, and samples generated by the second type of client devices.

13. The method according to claim 10, further comprising: A machine learning model is trained to map input from the first type of client device or the second type of client device to input from any of the plurality of client devices based on device attributes, samples generated by the first type of client device, and samples generated by the second type of client device.

14. The method according to claim 1, further comprising: calculating a common tag based on the data from the client device and the data from another client device; updating the task-specific model based on a first pairing between the data from the client device and the common tag and a second pairing between the data from the other client device and the common tag; as well as The updated task-specific model is deployed to the client device and the another client device.

15. The method of claim 14, wherein the data from the client device comprises a label generated for an input to the client device and a confidence level associated with the label.

16. A method according to claim 15, wherein the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with the multiple client devices, or a signal-to-noise ratio (SNR) associated with the input to the client device.

17. The method of claim 1, wherein the plurality of client devices comprises the client device.

18. The method according to claim 1, further comprising: The plurality of client devices are registered into a trusted ecosystem, wherein only devices in the trusted ecosystem are allowed to participate in training the task-specific model.

19. A computer-implemented method for network communications by a client device, the method comprising: sending a data set associated with the client device to a host device; receiving a version of the task-specific model trained based on at least a set of features extracted from the sent dataset; receiving input for processing; generating inferences based on the received input and the received version of the task-specific model; as well as One or more actions are performed based on the inference.

20. The method of claim 19, wherein the version of the task-specific model is based on capability information and target performance associated with the client device.

21. The method of claim 20, wherein the version of the task-specific model comprises a pruned version of the task-specific model such that the pruned model comprises a portion of the task-specific model.

22. The method of claim 20, wherein the version of the task-specific model comprises a version of the task-specific model having quantization intervals associated with an interval size that is larger than an interval size associated with the task-specific model.

23. The method according to claim 20, further comprising: sending an indication to the host device that the performance of the version of the task-specific model for the client device fails to meet the target performance for the client device; as well as A revised version of the task-specific model is received from the host device based on the indication.

24. The method of claim 23, wherein the revised version of the task-specific model comprises a pruned version of the task-specific model.

25. The method of claim 23, wherein the revised version of the task-specific model comprises another version of the task-specific model having quantization intervals associated with an interval size that is larger than the interval size associated with the version of the task-specific model.

26. The method of claim 19, further comprising receiving an updated task-specific model, wherein the updated task-specific model is based on: A common tag based on data from the client device and data from another client device; and A first pairing between the data from the client device and the public tag and a second pairing between the data from the other client device and the public tag.

27. The method of claim 26, wherein the data from a client device comprises a label generated for an input to the client device and a confidence level associated with the label.

28. A method according to claim 27, wherein the confidence level associated with the tag is based on at least one of: a distance between the client device and the host device, a model size associated with the client device, a model size associated with multiple client devices, or a signal-to-noise ratio (SNR) associated with the input to the client device.

29. A processing system, the processing system comprising: A memory, wherein executable instructions are stored in the memory; and a processor configured to execute the executable instructions so as to cause the processing system to: extracting a feature set from a data set associated with the client device using a client device specific feature extractor, wherein the feature set comprises a subset of features in a common feature space; training a task-specific model based on the extracted feature set and one or more other feature sets associated with other client devices, wherein the other feature sets associated with the other client devices include one or more feature subsets in the common feature space; as well as A respective version of the task-specific model is deployed to each respective client device of a plurality of client devices.

30. A processing system, the processing system comprising: A memory, wherein executable instructions are stored in the memory; and a processor configured to execute the executable instructions so as to cause the processing system to: sending a data set associated with the client device to the host device; receiving a version of the task-specific model trained based on at least a set of features extracted from the sent dataset; receiving input for processing; generating inferences based on the received input and the received version of the task-specific model; as well as One or more actions are performed based on the inference.