On-Device Activity Recognition

By using neural networks trained with differential privacy for on-device activity recognition, mobile and wearable devices can accurately recognize user activities while protecting privacy and reducing computational demands.

JP7691334B2Active Publication Date: 2025-06-11GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021164946
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-05
Filing Date
2021-10-06
Publication Date
2025-06-11
Estimated Expiration
2041-10-06

AI Technical Summary

Technical Problem

Existing mobile and wearable computing devices rely on external cloud systems for activity recognition, which raises privacy concerns and requires significant computing resources.

Method used

Implementing on-device activity recognition using neural networks trained with differential privacy, allowing the device to recognize user activities without transmitting sensor data to external systems.

Benefits of technology

This approach enhances user privacy by keeping sensor data local, reduces the need for extensive computing resources, and improves device performance by enabling faster and more accurate activity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691334000001
    Figure 0007691334000001
  • Figure 0007691334000002
    Figure 0007691334000002
  • Figure 0007691334000003
    Figure 0007691334000003
Patent Text Reader

Abstract

To provide a method, computing device, and storage medium, which allow for recognizing an activity of a user based on the sensor data provided by one or more sensor components without having to send the sensor data to an external computing system (e.g., a cloud computing system).SOLUTION: A computing device is configured to: receive motion data generated by one or more motion sensors that correspond to movement sensed by the one or more motion sensors; perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data; and, in response to recognizing the physical activity corresponding to the motion data, perform an operation associated with the physical activity.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background Some mobile computing devices and wearable computing devices can track user activities to assist a user in maintaining a healthier and more active lifestyle. For example, a mobile computing device may include one or more sensor components that provide sensor data indicative of a user engaged in physical activity. The mobile computing device may transmit data provided by one or more sensor components outside the device to a cloud computing system or the like for processing to identify the physical activity being performed by the user based on the sensor data.

Summary of the Invention

Means for Solving the Problems

[0002] Summary Generally, the techniques of the present disclosure are directed to performing on-device recognition of activities in which a user of a computing device is engaged using one or more neural networks trained with differential privacy. The computing device may recognize the user's activities based on sensor data provided by one or more sensor components without transmitting the sensor data to an external computing system (e.g., the cloud). Rather, the computing device may perform on-device activity recognition based on sensor data provided by one or more sensor components using one or more neural networks trained to perform activity recognition.

[0003] One or more neural networks can be trained outside the device in such a way that they perform activity recognition using fewer computing resources (e.g., fewer processing cycles and memory used) compared to the neural network that performs server-side activity recognition. As a result, the computing device can become capable of performing on-device activity recognition using one or more neural networks. By performing on-device activity recognition using one or more neural networks, the computing device can accurately recognize the user's activity without the need to send and receive data to and from an external computing system that performs server-side activity recognition. Rather, the user's privacy can be protected because sensor data provided by one or more sensor components can be retained on the computing device. Additionally, by performing on-device activity recognition, the performance of the computing device can be improved, as further explained below.

[0004] By training one or more neural networks with differential privacy, noise is added so that individual examples in the training dataset of the one or more neural networks are hidden. By training one or more neural networks with differential privacy, the one or more neural networks provide strong mathematical guarantees. This strong mathematical guarantee means that the one or more neural networks do not learn or remember details about any particular user whose data was used to train the one or more neural networks. For example, differential privacy can prevent a malicious person from accurately determining whether specific data was used during the training of one or more neural networks, thereby protecting the privacy of the users whose data was used to train the one or more neural networks. Therefore, by training one or more neural networks with differential privacy, it may be possible to train one or more neural networks based on free-living data from users by protecting the privacy of the users whose data was used to train the one or more neural networks.

[0005] In some examples, the method includes receiving, by a computing device, motion data generated by and corresponding to movement sensed by one or more motion sensors, performing, by the computing device, on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and in response to recognizing the physical activity corresponding to the motion data, performing, by the computing device, an action associated with the physical activity.

[0006] In some examples, a computing device includes memory and one or more processors, the one or more processors being configured to receive motion data corresponding to and generated by movement sensed by one or more motion sensors, perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and in response to recognizing the physical activity corresponding to the motion data, perform an action associated with the physical activity.

[0007] In some examples, a computer-readable storage medium stores instructions that, when executed, cause one or more processors of a computing device to receive motion data corresponding to and generated by movement sensed by one or more motion sensors, perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and in response to recognizing the physical activity corresponding to the motion data, perform an action associated with the physical activity.

[0008] In some examples, an apparatus includes means for receiving motion data corresponding to and generated by movement sensed by one or more motion sensors, means for performing on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and means for performing an action associated with the physical activity in response to recognizing the physical activity corresponding to the motion data.

[0009] The details of one or more examples are illustrated in the accompanying drawings and described in the following description. Other features, objects, and advantages of the present disclosure will become apparent from the following description, the accompanying drawings, and the appended claims.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 4

Best Mode for Carrying Out the Invention

[0011] Detailed Description FIG. 1 is a conceptual diagram showing a computing device 110 that can perform on-device recognition of physical activity using one or more neural networks trained using differential privacy, in accordance with one or more aspects of the present disclosure. As shown in FIG. 1, the computing device 110 can be a mobile computing device such as a mobile phone (including a smartphone), a laptop computer, a tablet computer, a wearable computing device, a personal digital assistant (PDA), or any other computing device suitable for detecting a user's activity. In some examples, the computing device 110 can be a wearable computing device such as a computerized watch, a computerized fitness band / tracker, computerized eyewear, computerized headgear, computerized gloves, or any other type of mobile computing device that can be attached to and worn on a person's body or clothing.

[0012] In some examples, computing device 110 may include a presence-sensing display 112. The presence-sensing display 112 of computing device 110 may function as an input device for computing device 110 and as an output device. The presence-sensing display 112 may be implemented using various technologies. For example, the presence-sensing display 112 may function as an input device that uses a presence-sensing input component such as a resistive touch screen, a surface acoustic wave touch screen, a capacitive touch screen, a projected capacitance touch screen, a pressure-sensing screen, an acoustic pulse recognition touch screen, or another presence-sensing display technology. The presence-sensing display 112 may function as an output (e.g., display) device that uses any one or more display components such as a liquid crystal display (LCD), a dot matrix display, a light emitting diode (LED) display, a micro LED, an organic light-emitting diode (OLED) display, electronic ink, or a similar monochrome or color display capable of outputting visible information to a user of computing device 110.

[0013] Computing device 110 may also include one or more sensor components 114. In some examples, the sensor component may be an input component that obtains environmental information about the environment including the computing device 110. The sensor component may be an input component that obtains physiological information of the user of the computing device 110. In some examples, the sensor component may be an input component that obtains physical location information, motion information, and / or location information of the computing device 110. For example, the sensor component 114 may include, but is not limited to, a motion sensor (e.g., an accelerometer, a gyroscope, etc.), a heart rate sensor, a temperature sensor, a position sensor, a pressure sensor (e.g., a barometer), a proximity sensor (e.g., an infrared sensor), an ambient light detector, a location sensor (e.g., a global positioning system sensor), or any other type of sensing component. As further described in the present disclosure, the activity recognition module 118 may determine one or more physical activities performed by the user of the computing device 110 based on sensor data generated by one or more sensor components 114 and / or one or more sensor components 108 of the wearable computing device 100.

[0014] In some examples, the computing device 110 may be communicatively coupled to one or more wearable computing devices 100. For example, the computing device 110 may transmit and receive data to and from the wearable computing device 100 using one or more communication protocols. In some examples, the communication protocol may include Bluetooth®, near field communication, WiFi®, or any other suitable communication protocol.

[0015] In the example of FIG. 1, the wearable computing device 100 is a computerized watch. However, in other examples, the wearable computing device 100 can be a computerized fitness band / tracker, computerized eyewear, computerized headwear, computerized gloves, etc. In other examples, the wearable computing device 100 can be any type of mobile computing device that can be attached to and worn on a person's body or clothing.

[0016] As shown in FIG. 1, in some examples, the wearable computing device 100 can include an attachment component 102 and an electrical housing 104. The housing 104 of the wearable computing device 100 includes the physical portion of the wearable computing device that houses the combination of the hardware, software, firmware, and / or other electrical components of the wearable computing device 100. For example, FIG. 1 shows that within the housing 104, the wearable computing device 100 can include a sensor component 108 and a presence-aware display 106.

[0017] The presence-aware display 106 can be a presence-aware display as described with respect to the presence-aware display 112, and the sensor component 108 can be a sensor component as described with respect to the sensor component 114. The housing 104 can also include other hardware components and / or software modules not shown in FIG. 1, such as one or more processors, memory, an operating system, applications, etc.

[0018] The attachment component 102 can include a physical part of the wearable computing device that contacts the user's body (e.g., tissue, muscle, skin, hair, clothing, etc.) when the user is wearing the wearable computing device 100 (however, in some examples, a portion of the housing 104 may also contact the user's body). For example, if the wearable computing device 100 is a watch, the attachment component 102 can be a watch band that fits around the user's wrist and contacts the user's skin. In an example where the wearable computing device 100 is eyewear or headwear, the attachment component 102 can be a part of the frame of the eyewear or headwear that fits around the user's head, and when the wearable computing device 100 is gloves, the attachment component 102 can be the material of the gloves that adheres to the user's fingers and hands. In some examples, the wearable computing device 100 can be gripped and held from the housing 104 and / or the attachment component 102.

[0019] As shown in FIG. 1, computing device 110 may include an activity recognition module 118. The activity recognition module 118 may determine one or more activities of a user based on sensor data generated by one or more of the sensor components 114, or, if the computing device 110 is communicatively coupled to the wearable computing device 100, based on sensor data generated by one or more of the sensor components 108 of the wearable computing device 100, or based on sensor data generated by a combination of one or more of the sensor components 114 and one or more of the sensor components 108. Activities detected by the activity recognition module 118 may include, but are not limited to, cycling, running, a stationary state (e.g., sitting or standing still), ascending or descending stairs, walking, swimming, yoga, weightlifting, etc.

[0020] For purposes of illustration, the activity recognition module 118 is described as being implemented in and operating the computing device 110, but the wearable computing device 100 may also implement and / or operate an activity recognition module that includes the functionality described with respect to the activity recognition module 118. In some examples, when a user of the computing device 110 is wearing the wearable computing device 100, the activity recognition module 118 may determine one or more activities of the user based on sensor data generated by one or more of the sensor components 108 of the wearable computing device 100.

[0021] Generally, when the activity recognition module 118 of the computing device 110 determines one or more activities of a user based on sensor data generated by one or more sensor components 108 of the wearable computing device 100, both the computing device 110 and the wearable computing device 100 can be under the control of the same user. That is, the user wearing the wearable computing device 100 can be the same user who uses and / or controls the computing device 110. For example, the user wearing the wearable computing device 100 can also carry or hold the computing device 110, or can have the computing device 110 in the physical vicinity of the user (e.g., in the same room). Thus, the computing device 110 may not be a remote computing server or a cloud computing system that communicates with the wearable computing device 100 via, for example, the Internet to perform activity recognition remotely based on sensor data generated by one or more sensor components 108 of the wearable computing device 100.

[0022] In some examples, the activity recognition module implemented and / or operated in the wearable computing device 100 can transmit and receive information to and from the activity recognition module 118 via wired or wireless communication. The activity recognition module 118 may use such information received from the wearable computing device 100 in accordance with the techniques of the present disclosure as if the information was locally generated in the computing device 110.

[0023] The activity recognition module 118 may receive sensor data corresponding to one or more of the sensor components 114 and / or one or more sensor components 108, and may determine one or more physical activities in which the user is engaged. In some examples, the activity recognition module 118 may receive sensor data from the sensor processing module (as further described, for example, in FIG. 2). The sensor processing module may provide an interface between the hardware implementing the sensor component 114 and modules such as the activity recognition module 118 that further processes the sensor data. For example, the sensor processing module may generate sensor data representative of or corresponding to the output of the hardware implementing a particular sensor component. As an example, a sensor processing module for an accelerometer sensor component may generate sensor data including acceleration values along various axes of a coordinate system (e.g., the x-axis, y-axis, and z-axis).

[0024] The activity recognition module 118 may perform on-device recognition of a user's physical activity based on sensor data. That is, the activity recognition module 118 may identify the user's activity without transmitting information such as sensor data outside the device to a cloud computing system or the like. Instead, the activity recognition module 118 may implement one or more neural networks and use the one or more neural networks to determine the user's activity based on sensor data. Examples of sensor data that the activity recognition module 118 may use to perform on-device recognition of the user's physical activity may include motion data generated by one or more motion sensors. As described herein, the motion data may include acceleration values along various axes of a coordinate system generated by one or more multi-axis accelerometers, heart rate data generated by a heart rate sensor, location data generated by a location sensor (e.g., a global positioning system (GPS) sensor), oxygen saturation data (e.g., peripheral oxygen saturation) generated by an oxygen saturation sensor, and the like. As described herein, the motion sensors for generating the motion data may include such multi-axis accelerometers, heart rate sensors, location sensors, oxygen saturation sensors, gyroscopes, and the like.

[0025] Generally, the one or more neural networks implemented by the activity recognition module 118 may include a plurality of interconnected nodes. Each node may apply one or more functions to a set of input values corresponding to one or more features and provide one or more corresponding output values. The one or more features may be sensor data, and the one or more corresponding output values of the one or more neural networks may be a representation of the user's activity corresponding to the sensor data.

[0026] In some examples, one or more corresponding output values may include the probability of the user's activity. Thus, the activity recognition module 118 may determine the probability of the user's activity based on the characteristics of the user input using one or more neural networks, and based on the corresponding probability, may determine and output a display of the user's activity with the highest probability among the user's activities.

[0027] In some examples, one or more neural networks may be trained on-device by the activity recognition module 118 to more accurately determine the body activity with the highest probability among the user's body activities based on the characteristics. For example, one or more neural networks may include one or more learnable parameters or "weights" applied to the characteristics. The activity recognition module 118 may adjust these learnable parameters during training to improve the accuracy with which one or more neural networks determine the user's body activity corresponding to the sensor data and / or for any other suitable purpose, such as learning secondary characteristics of the way the activity is performed. For example, the activity recognition module 118 may adjust the learnable parameters based on whether the user provides user input to indicate the user's actual body activity.

[0028] In some examples, one or more neural networks can be trained outside of the device and further downloaded or installed on the computing device 110. In particular, one or more neural networks can be trained using differential privacy. That is, one or more neural networks can be trained in a way such as adding noise to the training data of the one or more neural networks to hide individual examples that provide strong mathematical guarantees, also referred to as privacy guarantees, from the training data. The strong mathematical guarantees mean that the one or more neural networks do not learn or remember details about any particular user for whom data was used to train the one or more neural networks. By training one or more neural networks using differential privacy, it is prevented that a malicious person can accurately determine whether specific data was used during the training of the one or more neural networks or how the one or more neural networks were trained, thereby protecting the privacy of the users for whom data was used to train the one or more neural networks.

[0029] One or more neural networks can be trained using training data that may include sensor data provided by a sensor component of a computing device used by a group of users. For example, the training data may include sensor data provided by a sensor component of a computing device used by a user while performing various physical activities such as cycling, running, stationary (e.g., not moving), walking, etc. One or more resulting trained neural networks can be quantized and compressed so that they can be installed on a mobile computing device such as computing device 110 for the one or more neural networks to perform activity recognition. For example, the model weights in one or more neural networks can be compressed to 8-bit integers for more efficient on-device inference.

[0030] Based on the sensor data, the activity recognition module 118 can generate one or more probabilities, such as a probability distribution corresponding to one or more physical activities, using one or more neural networks. That is, the activity recognition module 118 can generate respective probabilities for respective physical activities with respect to the sensor data over a particular period. The activity recognition module 118 can use sensor data generated by one or more of the sensor components 114, one or more of the sensor components 108, and / or any combination of the sensor component 114 and the sensor component 108. As an example, the activity recognition module 118 can generate, with respect to the sensor data, a probability of 0.6 that the user is walking, a probability of 0.2 that the user is running, a probability of 0.2 that the user is cycling, and a probability of 0.0 that the user is stationary (i.e., not moving).

[0031] The activity recognition module 118 may determine the physical activity corresponding to the sensor data based at least in part on one or more probabilities corresponding to one or more physical activities. For example, the activity recognition module 118 may determine whether the corresponding probability satisfies a threshold (e.g., whether the probability is greater than, greater than or equal to, less than, less than or equal to, or equal to the threshold) for the activity having the highest probability among the one or more probabilities (e.g., in the above example, the physical activity of walking with a probability of 0.6). The threshold may be a hard-coded value, a value set by the user, or a dynamically changing value. In some examples, the activity recognition module 118 may store or use different thresholds for different activities. In any case, the activity recognition module 118 may determine whether the probability for the threshold activity satisfies the threshold. In some examples, if the activity recognition module 118 determines that the probability for a physical activity satisfies the threshold, the activity recognition module 118 may determine that the user is likely to be involved in a specific physical activity, and thereby recognize the specific physical activity as the physical activity corresponding to the sensor data.

[0032] In some examples, in response to recognizing a physical activity, the activity recognition module 118 may cause the computing device 110 to perform one or more operations associated with the physical activity. The one or more operations may include operations such as collecting specific data from a specific sensor for the associated activity. In some examples, the activity recognition module 118 may collect data from a set of sensors and store the data as physical activity information associated with the recognized physical activity (e.g., in the physical activity information data store 228 shown in FIG. 2). The physical activity information may include data describing the specific physical activity. As examples of such data, but not limited to, a few examples of physical activity information include the time of the physical activity, the geographical location where the user performs the physical activity, the heart rate, the number of steps the user has walked, the speed or rate of change of the user's movement, the user's body temperature or the user's environment, the altitude or height at which the user performs the activity.

[0033] In some examples, the activity recognition module 118 may output to the user, in a graphical user interface, physical activity information associated with the recognized physical activity. In some examples, the activity recognition module 118 may perform an analysis on the physical activity information, such as determining various statistical metrics including a total value, an average value, etc. In some examples, the activity recognition module 118 may transmit the physical activity information that associates the physical activity information with the user's user account to a remote server. In still other examples, the activity recognition module 118 may notify one or more third-party applications. For example, a third-party fitness application can register with the activity recognition module 118, and the third-party fitness application can record the physical activity information about the user.

[0034] FIG. 2 is a block diagram showing further details of a computing device 210 that performs on-device activity recognition in accordance with one or more aspects of the present disclosure. The computing device 210 of FIG. 2 is described below as an example of the computing device 110 shown in FIG. 1. FIG. 2 shows only one particular example of the computing device 210, and many other examples of the computing device 210 may be used in other scenarios, may include a subset of the components included in the exemplary computing device 210, or may include additional components not shown in FIG. 2.

[0035] As shown in the example of FIG. 2, computing device 210 includes a presence - sensitive display 212, one or more processors 240, one or more input components 242, one or more communication units 244, one or more output components 246, and one or more storage components 248. The presence - sensitive display (PSD) 212 includes a display component 202 and a presence - sensitive input component 204. The input component 242 includes a sensor component 214. The storage component 248 of the computing device 210 also includes an activity recognition module 218, an activity detection model 220, an application module 224, a sensor processing module 226, and a physical activity information data store 228.

[0036] Communication channel 250 can interconnect each of components 240, 212, 202, 204, 244, 246, 242, 214, 248, 218, 220, 224, 226, and 228 for communication between (physical, communicable, and / or operable) components. In some examples, communication channel 250 can include a system bus, a network connection, an inter - process communication data structure, or any other means for communicating data.

[0037] One or more input components 242 of the computing device 210 can receive inputs. Examples of inputs include tactile input, voice input, and video input. The input component 242 of the computing device 210, in one example, includes a presence - sensitive display, a touch - screen, a mouse, a keyboard, a voice response system, a video camera, a microphone, or any other type of device for detecting input from a person or a machine.

[0038] One or more input components 242 include one or more sensor components 214. There are numerous examples of sensor components 214, including any input component configured to obtain environmental information regarding the situation surrounding computing device 210, and / or physiological information that defines the activity state and / or physical health of a user of computing device 210. In some examples, the sensor component can be an input component that obtains information about the physical location, movement, and / or location of computing device 210. For example, sensor component 214 can include one or more location sensors 214A (GPS component, Wi-Fi component, cellular component), one or more temperature sensors 214B, one or more motion sensors 214C (e.g., multi-axis accelerometer, gyro), one or more pressure sensors 214D (e.g., barometer), one or more ambient light sensors 214E, and one or more other sensors 214F (e.g., microphone, camera, infrared proximity sensor, hygrometer, etc.). Other sensors can include, by way of several other non-limiting examples, a heart rate sensor, a magnetometer, a glucose sensor, a humidity sensor, an olfactory sensor, a compass sensor, and a pedometer sensor.

[0039] One or more output components 246 of computing device 210 can generate an output. Examples of outputs include haptic output, audio output, and video output. The output component 246 of computing device 210, in one example, includes a presence-sensing display, a sound card, a video graphics adapter card, a speaker, a cathode ray tube (CRT) monitor, a liquid crystal display (LCD), or any other type of device for generating an output to a person or a machine.

[0040] One or more communication units 244 of the computing device 210 can communicate with external devices via one or more wired networks and / or wireless networks by transmitting and / or receiving network signals on one or more networks. Examples of communication unit 244 include network interface cards (e.g., Ethernet® cards, etc.), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device capable of transmitting and / or receiving information. Other examples of communication unit 244 can include shortwave radio waves, cellular data radio waves, wireless network radio waves, and universal serial bus (USB) controllers.

[0041] The presence sensing display (PSD) 212 of the computing device 200 includes a display component 202 and a presence sensing input component 204. The display component 202 may be a screen on which information is displayed by the PSD 212, and the presence sensing input component 204 may detect objects in the display component 202 and / or objects near the display component 202. As one exemplary range, the presence sensing input component 204 may detect an object such as a finger or a stylus within 2 inches from the display component 202. The presence sensing input component 204 may determine the location (e.g., (x,y) coordinates) of the display component 202 where the object is detected. In another exemplary range, the presence sensing input component 204 may detect an object within 6 inches from the display component 202 and may also be detectable in other ranges. The presence sensing input component 204 may use capacitive, inductive, and / or optical recognition techniques to determine the position of the display component 202 selected by the user's finger. In some examples, the presence sensing input component 204 may also provide an output to the user using tactile, audio, or visual stimuli, as described with respect to the display component 202. In the example of FIG. 2, the PSD 212 presents a user interface.

[0042] Although the presence sensing display 212 is shown as an internal component of the computing device 210, it may represent an external component that shares a data path with the computing device 210 to send and / or receive inputs and outputs. For example, in one instance, the PSD 212 represents a built-in component of the computing device 210 that is disposed within and physically connected to the external packaging of the computing device 210 (e.g., a screen on a mobile phone). In another example, the PSD 212 represents an external component of the computing device 210 that is disposed external to the computing device 210 and physically separated from the packaging of the computing device 210 (e.g., a monitor, projector, etc. that shares a wired and / or wireless data path with a tablet computer).

[0043] The PSD212 of the computing device 210 can receive haptic input from a user of the computing device 110. The PSD212 can receive a haptic input presentation by detecting one or more gestures from a user of the computing device 210 (e.g., the user touches or points to one or more locations of the PSD212 with a finger or a stylus pen). The PSD212 can present an output to the user. The PSD212 can present the output as a graphical user interface that can be associated with a function provided by the computing device 210. For example, the PSD212 can present various user interfaces (e.g., an electronic messaging application, a navigation application, an internet browser application, a mobile operating system, etc.) of components of a computing platform, an operating system, an application, or a service that is executed on or accessible by the computing device 210. The user can interact with each user interface to cause the computing device 210 to perform operations related to the function.

[0044] The PSD212 of the computing device 210 can detect two-dimensional and / or three-dimensional gestures as input from a user of the computing device 210. For example, the sensors of the PSD212 can detect the movement of a user (such as moving a hand, arm, pen, stylus, etc.) within a threshold distance of the sensors of the PSD212. The PSD212 can determine a two-dimensional or three-dimensional vector representation of the movement and correlate the vector representation with a gesture input having multiple dimensions (such as a wave, pinch, clap, pen stroke, etc.). In other words, the PSD212 can detect a multi-dimensional gesture without the user having to perform the gesture on or near the screen or surface on which the PSD212 outputs information for display. Instead, the PSD212 can detect a multi-dimensional gesture that is performed in or near a sensor that may or may not be disposed near the screen or surface on which the PSD212 outputs information for display.

[0045] One or more processors 240 may implement functions and / or execute instructions within computing device 210. For example, a processor 240 on computing device 210 may receive and execute instructions stored by storage component 248 that execute the functions of modules 218, 224, 226, and / or 228 and model 220. Instructions executed by processor 240 may cause computing device 210 to store information in storage component 248 during program execution. Examples of processor 240 include an application processor, a display controller, a sensor hub, and any other hardware configured to function as a processing unit. Processor 240 may execute instructions of modules 218, 224, 226, and / or 228 and model 220 to cause PSD 212 to render a portion of the content of the display data as one of the user interface screenshots in PSD 212. That is, modules 218, 224, 226, and / or 228 may be operable by processor 240 to perform various actions or functions of computing device 210.

[0046] One or more storage components 248 within computing device 210 may store information for processing during operation of computing device 210 (e.g., computing device 210 may store data accessed by modules 218, 224, 226, and / or 228 and model 220 during execution on computing device 210). In some examples, storage component 248 is temporary memory, meaning that the primary purpose of storage component 248 is not long-term storage. Storage component 248 on computing device 210 may be configured to store information temporarily as volatile memory and thus may not retain stored content when the power is turned off. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art.

[0047] In some examples, storage component 248 also includes one or more computer-readable storage media. The storage component 248 can be configured to store more information than volatile memory. The storage component 248 can further be configured to store information as non-volatile memory space for a long period of time and retain the information after power-on / off cycles. Examples of non-volatile memory include magnetic hard disks, optical disks, floppy (registered trademark) disks, flash memory, or forms of electrically programmable memory (EPROM) or electrically erasable and programmable (EEPROM) memory. The storage component 248 can store program instructions and / or information (e.g., data) associated with modules 218, 224, 226, and / or 228, model 220, and data store 280.

[0048] The application module 224 represents all the various individual applications and services that are executed in the computing device 210. A user of the computing device 210 can interact with an interface (e.g., a graphical user interface) associated with one or more application modules 224 to cause the computing device 210 to perform functions. There can be numerous examples of the application module 224, and it can include a fitness application, a calendar application, a personal assistant or prediction engine, a search application, a map or navigation application, a transportation service application (e.g., a bus or train tracking application), a social media application, a game application, an email application, a messaging application, an Internet browser application, or any other arbitrary application that can be executed in the computing device 210. The activity recognition module 218 is shown separately from the application module 224 but can be included within one or more of the application modules 224 (e.g., can be included within a fitness application).

[0049] As shown in FIG. 2, the computing device 210 may include a sensor processing module 226. In some examples, the sensor processing module 226 may receive the output from the sensor component 214 and generate sensor data representing the output. For example, each of the sensor components 214 may have a corresponding sensor processing module. As an example, the sensor processing module for the location sensor component 214A may generate GPS coordinate values within the sensor data (e.g., an object). Here, the GPS coordinates are based on the hardware output of the location sensor component 214A. As another example, the sensor processing module for the motion sensor component 214C may generate motion and / or acceleration values along various axes of the coordinate system in the motion data. Here, the motion and / or acceleration values are based on the hardware output of the motion sensor component 214C.

[0050] In FIG. 2, the activity recognition module 218 may receive motion data generated by one or more motion sensors. For example, the activity recognition module 218 may receive the motion data generated by the motion sensor 214C, and the motion data may be the sensor data generated by the motion sensor 214C, or the motion data generated by the sensor processing module 226 based on the sensor data generated by the motion sensor 214C. In some examples, the activity recognition module 218 may receive motion data generated by one or more motion sensors of a wearable computing device communicatively coupled to the computing device 210, such as the wearable computing device 100 shown in FIG. 1.

[0051] Motion data generated by one or more motion sensors, such as motion sensor 214C of a wearable computing device or one or more motion sensors, and received by activity recognition module 218 may be in the form of multi-axis accelerometer data, such as tri-axis accelerometer data. For example, the tri-axis accelerometer data may be in the form of a three-channel floating-point number that identifies the accelerations measured by one or more motion sensors along the x-axis, y-axis, and z-axis.

[0052] Motion data generated by one or more motion sensors and received by activity recognition module 218 can be motion data sensed by one or more motion sensors over a period of time. For example, the motion data can be motion data sensed by one or more sensors over a period of about 10 seconds. Assuming a sampling rate of 25 Hertz, the motion data can include a total of 256 samples of multi-axis accelerometer data.

[0053] In response to receiving motion data generated by one or more sensors, activity recognition module 218 can recognize the physical activity corresponding to the motion data generated by one or more motion sensors using one or more neural networks trained with differential privacy. To that end, activity recognition module 218 can use activity recognition model 220 to recognize the physical activity corresponding to the motion data generated by one or more motion sensors.

[0054] The activity recognition module 218 may include an activity recognition model 220 trained outside the device with differential privacy to recognize physical activities corresponding to motion data generated by one or more motion sensors. Examples of the activity recognition model 220 may include one or more convolutional neural networks, regression neural networks, or any other suitable artificial neural network trained with differential privacy. The activity recognition module 218 may take in, as input, motion data generated by one or more motion sensors, such as multi-axis accelerometer data sensed by one or more sensors over a period of time, and may output, in a form such as a probability distribution, one or more probabilities corresponding to one or more physical activities. That is, the activity recognition model 220 may generate, for each physical activity, a respective probability for the motion data regarding a plurality of physical activities over a period of time, such as a probability distribution of the motion data regarding each physical activity.

[0055] In some examples, in addition to recognizing physical activities corresponding to motion data generated by one or more motion sensors using the motion data generated by one or more motion sensors, the activity recognition module 218 and / or the activity recognition model 220 may use additional sensor data generated by one or more sensor components 214 to recognize physical activities corresponding to motion data generated by one or more motion sensors. For example, the activity recognition module 218 may use heart rate data generated by a heart rate sensor that measures the user's heart rate to reinforce the use of motion data generated by one or more motion sensors, thereby recognizing physical activities corresponding to motion data generated by one or more motion sensors.

[0056] For example, the activity recognition module 218 can receive corresponding heart rate data generated by a heart rate sensor during the same given period for motion data generated by one or more motion sensors over a given period, and based on the corresponding heart rate data generated by the heart rate sensor during the same given period, can adjust one or more probabilities corresponding to one or more physical activities determined by the activity recognition model 220. For example, when the activity recognition module 218 determines that the user's heart rate during a given period is within a specific range of the user's resting heart rate, it can increase the probability that the user is in a stationary state and decrease the probability that the user is active, such as by decreasing the probability of the user walking, cycling, or running. In another example, when the activity recognition module 218 determines that the user's heart rate during a given period is outside a specific range of the user's resting heart rate, it can increase the probability that the user is active and decrease the probability that the user is in a stationary state, such as by increasing the probability that the user is walking, cycling, or running.

[0057] The activity recognition module 218 can recognize a physical activity corresponding to motion data generated by one or more motion sensors, based at least in part on one or more determined probabilities corresponding to one or more physical activities. In some examples, the activity recognition module 218 can determine that the physical activity associated with the highest probability among one or more probabilities corresponding to one or more physical activities determined by the activity recognition model 220 is the physical activity corresponding to the motion data generated by one or more motion sensors. For example, if the activity recognition model 220 determines that the probability of the physical activity being walking is 0.6, the probability of the physical activity being running is 0.2, the probability of the physical activity being cycling is 0.2, and the probability of the physical activity being in a stationary state (i.e., not moving) is 0.0, the activity recognition module 218 can recognize that the physical activity corresponding to the motion data is walking.

[0058] In some examples, when the probability associated with a physical activity meets a threshold (e.g., when the probability is greater than the threshold, above the threshold, below the threshold, less than the threshold, or equal to the threshold), the activity recognition module 218 may determine that the physical activity associated with the highest probability among one or more probabilities corresponding to one or more physical activities determined by the activity recognition model 220 corresponds to the physical activity generated by one or more motion sensors. If the activity detection module 218 determines that the probability regarding the physical activity meets the threshold, the activity recognition module 218 may determine that the user is likely to be involved in a specific physical activity, and thereby may recognize the specific physical activity as the physical activity corresponding to the motion data.

[0059] In response to recognizing the physical activity corresponding to the motion data, the activity recognition module 218 may perform one or more operations associated with the physical activity. In some examples, in response to recognizing the physical activity corresponding to the motion data, the activity recognition module 218 may perform fitness tracking of the user, such as by tracking physical activity information associated with the physical activity, including the time of the physical activity, the geographical location where the user performs the physical activity, the heart rate, the number of steps the user has walked, the speed or rate of change of the user's movement, the user's body temperature or the user's environment, the altitude or height at which the user is performing the activity.

[0060] In some examples, the activity recognition module 218 may store the display of the physical activity together with the physical activity information associated with the physical activity, such as the time of the physical activity, the geographical location where the user performs the physical activity, the heart rate, the number of steps the user has walked, the speed or rate of change of the user's movement, the user's body temperature or the user's environment, the altitude or height at which the user is performing the activity, in a storage component 248 such as the physical activity information data store 228.

[0061] In some examples, the activity recognition module 218 may perform analyses such as determining various statistical metrics including a sum value, an average value, etc. on the physical activity information associated with the physical activity. In some examples, the activity recognition module 218 may send the physical activity information associated with the physical activity to a remote server that associates the physical activity information with the user's user account. In still other examples, the activity recognition module 218 may notify one or more third-party applications. For example, a third-party fitness application can register with the activity recognition module 218, and the third-party fitness application can record the physical activity information about the user.

[0062] Figures 3A-3E are conceptual diagrams showing aspects of an exemplary machine learning model according to an exemplary implementation of the present disclosure. Figures 3A-3E are described below in the context of the activity recognition module 218 of FIG. 2. For example, in some instances, the machine learning model 300 can be an example of the activity recognition model 220 of FIG. 2 as referred to below.

[0063] FIG. 3A shows a conceptual diagram of an exemplary machine learning model according to an exemplary implementation of the present disclosure. As shown in FIG. 3A, in some implementations, the machine learning model 300 is trained to receive one or more types of input data and, in response thereto, provide one or more types of output data. Thus, FIG. 3A shows the machine learning model 300 that performs inference. For example, the input data received by the machine learning model 300 may be motion data such as sensor data generated by a multi-axis accelerometer, and the output data provided by the machine learning model 300 may be the user's activity corresponding to the motion data.

[0064] The input data may include one or more features associated with an instance or an example. In some implementations, one or more features associated with an instance or an example may be encoded into a feature vector. In some implementations, the output data may include one or more predictions. A prediction may also be referred to as an inference. Thus, when there are features associated with a particular instance, the machine learning model 300 may output a prediction regarding such an instance based on those features.

[0065] The machine learning model 300 may be or include one or more of various different types of machine learning models. In particular, in some implementations, the machine learning model 300 can perform classification, regression, clustering, anomaly detection, recommendation generation, and / or other tasks.

[0066] In some implementations, the machine learning model 300 can perform various types of classification based on the input data. For example, the machine learning model 300 can perform binary classification or multi-class classification. In binary classification, the output data may include classifying the input data into one of two different classes. In multi-class classification, the output data may include classifying the input data into one (or more) of three or more classes. These classifications can be single-label or multi-label. The machine learning model 300 can perform discrete categorical classification where the input data is simply classified into one or more classes or categories.

[0067] In some implementation examples, in the classification that the machine learning model 300 can perform, the machine learning model 300 provides a numerical value that describes the degree to which it is considered that the input data should be classified into the class corresponding to each one or more classes. In some cases, the numerical values provided by the machine learning model 300 may be referred to as "reliability scores" that indicate the respective reliabilities associated with the classification of the input into each class. In some implementation examples, the reliability scores can be compared with one or more thresholds to render discrete category predictions. In some implementation examples, only a specific number (e.g., one) of classes with the relatively maximum reliability scores can be selected to render discrete category predictions.

[0068] The machine learning model 300 can output probabilistic classification. For example, for a sample input, the machine learning model 300 can predict a probability distribution over a set of classes. Therefore, instead of only outputting the class with the highest likelihood that the sample input should belong to, for each class, the machine learning model 300 can output the probability that the sample input belongs to such a class. In some implementation examples, the probability distributions over all possible classes can be summed to 1. In some implementation examples, a Softmax function, or other types of functions or layers, can be used to squash a set of real values respectively associated with the possible classes into a set of real values within a range (0, 1) that sums to 1.

[0069] In some examples, the probabilities provided by the probability distribution can be compared with one or more thresholds to render discrete category predictions. In some implementation examples, only a specific number (e.g., one) of classes with the relatively maximum predicted probabilities can be selected to render discrete category predictions.

[0070] When the machine learning model 300 performs classification, the machine learning model 300 can be trained using supervised learning techniques. For example, the machine learning model 300 can be trained on a training dataset that includes training examples labeled as belonging to (or not belonging to) one or more classes. Further details regarding supervised training techniques are provided below in the description of FIGS. 3B-3E.

[0071] In some embodiments, the machine learning model 300 can perform regression to provide output data in the form of continuous numerical values. The continuous numerical values can correspond to any number of different metrics or numerical representations, including, for example, currency values, scores, or other numerical representations. As an example, the machine learning model 300 can perform linear regression, polynomial regression, or non-linear regression. As an example, the machine learning model 300 can perform simple regression or multiple regression. As described above, in some embodiments, a Softmax function or other function or layer can be used to map a set of real values associated with two or more possible classes to a set of real values within a range that sums to 1 (0, 1).

[0072] The machine learning model 300 can, in some cases, function as an agent within an environment. For example, the machine learning model 300 can be trained using reinforcement learning, which is described in more detail below.

[0073] In some embodiments, the machine learning model 300 can be a parametric model, while in other embodiments, the machine learning model 300 can be a non-parametric model. In some embodiments, the machine learning model 300 can be a linear model, while in other embodiments, the machine learning model 300 can be a non-linear model.

[0074] As described above, the machine learning model 300 can be or include one or more of various different types of machine learning models. Examples of such various types of machine learning models are provided below for illustration. One or more of the model examples described below can be used (e.g., in combination) to provide output data in response to input data. Additional models other than the exemplary models provided below can also be used in the same way.

[0075] In some implementations, the machine learning model 300 can be or include one or more classifier models, such as a linear classification model, a quadratic classification model, etc. The machine learning model 300 can be or include one or more regression models, such as a simple linear regression model, a multiple linear regression model, a logistic regression model, a stepwise regression model, a multivariate adaptive regression spline, a local estimation dispersion smoothing model, etc.

[0076] In some implementations, the machine learning model 300 can be or include one or more artificial neural networks (also simply referred to as neural networks). A neural network can include a group of connected nodes that can also be referred to as neurons or perceptrons. A neural network can be organized into one or more layers. A neural network that includes multiple layers can be referred to as a "deep" network. A deep network can include an input layer, an output layer, and one or more hidden layers disposed between the input layer and the output layer. The multiple nodes of a neural network can be fully connected or partially connected.

[0077] The machine learning model 300 can be or include one or more feedforward neural networks. In a feedforward network, the connections between nodes do not form cycles. For example, each connection can connect a node from a previous layer to a node from a subsequent layer.

[0078] In some cases, the machine learning model 300 can be or include one or more recurrent neural networks. In some cases, at least some of the plurality of nodes of the recurrent neural network can form a cycle. The recurrent neural network can be particularly useful for processing essentially continuous input data. In particular, in some cases, the recurrent neural network can pass or retain information from a previous part of an input data sequence to a later part of the input data sequence by using recurrent or directed cyclic node connections.

[0079] In some examples, the continuous input data can include time series data (e.g., sensor data over time, or images captured at various times). For example, a recurrent neural network can analyze sensor data over time to detect or predict a swipe direction, recognize handwriting, etc. The continuous input data can include words in a text (e.g., for natural language processing, speech detection or processing, etc.), tones in a piece of music, continuous actions taken by a user (e.g., to detect or predict continuous application usage), continuous object states, etc.

[0080] Examples of recurrent neural networks include long short-term (LSTM) recurrent neural networks, gated recurrent units, bidirectional recurrent neural networks, continuous-time recurrent neural networks, neural history compressors, echo state networks, Elman networks, Jordan networks, recurrent neural networks, Hopfield networks, fully recurrent networks, sequence-to-sequence architectures, etc.

[0081] In some implementation examples, the machine learning model 300 can be or include one or more convolutional neural networks. In some cases, the convolutional neural network can include one or more convolutional layers that perform convolutions over the input data using learned filters.

[0082] The filter can also be referred to as a kernel. The convolutional neural network can be particularly useful for vision-related problems such as when the input data includes images such as still images or videos. However, convolutional neural networks are also applicable to natural language processing.

[0083] In some examples, the machine learning model 300 can be or include one or more generative networks such as, for example, a generative adversarial network. New data such as new images or other content can be generated using the generative network.

[0084] The machine learning model 300 can be or include an autoencoder. In some cases, the purpose of the autoencoder is typically to learn a representation (e.g., a lower-dimensional encoding) for a set of data, for the purpose of dimensionality reduction. For example, in some cases, the autoencoder can attempt to encode the input data and then provide output data that reconstructs the input data from this encoding. In recent years, the concept of the autoencoder has been more widely used to learn generative models of data. In some cases, the autoencoder can include an additional loss beyond the reconstruction of the input data.

[0085] The machine learning model 300 can be, or can include, one or more other forms of artificial neural networks, such as, for example, a deep Boltzmann machine, a deep belief network, a stacked autoencoder, etc. By combining (e.g., stacking) any of the neural networks described herein, a more complex network can be formed.

[0086] One or more neural networks can be used to provide an embedding based on input data. For example, the embedding can be a representation of knowledge abstracted from the input data into one or more learned dimensions. In some cases, the embedding can be a useful source for identifying related entities. In some cases, the embedding can be extracted from the output of the network, and in other cases, the embedding can be extracted from any hidden node or layer of the network (e.g., a layer close to but not the final layer of the network). The embedding can be useful for performing, for example, auto - suggestions for the next video, product recommendations, entity or object recognition, etc. In some cases, the embedding is a useful input for a downstream model. For example, the embedding can be useful for generalizing input data (e.g., a search query) for a downstream model or processing system.

[0087] In some implementations, the machine learning model 300 can perform one or more dimensionality reduction techniques such as, for example, principal component analysis, kernel principal component analysis, graph - based kernel principal component analysis, principal component regression, partial least squares regression, Sammon mapping, multidimensional scaling, projection pursuit, linear discriminant analysis, mixture discriminant analysis, quadratic discriminant analysis, generalized discriminant analysis, flexible discriminant analysis, auto - encoding, etc.

[0088] In some implementation examples, the machine learning model 300 can execute or apply one or more reinforcement learning techniques such as Markov decision processes, dynamic programming, Q - functions or Q - learning, value - function approaches, deep Q - networks, differentiable neural computers, asynchronous advantage actor - critics, deterministic policy gradients, etc.

[0089] In some implementation examples, the machine learning model 300 can be an autoregressive model. In some cases, the autoregressive model can define that the output data linearly depends on its own previous values and probabilistic terms. In some cases, the autoregressive model can take the form of a probabilistic difference equation. As an example of an autoregressive model, there is WaveNet, a generative model for raw audio.

[0090] In some implementation examples, the machine learning model 300 can include or form part of an ensemble of multiple models. As an example, bootstrap aggregating, which may also be referred to as "bagging", can be performed. In bootstrap aggregating, the training data set is split into several subsets (e.g., by random sampling with replacement), and multiple models are trained for each number of subsets. At inference time, the outputs of each of the multiple models can be combined (e.g., by averaging, voting, or other techniques) and used as the output of the ensemble.

[0091] One example of an ensemble is a random forest, which may also be referred to as a random decision forest. A random forest is an ensemble learning method for classification, regression, and other tasks. A random forest is generated by creating multiple decision trees during training. In some cases, at inference time, the class that is the mode (for classification) or average prediction (for regression) of the individual trees can be used as the output of the forest. Random decision forests can correct the tendency of decision trees to overfit their training sets.

[0092] Another exemplary ensemble technique is stacking, which in some instances may be referred to as stacked generalization. Stacking involves training a combiner model to blend or combine the predictions of several other machine learning models. Thus, multiple (e.g., of the same type or different types) machine learning models can be trained based on training data. Additionally, the combiner model can be trained to take as input the predictions from other machine learning models and generate a final inference or prediction accordingly. In some cases, a single-layer logistic regression model can be used as the combiner model.

[0093] Another exemplary ensemble technique is boosting. Boosting can involve iteratively training weak models and then gradually building an ensemble by adding them to a final strong model. For example, in some cases, each new model can be trained to emphasize training examples that were misinterpreted (e.g., misclassified) by the previous model. For example, the weights associated with each such misinterpreted example can be increased. One common implementation of boosting is AdaBoost, which may also be referred to as Adaptive Boosting. Examples of other boosting techniques include LPBoost, TotalBoost, BrownBoost, xgboost, MadaBoost, LogitBoost, gradient boosting, and the like. Additionally, any of the above-described models (e.g., regression models and artificial neural networks) can be combined to form an ensemble. As an example, the ensemble can include a top-level machine learning model or heuristic function for combining and / or weighting the outputs of the models that form the ensemble.

[0094] In some implementations, multiple machine learning models (e.g., that form an ensemble) can be linked and trained together (e.g., by continuously backpropagating errors through the model ensemble). However, in some implementations, only a subset (e.g., one) of the models trained together is used for inference.

[0095] In some implementations, the machine learning model 300 can be used to preprocess input data for subsequent input to another model. For example, the machine learning model 300 can perform dimensionality reduction techniques and embeddings (e.g., matrix factorization, principal component analysis, singular value decomposition, word2vec / GLOVE, and / or related approaches), clustering, and even classification and regression for downstream consumption. Many of these techniques have been described above, but are further described below.

[0096] As described above, the machine learning model 300 can be trained or configured to receive input data and provide output data accordingly. The input data can include various types, forms, or variations of input data. By way of example, in various implementations, the input data can include content (or a portion of the content) initially selected by the user, such as the content of a document or image selected by the user, a link indicating a user selection, a link within a user selection related to other files available on a device or in the cloud, features describing metadata of the user selection, and the like. Additionally, with user permission, the input data can include the context of user usage obtained from the app itself or other sources. Examples of usage context can include the scope of sharing (whether to share publicly, with a large group, individually, or with a specific person), the context of sharing, and the like. Additional input data can include the state of the device, such as the device's location, the apps operating on the device, etc., if permitted by the user.

[0097] Additionally, with user permission, the input data can include the context of user usage obtained from the app itself or other sources. Examples of usage context can include the scope of sharing (whether to share publicly, with a large group, individually, or with a specific person), the context of sharing, and the like. Additional input data can include the state of the device, such as the device's location, the apps operating on the device, etc., if permitted by the user.

[0098] In some implementations, the machine learning model 300 can receive input data and use the input data in its raw form. In some implementations, the raw input data can be preprocessed. Thus, in addition to or instead of the raw input data, the machine learning model 300 can receive and use preprocessed input data.

[0099] In some implementation examples, preprocessing the input data may include extracting one or more additional features from the raw input data. For example, feature extraction techniques can be applied to the input data to generate one or more new additional features. Examples of feature extraction techniques include edge detection, corner detection, blob detection, ridge detection, scale-invariant feature transform, motion detection, optical flow, Hough transform, and the like.

[0100] In some implementation examples, the extracted features may include or be derived from a transformation of the input data to another domain and / or dimension. As an example, the extracted features may include or be derived from a transformation of the input data to the frequency domain. For example, wavelet transform and / or fast Fourier transform can be performed on the input data to generate additional features.

[0101] In some implementation examples, the extracted features may include statistics calculated from the input data or a specific part or dimension of the input data. Exemplary statistics include mode, mean, maximum, minimum, or other metrics for the input data or some of its parts.

[0102] In some implementation examples, as described above, the input data can be essentially continuous. In some cases, continuous input data can be generated by sampling or segmenting a stream of input data. As an example, frames can be extracted from a video. In some implementation examples, continuous data can be made discontinuous by summarization.

[0103] As another exemplary preprocessing technique, a part of the input data can be attributed. For example, additional synthetic input data can be generated by interpolation and / or extrapolation.

[0104] As another exemplary preprocessing technique, some or all of the input data can be scaled, normalized, regularized, generalized, and / or standardized. Exemplary regularization techniques include ridge regression, least absolute shrinkage and selection operator (LASSO), elastic net, least angle regression, cross-validation, L1 regularization, L2 regularization, and the like. As an example, some or all of the input data can be normalized by subtracting the mean across the eigenvalues of a given dimension from each of the individual eigenvalues and then dividing by the standard deviation or other metric.

[0105] As another exemplary preprocessing technique, some or all of the input data can be quantized or discretized. In some cases, qualitative features or variables included in the input data can be converted to quantitative features or variables. For example, one-hot encoding can be performed.

[0106] In some examples, dimensionality reduction techniques can be applied to the input data prior to input to the machine learning model 300. Some examples of the dimensionality reduction techniques described above include, for example, principal component analysis, kernel principal component analysis, graph-based kernel principal component analysis, principal component regression, partial least squares regression, Sammon mapping, multidimensional scaling, projection pursuit, linear discriminant analysis, mixture discriminant analysis, quadratic discriminant analysis, generalized discriminant analysis, flexible discriminant analysis, autoencoding, and the like.

[0107] In some implementations, during training, the input data can be intentionally distorted in a number of ways to enhance the robustness, generalization, or other qualities of the model. Exemplary techniques for distorting the input data include adding noise, changing color, shadow or hue, magnification, segmentation, amplification, and the like.

[0108] The machine learning model 300 can provide output data in response to receiving input data. The output data can include various types, forms, or variations of output data. As an example, in various implementations, the output data can be appropriately shareable along with an initial content selection and can include content stored locally on a user device or in the cloud.

[0109] As described above, in some implementations, the output data can include various types of classification data (e.g., binary classification, multi-class classification, single-label, multi-label, discrete classification, regression classification, probabilistic classification, etc.), or can include various types of regression data (e.g., linear regression, polynomial regression, non-linear regression, simple regression, multiple regression, etc.). In other examples, the output data can include clustering data, anomaly detection data, recommendation data, or any other form of the output data described above.

[0110] In some implementations, the output data can affect downstream processes or decision-making. As an example, in some implementations, the output data can be interpreted and / or acted upon by a rule-based regulator.

[0111] The present disclosure provides a system and method that includes or utilizes one or more machine learning models to propose content stored locally on a user's device or in the cloud that is appropriately shareable along with an initial content selection based on characteristics of the initial content selection. By combining any of the various types or forms of the input data described above with any of the various types or forms of the machine learning models described above, any of the various types or forms of the output data described above can be provided.

[0112] The systems and methods of the present disclosure can be implemented by or executed on one or more computing devices. Exemplary computing devices include user computing devices (e.g., laptops, desktops, and mobile computing devices such as tablets, smartphones, wearable computing devices, etc.), embedded computing devices (e.g., devices embedded in vehicles, cameras, image sensors, industrial machinery, satellites, game consoles or controllers, or household appliances such as refrigerators, thermostats, energy meters, home energy managers, smart home assistants, etc.), other computing devices, or combinations thereof.

[0113] FIG. 3B shows a conceptual diagram of a computing device 310, which is an example of the computing device 110 of FIG. 1. The computing device 310 includes a processing component 302, a memory component 304, and a machine learning model 300. The computing device 310 can locally (i.e., on-device) store and implement the machine learning model 300. Thus, the machine learning model 300 can be stored in and / or locally implemented by an embedded device or a user computing device, such as a mobile device. By using the output data obtained by locally implementing the machine learning model 300 in an embedded device or a user computing device, the performance of the embedded device or the user computing device (e.g., an application implemented by the embedded device or the user computing device) can be improved.

[0114] FIG. 3C shows a conceptual diagram of an exemplary computing device that communicates with an exemplary training computing system including a model trainer. FIG. 3C includes a client device 310 that communicates with a training device 320 via a network 330. The client device 310 is an example of the computing device 110 of FIG. 1. The machine learning model 300 described herein can be stored and / or implemented in one or more computing devices such as the client device 310 by being trained and provided by a training computing system such as the training device 320. For example, the model trainer 372 is executed locally in the training device 320. In some examples, the training device 320 including the model trainer 372 can be included in or separate from the client device 310 or any other computing device that implements the machine learning model 300.

[0115] The computing device 310 that implements the machine learning model 300 or other aspects of the present disclosure and the training device 320 that trains the machine learning model 300 can include some hardware components that enable the execution of the techniques described herein. For example, the computing device 310 can include one or more memory devices that store part or all of the machine learning model 300. For example, the machine learning model 300 can be a structured numerical representation stored in memory. The one or more memory devices can also include instructions for implementing the machine learning model 300 or instructions for performing other operations. Examples of memory devices include RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof.

[0116] Computing device 310 may also include one or more processing devices that implement some or all of machine learning model 300 and / or perform other related operations. Exemplary processing devices include a central processing unit (CPU), a visual processing unit (VPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), a neural processing engine, a core of a CPU, VPU, GPU, TPU, NPU, or other processing device, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coprocessor, a controller, or a combination of the above-described processing devices. The processing device can be embedded, for example, within other hardware components such as an image sensor, an accelerometer.

[0117] Training device 320 may execute graph processing techniques or other machine learning techniques using one or more machine learning platforms, frameworks, and / or libraries, such as TensorFlow, Caffe / Caffe2, Theano, Torch / PyTorch, MXnet, CNTK, etc. In some implementations, machine learning model 300 can be trained offline or online. (Also known as batch learning) In offline training, machine learning model 300 is trained against an entire static set of training data. In online learning, machine learning model 300 is continuously trained (or retrained) as new training data becomes available (e.g., while the model is being used to perform inference).

[0118] The model trainer 372 can perform intensive training of the machine learning model 300 (e.g., based on an intensively stored dataset). In other examples, distributed training techniques such as distributed training and federated learning can be used to train, update, or personalize the machine learning model 300.

[0119] The machine learning model 300 described herein can be trained according to one or more of a variety of different training types or techniques. For example, in some examples, the machine learning model 300 can be trained by the model trainer 372 using supervised learning. Here, the machine learning model 300 is trained on a training dataset that includes cases or examples with labels. These labels can be applied manually by an expert, created by crowdsourcing, or provided by other techniques (e.g., physics-based or complex mathematical models). In some examples, training examples can be provided by a user computing device if the user consents. In some examples, this process can be referred to as model personalization.

[0120] When the training device 320 finishes training the machine learning model 300, the machine learning model 300 can be installed on the client device 310. For example, the training device 320 can transfer the machine learning model 300 to the client device 310 via the network 330, or the machine learning model 300 can be installed on the client device 310 during the manufacture of the client device 310. In some examples, when the machine learning model 300 is trained on the training device 320, the training device 320 can perform post-training weight quantization, such as by using the TensorFlow Lite library, to compress the model weights, such as by compressing the model weights to 8-bit integers, enabling the client device 310 to perform more efficient on-device inference using the machine learning model 300.

[0121] FIG. 3D shows a conceptual diagram of a training process 340 that is an exemplary training process in which a training device 340 can train a machine learning model 300 on training data 341 that includes exemplary input data 342 having a label 343. The training process 340 is an example of a training process, and other training processes may be used as well.

[0122] The training data 341 used by the training process 340 can include anonymized usage logs of the shared flow, such as bundled content portions that already belong together and have been identified as belonging together from shared content items, such as entities in a knowledge graph, if the user permits the use of such data for training. In some implementations, the training data 341 can include examples of input data 342 to which a label 343 corresponding to the output data 344 is assigned.

[0123] In accordance with aspects of the present disclosure, the machine learning model 300 is trained with differential privacy. That is, the machine learning model 300 can be trained in a way that provides strong mathematical guarantees. The strong mathematical guarantees mean that the machine learning model 300 does not learn or remember details about any particular user for whom data was used to train the machine learning model 300 (e.g., as part of the exemplary input data 342 and / or training data 341), thereby preventing a malicious person from accurately determining whether specific data was used during the training of the machine learning model 300, and thereby protecting the privacy of the users whose data was used to train the machine learning model 300.

[0124] To train the machine learning model 300 with differential privacy, the training device 320 can add noise to the training data 341 to hide individual examples in the training data 341 of the machine learning model 300. To this end, the training device 320 can train the machine learning model 300 using a differential privacy framework such as TensorFlow Privacy. By training the machine learning model 300 using a differential privacy framework, the training device 320 adds noise to the training data 341 to provide strong privacy guarantees. The strong privacy guarantees mean that the machine learning model 300 does not learn or remember details about any particular user for whom data was used to train the machine learning model 300 (e.g., as part of the input data 342 and / or training data 341).

[0125] To provide a strong mathematical guarantee that the machine learning model 300 does not learn or remember details about any particular user for whom data is used to train the machine learning model 300, the differential privacy framework used to train the machine learning model 300 using differential privacy can be described using an epsilon (ε) parameter and a delta (δ) parameter. The value of the epsilon parameter can be a measure of the strength of the privacy guarantee and can be an upper bound (i.e., an upper limit) on how much the probability of a particular output can be increased by including or removing a single training example from the training data 341. Generally, the value of the epsilon parameter can be less than 10, or less than 1, for a more stringent privacy guarantee. The value of the delta parameter can be a bound on the probability that the privacy guarantee is not maintained. In some examples, the delta parameter may be set to a value that is the reciprocal of the size of the training data 341, or to a value that is less than the reciprocal of the size of the training data 341.

[0126] The training device 320 can specify a target value for the delta parameter used by the differential privacy framework to train the machine learning model 300 using differential privacy. The training device 320 can determine the value of the epsilon parameter used by the differential privacy framework to train the machine learning model 300 using differential privacy based on the target value of the delta parameter specified for the machine learning model 300 and a given set of hyperparameters.

[0127] To determine the value of the epsilon parameter for the machine learning model 300, in some examples, the training device 320 can compute the Renyi divergence (also known as Renyi entropy) of the training data for the machine learning model 300 and the neighborhood of the training data. The neighborhood can be a distribution of training data that is very similar to the training data (e.g., having a Hamming distance of 1 or another similar value). The Renyi divergence of the data can be essentially a generalized Kullback-Leibler (KL) divergence parameterized by the alpha value. Here, the KL divergence can be exactly derived from the Renyi divergence when the alpha value is set to 1.

[0128] When the training data has a large Renyi divergence, the distribution of the training data can shift significantly by changing a single variable. Thus, while it may be possible to identify an extra example, which could potentially compromise overall privacy, it may be easier to search for information in such a manner. When the training data has a small Renyi divergence, the addition of an extra example may not result in a significant change to the data, making it more difficult to identify whether this extra example was added to the data. As described above, since the value of the delta parameter can be set to the reciprocal of the size of the training data 341, it may be possible for the training data for the machine learning model 300 to have a relatively small Renyi divergence.

[0129] As described above, the machine learning model 300 may include one or more neural networks that perform activity recognition. That is, the machine learning model 300 may obtain motion data corresponding to movement as input, and in response, may output a display of physical activity that matches the input motion data. The motion data input to the machine learning model 300 may include multi-axis accelerometer data generated by one or more motion sensors such as one or more accelerometers. In some examples, the one or more accelerometers may generate tri-axial accelerometer data. The tri-axial accelerometer data may be in the form of three channels of floating-point numbers specifying the acceleration forces measured by one or more accelerometers along the x-axis, y-axis, and z-axis.

[0130] The machine learning model 300 may be trained to perform activity recognition based on motion data over a period of time generated by one or more motion sensors. For example, the motion data of the machine learning model 300 may include tri-axial accelerometer data for about 10 seconds. Assuming a sampling rate of 25 Hz over about 10 seconds, the machine learning model 300 may receive a total of 256 samples of tri-axial accelerometer data. In this case, each sample of the tri-axial accelerometer data may be in the form of three channels of floating-point numbers specifying the acceleration forces measured by one or more accelerometers along the x-axis, y-axis, and z-axis, and may be trained to determine the physical activity corresponding to the 256 samples of tri-axial accelerometer data.

[0131] Accordingly, each exemplary input data of the exemplary input data 342 may be motion data over a specified period of time, such as 10 seconds of tri-axial accelerometer data including a total of 256 samples of tri-axial accelerometer data. The exemplary input data 342 may include millions of individual exemplary input data. Further, the delta value of the machine learning model 300 described above may be the reciprocal of the number of individual exemplary input data in the input data 342.

[0132] As described above, the training data 341 includes exemplary input data 342 having a label 343, and thus each individual exemplary input data of the exemplary input data 342 has the label of the label 343. In order to train the machine learning model 300 to perform activity recognition, each individual exemplary input data of the exemplary input data 342 has a label indicating the activity corresponding to the individual exemplary input data. For example, if the machine learning model 300 is trained to recognize motion data as one of cycling, running, walking, or a stationary state (e.g., sitting still, standing still, and not moving), each of the exemplary input data corresponding to cycling may have a label indicating cycling, each of the exemplary input data corresponding to running may have a label indicating running, each of the exemplary input data corresponding to walking may have a label indicating walking, and each of the exemplary input data corresponding to the stationary state may have a label indicating the stationary state.

[0133] Exemplary input data 342 can be motion data generated by a motion sensor of a computing device carried or worn by a user in response to the user performing various physical activities that the machine learning model 300 is trained to recognize. For example, exemplary input data corresponding to cycling can be motion data generated by a motion sensor of a wearable computing device worn by the user during cycling, and exemplary input data corresponding to walking can be motion data generated by a motion sensor of a wearable computing device worn by the user during walking, and exemplary input data corresponding to running can be motion data generated by a motion sensor of a wearable computing device worn by the user during running, and exemplary input data corresponding to a stationary state can be motion data generated by a motion sensor of a wearable computing device worn by the user while the user is in a stationary state.

[0134] The training data 341 may include approximately equal numbers of exemplary input data for each of a plurality of physical activities that are trained to be recognized by the machine learning model 300. For example, if the machine learning model 300 is trained to recognize motion data as one of cycling, running, walking, or a stationary state, approximately 1 / 4 of the exemplary input data 342 may be exemplary input data corresponding to cycling, and approximately 1 / 4 of the exemplary input data 342 may be exemplary input data corresponding to running, and approximately 1 / 4 of the exemplary input data 342 may be exemplary input data corresponding to walking, and approximately 1 / 4 of the exemplary input data 342 may be exemplary input data corresponding to being in a stationary state. The training device 320 can ensure that the training data 341 includes approximately equal numbers of exemplary input data for each of a plurality of physical activities that are trained to be recognized by the machine learning model 300 by resampling the training data 341 to address any imbalance in any of the exemplary input data 342.

[0135] In some examples, the training device 320 may transform the exemplary input data 342 in all three directions (i.e., the x-axis, y-axis, and z-axis) with random rotations without constraints to train the machine learning model 300 to provide rotational invariance when performing activity recognition. That is, the training device 320 may apply a random rotation to the exemplary input data, such as by applying a random rotation to the acceleration forces along the x-axis, y-axis, and z-axis specified by the exemplary input data for each exemplary input data of the exemplary input data, and may include the randomly rotated exemplary input data as part of the training data 341 used to train the machine learning model 300. Since the position and orientation of the motion sensor of the computing device may depend on how such a computing device is carried or worn by a user and may change over time, the machine learning model 300 trained to provide rotational invariance will be able to accurately perform activity recognition regardless of the position and orientation of the motion sensor that generates the motion data used by the machine learning model 300 to perform activity recognition. In some examples, the training device 320 may also perform a reflection transformation of the exemplary input data 342, whereby the machine learning model 300 can accurately perform activity recognition regardless of whether the motion sensor (e.g., on a smartwatch) that generates the motion data used by the machine learning model 300 to perform activity recognition is worn on the user's left wrist or right wrist.

[0136] In some examples, the exemplary input data corresponding to a stationary state may include motion data generated by a user who was in a vehicle, such as a user in a car, bus, or train, while the vehicle was in motion or the user was riding in the vehicle. Specifically, the exemplary input data may include motion data generated by a motion sensor of a computing device carried or worn by a user who was in a vehicle or riding in a vehicle and was labeled as being in a static state. By training the machine learning model 300 using such exemplary input data, the machine learning model 300 can be trained not to recognize such motion data generated by a motion sensor of a computing device carried or worn by a user who was in a vehicle or riding in a vehicle as walking, running, or any form of body movement.

[0137] In some examples, the exemplary input data corresponding to walking may include motion data generated by a motion sensor of a smartphone carried by a user who was walking and motion data generated by a motion sensor of a wearable computing device worn by a user who was walking. By including motion data generated by a motion sensor of a smartphone carried by a user who was walking and motion data generated by a motion sensor of a wearable computing device, the availability of training data regarding walking can be increased, as opposed to using only motion data generated by a motion sensor of a wearable computing device such as a smartwatch.

[0138] In some examples, the training device 320 may incorporate motion data from the free-living motion data 346 into the training data 341 used to train the machine learning model 300. The free-living motion data 346 may include motion data generated by a motion sensor of a computing device that a user carries or wears when engaging in daily activities over a plurality of consecutive days, such as for three or more days. Generally, the free-living motion data 346 may not be labeled. Thus, the training device 320 may determine a label for the free-living motion data 346 and may include at least a portion of the free-living motion data 346 and its associated label within the training data 341 used to train the machine learning model 300.

[0139] In some examples, the training device 320 may determine a label associated with the motion data within the free-living motion data 346 based at least in part on accelerometer measurements (i.e., acceleration forces along the x, y, and z axes) associated with a window of motion data within the free-living motion data 346 and / or average accelerometer measurements (i.e., the average of the total acceleration force combining the x, y, and z axes). In particular, the training device 320 may determine whether the motion data within the free-living motion data 346 should be labeled and whether to select the motion data within the free-living motion data 346 and include it in the training data 341 used to train the machine learning model 300 based at least in part on the accelerometer measurements and / or average accelerometer measurements of the motion data within the free-living motion data 346.

[0140] The training device 320 can divide the free-living motion data 346 into a plurality of windows of motion data. For example, each window can be about 5 seconds of motion data, including 512 samples of motion data such as triaxial accelerometer data when sampled at 100 Hertz. The training device 320 can determine the average (i.e., median) accelerometer measurement for the window and, for each sample of the motion data, subtract the average accelerometer measurement for the window from the sample of the motion data to generate an average accelerometer measurement with zero mean within the window for purposes such as removing the influence of gravity. The average accelerometer measurement of the clipped window can represent the overall score for the window.

[0141] When the training device 320 generates a window having an average accelerometer measurement with zero mean, the training device 320 can clip the accelerometer measurements within the window between two values, such as between 0 and 4, to remove motion below the average and reduce the influence of outlier accelerometer measurement values. Thus, the training device 320 can determine the clipped average accelerometer measurement in the window as the overall score for the window.

[0142] The overall score for a window can correspond to the level of activity detected within the window. In this case, a low overall score for a window can correspond to a low level of activity detected within the window, and a high overall score for a window can correspond to a high level of activity. For example, a window can be classified as one of four activity levels, such as stationary, low intensity, medium intensity, and high intensity, based on the overall score for the window. Thus, the training device 320 can select the motion data of the free-living motion data 346 within the window classified as stationary for inclusion in the exemplary input data of the exemplary input data 342 corresponding to the stationary state used to train the machine learning model 300.

[0143] In some examples, the training device 320 may also select motion data 346 of free-living motion within a window classified as low-intensity activity for inclusion in exemplary input data 342 corresponding to a stationary state used to train the machine learning model 300. By including such motion data corresponding to low-intensity activity in the exemplary input data 342 corresponding to the stationary state used to train the machine learning model 300, it may be possible for the machine learning model 300 to function more appropriately as a general activity classifier and to overall improve the stationary state detection of the machine learning model 300.

[0144] In some examples, the training device 320 may perform activity recognition of the free-living motion data 346 using a pre-trained activity recognition model such as the activity recognition model 348, label the free-living motion data 346, and include the labeled free-living motion data 346 in the training data 341 used to train the machine learning model 300. The activity recognition model 348 may perform activity recognition to determine possible activities associated with the motion data within the window, and may estimate an activity recognition probability for each of the possible activities determined by the activity recognition model 348. The activity recognition model 348 may filter and remove the possible activities having a related activity recognition probability below a specified reliability threshold, and label each of the remaining windows of the motion data in the free-living motion data 346 with the related possible activities having an activity recognition probability equal to or greater than the reliability threshold, thereby generating labeled motion data that can be incorporated into the training data 341 used to train the machine learning model 300.

[0145] Regarding a window of motion data associated with a possible activity of a stationary state having an activity recognition probability above a reliability threshold, such a window of motion data may also have to pass the heuristic of the stationary state (e.g., classified as a stationary state) described above. In this case, it should be noted that the stillness of the window is determined at least partially based on the measured values of the average accelerometer for the window in order to be labeled as a stationary state. For example, if a window of motion data associated with a possible activity of a stationary state has an activity recognition probability above the reliability threshold but does not pass the stillness heuristic described above (e.g., is not classified as a stationary state), the training device 320 may not include the window of motion data in the training data 341 used to train the machine learning model 300.

[0146] In some examples, the activity recognition model 348 may determine two or more possible activities from a window of motion data within the free-living motion data 346. For example, the activity recognition model 348 may determine the possible activities for each of a plurality of overlapping sets of motion data. In this case, each set of motion data includes at least some samples of motion data within the window of motion data. Thus, the activity recognition model 348 may determine a single possible activity for the window of motion data by performing a weighted average of the plurality of overlapping sets of motion data. In this case, the weights used to perform the weighted average may correspond to the percentage of samples of motion data in the set of motion data within the window of motion data.

[0147] In some examples, the activity recognition model 348 may perform debouncing of a window of motion data within the free-living motion data 346 before including such window motion data in the training data 341. For example, some motion data may include random spikes of high-confidence predictions such that the motion data corresponds to a cycling physical activity of several seconds in length (e.g., 10 seconds) surrounded by motion data that appears to correspond to physical activities other than cycling. Since it is highly unlikely that a user cycles for only a few seconds at a time, the activity recognition model 348 may perform debouncing of such window motion data to filter out and remove short bursts of motion data corresponding to cycling physical activity by considering the context of the window surrounding the motion data.

[0148] For the current window having a related possible activity for cycling, the activity recognition model 348 may determine whether it is likely that the related possible activity for cycling regarding the current window is accurate based on a configurable number (e.g., 5) of adjacent windows, such as the adjacent window to the left of the current window in the free-living motion data 346, the adjacent window to the right of the current window in the free-living motion data 346, or the adjacent windows to the right and left of the current window centered between the adjacent windows. If the activity recognition model 348 determines that the number of adjacent windows also associated with the possible activity of cycling is not above a specified threshold, the activity recognition model 348 may determine that the current window may be misclassified as being associated with the possible activity of cycling.

[0149] If activity recognition model 348 determines that a window may be misclassified as being associated with a possible activity of cycling, activity recognition model 348 may determine whether the window may be associated with a possible activity different from cycling. For example, activity recognition model 348 may set the cycling component of the activity recognition probability vector for the window to zero (i.e., set all to zero), and may re-normalize the activity recognition probability vector for the window to 1, so that if there is another possible activity that is likely to be associated with the window, activity recognition model 348 may have an opportunity to associate that possible activity with the window.

[0150] In some examples, machine learning model 300 can be trained by optimizing an objective function such as objective function 345. For example, in some examples, objective function 345 can be, or can include, a loss function that compares (e.g., determines the difference between) the output data generated by the model from the training data and the label associated with the training data (e.g., a ground-truth label). For example, the loss function can evaluate the sum or average of the squared differences between the output data and the label. In some examples, objective function 345 can be, or can include, a cost function that describes the cost of a certain outcome or output data. Other examples of objective function 345 can include, for example, margin-based techniques such as triplet loss or max-margin training.

[0151] To optimize objective function 345, one or more of various optimization techniques can be executed. For example, the optimization technique can minimize or maximize objective function 345. Exemplary optimization techniques include Hessian-based techniques and gradient-based techniques, such as coordinate descent, gradient descent (e.g., stochastic gradient descent), and conjugate gradient method. Other optimization techniques include black box optimization techniques and heuristic techniques.

[0152] In some implementations, (e.g., when the machine learning model is a multi-layer model such as an artificial neural network) error backpropagation can be used in conjunction with an optimization technique (e.g., a gradient-based technique) to train the machine learning model 300. For example, an iterative cycle of propagation and model parameter (e.g., weight) updates can be executed to train the machine learning model 300. Examples of backpropagation techniques include truncated backpropagation through time, Levenberg-Marquardt backpropagation, and the like.

[0153] In some implementations, the machine learning model 300 described herein can be trained using unsupervised learning techniques. Unsupervised learning can involve inferring a function for describing hidden structure from unlabeled data. For example, classification or categorization may not be included in the data. Using unsupervised learning techniques, a machine learning model can be generated that can perform clustering, anomaly detection, learning of latent variable models, or other tasks.

[0154] The machine learning model 300 can be trained using semi-supervised techniques that combine aspects of supervised and unsupervised learning. The machine learning model 300 can be trained or generated by evolutionary techniques or genetic algorithms. In some implementations, the machine learning model 300 described herein can be trained using reinforcement learning. In reinforcement learning, an agent (e.g., a model) can take actions in an environment and learn to maximize the reward resulting from such actions and / or minimize the penalty resulting from such actions. Reinforcement learning can be different from supervised learning problems in that correct input / output pairs are not presented and sub-optimal actions are not explicitly corrected.

[0155]

[0155] In some implementation examples, in order to improve the generalization of the machine learning model 300, one or more generalization techniques can be executed during training. The generalization techniques can help reduce the overfitting of the machine learning model 300 to the training data. Examples of generalization techniques include dropout techniques, weight decay techniques, batch normalization, early stopping, subset selection, stepwise selection, and the like.

[0156]

[0155] In some implementation examples, the machine learning model 300 described herein may include or be affected by some hyperparameters, such as learning rate, number of layers, number of nodes in each layer, number of leaves of a tree, number of clusters, and the like. Hyperparameters can affect model performance. Hyperparameters can be selected manually or automatically by applying techniques such as grid search, black box optimization techniques (e.g., Bayesian optimization, random search, etc.), gradient-based optimization, and the like. Exemplary techniques and / or tools for performing automatic hyperparameter optimization include Hyperopt, Auto-WEKA, Spearmint, Metric Optimization Engine (MOE), and the like.

[0157]

[0155] In some implementation examples, various techniques can be used to optimize and / or adapt the learning rate during model training. Exemplary techniques and / or tools for performing learning rate optimization or adaptation include Adagrad, Adaptive Moment Estimation (ADAM), Adadelta, RMSprop, and the like.

[0158]

[0155] In some implementation examples, transfer learning techniques can be used to provide an initial model for starting the training of the machine learning model 300 described herein.

[0159] In some implementation examples, the machine learning model 300 described herein may be included in various parts of computer-readable code on a computing device. In one example, the machine learning model 300 may be included in a specific application or program and may be used (e.g., exclusively) by such a specific application or program. Thus, in one example, a computing device may include several applications, and one or more of such applications may each include its own machine learning library and machine learning model.

[0160] In another example, the machine learning model 300 described herein may be included in (e.g., in the central intelligence layer of) the operating system of a computing device and may be invoked or used by one or more applications that interact with the operating system. In some implementation examples, each application may communicate with the central intelligence layer (and the models stored therein) using an application programming interface (API) (e.g., a common public API across all applications).

[0161] In some implementation examples, the central intelligence layer can communicate with a central device data layer. The central device data layer may be a centralized repository of data for the computing device. The central device data layer may communicate with some other components of the computing device, such as, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some implementation examples, the central device data layer may communicate with each device component using an API (e.g., a private API).

[0162] The technology described in this specification refers to servers, databases, software applications, and other computer-based systems, as well as actions taken with respect to such systems and information transmitted to and received from such systems. The inherent flexibility of computer-based systems allows for a wide variety of realizable configurations, combinations, and divisions of tasks and functions among components. For example, the processes described in this specification can be realized using a single device or component, or multiple devices or components that function in combination. Databases and applications can be realized on a single system or distributed across multiple systems. Distributed components can operate continuously or in parallel.

[0163] In addition, the machine learning techniques described in this specification are easily replaceable and combinable. Although some exemplary techniques have been described, many other techniques exist and can be used in conjunction with aspects of the present disclosure.

[0164] In addition to the above description, the user may be provided with control that enables the user to select whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., the user's social network, social behavior or activities, occupation, user preferences, or the user's current location), and whether the user is to receive content or communications sent from the server. Additionally, certain data may be processed in one or more ways before being stored or used, such that information that can identify an individual is removed. For example, the user's identity may be processed so as not to be able to identify information that can identify the individual for that user, or the user's geographical location may be generalized so as not to be able to identify the user's specific location if location information (such as city, postal code, or state level) is obtained. Thus, the user may control what information is collected about the user, how that information is used, and what information is provided to the user.

[0165] FIG. 3E shows a conceptual diagram illustrating an exemplary machine learning model. As shown in FIG. 3E, the machine learning model 300 can be a convolutional neural network including eight convolutional blocks 352A-352H (the "convolutional blocks 352"), followed by a bottleneck layer 354 and a SoftMax layer 356.

[0166] Each of the convolutional blocks 352 can include a convolutional filter, followed by a rectified linear unit (ReLU) activation function and a max pooling operation. The convolutional filters of the convolutional blocks can exponentially increase the capacity per layer to suppress the reduction in spatial resolution from the max pooling operation. As shown in FIG. 3E, the machine learning model 300 can implement dropout regularization to reduce overfitting.

[0167] Input 350 can be triaxial accelerometer data of 256 samples that is sampled at 25 Hertz over approximately 10 seconds. For this reason, the convolution filter of convolution block 352A can have a height of 256 to accept the 256 samples of input 350. The convolution filter of convolution block 352A can also have a depth of 12. Convolution block 352A can generate an output of 128 samples from the 256-sample input 350.

[0168] The convolution filter of convolution block 352B can have a height of 128 to accept the 128 samples output by convolution block 352A. The convolution filter of convolution block 352B can also have a depth of 14. Convolution block 352B can generate an output of 64 samples from the 128 samples output from convolution block 352A.

[0169] The convolution filter of convolution block 352C can have a height of 64 to accept the 64 samples output by convolution block 352B. The convolution filter of convolution block 352C can also have a depth of 17. Convolution block 352C can generate an output of 32 samples from the 64 samples output from convolution block 352B.

[0170] The convolution filter of convolution block 352D can have a height of 32 to accept the 32 samples output by convolution block 352C. The convolution filter of convolution block 352D can also have a depth of 20. Convolution block 352D can generate an output of 16 samples from the 32 samples output from convolution block 352C.

[0171] The convolutional filter of convolutional block 352E may have a height of 16 to receive the 16 samples output by convolutional block 352D. The convolutional filter of convolutional block 352E may also have a depth of 25. Convolutional block 352E may generate an output of 8 samples from the 16 samples output from convolutional block 352D.

[0172] The convolutional filter of convolutional block 352F may have a height of 8 to receive the 8 samples output by convolutional block 352E. The convolutional filter of convolutional block 352F may also have a depth of 30. Convolutional block 352F may generate an output of 4 samples from the 8 samples output from convolutional block 352E.

[0173] The convolutional filter of convolutional block 352G may have a height of 4 to receive the 4 samples output by convolutional block 352F. The convolutional filter of convolutional block 352G may also have a depth of 35. Convolutional block 352G may generate an output of 2 samples from the 4 samples output from convolutional block 352F.

[0174] The convolutional filter of convolutional block 352H may have a height of 2 to receive the 2 samples output by convolutional block 352G. The convolutional filter of convolutional block 352H may also have a depth of 43. Convolutional block 352H may generate an output that can be supplied to bottleneck layer 354 from the 2 samples output from convolutional block 352F.

[0175] The bottleneck layer 354 can be a fully-connected layer of 16 units that flattens the output of the convolutional block 352. The fully-connected SoftMax layer 356 can receive the output of the bottleneck layer 354 and determine the probabilities of the motion data of the input 350 corresponding to each of the plurality of classes of physical activities. For example, the fully-connected SoftMax layer 356 can determine a probability distribution that can be a floating-point number that sums to 1 across a plurality of classes of physical activities, such as a probability distribution regarding physical activities of walking, running, cycling, and being in a stationary state. Thus, the fully-connected SoftMax layer 356 can generate an output 358 that is a probability distribution of the motion data of the input 350 regarding a plurality of physical activities.

[0176] Figure 4 is a flow diagram illustrating an exemplary operation of a computing device that can perform on-device recognition of physical activity, according to one or more aspects of the present disclosure. For purposes of illustration only, the operation example will be described below in the context of the computing device 110 of FIG. 1.

[0177] As shown in FIG. 4, the computing device 110 can receive motion data corresponding to the movement sensed by one or more motion sensors and generated by one or more motion sensors, such as one or more sensor components 114 and / or one or more sensor components 108 (400). The computing device 110 can perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize the physical activity corresponding to the motion data (402). The computing device 110 can execute an operation associated with the physical activity in response to recognizing the physical activity corresponding to the motion data (406).

[0178] The present disclosure includes the following examples. Example 1: The method includes receiving, by a computing device, motion data generated by one or more motion sensors corresponding to motion sensed by the one or more motion sensors; performing, by the computing device, on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data; and performing, by the computing device, an action associated with the physical activity in response to recognizing the physical activity corresponding to the motion data.

[0179] Example 2: The one or more neural networks are trained using a differential privacy framework and using a delta parameter that bounds the probability that the privacy guarantee of the one or more neural networks is not maintained, the delta parameter having a value set to the reciprocal of the size of a training set for the one or more neural networks, the method of Example 1.

[0180] Example 3: The one or more neural networks are trained using a training set of motion data corresponding to a plurality of physical activities, the method of Example 1 or 2.

[0181] Example 4: The training set of the motion data is transformed in a plurality of directions with random rotation without constraints, the method of Example 3.

[0182] Example 5: The training set of the motion data is generated from a user who is driving a vehicle or is a passenger in a vehicle and includes a plurality of motion data classified as being in a stationary state, the method of Example 3 or 4.

[0183] Example 6: The method according to any one of Examples 3 to 5, wherein the training set of the motion data includes a plurality of motion data generated by a pre-trained activity recognition model based at least in part on the labeling of unlabeled free-living motion data.

[0184] Example 7: The method according to Example 6, wherein the plurality of motion data generated by the pre-trained activity recognition model based at least in part on the labeling of the unlabeled free-living motion data is further generated by the pre-trained activity recognition model that performs debouncing of the window of the motion data in the unlabeled free-living motion data to remove one or more short bursts of motion data corresponding to cycling based on the context of adjacent windows of the motion data for the window of the motion data.

[0185] Example 8: The method according to any one of Examples 3 to 7, wherein the training set of the motion data includes a plurality of motion data classified as a stationary state selected from the unlabeled free-living motion data based at least in part on the measured values of the average accelerometer associated with the window of the motion data in the unlabeled free-living motion data.

[0186] Example 9: The method according to any one of Examples 1 to 8, wherein the step of performing the on-device activity recognition to recognize the physical activity includes the step of determining a probability distribution of the motion data over a plurality of physical activities using the one or more neural networks.

[0187] Example 10: The step of receiving the motion data corresponding to the movement sensed by the one or more motion sensors and generated by the one or more motion sensors further includes the step of receiving, by the computing device, the motion data generated by the one or more motion sensors of the wearable computing device communicatively coupled to the computing device, where the motion data corresponds to the movement of the wearable computing device sensed by the one or more motion sensors. The method according to any one of Examples 1 to 9.

[0188] Example 11: The step of receiving the motion data corresponding to the movement sensed by the one or more motion sensors and generated by the one or more motion sensors includes the step of receiving, by the computing device, the motion data generated by the one or more motion sensors of the computing device, where the motion data corresponds to the movement of the computing device sensed by the one or more motion sensors. The method according to any one of Examples 1 to 10.

[0189] Example 12: A computing device includes a memory and one or more processors. The one or more processors are configured to receive motion data corresponding to a movement sensed by one or more motion sensors and generated by the one or more motion sensors, perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and execute an action associated with the physical activity in response to recognizing the physical activity corresponding to the motion data.

[0190] Example 13: The computing device according to Example 12, wherein the one or more neural networks are trained using a differential privacy library and using a delta parameter that bounds the probability that the privacy guarantee of the one or more neural networks is not maintained, and the delta parameter has a value set to the reciprocal of the size of the training set for the one or more neural networks.

[0191] Example 14: The computing device according to Example 12 or 13, wherein the one or more neural networks are trained using a training set of motion data corresponding to a plurality of physical activities.

[0192] Example 15: The computing device according to Example 14, wherein the training set of the motion data is transformed in a plurality of directions with random rotations without constraints.

[0193] Example 16: The computing device according to Example 14 or 15, wherein the training set of the motion data includes a plurality of motion data generated by a pre-trained activity recognition model based at least in part on the labeling of unlabeled free-living motion data.

[0194] Example 17: A computer-readable storage medium stores instructions that, when executed, cause one or more processors of a computing device to receive motion data corresponding to and generated by movements sensed by one or more motion sensors, perform on-device activity recognition using one or more neural networks trained with differential privacy to recognize a physical activity corresponding to the motion data, and in response to recognizing the physical activity corresponding to the motion data, execute an action associated with the physical activity.

[0195] Example 18: The computer-readable storage medium according to Example 17, wherein the one or more neural networks are trained using a differential privacy library and a delta parameter that bounds the probability that the privacy guarantee of the one or more neural networks is not maintained, and the delta parameter has a value set to the reciprocal of the size of the training set for the one or more neural networks.

[0196] Example 19: The computer-readable storage medium according to Example 17 or 18, wherein the one or more neural networks are trained using a training set of motion data corresponding to a plurality of physical activities.

[0197] Example 20: The computer-readable storage medium according to Example 19, wherein the training set of the motion data is transformed in a plurality of directions with random rotations without constraints.

[0198] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, these functions may be stored on a computer-readable medium or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium, or a communication medium that includes any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally corresponds to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0199] By way of example and not limitation, such a computer-readable storage medium may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or other media, which can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. Also, any connection is suitable to be called a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient tangible storage media. As used herein, disk (both "disk" and "disc") includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy (registered trademark) disk, and Blu-ray disc, where disk typically magnetically reproduces data, while disc optically reproduces data using a laser. The above combinations should also be included within the scope of computer-readable media.

[0200] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits, etc. Thus, the term "processor" as used herein may mean any of the above structures or other structures suitable for the implementation of the technologies described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules. Also, these technologies may be fully realized in one or more circuits or logic elements.

[0201] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses including wireless handsets, integrated circuits (ICs) or a set of ICs (e.g., a chipset). Although various components, modules or units are described in the present disclosure to emphasize the functional aspects of the devices configured to execute the disclosed techniques, it is not necessarily required to be implemented by various hardware units. Rather, as described above, the various units may be combined within a hardware unit or may be provided by an assembly of interoperable hardware units including the above one or more processors, along with suitable software and / or firmware.

[0202] Depending on the embodiment, any one or more of the acts or events of any of the methods described herein may be performed in a different order, may be added, may be combined, or may all be excluded (e.g., not all of the described acts or events are necessary for the implementation of the method). It should further be recognized that in some embodiments, the acts or events may be performed simultaneously rather than sequentially, for example, by multi-threaded processing, interrupt processing, or by a plurality of processors.

[0203] In some examples, a computer-readable storage medium includes a non-transitory medium. In some examples, the term "non-transitory" indicates that the storage medium is not embodied in a carrier wave or a propagated signal. In some examples, a non-transitory storage medium may store data that can change over time (e.g., in RAM or a cache). Although some examples are described as outputting various information for display, the techniques of this disclosure may output such information in other forms, such as audio, holographic, or tactile forms, to name just a few.

[0204] Various aspects of the present disclosure have been described. These and other aspects are within the scope of the appended claims.

Claims

Claim 1 A method comprising: one or more processors of a wearable computing device receive motion data generated by one or more motion sensors and corresponding to movement of the user sensed by the one or more motion sensors; the one or more processors execute on-device activity recognition processing to recognize a physical activity corresponding to the motion data using one or more neural networks trained with differential privacy, the one or more neural networks being trained using a training set of motion data corresponding to a plurality of physical activities, the training set of motion data including a plurality of motion data generated by a user who is driving a vehicle or riding in a vehicle and classified as being in a stationary state; the one or more processors execute an operation associated with the physical activity in response to recognizing the physical activity corresponding to the received motion data. Claim 2 The method of claim 1, wherein the training set of motion data is transformed in a plurality of directions with random rotation without constraints. Claim 3 The method of claim 1, wherein the training set of motion data includes a plurality of motion data generated by a pre-trained activity recognition model based at least in part on labeling of motion data that has not been labeled for classification among motion data generated by detecting the user's daily activities by the one or more motion sensors. Claim 4 The plurality of motion data generated by the pre-trained activity recognition model based at least in part on the labeling of the unlabeled motion data is further obtained by the pre-trained activity recognition model for performing debouncing of the window of motion data in the unlabeled motion data and removing one or more short bursts of motion data corresponding to cycling based on the context of adjacent windows of motion data for the window of motion data, in the unlabeled motion data. The method according to claim 3.

5. The training set of the motion data includes a plurality of motion data classified as a stationary state selected from the unlabeled motion data based at least in part on measured values of an average accelerometer associated with a window of motion data in the unlabeled motion data. The method according to claim 1.

6. The step of performing the on-device activity recognition process to recognize the physical activity includes generating, using the one or more neural networks, a probability distribution of the motion data composed of probabilities of each of a plurality of physical activities; and recognizing a physical activity corresponding to the motion data based on a comparison result between the probability distribution and a preset threshold. The method according to any one of claims 1 to 5.

7. The step of receiving the motion data corresponding to the movement sensed by the one or more motion sensors and generated by the one or more motion sensors further includes the one or more processors receiving the motion data generated by the one or more motion sensors of the wearable computing device communicatively coupled to the computing device, the motion data corresponding to the movement of the wearable computing device sensed by the one or more motion sensors. The method according to any one of claims 1 to 6.

8. The step of receiving the motion data corresponding to the movement sensed by the one or more motion sensors and generated by the one or more motion sensors further includes the one or more processors receiving the motion data generated by the one or more motion sensors of the computing device, the motion data corresponding to the movement of the computing device sensed by the one or more motion sensors, the method according to any one of claims 1 to 6. **Claim 9** A computing device wearable by a user, a memory, and one or more processors, the one or more processors receiving motion data corresponding to a movement sensed by one or more motion sensors and generated by the one or more motion sensors, performing on-device activity recognition processing to recognize a physical activity corresponding to the motion data using one or more neural networks trained with differential privacy, the one or more neural networks being trained using a training set of motion data corresponding to a plurality of physical activities, the training set of motion data including a plurality of motion data generated from a user who is driving a vehicle or riding in a vehicle and classified as being in a stationary state, a computing device configured to perform an operation associated with the physical activity in response to recognizing the physical activity corresponding to the received motion data. **Claim 10** The computing device according to claim 9, wherein the training set of the motion data is transformed in a plurality of directions with random rotation without constraints. **Claim 11** The computing device according to claim 10, wherein the training set of the motion data includes a plurality of motion data generated by a pre-trained activity recognition model based at least in part on labeling of motion data that has not been labeled for classification among the motion data generated by detecting the daily activities of the user by the one or more motion sensors. **Claim 12** A program for causing one or more processors to execute the method according to any one of claims 1 to 8. **Claim 13** A memory storing the program according to claim 12, and one or more processors configured to execute the program, the apparatus comprising.

Citation Information

Patent Citations

  • Discrimination device, information provision system, discrimination method, and information provision method

    JP2020124389A