Wearable multimedia device and cloud computing platform with laser projection system

The wearable multimedia device with a camera, depth sensor, and laser projection system addresses the challenge of capturing spontaneous moments by automatically editing and projecting data, enhancing user interaction and data processing.

JP2025169324APending Publication Date: 2025-11-12HUMANE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025134330
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-17
Filing Date
2025-08-12
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Modern mobile devices often fail to capture important spontaneous moments due to their limitations in capturing events too quickly or users forgetting to take pictures or videos.

Method used

A wearable multimedia device equipped with a camera, depth sensor, and laser projection system that captures and processes multimedia data, allowing for automatic editing and formatting on a cloud computing platform, and projects data onto surfaces using laser projections.

Benefits of technology

The device captures spontaneous moments with minimal user interaction, automatically edits and formats multimedia data based on user preferences, and provides interactive projections for enhanced user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169324000001_ABST
    Figure 2025169324000001_ABST
Patent Text Reader

Abstract

To provide systems for a wearable multimedia device and a cloud computing platform with an application ecosystem for processing multimedia data captured by the wearable multimedia device.SOLUTION: A body-worn device disclosed herein comprises a camera, a depth sensor, a laser projection system, and one or more processors. The one or more processors are configured to capture a set of digital images using the camera, identify an object in the set of digital images, capture depth data using the depth sensor, identify a gesture of a user wearing the device in the depth data, associate the object with the gesture, obtain data associated with the object, and project a laser projection of the data on a surface using the laser projection system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority from U.S. Provisional Patent Application No. 62 / 863,222, filed June 18, 2019, for "Wearable Multimedia Device and Cloud Computing Platform With Application Ecosystem," and to U.S. Patent Application No. 16 / 904,544, filed June 17, 2020, for "Wearable Multimedia Device and Cloud Computing Platform With Laser Projection System," which is a continuation-in-part of U.S. Patent Application No. 15 / 976,632, filed May 20, 2018, for "Wearable Multimedia Device and Cloud Computing Platform With Application Ecosystem," each of which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates generally to cloud computing and multimedia editing. [Background technology]

[0003] Modern mobile devices (e.g., smartphones, tablet computers) often include built-in cameras that allow users to take digital images or videos of spontaneous events. These digital images and videos can be stored in an online database associated with a user account to free up memory on the mobile device. Users can share their images and videos with friends and family and download or stream images and videos on demand using their various playback devices. These built-in cameras offer significant advantages over traditional digital cameras, which are bulky and often require more time to prepare for capture.

[0004] Despite the convenience of mobile device built-in cameras, there are many important moments that are not captured by these devices because they occur too quickly or the user is swept up in the moment and simply forgets to take a picture or video. Summary of the Invention [Means for solving the problem]

[0005] A system, method, device and non-transitory computer-readable storage medium for a wearable multimedia device and a cloud computing platform with an application ecosystem for processing multimedia data captured by the wearable multimedia device are disclosed.

[0006] In one embodiment, a body-worn device comprises a camera, a depth sensor, a laser projection system, one or more processors, and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including capturing a set of digital images using the camera; identifying an object within the set of digital images; capturing depth data using the depth sensor; identifying a gesture of a user wearing the device within the depth data; associating the object with the gesture; obtaining data associated with the object; and projecting a laser projection of the data onto a surface using the laser projection system.

[0007] In one embodiment, the laser projection includes a text label for the object.

[0008] In one embodiment, the laser projection includes a size template for the object.

[0009] In one embodiment, the laser projection includes instructions to perform an action on the object.

[0010] In one embodiment, a body-worn device comprises a camera, a depth sensor, a laser projection system, one or more processors, and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: capturing depth data using the sensor; identifying a first gesture in the depth data, the first gesture being made by a user wearing the device; associating the first gesture with a request or command; and projecting, using the laser projection system, a laser projection associated with the request or command onto a surface.

[0011] In one embodiment, the operations further include obtaining a second gesture associated with the laser projection using the depth sensor, determining a user input based on the second gesture, and initiating one or more actions according to the user input.

[0012] In one embodiment, the operations further include masking the laser projection to prevent projecting data onto the user's hand making the second gesture.

[0013] In one embodiment, the operations further include obtaining depth or image data indicative of the geometry, material, or texture of the surface using a depth sensor or camera, and adjusting one or more parameters of the laser projection system based on the geometry, material, or texture of the surface.

[0014] In one embodiment, the operations further include using a camera to capture a reflection of the laser projection from the surface, and automatically adjusting the intensity of the laser projection to compensate for the different refractive indices so that the laser projection has uniform brightness.

[0015] In one embodiment, the device includes a magnetic attachment mechanism configured to magnetically couple to the battery pack through a user's clothing, the magnetic attachment mechanism further configured to receive inductive charging from the battery pack.

[0016] In one embodiment, a method includes capturing depth data using a depth sensor of a body-worn device; identifying, using one or more processors of the device, a first gesture in the depth data, the first gesture being made by a user wearing the device; associating, using the one or more processors, the first gesture with a request or command; and projecting, using a laser projection system of the device, a laser projection associated with the request or command onto a surface.

[0017] In one embodiment, the method further includes acquiring a second gesture by the user using the depth sensor, the second gesture being associated with the laser projection; determining a user input based on the second gesture; and initiating one or more actions according to the user input.

[0018] In one embodiment, the one or more actions include controlling another device.

[0019] In one embodiment, the method further includes masking the laser projection to prevent projecting data onto the user's hand making the second gesture.

[0020] In one embodiment, the method further includes using a depth sensor or camera to acquire depth or image data indicative of the surface geometry, material or texture, and adjusting one or more parameters of the laser projection system based on the surface geometry, material or texture.

[0021] In one embodiment, a method includes receiving, by one or more processors of a cloud computing platform, context data from a wearable multimedia device including at least one data capture device for capturing the context data; creating, by the one or more processors, a data processing pipeline in one or more applications based on one or more characteristics of the context data and a user request; processing, by the one or more processors, the context data through the data processing pipeline; and sending, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or other device for presentation of the output.

[0022] In one embodiment, a system includes one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including receiving, by one or more processors of a cloud computing platform, context data from a wearable multimedia device including at least one data capture device for capturing the context data; creating, by the one or more processors, a data processing pipeline with one or more applications based on one or more characteristics of the context data and a user request; processing, by the one or more processors, the context data through the data processing pipeline; and sending, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or other device for presentation of the output.

[0023] In one embodiment, a non-transitory computer-readable storage medium includes instructions for receiving, by one or more processors of a cloud computing platform, contextual data from a wearable multimedia device including at least one data capture device for capturing the contextual data; creating, by the one or more processors, a data processing pipeline with one or more applications based on one or more characteristics of the contextual data and a user request; processing, by the one or more processors, the contextual data through the data processing pipeline; and sending, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or other device for presentation of the output.

[0024] In one embodiment, a method includes receiving, by a controller of a wearable multimedia device, depth or image data indicative of surface geometry, material, or texture provided by one or more sensors of the wearable multimedia device; adjusting, by the controller, one or more parameters of a projector of the wearable multimedia device based on the surface geometry, material, or texture; projecting, by the projector of the wearable multimedia device, text or image data onto the surface; receiving, by the controller, depth or image data from one or more sensors indicative of user interaction with the text or image data projected onto the surface; determining, by the controller, user input based on the user interaction; and initiating, by a processor of the wearable multimedia device, one or more actions according to the user input.

[0025] In one embodiment, a wearable multimedia device comprises one or more sensors; a projector; and a controller configured to receive depth or image data from the one or more sensors, the depth or image data being indicative of a surface geometry, material, or texture, the depth or image data being provided by the one or more sensors of the wearable multimedia device; adjust one or more parameters of the projector based on the surface geometry, material, or texture; project text or image data onto a surface using the projector; receive depth or image data from the one or more sensors indicative of user interaction with the text or image data projected onto the surface; determine user input based on the user interaction; and initiate one or more actions according to the user input. [Effects of the Invention]

[0026] Certain embodiments disclosed herein provide one or more of the following advantages: A wearable multimedia device captures multimedia data of spontaneous moments and transactions with minimal user interaction. The multimedia data is automatically edited and formatted on a cloud computing platform based on user preferences and then made available to the user for playback on various user playback devices. In one embodiment, data editing and / or processing is performed by an ecosystem of applications that are proprietary and / or provided / licensed from third-party developers. The application ecosystem provides various access points (e.g., websites, portals, APIs) that allow third-party developers to upload, verify, and update their applications. The cloud computing platform automatically builds a custom processing pipeline for each multimedia data stream using one or more of the ecosystem applications, user preferences, and other information (e.g., data type or format, data quantity and quality).

[0027] Additionally, wearable multimedia devices include cameras and depth sensors that can detect objects and air gestures by the user and then perform or infer various actions based on the detection, such as labeling objects in camera images or controlling other devices. In one embodiment, the wearable multimedia device does not include a display, allowing the user to continue interacting with friends, family, and colleagues without being immersed in a display, a current challenge for smartphone and tablet computer users. As such, wearable multimedia devices take a different technological approach than, for example, smart goggles or smart glasses for augmented reality (AR) and virtual reality (VR), which further remove the user from the real-world environment. To facilitate collaboration with others and compensate for the lack of a display, the wearable multimedia computer includes a laser projection system that projects a laser projection onto any surface, including tables, walls, and even the palm of the user's hand. Laser projection can label objects, provide text or instructions related to objects, and provide an ephemeral user interface (e.g., keyboard, numeric keypad, device controller) that allows users to compose messages, control other devices, or simply share and discuss content with others.

[0028] The details of the disclosed embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0029] [Figure 1] FIG. 1 is a block diagram of an operating environment for a cloud computing platform with a wearable multimedia device and an application ecosystem for processing multimedia data captured by the wearable multimedia device, according to one embodiment. [Figure 2]2 is a block diagram of a data processing system implemented by the cloud computing platform of FIG. 1 according to one embodiment. [Figure 3] FIG. 2 is a block diagram of a data processing pipeline for processing a contextual data stream according to one embodiment. [Figure 4] FIG. 10 is a block diagram of another data process for processing contextual data streams for transportation applications according to one embodiment. [Figure 5] 3 illustrates an example data object used by the data processing system of FIG. 2 according to one embodiment. [Figure 6] FIG. 1 is a flow diagram of a data pipeline process according to one embodiment. [Figure 7] 1 is an architecture for a cloud computing platform, according to one embodiment. [Figure 8] 1 is an architecture for a wearable multimedia device, according to one embodiment. [Figure 9] 4 is a screenshot of an example graphical user interface (GUI) for the scene identification application described with respect to FIG. 3, according to one embodiment. [Figure 10] 10 illustrates a classifier framework for classifying raw or pre-processed context data into objects and metadata that can be searched using the GUI of FIG. 9 according to one embodiment. [Figure 11] FIG. 1 is a system block diagram illustrating a hardware architecture for a wearable multimedia device according to one embodiment. [Figure 12] FIG. 1 is a system block diagram illustrating a processing framework implemented in a cloud computing platform for processing raw or pre-processed context data received from a wearable multimedia device, according to one embodiment. [Figure 13] 1 illustrates software components for a wearable multimedia device according to one embodiment. [Figure 14A] 1 illustrates the use of a projector in a wearable multimedia device to project various types of information onto the palm of a user's hand, according to one embodiment. [Figure 14B] 1 illustrates the use of a projector in a wearable multimedia device to project various types of information onto the palm of a user's hand, according to one embodiment. [Figure 14C] 1 illustrates the use of a projector in a wearable multimedia device to project various types of information onto the palm of a user's hand, according to one embodiment. [Figure 14D] 1 illustrates the use of a projector in a wearable multimedia device to project various types of information onto the palm of a user's hand, according to one embodiment. [Figure 15A] 1 illustrates an application of a projector in which information is projected onto a car engine to assist a user in checking their engine oil, according to one embodiment. [Figure 15B] 1 illustrates an application of a projector in which information is projected onto a car engine to assist a user in checking their engine oil, according to one embodiment. [Figure 16] 1 illustrates an application of a projector where information is projected onto a cutting board to assist a home cook in chopping vegetables, according to one embodiment. [Figure 17] FIG. 1 is a system block diagram of a projector architecture according to one embodiment. [Figure 18] 10 illustrates the adjustment of laser parameters based on different surface geometries or materials, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0030] The same reference symbols used in the various drawings indicate like elements.

[0031] Overview A wearable multimedia device is a lightweight, compact, battery-powered device that can be attached to a user's clothing or an object using a tension clasp, interlocking pinback, magnets, or any other attachment mechanism. The wearable multimedia device includes a digital image capture device (e.g., 180° FOV with optical image stabilizer (OIS)) that allows a user to spontaneously capture multimedia data (e.g., video, audio, depth data) of life events ("moments") and document transactions (e.g., financial transactions) with minimal user interaction or device setup. The multimedia data ("context data") captured by the wireless multimedia device is uploaded to a cloud computing platform with an application ecosystem that enables the context data to be processed, edited, and formatted by one or more applications (e.g., artificial intelligence (AI) applications) into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played on the wearable multimedia device and / or any other playback device. For example, the cloud computing platform can convert video and audio data into any desired shooting style (eg, documentary, lifestyle, snapshot, photojournalism, sports, street) specified by the user.

[0032] In one embodiment, the contextual data is processed by a server computer of a cloud computing platform based on user preferences. For example, based on the user preferences, an image can be color-graded, stabilized, and cropped to perfectly match the moment the user wants to recreate. The user preferences can be stored in a user profile created by the user through an online account accessible through a website or portal, or the user preferences can be learned by the platform over time (e.g., using machine learning). In one embodiment, the cloud computing platform is a scalable distributed computing environment. For example, the cloud computing platform can be a distributed streaming platform (e.g., Apache Kafka™) with real-time streaming data pipelines and streaming applications that transform or react to streams of data.

[0033] In one embodiment, a user can start and stop a context data capture session on the wearable multimedia device with a simple touch gesture (e.g., a tap or swipe), by issuing a command, or by any other input mechanism. All or part of the wearable multimedia device can automatically power down when it detects that the device is not being worn by the user using one or more sensors (e.g., a proximity sensor, an optical sensor, an accelerometer, a gyroscope).

[0034] The context data can be encrypted and compressed using any desired encryption or compression technique and stored in an online database associated with the user account. The context data can be stored for a specified period of time that can be set by the user. Users can be provided with opt-in mechanisms and other tools to manage their data and data privacy through a website, portal, or mobile application.

[0035] In one embodiment, the context data includes point cloud data to provide three-dimensional (3D) surface mapping objects that can be processed using, for example, augmented reality (AR) and virtual reality (VR) applications in the application ecosystem. The point cloud data can be generated by a depth sensor (e.g., lidar or time-of-flight (TOF)) integrated on the wearable multimedia device.

[0036] In one embodiment, the wearable multimedia device includes a Global Navigation Satellite System (GNSS) receiver (e.g., a Global Positioning System (GPS)) and one or more inertial sensors (e.g., an accelerometer, a gyroscope) to determine the position and orientation of a user wearing the device at the time the context data was captured. In one embodiment, one or more images in the context data can be used by a location application, such as a visual odometry application, in the application ecosystem to determine the location and orientation of the user.

[0037] In one embodiment, the wearable multimedia device may also include one or more environmental sensors, including but not limited to ambient light sensors, magnetometers, pressure sensors, voice activity detectors, etc. This sensor data may be included in the context data to enhance content presentation with additional information that can be used to capture moments.

[0038] In one embodiment, the wearable multimedia device may include one or more biometric sensors, such as a heart rate sensor, a fingerprint scanner, etc. This sensor data may be included in the context data to document a transaction or indicate the user's emotional state during a moment (e.g., a high heart rate may indicate excitement or fear).

[0039] In one embodiment, the wearable multimedia device includes a headphone jack for connecting a headset or earbuds and one or more microphones for receiving voice commands and capturing ambient audio. In an alternative embodiment, the wearable multimedia device includes short-range communication technologies, including but not limited to Bluetooth, IEEE 802.15.4 (ZigBee™), and Near Field Communication (NFC). Short-range communication technologies can be used in addition to or instead of a headphone jack to wirelessly connect to a wireless headset or earbuds and / or to any other external device (e.g., a computer, printer, projector, television, and other wearable devices).

[0040] In one embodiment, the wearable multimedia device includes wireless transceivers and communication protocol stacks for various communication technologies, including WiFi, 3G, 4G, and 5G communication technologies. In one embodiment, the headset or earbuds also include sensors (e.g., biometric sensors, inertial sensors) that provide information about the direction the user is facing in order to provide commands, such as with head gestures. In one embodiment, the camera direction can be controlled by head gestures so that the camera field of view follows the user's direction of view. In one embodiment, the wearable multimedia device can be integrated into or attached to the user's eyeglasses.

[0041] In one embodiment, the wearable multimedia device includes a projector (e.g., laser projector, LCoS, DLP, LCD) or can be hardwired or wirelessly coupled to an external projector, allowing the user to play back moments on a surface such as a wall or tabletop. In another embodiment, the wearable multimedia device includes an output port that can be connected to a projector or other output device.

[0042] In one embodiment, the wearable multimedia capturing device includes a touch surface that responds to touch gestures (e.g., tap, multi-tap, or swipe gestures). The wearable multimedia device may include a small display for presenting information and one or more light indicators to show on / off status, power status, or any other desired status.

[0043] In one embodiment, the cloud computing platform can be driven by context-based gestures (e.g., air gestures) in combination with voice queries, such as a user pointing at an object in their environment and saying, "What's that building?" The cloud computing platform uses the air gestures to narrow the camera's viewport to isolate the building. One or more images of the building are captured and sent to the cloud computing platform, where an image recognition application can run the image query and store or return the results to the user. Air and touch gestures can also be made on a projected ephemeral display, for example, in response to user interface elements.

[0044] In one embodiment, the context data can be encrypted on the device and on a cloud computing platform so that only the user or any authorized viewer can replay the moment on a connected screen (e.g., smartphone, computer, television, etc.) or as a projection on a surface. An example architecture for a wearable multimedia device is described with respect to FIG.

[0045] In addition to personal life events, wearable multimedia devices simplify the capture of financial transactions that are now handled by smartphones. Capturing everyday transactions (e.g., business transactions, microtransactions) is made easier, faster, and more fluid by using the visually assisted contextual awareness provided by wearable multimedia devices. For example, when a user engages in a financial transaction (e.g., making a purchase), the wearable multimedia device will generate data that stores the financial transaction, including the date, time, amount, digital images or video of the parties involved, audio (e.g., user annotations describing the transaction), and environmental data (e.g., location data). The data can be included in a multimedia data stream sent to a cloud computing platform, where it can be stored online and / or processed by one or more financial applications (e.g., financial management, accounting, budgeting, tax preparation, inventory, etc.).

[0046] In one embodiment, the cloud computing platform provides a graphical user interface on a website or portal that enables various third-party application developers to upload, update, and manage their applications in the application ecosystem. Some example applications can include, but are not limited to, personal live streaming (e.g., Instagram™ Live, Snapchat™), elderly monitoring (e.g., to ensure a loved one has taken their medication), recall (e.g., showing a child's soccer game from last week), and personal guides (e.g., an AI-enabled personal guide that knows the user's location and guides the user to take actions).

[0047] In one embodiment, the wearable multimedia device includes one or more microphones and a headset. In some embodiments, the headset wire includes the microphone. In one embodiment, a digital assistant is implemented on the wearable multimedia device that responds to user queries, requests, and commands. For example, a wearable multimedia device worn by a parent captures instantaneous context data for a child's soccer game, particularly the "moment" when the child scores a goal. The user can request (e.g., using a voice command) that the platform create a video clip of the goal and store it in the user's user account. Without any further action by the user, the cloud computing platform identifies the correct portion of the instantaneous context data when the goal is scored (e.g., using facial recognition, visual, or audio cues), compiles the instantaneous context data into a video clip, and stores the video clip in a database associated with the user account.

[0048] In one embodiment, the device can include photovoltaic surface technology to extend battery life, as well as inductive charging circuitry (e.g., Qi) to enable inductive charging on a charging mat and wireless over-the-air (OTA) charging.

[0049] In one embodiment, the wearable multimedia device is configured to magnetically couple or mate with a rechargeable portable battery pack. The portable battery pack includes a mating surface on which a permanent magnet (e.g., north pole) is disposed, and the wearable multimedia device has a corresponding mating surface on which a permanent magnet (e.g., south pole) is disposed. Any number of permanent magnets having any desired shape or size can be arranged in any desired pattern on the mating surface.

[0050] The permanent magnets hold the portable battery pack and the wearable multimedia device together in a mated configuration with the garment (e.g., the user's shirt) in between. In one embodiment, the portable battery pack and the wearable multimedia device have the same mating surface dimensions, so that there are no overhanging portions when in the mated configuration. The user places the portable battery pack on the back of their garment and then places the wearable multimedia device on top of the portable battery pack on the outside of their garment, so that the permanent magnets attract each other through the garment, magnetically fastening the wearable multimedia device to their garment. In one embodiment, the portable battery pack has a built-in wireless power transmitter that is used to wirelessly power the wearable multimedia device while in the mated configuration using the principle of resonant inductive coupling. In one embodiment, the wearable multimedia device includes a built-in wireless power receiver that is used to receive power from the portable battery pack while in the mated configuration.

[0051] Operating environment example 1 is a block diagram of an operating environment for a wearable multimedia device and a cloud computing platform with an application ecosystem for processing multimedia data captured by the wearable multimedia device, according to one embodiment. Operating environment 100 includes wearable multimedia device 101, cloud computing platform 102, network 103, application (“app”) developer 104, and third-party platform 105. Cloud computing platform 102 is coupled to one or more databases 106 for storing contextual data uploaded by wearable multimedia device 101.

[0052] As described above, the wearable multimedia device 101 is a lightweight, compact, battery-powered device that can be attached to a user's clothing or an object using a tension clasp, an interlocking pinback, a magnet, or any other attachment mechanism. The wearable multimedia device 101 includes a digital image capture device (e.g., 180° FOV with OIS) that allows a user to spontaneously capture "instant" multimedia data (e.g., video, audio, depth data) and document everyday transactions (e.g., financial transactions) with minimal user interaction or device setup. Contextual data captured by the wireless multimedia device 101 is uploaded to a cloud computing platform 102. The cloud computing platform 102 includes an application ecosystem that allows the contextual data to be processed, edited, and formatted by one or more server-side applications into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played on the wearable multimedia device and / or other playback devices.

[0053] As an example, at a child's birthday party, a parent can clip a wearable multimedia device onto their clothing (or attach the device to a necklace or chain and wear it around their neck) with the camera lens facing their line of sight. The camera includes a 180° FOV, allowing the camera to capture almost everything the user is currently viewing. The user can start recording by simply tapping the surface of the device or pressing a button. No additional setup is required. A multimedia data stream (e.g., video with audio) capturing special moments of the birthday (e.g., blowing out the candles) is recorded. This "context data" is sent in real time over a wireless network (e.g., WiFi, cellular) to the cloud computing platform 102. In one embodiment, the context data is stored on the wearable multimedia device so that it can be uploaded at a later time. In another embodiment, the user can transfer the context data to another device (e.g., a personal computer hard drive, a smartphone, a tablet computer, a thumb drive) and later use an application to upload the context data to the cloud computing platform 102.

[0054] In one embodiment, the context data is processed by one or more applications in an application ecosystem hosted and managed by the cloud computing platform 102. The applications are accessible through their respective application programming interfaces (APIs). A custom distributed streaming pipeline is created by the cloud computing platform 102 to process the context data based on one or more of data type, data volume, data quality, user preferences, templates, and / or any other information to generate a desired presentation based on user preferences. In one embodiment, machine learning techniques can be used to automatically select appropriate applications to include in the data processing pipeline with or without user preferences. For example, historical user context data stored in a database (e.g., a NoSQL database) can be used to determine user preferences for data processing using any suitable machine learning technique (e.g., deep learning or convolutional neural networks).

[0055] In one embodiment, the application ecosystem can include a third-party platform 105 that processes the context data. A secure session is established between the cloud computing platform 102 and the third-party platform 105 to send and receive the context data. This design allows third-party app providers to control access to their own applications and provide updates. In another embodiment, the applications run on the cloud computing platform 102's servers, and updates are sent to the cloud computing platform 102. In the latter embodiment, app developers 104 can use APIs provided by the cloud computing platform 102 to upload and update applications to be included in the application ecosystem.

[0056] Data Processing System Example Figure 2 is a block diagram of a data processing system implemented by the cloud computing platform of Figure 1, according to one embodiment. Data processing system 200 includes recorder 201, video buffer 202, audio buffer 203, photo buffer 204, ingestion server 205, data store 206, video processor 207, audio processor 208, photo processor 209, and third-party processor 210.

[0057] A recorder 201 (e.g., a software application) running on the wearable multimedia device records video, audio, and photo data (“context data”) captured by the camera and audio subsystems and stores the data in buffers 202, 203, and 204, respectively. This context data is then sent (e.g., using wireless OTA technology) to an ingestion server 205 of the cloud computing platform 102. In one embodiment, the data can be sent in separate data streams, each with a unique stream identifier (streamid). Streams are distinct pieces of data that can include, as example attributes, location (e.g., latitude, longitude), user, audio data, video streams of various durations, and N photos. Streams can have durations from 1 to MAXSTREAM_LEN seconds, where MAXSTREAM_LEN=20 seconds in this example.

[0058] The ingestion server 205 ingests the streams and stores the results of the processors 207-209 by creating stream records in the data store 206. In one embodiment, the audio stream is processed first and used to determine what other streams are needed. The ingestion server 205 routes the streams to the appropriate processors 207-209 based on the streamid. For example, a video stream is routed to the video processor 207, an audio stream to the audio processor 208, and a photo stream to the photo processor 209. In one embodiment, at least a portion of the data collected from the wearable multimedia device (e.g., image data) is processed and encrypted into metadata so that it can be further processed by a given application, and then sent back to the wearable multimedia device or other device.

[0059] Processors 207-209 can execute proprietary or third-party applications, as described above. For example, video processor 207 can be a video processing server that sends raw video data stored in video buffer 202 to a set of one or more image processing / editing applications 211, 212 based on user preferences or other information. Processor 207 sends requests to applications 211, 212 and returns results to ingestion server 205. In one embodiment, third-party processor 210 can process one or more of the streams using its own processor and application. In another example, audio processor 208 can be an audio processing server that sends audio data stored in audio buffer 203 to speech-to-text application 213.

[0060] Scene Recognition Application Example 3 is a block diagram of a data processing pipeline for processing a contextual data stream, according to one embodiment. In this embodiment, the data processing pipeline 300 is created and configured to determine what a user is viewing based on contextual data captured by a wearable multimedia device worn by the user. The ingestion server 301 receives an audio stream (e.g., including user annotations) from the audio buffer 203 of the wearable multimedia device and sends the audio stream to an audio processor 305. The audio processor 305 sends the audio stream to an app 306, which performs speech-to-text conversion and returns analyzed text to the audio processor 305. The audio processor 305 returns the analyzed text to the ingestion server 301.

[0061] The video processor 302 receives the parsed text from the capture server 301 and sends a request to the video processing app 307. The video processing app 307 identifies objects in the video scene and labels the objects using the parsed text. The video processing app 307 sends a response describing the scene (e.g., labeled objects) to the video processor 302. The video processor then forwards the response to the capture server 301. The capture server 301 sends the response to the data merging process 308, which merges the response with the user's location, orientation, and map data. The data merging process 308 returns a response with a scene description to the recorder 304 on the wearable multimedia device. For example, the response could include text describing the scene as a child's birthday party, including a map location and a description of the objects in the scene (e.g., identifying people in the scene). The recorder 304 associates the scene description with the multimedia data stored on the wearable multimedia device (e.g., using a stream ID). When the user recalls the data, the data is enriched with the scene description.

[0062] In one embodiment, the data merge process 308 may use more than just location and map data. There can also be the concept of ontologies. For example, the facial features of a user's father captured in an image can be recognized by a cloud computing platform and returned as "Dad" rather than the user's name, and an address such as "555 Main Street, San Francisco, CA" can be returned as "Home." Ontologies can be specific to a user and can grow and learn from user input.

[0063] Transportation Application Examples FIG. 4 is another data processing block diagram for processing a contextual data stream for a transportation application, according to one embodiment. In this embodiment, a data processing pipeline 400 is created to call a transportation company (e.g., Uber®, Lyft®) for a ride home. Contextual data from the wearable multimedia device is received by an ingestion server 401, and an audio stream from the audio buffer 203 is sent to an audio processor 405. The audio processor 405 sends the audio stream to an app 406, which converts the speech to text. The parsed text is returned to the audio processor 405, which returns the parsed text (e.g., the user's spoken request for transportation) to the ingestion server 401. The processed text is sent to a third-party processor 402. The third-party processor 402 sends the user location and a token to a third-party application 407 (e.g., an Uber® or Lyft® application). In one embodiment, the token is an API and authorization token used to broker requests on behalf of the user. The application 407 returns a response data structure to the third-party processor 402, which forwards the response data structure to the ingestion server 401. The ingestion server 401 checks the ride arrival status (e.g., ETA) in the response data structure and sets up a callback to the user in the user callback queue 408. The ingestion server 401 returns a response to the recorder 404 with a vehicle description, which can be played to the user through loudspeakers on the wearable multimedia device or by the digital assistant through the user's headphones or earbuds via a wired or wireless connection.

[0064] Figure 5 illustrates data objects used by the data processing system of Figure 2, according to one embodiment. The data objects are part of a software component infrastructure instantiated on a cloud computing platform. The "Streams" object includes data streamid, deviceid, start, end, lat, lon, attributes, and entities. The "streamid" identifies the stream (e.g., video, audio, photo), the "deviceid" identifies the wearable multimedia device (e.g., mobile device ID), the "start" is the start time of the context data stream, the "end" is the end time of the context data stream, the "lat" is the latitude of the wearable multimedia device, the "lon" is the longitude of the wearable multimedia device, the "attributes" include, for example, birthday, facial feature points, skin color, audio characteristics, address, phone number, etc., and the "entities" constitute an ontology. For example, the name "John Do" may be mapped to "Dad" or "Brother" depending on the user.

[0065] The "Users" object contains the data userid, deviceid, email, fname, and lname. userid identifies the user by a unique identifier, deviceid identifies the wearable device by a unique identifier, email is the user's registered email address, fname is the user's first name, and lname is the user's last name. The "Userdevices" object contains the data userid and deviceid. The "Devices" object contains the data deviceid, started, state, modified, and created. In one embodiment, deviceid is a unique identifier for the device (e.g., separate from the MAC address). started is when the device was first started. state is on / off / sleep. modified is the last modification date, reflecting the last state change or operating system (OS) change. created is the first time the device was turned on.

[0066] The "ProcessingResults" object contains the data streamid, ai, result, callback, duration, and accuracy. In one embodiment, streamid is each user stream as a universally unique identifier (UUID). For example, a stream that started from 8:00 AM to 10:00 AM would have id:15h158dhb4, and a stream that started from 10:15 AM to 10:18 AM would have the UUID contacted for this stream. ai is the identifier for the platform application that was contacted for this stream. result is the data sent from the platform application. callback is the callback used (the callback is tracked in case the platform needs to replay the request, as versions may change). accuracy is a score for how accurate the result set was. In one embodiment, the processing results can be used for multiple tasks, such as 1) notifying a merge server of the full set of results, 2) determining the fastest AI so that the user experience can be improved, and 3) determining the most accurate AI. Depending on your use case, you may want to prioritize speed over accuracy, or vice versa.

[0067] An "Entities" object contains the data entityID, userID, entityName, entityType, and entityAttribute. The entityID is the UUID for the entity, and entities have multiple fields where the entityID references that single entity. For example, "Barack Obama" might have an entityID of 144, which could be linked to POTUS44 or "Barack Hussein Obama" or "President Obama" in an association table. The userID identifies the user for whom the entity record was created. The entityName is the name the userID calls the entity. For example, Malia Obama's entityName for entityID 144 could be "Dad" or "Papa." The entityType is a person, place, or thing. The entityAttribute is an array of attributes about that entity that are specific to the userID's understanding of the entity. So, for example, when Malia utters a voice query: "Can you see Dad?", the cloud computing platform translates the query to Barack Hussein Obama and maps the entities together so that it can be used to broker requests to third parties or look up information in the system.

[0068] Process Example 6 is a flow diagram of a data pipeline process according to one embodiment. The process 600 can be implemented using the wearable multimedia device 101 and cloud computing platform 102 described with respect to FIGS.

[0069] Process 600 can begin by receiving context data from a wearable multimedia device 601. For example, the context data can include video, audio, and still images captured by the camera and audio subsystem of the wearable multimedia device.

[0070] Process 600 can continue by creating (e.g., instantiating) 602 a data processing pipeline with applications based on the context data and user requests / preferences. For example, based on user requests or preferences, and also based on data type (e.g., audio, video, photo), one or more applications can be logically connected to form a data processing pipeline that processes the context data into a presentation to be played on the wearable multimedia device or another device.

[0071] Process 600 can continue by processing 603 the contextual data in a data processing pipeline. For example, audio from user annotations during a moment or transaction can be converted to text, which can then be used to label objects in a video clip.

[0072] Process 600 may continue by sending (604) the output of the data processing pipeline to a wearable multimedia device and / or other playback device.

[0073] Cloud Computing Platform Architecture Example 7 illustrates an example architecture 700 for the cloud computing platform 102 described with respect to FIGS. 1-6 and 9 , according to one embodiment. Other architectures are possible, including architectures with more or fewer components. In some implementations, the architecture 700 includes one or more processors 702 (e.g., dual-core Intel® Xeon® processors), one or more network interfaces 706, one or more storage devices 704 (e.g., hard disks, optical disks, flash memory), and one or more computer-readable media 708 (e.g., hard disks, optical disks, flash memory, etc.). These components can communicate and exchange data through one or more communication channels 710 (e.g., buses), which can utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0074] The term "computer-readable medium" refers to any medium that participates in providing instructions to processor 702 for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media include, but are not limited to, coaxial cables, copper wire, and fiber optics.

[0075] The computer readable medium 708 may further include an operating system 712 (eg, Mac OS Server, Windows NT Server, Linux Server), a network communication module 714, interface instructions 718, and data processing instructions 716.

[0076] Operating system 712 can be multi-user, multi-processing, multi-tasking, multi-threading, real-time, etc. Operating system 712 performs basic tasks, including, but not limited to, recognizing input from and providing output to devices 702, 704, 706, and 708, tracking and managing files and directories on computer-readable medium 708 (e.g., memory or storage device), controlling peripheral devices, and managing traffic on one or more communication channels 710. Network communication module 714 includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.), as well as various components for creating a distributed streaming platform using, for example, Apache Kafka™. Data processing instructions 716 include server-side or back-end software for implementing server-side operations, as described with respect to FIGS. 1-6. The interface instructions 718 include software for implementing a web server and / or portal for sending and receiving data to / from the wearable multimedia device 101, third party application developers 104 and third party platforms 105, as described with respect to FIG. 1.

[0077] The architecture 700 can be included in any computing device, including one or more server computers in a local or distributed network, each having one or more processing cores. The architecture 700 can be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software can include multiple software components or can be a single body of code.

[0078] Wearable multimedia device architecture example Figure 8 is a block diagram of an example architecture 800 for a wearable multimedia device that implements the features and processes described with respect to Figures 1-6 and 9. The architecture 800 may include a memory interface 802, a data processor, image processor, or central processing unit 804, and a peripherals interface 806. The memory interface 802, the processor 804, or the peripherals interface 806 may be separate components or may be integrated into one or more integrated circuits. One or more communication buses or signal lines may couple the various components.

[0079] Sensors, devices, and subsystems may be coupled to the peripherals interface 806 to facilitate multiple functions. For example, a motion sensor 810, a biometric sensor 812, and a depth sensor 814 may be coupled to the peripherals interface 806 to facilitate motion, orientation, biometric, and depth sensing functions. In some implementations, the motion sensor 810 (e.g., accelerometer, rate gyroscope) may be utilized to detect movement and orientation of the wearable multimedia device.

[0080] Other sensors, such as environmental sensors (e.g., temperature sensors, barometers, ambient light), may also be connected to the peripherals interface 806 to facilitate environmental sensing functions. For example, biometric sensors may detect fingerprints, facial recognition, heart rate, and other fitness parameters. In one embodiment, a haptic motor (not shown) may be coupled to the peripherals interface and may provide vibration patterns as haptic feedback to the user.

[0081] A position processor 815 (e.g., a GNSS receiver chip) may be connected to the peripherals interface 806 to provide georeferencing. An electronic magnetometer 816 (e.g., an integrated circuit chip) may also be connected to the peripherals interface 806 to provide data that can be used to determine the direction of magnetic north. In this manner, the electronic magnetometer 816 may be used by an electronic compass application.

[0082] A camera subsystem 820 and light sensor 822, such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) light sensor, may be utilized to facilitate camera functions such as recording photographs and video clips. In one embodiment, the camera has a 180° FOV and OIS. The depth sensor may include an infrared emitter that projects dots in a known pattern onto the object / subject. The dots are then photographed by a dedicated infrared camera and analyzed to determine depth data. In one embodiment, a time-of-flight (TOF) camera can be used to determine distance based on the known speed of light, measuring the time of flight of the light signal between the camera and the object / subject for each point in the image.

[0083] Communication functions may be facilitated through one or more communications subsystems 824. The communications subsystems 824 may include one or more wireless communications subsystems. The wireless communications subsystems 824 may include radio frequency receivers and transmitters and / or optical (e.g., infrared) receivers and transmitters. The wired communications system may include a port device, such as a universal serial bus (USB) port or some other wired port connection, that can be used to establish a wired connection to other computing devices, such as other communications devices, network access devices, personal computers, printers, display screens, or other processing devices capable of receiving or transmitting data (e.g., projectors).

[0084] The specific design and implementation of the communications subsystem 824 may depend on the communications network or medium over which the device is intended to operate. For example, the device may include a wireless communications subsystem designed to operate over a Global System for Mobile Communications (GSM) network, a GPRS network, an Enhanced Data GSM Environment (EDGE) network, an IEEE 802.xx communications network (e.g., WiFi, WiMax, ZigBee™), 3G, 4G, 4G LTE, a code division multiple access (CDMA) network, a near field communication (NFC), Wi-Fi Direct, and a Bluetooth™ network. The wireless communications subsystem 824 may include a hosting protocol so that the device can be configured as a base station for other wireless devices. As another example, the communications subsystem may enable the device to synchronize with a host device using one or more protocols or communications technologies, such as the TCP / IP protocol, the HTTP protocol, the UDP protocol, the ICMP protocol, the POP protocol, the FTP protocol, the IMAP protocol, the DCOM protocol, the DDE protocol, the SOAP protocol, HTTP Live Streaming, MPEG Dash, and any other known communications protocol or technology.

[0085] The audio subsystem 826 may be coupled to a speaker 828 and one or more microphones 830 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, telephony, and beamforming.

[0086] The I / O subsystem 840 may include a touch controller 842 and / or another input controller 844. The touch controller 842 may be coupled to a touch surface 846. The touch surface 846 and touch controller 842 may detect contact and movement or disruption using any of several touch sensitivity technologies, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements to determine one or more points of contact with the touch surface 846. In one implementation, the touch surface 846 may display virtual or soft buttons, which may be used as input / output devices by a user.

[0087] Other input controller 844 may be coupled to other input / control devices 848, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses. One or more buttons (not shown) may include up / down buttons for adjusting the volume of speaker 828 and / or microphone 830.

[0088] In some implementations, device 800 plays user-recorded audio and / or video files, such as MP3, AAC, and MPEG video files. In some implementations, device 800 may include MP3 player functionality and may include pin connectors or other ports for tethering to other devices. Other input / output and control devices may also be used. In one embodiment, device 800 may include an audio processing unit for streaming audio to accessory devices over a direct or indirect communication link.

[0089] The memory interface 802 may be coupled to memory 850. The memory 850 may include high-speed random-access memory or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, or flash memory (e.g., NAND, NOR). The memory 850 may store an operating system 852, such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks. The operating system 852 may include instructions for handling basic system services and for performing hardware-dependent tasks. In some implementations, the operating system 852 may include a kernel (e.g., a UNIX kernel).

[0090] The memory 850 may also store communication instructions 854 that facilitate communication with one or more additional devices, one or more computers, or servers, including peer-to-peer communication with wireless accessory devices, as described with respect to Figures 1-6. The communication instructions 854 may also be used to select an operating mode or communication medium for use by the device based on the geographic location of the device.

[0091] Memory 850 may include sensor processing instructions 858 that facilitate sensor-related processing and functions, and recorder instructions 860 that facilitate recording functions, as described with respect to Figures 1-6. Other instructions may include GNSS / navigation instructions that facilitate GNSS and navigation-related processes, camera instructions that facilitate camera-related processes, and user interface instructions that facilitate user interface processing, including touch models for interpreting touch input.

[0092] Each of the above-identified instructions and applications may correspond to a set of instructions for performing one or more functions described above. These instructions need not be implemented as separate software programs, procedures, or modules. Memory 850 may include additional or fewer instructions. Furthermore, various functions of the device may be implemented in hardware and / or software, including in one or more signal processing and / or application-specific integrated circuits (ASICs).

[0093] Graphical User Interface Example 9 is a screenshot of an example graphical user interface (GUI) 900 for use with the scene identification application described with respect to FIG. 3 , according to one embodiment. GUI 900 includes a video pane 901, time / location data 902, objects 903, 906a, 906b, 906c, a search button 904, a menu of categories 905, and thumbnail images 907. GUI 900 can be presented on a user device (e.g., a smartphone, tablet computer, wearable device, desktop computer, notebook computer), for example, through a client application or through a web page provided by a web server of cloud computing platform 102. In this example, a user captured a digital image of a young man in video pane 901 standing on Orchard Street, New York, New York, at 12:45 PM on October 18, 2018, as indicated by time / location data 902.

[0094] In one embodiment, the image is processed through an object detection framework implemented on the cloud computing platform 102, such as a Viola-Jones object detection network. For example, a model or algorithm is used to generate a region of interest or region proposal, which includes a set of bounding boxes spanning the entire digital image. Visual features are extracted for each of the bounding boxes and evaluated to determine whether and which objects are present in the region proposal based on the visual features. Overlapping boxes are combined (e.g., using non-maximum suppression) into a single bounding box. In one embodiment, overlapping boxes are also used to organize objects into categories for big data storage. For example, object 903 (the young man) is considered a parent object, and objects 906a-906c (the clothing items he is wearing) are considered child objects (shoes, shirt, pants) of object 903 due to their overlapping bounding boxes. Thus, a search for "person" using a search engine results in all objects labeled "person" and their child objects, if any.

[0095] In an embodiment, complex polygons are used to identify objects in an image rather than bounding boxes. Complex polygons are used to determine highlight / hotspot areas in an image, for example, where a user is pointing. Only complex polysegmentation portions (rather than the entire image) are sent to the cloud computing platform, improving privacy, security, and speed.

[0096] Other examples of object detection frameworks that may be implemented by the cloud computing platform 102 to detect and label objects in digital images include, but are not limited to, Region Convolutional Neural Network (R-CNN), Fast R-CNN, and Faster R-CNN.

[0097] In this example, objects identified in the digital image include people, cars, buildings, roads, windows, doors, stairs, signs, and text. The identified objects are organized and presented as categories for the user to search. The user selects the category "People" using a cursor or a finger (if using a touch-sensitive screen). By selecting the category "People," object 903 (i.e., the young man in the image) is separated from the remaining objects in the digital image, and a subset of objects 906a-906c are displayed in thumbnail image 907 along with their respective metadata. Object 906a is labeled "Orange, Shirt, Button, Short Sleeves," object 906b is labeled "Blue, Jeans, Ripped, Denim, Pocket, Phone," and object 906c is labeled "Blue, Nike, Shoe, Left, Air Max, Red Sock, White Swoosh, Logo."

[0098] Pressing the search button 904 initiates a new search based on the category selected by the user and the particular image in the video pane 901. The search results include thumbnail images 907. Similarly, if the user selects the category "cars" and then presses the search button 904, a new collection of thumbnails 907 will be displayed showing all the cars captured in the images along with their respective metadata.

[0099] FIG. 10 illustrates a classifier framework 1000 for classifying raw or pre-processed context data into objects and metadata that can be searched using the GUI 900 of FIG. 9 , according to one embodiment. The framework 1000 includes an API 1001, classifiers 1002a-1002n, and a data store 1005. Raw or pre-processed context data captured on a wearable multimedia device is uploaded through the API 1001. The context data is passed through the classifiers 1002a-1002n (e.g., neural networks). In one embodiment, the classifiers 1002a-1002n are trained using context data crowd-sourced from a large number of wearable multimedia devices. The output of the classifiers 1002a-1002n are objects and metadata (e.g., labels) that are stored in the data store 1005. A search index is generated for the objects / metadata in the data store 1005, which can be used by a search engine to search for objects / metadata that satisfy a search query entered using the GUI 900. Various types of search indexes can be used, including but not limited to tree indexes, suffix tree indexes, inverted indexes, citation index analysis, and n-gram indexes.

[0100] Classifiers 1002a-1002n are selected and added to the dynamic data processing pipeline to generate the desired presentation based on one or more of data type, data volume, data quality, user preferences, user-initiated or application-initiated search queries, voice commands, application requirements, templates, and / or any other information. Any known classifier can be used, including neural networks, support vector machines (SVMs), random forests, boosted decision trees, and any combination of these individual classifiers using voting, stacking, and grading techniques. In one embodiment, some of the classifiers are user-personalized, i.e., the classifiers are trained solely on contextual data from a particular user device. Such classifiers can be trained to detect and label people and objects that are personal to the user. For example, one classifier can be used for face detection, which detects faces of individuals known to the user (e.g., family members, friends) in images labeled, for example, by user input.

[0101] As an example, a user can utter multiple phrases such as "Make a movie from my videos that includes my mom and dad in New Orleans," "Add jazz music as a soundtrack," "Send me a drink recipe for making a Hurricane cocktail," and "Send me directions to the nearest liquor store." The speech phrases are analyzed by the cloud computing platform 102, and the words are used to assemble a personalized processing pipeline that performs the requested task, including adding a classifier to detect the faces of the user's mother and father.

[0102] In one embodiment, AI is used to determine how a user interacts with the cloud computing platform during a messaging session. For example, if a user issues the message "Bob, have you seen 'Toy Story 4'?", the cloud computing platform determines who Bob is and parses "Bob" from the string sent to a message relay server on the cloud computing platform. Similarly, if the message says "Bob, look at this," the platform device sends the image with the message in one step without having to attach the image as a separate transaction. The image can be visually confirmed by the user using projector 1115 and any desired surface before sending it to Bob. Additionally, the platform maintains a persistent and personal communication channel with Bob for some time, eliminating the need to precede each communication with the name "Bob" during a messaging session.

[0103] Context Data Broker Service In one embodiment, a context data brokerage service is provided through a cloud computing platform 102. The service allows users to sell their private raw or processed context data to entities of their choice. The platform 102 hosts the context data brokerage service and provides the security protocols needed to protect the privacy of user context data. The platform 102 also facilitates transactions and monetary transactions or settlements between entities and users.

[0104] The raw and pre-processed context data can be stored using big data storage. Big data storage supports storage and I / O operations on storage involving large numbers of data files and objects. In one embodiment, big data storage includes an architecture consisting of a redundant and scalable supply of infrastructure based on direct-attached storage (DAS) pools, scale-out or clustered network-attached storage (NAS), or object storage formats. The storage infrastructure is connected to computing server nodes that enable rapid processing and retrieval of large amounts of data. In one embodiment, the big data storage architecture includes native support for big data analytics solutions such as Hadoop™, Cassandra™, and NoSQL™.

[0105] In one embodiment, entities interested in purchasing raw or processed context data subscribe to the context data broker service through a registration GUI or web page of the cloud computing platform 102. Once registered, the entities (e.g., companies, advertising agencies) are enabled to transact directly or indirectly with users through one or more GUIs tailored to facilitate data brokerage. In one embodiment, the platform 102 can match an entity's request for a particular type of context data with users who can provide the context data. For example, a clothing company may be interested in all images in which the logos of its clothing or competitors' companies are detected. The clothing company can then use the context data to better identify the demographics of its customers. In another example, a news organization or political campaign may be interested in video footage of newsworthy events to use in a cover story or feature article. Various companies may be interested in users' search or purchase histories for improved ad targeting or other marketing projects. Various entities may be interested in purchasing context data for use as training data for other object detectors, such as object detectors for autonomous vehicles.

[0106] In one embodiment, a user's raw or processed contextual data is made available in a secure format to protect the user's privacy. Both users and entities can have their own online accounts to deposit and withdraw money resulting from brokered transactions. In one embodiment, the data broker service collects transaction fees based on a pricing model. Fees can also be earned through traditional online advertising (e.g., click-throughs on banner ads, etc.).

[0107] In one embodiment, individuals can create their own metadata using the wearable multimedia device. For example, a celebrity chef may wear the wearable multimedia device 101 while preparing a meal. Objects in the image are labeled using metadata provided by the chef. Users can obtain access to the metadata from a broker service. When a user attempts to prepare a dish while wearing their wearable multimedia device 101, objects are detected, and metadata provided by the chef to help the user recreate something (e.g., timing, amount, order, scale) is projected onto the user's work surface (e.g., cutting board, countertop, range, oven, etc.), such as dimensions, cooking time, and additional tips. For example, a piece of meat may be detected on a cutting board, and text may be projected onto the cutting board by projector 1115 reminding the user to cut the meat perpendicular to the grain, as well as a dimensional guide projected onto the surface of the meat to guide the user in cutting slices of uniform thickness according to the chef's metadata. Laser projection guides can also be used to cut vegetables (e.g., julienne, mince) to uniform thicknesses. Users can upload their metadata to the cloud service platform, create their own channels, and earn revenue through subscriptions and advertisements, similar to the YouTube® platform.

[0108] 11 is a system block diagram illustrating a hardware architecture 1100 for a wearable multimedia device, according to one embodiment. The architecture 1100 includes a system-on-chip (SoC) 1101 (e.g., a Qualcomm Snapdragon® chip), a main camera 1102, a 3D camera 1103, a capacitance sensor 1104, a motion sensor 1105 (e.g., accelerometer, gyro, magnetometer), a microphone 1106, a memory 1107, a global navigation satellite system receiver (e.g., GPS receiver) 1108, a WiFi / Bluetooth chip 1109, a wireless transceiver chip 1110 (e.g., 4G, 5G), a radio frequency (RF) transceiver chip 1112, an RF front-end electronics (RFFE) 1113, an LED 1114, a projector 1115 (e.g., laser projection, pico projector, LCoS, DLP, LCD), an audio amplifier 1116, a speaker 1117, an external battery 1118 (e.g., battery pack), a magnetic inductance network 1119, a power management chip (PMIC) 1120, and an internal battery 1121. All of these components work together to facilitate the various tasks described herein.

[0109] FIG. 12 is a system block diagram illustrating an alternative cloud computing platform 1200 for processing raw or pre-processed context data received from a wearable multimedia device, according to one embodiment. An edge server 1201 receives raw or pre-processed context data from a wearable multimedia device 1202 over a wireless communication link. The edge server 1201 provides limited local pre-processing, such as AI or camera video (CV) processing and gesture detection. In the edge server 1201, a dispatcher 1203 directs the raw or pre-processed context data to a state / context detector 1204, a first-party handler 1205, and / or a limited AI resolver 1206 for performing limited AI tasks. The state / context detector 1204 determines the location where the context data was captured, using GNSS data provided, for example, by the wearable multimedia device's 1202 GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometry). The state / context detector 1204 also uses image and audio technology and AI to analyze the images, audio and sensor data (e.g., motion sensor data, biometric data) contained in the context data to determine user activity, mood and interests.

[0110] The edge server 1201 is coupled by fiber and routers to a regional data center 1207. The regional data center 1207 performs full AI and / or CV processing of the pre-processed or raw context data. In the regional data center 1207, a dispatcher 1208 directs the raw or pre-processed context data to a state / context detector 1209, a full AI resolver 1210, a first handler 1211, and / or a second handler 1212. The state / context detector 1209 determines the location where the context data was captured, for example, using GNSS data provided by the wearable multimedia device 1202's GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometry). The state / context detector 1209 also uses image and voice recognition technology and AI to analyze the images, audio, and sensor data (e.g., motion sensor data, biometric data) included in the context data to determine user activity, mood, and interests.

[0111] FIG. 13 illustrates software components 1300 for a wearable multimedia device, according to one embodiment. For example, the software components include daemons 1301 for AI and CV, gesture recognition, messaging, media capture, server connectivity, and accessory connectivity. The software components further include libraries 1302 for graphics processing units (GPUs), machine learning (ML), camera video (CV), and network services. The software components include an operating system 1303, such as the Android Native Development Kit (NDK), which includes hardware abstraction and a Linux kernel. Other software components 1304 include components for power management, connectivity, security and encryption, and software updates.

[0112] 14A-14D illustrate the use of a projector 1115 of a wearable multimedia device to project various types of information onto a user's palm, according to one embodiment. In particular, FIG. 14A illustrates a laser projection of a number pad onto the user's palm for use in dialing phone numbers and other tasks requiring number entry. A 3D camera 1103 (depth sensor) is used to determine the location of the user's fingers on the number pad. A user can interact with the number pad, such as by dialing a phone number. FIG. 14B illustrates turn-by-turn directions projected onto the user's palm. FIG. 14C illustrates a clock projected onto the user's palm. FIG. 14D illustrates a temperature reading on the user's palm. Various one- or two-finger gestures (e.g., tap, long press, swipe, pinch / de-pinch) can be detected by the 3D camera 1103, resulting in different actions being triggered on the wearable multimedia device.

[0113] While Figures 14A-14D illustrate laser projection onto the palm of a user's hand, any projection surface can be used, including but not limited to walls, floors, ceilings, curtains, clothing, projection screens, tables / desks / countertops, appliances (e.g., ranges, washer / dryer), and devices (e.g., car engines, electronic circuit boards, cutting boards).

[0114] 15A and 15B illustrate an application of projector 1115 in which information to assist a user in checking their engine oil is projected onto a car engine, according to one embodiment. FIGS. 15A and 15B illustrate before and after images of the image. The user utters the words: "How do I check the oil?" In this example, audio is received by microphone 1106, and main camera 1102 and / or 3D camera 1103 of wearable multimedia device 1202 capture images of the engine. The images and audio are compressed and sent to edge server 1201. Edge server 1201 sends the images and audio to regional data center 1207. At regional data center 1207, the images and audio are decompressed and one or more classifiers are used to detect and label the locations of the dipstick and fuel filler cap in the images. The labels and their image coordinates are sent back to wearable multimedia device 1202. Projector 1115 projects a label onto the car's engine based on image coordinates.

[0115] FIG. 16 illustrates an application of a projector in which information is projected onto a cutting board to assist a home cook in cutting vegetables, according to one embodiment. The user speaks the phrase: "How big should I cut this into?" In this example, audio is received by microphone 1106, and main camera 1102 and / or 3D camera 1103 of wearable multimedia device 1202 capture images of the cutting board and vegetables. The images and audio are compressed and sent to edge server 1201. Edge server 1201 sends the images and audio to regional data center 1207. In regional data center 1207, the images and audio are decompressed and one or more classifiers are used to detect the type of vegetable (e.g., cauliflower), its size, and its location in the image. Based on the image information and audio, cutting instructions (e.g., obtained from a database or other data source) and image coordinates are determined and sent back to wearable multimedia device 1202. Projector 1115 uses the information and image coordinates to project a size template onto the cutting board around the vegetable.

[0116] 17 is a system block diagram of a projector architecture 1700 according to one embodiment. The projector 1115 scans pixels in two dimensions, images a 2D array of pixels, or combines imaging and scanning. A scanning projector directly leverages the narrow divergence and two-dimensional (2D) scanning of a laser beam to "paint" an image pixel by pixel. In some embodiments, separate scanners are used for the horizontal and vertical scan directions. In other embodiments, a single two-axis scanner is used. The specific beam trajectory also varies depending on the type of scanner used.

[0117] In the illustrated example, projector 1700 is a scanning picoprojector that includes a controller 1701, a battery 1118 / 1121, a power management chip (PMIC) 1120, a solid-state laser 1704, an XY scanner 1705, a driver 1706, a memory 1707, a digital-to-analog converter (DAC) 1708, and an analog-to-digital converter (ADC) 1709.

[0118] The controller 1701 provides control signals to the XY scanner 1705. The XY scanner 1705 uses movable mirrors to steer the laser beam generated by the solid-state laser 1704 in two dimensions in response to the control signals. The XY scanner 1705 includes one or more microelectromechanical systems (MEMS) micromirrors with tilt angles controllable in one or two dimensions. The driver 1706 includes power amplifiers and other electronic circuitry (e.g., filters, switches) that provide control signals (e.g., voltage or current) to the XY scanner 1705. The memory 1707 stores various data used by the projector, including laser patterns for the text and images to be projected. The DAC 1708 and ADC 1709 provide data conversion between the digital and analog domains. The PMIC 1120 manages the power and duty cycle of the solid-state laser 1704, including turning it on and off and adjusting the amount of power supplied to it. The solid-state laser 1704 may be, for example, a vertical-cavity surface-emitting laser (VCSEL).

[0119] In one embodiment, the controller 1701 uses image data from the main camera 1102 and depth data from the 3D camera 1103 to recognize and track the user's hand and / or finger position on the laser projection, such that user input is received by the wearable multimedia device 101 using the laser projection as an input interface.

[0120] In another embodiment, the projector 1115 uses a vector graphics projection display and low-power fixed MEMS micromirrors to conserve power. Because the projector 1115 includes a depth sensor, the projection area can be masked as needed to prevent projection onto fingers / hands interacting with the laser projected image. In one embodiment, the depth sensor can also track gestures to control input on another device (e.g., swiping an image on a television screen, interacting with a computer, smart speaker, etc.).

[0121] In other embodiments, liquid crystal on silicon (LCoS or LCOS), digital light processing (DLP) or liquid crystal display (LCD) digital projection technologies can be used instead of a picoprojector.

[0122] FIG. 18 illustrates adjusting laser parameters based on the amount of light reflected by a surface. To ensure that the projection is clear and legible on a wide variety of surfaces, data from the 3D camera 1103 is used to adjust one or more parameters of the projector 1115 based on the surface reflection. In one embodiment, reflections of the laser beam from the surface are used to automatically adjust the intensity of the laser beam to compensate for different refractive indices to create a projection of uniform brightness. The intensity can be adjusted, for example, by adjusting the power supplied to the solid-state laser 1115. The amount of adjustment can be calculated by the controller 1701 based on the energy level of the reflected laser beam.

[0123] In the illustrated example, a circular pattern 1800 is projected onto surface 1801, including region 1802 having a first surface reflectance and region 1803 having a second surface reflectance different from the first surface reflectance. As a result of the difference in surface reflectance in regions 1802, 1803 (e.g., due to different refractive indices), circular pattern 1800 is less bright in region 1802 than in region 1803. To generate circular pattern 1800 of uniform intensity, solid-state laser 1704 is commanded by controller 1701 (through PMIC 1120) to increase or decrease the power supplied to solid-state laser 1704 to increase the intensity of the laser beam as it scans across region 1802. The result is circular pattern 1800 of uniform brightness. If regions 1802 and 1803 have different surface geometries, one or more lenses can be used to adjust the size of the projected text or image on the surface. For example, region 1802 can be curved, while region 1803 can be flat. In this scenario, the size of the text or image in region 1802 can be adjusted to compensate for the curvature of the surface in region 1802 .

[0124] In one embodiment, laser projection is automatically or manually requested by a user's air gestures (e.g., pointing to identify an object of interest, swiping to indicate an action on data, holding up a finger to indicate a number, thumbs up or down to indicate a preference, etc.) and can be projected onto any surface or object in the environment. For example, a user can point at the thermostat in their home and temperature data is projected onto the palm of their hand or other surface. Cameras and depth sensors detect where the user is pointing, identify the object being pointed at as the thermostat, run a thermostat application on a cloud computing platform, and stream application data to a wearable multimedia device, which is displayed on the surface (e.g., the user's palm, a wall, a table). In another example, if a user is standing in front of the smart lock on their front door and all locks in their home are interlocked, controls for the smart lock are projected onto the surface of the smart lock or door for accessing that lock or other locks in their home.

[0125] In one embodiment, images captured by the camera and its large field of view (FOV) can be presented to the user in a "contact sheet" using an AI-driven virtual photographer running on a cloud computing platform. For example, various representations of the image are created with different cropping and processing using machine learning (e.g., neural networks) trained on the images / metadata created by the professional photographer. With this feature, every captured image can have multiple "looks" with multiple image processing operations on the original image, including operations informed by sensor data (e.g., depth, ambient light, accelerometer, gyro).

[0126] The described features can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations of them. The features can also be implemented in a computer program product tangibly embodied in an information carrier, for example a machine-readable storage device, for execution by a programmable processor. Method steps can be performed by a programmable processor executing a program of instructions that performs the functions of the described implementations by operating on input data and generating output.

[0127] The described features may be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform an activity or bring about a result. A computer program may be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0128] Suitable processors for executing a program of instructions include, by way of example, both general-purpose and special-purpose microprocessors, and the sole processor or one of multiple processors or cores of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may be in communication with mass storage devices for storing data files. These mass storage devices may include magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Suitable storage devices for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices, such as EPROMs, EEPROMs, and flash memory devices, magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, ASICs (application-specific integrated circuits). To provide for user interaction, the above features may be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light emitting diode) or LCD (liquid crystal display) display or monitor for displaying information to the creator, a keyboard by which the creator may provide input to the computer, and a pointing device such as a mouse or trackball.

[0129] One or more features or steps of the disclosed embodiments may be implemented using an application programming interface (API). An API may define one or more parameters passed between a calling application and other software code (e.g., an operating system, library routine, function) that provides a service, provides data, or performs an operation or calculation. An API may be implemented as one or more calls to program code that send or receive one or more parameters through a parameter list or other structure based on a calling convention defined in an API specification. A parameter may be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters may be implemented in any programming language. The programming language may define the vocabulary and calling conventions that a programmer will utilize to access functions that support the API. In some implementations, API calls may report capabilities, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, etc., of the device on which the application is running to the application.

[0130] Several example implementations have been described. Nevertheless, it will be understood that various modifications may be made. Elements of one or more example implementations may be combined, deleted, modified, or supplemented to form further example implementations. In yet another example, the logic flow depicted in the figures does not require the particular order or sequence shown to achieve desirable results. Additionally, other steps may be provided or steps may be eliminated from the described flow, and other components may be added to or deleted from the described system. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]

[0131] 100 Operating environment 101 Wearable Multimedia Devices 102 Cloud Computing Platform 103 Network 104 Application Developer 105 Third Party Platforms 106 databases 200 Data Processing System 201 Recorder 202 Video Buffer 203 Audio Buffer 204 Photo Buffer 205 Import Server 206 Datastore 207 Video Processor 208 Audio Processor 209 Photo Processor 210 Third Party Processors 211 App 1 212 App 2 213 App 1 214 App 2 215 App 1 216 App 2 217 App 300 Data Processing Pipeline 301 Import Server 302 Video Processor 304 Recorder 305 Audio Processor 306 App 307 Video Processing Apps 308 Data Merge Process 400 Data Processing Pipelines 401 Import Server 402 Third Party Processors 404 Recorder 405 Audio Processor 406 App 407 Third Party Application 408 User Callback Queue 700 Architecture 702 processor 704 Storage Devices 706 Network Interface 708 Computer-Readable Medium 710 Communication Channel 712 Operating Systems 714 Network Communication Module 716 Data Processing Instructions 718 Interface Instructions 800 Architecture 802 memory interface 804 processor 806 Peripheral Interface 810 Motion Sensor 812 Biometric Sensors 814 Depth Sensor 815 Position Processor 816 Electronic Magnetometer 817 Environmental Sensor 820 Camera Subsystem 822 Optical Sensor 824 Communication Subsystem 826 Audio Subsystem 828 Speaker 830 Microphone 840 I / O Subsystem 842 Touch Controller 844 Other Input Controllers 846 Touch Surface 848 Other Input / Control Devices 850 memory 852 Operating Systems 854 Communication Order 858 Sensor Processing Instructions 860 Recorder Instructions 900 Graphical User Interface 901 Video Pane 902 Time / Location Data 903 Objects 904 Search button 905 Categories 906a, 906b, 906c Objects 907 thumbnail images 1000 Classifier Framework 1001 API 1002a~1002n classifier 1005 Datastore 1100 Hardware Architecture 1101 System on Chip 1102 Main Camera 1103 3D Camera 1104 Capacitive Sensor 1105 Motion Sensor 1106 Microphone 1107 Memory 1108 Global Navigation Satellite System Receiver 1109 WiFi / Bluetooth chip 1110 wireless transmitter / receiver chip 1112 Radio Frequency Transceiver Chip 1113 RF Front-End Electronic Circuit 1114 LED 1115 Projector 1116 Audio Amplifier 1117 Speaker 1118 External Battery 1119 Magnetic Inductance Network 1120 Power Management Chip 1121 Internal Battery 1200 Cloud Computing Platform 1201 Edge Server 1202 Wearable Multimedia Devices 1203 Dispatcher 1204 State / Context Detector 1205 First Party Handlers 1206 Limited AI Resolver 1207 Area Data Center 1208 Dispatcher 1209 State / Context Detector 1210 Fully AI Resolver 1211 First Handler 1212 Second Handler 1300 Software Components 1301 Demon 1302 Library 1303 Operating System 1304 Other Software Components 1700 Projector, Projector Architecture 1701 Controller 1704 Solid-state laser 1705 XY scanner 1706 Driver 1707 Memory 1708 Digital-to-Analog Converter 1709 Analog-to-Digital Converter 1800 yen pattern 1801 Surface 1802, 1803 area

Claims

1. 1. A body worn device comprising: A camera and A depth sensor; a laser projection system; one or more processors; When executed by the one or more processors, it causes the one or more processors to: capturing a set of digital images using said camera; identifying an object within the set of digital images; capturing depth data using the depth sensor; identifying a gesture of a user wearing the device within the depth data; associating the object with the gesture; obtaining data associated with the object; and a memory having stored thereon instructions for performing operations including projecting a laser projection of the data onto a surface using the laser projection system.

2. The apparatus of claim 1 , wherein the laser projection includes a text label for the object.

3. The apparatus of claim 1 or 2, wherein the laser projection includes a size template for the object.

4. The apparatus of claim 1 , wherein the laser projection includes instructions for performing an action on the object.

5. 1. A body worn device comprising: A camera and A depth sensor; a laser projection system; one or more processors; When executed by the one or more processors, it causes the one or more processors to: capturing depth data using the sensor; identifying a first gesture in the depth data, the gesture being made by a user wearing the device; Associating the first gesture with a request or command; and a memory having stored thereon instructions for causing the laser projection system to perform an action including projecting a laser projection associated with the request or command onto a surface using the laser projection system.

6. The operation is acquiring a second gesture associated with the laser projection using the depth sensor; and determining a user input based on the second gesture; and and initiating one or more actions according to the user input.

7. The operation is 7. The apparatus of claim 6, further comprising masking the laser projection to prevent projecting the data onto the user's hand making the second gesture.

8. The operation is using said depth sensor or camera to acquire depth or image data indicative of the geometry, material or texture of said surface; and adjusting one or more parameters of the laser projection system based on the geometry, material, or texture of the surface.

9. The operation is using the camera to capture a reflection of the laser projection from the surface; 9. The apparatus of claim 1, further comprising automatically adjusting the intensity of the laser projection to compensate for different refractive indices so that the laser projection has uniform brightness.

10. a magnetic attachment mechanism configured to magnetically couple to a battery pack through a user's clothing, the magnetic attachment mechanism being further configured to receive inductive charging from the battery pack; The apparatus of claim 1 , further comprising:

11. capturing depth data using a depth sensor in a body worn device; using one or more processors of the device to identify a first gesture in the depth data, the first gesture being made by a user wearing the device; using the one or more processors to associate the first gesture with a request or command; projecting a laser projection associated with the request or command onto a surface using a laser projection system of the device; A method comprising:

12. using the depth sensor to capture a second gesture by the user, the second gesture associated with the laser projection; determining a user input based on the second gesture; initiating one or more actions according to the user input; The method of claim 11 further comprising:

13. The method of claim 12 , wherein the one or more actions include controlling another device.

14. masking the laser projection to prevent projecting the data onto the user's hand making the second gesture.

14. The method of any one of claims 11 to 13, further comprising:

15. using said depth sensor or camera to acquire depth or image data indicative of the geometry, material or texture of said surface; adjusting one or more parameters of the laser projection system based on the geometry, material, or texture of the surface; 15. The method of any one of claims 11 to 14, further comprising: