Cloud computing platform with wearable multimedia devices and laser projection systems
The wearable multimedia device with a cloud computing platform addresses the issue of missed spontaneous moments by automatically capturing and processing multimedia data for seamless interaction and editing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HEWLETT PACKARD DEVELOPMENT COMPANY LP
- Filing Date
- 2024-01-31
- Publication Date
- 2026-04-30
AI Technical Summary
Modern mobile devices often fail to capture important spontaneous moments due to their bulkiness and the need for user intervention, leading to missed photo or video opportunities.
A wearable multimedia device equipped with a camera, depth sensor, laser projection system, and cloud computing platform that automatically captures, processes, and projects multimedia data, allowing for minimal user interaction and enhanced data editing.
The device captures and edits multimedia data efficiently, enabling spontaneous event documentation and interaction with others without immersion, providing a seamless user experience.
Smart Images

Figure 0007853530000001 
Figure 0007853530000002 
Figure 0007853530000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Patent Application No. 16 / 904,544, filed Jun. 17, 2020, for “Wearable Multimedia Device and Cloud Computing Platform With Laser Projection System”, which is a continuation - in - part of U.S. Provisional Patent Application No. 62 / 863,222, filed Jun. 18, 2019, for “Wearable Multimedia Device and Cloud Computing Platform With Application Ecosystem” and of U.S. Patent Application No. 15 / 976,632, filed May 20, 2018, for “Wearable Multimedia Device and Cloud Computing Platform With Application Ecosystem”, each of which is hereby incorporated by reference in its entirety.
[0002] This disclosure generally relates to cloud computing and multimedia editing.
Background Art
[0003] Modern mobile devices (e.g., smartphones, tablet computers) often include built - in cameras that enable users to capture digital images or videos of spontaneous events. These digital images and videos can be stored in an online database associated with the user account to free up memory on the mobile device. Users can share their images and videos with friends and family and can also download or stream images and videos on demand using their various playback devices. These built - in cameras provide significant advantages compared to conventional digital cameras, which are often bulky and require more time to prepare for shooting.
[0004] Despite the convenience of built-in cameras on mobile devices, many important moments are missed by these devices because they occur too suddenly, or users are so caught up in the moment that they simply forget to take a picture or video. [Overview of the Initiative] [Means for solving the problem]
[0005] Disclosed are systems, methods, devices, and non-temporary computer-readable storage media for a cloud computing platform with a wearable multimedia device and an application ecosystem for processing multimedia data captured by the wearable multimedia device.
[0006] In one embodiment, the body-worn device comprises a camera, a depth sensor, a laser projection system, one or more processors, and a memory that stores instructions causing one or more processors to perform operations, when executed by the one or more processors, including capturing a set of digital images using the camera, identifying objects within the set of digital images, capturing depth data using the depth sensor, identifying the gestures of the user wearing the device within the depth data, associating the objects with the gestures, acquiring data associated with the objects, and projecting a laser projection of the data onto a surface using the laser projection system.
[0007] In one embodiment, the laser projection includes a text label for an object.
[0008] In one embodiment, the laser projection includes a size template for the object.
[0009] In one embodiment, the laser projection includes instructions for performing an action on an object.
[0010] In one embodiment, the body-worn device includes a camera, a depth sensor, a laser projection system, one or more processors, and a memory that stores instructions causing one or more processors to perform actions including: capturing depth data using the sensor; identifying a first gesture in the depth data made by a user wearing the device; associating the first gesture with a request or command; and projecting a laser projection associated with the request or command onto a surface using the laser projection system.
[0011] In one embodiment, the operation further includes using a depth sensor to acquire a second gesture associated with laser projection, determining user input based on the second gesture, and initiating one or more actions according to the user input.
[0012] In one embodiment, the operation further includes masking the laser projection to prevent it from projecting data onto the user's hand making a second gesture.
[0013] In one embodiment, the operation further includes using a depth sensor or camera to acquire depth or image data indicating the geometry, material, or texture of a surface, and adjusting one or more parameters of a laser projection system based on the geometry, material, or texture of the surface.
[0014] In one embodiment, the operation further includes using a camera to capture the reflection of the laser projection from the surface and automatically adjusting the intensity of the laser projection to compensate for different refractive indices so that the laser projection has a uniform brightness.
[0015] In one embodiment, the device includes a magnetic mounting mechanism configured to magnetically couple to a battery pack through the user's clothing, and further configured to receive inductive charging from the battery pack.
[0016] In one embodiment, the method includes the steps of: capturing depth data using a depth sensor of a body-worn device; identifying a first gesture in the depth data, made by a user wearing the device, using one or more processors of the device; associating the first gesture with a request or command using one or more processors; and projecting a laser projection associated with the request or command onto a surface using a laser projection system of the device.
[0017] In one embodiment, the method further includes the steps of: using a depth sensor to acquire a second gesture made by a user, which is associated with laser projection; determining a user input based on the second gesture; and initiating one or more actions according to the user input.
[0018] In one embodiment, one or more actions include controlling another device.
[0019] In one embodiment, the method further includes the step of masking the laser projection to prevent it from projecting data onto the user's hand making a second gesture.
[0020] In one embodiment, the method further includes the steps of using a depth sensor or camera to acquire depth or image data indicating the geometry, material, or texture of a surface, and adjusting one or more parameters of a laser projection system based on the geometry, material, or texture of the surface.
[0021] In one embodiment, the method includes the steps of: receiving contextual data from a wearable multimedia device, which includes at least one data capture device for capturing contextual data, by one or more processors of a cloud computing platform; creating a data processing pipeline in one or more applications based on one or more characteristics of the contextual data and user requests by one or more processors; processing the contextual data through the data processing pipeline by one or more processors; and sending the output of the data processing pipeline to a wearable multimedia device or another device for presenting the output by one or more processors.
[0022] In one embodiment, the system comprises one or more processors and a memory storing instructions that, when executed by one or more processors, cause one or more processors to perform operations including receiving context data from a wearable multimedia device, which includes at least one data capture device for capturing context data, by one or more processors of a cloud computing platform; creating a data processing pipeline in one or more applications based on one or more characteristics of the context data and user requests; processing the context data through the data processing pipeline; and sending the output of the data processing pipeline to a wearable multimedia device or another device for presenting the output.
[0023] In one embodiment, a non-temporary computer-readable storage medium receives contextual data from a wearable multimedia device, which includes at least one data capture device for capturing contextual data, by one or more processors of a cloud computing platform; by one or more processors, creates a data processing pipeline in one or more applications based on one or more characteristics of the contextual data and user requests; by one or more processors, processes the contextual data through the data processing pipeline; and by one or more processors, includes instructions for sending the output of the data processing pipeline to the wearable multimedia device or another device for presenting the output.
[0024] In one embodiment, the method includes the steps of: receiving depth or image data indicating surface geometry, material or texture provided by one or more sensors of the wearable multimedia device via a controller of the wearable multimedia device; adjusting one or more parameters of the projector of the wearable multimedia device based on the surface geometry, material or texture via the controller; projecting text or image data onto a surface via the projector of the wearable multimedia device; receiving depth or image data from one or more sensors indicating user interaction with the text or image data projected onto the surface via the controller; determining user input based on user interaction via the controller; and initiating one or more actions in accordance with the user input via the processor of the wearable multimedia device.
[0025] In one embodiment, a wearable multimedia device includes one or more sensors, a projector, and depth or image data indicative of a surface geometry, material, or texture, which is received from the one or more sensors of the wearable multimedia device, adjusts one or more parameters of the projector based on the surface geometry, material, or texture, projects text or image data onto the surface using the projector, receives depth or image data from the one or more sensors indicative of user interaction with the text or image data projected onto the surface, determines a user input based on the user interaction, and a controller configured to initiate one or more actions in accordance with the user input.
Advantages of the Invention
[0026] Certain embodiments disclosed herein provide one or more of the following advantages. The wearable multimedia device captures multimedia data of spontaneous moments and transactions with minimal interaction by the user. The multimedia data is automatically edited and formatted on a cloud computing platform based on user preferences and then made available to the user for playback on various user playback devices. In one embodiment, the data editing and / or processing is performed by an ecosystem of applications that are proprietary and / or provided / licensed from third-party developers. The application ecosystem provides various access points (e.g., websites, portals, APIs) that enable third-party developers to upload, verify, and update their applications. The cloud computing platform automatically constructs a custom processing pipeline for each multimedia data stream using one or more of the ecosystem applications, user preferences, and other information (e.g., data type or format, data volume and quality).
[0027] Additionally, the wearable multimedia device includes cameras and depth sensors that can detect object and user air gestures and then perform or infer various actions based on the detection, such as labeling objects in camera images or controlling other devices. In one embodiment, the wearable multimedia device does not include a display, allowing the user to continue interacting with friends, family, and colleagues without immersing in a display as is an issue for current smartphone and tablet computer users. Therefore, the wearable multimedia device takes a different technical approach from smart goggles or smart glasses for augmented reality (AR) and virtual reality (VR), for example, where the user is further removed from the real-world environment. To facilitate collaboration with others and compensate for the lack of a display, the wearable multimedia computer includes a laser projection system that projects laser projections onto any surface, including tables, walls, and even the palms of the user's hands. The laser projection can label objects, provide text or instructions related to the objects, and provide an ephemeral user interface (e.g., keyboard, numeric keypad, device controller) that allows the user to create messages, control other devices, or simply share and discuss content with others.
[0028] Details of the disclosed embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0029] [Figure 1] FIG. 1 is a block diagram of an operating environment for a cloud computing platform with an application ecosystem for processing multimedia data captured by a wearable multimedia device and a wearable multimedia device according to one embodiment. [Figure 2]This is a block diagram of a data processing system implemented by the cloud computing platform shown in Figure 1, according to one embodiment. [Figure 3] This is a block diagram of a data processing pipeline for processing a contextual data stream according to one embodiment. [Figure 4] This is a block diagram of another data processing method for processing a contextual data stream for a transport application, according to one embodiment. [Figure 5] An example of a data object used by the data processing system shown in Figure 2, according to one embodiment, is provided. [Figure 6] This is a flowchart of a data pipeline process according to one embodiment. [Figure 7] This is an architecture for a cloud computing platform according to one embodiment. [Figure 8] This is an architecture for a wearable multimedia device according to one embodiment. [Figure 9] This is a screenshot of an example graphical user interface (GUI) for a scene identification application described with respect to Figure 3, according to one embodiment. [Figure 10] One embodiment illustrates a classifier framework for classifying raw or preprocessed context data into objects and metadata that can be searched using the GUI in Figure 9. [Figure 11] This is a system block diagram illustrating a hardware architecture for a wearable multimedia device according to one embodiment. [Figure 12] This is a system block diagram illustrating a processing framework implemented on a cloud computing platform for processing raw or preprocessed context data received from a wearable multimedia device, according to one embodiment. [Figure 13] An example of a software component for a wearable multimedia device according to one embodiment is provided. [Figure 14A] One embodiment illustrates the use of a projector in a wearable multimedia device that projects various types of information onto the user's palm. [Figure 14B] One embodiment illustrates the use of a projector in a wearable multimedia device that projects various types of information onto the user's palm. [Figure 14C] One embodiment illustrates the use of a projector in a wearable multimedia device that projects various types of information onto the user's palm. [Figure 14D] One embodiment illustrates the use of a projector in a wearable multimedia device that projects various types of information onto the user's palm. [Figure 15A] One embodiment illustrates a projector application in which information is projected onto a car engine to assist a user in checking their engine oil. [Figure 15B] One embodiment illustrates a projector application in which information is projected onto a car engine to assist a user in checking their engine oil. [Figure 16] One embodiment illustrates the application of a projector in which information to assist a home cook in cutting vegetables is projected onto a cutting board. [Figure 17] This is a system block diagram of a projector architecture according to one embodiment. [Figure 18] An example of adjusting laser parameters based on different surface geometries or materials, according to one embodiment, is provided. [Modes for carrying out the invention]
[0030] The same reference symbols used in various drawings represent similar elements.
[0031] Overview A wearable multimedia device is a lightweight, compact, battery-powered device that can be attached to a user's clothing or object using a tension clasp, interlock pinback, magnet, or any other mounting mechanism. A wearable multimedia device includes a digital image capture device (e.g., 180° FOV with optical image stabilization (OIS)) that enables a user to spontaneously capture multimedia data (e.g., video, audio, depth data) of life events ("moments") and document transactions (e.g., financial transactions) with minimal user interaction or device setup. Multimedia data captured by a wireless multimedia device ("context data") is uploaded to a cloud computing platform with an application ecosystem that allows one or more applications (e.g., artificial intelligence (AI) applications) to process, edit, and format the context data into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played on the wearable multimedia device and / or any other playback device. For example, a cloud computing platform can convert video and audio data into any desired shooting style specified by the user (e.g., documentary, lifestyle, snapshot, photojournalism, sports, street).
[0032] In one embodiment, contextual data is processed by server computers of a cloud computing platform based on user preferences. For example, based on user preferences, images can be color graded, stabilized, and cropped to perfectly match the moment the user wants to recreate. User preferences can be stored in user profiles created by the user through online accounts accessible through a website or portal, or user preferences can be learned by the platform over time (e.g., using machine learning). In one embodiment, the cloud computing platform is a scalable distributed computing environment. For example, the cloud computing platform could be a distributed streaming platform (e.g., Apache Kafka®) with real-time streaming data pipelines and streaming applications that transform or react to streams of data.
[0033] In one embodiment, a user can start and stop a context data capture session on a wearable multimedia device by issuing a command with a simple touch gesture (e.g., tap or swipe) or by any other input mechanism. All or part of the wearable multimedia device can automatically power off when it detects that the device is not being worn by a user using one or more sensors (e.g., proximity sensor, optical sensor, accelerometer, gyroscope).
[0034] Contextual data can be encrypted and compressed using any desired encryption or compression technology and stored in an online database associated with the user account. Contextual data can be stored for a specified period, which can be set by the user. Users can be provided with opt-in mechanisms and other tools to manage their data and data privacy through a website, portal, or mobile application.
[0035] In one embodiment, context data includes point cloud data to provide a three-dimensional (3D) surface mapping object that can be processed using, for example, augmented reality (AR) and virtual reality (VR) applications in an application ecosystem. Point cloud data can be generated by a depth sensor (e.g., a lidar or time-of-flight (TOF)) embedded on a wearable multimedia device.
[0036] In one embodiment, a wearable multimedia device includes a Global Navigation Satellite System (GNSS) receiver (e.g., Global Positioning System (GPS)) and one or more inertial sensors (e.g., accelerometer, gyroscope) for determining the location and orientation of the user wearing the device when context data is captured. In one embodiment, one or more images in the context data can be used by a location-finding application, such as a visual odometry application, in the application ecosystem to determine the user's location and orientation.
[0037] In one embodiment, the wearable multimedia device may also include one or more environmental sensors, including but not limited to ambient light sensors, magnetometers, pressure sensors, and sound activity detectors. This sensor data can be included in contextual data to enhance content presentation with additional information that can be used to capture moments.
[0038] In one embodiment, a wearable multimedia device may include one or more biometric sensors, such as a heart rate sensor or a fingerprint scanner. This sensor data can be included in contextual data to document a transaction or indicate the user's emotional state at a given moment (for example, a high heart rate may indicate excitement or fear).
[0039] In one embodiment, the wearable multimedia device includes a headphone jack for connecting a headset or earbuds, and one or more microphones for receiving voice commands and capturing ambient audio. In an alternative embodiment, the wearable multimedia device includes, but is not limited to, Bluetooth, IEEE 802.15.4 (ZigBee®), and Near Field Communication (NFC). The NFC technology can be used in addition to, or in place of, the headphone jack to wirelessly connect to a wireless headset or earbuds and / or to any other external device (e.g., a computer, printer, projector, television, and other wearable devices).
[0040] In one embodiment, the wearable multimedia device includes a wireless transceiver and a communication protocol stack for various communication technologies, including WiFi, 3G, 4G, and 5G communication technologies. In one embodiment, the headset or earbuds also include sensors (e.g., biometric sensors, inertial sensors) that provide information about the direction the user is facing in order to provide commands via head gestures, etc. In one embodiment, the camera direction can be controlled by head gestures so that the camera field of view follows the user's field of view. In one embodiment, the wearable multimedia device can be incorporated into or attached to the user's eyeglasses.
[0041] In one embodiment, the wearable multimedia device includes a projector (e.g., a laser projector, LCoS, DLP, LCD) that allows the user to project moments onto a surface such as a wall or tabletop, or can be wired or wirelessly coupled to an external projector. In another embodiment, the wearable multimedia device includes an output port that can be connected to a projector or other output device.
[0042] In one embodiment, the wearable multimedia capture device includes a touch surface that responds to touch gestures (e.g., tap, multitap, or swipe gestures). The wearable multimedia device may include a small display for presenting information and one or more light indicators that show on / off status, power status, or any other desired status.
[0043] In one embodiment, the cloud computing platform can be driven by context-based gestures (e.g., aerial gestures) combined with voice queries, such as when a user points to an object in their environment and says, "What is that building?". The cloud computing platform uses aerial gestures to narrow the camera's viewport range and isolate the building. One or more images of the building are captured and sent to the cloud computing platform, where an image recognition application can perform image queries and store or return the results to the user. Aerial and touch gestures can also be performed on a projected ephemeral display, for example, in response to user interface elements.
[0044] In one embodiment, context data can be encrypted on the device and on a cloud computing platform so that only the user or any authorized viewer can reproduce the moment as a projection on a connected screen (e.g., a smartphone, computer, television, etc.) or surface. An example architecture for a wearable multimedia device is described with respect to Figure 8.
[0045] In addition to personal life events, wearable multimedia devices simplify the capture of financial transactions currently handled by smartphones. Capturing daily transactions (e.g., commercial transactions, microtransactions) becomes easier, faster, and smoother by using the visually aided contextual awareness provided by wearable multimedia devices. For example, when a user engages in a financial transaction (e.g., making a purchase), the wearable multimedia device will generate data that records the transaction, including the date, time, amount, digital images or videos of the parties, audio (e.g., user annotations describing the transaction), and environmental data (e.g., location data). The data can be included in a multimedia data stream sent to a cloud computing platform, where it can be stored online and / or processed by one or more financial applications (e.g., financial management, accounting, budgeting, tax return preparation, inventory, etc.).
[0046] In one embodiment, the cloud computing platform provides a graphical user interface on a website or portal that enables various third-party application developers to upload, update, and manage their applications within an application ecosystem. Some application examples may include, but are not limited to, personal live streaming (e.g., Instagram® Live, Snapchat®), elderly monitoring (e.g., to confirm that a loved one has taken their medication), recollection (e.g., showing a child's soccer game from last week), and personal guidance (e.g., an AI-enabled personal guide that knows the user's location and guides the user to take action).
[0047] In one embodiment, the wearable multimedia device includes one or more microphones and a headset. In some embodiments, the headset wire includes a microphone. In one embodiment, a digital assistant that responds to user queries, requests, and commands is implemented on the wearable multimedia device. For example, a wearable multimedia device worn by a parent captures instantaneous contextual data of a child's soccer match, specifically the "moment" when the child scores a goal. The user can request (e.g., using a voice command) that the platform create a video clip of the goal and store it in the user's user account. Without any further action from the user, the cloud computing platform identifies the correct portion of the instantaneous contextual data when the goal is scored (e.g., using facial recognition, visual or audio cues), edits the instantaneous contextual data into a video clip, and stores the video clip in a database associated with the user account.
[0048] In one embodiment, the device may include photovoltaic surface technology for sustained battery life, as well as an inductive charging network (e.g., Qi) that enables inductive charging on a charging mat and wireless over-the-air (OTA) charging.
[0049] In one embodiment, a wearable multimedia device is configured to magnetically couple or mate with a rechargeable portable battery pack. The portable battery pack includes a mating surface on which a permanent magnet (e.g., north pole) is provided, and the wearable multimedia device has a corresponding mating surface on which a permanent magnet (e.g., south pole) is provided. Any number of permanent magnets having any desired shape or size can be arranged on the mating surface in any desired pattern.
[0050] Permanent magnets hold the portable battery pack and the wearable multimedia device together in a joined configuration with a garment (e.g., the user's shirt) in between. In one embodiment, the portable battery pack and the wearable multimedia device have the same mating surface dimensions, resulting in no protruding parts when in the joined configuration. The user places the portable battery pack on the inside of their garment and the wearable multimedia device on top of the portable battery pack on the outside of their garment, resulting in the wearable multimedia device magnetically attaching to their garment by the permanent magnets attracting each other through the garment. In one embodiment, the portable battery pack has an embedded wireless transmitter used to wirelessly power the wearable multimedia device while in the joined configuration using the principle of resonant inductive coupling. In one embodiment, the wearable multimedia device includes an embedded wireless receiver used to receive power from the portable battery pack while in the joined configuration.
[0051] Example operating environment Figure 1 is a block diagram of an operating environment for a cloud computing platform with a wearable multimedia device and an application ecosystem for processing multimedia data captured by the wearable multimedia device, according to one embodiment. The operating environment 100 includes a wearable multimedia device 101, a cloud computing platform 102, a network 103, an application ("app") developer 104, and a third-party platform 105. The cloud computing platform 102 is coupled to one or more databases 106 for storing contextual data uploaded by the wearable multimedia device 101.
[0052] As described above, the wearable multimedia device 101 is a lightweight, compact, battery-powered device that can be attached to the user's clothing or object using a tension clasp, interlock pinback, magnet, or any other attachment mechanism. The wearable multimedia device 101 includes a digital image capture device (e.g., 180° FOV with OIS) that enables the user to spontaneously capture "instant" multimedia data (e.g., video, audio, depth data) and document daily transactions (e.g., financial transactions) with minimal user interaction or device setup. Contextual data captured by the wireless multimedia device 101 is uploaded to the cloud computing platform 102. The cloud computing platform 102 includes an application ecosystem that enables one or more server-side applications to process, edit, and format the contextual data into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played on the wearable multimedia device and / or other playback devices.
[0053] For example, at a child's birthday party, a parent can attach a wearable multimedia device to their clothing (or attach the device to a necklace or chain and wear it around their neck) so that the camera lens is facing their field of view. The camera includes a 180° FOV so that it can capture almost everything the user is currently seeing. The user can start recording simply by tapping the surface of the device or pressing a button. No additional setup is required. A multimedia data stream (e.g., video with audio) capturing special moments of the birthday (e.g., blowing out the candles) is recorded. This “context data” is sent in real time to the cloud computing platform 102 via a wireless network (e.g., WiFi, cellular). In one embodiment, the context data is stored on the wearable multimedia device so that it can be uploaded later. In another embodiment, the user can transfer the context data to another device (e.g., a personal computer hard drive, smartphone, tablet computer, thumb drive) and later upload the context data to the cloud computing platform 102 using an application.
[0054] In one embodiment, context data is processed by one or more applications in an application ecosystem hosted and managed by the cloud computing platform 102. These applications can access the context data through their individual application programming interfaces (APIs). The cloud computing platform 102 creates a custom distributed streaming pipeline to process the context data based on one or more of the following: data type, data volume, data quality, user preferences, templates, and / or any other information, in order to generate desired presentations based on user preferences. In one embodiment, machine learning techniques can be used to automatically select appropriate applications to include in the data processing pipeline, with or without user preferences. For example, historical user context data stored in a database (e.g., a NoSQL database) can be used to determine user preferences for data processing using any suitable machine learning technique (e.g., deep learning or convolutional neural networks).
[0055] In one embodiment, the application ecosystem may include a third-party platform 105 for processing contextual data. A secure session is established between the cloud computing platform 102 and the third-party platform 105 for sending and receiving contextual data. This design allows third-party app providers to control access to their own applications and provide updates. In another embodiment, the application runs on a server of the cloud computing platform 102, and updates are sent to the cloud computing platform 102. In the latter embodiment, an app developer 104 can upload and update applications to be included in the application ecosystem using APIs provided by the cloud computing platform 102.
[0056] Example of a data processing system Figure 2 is a block diagram of a data processing system implemented by the cloud computing platform of Figure 1, according to one embodiment. The data processing system 200 includes a recorder 201, a video buffer 202, an audio buffer 203, a photo buffer 204, an ingestion server 205, a data store 206, a video processor 207, an audio processor 208, a photo processor 209, and a third-party processor 210.
[0057] A recorder 201 (e.g., a software application) running on a wearable multimedia device records video, audio, and photographic data ("context data") captured by the camera and audio subsystems, and stores the data in buffers 202, 203, and 204, respectively. This context data is then sent to an ingestion server 205 on a cloud computing platform 102 (e.g., using wireless OTA technology). In one embodiment, the data can be sent in separate data streams, each with a unique stream identifier (streamid). A stream is a distinct data fragment that may include, as an example of attributes, location (e.g., latitude, longitude), user, audio data, video streams of various durations, and N photographs. A stream can have a duration of 1 to MAXSTREAM_LEN seconds, where in this example MAXSTREAM_LEN = 20 seconds.
[0058] The ingest server 205 ingests the streams and stores the results of processors 207-209 by creating stream records in the data store 206. In one embodiment, an audio stream is processed first and used to determine any other streams that may be needed. The ingest server 205 sends the streams to the appropriate processors 207-209 based on the streamid. For example, a video stream is sent to the video processor 207, an audio stream to the audio processor 208, and a photo stream to the photo processor 209. In one embodiment, at least a portion of the data collected from the wearable multimedia device (e.g., image data) is processed into metadata, encrypted, and sent back to the wearable multimedia device or other device so that it can be further processed by a given application.
[0059] Processors 207-209 can run proprietary or third-party applications as described above. For example, video processor 207 may be a video processing server that sends raw video data stored in video buffer 202 to a set of one or more image processing / editing applications 211, 212 based on user preferences or other information. Processor 207 sends requests to applications 211, 212 and returns results to capture server 205. In one embodiment, a third-party processor 210 can process one or more streams using its own processor and applications. In another example, audio processor 208 may be an audio processing server that sends audio data stored in audio buffer 203 to speech-to-text application 213.
[0060] Example of a scene recognition application Figure 3 is a block diagram of a data processing pipeline for processing a context data stream according to one embodiment. In this embodiment, the data processing pipeline 300 is created and configured to determine what the user is looking at based on context data captured by a wearable multimedia device worn by the user. The capture server 301 receives an audio stream (including, for example, user annotations) from the audio buffer 203 of the wearable multimedia device and sends the audio stream to the audio processor 305. The audio processor 305 sends the audio stream to an application 306 that performs speech-to-text conversion and returns the parsed text to the audio processor 305. The audio processor 305 returns the parsed text to the capture server 301.
[0061] The video processor 302 receives the parsed text from the capture server 301 and sends a request to the video processing application 307. The video processing application 307 identifies objects in the video scene and labels the objects using the parsed text. The video processing application 307 sends a response describing the scene (e.g., labeled objects) to the video processor 302. The video processor then forwards the response to the capture server 301. The capture server 301 sends the response to the data merging process 308, which merges the response with the user's location, orientation, and map data. The data merging process 308 returns the response with the scene description to the recorder 304 on the wearable multimedia device. For example, the response may include map location and a description of objects in the scene (e.g., identifying people in the scene), and may include text describing the scene as a child's birthday party. The recorder 304 associates the scene description with multimedia data stored on the wearable multimedia device (e.g., using streamid). When the user retrieves the data, it is enhanced with the scene description.
[0062] In one embodiment, the data merging process 308 may use more than just location and map data. The concept of an ontology may also be used. For example, facial features of the user's father captured in an image can be recognized by a cloud computing platform and returned as "Dad" rather than the user's name, and an address such as "555 Main Street, San Francisco, CA" can be returned as "Home." The ontology can be unique to the user and can grow and learn from user input.
[0063] Examples of transportation applications Figure 4 is a block diagram of another data processing configuration for processing a context data stream for a transportation application according to one embodiment. In this embodiment, the data processing pipeline 400 is configured to call a transportation company (e.g., Uber®, Lyft®) to get a ride home. Context data from a wearable multimedia device is received by an ingestion server 401, and an audio stream from an audio buffer 203 is sent to an audio processor 405. The audio processor 405 sends the audio stream to an application 406, which converts the speech to text. The parsed text is returned to the audio processor 405, which returns the parsed text to the ingestion server 401 (e.g., the user's voice request for transportation). The processed text is sent to a third-party processor 402. The third-party processor 402 sends the user location and token to a third-party application 407 (e.g., an Uber® or Lyft® application). In one embodiment, the token is an API and authorization token used to mediate the request on behalf of the user. Application 407 returns a response data structure to third-party processor 402, which is then forwarded to capture server 401. Capture server 401 checks the boarding arrival status (e.g., ETA) in the response data structure and sets up a callback to the user in user callback queue 408. Capture server 401 returns a response with a vehicle description to recorder 404, which can be broadcast to the user by a digital assistant through a loudspeaker on a wearable multimedia device, or through the user's headphones or earbuds via a wired or wireless connection.
[0064] Figure 5 illustrates a data object used by the data processing system of Figure 2 according to one embodiment. The data object is part of the software component infrastructure instantiated on a cloud computing platform. The "Streams" object includes data streamid, deviceid, start, end, lat, lon, attributes, and entities. "streamid" identifies the stream (e.g., video, audio, photo), "deviceid" identifies the wearable multimedia device (e.g., mobile device ID), "start" is the start time of the context data stream, "end" is the end time of the context data stream, "lat" is the latitude of the wearable multimedia device, "lon" is the longitude of the wearable multimedia device, "attributes" include, for example, birthday, facial feature points, skin tone, audio characteristics, address, phone number, etc., and "entities" constitute the ontology. For example, the name "John Do" is mapped to "Dad" or "Brother (or Younger Brother)" depending on the user.
[0065] The "Users" object contains the data userid, deviceid, email, fname, and lname. userid is a unique identifier for the user, deviceid is a unique identifier for the wearable device, email is the user's registered email address, fname is the user's first name, and lname is the user's last name. The "Userdevices" object contains the data userid and deviceid. The "Devices" object contains the data deviceid, started, state, modified, and created. In one embodiment, deviceid is a unique identifier for the device (e.g., separate from the MAC address). started is when the device was first started. state is on / off / sleep. modified is the last modification date and reflects the last state change or operating system (OS) change. created is the first time the device was turned on.
[0066] The "ProcessingResults" object contains the data streamid, ai, result, callback, duration, and accuracy. In one embodiment, streamid is each user stream as a universally unique identifier (UUID). For example, a stream started from 8:00 AM to 10:00 AM would have id:15h158dhb4, and a stream started from 10:15 AM to 10:18 AM would have the UUID contacted for this stream. ai is the identifier for the platform application contacted for this stream. result is the data sent from the platform application. callback is the callback used (the callback is tracked in case the platform needs to replay the request as the version may change). accuracy is a score of how accurate the result set is. In one embodiment, processing results can be used for multiple tasks, such as 1) notifying a merge server of the results of the entire set, 2) determining the fastest AI to improve the user experience, and 3) determining the most accurate AI. Depending on the use case, you can prioritize speed over accuracy, or vice versa.
[0067] An "Entities" object contains the data entityID, userID, entityName, entityType, and entityAttribute. entityID is the UUID for an entity, and an entity may have multiple fields where entityID refers to that single entity. For example, if "Barack Obama" has entityID 144, this could link to POTUS44 or "Barack Hussein Obama" or "President Obama" in related tables. userID identifies the user who created the entity record. entityName is the name that userID calls the entity. For example, for entityID 144, Malia Obama's entityName could be "Dad" or "Papa". entityType is a person, place, or thing. entityAttribute is an array of attributes about that entity that are specific to userID's understanding of the entity. This allows, for example, when Malia makes a voice query: "Can you see Dad?", the cloud computing platform translates the query to Barack Hussein Obama and maps the entities together so that it can be used when mediating the request to a third party or looking up information in the system.
[0068] Process example Figure 6 is a flowchart of a data pipeline process according to one embodiment. Process 600 can be implemented using the wearable multimedia device 101 and cloud computing platform 102 described in relation to Figures 1 to 5.
[0069] Process 600 can begin by receiving context data from a wearable multimedia device (601). For example, the context data may include video, audio, and still images captured by the camera and audio subsystems of the wearable multimedia device.
[0070] Process 600 may be followed by the application creating (e.g., instantiating) a data processing pipeline based on contextual data and user requirements / preferences (602). For example, based on user requirements or preferences, and also based on data types (e.g., audio, video, photos), one or more applications may be logically connected to form a data processing pipeline that processes contextual data into presentations to be played on a wearable multimedia device or another device.
[0071] Process 600 can be followed by processing contextual data in a data processing pipeline (603). For example, audio from user annotations during a moment or transaction can be converted to text, which can then be used to label objects within a video clip.
[0072] Process 600 may be followed by sending the output of the data processing pipeline to a wearable multimedia device and / or other playback device (604).
[0073] Cloud Computing Platform Architecture Examples Figure 7 shows an example architecture 700 for a cloud computing platform 102 described with respect to Figures 1 to 6 and 9, according to one embodiment. Other architectures are possible, including architectures with more or fewer components. In some implementation examples, architecture 700 includes one or more processors 702 (e.g., dual-core Intel® Xeon® processors), one or more network interfaces 706, one or more storage devices 704 (e.g., hard disks, optical disks, flash memory), and one or more computer-readable media 708 (e.g., hard disks, optical disks, flash memory, etc.). These components can communicate and exchange data through one or more communication channels 710 (e.g., buses), which can leverage various hardware and software to facilitate the transfer of data and control signals between components.
[0074] The term “computer-readable medium” refers to any medium involved in providing instructions to the processor 702 for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media include, but not limited to, coaxial cables, copper wires, and optical fibers.
[0075] The computer-readable medium 708 may further include an operating system 712 (e.g., Mac OS® Server, Windows® NT Server, Linux Server), a network communication module 714, interface instructions 718, and data processing instructions 716.
[0076] The operating system 712 can be multi-user, multi-processing, multi-tasking, multi-threading, real-time, etc. The operating system 712 performs basic tasks including, but not limited to, recognizing inputs from devices 702, 704, 706, and 708 and providing outputs to them, tracking and managing files and directories on computer-readable media 708 (e.g., memory or storage devices), controlling peripheral devices, and managing traffic on one or more communication channels 710. The network communication module 714 includes various components for establishing and maintaining network connectivity (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.), as well as various components for creating a distributed streaming platform using, for example, Apache Kafka®. The data processing instructions 716 include server-side or backend software for implementing server-side operations, as described with respect to Figures 1 to 6. Interface instruction 718 includes software for implementing a web server and / or portal for sending and receiving data to and from a wearable multimedia device 101, a third-party application developer 104, and a third-party platform 105, as described with respect to Figure 1.
[0077] Architecture 700 can be included in any computer device, including one or more server computers in a local or distributed network, each having one or more processing cores. Architecture 700 can be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device with one or more processors. The software may include multiple software components or may be a single body of code.
[0078] Example of a wearable multimedia device architecture Figure 8 is a block diagram of an example architecture 800 for a wearable multimedia device that implements the features and processes described in Figures 1-6 and 9. Architecture 800 may include a memory interface 802, a data processor, an image processor or central processing unit 804, and a peripheral interface 806. The memory interface 802, the processor 804, or the peripheral interface 806 may be separate components or integrated into one or more integrated circuits. One or more communication buses or signal lines may connect the various components.
[0079] Sensors, devices, and subsystems can be coupled to the peripheral interface 806 to facilitate multiple functions. For example, motion sensors 810, biometric sensors 812, and depth sensors 814 can be coupled to the peripheral interface 806 to facilitate motion, orientation, biometric, and depth detection functions. In some implementations, motion sensors 810 (e.g., accelerometers, rate gyroscopes) may be used to detect the movement and orientation of a wearable multimedia device.
[0080] Other sensors, such as environmental sensors (e.g., temperature sensors, barometers, ambient light sensors), can also be connected to the peripheral interface 806 to facilitate environmental sensing functions. For example, biometric sensors can detect fingerprints, facial recognition, heart rate, and other fitness parameters. In one embodiment, a tactile motor (not shown) can be coupled to the peripheral interface to provide vibration patterns as tactile feedback to the user.
[0081] A position processor 815 (e.g., a GNSS receiver chip) may be connected to the peripheral interface 806 to provide georeferencing. An electronic magnetometer 816 (e.g., an integrated circuit chip) may also be connected to the peripheral interface 806 to provide data that can be used to determine the direction of magnetic north. Thus, the electronic magnetometer 816 may be used by electronic compass applications.
[0082] A camera subsystem 820 and an optical sensor 822, such as a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) optical sensor, may be utilized to facilitate camera functions such as recording photographs and video clips. In one embodiment, the camera has a 180° FOV and OIS. The depth sensor may include an infrared emitter that projects dots onto an object / subject in a known pattern. The dots are then captured by a dedicated infrared camera and analyzed to determine depth data. In one embodiment, a time-of-flight (TOF) camera may be used to determine the distance based on the known speed of light and measure the time of flight of the optical signal between the camera and the object / subject for each point in the image.
[0083] Communication functions may be facilitated through one or more communication subsystems 824. The communication subsystem 824 may include one or more wireless communication subsystems. The wireless communication subsystem 824 may include radio frequency receivers and transmitters and / or optical (e.g., infrared) receivers and transmitters. The wired communication system may include port devices, such as a Universal Serial Bus (USB) port or some other wired port connection, which can be used to establish wired connections to other computing devices, such as other communication devices, network access devices, personal computers, printers, display screens, or other processing devices capable of receiving or transmitting data (e.g., projectors).
[0084] The specific design and implementation of the communication subsystem 824 may depend on the communication network or medium on which the device is intended to operate. For example, the device may include a wireless communication subsystem designed to operate over Mobile Communications Global System (GSM) networks, GPRS networks, Enhanced Data GSM Environment (EDGE) networks, IEEE 802.xx communication networks (e.g., WiFi, WiMAX, ZigBee®), 3G, 4G, 4G LTE, Code Division Multiple Access (CDMA) networks, Near Field Communication (NFC), Wi-Fi Direct, and Bluetooth® networks. The wireless communication subsystem 824 may include a hosting protocol so that the device can be configured as a base station for other wireless devices. As another example, the communication subsystem may enable the device to synchronize with a host device using one or more protocols or communication technologies, such as TCP / IP, HTTP, UDP, ICMP, POP, FTP, IMAP, DCOM, DDE, SOAP, HTTP Live Streaming, MPEG Dash, and any other known communication protocols or technologies.
[0085] The audio subsystem 826 can be coupled to the speaker 828 and one or more microphones 830 to facilitate voice-enabled functions such as speech recognition, voice duplication, digital recording, telephone functions, and beamforming.
[0086] The I / O subsystem 840 may include a touch controller 842 and / or another input controller 844. The touch controller 842 may be coupled to the touch surface 846. The touch surface 846 and the touch controller 842 may detect their contact and movement or interruption using any of the following: several touch sensitivity technologies, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more contact points with the touch surface 846. In one implementation example, the touch surface 846 may display virtual or soft buttons, which may be used as user input / output devices.
[0087] Other input controllers 844 may be coupled to other input / control devices 848, such as one or more buttons, rocker switches, thumbwheels, infrared ports, USB ports, and / or pointer devices such as styluses. One or more buttons (not shown) may include up / down buttons for adjusting the volume of speaker 828 and / or microphone 830.
[0088] In some implementations, device 800 plays user-recorded audio and / or video files, such as MP3, AAC, and MPEG video files. In some implementations, device 800 may include MP3 player functionality and may include pin connectors or other ports for tethering to other devices. Other input / output and control devices may also be used. In one embodiment, device 800 may include an audio processing unit for streaming audio to an accessory device via a direct or indirect communication link.
[0089] The memory interface 802 may be coupled to memory 850. Memory 850 may include high-speed random-access memory or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, or flash memory (e.g., NAND, NOR). Memory 850 may store an operating system 852, such as an embedded operating system like Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks. Operating system 852 may include instructions for handling basic system services and performing hardware-dependent tasks. In some implementations, operating system 852 may include a kernel (e.g., a UNIX kernel).
[0090] Memory 850 may also store communication instructions 854 that facilitate communication with one or more additional devices, one or more computers or servers, including peer-to-peer communication with wireless accessory devices, as described with respect to Figures 1 to 6. Communication instructions 854 may also be used to select an operating mode or communication medium for use by the device based on the device's geographical location.
[0091] Memory 850 may include sensor processing instructions 858 to facilitate sensor-related processing and functions, and recorder instructions 860 to facilitate recording functions, as described with respect to Figures 1 to 6. Other instructions may include GNSS / navigation instructions to facilitate GNSS and navigation-related processes, camera instructions to facilitate camera-related processes, and user interface instructions to facilitate user interface processing, including touch models for interpreting touch input.
[0092] Each of the identified instructions and applications described above may correspond to a set of instructions for performing one or more of the functions described above. These instructions do not need to be implemented as separate software programs, procedures, or modules. Memory 850 may contain additional or fewer instructions. Furthermore, various functions of the device may be implemented in hardware and / or software, including one or more signal processing and / or application-specific integrated circuits (ASICs).
[0093] Examples of graphical user interfaces Figure 9 is a screenshot of an example of a graphical user interface (GUI) 900 for use with the scene recognition application described in relation to Figure 3, according to one embodiment. The GUI 900 includes a video pane 901, time / location data 902, objects 903, 906a, 906b, 906c, a search button 904, a menu of categories 905, and thumbnail images 907. The GUI 900 can be presented on a user device (e.g., a smartphone, tablet computer, wearable device, desktop computer, notebook computer) for example, through a client application or through a web page provided by a web server of a cloud computing platform 102. In this example, the user captured a digital image of a young man in the video pane 901 standing in Orchard Street, New York, New York, at 12:45 PM on October 18, 2018, as indicated by the time / location data 902.
[0094] In one embodiment, the image is processed through an object detection framework implemented on a cloud computing platform 102, such as a Viola-Jones object detection network. For example, a model or algorithm is used to generate a region of interest or region proposal containing a set of bounding boxes that extend across the entire digital image. Visual features are extracted for each bounding box and evaluated based on these features to determine whether and which objects exist in the region proposal. Overlapping boxes are merged into a single bounding box (e.g., using non-maximal suppression). In one embodiment, overlapping boxes are also used to organize objects into categories in big data storage. For example, object 903 (young man) is considered a parent object, and objects 906a-906c (clothing he is wearing) are considered child objects of object 903 (shoes, shirt, pants) due to overlapping bounding boxes. Thus, as a result of searching for "person" using a search engine, all objects labeled "person" and their child objects, if any, are included in the search results.
[0095] In some embodiments, composite polygons are used rather than bounding boxes to identify objects within an image. These composite polygons are used, for example, to determine a highlight / hotspot area in the image that the user is pointing to. Only the composite polysegmentation portion (rather than the entire image) is sent to the cloud computing platform, resulting in improved privacy, security, and speed.
[0096] Other examples of object detection frameworks that can be implemented by the cloud computing platform 102 to detect and label objects in digital images include, but are not limited to, regional convolutional neural networks (R-CNNs), Fast R-CNN, and Faster R-CNN.
[0097] In this example, the objects identified within the digital image include people, cars, buildings, roads, windows, doors, stairs, signs, and text. The identified objects are organized and presented as categories for the user to search. The user selects the category "People" using the cursor or (if using a touch-sensitive screen) their finger. By selecting the category "People," object 903 (i.e., the young man in the image) is separated from the rest of the objects in the digital image, and the subset of objects 906a-906c is displayed as thumbnail image 907 along with their respective metadata. Object 906a is labeled "orange, shirt, buttons, short sleeves," object 906b is labeled "blue, jeans, ripped, denim, pocket, phone," and object 906c is labeled "blue, Nike, shoes, left, Air Max, red socks, white Swoosh, logo."
[0098] When the search button 904 is pressed, a new search is initiated based on the category selected by the user and a specific image within the video pane 901. The search results include thumbnail images 907. Similarly, if the user selects the category "cars" and then presses the search button 904, a new set of thumbnails 907 will be displayed, showing all the cars captured in the images, along with their respective metadata.
[0099] Figure 10 illustrates a classifier framework 1000 for classifying raw or preprocessed context data into objects and metadata that can be searched using the GUI 900 of Figure 9, according to one embodiment. The framework 1000 includes an API 1001, classifiers 1002a-1002n, and a data store 1005. Raw or preprocessed context data captured on a wearable multimedia device is uploaded via API 1001. The context data is passed through classifiers 1002a-1002n (e.g., neural networks). In one embodiment, classifiers 1002a-1002n are trained using context data crowdsourced from a large number of wearable multimedia devices. The output of classifiers 1002a-1002n is objects and metadata (e.g., labels) stored in the data store 1005. A search index is generated for the objects / metadata in the data store 1005, which can be used by a search engine to find objects / metadata that satisfy a search query entered using the GUI 900. Various types of search indexes can be used, including but not limited to tree indexes, suffix tree indexes, inverted indexes, citation index analysis, and n-gram indexes.
[0100] Classifiers 1002a-1002n are selected and added to the dynamic data processing pipeline based on one or more of the following: data type, data volume, data quality, user preferences, user-initiated or application-initiated search queries, voice commands, application requirements, templates, and / or any other information, in order to generate the desired presentation. Any known classifier can be used, including neural networks, support vector machines (SVMs), random forests, boosted decision trees, and any combination of these individual classifiers using voting, stacking, and grading techniques. In one embodiment, a portion of the classifiers is user-specific, i.e., the classifiers are trained solely on contextual data from a particular user device. Such classifiers can be trained to detect and label people and objects that are personal to the user. For example, one classifier could be used for face detection to detect faces of individuals known to the user (e.g., family, friends) in images labeled, for example, by user input.
[0101] For example, a user might utter multiple phrases such as, "Create a movie from my videos that includes my mom and dad in New Orleans," "Add jazz music as a soundtrack," "Send me a drink recipe to make a 'Hurricane' cocktail," and "Send me directions to the nearest liquor store." The cloud computing platform 102 analyzes the voice phrases, uses the words, and assembles a personalized processing pipeline to perform the requested tasks, including adding a classifier to detect the faces of the user's parents.
[0102] In one embodiment, AI is used to determine how a user interacts with the cloud computing platform during a message session. For example, if the user sends the message "Bob, have you seen 'Toy Story 4'?", the cloud computing platform determines who Bob is and parses "Bob" from the string sent to the message relay server on the cloud computing platform. Similarly, if the message says "Bob, look at this", the platform device sends the image along with the message in one step without needing to attach the image as a separate transaction. The image can be visually confirmed by the user before being sent to Bob using the projector 1115 and any desired surface. The platform also maintains a persistent and personal communication channel with Bob for a period of time, eliminating the need to precede each communication with the name "Bob" during the message session.
[0103] Context Data Broker Service In one embodiment, a context data broker service is provided through a cloud computing platform 102. This service allows users to sell their private raw or processed context data to entities of their choice. Platform 102 hosts the context data broker service and provides the security protocols necessary to protect the privacy of user context data. Platform 102 also facilitates transactions and the exchange or settlement of money between entities and users.
[0104] Raw and preprocessed context data can be stored using big data storage. Big data storage supports storage and I / O operations on storage with a large number of data files and objects. In one embodiment, big data storage includes an architecture consisting of a redundant and scalable supply of infrastructure based on direct-attached storage (DAS) pools, scale-out or cluster network-attached storage (NAS), or object storage formats. The storage infrastructure is connected to compute server nodes that enable rapid processing and retrieval of large amounts of data. In one embodiment, the big data storage architecture includes native support for big data analytics solutions such as Hadoop®, Cassandra®, and NoSQL®.
[0105] In one embodiment, entities interested in purchasing raw or processed contextual data subscribe to the contextual data broker service through a registration GUI or webpage on the cloud computing platform 102. Once registered, entities (e.g., companies, advertising agencies) are enabled to transact directly or indirectly with users through one or more GUIs tailored to facilitate data brokerage. In one embodiment, platform 102 can match an entity's request for a particular type of contextual data with a user who can provide such data. For example, a clothing company might be interested in all images in which its clothing or a competitor's logo is detected. The clothing company can then use the contextual data to better identify its customer demographics. In another example, a news organization or political movement might be interested in video footage of newsworthy events to use in a cover story or feature article. Various companies might be interested in a user's search or purchase history for improved ad targeting or other marketing projects. Various entities might be interested in purchasing contextual data to use as training data for other object detectors, such as object detectors for autonomous vehicles.
[0106] In one embodiment, the user's raw or processing context data is made available in a secure format to protect the user's privacy. Both users and entities may have their own online accounts for depositing and withdrawing money arising from brokered transactions. In one embodiment, the data broker service collects transaction fees based on a pricing model. Fees may also be collected through conventional online advertising (e.g., click-throughs of banner ads).
[0107] In one embodiment, an individual can create their own metadata using a wearable multimedia device. For example, a renowned chef may wear the wearable multimedia device 101 while preparing a meal. Objects in the image are labeled using metadata provided by the chef. Users can obtain access to the metadata from a broker service. When a user attempts to prepare a dish while wearing their own wearable multimedia device 101, objects are detected, and metadata provided by the chef to help the user recreate something (e.g., timing, quantity, order, scale), including dimensions, cooking time, and additional hints, is projected onto the user's work surface (e.g., cutting board, countertop, range, oven, etc.). For example, if a piece of meat is detected on a cutting board, text instructing the user to cut the meat perpendicular to the grain is projected onto the cutting board by projector 1115, and further, a dimension guide is projected onto the meat surface to guide the user when slicing to a uniform thickness according to the chef's metadata. The laser projection guide can also be used to cut vegetables to a uniform thickness (e.g., julienne, mince). Users can upload their metadata to the cloud service platform, create their own channels, and earn revenue through subscriptions and advertising, similar to the YouTube® platform.
[0108] Figure 11 is a system block diagram illustrating a hardware architecture 1100 for a wearable multimedia device according to one embodiment. The architecture 1100 includes a system-on-a-chip (SoC) 1101 (e.g., Qualcomm Includes a Snapdragon® chip, a main camera 1102, a 3D camera 1103, a capacitive sensor 1104, a motion sensor 1105 (e.g., accelerometer, gyroscope, magnetometer), a microphone 1106, memory 1107, a global navigation satellite system receiver (e.g., GPS receiver) 1108, a WiFi / Bluetooth chip 1109, a wireless transceiver chip 1110 (e.g., 4G, 5G), a radio frequency (RF) transceiver chip 1112, an RF front-end electronic circuit (RFFE) 1113, an LED 1114, a projector 1115 (e.g., laser projection, pico projector, LCoS, DLP, LCD), an audio amplifier 1116, a speaker 1117, an external battery 1118 (e.g., battery pack), a magnetic inductance network 1119, a power management chip (PMIC) 1120, and an internal battery 1121. All of these components work together to facilitate the various tasks described herein.
[0109] Figure 12 is a system block diagram illustrating an alternative cloud computing platform 1200 for processing raw or preprocessed context data received from a wearable multimedia device, according to one embodiment. An edge server 1201 receives raw or preprocessed context data from a wearable multimedia device 1202 via a wireless communication link. The edge server 1201 provides limited local preprocessing, such as AI or camera video (CV) processing and gesture detection. In the edge server 1201, a dispatcher 1203 directs the raw or preprocessed context data to a state / context detector 1204, a first-party handler 1205, and / or a limited AI resolver 1206 for performing limited AI tasks. The state / context detector 1204 determines the location where the context data was captured, using GNSS data provided by, for example, a GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometry) of the wearable multimedia device 1202. The state / context detector 1204 uses image and audio technologies as well as AI to analyze image, audio, and sensor data (e.g., motion sensor data, biometric data) contained in the context data to determine user activity, mood, and interests.
[0110] The edge server 1201 is connected to the regional data center 1207 by fiber and routers. The regional data center 1207 performs full AI and / or CV processing on the preprocessed or raw context data. In the regional data center 1207, the dispatcher 1208 directs the raw or preprocessed context data to the state / context detector 1209, the full AI resolver 1210, the first handler 1211, and / or the second handler 1212. The state / context detector 1209 determines the location where the context data was captured, using GNSS data provided by, for example, a GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometry) of a wearable multimedia device 1202. The state / context detector 1209 also uses image and speech recognition technology as well as AI to analyze the images, audio, and sensor data (e.g., motion sensor data, biometric data) contained in the context data to determine user activity, mood, and interests.
[0111] Figure 13 illustrates a software component 1300 for a wearable multimedia device according to one embodiment. For example, the software component includes a daemon 1301 for AI and CV, gesture recognition, messaging, media capture, server connectivity, and accessory connectivity. The software component further includes a library 1302 for graphics processing units (GPU), machine learning (ML), camera video (CV), and network services. The software component includes an operating system 1303, such as an Android® Native Development Kit (NDK), which includes hardware abstractions and a Linux kernel. Other software components 1304 include components for power management, connectivity, security + encryption, and software updates.
[0112] Figures 14A to 14D illustrate the use of the projector 1115 in a wearable multimedia device that projects various types of information onto the user's palm, according to one embodiment. In particular, Figure 14A illustrates the laser projection of a numeric keypad onto the user's palm for use when dialing a phone number and other tasks requiring number input. A 3D camera 1103 (depth sensor) is used to determine the position of the user's fingers on the numeric keypad. The user can interact with the numeric keypad, for example, by dialing a phone number. Figure 14B illustrates turn-by-turn directions projected onto the user's palm. Figure 14C illustrates a clock projected onto the user's palm. Figure 14D illustrates the temperature reading on the user's palm. Different actions are triggered on the wearable multimedia device as a result of the 3D camera 1103 being able to detect various one- or two-finger gestures (e.g., tap, long press, swipe, pinch / depinch).
[0113] Figures 14A to 14D illustrate laser projection onto the user's palm, but any projection surface can be used, including but not limited to walls, floors, ceilings, curtains, clothing, projection screens, tables / desks / countertops, electrical appliances (e.g., ranges, washing machines / dryers), and devices (e.g., car engines, electronic circuit boards, cutting boards).
[0114] Figures 15A and 15B illustrate an application of projector 1115 in which information assisting a user in checking their engine oil is projected onto a car engine, according to one embodiment. Figures 15A and 15B show images before and after the image. The user says the words: "How do I check the oil?" In this example, the voice is received by microphone 1106, and the main camera 1102 and / or 3D camera 1103 of the wearable multimedia device 1202 capture an image of the engine. The image and audio are compressed and sent to edge server 1201. Edge server 1201 sends the image and audio to regional data center 1207. In regional data center 1207, the image and audio are decompressed, and one or more classifiers are used to detect and label the positions of the oil test stick and oil filler cap in the image. The labels and their image coordinates are sent back to wearable multimedia device 1202. The projector 1115 projects a label onto the car's engine based on image coordinates.
[0115] Figure 16 illustrates a projector application according to one embodiment, in which information is projected onto a cutting board to assist a home cook in cutting vegetables. The user speaks aloud: "How big should I cut this?" In this example, the voice is received by a microphone 1106, and the main camera 1102 and / or 3D camera 1103 of a wearable multimedia device 1202 capture images of the cutting board and vegetables. The images and audio are compressed and sent to an edge server 1201. The edge server 1201 sends the images and audio to a regional data center 1207. In the regional data center 1207, the images and audio are decompressed, and one or more classifiers are used to detect the type of vegetable in the image (e.g., cauliflower), its size, and its location. Based on the image information and audio, a cut command (e.g., obtained from a database or other data source) and image coordinates are determined and sent back to the wearable multimedia device 1202. Projector 1115 uses information and image coordinates to project a size template onto the cutting board, surrounding the vegetables.
[0116] Figure 17 is a system block diagram of a projector architecture 1700 according to one embodiment. The projector 1115 scans pixels in two dimensions, images a 2D array of pixels, or combines imaging and scanning. The scanning projector "paints" the image pixel by pixel by directly utilizing the narrow divergence of the laser beam and two-dimensional (2D) scanning. In some embodiments, separate scanners are used for the horizontal and vertical scanning directions. In other embodiments, a single two-axis scanner is used. The specific beam trajectory also varies depending on the type of scanner used.
[0117] In the illustrated example, the projector 1700 is a scanning pico projector that includes a controller 1701, a battery 1118 / 1121, a power management chip (PMIC) 1120, a solid-state laser 1704, an XY scanner 1705, a driver 1706, a memory 1707, a digital-to-analog converter (DAC) 1708, and an analog-to-digital converter (ADC) 1709.
[0118] The controller 1701 provides control signals to the XY scanner 1705. The XY scanner 1705 uses movable mirrors to steer the laser beam generated by the solid-state laser 1704 in two dimensions in response to the control signals. The XY scanner 1705 includes one or more microelectromechanical (MEMS) micromirrors with tilt angles controllable in one or two dimensions. The driver 1706 includes a power amplifier and other electronic circuits (e.g., filters, switches) that provide control signals (e.g., voltage or current) to the XY scanner 1705. The memory 1707 stores various data used by the projector, including laser patterns for text and images to be projected. The DAC 1708 and ADC 1709 provide data conversion between the digital and analog domains. The PMIC 1120 manages the power and duty cycle of the solid-state laser 1704, including turning the solid-state laser 1704 on and adjusting the amount of power supplied to the solid-state laser 1704. The solid-state laser 1704 is, for example, a vertical-cavity surface-emitting laser (VCSEL).
[0119] In one embodiment, the controller 1701 uses image data from the main camera 1102 and depth data from the 3D camera 1103 to recognize and track the position of the user's hand and / or fingers on the laser projection, and as a result, user input is received by the wearable multimedia device 101 using the laser projection as an input interface.
[0120] In another embodiment, the projector 1115 uses a vector graphics projection display and a low-power fixed MEMS micromirror to conserve power. Since the projector 1115 includes a depth sensor, the projection range can be masked as needed to prevent projection onto fingers / hands interacting with the laser projection image. In one embodiment, the depth sensor can also track gestures to control input on another device (e.g., swiping images on a TV screen, interacting with a computer, smart speaker, etc.).
[0121] In other embodiments, instead of a pico projector, liquid crystal on silicon (LCoS or LCOS), digital light processing (DLP), or liquid crystal display (LCD) digital projection technology can be used.
[0122] Figure 18 illustrates the adjustment of laser parameters based on the amount of light reflected by a surface. To ensure that the projection is clear and easy to read on a wide variety of surfaces, data from the 3D camera 1103 is used to adjust one or more parameters of the projector 1115 based on surface reflection. In one embodiment, the reflection of the laser beam from the surface is used to automatically adjust the intensity of the laser beam to compensate for different refractive indices, thereby creating a projection of uniform brightness. The intensity can be adjusted, for example, by adjusting the power supplied to the solid-state laser 1115. The amount of adjustment can be calculated by the controller 1701 based on the energy levels of the reflected laser beam.
[0123] In the illustrated example, a circular pattern 1800 is projected onto a surface 1801, which includes a region 1802 having a first surface reflection and a region 1803 having a second surface reflection different from the first. As a result of the difference in surface reflections in regions 1802 and 1803 (e.g., due to different refractive indices), the circular pattern 1800 is no brighter in region 1802 than in region 1803. To generate a circular pattern 1800 of uniform intensity, the solid-state laser 1704 is instructed by the controller 1701 (through the PMIC 1120) to increase or decrease the power supplied to the solid-state laser 1704, thereby increasing the intensity of the laser beam when scanning region 1802. The result is a circular pattern 1800 of uniform brightness. If the surface geometries of regions 1802 and 1803 are different, one or more lenses can be used to adjust the size of the projected text or image on the surface. For example, region 1802 can be curved while region 1803 can be flat. In this scenario, the size of the text or image in region 1802 can be adjusted to compensate for the curvature of the surface in region 1802.
[0124] In one embodiment, laser projection can be automatically or manually requested by a user through aerial gestures (e.g., pointing to identify an object of interest, swiping to indicate action on data, holding up a finger to indicate a number, giving or giving a thumbs-up to indicate preference, etc.) and can be projected onto any surface or object in the environment. For example, a user could point to a thermostat in their home, and temperature data would be projected onto their palm or other surface. A camera and depth sensor would detect where the user is pointing, identify the object being pointed to as a thermostat, run a thermostat application on a cloud computing platform, and stream the application data to a wearable multimedia device, which would then be displayed on a surface (e.g., the user's palm, a wall, or a table). In another example, if a user is standing in front of a smart lock on the front door of their home, and all the locks in their home are interconnected, a control mechanism for the smart lock would be projected onto the smart lock or the door surface to access that lock or any other locks in their home.
[0125] In one embodiment, images captured by a camera and its large field of view (FOV) can be presented to the user in a "contact sheet" using an AI-driven virtual photographer running on a cloud computing platform. For example, different cropping and processing are used to create various presentations of the image using machine learning (e.g., a neural network) trained on images / metadata created by professional photographers. This feature allows any captured image to have multiple "looks" involving multiple image processing operations on the original image, including operations notified by sensor data (e.g., depth, ambient light, accelerometer, gyroscope).
[0126] The described features can be implemented in digital electronic networks, or in computer hardware, firmware, software, or a combination thereof. The above features can also be implemented in computer program products tangibly embodied in information carriers for execution by a programmable processor, for example, in machine-readable memory devices. The method steps can be performed by a programmable processor executing a program of instructions that perform the functions of the described implementation example by acting on input data to produce an output.
[0127] The described features can be advantageously implemented in one or more computer programs executable on a programmable system including a data storage system and at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to at least one input device and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform an activity or produce a result. A computer program can be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0128] Appropriate processors for executing instruction programs include, for example, both general-purpose and dedicated microprocessors, as well as a single processor or a multiprocessor or core of any type of computer. Generally, a processor will receive instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may communicate with mass storage devices for storing data files. These mass storage devices may include magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Appropriate storage devices for tangibly embodying computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and all forms of non-volatile memory, including CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into ASICs (Application-Specific Integrated Circuits). To provide interaction with the user, the above features may be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light-emitting diode), or LCD (liquid crystal display) display or monitor for displaying information to the creator, a keyboard that allows the creator to provide input to the computer, and a pointing device such as a mouse or trackball.
[0129] One or more features or steps of the disclosed embodiments may be implemented using an Application Programming Interface (API). An API may define one or more parameters passed between a calling application and other software code (e.g., an operating system, library routines, or functions) that provides services, provides data, or performs operations or calculations. An API may be implemented as one or more calls of program code that send or receive one or more parameters through a parameter list or other structure based on a calling convention defined in the API specification. Parameters may be constants, keys, data structures, objects, object classes, variables, data types, pointers, arrays, lists, or other calls. API calls and parameters may be implemented in any programming language. The programming language may define vocabulary and calling conventions that a programmer would use to access the functionality supporting the API. In some implementations, an API call may report to the application the capabilities of the device running the application, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, etc.
[0130] Several implementation examples have been provided. Nevertheless, it will be understood that various modifications may be made. Elements of one or more implementation examples may be combined, deleted, modified, or supplemented to form further implementation examples. In yet another example, the logical flow depicted in the diagram does not require a specific order or sequence to be illustrated in order to achieve the desired result. In addition, other steps may be added or steps may be removed from the described flow, and other components may be added to or components removed from the described system. Accordingly, other implementation examples are within the scope of the following claims. [Explanation of symbols]
[0131] 100 Operating environment 101 Wearable Multimedia Devices 102 Cloud Computing Platform 103 Network 104 Application Developers 105 Third-Party Platforms 106 Databases 200 Data Processing Systems 201 Recorder 202 video buffer 203 Audio Buffer 204 Photo Buffer 205 Ingestion Server 206 Datastores 207 Video Processors 208 Audio Processors 209 Photo Processors 210 Third-Party Processors 211 App 1 212 App 2 213 App 1 214 App 2 215 App 1 216 App 2 217 apps 300 data processing pipelines 301 Intake Server 302 Video Processors 304 Recorder 305 Audio Processor 306 Apps 307 Video Processing Apps 308 Data Merge Process 400 data processing pipelines 401 Inbound Server 402 Third-party processors 404 Recorder 405 Audio Processor 406 App 407 Third-party applications 408 User Callback Queue 700 Architecture 702 Processor 704 Storage Devices 706 Network Interface 708 Computer-readable media 710 Communication Channel 712 Operating Systems 714 Network Communication Module 716 Data Processing Instructions 718 Interface Instructions 800 Architecture 802 Memory Interface 804 Processor 806 Peripheral Interface 810 Motion Sensor 812 Biometric Sensor 814 Depth Sensor 815 Position processor 816 Electronic magnetometer 817 Environmental Sensor 820 Camera Subsystem 822 Light Sensor 824 Communication Subsystem 826 Audio Subsystem 828 Speakers 830 Microphone 840 I / O subsystems 842 Touch Controller 844 Other Input Controllers 846 Touch surface 848 Other input / control devices 850 memory 852 Operating Systems 854 Communication Order 858 Sensor processing command 860 Recorder command 900 Graphical User Interface 901 Video Pane 902 hours / location data 903 Object 904 Search button 905 Categories 906a, 906b, 906c objects 907 Thumbnail Images 1000 Classifier Frameworks 1001 API 1002a~1002n classifier 1005 Datastore 1100 Hardware Architectures 1101 System-on-a-Chip 1102 Main Camera 1103 3D camera 1104 Capacitive Sensor 1105 Motion Sensor 1106 Microphone 1107 memory 1108 Global Navigation Satellite System Receiver 1109 WiFi / Bluetooth Chip 1110 Wireless Transceiver Chip 1112 Radio frequency transceiver chip 1113 RF Front-End Electronic Circuit 1114 LED 1115 Projector 1116 Audio Amplifier 1117 Speaker 1118 External Battery 1119 Magnetic Inductance Network 1120 Power Management Chip 1121 Internal Battery 1200 Cloud Computing Platforms 1201 Edge Server 1202 Wearable Multimedia Devices 1203 Dispatcher 1204 State / Context Detector 1205 First-Party Handler 1206 Limited AI Resolver 1207 Area Data Center 1208 Dispatcher 1209 State / Context Detector 1210 Fully AI Resolver 1211 First Handler 1212 Second Handler 1300 Software Components 1301 Daemon 1302 Library 1303 Operating Systems 1304 Other software components 1700 projectors, projector architecture 1701 Controller 1704 Solid-state laser 1705 XY Scanner 1706 Driver 1707 memory 1708 Digital-to-Analog Converter 1709 Analog-to-Digital Converter 1800 yen pattern 1801 Surface 1802, 1803 area
Claims
1. A body-worn device, Camera and, Depth sensor and, Laser projection system and One or more processors, When executed by the one or more processors, the one or more processors will Using the aforementioned camera, capture a first set of digital images, Identifying real-world objects within the first set of digital images, Using the aforementioned depth sensor, first depth data is captured, The system includes a memory that stores instructions for performing an action that includes identifying a first set of digital images and a first gesture of a user wearing the body-worn device within the first depth data, Identifying the aforementioned real-world object and the first gesture, The method involves processing the first set of digital images through an object detection framework using composite polygons to identify hotspot regions in the first set of digital images, wherein the hotspot regions are smaller than the entire image, and while excluding all other objects in the first set of digital images, the method captures and identifies the first gesture and the real-world objects. Sending the aforementioned hotspot area to a cloud computing platform, Receiving information related to the real-world object from the aforementioned cloud computing platform, The laser projection system includes projecting at least some of the information onto a surface using a laser, Body-worn device.
2. The apparatus according to claim 1, wherein the information includes text labels for the real-world objects.
3. The apparatus according to claim 1, wherein the information includes instructions for performing an action on the real-world object.
4. The apparatus according to claim 1, wherein the information includes commands for controlling the real-world object.
5. The aforementioned operation, Using the depth sensor, a second gesture associated with the laser projection is acquired, Determining user input based on the second gesture described above, The apparatus according to claim 1, further comprising initiating one or more actions in accordance with the user input.
6. The aforementioned operation, The apparatus according to claim 5, further comprising masking the laser projection to prevent the information from being projected onto the user's hand making the second gesture.
7. The aforementioned operation, Using the aforementioned camera, capture the reflection of the laser projection from the surface, The apparatus according to claim 1, further comprising automatically adjusting the intensity of the laser projection to compensate for different refractive indices so that the laser projection has a uniform brightness.
8. A magnetic mounting mechanism configured to magnetically couple to a battery pack through the user's clothing, and further configured to receive inductive charging from the battery pack. The apparatus according to claim 1, further comprising the following:
9. The steps involve using a camera to capture a set of digital images, The steps include capturing depth data using a depth sensor on a body-worn device, The first step is to identify real-world objects within a set of digital images, The process includes the step of identifying a first set of digital images and a first gesture in the depth data, which was made by a user wearing the body-worn device, The step of identifying the real-world object and the first gesture is: A step of processing a first set of digital images through an object detection framework using composite polygons to identify a hotspot region of the first set of digital images, wherein the hotspot region is smaller than the entire image and captures and identifies the first gesture and the real-world object while excluding all other objects in the first set of digital images. The steps include sending the aforementioned hotspot area to a cloud computing platform, Receiving information related to the real-world object from the aforementioned cloud computing platform, A laser projection system is used to project at least some of the aforementioned information onto a surface. A method that includes this.
10. A step of using the depth sensor to acquire a second gesture by the user, which is associated with the laser projection. The steps include determining user input based on the second gesture described above, The steps include: initiating one or more actions in accordance with the user input; The method according to claim 9, further comprising:
11. The method according to claim 10, wherein one or more of the actions include controlling the real-world object.
12. The step of masking the laser projection to prevent the information from being projected onto the user's hand while the second gesture is being performed. The method according to claim 10, further comprising:
13. The method according to claim 11, wherein the controllable real-world object is a television or computer screen, and the one or more actions include swiping through one or more images displayed by the television or computer screen.
14. The method according to claim 11, wherein the controllable real-world object is a thermostat, and the one or more actions include changing the temperature setting of the thermostat.
15. The method according to claim 9, wherein the laser projection includes a size template for measuring the real-world object.
16. The steps include receiving voice input from the user, The method according to claim 9, further comprising the step of using one or more processors to associate the first gesture and voice input with a request or command for controlling the real-world object.
17. The method according to claim 9, wherein the surface is the palm of the user's hand.
Citation Information
Patent Citations
Local advertising content on an interactive head-mounted eyepiece
JP2013521576A
Display control device, display control method, and program
JP2014132478A
Head-mounted display device, control system, method for controlling head-mounted display device, and computer program
JP2016148968A
Human body gesture-based region and volume selection for hmd
JP2016514298A
Head-mounted display for cooking and program for head-mounted display for cooking
JP2017120329A