Wearable multimedia device and cloud computing platform with an application ecosystem
The wearable multimedia device and cloud computing platform address the limitations of mobile devices in capturing spontaneous moments by automatically processing and editing multimedia data, ensuring that important events are preserved and easily shareable.
Patent Information
- Application Number
- JP2023186584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-05-10
- Filing Date
- 2023-10-31
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2038-05-10
AI Technical Summary
Modern mobile devices with embedded cameras often fail to capture important moments due to their limitations in capturing spontaneous events, especially when users are emotionally involved or when moments occur too quickly.
A wearable multimedia device and a cloud computing platform with an application ecosystem that processes multimedia data captured by the device, allowing for automatic editing and formatting based on user preferences, and enabling seamless playback on various devices.
The wearable multimedia device captures multimedia data with minimal user interaction, and the cloud computing platform automatically edits and formats it, ensuring that important moments are preserved and can be easily shared or replayed on various devices.
Smart Images

Figure 0007696976000001 
Figure 0007696976000002 
Figure 0007696976000003
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to cloud computing and multimedia editing.
Background Art
[0002] Modern mobile devices (e.g., smartphones, tablet computers) often include embedded cameras that allow users to capture digital images or videos of spontaneous events. These digital images and videos can be stored in an online database associated with the user account to free up memory on the mobile device. Users can share those images and videos with friends and family and download or stream the images and videos on demand using their various playback devices. These embedded cameras offer significant advantages compared to conventional digital cameras, which are larger and more cumbersome and often require more time to set up shots.
[0003] Despite the convenience of mobile devices with embedded cameras, there are many important moments that are not captured by these devices or that the user simply forgets to take an image or video of because the moment occurs too quickly or the user becomes emotionally involved in the moment.
Summary of the Invention
[0004] Disclosed are a system, method, device, and non-transitory computer-readable storage medium for a wearable multimedia device and a cloud computing platform that includes an application ecosystem for processing multimedia data captured by the wearable multimedia device.
[0005] In one embodiment, the method comprises receiving, by one or more processors of a cloud computing platform, context data from a wearable multimedia device, the wearable multimedia device including at least one data capture device for capturing the context data; creating, by the one or more processors, a data processing pipeline including one or more applications based on the context data and one or more characteristics of a user request; processing, by the one or more processors, the context data through the data processing pipeline; and transmitting, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or to another device for presenting the output.
[0006] In one embodiment, the system comprises one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including receiving, by one or more processors of a cloud computing platform, context data from a wearable multimedia device, the wearable multimedia device including at least one data capture device for capturing the context data; creating, by the one or more processors, a data processing pipeline including one or more applications based on the context data and one or more characteristics of a user request; processing, by the one or more processors, the context data through the data processing pipeline; and transmitting, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or to another device for presenting the output.
[0007] In one embodiment, the non - transitory computer - readable storage medium is to receive context data from a wearable multimedia device by one or more processors of a cloud computing platform, wherein the wearable multimedia device includes at least one data capture device for capturing the context data; to create, by one or more processors, a data processing pipeline including one or more applications based on the context data and one or more characteristics of user requirements; to process the context data through the data processing pipeline by one or more processors; and to transmit, by one or more processors, the output of the data processing pipeline to the wearable multimedia device or another device for presenting the output, and includes instructions for performing the above.
[0008] Certain embodiments disclosed herein provide one or more of the following advantages. The wearable multimedia device captures multimedia data of spontaneous moments and transactions with minimal user interaction. The multimedia data is automatically edited and formatted on a cloud computing platform based on user preferences and then made available to the user for playback on various user playback devices. In one embodiment, the data editing and / or processing is proprietary and / or is performed by an ecosystem of applications provided / authorized by third - party developers. The application ecosystem provides various access points (e.g., websites, portals, APIs) that enable third - party developers to upload, verify, and update their applications. The cloud computing platform automatically constructs a custom processing pipeline for each multimedia data stream using one or more of the ecosystem applications, user preferences, and other information (e.g., data type or format, data volume and quality).
[0009] The details of the disclosed embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0011] Like reference numerals used in the various drawings refer to like elements.
Best Mode for Carrying Out the Invention
[0012] Overview A wearable multimedia device is a lightweight, small form factor battery-powered device that can be attached to a user's clothing or object using a tension clasp, an interlock pin back, a magnet, or any other attachment mechanism. The wearable multimedia device includes a digital image capture device (e.g., 180° FOV with optical image stabilizer (OIS)) that enables a user to spontaneously capture multimedia data (e.g., video, audio, depth data) and document transactions (e.g., financial transactions) of life events ("moments") with minimal user interaction or device setup. The multimedia data ("context data") captured by the wireless multimedia device is uploaded to a cloud computing platform with an application ecosystem that can process, edit, and format the context data into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played back on the wearable multimedia device and / or any other playback device by one or more applications (e.g., Artificial Intelligence (AI) applications). For example, the cloud computing platform can transform video data and audio data into any desired movie-making style (e.g., documentary, lifestyle, candid, photojournalism, sports, street) specified by the user.
[0013] In one embodiment, the context data is processed by a server computer(s) of a cloud computing platform based on user preferences. For example, based on user preferences, an image can be fully color-graded, stabilized, and trimmed at the moment when the user desires an immersive experience. The user preferences can be stored in a user profile created by the user via an online account accessible through a website or portal, or the user preferences can be learned by the platform over time (e.g., using machine learning). In one embodiment, the cloud computing platform is an extensible distributed computing environment. For example, the cloud computing platform can be a distributed streaming platform (e.g., Apache Kafka™) that includes a real-time streaming data pipeline and streaming applications that transform or react to data streams.
[0014] In one embodiment, the user can start and stop a context data capture session on a wearable multimedia device by uttering a command or any other input mechanism, or with a simple touch gesture (e.g., tap or swipe). All or part of the wearable multimedia device can automatically power down when it detects, using one or more sensors (e.g., proximity sensor, light sensor, accelerometer, gyroscope), that the device is not being worn by the user. In one embodiment, the device can include a photovoltaic surface technology to enable inductive charging for a charging mat and wireless over-the-air (OTA) charging while maintaining battery life and an inductive charging circuit configuration (e.g., Qi).
[0015] The context data is encrypted and compressed and can be stored in an online database associated with the user account using any desired encryption or compression technology. The context data can be stored for a specified period that can be set by the user. The user can be provided with an opt-in mechanism and other tools for managing those data and data privacy via a website, portal, or mobile application.
[0016] In one embodiment, the context data includes point cloud data for providing a mapped object of a three-dimensional (3D) surface that can be processed, for example, using augmented reality (AR) and virtual reality (VR) applications within an application ecosystem. The point cloud data can be generated by a depth sensor (e.g., LiDAR or Time of Flight (TOF)) embedded in a wearable multimedia device.
[0017] In one embodiment, the wearable multimedia device includes a Global Navigation Satellite System (GNSS) receiver (e.g., Global Positioning System (GPS)) and one or more inertial sensors (e.g., accelerometer, gyroscope) for determining the position and orientation of the user wearing the device when the context data is captured. In one embodiment, one or more images within the context data can be used by a location-specific application, such as a visual odometry application, within the application ecosystem to determine the position and orientation of the user.
[0018] In one embodiment, the wearable multimedia device may also include one or more environmental sensors, including but not limited to an ambient light sensor, a magnetometer, a pressure sensor, a voice activity detector, etc. The sensor data can be included in the context data to enhance content presentation that includes additional information for use in capturing moments.
[0019] In one embodiment, the wearable multimedia device may include one or more biometric sensors such as a heart rate sensor, a fingerprint scanner, etc. The sensor data can be included in the context data to document a transaction or to indicate the user's emotional state during a moment (for example, an increase in heart rate may indicate excitement or fear).
[0020] In one embodiment, the wearable multimedia device includes a headset or earphone for receiving voice commands and capturing ambient audio, and a headphone jack for connecting one or more microphones. In an alternative embodiment, the wearable multimedia device includes short-range communication technologies including but not limited to Bluetooth®, IEEE802.15.4 (ZigBee™), and near field communication (NFC). The short-range communication technologies may be used to wirelessly connect to a wireless headset or earphone in addition to or instead of the headphone jack, and / or to wirelessly connect to any other external device (for example, a computer, a printer, a laser projector, a television, and other wearable devices).
[0021] In one embodiment, the wearable multimedia device includes a wireless transceiver and a communication protocol stack for various communication technologies including WiFi, 3G, 4G, and 5G communication technologies. In one embodiment, the headset or earphone also includes sensors (e.g., biometric sensors, inertial sensors) that provide information about the direction the user is facing in order to provide commands using head gestures. In one embodiment, the camera direction can be controlled by head gestures so that the camera field of view follows the direction of the user's field of view. In one embodiment, the wearable multimedia device can be embedded in or attached to the user's glasses.
[0022] In one embodiment, the wearable multimedia device includes a laser projector (or can be wired or wirelessly coupled to an external laser projector) that enables the user to play back moments on a surface such as a wall or tabletop. In another embodiment, the wearable multimedia device includes an output port that can be connected to a laser projector or other output device.
[0023] In one embodiment, the wearable multimedia capture device includes a touch surface that responds to touch gestures (e.g., tap, multi - tap, or swipe gestures). The wearable multimedia device can include a small display for presenting information and one or more optical indicators for indicating the on / off status, power state, or any other desired status.
[0024] In one embodiment, the cloud computing platform can be driven by context-based gestures (e.g., air gestures) in combination with spoken queries such as when a user points to an object in their environment and says "What is that building?" The cloud computing platform uses air gestures to narrow the field of view port of the camera and isolate the building. One or more images of the building are captured and sent to the cloud computing platform where an image recognition application can operate on the image query and store the results or return them to the user. Air gestures and touch gestures can also be performed on a projected transient display, for example, in response to user interface elements.
[0025] In one embodiment, context data can be encrypted on the device and the cloud computing platform such that only the user or any authorized viewer can experience the moment on a connected screen (e.g., smartphone, computer, television, etc.) or as a laser projection onto a surface. Referring to FIG. 8, an exemplary architecture for a wearable multimedia device will be described.
[0026] In addition to personal life events, wearable multimedia devices simplify the capture of financial transactions currently handled by smartphones. The capture of daily transactions (e.g., business transactions, microtransactions) is made simpler, faster, and more fluid by using the vision-assisted context awareness provided by wearable multimedia devices. For example, when a user is engaged in a financial transaction (e.g., making a purchase), the wearable multimedia device will generate data that memorializes the financial transaction, including the date, time, amount, digital image or video of the parties, audio (e.g., user commentary explaining the transaction), and environmental data (e.g., location data). The data may be included in a multimedia data stream transmitted to a cloud computing platform where the data can be stored online and / or processed by one or more financial applications (e.g., financial management, accounting, budgeting, tax filing, inventory, etc.).
[0027] In one embodiment, the cloud computing platform provides a graphical user interface on a website or portal, thereby enabling various third-party application developers to upload, update, and manage their own applications within the application ecosystem. Some exemplary applications include, but are not limited to, personal live broadcasts (e.g., Instagram™ Life, Snapchat™), senior monitoring (e.g., to ensure that a loved one has taken their medication), memory recall (e.g., showing last week's children's soccer game), and personal guides (e.g., an AI-enabled personal guide that recognizes the user's location and guides the user to take action).
[0028] In one embodiment, a wearable multimedia device includes one or more microphones and a headset. In some embodiments, the headset wire includes a microphone. In one embodiment, a digital assistant is implemented in the wearable multimedia device, which responds to user queries, requests, and commands. For example, a wearable multimedia device worn by a parent captures context data at the moment of a child's soccer game, particularly the "moment" when the child scores a goal. The user can request (e.g., using a voice command) that the platform create a video clip of the goal and store it in the user's own user account. Without any further action by the user, the cloud computing platform identifies the correct portion of the context data at the moment the goal is scored (e.g., using face recognition, visual cues, or audible cues), edits the context data of the moment into a video clip, and stores the video clip in a database associated with the user account. Exemplary Operating Environment
[0029] FIG. 1 is a block diagram of an operating environment for a wearable multimedia device and a cloud computing platform for processing multimedia data captured by the wearable multimedia device according to one embodiment. The operating environment 100 includes a wearable multimedia device 101, a cloud computing platform 102, a network 103, an application ( "app") developer 104, and a third-party platform 105. The cloud computing platform 102 is coupled to one or more databases 106 for storing context data uploaded by the wearable multimedia device 101.
[0030] As described above, the wearable multimedia device 101 is a lightweight and small form factor battery-powered device that can be attached to the user's clothing or objects using a tension clasp, an interlock pin back, a magnet, or any other attachment mechanism. The wearable multimedia device 101 includes a digital image capture device (e.g., 180° FOV with OIS) that enables the user to spontaneously capture "instant" multimedia data (e.g., video, audio, depth data) with minimal user interaction or device setup and document daily transactions (e.g., financial transactions). The context data captured by the wireless multimedia device 101 is uploaded to a cloud computing platform 102. The cloud computing platform 101 includes an application ecosystem that can process, edit, and format the context data into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image gallery) that can be downloaded and played back on the wearable multimedia device and / or other playback devices by one or more server-side applications.
[0031] As an example, at a child's birthday party, the parent can clip the wearable multimedia device to their own clothing (or attach the device to a necklace or chain and wear it around their own neck) so that the camera lens faces in the parent's line of sight. The camera includes a 180° FOV that enables the camera to capture substantially all that the user is currently viewing. The user can start recording by simply tapping the surface of the device or pressing a button. No additional setup is required. A multimedia data stream (e.g., a video including audio) that captures special moments of the birthday (e.g., the moment of blowing out the candles) is recorded. This “context data” is transmitted in real time via a wireless network (e.g., WiFi, cellular) to the cloud computing platform 102. In one embodiment, the context data is stored on the wearable multimedia device so that it can be uploaded later. In another embodiment, the user can transfer the context data to another device (e.g., a personal computer hard drive, smartphone, tablet computer, thumb drive) and later use an application to upload the context data to the cloud computing platform 102.
[0032] In one embodiment, the context data is processed by one or more applications of an application ecosystem hosted and managed by a cloud computing platform 102. The applications can be accessed through their respective application programming interfaces (APIs). A custom distributed streaming pipeline is created by the cloud computing platform 102 to process the context data based on one or more of the type of data, the amount of data, the quality of data, user preferences, templates, and / or any other information, and generate a desired presentation based on user preferences. In one embodiment, machine learning techniques can be used to automatically select suitable applications for inclusion within the data processing pipeline, regardless of the presence or absence of user preferences. For example, historical user context data stored in a database (e.g., a NoSQL database) can be used to determine user preferences for data processing using any suitable machine learning technique (e.g., deep learning or convolutional neural networks).
[0033] In one embodiment, the application ecosystem can include a third-party platform 105 that processes the context data. A secure session is set up between the cloud computing platform 102 and the third-party platform 105 to send / receive the context data. This design enables the third-party app provider to control access to its own applications and provide updates. In other embodiments, the applications operate on the servers of the cloud computing platform 102 and updates are sent to the cloud computing platform 102. In the latter embodiments, the app developer 104 can upload and update the applications included within the application ecosystem using the APIs provided by the cloud computing platform 102. Exemplary Data Processing System
[0034] FIG. 2 is a block diagram of a data processing system implemented by the cloud computing platform of FIG. 1 according to one embodiment. The data processing system 200 includes a recorder 201, a video buffer 202, an audio buffer 203, a photo buffer 204, an absorption server 205, a data store 206, a video processor 207, an audio processor 208, a photo processor 209, and a third-party processor 210.
[0035] The recorder 201 (e.g., a software application) operating in the wearable multimedia device records video, audio, and photo data (the "context data") captured by the camera and audio subsystem, and stores the data in the buffers 202, 203, 204 respectively. This context data is then sent (e.g., using wireless OTA technology) to the absorption server 205 of the cloud computing platform 102. In one embodiment, the data can be sent in separate data streams each having a unique stream identifier (stream ID). A stream is a separate piece of data that can include the following exemplary attributes: location (e.g., latitude, longitude), user, audio data, video streams of various durations, and N photos. A stream can have a duration of 1 to MAXSTREAM_LEN seconds, where in this example, MAXSTREAM_LEN = 20 seconds.
[0036] Absorption server 205 absorbs the stream, creates stream records in data store 206, and stores the results of processors 207-209. In one embodiment, the audio stream is processed first and used to determine the other streams required. Absorption server 205 sends the stream to the appropriate processor 207-209 based on the stream ID. For example, the video stream is sent to video processor 207, the audio stream is sent to audio processor 208, and the photo stream is sent to photo processor 209. In one embodiment, at least a portion of the data collected from the wearable multimedia device (e.g., image data) is processed into metadata and encrypted so that it can be further processed by a given application and sent back to the wearable multimedia device or other device.
[0037] Processors 207-209 can operate proprietary applications or third-party applications as described above. For example, video processor 207 can be a video processing server that sends the raw video data stored in video buffer 202 to a set of one or more image processing / editing applications 211, 212 based on user preferences or other information. Processor 207 sends the request to applications 211, 212 and returns the results to absorption server 205. In one embodiment, third-party processor 210 can process one or more of the streams using its own processors and applications. In another example, audio processor 208 can be an audio processing server that sends the speech data stored in audio buffer 203 to a speech-to-text converter application 213. Exemplary scene identification application
[0038] FIG. 3 is a block diagram of a data processing pipeline for processing a context data stream according to an embodiment. In this embodiment, the data processing pipeline 300 is created and configured to determine what the user is looking at based on context data captured by a wearable multimedia device worn by the user. The absorption server 301 receives an audio stream (e.g., including user commentary) from the audio buffer 203 of the wearable multimedia device and transmits the audio stream to the audio processor 305. The audio processor 305 transmits the audio stream to the app 306, and the app 306 performs speech-to-text conversion and returns the analyzed text to the audio processor 305. The audio processor 305 returns the analyzed text to the absorption server 301.
[0039] The video processor 302 receives the analyzed text from the absorption server 301 and transmits the request to the video processing app 307. The video processing app 307 identifies objects within the video scene and uses the analyzed text to label the objects. The video processing app 307 transmits a response that describes the scene (e.g., the labeled objects) to the video processor 302. The video processor then transfers the response to the absorption server 301. The absorption server 301 transmits the response to the data merge process 308, and the data merge process 308 merges the response with the user's position, orientation, and map data. The data merge process 308 returns the response with the scene description to the recorder 304 on the wearable multimedia device. For example, the response may include text that describes a scene such as a child's birthday party, including the position on the map and descriptions of the objects within the scene (e.g., identifying the people within the scene). The recorder 304 associates the scene description with the multimedia data stored on the wearable multimedia device (e.g., using a stream ID). When the user calls up the data, the data is enhanced with the scene description.
[0040] In one embodiment, the data merge process 308 may use more than just location and map data. The concept of ontology may also exist. For example, the features of the face of the user's father captured in an image can be recognized by a cloud computing platform and returned as "father" instead of the user's name, and an address such as "555 Main Street, San Francisco, CA" can be returned as "home". The ontology can be unique to the user and can grow and learn from the user's input. Exemplary transportation application
[0041] Figure 4 is another block diagram of data processing for processing the context data stream of a transportation application according to an embodiment. In this embodiment, the data processing pipeline 400 is created to call a transportation company (e.g., Uber (registered trademark), Lyft (registered trademark)) to obtain transportation means to the user's home. Context data from the wearable multimedia device is received by the absorption server 401, and the audio stream from the audio buffer 203 is sent to the audio processor 405. The audio processor 405 sends the audio stream to an application 406 that converts speech to text. The analyzed text is returned to the audio processor 405 that returns the analyzed text to the absorption server 401 (e.g., the user's speech request for transportation). The processed text is sent to a third-party processor 402. The third-party processor 402 sends the user's location and a token to a third-party application 407 (e.g., the Uber (registered trademark) or Lyft (trademark) (registered trademark) application). In one embodiment, the token is an API and an authorization token used to mediate requests on behalf of the user. The application 407 returns a response data structure to the third-party processor 402, which is transferred to the absorption server 401. The absorption server 401 checks the transportation arrival status (e.g., ETA) in the response data structure and sets up a callback to the user in the user callback queue 408. The absorption server 401 returns a response with vehicle description to a recorder 404 that can speak to the user through a loudspeaker on the wearable multimedia device or through the user's headphones or earphones via a wired or wireless connection by a digital assistant.
[0042] FIG. 5 illustrates a data object used by the data processing system of FIG. 2 according to one embodiment. The data object is part of a software component infrastructure instantiated on a cloud computing platform. The "stream" object includes a stream ID of data, a device ID, a start, an end, a latitude, a longitude, an attribute, and an entity. The "stream ID" identifies a stream (e.g., video, audio, photo), the "device ID" identifies a wearable multimedia device (e.g., a mobile device ID), the "start" is the start time of the context data stream, the "end" is the end time of the context data stream, the "latitude" is the latitude of the wearable multimedia device, the "longitude" is the longitude of the wearable multimedia device, the "attribute" includes, for example, a birthday, facial points, skin tone, audio characteristics, an address, a phone number, etc., and the "entity" constitutes an ontology. For example, the name "John Do" will be mapped to "father" or "brother" depending on the user.
[0043] The "user" object includes a user ID of data, a device ID, an email, an f name, and an l name. The user ID identifies a user having a unique identifier, the device ID identifies a wearable device having a unique identifier, the email is the registered email address of the user, the f name is the first name of the user, and the l name is the last name of the user. The "user device" object includes a user ID and a device ID of data. The "device" object includes a device ID, a start, a state, a modification, and a creation of data. In one embodiment, the device ID is a unique identifier of the device (e.g., separate from the MAC address). The start is when the device was first started. The state is on / off / sleep. The modification is the last change in state or the last modification date reflecting a change in the operating system (OS). The creation is the first time the device was turned on.
[0044] The "processing result" object includes the data stream ID, ai, result, callback, and duration accuracy. In one embodiment, the stream ID is each user stream as a Universally Unique Identifier (UUID). For example, a stream started from 8:00 AM to 10:00 AM has id: 15h158dhb4, and a stream started from 10:15 AM to 10:18 AM will have the UUID that contacted this stream. The AI is the identifier of the platform application that contacted this stream. The result is the data sent from the platform application. The callback is the callback used (the version can be changed, so the callback is tracked if the platform needs to reproduce the request). The accuracy is a score regarding how accurate the result set is. In one embodiment, the processing result can be used for many things, such as 1) notifying the merge server of the entire set of results, 2) determining the fastest ai to enhance the user experience, and 3) determining the most accurate ai. Depending on the use case, speed may be prioritized over accuracy, or vice versa.
[0045] The "entity" object includes an entity ID of data, a user ID, an entity name, an entity type, and entity attributes. The entity ID is the UUID of the entity, and the entity ID is an entity that has multiple inputs referring to one entity. For example, "Barack Obama" may have 144 entity IDs, which may be linked to POTUS44 or "Barack Hussein Obama" or "President Obama" in the related table. The user ID identifies the user by whom the entity record was created. The entity name is the name by which the user ID calls the entity. For example, the entity name of Malia Obama with entity ID 144 may be "father" or "dad". The entity type is a person, a place, or a thing. The entity attributes are an array of attributes regarding the entity that are specific to the understanding of the user ID of that entity. This mapping enables, for example, when Malia makes a speech query such as "Can you see my father?", the cloud computing platform to translate the query to Barack Hussein Obama and use it when mediating requests to third parties and searching for information within the system to map the entities together. Exemplary processing
[0046] FIG. 6 is a flowchart of data pipeline processing according to an embodiment. The processing 600 can be implemented using the wearable multimedia device 101 and the cloud computing platform 102 described with reference to FIGS. 1-5.
[0047] The processing 600 can start by receiving context data from a wearable multimedia device (601). For example, the context data may include video, audio, and still images captured by the camera and audio subsystem of the wearable multimedia device.
[0048] Process 600 may continue (602) by creating (e.g., instantiating) a data processing pipeline that includes an application based on context data and the user's requests / preferences. For example, based on the user's requests or preferences and also based on the type of data (e.g., audio, video, photo), one or more applications can be logically connected to form a data processing pipeline, and the context data can be processed into a presentation to be played back on a wearable multimedia device or another device.
[0049] Process 600 may continue (603) by processing context data within the data processing pipeline. For example, utterances from an in - moment or transactional user commentary can be converted to text and then used to label objects within a video clip.
[0050] Process 600 may continue (604) by sending the output of the data processing pipeline to a wearable multimedia device and / or other playback device. Exemplary Cloud Computing Platform Architecture
[0051] FIG. 7 is an exemplary architecture 700 of a cloud computing platform 102 described with reference to FIGS. 1-6 according to one embodiment. Other architectures are possible, including architectures having more or fewer components. In some embodiments, architecture 700 includes one or more processors 702 (e.g., dual-core Intel® Xeon® processors), one or more network interfaces 706, one or more storage devices 704 (e.g., hard disk, optical disk, flash memory), and one or more computer-readable media 708 (e.g., hard disk, optical disk, flash memory, etc.). These components can communicate and exchange data via one or more communication channels 710 (e.g., a bus), which can utilize various hardware and software to facilitate the transfer of data and control signals between components.
[0052] The term "computer-readable media" refers to any media that participates in providing instructions to processor(s) 702 for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media includes, but is not limited to, coaxial cables, copper wire, and fiber optics.
[0053] Computer-readable media 708 may further include an operating system 712 (e.g., Mac OS® Server, Windows® NT Server, Linux® Server), a network communication module 714, interface instructions 716, and data processing instructions 718.
[0054] The operating system 712 can be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system 712 recognizes inputs from devices 702, 704, 706, and 708, provides outputs to these devices, tracks and manages files and directories of computer-readable medium(s) 708 (e.g., memory or storage device), controls peripheral devices, and manages traffic on one or more communication channels 710, and performs basic tasks including but not limited to these. The network communication module 714 includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.) and for creating a distributed streaming platform using, for example, Apache Kafka (trademark). The data processing instructions 716 include server-side or backend software for implementing server-side operations as described with reference to FIGS. 1 - 6. The interface instructions 718 include software for implementing a web server and / or portal for transmitting and receiving data with the wearable multimedia device 101, third-party application developers 104, and third-party platforms 105 as described with reference to FIG. 1.
[0055] The architecture 700 can be included within any computer device that includes one or more server computers within a local network or distributed network, each having one or more processing cores. The architecture 700 can be implemented in a parallel processing or peer-to-peer infrastructure or in a single device having one or more processors. The software can include multiple software components or can be a single code body. Exemplary Wearable Multimedia Device Architecture
[0056] FIG. 8 is a block diagram of an exemplary architecture 800 for a wearable multimedia device that implements the features and processes described with reference to FIGS. 1-6. Architecture 800 can be implemented within any wearable multimedia device 101 that implements the features and processors described with reference to FIGS. 1-6. Architecture 800 can include a memory interface 802, a data processor(s), an image processor(s) or a central processing unit(s) 804, and a peripheral interface 806. The memory interface 802, the processor(s) 804, or the peripheral interface 806 can be separate components or can be integrated within one or more integrated circuits. One or more communication buses or signal lines can couple the various components.
[0057] Sensors, devices, and subsystems can be coupled to the peripheral interface 806 to facilitate a plurality of functions. For example, one or more motion sensors 810, one or more biometric sensors 812, and a depth sensor 814 can be coupled to the peripheral interface 806 to facilitate motion, orientation, biometric, and depth detection functions. In some implementations, one or more motion sensors 810 (e.g., accelerometers, rate gyroscopes) can be utilized to detect the movement and orientation of the wearable multimedia device.
[0058] Other sensors such as one or more environmental sensors (e.g., temperature sensors, barometers, ambient light) can also be connected to the peripheral interface 806 to facilitate environmental sensing functions. For example, biometric sensors can detect fingerprints, face recognition, heart rate, and other fitness parameters. In one embodiment, a tactile motor (not shown) can be coupled to the peripheral interface, thereby providing a vibration pattern to the user as tactile feedback.
[0059] A position processor 815 (e.g., a GNSS receiver chip) can be connected to the peripheral device interface 806 to provide georeferencing. An electronic magnetometer 816 (e.g., an integrated circuit chip) can also be connected to the peripheral device interface 806 to provide data that can be used to determine the magnetic north direction. Thus, the electronic magnetometer 816 can be used by an electronic compass application.
[0060] The camera subsystem 820 and the optical sensor 822, e.g., a charged coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) optical sensor, can be utilized to facilitate camera functions such as recording photos and video clips. In one embodiment, the camera has a 180° FOV and OIS. The depth sensor can include an infrared emitter that projects dots in a known pattern onto the object / subject. The dots are then captured by a dedicated infrared camera and analyzed to determine depth data. In one embodiment, a time-of-flight (TOF) camera can be used to determine distances based on the known speed of light and measure the time of flight of the optical signal between the camera and the object / subject for each point in the image.
[0061] The communication function can be facilitated via one or more communication subsystems 824. The communication subsystem(s) 824 can include one or more wireless communication subsystems. The wireless communication subsystem 824 can include a radio frequency receiver and transmitter, and / or an optical (e.g., infrared) receiver and transmitter. The wired communication system can include connections to port devices, such as Universal Serial Bus (USB) ports, or some other wired connection ports, which can be used to establish a wired connection to other computing devices, such as other communication devices, network access devices, personal computers, printers, display screens, or other processing devices (e.g., laser projectors) that can receive or transmit data.
[0062] The specific design and implementation example of the communication subsystem 824 may depend on the communication network(s) or medium(s) on which the device is intended to operate. For example, the device may include a wireless communication subsystem designed to operate via a global system for mobile communications (GSM (registered trademark, the same hereinafter)) network, a GPRS network, an enhanced data GSM environment (EDGE) network, an IEEE802.xx communication network (e.g., WiFi, WiMax, ZigBee (trademark)), 3G, 4G, 4G LTE, a code division multiple access (CDMA) network, near field communication (NFC), Wi-Fi Direct, and a Bluetooth (registered trademark) network. The wireless communication subsystem 824 may include a hosting protocol so that the device can be configured as a base station for other wireless devices. As another example, the communication subsystem may enable the device to synchronize with a host device using one or more protocols or communication technologies such as, for example, the TCP / IP protocol, the HTTP protocol, the UDP protocol, the ICMP protocol, the POP protocol, the FTP protocol, the IMAP protocol, the DCOM protocol, the DDE protocol, the SOAP protocol, HTTP live streaming, MPEG Dash, and any other known communication protocol or technology.
[0063] Coupling the audio subsystem 826 to the speaker 828 and one or more microphones 830 can facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, telephone functionality, and beamforming.
[0064] The I / O subsystem 840 may include a touch controller 842 and / or one or more other input controllers 844. The touch controller 842 can be coupled to a touch surface 846. The touch surface 846 and the touch controller 842 can use any of several touch sensitivity technologies, including, for example, capacitive technology, resistive technology, infrared technology, and surface acoustic wave technology, as well as other proximity sensor arrays or other elements for determining one or more contact points with the touch surface 846, to detect contact, movement, or failure of the touch surface 846 and the touch controller 842. In one implementation, the touch surface 846 can display virtual buttons or soft buttons that can be used by a user as input / output devices.
[0065] The one or more other input controllers 844 can be coupled to other input / control devices 848 such as one or more buttons, rocker switches, thumb wheels, infrared ports, USB ports, and / or pointer devices such as a stylus. The one or more buttons (not shown) can include up / down buttons for volume control of the speaker 828 and / or the microphone 830.
[0066] In some implementations, the device 800 plays back audio files and / or video files recorded by a user, such as MP3, AAC, and MPEG video files. In some implementations, the device 800 can include the functionality of an MP3 player and can include a pin connector or other port for tethering to other devices. Other input / output and control devices can be used. In one embodiment, the device 800 can include an audio processing unit for streaming audio to an accessory device via a direct communication link or an indirect communication link.
[0067] The memory interface 802 can be coupled to the memory 850. The memory 850 can include high-speed random access memory or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, or flash memory (e.g., NAND, NOR). The memory 850 can store an operating system 852, such as an embedded operating system like Darwin, RTXC, LINUX®, UNIX®, OS X, iOS, WINDOWS®, or VxWorks. The operating system 852 can include instructions for handling basic system services and performing hardware-dependent tasks. In some implementations, the operating system 852 can include a kernel (e.g., a UNIX® kernel).
[0068] The memory 850 can also store communication instructions 854 for facilitating communication with one or more additional devices, one or more computers or servers, including peer-to-peer communication with wireless accessory devices, as described with reference to FIGS. 1-6. The communication instructions 854 can also be used to select an operating mode or communication medium for use by the device based on the geographical location of the device.
[0069] The memory 850 can include sensor processing instructions 858 for facilitating sensor-related processing and functions, and recorder instructions 860 for facilitating a recording function, as described with reference to FIGS. 1-6. Other instructions can include GNSS / navigation instructions for facilitating GNSS and navigation-related processing, camera instructions for facilitating camera-related processing, and user interface instructions including a touch model for interpreting touch input to facilitate user interface processing.
[0070] Each of the instructions and applications identified above may correspond to a set of instructions for performing one or more of the functions described above. These instructions need not be implemented as separate software programs, procedures, or modules. Memory 850 may contain additional instructions or fewer instructions. Further, the various functions of the device may be implemented in hardware and / or software, including one or more signal processing and / or application specific integrated circuits (ASICs).
[0071] The described features may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or combinations thereof. The features may be implemented in a computer program product tangibly embodied in an information carrier, e.g., in a machine-readable storage device, for execution by a programmable processor. Method steps may be performed by a programmable processor executing a program of instructions to perform functions of the described implementations by operating on input data and generating output.
[0072] The features described can be advantageously implemented in one or more computer programs executable in a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a particular activity or to cause a particular result. The computer program may be written in any form of programming language including compiled or interpreted languages (e.g., Objective-C, Java (registered trademark)), and may be deployed in any form, including as a stand-alone program, or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0073] Suitable processors for executing the program of the command include, for example, both general-purpose microprocessors and dedicated microprocessors, as well as a single processor of any type of computer, or one of a plurality of processors or cores. Generally, the processor will receive instructions and data from read-only memory or random access memory or both. Essential elements of a computer are a processor for executing instructions, and one or more memories for storing instructions and data. Generally, a computer can communicate with a mass storage device for storing data files. These mass storage devices can include magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Suitable storage devices for tangibly embodying computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and all forms of non-volatile memory including CD-ROM and DVD-ROM disks. The processor and memory can be complemented by, or incorporated within, an ASIC (application specific integrated circuit). To provide interaction with the user, the feature can be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light emitting diode), or LCD (liquid crystal display) display or monitor for displaying information to the creator, a keyboard, and a pointing device such as a mouse or trackball by which the creator can provide input to the computer.
[0074] One or more features or steps of the disclosed embodiments may be implemented using an application programming interface (API). The API may define one or more parameters that are passed between a calling application and other software code (e.g., an operating system, library routine, function) that provides a service that provides data or performs an operation or calculation. The API may be implemented as one or more calls within program code that send or receive one or more parameters via a parameter list or other structure based on a calling convention defined in the API specification. The parameters may be constants, keys, data structures, objects, object classes, variables, data types, pointers, arrays, lists, or another call. The API calls and parameters may be implemented in any programming language. The programming language may define a vocabulary and a calling convention that a programmer will use to access the features supported by the API. In some implementation examples, the API calls may report to the application the capabilities of the device on which the application operates, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, and the like.
[0075] Some implementation examples have been described. Nevertheless, it will be understood that various modifications can be made. The elements of one or more implementation examples may be combined, deleted, modified, or supplemented to form further implementation examples. In yet another example, the logical flow depicted in the figures does not require the particular order, or sequential order, shown to achieve the desired result. Additionally, other steps may be provided, or steps may be excluded from the flow described, or other components may be added to or removed from the system described. Accordingly, other implementation examples are within the scope of the following claims.
Description of the Reference Numerals
[0076] 100 Operating environment 101 Wearable multimedia device 102 Cloud computing platform 103 Network 104 App developer 105 Third-party platform 106 Database 200 Data processing system 201 Recorder 202 Video buffer 203 Audio buffer 204 Photo buffer 205 Absorption server 206 Data store 207 Video processor 208 Audio processor 209 Photo processor 210 Third-party processor 211, 212, 213 Applications 300 Data processing pipeline 301 Absorption server 302 Video processor 304 Recorder 305 Audio processor 306 App 307 Video processing app 308 Data merge processing 400 Data processing pipeline 401 Absorption server 402 Third-party processor 404 Recorder 405 Audio processor 406, 407 Applications 408 User callback waiting queue 700 Architecture 702 Processor(s) 704 Memory device(s) 706 Network interface(s) 708 Computer-readable medium(s) 710 Communication channel(s) 712 Operating system 714 Network communication module 716 Data processing instruction 718 Interface instruction 800 Device 802 Memory interface 804 Processor(s) 806 Peripheral device interface 810 Motion sensor(s) 812 Biometric sensor(s) 814 Depth sensor 815 Position processor 816 Electromagnetic compass 820 Camera subsystem 822 Optical sensor 824 Wireless communication subsystem 826 Audio subsystem 828 Speaker 830 Microphone 840 Subsystem 842 Touch controller 844 Other input controller(s) 846 Touch surface 846 Touch surface 848 Control device 850 Memory 852 Operating system instruction 854 Communication instruction 858 Sensor processing instruction 860 Recorder instruction
Claims
1. Receiving, using one or more processors of a cloud computing platform, two or more data streams, each data stream including a unique identifier and context data captured by a wearable multimedia device in the real world, the context data including one or more digital images and depth data; For each data stream Identifying, using the one or more processors, a real-world object and one or more gestures associated with the real-world object based on the depth data of the context data; Creating, using the one or more processors, a data processing pipeline comprising one or more applications based on one or more characteristics of the context data and the unique identifier; Generating, using the data processing pipeline, a description of the identified real-world object, the description including a label of the real-world object; Transmitting, using the one or more processors, the description to the wearable multimedia device or another device; A method comprising.
2. The context data further includes audio, and the method includes Determining that the audio includes a user request in the form of speech; Converting the speech to text; Identifying the real-world object using at least a portion of the text; The method according to claim 1, further comprising.
3. The context data further includes audio, and the method includes Determining that the audio includes a user request in the form of speech; The step of converting the speech into text; The step of sending the text to a transportation service processor; The step of receiving a transportation status and an explanation of a vehicle in response to the text; The step of sending the transportation status and the explanation of the vehicle to the wearable multimedia device or other device; Further comprising: The method according to claim 1.
4. The one or more applications include a virtual reality (VR) or augmented reality (AR) application, and the method includes: The step of generating VR or AR content by using at least one of the one or more digital images or the depth data by the VR or AR application; The step of sending the VR or AR content to the wearable multimedia device or other device; Further comprising: The method according to claim 1.
5. The one or more applications include an artificial intelligence application, and the method includes: The step of determining a user's preference from a past history of user requests by using the artificial intelligence application; The step of processing the context data according to the user's preference by using the one or more processors, further comprising: The method according to claim 1.
6. The one or more applications include a location identification application, and the method includes: The step of determining the location of the wearable multimedia device based on at least one of the one or more digital images or the depth data by using the location identification application; The step of transmitting the position to the wearable multimedia device or other device, further comprising The method according to claim 1. **Claim 7** The context data includes financial transaction data of a financial transaction, biometric data, and position data indicating the position of the financial transaction, the one or more applications include a financial application, and the method Using the financial application to create a financial record of the financial transaction based on the financial transaction data, biometric data, and position data, and The step of transmitting the financial record to the wearable multimedia device or other device, further comprising the method according to claim 1. **Claim 8** The context data includes environmental sensor data, the one or more applications include an environmental application, and the method Using the environmental application to generate content associated with the operating environment of the wearable multimedia device based on the environmental sensor data, and The step of transmitting the content to the wearable multimedia device or other device, further comprising the method according to claim 1. **Claim 9** The context data includes video and audio, the one or more applications include a video editing application, and the method Editing the video and audio according to user requirements or user preferences for a specific movie style, and The step of transmitting the edited video and audio to the wearable multimedia device or other device, further comprising The method according to claim 1. **Claim 10** A system, comprising One or more processors, and including a memory for storing instructions, when the instructions are executed by the one or more processors, causing the one or more processors to receive two or more data streams, each data stream including a unique identifier and context data captured by a wearable multimedia device in the real world, the context data including one or more digital images and depth data, receiving; for each data stream identifying, based on the depth data of the context data, a real-world object and one or more gestures associated with the real-world object; creating a data processing pipeline with one or more applications based on one or more characteristics of the context data and the unique identifier; using the data processing pipeline to generate a description of the identified real-world object, the description including a label of the real-world object, generating; transmitting the description to the wearable multimedia device or another device; performing operations including a system.
11. the context data further includes audio, and the operations include determining that the audio includes a user request in the form of speech; converting the speech to text; further including using at least a portion of the text to identify the real-world object, the system according to claim 10.
12. the context data further includes audio, and the operations include determining that the audio includes a user request in the form of speech; converting the speech to text; Sending the text to a transport service processor; Receiving a transport status and an explanation of the vehicle in response to the text; Further comprising sending the transport status and the explanation of the vehicle to the wearable multimedia device or another device. The system according to claim 10.
13. The one or more applications include a virtual reality (VR) or augmented reality (AR) application, and the operations include: Generating VR or AR content based on at least one of the one or more digital images or the depth data by the VR or AR application; Sending the VR or AR content to the wearable multimedia device or another device. Further comprising: The system according to claim 10.
14. The one or more applications include an artificial intelligence application, and the operations include: Determining the user's preferences from a past history of user requests using the artificial intelligence application; Further comprising processing the context data according to the user's preferences by the one or more processors. The system according to claim 10.
15. The one or more applications include a location identification application, and the operations include: Determining the location of the wearable multimedia device based on at least one of the one or more digital images or the depth data using the location identification application; Further comprising sending the location to the wearable multimedia device or another device. The system according to claim 10. **Claim 16** wherein the context data includes financial transaction data of a financial transaction, biometric data, and location data indicating the location of the financial transaction, the one or more applications include a financial application, and the operation creating, by the financial application, a financial record of the financial transaction based on the financial transaction data, biometric data, and location data; and transmitting the financial record to the wearable multimedia device or another device, the system according to claim 10. **Claim 17** wherein the context data includes environmental sensor data, the one or more applications include an environmental application, and the operation generating, by the environmental application, content associated with an operating environment of the wearable multimedia device based on the environmental sensor data; and transmitting the content to the wearable multimedia device or another device, the system according to claim 10. **Claim 18** The system according to claim 10, wherein the system is a distributed streaming platform. **Claim 19** wherein the context data includes video and audio, the one or more applications include a video editing application, and the operation editing the video and audio according to a user request or a user preference for a specific movie style; and transmitting the edited video and audio to the wearable multimedia device or another device, the system according to claim 10. **Claim 20** A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to receive two or more data streams, each data stream including a unique identifier and context data captured by a wearable multimedia device in the real world, the context data including one or more digital images and depth data; for each data stream identify a real-world object and one or more gestures associated with the real-world object based on the depth data of the context data; create a data processing pipeline with one or more applications based on one or more characteristics of the context data and the unique identifier; generate a description of the identified real-world object using the data processing pipeline, the description including a label of the real-world object; send the description to the wearable multimedia device or another device; perform operations including A non-transitory computer-readable storage medium. **Claim 21** Capturing a first set of digital images using a camera of a wearable multimedia device, the first set of digital images including two or more objects; Capturing first depth data using a depth sensor of the wearable multimedia device, the first depth data indicating a first gesture by a user wearing the wearable multimedia device; Based on the first gesture, separating a part of at least one digital image within the first set of digital images, wherein the separated part includes one of the two or more objects, the separating step; Using the wireless transceiver of the wearable multimedia device to transmit the separated part to a network-based data processing pipeline; Using the wireless transceiver to receive first data from the network-based data processing pipeline, wherein the first data is associated with the object and the first data includes a label of the object, the receiving step; Using the laser projection system of the wearable multimedia device to project a transient display onto a surface, wherein the transient display includes at least a part of the first data, the projecting step; A method comprising. Claim 22 Using the camera and depth sensor to receive a second set of digital images and second depth data, wherein the second depth data indicates a second gesture of the user associated with the transient display, the receiving step; Using the wireless transceiver to transmit the second set of digital images to the network-based data processing pipeline; Using the wireless transceiver to receive second data in response to the second gesture from the network-based data processing pipeline; Using the laser projection system to project at least a part of the second data within the transient display onto the surface, further comprising. The method according to claim 21. Claim 23 Using one or more microphones of the wearable multimedia device to capture audio including speech; Using the wireless transceiver to transmit the utterance to the network-based data processing pipeline; Using the wireless transceiver to receive second data in response to the utterance from the network-based data processing pipeline; and The method according to claim 21.
24. Using the wearable multimedia device's global navigation satellite system receiver or the wireless transceiver to obtain the geographical location of the wearable multimedia device; Using the wireless transceiver to transmit the geographical location of the wearable multimedia device to the network-based data processing pipeline; Using the wireless transceiver to receive second data in response to a second gesture and the geographical location from the network-based data processing pipeline; and The method according to claim 21.
25. A camera; A laser projection system; A wireless transceiver; One or more processors; A memory storing instructions, wherein when the instructions are executed by the one or more processors, the one or more processors are caused to Use the camera to capture a first set of digital images, the first set of digital images including two or more real-world objects; Separate a first portion of at least one digital image within the first set of digital images, the separated first portion including a first object among the two or more real-world objects; Separating a second portion of at least one digital image within the first set of digital images, wherein the separated second portion includes the second object among the two or more real-world objects; Transmitting, using the wireless transceiver, a first data stream having a first unique identifier to a network-based data processing pipeline; Transmitting, using the wireless transceiver, a second data stream having a second unique identifier to the network-based data processing pipeline; Receiving, using the wireless transceiver, first data from the network-based data processing pipeline, wherein the first data is associated with the first object; Receiving, using the wireless transceiver, second data from the network-based data processing pipeline, wherein the second data is associated with the second object; Projecting, using the laser projection system, a transient display including at least a portion of the first data or the second data onto a surface; An apparatus.
26. The operations further include: Capturing, using the camera, a second set of digital images including gestures that interact with the transient display; Transmitting, using the wireless transceiver, the second set of digital images to the network-based data processing pipeline; Receiving, using the wireless transceiver, third data responsive to the gestures from the network-based data processing pipeline; Projecting, using the laser projection system, at least a portion of the third data within the transient display onto the surface. The apparatus according to claim 25. **Claim 27** Further comprising one or more microphones of the apparatus, wherein the operations comprise capturing an audio input using the one or more microphones; transmitting the audio input to the network-based data processing pipeline using the wireless transceiver; and receiving, using the wireless transceiver, third data in response to the audio input from the network-based data processing pipeline. The apparatus according to claim 25. **Claim 28** Further comprising a global navigation satellite system receiver, wherein the operations comprise obtaining a geographical location of the apparatus using the global navigation satellite system receiver or the wireless transceiver; transmitting the geographical location of the apparatus to the network-based data processing pipeline using the wireless transceiver; and receiving, using the wireless transceiver, third data associated with the geographical location from the network-based data processing pipeline. The apparatus according to claim 25. **Claim 29** Further comprising a magnetic attachment mechanism for attaching the apparatus to a user's clothing. The apparatus according to claim 25. **Claim 30** Further comprising an inductive charging circuit for inductive charging or wireless over-air charging. The apparatus according to claim 25. **Claim 31** An apparatus comprising: an attachment mechanism for attaching the apparatus to a user's clothing; a camera; a depth sensor; a laser projection system; one or more processors; It includes a memory for storing instructions, and when the instructions are executed by the one or more processors, the one or more processors are caused to use the camera to capture a set of digital images; use the depth sensor to capture depth data; use the set of digital images and the depth data to identify one or more real-world objects, where identifying the one or more real-world objects includes determining a label for each of the one or more real-world objects; obtain data associated with the identified one or more real-world objects; use the laser projection system to project the associated data onto a surface, and perform operations including this. An apparatus.
32. A screenless apparatus, having an attachment mechanism for attaching the apparatus to a user's clothing, a camera, a laser projection system, one or more processors, and a memory for storing instructions, and when the instructions are executed by the one or more processors, the one or more processors are caused to use the camera to capture a set of digital images; use the one or more processors to label one or more real-world objects within the set of digital images; use the one or more processors to obtain data associated with the one or more real-world objects based on at least a portion of the labeling; use the laser projection system to project the data onto a surface, and perform operations including this. An apparatus.
33. The wearable multimedia device includes a magnet attachment mechanism for attaching the wearable multimedia device to a user's clothing. The method according to claim 21.
34. Further comprising an inductive charging circuit for inductive charging or wireless over-air charging. The apparatus according to claim 32.
35. The method according to claim 21, wherein the first data is associated with the user's personal ontology.
36. The method according to claim 35, wherein one of the two or more real-world objects is a person, and the first data includes the name of the person retrieved from the ontology.
37. The method according to claim 21, wherein the wearable multimedia device does not include a screen for viewing the set of digital images.
38. The method according to claim 23, wherein the audio including the two or more objects and utterances is transmitted to the network-based data processing pipeline as separate data streams, each data stream including a unique identifier.
39. The method according to claim 21, wherein the first data includes location data and a description of the two or more objects.
40. The apparatus according to claim 25, wherein the first data is associated with the user's personal ontology.
41. The apparatus according to claim 40, wherein one of the one or more real-world objects is a person, and the first data includes the name of the person retrieved from the ontology.
42. The apparatus according to claim 25, wherein the apparatus does not include a screen for viewing the set of digital images.
Citation Information
Patent Citations
Taxi allocation operation system and allocation method, and storage medium with allocation program stored therein
JP2002133588A
Telephone receiving system, telephone receiving method, program, and recording medium
JP2009239466A
Information processing method, information processor, scenery metadata extraction device, lack complementary information generating device and program
JP2011239141A
Display system, display device, display device control method, and program
JP2017016056A
Intelligent Automated Assistant
US20120245944A1