Wearable multimedia device with laser projection system and cloud computing platform
By integrating laser projection systems and depth sensors in wearable multimedia devices, the problem of difficulty in capturing important moments is solved, and the function of automatically capturing and editing multimedia data is realized, making it possible to playback of multiple devices.
Patent Information
- Application Number
- CN202411731034.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-17
- Filing Date
- 2020-06-18
- Publication Date
- 2025-05-13
AI Technical Summary
Existing mobile embedded cameras have difficulty capturing some important moments because these moments happen too quickly or the user forgets to shoot.
A wearable multimedia device with a laser projection system is designed, equipped with a camera, depth sensor and processor, which can automatically capture multimedia data and project data on the surface through a laser projection system, providing the object's text label, dimension template or instructions to perform actions.
Multimedia data of spontaneous moments and transactions is realized with minimal user interaction, and automatically editing and formatting of these data through a cloud computing platform, making it possible to playback of various user playback devices.
Smart Images

Figure CN119996797A_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 202080057823.1, filed on June 18, 2020, and entitled “Wearable Multimedia Device and Cloud Computing Platform with Laser Projection System”. Technical Field
[0002] The present disclosure relates generally to cloud computing and multimedia editing. Background Art
[0003] Modern mobile devices (e.g., smart phones, tablet computers) often include embedded cameras that allow users to capture digital images or videos of spontaneous events. These digital images and videos can be stored in an online database associated with a user account to free up memory on the mobile device. Users can share their images and videos with friends and family and download or stream images and videos on demand using their various playback devices. These embedded cameras offer significant advantages over conventional digital cameras that are bulky and typically require more time to set up the shot.
[0004] Despite the convenience of mobile device embedded cameras, there are still many important moments that are not captured by these devices because they happen too quickly or the user simply forgets to take an image or video because of the excitement at the time. Summary of the invention
[0005] Disclosed are systems, methods, devices, and non-transitory computer-readable storage media for a wearable multimedia device and a cloud computing platform having an application ecosystem for processing multimedia data captured by the wearable multimedia device.
[0006] In an embodiment, a body-worn device includes: a camera; a depth sensor; a laser projection system; one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: capturing a set of digital images using a camera; identifying objects in the set of digital images; capturing depth data using a depth sensor; identifying a posture of a user wearing the device in the depth data; associating an object with a posture; obtaining data associated with the object; and laser projection of projecting data onto a surface using a laser projection system.
[0007] In an embodiment, the laser projection includes a text label of the object.
[0008] In an embodiment, the laser projection comprises a size template of the object.
[0009] In an embodiment, the laser projection includes instructions for performing an action on the object.
[0010] In an embodiment, a body-worn device includes: a camera; a depth sensor; a laser projection system; one or more processors; a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: capturing depth data using a sensor; identifying a first gesture in the depth data, the gesture being made by a user wearing the device; associating the first gesture with a request or command; and projecting a laser projection on a surface using the laser projection system, the laser projection being associated with the request or command.
[0011] In an embodiment, the operations further include: obtaining, using the depth sensor, a second gesture associated with the laser projection; determining a user input based on the second gesture; and initiating one or more actions based on the user input.
[0012] In an embodiment, the operations further include shielding the laser projection to prevent data from being projected onto a hand of the user making the second gesture.
[0013] In an embodiment, the operations further include: obtaining depth or image data indicative of the geometry, material, or texture of the surface using a depth sensor or camera; and adjusting one or more parameters of the laser projection system based on the geometry, material, or texture of the surface.
[0014] In an embodiment, the operations further include: capturing a reflection of a laser projection from the surface using a camera; and automatically adjusting an intensity of the laser projection to compensate for different refractive indices so that the laser projection has uniform brightness.
[0015] In an embodiment, the device comprises: a magnetic attachment mechanism configured to magnetically couple to a battery pack through a user's clothing, the magnetic attachment mechanism further configured to receive an inductive charge from the battery pack.
[0016] In an embodiment, a method includes: capturing depth data using a depth sensor of a body-worn device; identifying a first gesture in the depth data using one or more processors of the device, the first gesture being made by a user wearing the device; associating the first gesture with a request or command using the one or more processors; and projecting a laser projection on a surface using a laser projection system of the device, the laser projection being associated with the request or command.
[0017] In an embodiment, the method further comprises: obtaining a second posture of the user using a depth sensor, the second posture being associated with the laser projection; determining a user input based on the second posture; and initiating one or more actions according to the user input.
[0018] In an embodiment, the one or more actions include controlling another device.
[0019] In an embodiment, the method further comprises shielding the laser projection to prevent data from being projected onto a hand of the user making the second gesture.
[0020] In an embodiment, the method further comprises obtaining depth or image data indicating the geometry, material or texture of the surface using a depth sensor or camera; and adjusting one or more parameters of the laser projection system based on the geometry, material or texture of the surface.
[0021] In an embodiment, a method includes: receiving, by one or more processors of a cloud computing platform, context data from a wearable multimedia device, the wearable multimedia device including at least one data capture device for capturing context data; creating, by the one or more processors, a data processing pipeline having one or more applications based on one or more characteristics of the context data and a user request; processing, by the one or more processors, the context data through the data processing pipeline; and sending, by the one or more processors, an output of the data processing pipeline to the wearable multimedia device or other device for presenting the output.
[0022] In an embodiment, a system includes: one or more processors; a memory storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations including the following: receiving context data from a wearable multimedia device by one or more processors of a cloud computing platform, the wearable multimedia device including at least one data capture device for capturing context data; creating a data processing pipeline with one or more applications based on one or more characteristics of the context data and a user request by the one or more processors; processing the context data through the data processing pipeline by the one or more processors; and sending the output of the data processing pipeline to the wearable multimedia device or other device for presenting the output by the one or more processors.
[0023] In an embodiment, a non-transitory computer-readable storage medium includes instructions for: receiving context data from a wearable multimedia device by one or more processors of a cloud computing platform, the wearable multimedia device including at least one data capture device for capturing context data; creating a data processing pipeline with one or more applications based on one or more characteristics of the context data and a user request by the one or more processors; processing the context data through the data processing pipeline by the one or more processors; and sending the output of the data processing pipeline to the wearable multimedia device or other device for presenting the output by the one or more processors.
[0024] In an embodiment, a method includes: receiving, by a controller of a wearable multimedia device, depth or image data indicating a surface geometry, material, or texture, the depth or image data being provided by one or more sensors of the wearable multimedia device; adjusting, by the controller, one or more parameters of a projector of the wearable multimedia device based on the surface geometry, material, or texture; projecting, by the projector of the wearable multimedia device, text or image data onto the surface; receiving, by the controller, depth or image data indicating a user interaction with the text or image data projected onto the surface from the one or more sensors; determining, by the controller, a user input based on the user interaction; and initiating, by a processor of the wearable multimedia device, one or more actions based on the user input.
[0025] In an embodiment, a wearable multimedia device comprises: one or more sensors; a projector; and a controller configured to: receive depth or image data from the one or more sensors, the depth or image data indicating a surface geometry, material, or texture, the depth or image data being provided by the one or more sensors of the wearable multimedia device; adjust one or more parameters of the projector based on the surface geometry, material, or texture; project text or image data onto a surface using the projector; receive depth or image data from the one or more sensors indicating user interaction with the text or image data projected onto the surface; and determine user input based on the user interaction; and initiate one or more actions based on the user input.
[0026] Certain embodiments disclosed herein provide one or more of the following advantages. The wearable multimedia device captures multimedia data of spontaneous moments and transactions with minimal user interaction. The multimedia data is automatically edited and formatted on a cloud computing platform based on user preferences and then available for user replay on various user playback devices. In an embodiment, data editing and / or processing is performed by an application ecosystem that is proprietary and / or provided / licensed by third-party developers. The application ecosystem provides various access points (e.g., websites, portals, APIs) that allow third-party developers to upload, verify, and update their applications. The cloud computing platform automatically builds a custom processing pipeline for each multimedia data stream using one or more of the ecosystem applications, user preferences, and other information (e.g., the type or format of the data, the quantity and quality of the data).
[0027] In addition, the wearable multimedia device includes a camera and a depth sensor that can detect the air gestures of objects and users, and then perform or infer various actions based on the detection, such as marking objects in camera images or controlling other devices. In an embodiment, the wearable multimedia device does not include a display, thereby allowing the user to continue to interact with friends, family, and colleagues without being immersed in a display, which is a current problem for smartphone and tablet computer users. As a result, the wearable multimedia device adopts a technical approach different from, for example, smart goggles or glasses for augmented reality (AR) and virtual reality (VR) (where the user is further separated from the real world environment). In order to promote collaboration with others and make up for the absence of a display, the wearable multimedia computer includes a laser projection system that projects laser projections onto any surface, including tables, walls, and even the palm of the user. The laser projection can mark objects, provide text or instructions related to the object, and provide a temporary user interface (e.g., keyboard, numeric keypad, device controller) that allows the user to write messages, control other devices, or simply share and discuss content with others.
[0028] The details of the disclosed embodiments are set forth in the accompanying drawings and the description that follows. Other features, objects, and advantages are apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a block diagram of an operating environment of a wearable multimedia device and a cloud computing platform having an application ecosystem for processing multimedia data captured by the wearable multimedia device according to an embodiment.
[0030] Figure 2 According to the embodiment, Figure 1 Block diagram of the data processing system implemented on the cloud computing platform.
[0031] Figure 3 is a block diagram of a data processing pipeline for processing a context data stream, according to an embodiment.
[0032] Figure 4 is a block diagram of another data process for processing a context data flow for a transportation application according to an embodiment.
[0033] Figure 5 The diagram shows a method according to an embodiment of the present invention. Figure 2 Data objects used by a data processing system.
[0034] Figure 6 is a flow chart of a data pipeline process according to an embodiment.
[0035] Figure 7 The present invention is an architecture of a cloud computing platform according to an embodiment.
[0036] Figure 8 is a system architecture of a wearable multimedia device according to an embodiment.
[0037] Fig. 9 is for reference according to the embodiment Figure 3 Screen shot of an example graphical user interface (GUI) for the described scene recognition application.
[0038] Fig.10 Illustrated is a method for classifying raw or pre-processed context data into Fig. 9 A GUI for searching objects and metadata in the classifier framework.
[0039] Fig.11 is a system block diagram illustrating a hardware architecture of a wearable multimedia device according to an embodiment.
[0040] Fig.12 is a system block diagram illustrating a processing framework implemented in a cloud computing platform for processing raw or pre-processed context data received from a wearable multimedia device according to an embodiment.
[0041] Fig.13 Software components for a wearable multimedia device according to an embodiment are illustrated.
[0042] Figures 14A-14D It is illustrated that various types of information are projected on the palm of a user using a projector of a wearable multimedia device according to an embodiment.
[0043] Fig.15A and 15B An application of a projector according to an embodiment is illustrated, where information is projected onto a car engine to help a user check his engine oil.
[0044] Fig.16 An application of a projector according to an embodiment is illustrated, wherein information for assisting a home cook in chopping vegetables is projected onto a cutting board.
[0045] Fig.17 is a system block diagram of a projector architecture according to an embodiment.
[0046] Fig.18 Laser parameter adjustments based on different surface geometries or materials according to an embodiment are illustrated.
[0047] The same reference numbers used in different drawings denote the same elements. DETAILED DESCRIPTION
[0048] Overview
[0049] Wearable multimedia devices are lightweight, small form factor, battery-powered devices that can be attached to a user's clothing or object using tension buckles, interlocking pin backs, magnets, or any other attachment mechanism. Wearable multimedia devices include digital image capture devices (e.g., 180° FOV with optical image stabilizer (OIS)) that allow users to spontaneously capture multimedia data (e.g., video, audio, depth data) of life events ("moments") and record transactions (e.g., financial transactions) with minimal user interaction or device setup. The multimedia data ("context data") captured by the wireless multimedia device is uploaded to a cloud computing platform with an application ecosystem that allows the context data to be processed, edited, and formatted by one or more applications (e.g., artificial intelligence (AI) applications) into any desired presentation format (e.g., single image, image stream, video clip, audio clip, multimedia presentation, image library) that can be downloaded and replayed on the wearable multimedia device and / or any other playback device. For example, the cloud computing platform can transform the video data and audio data into any desired filmmaking style specified by the user (e.g., documentary, lifestyle, candid, photojournalism, sports, street).
[0050] In an embodiment, the contextual data is processed by (one or more) server computers of the cloud computing platform based on user preferences. For example, an image may be perfectly color graded, stabilized, and cropped to a moment the user wants to relive based on the user preferences. The user preferences may be stored in a user profile created by the user through an online account accessible through a website or portal, or the user preferences may be learned by the platform over time (e.g., using machine learning). In an embodiment, the cloud computing platform is a scalable distributed computing environment. For example, the cloud computing platform may be a distributed streaming platform (e.g., Apache Kafka) with a real-time streaming data pipeline and streaming applications that transform or react to the data streams. TM ).
[0051] In an embodiment, a user may start and stop a contextual data capture session on a wearable multimedia device by speaking a command or any other input mechanism through a simple touch gesture (e.g., tap or swipe). When the wearable multimedia device detects that it is not being worn by a user using one or more sensors (e.g., proximity sensor, optical sensor, accelerometer, gyroscope), all or part of the wearable multimedia device may be automatically powered off.
[0052] The context data may be encrypted and compressed using any desired encryption or compression technique and stored in an online database associated with the user account. The context data may be stored for a specified period of time that may be set by the user. Opt-in mechanisms and other tools may be provided to users through a website, portal, or mobile application to manage their data and data privacy.
[0053] In an embodiment, the context data includes point cloud data to provide a three-dimensional (3D) surface mapping object that can be processed using, for example, augmented reality (AR) and virtual reality (VR) applications in an application ecosystem. The point cloud data may be generated by a depth sensor (e.g., LiDAR or time of flight (TOF)) embedded on the wearable multimedia device.
[0054] In an embodiment, the wearable multimedia device includes a global navigation satellite system (GNSS) receiver (e.g., a global positioning system (GPS)) and one or more inertial sensors (e.g., an accelerometer, a gyroscope) for determining the location and orientation of a user wearing the device when capturing contextual data. In an embodiment, one or more images in the contextual data can be used by a positioning application such as a visual odometry application in an application ecosystem to determine the user's position and orientation.
[0055] In an embodiment, the wearable multimedia device may also include one or more environmental sensors, including but not limited to: ambient light sensors, magnetometers, pressure sensors, voice activity detectors, etc. This sensor data may be included in the contextual data to enrich the content presentation with additional information that may be used to capture the moment.
[0056] In an embodiment, the wearable multimedia device may include one or more biometric sensors, such as a heart rate sensor, a fingerprint scanner, etc. Such sensor data may be included in the contextual data to record a transaction or indicate the user's emotional state during that moment (e.g., an elevated heart rate may indicate excitement or fear).
[0057] In an embodiment, the wearable multimedia device includes a headphone jack for connecting headphones or earbuds, and one or more microphones for receiving voice commands and capturing ambient audio. In an alternative embodiment, the wearable multimedia device includes short-range communication technology, including but not limited to Bluetooth, IEEE 802.15.4 (ZigBee TM ) and Near Field Communication (NFC). In addition to or in lieu of the headphone jack, short-range communication technology can be used to wirelessly connect to wireless headphones or earbuds, and / or can wirelessly connect to any other external device (e.g., computers, printers, projectors, televisions, and other wearable devices).
[0058] In an embodiment, the wearable multimedia device includes a wireless transceiver and a communication protocol stack for a variety of communication technologies including WiFi, 3G, 4G, and 5G communication technologies. In an embodiment, the headset or earbuds also include sensors (e.g., biometric sensors, inertial sensors) that provide information about the direction the user is facing to provide commands with head gestures, etc. In an embodiment, the camera direction can be controlled by the head gesture so that the camera view follows the user's line of sight. In an embodiment, the wearable multimedia device can be embedded or attached to the user's glasses.
[0059] In an embodiment, the wearable multimedia device includes a projector (e.g., laser projector, LCoS, DLP, LCD), or can be wired or wirelessly coupled to an external projector that allows the user to replay the moment on a surface such as a wall or desktop. In another embodiment, the wearable multimedia device includes an output port that can be connected to a projector or other output device.
[0060] In an embodiment, the wearable multimedia capture device includes a touch surface that responds to touch gestures (e.g., taps, multi-tap, or swipe gestures). The wearable multimedia device may include a small display for presenting information and one or more indicator lights to indicate on / off status, power status, or any other desired status.
[0061] In an embodiment, the cloud computing platform can be driven by context-based gestures (e.g., air gestures) in conjunction with voice queries, such as a user pointing to an object in their environment and saying, "What building is that?" The cloud computing platform uses the air gesture to narrow the camera viewport and isolate the building. One or more images of the building are captured and sent to the cloud computing platform, where the image recognition application can run the image query and store the results or return the results to the user. Air and touch gestures can also be performed on a projected temporary display, such as in response to a user interface element.
[0062] In an embodiment, the contextual data may be encrypted on the device and cloud computing platform so that only the user or any authorized audience can relive the moment on a connected screen (e.g., smartphone, computer, TV, etc.) or projection on a surface. Figure 8 An example architecture for a wearable multimedia device is described.
[0063] In addition to personal life events, wearable multimedia devices also simplify the capture of financial transactions currently handled by smart phones. By using visually assisted contextual awareness provided by wearable multimedia devices, the capture of daily transactions (e.g., business transactions, micro-transactions) becomes simpler, faster, and more streamlined. For example, when a user conducts a financial transaction (e.g., makes a purchase), the wearable multimedia device will generate data that remembers the financial transaction, including the date, time, amount, digital images or videos of the parties, audio (e.g., user comments describing the transaction), and environmental data (e.g., location data). The data can be included in a multimedia data stream sent to a cloud computing platform, where the data can be stored online and / or processed by one or more financial applications (e.g., financial management, accounting, budgeting, tax preparation, inventory, etc.).
[0064] In an embodiment, the cloud computing platform provides a graphical user interface on a website or portal that allows various third-party application developers to upload, update, and manage their applications in the application ecosystem. Some example applications may include, but are not limited to: personal live broadcast (e.g., Instagram TM Life, Snapchat TM ), elder monitoring (e.g., to make sure a loved one has taken their medication), memory review (e.g., showing a child’s soccer game last week), and personal guidance (e.g., an AI-enabled personal guide that knows the user’s location and guides the user through actions).
[0065] In an embodiment, the wearable multimedia device includes one or more microphones and headphones. In some embodiments, the headphone cord includes a microphone. In an embodiment, a digital assistant is implemented on a wearable multimedia device that responds to user queries, requests, and commands. For example, a wearable multimedia device worn by a parent captures moment context data of a child's soccer game, particularly the "moment" when the child scores a goal. The user can request (e.g., using a voice command) that the platform create a video clip of the goal and store it in their user account. Without any further action by the user, the cloud computing platform identifies the correct portion of the moment context data when a goal is scored (e.g., using facial recognition, visual or audio cues), compiles the moment context data into a video clip, and stores the video clip in a database associated with the user account.
[0066] In an embodiment, the device may include photovoltaic surface technology to maintain battery life and inductive charging circuitry (eg, Qi) to allow for inductive charging on a charging pad as well as wireless over-the-air (OTA) charging.
[0067] In an embodiment, the wearable multimedia device is configured to magnetically couple or mate with a rechargeable portable battery pack. The portable battery pack includes a mating surface on which a permanent magnet (e.g., an N pole) is disposed, and the wearable multimedia device has a corresponding mating surface on which a permanent magnet (e.g., an S pole) is disposed. Any number of permanent magnets having any desired shape or size can be arranged on the mating surface in any desired pattern.
[0068] The permanent magnets hold the portable battery pack and the wearable multimedia device together in a paired configuration with clothing (e.g., a user's shirt) located between them. In an embodiment, the portable battery pack and the wearable multimedia device have the same mating surface dimensions so that there are no overhanging portions when in the mated configuration. The user magnetically fastens the wearable multimedia device to their clothing by placing the portable battery pack under their clothing and placing the wearable multimedia device on top of the portable battery pack outside their clothing so that the permanent magnets attract each other through the clothing. In an embodiment, the portable battery pack has a built-in wireless power transmitter that is used to wirelessly power the wearable multimedia device in a paired configuration using the principles of resonant inductive coupling. In an embodiment, the wearable multimedia device includes a built-in wireless power receiver that is used to receive power from the portable battery pack in the mated configuration.
[0069] Example operating environment
[0070] Figure 1 1 is a block diagram of an operating environment of a wearable multimedia device and a cloud computing platform having an application ecosystem for processing multimedia data captured by the wearable multimedia device according to an embodiment. The operating environment 100 includes a wearable multimedia device 101, a cloud computing platform 102, a network 103, an application ("app") developer 104, and a third-party platform 105. The cloud computing platform 102 is coupled to one or more databases 106 for storing context data uploaded by the wearable multimedia device 101.
[0071] As previously described, the wearable multimedia device 101 is a lightweight, small form factor, battery-powered device that can be attached to a user's clothing or object using a tension buckle, interlocking pin back, magnets, or any other attachment mechanism. The wearable multimedia device 101 includes a digital image capture device (e.g., 180° FOV with OIS) that allows the user to spontaneously capture multimedia data (e.g., video, audio, depth data) of the "moment" and record daily transactions (e.g., financial transactions) with minimal user interaction or device setup. The contextual data captured by the wireless multimedia device 101 is uploaded to the cloud computing platform 102. The cloud computing platform 101 includes an application ecosystem that allows the contextual data to be processed, edited, and formatted by one or more server-side applications into any desired presentation format (e.g., a single image, image stream, video clip, audio clip, multimedia presentation, photo gallery) that can be downloaded and replayed on the wearable multimedia device and / or other playback devices.
[0072] As an example, at a child's birthday party, a parent can clip the wearable multimedia device to their clothing (or attach the device to a necklace or chain and wear it around the neck) so that the camera lens is facing the direction of their line of sight. The camera includes a 180° FOV, which allows the camera to capture almost everything the user is currently seeing. The user simply taps the surface of the device or presses a button to start recording. No additional setup is required. A multimedia data stream (e.g., video with audio) is recorded that captures the special moments of the birthday (e.g., blowing out the candles). This "context data" is sent to the cloud computing platform 102 in real time via a wireless network (e.g., WiFi, cellular). In an embodiment, the context data is stored on the wearable multimedia device so that it can be uploaded later. In another embodiment, the user can transfer the context data to another device (e.g., a personal computer hard drive, a smart phone, a tablet computer, a thumb drive) and later upload the context data to the cloud computing platform 102 using an application.
[0073] In an embodiment, the context data is processed by one or more applications of an application ecosystem hosted and managed by the cloud computing platform 102. The applications can be accessed through their respective application programming interfaces (APIs). The cloud computing platform 102 creates a custom distributed streaming pipeline to process the context data based on one or more of the data type, data quantity, data quality, user preferences, templates, and / or any other information to generate the desired presentation based on the user preferences. In an embodiment, machine learning techniques can be used to automatically select appropriate applications to be included in the data processing pipeline with or without user preferences. For example, historical user context data stored in a database (e.g., a NoSQL database) can be used to determine user preferences for data processing using any suitable machine learning technique (e.g., deep learning or convolutional neural networks).
[0074] In an embodiment, the application ecosystem may include a third-party platform 105 that processes context data. A secure session is established between the cloud computing platform 102 and the third-party platform 105 to send / receive context data. This design allows third-party application providers to control access to their applications and provide updates. In other embodiments, the applications run on the servers of the cloud computing platform 102 and updates are sent to the cloud computing platform 102. In the latter embodiment, the application developer 104 can use the API provided by the cloud computing platform 102 to upload and update the applications to be included in the application ecosystem.
[0075] Example Data Processing System
[0076] Figure 2 According to the embodiment, Figure 1 The data processing system 200 includes a recorder 201, a video buffer 202, an audio buffer 203, a photo buffer 204, an ingest server 205, a data repository 206, a video processor 207, an audio processor 208, a photo processor 209, and a third-party processor 210.
[0077] A recorder 201 (e.g., a software application) running on the wearable multimedia device records video, audio, and photo data ("context data") captured by the camera and audio subsystems, and stores the data in buffers 202, 203, 204, respectively. The context data is then sent to an ingest server 205 of the cloud computing platform 102 (e.g., using wireless OTA technology). In an embodiment, the data may be sent in separate data streams, each with a unique stream identifier (streamid). A stream is a discrete piece of data that may contain the following example attributes: location (e.g., latitude, longitude), user, audio data, video streams of varying duration, and N photos. The duration of a stream may be from 1 to MAXSTREAM_LEN seconds, in this example, MAXSTREAM_LEN = 20 seconds.
[0078] The ingestion server 205 ingests the streams and creates a stream record in the data repository 206 to store the results of the processors 207-209. In an embodiment, the audio stream is processed first and used to determine the other streams needed. The ingestion server 205 sends the streams to the appropriate processors 207-209 based on the streamid. For example, the video stream is sent to the video processor 207, the audio stream is sent to the audio processor 208 and the photo stream is sent to the photo processor 209. In an embodiment, at least a portion of the data (e.g., image data) collected from the wearable multimedia device is processed into metadata and encrypted so that it can be further processed by a given application and sent back to the wearable multimedia device or other device.
[0079] As previously described, processors 207-209 can run proprietary or third-party applications. For example, video processor 207 can be a video processing server that sends raw video data stored in video buffer 202 to a collection of one or more image processing / editing applications (211, 212) based on user preferences or other information. Processor 207 sends requests to applications 211, 212 and returns the results to ingest server 205. In an embodiment, third-party processor 210 can use its own processor and application to process one or more streams. In another example, audio processor 208 can be an audio processing server that sends voice data stored in audio buffer 203 to voice-to-text converter application 213.
[0080] Example scene recognition application
[0081] Figure 3300 is a block diagram of a data processing pipeline for processing a context data stream according to an embodiment. In this embodiment, a data processing pipeline 300 is created and configured to determine what a user is seeing based on context data captured by a wearable multimedia device worn by the user. An ingestion server 301 receives an audio stream (e.g., including user comments) from an audio buffer 203 of the wearable multimedia device and sends the audio stream to an audio processor 305. The audio processor 305 sends the audio stream to an application 306, which performs speech-to-text conversion and returns the parsed text to the audio processor 305. The audio processor 305 returns the parsed text to the ingestion server 301.
[0082] The video processor 302 receives the parsed text from the ingest server 301 and sends a request to the video processing application 307. The video processing application 307 identifies objects in the video scene and tags the objects using the parsed text. The video processing application 307 sends a response describing the scene (e.g., the tagged objects) to the video processor 302. The video processor then forwards the response to the ingest server 301. The ingest server 301 sends the response to the data merge processing 308, which merges the response with the user's location, position, and map data. The data merge processing 308 returns a response with a description of the scene to the recorder 304 on the wearable multimedia device. For example, the response may include text describing the scene as a child's birthday party, including a map location and a description of the objects in the scene (e.g., identifying the people in the scene). The recorder 304 associates the scene description with the multimedia data stored on the wearable multimedia device (e.g., using a streamid). When the user reviews the data, the data is enriched with the scene description.
[0083] In an embodiment, the data merging process 308 may use more than just location and map data. There may also be a concept of ontology. For example, the facial features of the user's dad captured in an image may be recognized by the cloud computing platform and returned as "Dad" instead of the user's name, and an address such as "555 Main Street, San Francisco, CA" may be returned as "Home". Ontologies may be user-specific and may grow and learn based on the user's input.
[0084] Sample Transportation Application
[0085] Figure 44 is a block diagram of another data processing for processing a contextual data stream for a transportation application according to an embodiment. In this embodiment, a data processing pipeline 400 is created to call a transportation company (e.g., Uber®, Lyft®) for a ride home. An ingestion server 401 receives contextual data from a wearable multimedia device, and an audio stream from an audio buffer 203 is sent to an audio processor 405. The audio processor 405 sends the audio stream to an application 406, which converts speech to text. The parsed text is returned to the audio processor 405, which returns the parsed text to the ingestion server 401 (e.g., a user voice request for transportation). The processed text is sent to a third-party processor 402. The third-party processor 402 sends the user location and token to a third-party application 407 (e.g., Uber® or Lyft TM ® application). In an embodiment, the token is an API and authorization token for brokering requests on behalf of the user. The application 407 returns a response data structure to the third-party processor 402, which is forwarded to the ingestion server 401. The ingestion server 401 checks the ride arrival status (e.g., ETA) in the response data structure and establishes a callback to the user in the user callback queue 408. The ingestion server 401 returns a response with the vehicle description to the recorder 404, which the digital assistant can speak to the user through the speaker on the wearable multimedia device or through the user's headphones or earbuds via a wired or wireless connection.
[0086] Figure 5 The diagram shows a method according to an embodiment of the present invention. Figure 2 The data objects used by the data processing system of the cloud computing platform. The data objects are part of the software component infrastructure instantiated on the cloud computing platform. The "streams" object includes data streamid, deviceid, start, end, lat, lon, attributes, and entities. "streamid" identifies the stream (e.g., video, audio, photo), "deviceid" identifies the wearable multimedia device (e.g., mobile device ID), "start" is the start time of the context data stream, "end" is the end time of the context data stream, "lat" is the latitude of the wearable multimedia device, "lon" is the longitude of the wearable multimedia device, "attributes" include, for example, birthday, facial points, skin color, voice characteristics, address, phone number, etc., and "entities" constitute the ontology. For example, the name "JohnDo" will be mapped to "Dad" or "Brother" depending on the user.
[0087] The "Users" object includes the data userid, deviceid, email, fname, and lname. The userid identifies the user with a unique identifier, the deviceid identifies the wearable device with a unique identifier, email is the user's registered email address, fname is the user's first name, and lname is the user's last name. The "Userdevices" object includes the data userid and deviceid. The "Devices" object includes the data deviceid, started, state, modified, and created. In an embodiment, deviceid is a unique identifier for the device (for example, different from a MAC address). started is the time when the device was first started. state is on / off / sleep. modified is the last modified date reflecting the last state change or operating system (OS) change. created is the time when the device was first turned on.
[0088] The "ProcessingResults" object includes streamid, ai, result, callback, duration, and accuracy. In an embodiment, streamid is each user stream as a universal unique identifier (UUID). For example, a stream starting at 8:00 am to 10:00 am will have id:15h158dhb4, and a stream starting at 10:15 am to 10:18 am will have a UUID associated with the stream. ai is the identifier of the platform application associated with the stream. result is the data sent from the platform application. callback is the callback used (versions can change, so keep track of callbacks in case the platform needs to replay requests). accuracy is a score of how accurate the result set is. In an embodiment, the processing results can be used for multiple tasks, such as 1) notifying the merge server of the complete result set, 2) determining the fastest AI so that the user experience can be enhanced, and 3) determining the most accurate ai. Depending on the use case, speed may be preferred over accuracy, or vice versa.
[0089] The "Entities" object includes the data entityID, userID, entityName, entityType, and entityAttribute. EntityID is the UUID of the entity and entities with multiple entries, where entityID refers to one entity. For example, "Barack Obama" has an entityID of 144, which can be linked to POTUS44 or "Barack Hussein Obama" or "President Obama" in an association table. UserID identifies the user for whom the entity record is created. EntityName is the name that the userID will call the entity. For example, Malia Obama's entityName with an entityID of 144 can be "Dad" or "Daddy". EntityType is a person, place, or thing. EntityAttribute is an array of properties about the entity that are specific to the userID's understanding of that entity. This maps entities together so that, for example, when Malia makes a voice query: "Can you see Dad?", the cloud computing platform can translate the query to BarackHussein Obama and use it to proxy the request to a third party or look up information in the system.
[0090] Example Processing
[0091] Figure 6 is a flow chart of data pipeline processing according to an embodiment. Figure 1-5 The wearable multimedia device 101 and the cloud computing platform 102 are described to implement the process 600 .
[0092] Process 600 may begin by receiving contextual data from a wearable multimedia device ( 601 ). For example, the contextual data may include video, audio, and still images captured by a camera and audio subsystem of the wearable multimedia device.
[0093] Process 600 may continue by creating (e.g., instantiating) a data processing pipeline with applications based on the context data and the user request / preference (602). For example, based on the user request or preference, and also based on the data type (e.g., audio, video, photo), one or more applications may be logically connected to form a data processing pipeline to process the context data into a presentation played on the wearable multimedia device or another device.
[0094] Process 600 may continue by processing the contextual data in a data processing pipeline (603). For example, a user's voice commentary during a moment or transaction may be converted to text, which may then be used to tag objects in a video clip.
[0095] Process 600 may continue by sending the output of the data processing pipeline to the wearable multimedia device and / or other playback device ( 604 ).
[0096] Example cloud computing platform architecture
[0097] Figure 7 is a reference according to the embodiment Figure 1-6 and Fig. 9 102. Other architectures are possible, including architectures with more or fewer components. In some embodiments, the architecture 700 includes one or more processors 702 (e.g., dual-core Intel® Xeon® processors), one or more network interfaces 706, one or more storage devices 704 (e.g., hard disks, optical disks, flash memory), and one or more computer-readable media 708 (e.g., hard disks, optical disks, flash memory, etc.). These components can exchange communications and data through one or more communication channels 710 (e.g., buses), which can utilize various hardware and software to facilitate the transmission of data and control signals between components.
[0098] The term "computer-readable medium" refers to any medium that participates in providing instructions to processor(s) 702 for execution, including but not limited to non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media. Transmission media include but are not limited to coaxial cables, copper wire, and fiber optics.
[0099] The computer-readable medium(s) 708 may also include an operating system 712 (eg, Mac OS® Server, Windows® NT Server, Linux Server), a network communication module 714 , interface instructions 716 , and data processing instructions 718 .
[0100] The operating system 712 may be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system 712 performs basic tasks including, but not limited to: recognizing input from and providing output to devices 702, 704, 706, and 708; tracking and managing files and directories on computer-readable medium(s) 708 (e.g., memory or storage devices); controlling peripheral devices; and managing traffic on one or more communication channels 710. The network communication module 714 includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, etc.) and for using, for example, Apache Kafka. TMThe various components of the distributed streaming platform are created. The data processing instructions 716 include instructions for implementing the Figure 1-6 The interface instructions 718 include server-side or back-end software for implementing the server-side operations described in the reference Figure 1 The described wearable multimedia device 101, third-party application developers 104, and third-party platforms 105 send and receive data to and from a web server and / or portal software.
[0101] The architecture 700 may be included in any computer device including one or more server computers in a local network or a distributed network, each server computer having one or more processing cores. The architecture 700 may be implemented in a parallel processing or peer-to-peer infrastructure or on a single device having one or more processors. The software may include multiple software components or may be a single body of code.
[0102] Example Wearable Multimedia Device Architecture
[0103] Figure 8 It is used to implement reference Figure 1-6 and Fig. 9 800 is a block diagram of an example architecture 800 of a wearable multimedia device that describes features and processes. The architecture 800 may include a memory interface 802, (one or more) data processors, (one or more) image processors or (one or more) central processing units 804, and a peripheral interface 806. The memory interface 802, (one or more) processors 804, or the peripheral interface 806 may be separate components or may be integrated in one or more integrated circuits. One or more communication buses or signal lines may couple the various components.
[0104] Sensors, devices, and subsystems can be coupled to the peripheral interface 806 to facilitate a variety of functions. For example, motion sensor(s) 810, biometric sensor(s) 812, depth sensor(s) 814 can be coupled to the peripheral interface 806 to facilitate motion, orientation, biometric, and depth detection functions. In some embodiments, motion sensor(s) 810 (e.g., accelerometer, rate gyroscope) can be used to detect movement and orientation of the wearable multimedia device.
[0105] Other sensors may also be connected to the peripheral interface 806, such as (one or more) environmental sensors (e.g., temperature sensor, barometer, ambient light) to facilitate environmental sensing functions. For example, biometric sensors can detect fingerprints, facial recognition, heart rate, and other fitness parameters. In an embodiment, a haptic motor (not shown) can be coupled to the peripheral interface, which can provide a vibration pattern as tactile feedback to the user.
[0106] A location processor 815 (e.g., a GNSS receiver chip) can be connected to the peripheral interface 806 to provide geographic reference. An electronic magnetometer 816 (e.g., an integrated circuit chip) can also be connected to the peripheral interface 806 to provide data that can be used to determine the direction of magnetic north. Thus, the electronic magnetometer 816 can be used by an electronic compass application.
[0107] The camera subsystem 820 and optical sensor 822 (e.g., a charge coupled device (CCD) or complementary metal oxide semiconductor (CMOS) optical sensor) can be used to facilitate camera functions, such as recording photos and video clips. In an embodiment, the camera has a 180° FOV and OIS. The depth sensor can include an infrared emitter that projects points in a known pattern onto an object / subject. These points are then photographed by a dedicated infrared camera and analyzed to determine depth data. In an embodiment, a time of flight (TOF) camera can be used to resolve distances based on the known speed of light and measure the time of flight of the light signal between the camera and the object / subject for each point of the image.
[0108] The communication functions may be facilitated by one or more communication subsystems 824. The communication subsystem(s) 824 may include one or more wireless communication subsystems. The wireless communication subsystem 824 may include a radio frequency receiver and transmitter and / or an optical (e.g., infrared) receiver and transmitter. A wired communication system may include a port device (e.g., a universal serial bus (USB) port or some other wired port connection) that may be used to establish a wired connection to other computing devices such as other communication devices, network access devices, personal computers, printers, display screens, or other processing devices capable of receiving or sending data (e.g., a projector).
[0109] The specific design and implementation of the communication subsystem 824 may depend on the communication network(s) or medium(s) over which the device is intended to operate. For example, the device may include a device designed to operate over a Global System for Mobile Communications (GSM) network, a GPRS network, an Enhanced Data GSM Environment (EDGE) network, an IEEE 802.xx communication network (e.g., WiFi, WiMax, ZigBee TM ), 3G, 4G, 4G LTE, Code Division Multiple Access (CDMA) networks, Near Field Communication (NFC), Wi-Fi Direct, and Bluetooth TMThe wireless communication subsystem 824 may include a wireless communication subsystem that operates on a (Bluetooth) network. The wireless communication subsystem 824 may include a hosting protocol so that the device can be configured as a base station for other wireless devices. As another example, the communication subsystem may allow the device to synchronize with a host device using one or more protocols or communication technologies, such as, for example, TCP / IP, HTTP, UDP, ICMP, POP, FTP, IMAP, DCOM, DDE, SOAP, HTTP Live Streaming, MPEG Dash, and any other known communication protocols or technologies.
[0110] The audio subsystem 826 may be coupled to a speaker 828 and one or more microphones 830 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, telephony functions, and beamforming.
[0111] The I / O subsystem 840 may include a touch controller 842 and / or additional input controller(s) 844. The touch controller 842 may be coupled to a touch surface 846. The touch surface 846 and touch controller 842 may, for example, use any of a variety of touch sensitivity technologies (including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies) to detect contact and movement or disconnection thereof, as well as other proximity sensor arrays or other elements for determining one or more contact points with the touch surface 846 to detect contact and movement or disconnection thereof. In one embodiment, the touch surface 846 may display virtual or soft buttons that may be used by a user as an input / output device.
[0112] Other input controller(s) 844 may be coupled to other input / control devices 848, such as one or more buttons, rocker switches, thumb wheels, infrared ports, USB ports, and / or pointing devices such as a stylus. The one or more buttons (not shown) may include up / down buttons for volume control of the speaker 828 and / or microphone 830.
[0113] In some embodiments, the device 800 plays back audio and / or video files recorded by the user, such as MP3, AAC, and MPEG video files. In some embodiments, the device 800 may include the functionality of an MP3 player and may include a pin connector or other ports for tethering with other devices. Other input / output devices and control devices may be used. In an embodiment, the device 800 may include an audio processing unit for transmitting an audio stream to an attached device via a direct or indirect communication link.
[0114] The memory interface 802 may be coupled to a memory 850. The memory 850 may include a high-speed random access memory or non-volatile memory, such as one or more disk storage devices, one or more optical storage devices, or flash memory (e.g., NAND, NOR). The memory 850 may store an operating system 852, such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks. The operating system 852 may include instructions for handling basic system services and for performing hardware-related tasks. In some embodiments, the operating system 852 may include a kernel (e.g., a UNIX kernel).
[0115] The memory 850 may also store communication instructions 854 for facilitating communication with one or more additional devices, one or more computers or servers, including peer-to-peer communication with wireless accessory devices, such as reference Figure 1-6 Communications instructions 854 may also be used to select an operating mode or communications medium to be used by a device based on the geographic location of the device.
[0116] The memory 850 may include sensor processing instructions 858 for facilitating sensor-related processing and functions and recorder instructions 860 for facilitating recording functions, as described in reference to FIG. Figure 1-6 Other instructions may include GNSS / navigation instructions for facilitating GNSS and navigation related processing, camera instructions for facilitating camera related processing, and user interface instructions for facilitating user interface processing, including a touch model for interpreting touch input.
[0117] Each of the instructions and applications identified above may correspond to an instruction set for performing one or more of the functions described above. These instructions need not be implemented as separate software programs, processes, or modules. The memory 850 may include additional instructions or fewer instructions. In addition, the various functions of the device may be implemented in hardware and / or software, including using one or more signal processing and / or application specific integrated circuits (ASICs).
[0118] Example GUI
[0119] Fig. 9 is used according to the embodiment with reference Figure 3900 is a screenshot of an example graphical user interface (GUI) 900 used with the described scene recognition application. The GUI 900 includes a video pane 901, time / location data 902, objects 903, 906a, 906b, 906c, a search button 904, a category menu 905, and thumbnails 907. The GUI 900 can be presented on a user device (e.g., a smart phone, a tablet computer, a wearable device, a desktop computer, a notebook computer) by, for example, a client application or by a web page provided by a web server of the cloud computing platform 102. In this example, the user has captured a digital image of a young person standing on Orchard Street in New York City, New York at 12:45 p.m. on October 18, 2018 in the video pane 901, as indicated by the time / location data 902.
[0120] In an embodiment, the image is processed by an object detection framework (such as a Viola-Jones object detection network) implemented on the cloud computing platform 102. For example, a model or algorithm is used to generate a region of interest or region proposal that includes a set of bounding boxes spanning the entire digital image. Visual features are extracted and evaluated for each bounding box to determine whether an object exists in the region proposal and which objects exist based on the visual features. Overlapping boxes are combined into a single bounding box (e.g., using non-maximum suppression). In an embodiment, overlapping boxes are also used to organize objects into categories in a big data store. For example, object 903 (a young man) is considered a parent object, and objects 906a-906c (clothes he is wearing) are considered child objects (shoes, shirts, pants) of object 903 due to overlapping bounding boxes. Therefore, a search for "people" using a search engine produces all objects labeled "people" and their child objects (if any, all included in the search results).
[0121] In an embodiment, instead of using bounding boxes, complex polygons are used to identify objects in an image. The complex polygons are used to determine highlight / hotspot areas in the image, such as where a user is pointing. Because only a complex polygon segment is sent to the cloud computing platform (rather than the entire image), privacy, security, and speed are improved.
[0122] Other examples of object detection frameworks that may be implemented by the cloud computing platform 102 to detect and label objects in digital images include, but are not limited to: Regional Convolutional Neural Network (R-CNN), Fast R-CNN, and Faster R-CNN.
[0123] In this example, the objects identified in the digital image include people, cars, buildings, roads, windows, doors, stair sign text. The identified objects are organized and presented as categories for the user to search. The user has selected the category "people" using a cursor or finger (if a touch-sensitive screen is used). By selecting the category "people", object 903 (i.e., the young man in the image) is isolated from the rest of the objects in the digital image, and a subset of objects 906a-906c are displayed in thumbnails 907 with their corresponding metadata. Object 906a is labeled "orange, shirt, button, short sleeve". Object 906b is labeled "blue, jeans, ripped, denim, pocket, phone", and object 906c is labeled "blue, Nike, shoes, left, air max, red socks, white logo, logo".
[0124] Search button 904, when pressed, initiates a new search based on the category selected by the user and the particular image in video pane 901. The search results include thumbnails 907. Similarly, if the user selects the category "Cars" and then presses search button 904, a new collection of thumbnails 907 is displayed, showing all of the cars captured in the image along with their respective metadata.
[0125] Fig.10 Illustrated is a method for classifying raw or pre-processed context data into Fig. 9 The classifier framework 1000 for searching objects and metadata by the GUI 900 of the present invention. The framework 1000 includes an API 1001, classifiers 1002a-1002n, and a data repository 1005. The raw or pre-processed context data captured on the wearable multimedia device is uploaded through the API 1001. The context data is run through the classifiers 1002a-1002n (e.g., neural networks). In an embodiment, the classifiers 1002a-1002n are trained using crowd-sourced context data from a large number of wearable multimedia devices. The output of the classifiers 1002a-1002n is objects and metadata (e.g., tags), which are stored in the data repository 1005. A search index is generated for the objects / metadata in the data repository 1005, which can be used by a search engine to search for objects / metadata that meet the search query entered using the GUI 900. Various types of search indexes can be used, including but not limited to: tree indexes, suffix tree indexes, inverted indexes, citation index analysis, and n-gram indexes.
[0126] Classifiers 1002a-1002n are selected and added to the dynamic data processing pipeline based on one or more of the data type, data volume, data quality, user preferences, user-initiated or application-initiated search queries, voice commands, (one or more) application requirements, templates, and / or any other information used to generate the desired presentation. Any known classifier can be used, including neural networks, support vector machines (SVMs), random forests, enhanced decision trees, and any combination of these individual classifiers using voting, stacking, and grading techniques. In an embodiment, some classifiers are personal to the user, that is, the classifier is trained only on contextual data from a specific user device. Such classifiers can be trained to detect and mark people and objects that are personal to the user. For example, a classifier can be used for face detection to detect faces in personal images that are known to the user (e.g., family members, friends) and have been marked by, for example, user input.
[0127] As an example, a user may speak multiple phrases such as: "Create a movie from my videos that includes my mom and dad in New Orleans"; "Add jazz music as a soundtrack"; "Send me a drink recipe to make Hurricane"; and "Send me directions to the nearest liquor store." The voice phrases are parsed, and the cloud computing platform 102 uses the words to assemble a personalized processing pipeline to perform the requested task, including adding a classifier for detecting the faces of the user's mom and dad.
[0128] In an embodiment, AI is used to determine how a user interacts with the cloud computing platform during a messaging session. For example, if a user says the message: "Bob, have you seen Toy Story 4?", the cloud computing platform determines who Bob is and parses "Bob" from the string sent to the message relay server on the cloud computing platform. Similarly, if the message says "Bob, look at this", the platform device sends an image with the message in one step without having to attach the image as a separate transaction. The user can visually confirm the image before sending it to Bob using the projector 1115 and any desired surface. In addition, the platform maintains a persistent personal communication channel with Bob for a period of time, so the name "Bob" does not have to appear before each communication during the messaging session.
[0129] Contextual Data Broker Service
[0130] In an embodiment, a context data broker service is provided by the cloud computing platform 102. The service allows users to sell their private raw or processed context data to entities of their choice. The platform 102 hosts the context data broker service and provides the security protocols required to protect the privacy of user context data. The platform 102 also facilitates transactions and transfers of currency or credit between entities and users.
[0131] A big data store may be used to store raw and pre-processed contextual data. The big data store supports storage and input / output operations for storage with large numbers of data files and objects. In an embodiment, the big data store includes an architecture consisting of a redundant and scalable provision of direct attached storage (DAS) pools, scale-out or clustered network attached storage (NAS), or object storage format-based infrastructure. The storage infrastructure is connected to compute server nodes that enable rapid processing and retrieval of large amounts of data. In an embodiment, the big data store architecture includes support for big data analytics solutions such as Hadoop. TM Cassandra TM and NoSQL TM ) native support.
[0132] In an embodiment, an entity interested in purchasing raw or processed contextual data subscribes to a contextual data broker service through a registration GUI or webpage of the cloud computing platform 102. Once registered, an entity (e.g., a company, an advertising agency) is allowed to trade directly or indirectly with a user through one or more GUIs customized to facilitate data brokering. In an embodiment, the platform 102 can match an entity's request for a particular type of contextual data with a user who can provide contextual data. For example, a clothing company may be interested in all images in which its clothing or competitor logos are detected. The clothing company can then use the contextual data to better identify the demographics of its customers. In another example, a news media or political campaign may be interested in video clips of newsworthy events for cover stories or feature articles. Various companies may be interested in a user's search history or purchase history to improve advertising targeting or other marketing projects. Various entities may be interested in purchasing contextual data to be used as training data for other object detectors (such as object detectors for autonomous vehicles).
[0133] In an embodiment, the user's raw or processed contextual data is provided in a secure format to protect the user's privacy. Both users and entities can have their own online accounts for depositing and withdrawing funds generated by proxy transactions. In an embodiment, the data proxy service collects transaction fees based on a pricing model. Fees can also be obtained through traditional online advertising (e.g., click-through rates of banner ads, etc.).
[0134] In an embodiment, individuals can create their own metadata using wearable multimedia devices. For example, a celebrity chef can wear a wearable multimedia device 102 while preparing a meal. Objects in an image are tagged using metadata provided by the chef. Users can obtain access to metadata from a proxy service. When a user attempts to prepare a dish while wearing their own wearable multimedia device 102, objects are detected and metadata (e.g., time, amount, sequence, proportion) provided by the chef to help the user reproduce something is projected onto the user's work surface (e.g., cutting board, countertop, stove, oven, etc.), such as measurements, cooking time, and additional tips. For example, a piece of meat is detected on a cutting board, and the projector 1115 projects text onto the cutting board, reminding the user to cut the meat against the grain, and also projects a measurement guide onto the meat surface to guide the user to cut slices of uniform thickness according to the chef's metadata. Laser projection guides can also be used to cut vegetables of uniform thickness (e.g., filaments, diced). Users can upload their own metadata to a cloud service platform, create their own channel, and earn revenue through subscriptions and advertising (similar to the YouTube platform).
[0135] Fig.11 is a system block diagram illustrating a hardware architecture 1100 for a wearable multimedia device according to an embodiment. The architecture 1100 includes a system on chip (SoC) 1101 (e.g., a Qualcomm Snapdragon chip), a main camera 1102, a 3D camera 1103, a capacitive sensor 1104, a motion sensor 1105 (e.g., an accelerometer, a gyroscope, a magnetometer), a microphone 1106, a memory 1107, a global navigation satellite system receiver (e.g., a GPS receiver) 1108, a WiFi / Bluetooth chip 1109, a wireless transceiver chip 1110 (e.g., 4G, 5G), a radio frequency (RF) transceiver chip 1112, an RF front-end electronics device (RFFE) 1113, an LED 1114, a projector 1115 (e.g., laser projection, micro projector, LCoS, DLP, LCD), an audio amplifier 1116, a speaker 1117, an external battery 1118 (e.g., a battery pack), a magnetic sensing circuit system 1119, a power management chip (PMIC) 1120, and an internal battery 1121. All of these components work together to facilitate the various tasks described herein.
[0136] Fig.121 is a system block diagram illustrating an alternative cloud computing platform 1200 for processing raw or pre-processed context data received from a wearable multimedia device according to an embodiment. An edge server 1201 receives raw or pre-processed context data from a wearable multimedia device 1202 via a wireless communication link. The edge server 1201 provides limited local pre-processing, such as AI or camera video (CV) processing and gesture detection. At the edge server 1201, a scheduler 1203 directs the raw or pre-processed context data to a state / context detector 1204, a first-party handler 1205, and / or a restricted AI parser 1206 to perform restricted AI tasks. The state / context detector 1204 uses GNSS data, for example, provided by a GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometry) of the wearable multimedia device 1202 to determine the location where the context data was captured. The state / context detector 1204 also uses image and voice technology and AI to analyze images, audio, and sensor data (e.g., motion sensor data, biometric data) contained in the context data to determine user activities, emotions, and interests.
[0137] The edge server 1201 is coupled to the regional data center 1207 via an optical fiber and a router. The regional data center 1207 performs full AI and / or CV processing on the pre-processed or raw context data. At the regional data center 1207, the scheduler 1208 directs the raw or pre-processed context data to the state / context detector 1209, the full AI parser 1210, the first processing program 1211, and / or the second processing program 1212. The state / context detector 1209 uses GNSS data provided by, for example, a GPS receiver or other positioning technology (e.g., Wi-Fi, cellular, visual odometer) of the wearable multimedia device 1202 to determine the location where the context data was captured. The state / context detector 1209 also uses image and speech recognition technology and AI to analyze images, audio, and sensor data (e.g., motion sensor data, biometric data) contained in the context data to determine user activities, emotions, and interests.
[0138] Fig.13The software components 1300 for a wearable multimedia device according to an embodiment are illustrated. For example, the software components include a daemon 1301 for AI and CV, gesture recognition, messaging, media capture, server connection, and accessory connection. The software components also include libraries 1302 for graphics processing units (GPUs), machine learning (ML), camera video (CV), and network services. The software components include an operating system 1303, such as an Android Native Development Kit (NDK), including hardware abstraction and a Linux kernel. Other software components 1304 include components for power management, connectivity, security + encryption, and software updates.
[0139] Figures 14A-14D It is illustrated that various types of information projected by the projector 1115 are projected onto the palm of the user using the projector 1115 of the wearable multimedia device according to an embodiment. In particular, Fig.14A A laser projection of a numeric keypad on the palm of a user is shown for dialing a phone number and other tasks requiring the input of numbers. A 3D camera 1103 (depth sensor) is used to determine the position of the user's fingers on the numeric keypad. The user can interact with the numeric keypad using functions such as dialing a phone number. Fig. 14B The turn-by-turn direction projected on the palm of the user is shown. Fig. 14C A clock is shown projected onto the palm of a user. Fig.14D A temperature reading on the user's palm is shown. The 3D camera 1103 can detect various one or two finger gestures (eg, tap, long press, swipe pinch / unpinch), resulting in triggering different actions on the wearable multimedia device.
[0140] Although Figures 14A-14D Laser projection onto the palm of a user's hand is shown, but any projection surface may be used, including but not limited to: walls, floors, ceilings, drapes, clothing, projection screens, tables / desks / countertops, appliances (e.g., stoves, washers / dryers), and devices (e.g., car engines, electronic circuit boards, cutting boards).
[0141] Fig.15A and 15B An application of projector 1115 is illustrated in accordance with an embodiment, where information to assist a user in checking their engine oil is projected onto a car engine. Fig.15A and 15BBefore and after images of an image are shown. The user speaks a voice: "How do I check the oil?" In this example, the voice is received by the microphone 1106, and the main camera 1102 and / or the 3D camera 1103 of the wearable multimedia device 1202 captures an image of the engine. The image and voice are compressed and sent to the edge server 1201. The edge server 1201 sends the image and audio to the regional data center 1207. At the regional data center 1207, the image and audio are decompressed, and one or more classifiers are used to detect and mark the location of the dipstick and the fuel filler cap in the image. The label and its image coordinates are sent back to the wearable multimedia device 1202. The projector 1115 projects the label onto the car engine based on the image coordinates.
[0142] Fig.16 The application of a projector according to an embodiment is illustrated, in which information for helping a home chef cut vegetables is projected onto a cutting board. The user speaks a voice: "How big should I cut this?" In this example, the voice is received by the microphone 1106, and the main camera 1102 and / or the 3D camera 1103 of the wearable multimedia device 1202 captures an image of the cutting board and the vegetable. The image and voice are compressed and sent to the edge server 1201. The edge server 1201 sends the image and audio to the regional data center 1207. At the regional data center 1207, the image and audio are decompressed, and one or more classifiers are used to detect the type of vegetable (e.g., broccoli), its size, and its location in the image. Based on the image information and audio, the cutting instructions (e.g., obtained from a database or other data source) and the image coordinates are determined and sent back to the wearable multimedia device 1202. The projector 1115 uses the information and the image coordinates to project a size template onto the cutting board around the vegetable.
[0143] Fig.17 1700 is a system block diagram of a projector architecture according to an embodiment. The projector 1115 scans pixels in two dimensions, images a 2D pixel array, or a hybrid of imaging and scanning. Scanning projectors directly use a narrow divergence of a laser beam and two-dimensional (2D) scanning to "draw" an image pixel by pixel. In some embodiments, separate scanners are used for horizontal and vertical scanning directions. In other embodiments, a single two-axis scanner is used. The specific beam trajectory also varies depending on the type of scanner used.
[0144] In the example shown, projector 1700 is a scanning micro-projector that includes a controller 1701 , a battery 1118 / 1121 , a power management chip (PMIC) 1120 , a solid-state laser 1704 , an XY scanner 1705 , a driver 1706 , a memory 1707 , a digital-to-analog converter (DAC) 1708 , and an analog-to-digital converter (ADC) 1709 .
[0145] The controller 1701 provides a control signal to the XY scanner 1705. The XY scanner 1705 uses a movable mirror to manipulate the laser beam generated by the solid-state laser 1704 in two dimensions in response to the control signal. The XY scanner 1705 includes one or more micro-electromechanical (MEMS) micro-mirrors having a controllable tilt angle in one or two dimensions. The driver 1706 includes a power amplifier and other electronic circuit systems (e.g., filters, switches) to provide a control signal (e.g., voltage or current) to the XY scanner 1705. The memory 1707 stores various data used by the projector, including laser patterns for text and images to be projected. The DAC 1708 and the ADC 1709 provide data conversion between the digital domain and the analog domain. The PMIC 1120 manages the power and duty cycle of the solid-state laser 1704, including turning the solid-state laser 1704 on and off and adjusting the amount of power supplied to the solid-state laser 1704. The solid-state laser 1704 is, for example, a vertical cavity surface emitting laser (VCSEL).
[0146] In an embodiment, the controller 1701 uses image data from the main camera 1102 and depth data from the 3D camera 1103 to identify and track the user's hand and / or finger positions on the laser projection, so that the wearable multimedia device 102 receives user input using the laser projection as an input interface.
[0147] In another embodiment, the projector 1115 uses a vector graphics projection display and a low-power fixed MEMS micromirror to save power. Because the projector 1115 includes a depth sensor, the projection area can be shielded if necessary to prevent projection onto fingers / hands interacting with the laser projected image. In an embodiment, the depth sensor can also track gestures to control input on other devices (e.g., scanning images on a TV screen, interacting with a computer, smart speaker, etc.).
[0148] In other embodiments, liquid crystal on silicon (LCoS or LCOS), digital light processing (DLP), or liquid crystal display (LCD) digital projection technology may be used instead of the pico projector.
[0149] Fig.18 Laser parameter adjustment based on the amount of light reflected by a surface is illustrated. To ensure that projections are clear and easy to read on a variety of surfaces, data from the 3D camera 1103 is used to adjust one or more parameters of the projector 1115 based on surface reflections. In an embodiment, reflections of the laser beam from a surface are used to automatically adjust the intensity of the laser beam to compensate for different refractive indices to produce a projection with uniform brightness. For example, the intensity can be adjusted by adjusting the power supplied to the solid-state laser 1115. The amount of adjustment can be calculated by the controller 1701 based on the energy level of the reflected laser beam.
[0150] In the example shown, a circular pattern 1800 is projected on a surface 1801 that includes an area 1802 having a first surface reflection and an area 1803 having a second surface reflection different from the first surface reflection. The difference in surface reflection in the areas 1802, 1803 (e.g., due to different refractive indices) causes the brightness of the circular pattern 1800 in the area 1802 to be lower than that in the area 1803. In order to generate a circular pattern 1800 with uniform intensity, the solid-state laser 1704 is commanded by the controller 1701 (via the PMIC 1120) to increase / decrease the power supplied to the solid-state laser 1704 to increase the intensity of the laser beam when scanning in the area 1802. The result is a circular pattern 1800 with uniform brightness. In the case where the surface geometries of the areas 1802 and 1803 are different, one or more lenses can be used to adjust the size of the text or image projected on the surface. For example, the area 1802 can be curved, while the area 1803 can be flat. In such a scenario, the size of the text or image in region 1802 may be adjusted to compensate for the curvature of the surface in region 1802 .
[0151] In an embodiment, laser projection can be automatically or manually requested by a user's air gesture (e.g., finger pointing to identify an object of interest, swiping to indicate an operation on data, raising a finger to indicate counting, thumbs up or down to indicate a preference, etc.), and projected onto any surface or object in the environment. For example, a user can point to a thermostat in their home and the temperature data is projected onto their palm or other surface. The camera and depth sensor detect where the user is pointing, identify the object as a thermostat, run the thermostat application on a cloud computing platform, and stream the application data to the wearable multimedia device, where the application data is displayed on a surface (e.g., the user's palm, wall, table). In another example, if a user is standing in front of a smart lock on the front door of their house, and all the locks in their home are linked, the controls for the smart lock are projected on the surface of the smart lock or door to access that lock or other locks in their home.
[0152] In an embodiment, images captured by a camera and its large field of view (FOV) can be presented to a user in a "contact sheet" using an AI-driven virtual photographer running on a cloud computing platform. For example, various renderings of the image are created using different crops and processing methods using machine learning (e.g., neural networks) trained with images / metadata created by professional photographers. With this feature, each image captured can have multiple "looks" involving multiple image processing operations on the original image, including operations informed by sensor data (e.g., depth, ambient light, accelerometer, gyroscope).
[0153] The features described may be implemented in digital electronic circuitry or in computer hardware, firmware, software, or a combination thereof. The features may be implemented in a computer program product tangibly embodied in an information carrier, such as in a machine-readable storage device, for execution by a programmable processor. The method steps may be performed by a programmable processor executing a program of instructions to perform the functions of the described embodiments by operating on input data and generating output.
[0154] The described features may advantageously be implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a collection of instructions that can be used, directly or indirectly, in a computer to perform a specific activity or produce a specific result. A computer program may be written in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0155] As examples, suitable processors for executing instruction programs include general and special purpose microprocessors, and the sole processor or one of multiple processors or cores of any type of computer. In general, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. In general, a computer can communicate with a mass storage device for storing data files. These mass storage devices may include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly implementing computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into an ASIC (Application Specific Integrated Circuit). To provide interaction with a user, these features may be implemented on a computer having a display device such as a CRT (cathode ray tube), LED (light emitting diode), or LCD (liquid crystal display) display or monitor for displaying information to the author, a keyboard, and a pointing device (such as a mouse or trackball) through which the author can provide input to the computer.
[0156] One or more features or steps of the disclosed embodiments may be implemented using an application programming interface (API). The API may define one or more parameters passed between a calling application and other software code (e.g., an operating system, a library routine, a function) that provides a service, provides data, or performs an operation or calculation. The API may be implemented as one or more calls in a program code that send or receive one or more parameters through a parameter list or other structure based on a calling convention defined in an API specification document. Parameters may be constants, keys, data structures, objects, object classes, variables, data types, pointers, arrays, lists, or other calls. API calls and parameters may be implemented in any programming language. A programming language may define the vocabulary and calling conventions that a programmer will use to access functions that support the API. In some embodiments, an API call may report to an application the capabilities of the device running the application, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, and the like.
[0157] Many embodiments have been described. Nevertheless, it should be understood that various modifications can be made. Elements of one or more embodiments can be combined, deleted, modified or supplemented to form other embodiments. In another example, the logic flow depicted in the figure does not require the specific order or sequence shown to achieve the desired result. In addition, other steps can be provided from the described process, or steps can be eliminated therefrom, and other components can be added to or removed from the described system. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A body-worn device comprising: camera; Depth sensor; Laser projection system; one or more processors; A memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: capturing a first set of digital images using a camera; identifying real-world objects in the first set of digital images; capturing first depth data using a depth sensor; identifying a first posture of a user wearing the device in the first set of digital images and the first depth data, Wherein identifying the real-world object and the first posture comprises: processing the first set of data images through an object detection framework that uses complex polygons to identify hotspot regions in the first set of images, wherein the hotspot regions are smaller than the entire image and capture the first pose and the real-world object but exclude all other objects in the first set of images; Sending the hot spot area to the cloud computing platform; receiving information related to the real-world object from a cloud computing platform; and At least some of the information is projected onto a surface using the laser projection system.
2. The apparatus of claim 1, wherein the information comprises a text label for the real-world object.
3. The apparatus of claim 1, wherein the information comprises instructions for performing an action on the real-world object.
4. The apparatus of claim 1, wherein the information includes commands for controlling the real-world object.
5. The apparatus of claim 1, wherein the operations further comprise: obtaining a second pose associated with the laser projection using the depth sensor; determining a user input based on the second gesture; as well as One or more actions are initiated based on the user input.
6. The apparatus of claim 1, wherein the operations further comprise: The laser projection is shielded to prevent the data from being projected onto the hand of the user making the second gesture.
7. The apparatus of claim 1, wherein the operations further comprise: capturing, using a camera, a reflection of the laser projection from the surface; The intensity of the laser projection is automatically adjusted to compensate for different refractive indices so that the laser projection has uniform brightness.
8. The apparatus of claim 1, further comprising: A magnetic attachment mechanism is configured to magnetically couple to a battery pack through a user's clothing, the magnetic attachment mechanism also being configured to receive an inductive charge from the battery pack.
9. A method comprising: capturing a collection of digital images using a camera; capturing depth data using a depth sensor of a body-worn device; identifying real-world objects in said first set of digital images; identifying a first gesture in the set of digital images and the depth data, the first gesture being performed by a user wearing the device; Wherein identifying the real-world object and the first posture comprises: processing the first set of data images through an object detection framework that uses complex polygons to identify hotspot regions in the first set of images, wherein the hotspot regions are smaller than the entire image and capture the first pose and the real-world object but exclude all other objects in the first set of images; Sending the hot spot area to the cloud computing platform; receiving information related to the real-world object from a cloud computing platform; and At least some of the information is projected onto a surface using the laser projection system.
10. The method of claim 9, further comprising: obtaining a second posture of the user using the depth sensor, the second posture being associated with the laser projection; determining a user input based on the second gesture; as well as One or more actions are initiated based on the user input. The method of claim 10 , wherein the one or more actions include controlling the real-world object.
12. The method of claim 10, further comprising: The laser projection is shielded to prevent the data from being projected onto the hand of the user making the second gesture.
13. The method of claim 11, wherein the controllable real-world object is a television or a computer screen, and the one or more actions include swiping one or more images displayed by the television or computer screen.
14. The method of claim 11, wherein the controllable real-world object is a thermostat and the one or more actions include changing a temperature setting of the thermostat.
15. The method of claim 9, wherein the laser projection includes a dimension template for measuring the real-world object.
16. The method of claim 9, further comprising: receiving audio input from a user; as well as Using the one or more processors, a first gesture and an audio input are associated with the request or command to control the real-world object.
17. The method of claim 9, wherein the surface is a palm of a user.