Real-time anonymization of private spaces
By using generative AI models in augmented reality and virtual reality devices to identify and obfuscate private information in real time, combined with a secure data vault system and sandbox environment, the problem of private information capture and leakage by devices is solved, achieving higher privacy protection and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2024-10-25
- Publication Date
- 2026-05-29
AI Technical Summary
When using augmented reality and virtual reality devices, users may inadvertently capture or hear private information, and there is a lack of effective privacy protection mechanisms.
Generative AI models are used to identify, blur, or cover up private information in real time. This information is then executed locally on the device through a secure data vault system. Combined with a sandbox environment and secure network services, the system provides isolation and authentication to ensure data security.
It enables real-time detection and protection of private information in augmented reality and virtual reality devices, improving user privacy and device security, and preventing malware attacks.
Smart Images

Figure CN122122583A_ABST
Abstract
Description
Priority Statement
[0001] This patent application claims the benefit of priority to U.S. Application No. 18 / 496,718, filed October 27, 2023, which is incorporated herein by reference in its entirety. Background Technology
[0002] Augmented reality (AR) systems include camera systems capable of capturing various electronic images and videos, such as cameras mounted on mobile devices. The popularity of image and video capture continues to grow. Image and video capture is used to provide AR visualizations. Additionally, users are increasingly sharing media content items such as electronic images and videos with each other. Users are also increasingly utilizing their mobile devices to communicate with each other using messaging apps. Attached Figure Description
[0003] In the accompanying drawings (which are not necessarily drawn to scale), similar reference numerals may describe similar parts in different views. For ease of identification of any particular element or action being discussed, one or more of the most significant digits in the reference numerals refer to the figure number in which that element is first introduced. Some non-limiting examples are shown in the accompanying drawings:
[0004] Figure 1 It is a diagrammatic representation of a networked environment in which the content of this disclosure can be deployed, based on some examples.
[0005] Figure 2 This is a block diagram illustrating details of an example of a secure data vault system based on some examples.
[0006] Figure 3 It is based on some examples stored in Figure 2 A diagram of various generative AI models in a secure data vault.
[0007] Figure 4 This is a diagram illustrating, based on some examples, the use of generative AI models to detect and block certain private information when using AR and / or VR devices.
[0008] Figure 5 This is a diagram illustrating, based on some examples, the use of generative AI models to detect and block certain private information based on location recognition when using AR and / or VR devices.
[0009] Figure 6 It is a graphical representation of a messaging system with both client-side and server-side functionalities, based on some examples.
[0010] Figure 7 It is a graphical representation based on examples such as data structures maintained in a database.
[0011] Figure 8 It is a graphical representation based on some example messages.
[0012] Figure 9 A system with a head-mounted device is shown according to some examples.
[0013] Figure 10 The following examples illustrate the processing applicable to retrieving and displaying certain group communication items, including media montages.
[0014] Figure 11 It is a graphical representation of a machine in the form of a computer system, based on some examples, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.
[0015] Figure 12 It is a block diagram showing the software architecture that can be implemented within it. Detailed Implementation
[0016] Camera systems and microphones are included in a variety of devices such as mobile devices, smartwatches, drones, etc. These systems enable users to capture images and videos and communicatively and / or operatively couple to applications, such as interactive clients. In some examples, interactive clients allow users to capture media content while using the client and to apply augmented reality (AR) and / or virtual reality (VR) content, including photographic filters and / or virtual lenses, to the media content. The obtained media content is used to interact with other users (e.g., group members) by transmitting media messages, who can then respond with their own media content. Users wearing AR and / or VR devices may be using the devices in a public space such as a hospital, hotel lobby, airplane, etc. Users may inadvertently enter private spaces, such as restrooms, while wearing AR and / or VR devices. Similarly, users may inadvertently overhear private conversations, such as conversations taking place in the same room where they are playing AR and / or VR games.
[0017] The techniques described herein protect privacy by identifying certain private information in real time and then blurring or otherwise obfuscating images and / or sounds deemed private. In some examples, AI models, such as generative AI models, are used to derive the private information. Generative AI models are a class of artificial intelligence models designed to generate data based on training data, typically in the form of text, images, and / or audio. In some examples, generative AI models are used to generate augmented reality (AR) and / or virtual reality (VR) content. Data (e.g., data provided via camera systems and microphones included in various devices such as mobile devices) is provided as input to the generative AI model. The AR and / or VR system then uses the output generated via the generative AI model to add certain virtual content, such as images, videos, audio, text, etc., to the corresponding AR and / or VR environment. Devices that support AR experiences in any of these methods are referred to herein as "AR devices," and devices that support VR experiences are referred to herein as "VR devices."
[0018] In some examples, a secure data vault system is an isolated and secure environment within an operating system (OS), such as AR OS. The secure data vault system maintains separation from user applications while allowing users / creators to retain control over their data. The secure data vault system includes a "sandbox" that allows users to perform certain computations (e.g., via generative AI models) on raw and computed data within a sandbox. Raw data includes data from cameras, microphones, and other sensors. Raw data is shared with the secure data vault system via privileged areas of the OS, allowing the secure data vault system to manage the use of the raw data. Computed data is data that has been generated from the raw data, for example, via a generative AI model, or data generated without using cameras, microphones, and other sensors. As further described below, raw data is stored within the secure data vault system, while computed data is allowed to leave the secure data vault system subject to certain policy enforcement. Networked computing environment
[0019] Figure 1This is a block diagram illustrating an example interactive system 100 for facilitating interactions on a network, such as exchanging text messages, making text audio and video calls, or playing games. Interactive system 100 includes multiple user systems 102, each of which hosts multiple applications including an interactive client 104 and other applications 106. Each interactive client 104 is communicatively coupled to other instances of the interactive client 104 (e.g., hosted on corresponding other user systems 102), an interactive server system 110, and a third-party server 112 via one or more communication networks including a network 108 (e.g., the Internet). The interactive client 104 may also communicate with the locally hosted applications 106 using an application programming interface (API).
[0020] Each user system 102 may include multiple user devices, such as mobile devices 114, head-mounted devices 116, and computer client devices 118, which can be communicatively connected to exchange data and messages.
[0021] Interactive client 104 interacts with other interactive clients 104 and with interactive server system 110 via network 108. The data exchanged between interactive clients 104 (e.g., interaction 120) and between interactive client 104 and interactive server system 110 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0022] Interactive server system 110 provides server-side functionality to interactive client 104 via network 108. While some functions of interactive system 100 are described herein as being performed by interactive client 104 or interactive server system 110, the location of certain functions within interactive client 104 or interactive server system 110 may be a design choice. For example, it may be technically preferred that specific technologies and functions are initially deployed within interactive server system 110, but that technology and functions are later migrated to interactive client 104 of user system 102, which has sufficient processing power.
[0023] The interactive server system 110 supports various services and operations provided to the interactive client 104. Such operations include sending data to and receiving data from the interactive client 104, and processing data generated by the interactive client 104. This data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, entity relationship information, and live event information. Data exchange within the interactive system 100 is activated and controlled via functions available through the user interface (UI) of the interactive client 104.
[0024] Specifically, the focus now shifts to interactive server system 110. Application Programming Interface (API) server 122 is coupled to interactive server 124 and provides it with a programming interface, making the functionality of interactive server 124 accessible to interactive client 104, other applications 106, and third-party server 112. Interactive server 124 is communicatively coupled to database server 126, thereby facilitating access to database 128, which stores data associated with the interactions processed by interactive server 124. Similarly, web server 130 is coupled to interactive server 124 and provides a web-based interface to interactive server 124. For this purpose, web server 130 handles incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0025] Application Programming Interface (API) server 122 receives and sends interactive data (e.g., command and message payloads) between interactive server 124 and user system 102 (as well as interactive client 104 and other applications 106) and third-party server 112. Specifically, API server 122 provides a set of interfaces (e.g., routines and protocols) that interactive client 104 and other applications 106 can invoke or query to activate the functionality of interactive server 124. Application Programming Interface (API) server 122 exposes various functions supported by interaction server 124, including account registration; login functionality; sending interactive data from one interaction client 104 to another interaction client 104 via interaction server 124; transferring media files (e.g., images or videos) from interaction client 104 to interaction server 124; setting up media data sets (e.g., stories); retrieving the friend list of users of user system 102; retrieving messages and content; adding and deleting entities (e.g., friends) against an entity relationship graph (e.g., entity graph 710); locating friends in the entity relationship graph; and opening (e.g., application events associated with interaction client 104).
[0026] Also shown is a secure data vault system 132 for processing raw data from devices 114, 116, and 118. As mentioned earlier, the raw data includes data generated via camera devices, microphones, and / or other sensors such as gyroscopes, navigation systems (e.g., GPS, inertial navigation systems), temperature sensors, etc., included in devices 114, 116, and 118. In some embodiments, the raw data is processed only via the secure data vault system 132, thereby providing data isolation and security.
[0027] In practice, the secure data vault system is an isolated and secure environment within an operating system (OS), such as an AR OS. The secure data vault system remains separate from application 106 while allowing users / creators to maintain control over their data. In some implementations, as further described below, the secure data vault system is part of the OS but includes hardware components for providing isolation and trust. The secure data vault system also enables the local execution of machine learning and generative AI models that use raw data as input. That is, certain generative AI models are trained to take raw data as input and then output AR “terms,” such as virtual content suitable for overlaying real-world images. These generative AI models are executed locally by the secure data vault system 132 only within the corresponding devices 114, 116, 118, and not on external systems such as external servers or cloud-based systems.
[0028] Generative AI model 208 is trained to recognize certain visual and / or audio information considered private. For example, a user wearing AR and / or VR devices may be playing a game while images and audio are captured by cameras and microphones included in the AR and / or VR device. During operation of the AR and / or VR device, certain private images and / or audio may be unintentionally captured. For example, a user may enter a public restroom or otherwise see the interior of a public restroom. Similarly, a user may hear certain private conversations or sounds. The techniques described herein not only use generative AI model 208 to generate new AR / VR content, but also additionally detect private information and then blur or otherwise “mask” certain portions of images and / or audio to protect privacy.
[0029] Figure 2 This is a block diagram illustrating further details of an example of a secure data vault system 132 according to some examples. In the depicted example, the secure data vault system 132 includes an application programming interface (API) 202, which is used by developers to interface with the secure data vault system 132 and perform certain functions of the secure data vault system 132. In some implementations, interface with the secure data vault system 132 via API 202 is performed using a developer-provided API (e.g., via the depicted public developer API 204).
[0030] API 202 includes function calls, methods, object-oriented classes, etc., that can execute certain functions provided by the secure data vault system 132. In some implementations, API 202 may include certain security functions, such as authenticating developer use of API 202 via receiving a security token, login / password combination, challenge / response authentication, secure handshake authentication, etc. Once the session is authenticated, the secure data vault system 132 can receive raw data via privileged access OS API 206. More specifically, privileged access OS API 206 is used solely by the secure data vault system 132 to receive raw data and is not used by any other system.
[0031] As mentioned above, the raw data includes camera device data, microphone data, and / or sensor data for various sensors included in devices such as devices 114, 116, and 118 (e.g., AR devices, VR devices). In some embodiments, the OS is also authorized to access the raw data, for example, to perform tests on the camera device, microphone, and / or other sensors, to calibrate the camera device, microphone, and / or other sensors, etc. By isolating the raw data when using the camera device, microphone, and / or other sensors so that it is processed only via the secure data vault system 132, the techniques described herein improve security and enhance privacy. For example, users of devices 114, 116, and 118 can enjoy AR experiences including new images and / or videos created via generative AI model 208. The secure data vault system 132 will fully execute the generative AI model 208 in execution environment 212 to enhance security and privacy.
[0032] That is, execution environment 212 is completely enclosed within sandbox environment 210. The sandbox environment provides increased isolation and security. For example, sandbox environment 210 ensures that any actions performed within it remain isolated from the rest of the system, thereby preventing potential harm from malware or unintended consequences from the code. In some examples, execution environment 212 restricts the execution of instructions or code to a predefined memory range (e.g., a start memory address and an end memory address). By isolating the application or code (e.g., the code used to execute generative AI model 208), sandbox environment 210 reduces the risk of security vulnerabilities such as buffer overflows or privilege escalation affecting the entire system. If one or more generative AI models 208 executing within sandbox environment 210 unexpectedly behave or become unresponsive, the generative AI model can be terminated without affecting the rest of the system. Sandbox environment 210 may additionally utilize virtualization techniques to create completely separate virtual machines (VMs) or containers to host the isolated execution environment 212. In some implementations, sandbox environment 210 is a separate hardware system. For example, field-programmable gate arrays (FPGAs), separate microprocessors, custom circuitry, etc., can be provided as part of devices 114, 116, 118 and used as sandbox environments 210. Thus, for example, attack pathways are minimized by minimizing the exposure of code that could be maliciously modified.
[0033] The secure data vault system 132 imposes constraints on data sharing both on-device and off-device (network). In some implementations, the secure data vault system 132 may communicate only with a trusted set of features. A list of permitted processes / features can be placed in a security compatibility matrix 214. For example, columns in the security compatibility matrix 214 list the various functions provided by the secure data vault system 132, and rows in the security compatibility matrix 214 list the processes authorized to access these functions. Therefore, a process can invoke API 202 and request the execution of a function, including functions related to generative AI models (e.g., providing output based on input, upgrading to a newer model version, etc.), and can then check the security compatibility matrix 214 to verify that the process has permission to execute that function.
[0034] Also shown is a secure network service 216 included in the secure data vault system 132. Secure network service 216 is used to communicate with external systems to update AI model 208, execution environment 212, and / or security compatibility matrix 214. Secure network service 216 can use technologies such as Transport Layer Security (TLS) in Hypertext Transfer Protocol Security (HTTPS) download-only mode. That is, in some examples, secure network service 216 can be used only to download information, such as a newer version of generative AI model 208. Authentication of downloads made via secure network service 216 can be provided via challenge / response hardware technology (e.g., via a hardware security token device that provides multi-factor authentication), key usage, time-slot downloads (e.g., where downloads only occur at a given time of day), etc.
[0035] Display and rendering system 218 is also included in secure data vault system 132, which is adapted to dynamically create and display 3D representations, for example, as an overlay on a real-time view of the surrounding environment provided by a camera device. Sound can also be provided via display and rendering system 218.
[0036] The display and rendering system 218 also includes rendering certain virtual content using the output from the generative AI model 208. For example, the generative AI model can be used in games to create filters, stickers, animations, etc., based on input data. For instance, camera data can include real-time images of people in a room, and the generative AI model 208 can then create virtual animations of the people. Similarly, gameplay can be generated in real-time via the generative AI model 208.
[0037] The display and rendering system 218 further includes security features, such as using a generative AI model 208 to detect and blur images and / or sounds that may have been found to be private. As mentioned earlier, certain documents (e.g., driver's licenses, credit cards, social security cards, passports, financial documents, etc.) may have been viewed during an AR session. The generative AI model 208 can detect the presence of private information and cooperate with the display and rendering system 218 to blur images and / or mute sounds, thereby protecting privacy. In some implementations, the display and rendering system 218 may also notify the user of the presence of private information. By providing isolated, secure, and local execution via the sandbox environment 210, the secure data vault system 132 enhances privacy and user security in various AR systems.
[0038] Figure 3This is a block diagram illustrating further details of a generative AI model 208 based on some examples. The artificial intelligence model is designed to generate new data or content similar to or analogous to a given training dataset. The generative AI model 208 is used to create data such as images, text, audio, video, animation, filters, and / or stickers based on patterns and structures it has learned from input data during training. The generative AI model 208 operates by learning complex statistical patterns and dependencies in the training data, and then uses this knowledge to generate new, coherent data consistent with these patterns.
[0039] In the depicted example, the generative AI model 208 includes various model types, such as a variational autoencoder (VAE) model 302, a generative adversarial network (GAN) model 304, a recurrent neural network (RNN) model 306, a Transformer model 308, an autoencoder model 310, and other models 312. The VAE model 302 learns a probabilistic mapping between data and a latent space and can be used for tasks such as image generation and data compression. The GAN model 304 includes a generator and a discriminator trained in a competitive manner.
[0040] RNN model 306 computation depends on the time step of the previous time step. RNN model 306 has self-looping connections, allowing it to maintain hidden states representing information about previous elements in the sequence it has processed. These hidden states enable RNN model 306 to exhibit dynamic temporal behavior, making them well-suited for tasks involving sequences. Transformer model 308 is based on a neural network architecture designed for processing sequential data, primarily focusing on natural language processing tasks such as machine translation and text generation. Transformer model 308 was introduced in the 2017 paper "Attention is All You Need" by Vaswani et al. Autoencoder model 310 is a model that learns to reconstruct input data from lower-dimensional representations. Other models 312 are also used, such as variants of the Large Language Model (LLM).
[0041] In the depicted implementation, training dataset 314 is used to train various generative AI models 208. In some examples, training dataset 314 is not part of the secure data vault system 132, but is only used within the secure data vault system 132 for the trained generative AI models 208. Training dataset 314 includes AR and VR data, such as images, videos, audio, text, stickers, animations, filters, etc., created during the use of AR and VR devices. Other data, such as “records” of AR and VR experiences while using AR and VR devices, are also part of training dataset 314. The training overview for generative AI model 208 is as follows:
[0042] Data collection and preprocessing: Collect diverse and extensive datasets containing data from areas of interest. These datasets can include VR / AR recordings, audio, images, videos, books, articles, websites, and other text sources. For detecting private information, the training dataset includes images and / or text of credit cards, passports, bank statements, loan documents, driver's licenses, social security cards, licensing documents (e.g., fishing licenses, hunting licenses), registration documents (e.g., vehicle registrations, boat registrations), deed documents (e.g., residential property deeds, commercial property deeds), liens, etc. Audio training data includes conversations discussing financial matters, such as calls to banks, credit card companies, loan companies, insurance companies, real estate agents, etc. Audio training data also includes sounds such as shouting, fighting, showering, bathroom sounds, bedroom sounds, etc. For detecting private locations, the training dataset includes images of public locations (e.g., hotels, airplanes, hospitals, schools, libraries, restaurants, bars, buses, cinemas, theaters, etc.) and private locations found within these public locations. For example, images of private locations include images of various restrooms and restroom areas (including doors), other rooms (e.g., other bedrooms in a hotel), portions of a room (e.g., a hospital room may include multiple areas for multiple patients, and each area can then be a private location, and images of these areas are part of the training dataset), images of meeting rooms, images of certain workplace areas (e.g., offices, shared spaces), and so on. Texts such as “restroom,” “meeting room,” “private,” “staff only,” and “library” can also be part of the training dataset to represent private locations. Similarly, icons representing private locations, such as restroom icons, “do not enter” icons, restaurant kitchen icons, etc., are used in the training dataset 314. The data is preprocessed by segmenting the text into smaller units such as words or sub-words, image fragments, video clips, etc., and any necessary cleaning and formatting is performed.
[0043] Word segmentation and vocabulary creation: Tokenize data (e.g., text) into units that the model can process, such as subwords or words. Tokenization is used to create a vocabulary that the model uses during training. A vocabulary is constructed based on the lexical units in the training data. This vocabulary defines the set of lexical units that the model can recognize and generate.
[0044] Model initialization: Generative AI models are initialized using random weights or pre-trained weights (if fine-tuning an existing model).208
[0045] Pre-training: The model is pre-trained on a large amount of data in an unsupervised manner. During pre-training, the model learns to predict the next lexical term in a sequence (e.g., autoregressive language modeling) or performs other unsupervised tasks, such as masked language modeling. Techniques such as batch processing, distributed computing, and parallel processing are used to efficiently process large amounts of data.
[0046] Loss function: Define an appropriate loss function for the pre-training task, such as cross-entropy loss.
[0047] optimization: Choose an optimization algorithm (e.g., Adam, SGD) and tune hyperparameters such as learning rate, batch size, and weight decay. Apply gradient clipping to prevent gradient explosion.
[0048] train: The model is trained on a pre-training task with numerous iterations or periods. This step takes some time. Monitor training progress, track loss values, and use evaluation metrics to assess model performance.
[0049] Fine-tuning: After pre-training, the model can be fine-tuned for domain-specific or downstream tasks by adding task-specific layers and training the model on labeled data. Fine-tuning allows the pre-trained generative AI model 208 to be adapted to specific applications, such as text classification, language translation, or question answering.
[0050] Regularization: Regularization techniques such as random deactivation or weight decay can be applied to prevent overfitting.
[0051] Verification and evaluation: Use a validation dataset to select the best model checkpoint based on performance metrics relevant to tasks such as displaying AR / VR content. The model was evaluated on a separate test dataset to assess its generalization performance.
[0052] The trained generative AI model 208 is then deployed as part of the secure data vault system 132. In operation, the secure data vault system 132 provides various inputs to the generative AI model 208. For example, during AR / VR activities, camera data, audio, positioning information (e.g., gyroscope information, geolocation information), etc., can be provided as input 316 to the generative AI model 208. The generative AI model 208 then produces output 318, such as images, videos, text, audio, animations, stickers, filters, etc. The output 318 is then provided via a display (including an AR / VR display).
[0053] In some examples, such as about Figure 4 Furthermore, the trained generative AI model 208 will detect private information (both private images and audio) and block the use of private information.
[0054] Now turn to Figure 4 The accompanying diagram illustrates, according to some examples, the use of generative AI models to detect and block certain private information when using AR and / or VR devices. In the depicted example, user 402 is portrayed as wearable device 404, such as an AR and / or VR device. Device 404 includes one or more camera devices and one or more microphones. The device also includes or is connected to one or more speakers.
[0055] In use, device 404 captures images and sounds from a real-world environment 406. For example, user 402 views various objects 408, 410, 412, 414 that can be placed on a table. The techniques described herein identify two objects containing private information, such as passport 408 and credit card 414. To identify objects with private information, device 404 sends images and / or audio as input to generative AI model 208. Generative AI model 208 then uses its training to identify certain image portions or segments containing private information, such as objects 408 and 414, in the images used as input. In some examples, the images additionally include readable text containing private information. For example, the text may include financial terms (e.g., "credit card," "passport," "savings account," etc.). In other examples, the images do not include readable text, but generative AI model 208 performs detection based on certain patterns, such as identifying a driver's license based on the position and size of the photo on the card, identifying a credit card based on the logo and size of the card, identifying a passport based on the document size and color, etc.
[0056] Generative AI model 208 also uses audio input to determine whether private information is being observed. For example, based on training data, it can identify and mute shouts, fighting, showering, bathroom sounds, bedroom sounds, etc. In the depicted implementation, the real-world environment 406 is then transformed into an AR and / or VR environment 416. Non-private portions of the captured image (e.g., objects 412, 410) are now displayed as objects 420, 422, while objects 408, 414 are now displayed as obscured or otherwise blurred objects 418, 424. Any audio containing private information is similarly detected and muted. Thus, environment 416 provides enhanced privacy when shared with other users. In some examples, private information is brought to the attention of user 402 by providing additional cues such as written text on a display, spoken text, flashing obscured portions, etc. Similarly, audio signals such as beeps or voice prompts (e.g., "Audio muting is enabled for privacy") can be used to inform users that audio muting is being performed for enhanced privacy.
[0057] Figure 5 This is a block diagram illustrating, according to some examples, the use of generative AI models to detect and block certain private information based on location recognition when using AR and / or VR devices. In the depicted examples, user 402 is depicted as wearable device 404, for example... Figure 4 The AR and / or VR devices shown are illustrated. As mentioned above, device 404 includes one or more camera devices and one or more microphones. The device also includes or is connected to one or more speakers.
[0058] In use, device 404 captures images and sounds from real-world environments with private locations 502, 504, and / or 506. For example, user 402 may be in a public location that includes some private areas (e.g., hotels, airplanes, hospitals, schools, libraries, restaurants, bars, buses, cinemas, theaters, etc.). Private locations 502, 504, and / or 506 may include restrooms, other rooms (e.g., other bedrooms in a hotel), portions of a room (e.g., a hospital room may include multiple areas for multiple patients, and each area may be a private location), meeting rooms, certain workplace areas (e.g., offices, shared spaces), etc. Typically, private locations are those where the use of camera devices and / or microphones is not permitted.
[0059] The technique described herein identifies locations 502, 504, and 506 as private locations where camera devices are not permitted. To identify private locations, device 404 transmits images and / or audio to generative AI model 208 as input. Generative AI model 208 then uses training to identify images used as input that contain certain image portions or segments representing private locations where camera devices and / or microphones are not permitted, such as text (e.g., “restroom,” “meeting room,” “private,” “staff only,” “library,” etc.) and icons (e.g., restroom icon, “do not enter” icon, restaurant kitchen icon, etc.).
[0060] Generative AI model 208 also uses audio input to determine whether private information is being observed. For example, based on training data, it can identify and mute sounds such as those from a commercial kitchen, shouting, fighting, showering, bathroom, or bedroom. In the depicted implementation, images from private locations 502, 504, and 506 are now displayed as obscured or otherwise blurred. Any audio containing private information will be similarly detected and muted. Thus, location recognition provides enhanced privacy. In some examples, private information is brought to the attention of user 402 by providing additional cues such as written text on a display, spoken text, a flashing obscured portion, etc., thereby warning the user that they are viewing or have inadvertently entered private locations 502, 504, and 506. Similarly, audio signals such as beeps, voice prompts (e.g., “Muting audio for privacy”) can be used to inform user 402 that audio muting is being performed for enhanced privacy. System Architecture
[0061] Figure 6This is a block diagram illustrating further details of the interactive system 100 according to some examples. Specifically, the interactive system 100 is shown as including an interactive client 104 and an interactive server 124. The interactive system 100 includes multiple subsystems supported on the client side by the interactive client 104 and on the server side by the interactive server 124. In some examples, these subsystems are implemented as microservices. A microservice subsystem (e.g., a microservice application) may have components that enable it to operate independently and communicate with other services. Example components of a microservice subsystem may include: Functional logic: Functional logic implements the functions of the microservice subsystem and represents the specific capabilities or functions provided by the microservice. API Interface: Microservices can communicate with each other using well-defined APIs or interfaces, employing lightweight protocols such as REST or messaging. The API interface defines the inputs and outputs of a microservice subsystem, and how it interacts with other microservice subsystems within the interactive system 100. Data storage device: The microservice subsystem can be responsible for its own data storage device, which can take the form of a database, cache, or other storage mechanism (e.g., using database server 126 and database 128). This allows the microservice subsystem to operate independently of other microservices in the interactive system 100. Service discovery: Microservice subsystems can find and communicate with other microservice subsystems in the interacting system 100. The service discovery mechanism enables microservice subsystems to locate and communicate with other microservice subsystems in a scalable and efficient manner. Monitoring and logging: Monitoring and logging of microservice subsystems may be necessary to ensure availability and performance. Monitoring and logging mechanisms enable the tracking of the health and performance of microservice subsystems.
[0062] In some examples, the interactive system 100 may employ a monolithic architecture, a service-oriented architecture (SOA), a function-as-a-service (FaaS) architecture, or a modular architecture:
[0063] The example subsystem is discussed below.
[0064] The image processing system 602 provides various functions that enable users to capture and enhance (e.g., annotate or otherwise modify or edit) media content associated with a message.
[0065] The camera device system 604 includes (e.g., in a camera device application) control software that (e.g., directly or via an operating system) interacts with and controls the hardware camera device hardware of the user system 102 to modify and enhance real-time images captured and displayed via the interactive client 104.
[0066] Enhancement system 606 provides functionality related to the generation and distribution of enhancements (e.g., media overlays) for images captured in real time by the camera device of user system 102 or retrieved from the memory of user system 102. For example, enhancement system 606 is operable to select, present, and display media overlays (e.g., image filters or image lenses) for interactive client 104 to enhance real-time images received via camera device system 604 or stored images retrieved from memory 902 of user system 102. These enhancements are selected by enhancement system 606 and presented to the user of interactive client 104 based on some inputs and data, such as: The geographical location of user system 102; and User entity relationship information of users in user system 102.
[0067] Enhancements may include audio and visual content and visual effects. Examples of audio and visual content include images, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos or videos) at user system 102 for transmission in messages, or to video content such as video content streams or feeds sent from interactive client 104. Therefore, image processing system 602 can interact with and support various subsystems of communication system 608, such as messaging system 610 and video communication system 612.
[0068] Media overlays may include text or image data that can be superimposed on photographs taken by user system 102 or video streams produced by user system 102. In some examples, media overlays may be location overlays (e.g., Venice Beach), names of live events, or names of businesses (e.g., beach cafes). In other examples, image processing system 602 uses the geographic location of user system 102 to identify media overlays that include the names of businesses located at the geographic location of user system 102. Media overlays may include additional tags associated with businesses. Media overlays may be stored in database 128 and accessed through database server 126.
[0069] Image processing system 602 provides a user-based publishing platform that allows users to select geographical locations on a map and upload content associated with those locations. Users can also specify which media overlays should be provided to other users. Image processing system 602 generates a media overlay that includes the uploaded content and associates it with the selected geographical location.
[0070] The augmented reality creation system 614 supports augmented reality developer platforms and includes applications for content creators (e.g., artists and developers) to create and publish interactive clients 104, such as augmented reality experiences. The augmented reality creation system 614 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking technologies, and templates.
[0071] In some examples, the enhancement creation system 614 provides a merchant-based publishing platform that allows merchants to select specific enhancements associated with geographic locations via a bidding process. For instance, the enhancement creation system 614 associates the media overlay of the highest-bidding merchant with a corresponding geographic location for a predefined amount of time.
[0072] Communication system 608 is responsible for enabling and processing various forms of communication and interaction within interactive system 100, and includes messaging system 610, audio communication system 616, and video communication system 612. Messaging system 610 is responsible for enabling temporary or time-limited access to content by interactive client 104. Messaging system 610 includes (e.g., in a short-lived timer system) multiple timers that selectively enable access (e.g., for presentation and display) of messages and associated content via interactive client 104 based on duration and display parameters associated with a message or set of messages (e.g., a story). Audio communication system 616 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 104. Similarly, video communication system 612 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 104.
[0073] The user management system 618 is responsible for managing user data and profiles, and maintaining entity information about users of the interactive system 100 and the relationships between users (e.g., stored in entity table 708, entity diagram 710, and profile data 702).
[0074] The collection management system 620 is operationally responsible for managing collections or sets of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into “event libraries” or “event stories.” Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be available as a “story” for the duration of the concert. The collection management system 620 can also be responsible for publishing icons that notify the user interface of the interactive client 104 of the availability of specific collections. The collection management system 620 includes curation functions that enable collection managers to manage and curate specific content collections. For example, a curation interface enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 620 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users may be compensated for including user-generated content in a collection. In such cases, the collection management system 620 operates to automatically pay such users for using their content.
[0075] Map system 622 provides various geographic location (e.g., geolocation) functions and supports interactive client 104 in presenting map-based media content and messages. For example, map system 622 enables the display of user icons or avatars (e.g., stored in profile data 702) on the map to indicate the current or past locations of the user's "friends," as well as media content (e.g., a collection of messages including photos and videos) generated by such friends within the context of the map. For example, a message posted by a user from a specific geographic location to interactive system 100 can be displayed to the specific user's "friends" on the map interface of interactive client 104 within the context of the map at that specific location. The user can also share his or her location and status information (e.g., using an appropriate status avatar) with other users of interactive system 100 via interactive client 104, which is similarly displayed to the selected user within the context of the map interface of interactive client 104.
[0076] Game system 624 provides various game functions within the context of interactive client 104. Interactive client 104 provides a game interface with a list of available games that can be launched by a user within the context of interactive client 104 and played with other users of interactive system 100. Interactive system 100 also enables specific users to invite other users to participate in specific games by sending invitations from interactive client 104. Interactive client 104 also supports the sending and receiving of audio, video, and text messages (e.g., chat) within the game context, provides leaderboards for the game, and also supports providing in-game rewards (e.g., game currency and items).
[0077] External resource system 626 provides interactive client 104 with an interface for communicating with remote servers (e.g., third-party server 112) to launch or access external resources (i.e., applications or applets). Each third-party server 112 hosts applications or smaller versions of applications (e.g., game applications, utility applications, payment applications, or ride-sharing applications) based on markup languages (e.g., HTML5). Interactive client 104 can launch web-based resources (e.g., applications) by accessing HTML5 files from the third-party server 112 associated with the web-based resource. The application hosted by third-party server 112 is programmed in JavaScript using a software development kit (SDK) provided by interactive server 124. The SDK includes application programming interfaces (APIs) with functions that can be called or activated by the web-based application. Interactive server 124 hosts a JavaScript library that provides access to a given external resource for specific user data of interactive client 104. HTML5 is an example of a technology for programming games, but applications and resources programmed using other technologies can be used.
[0078] To integrate the SDK's functionality into the web-based resource, the third-party server 112 downloads the SDK from the interactive server 124, or the third-party server 112 otherwise receives the SDK. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interactive client 104 into the web-based resource.
[0079] The SDK stored on the interactive server system 110 effectively bridges the gap between external resources (e.g., application 106 or applet) and the interactive client 104. This provides users with a seamless experience communicating with other users on the interactive client 104 while preserving the appearance of the interactive client 104. To bridge communication between the external resources and the interactive client 104, the SDK facilitates communication between the third-party server 112 and the interactive client 104. A bridging script running on the user system 102 establishes two unidirectional communication channels between the external resources and the interactive client 104. Messages are sent asynchronously between the external resources and the interactive client 104 via these communication channels. Each SDK function activation is sent as a message and callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.
[0080] By using the SDK, not all information from the interactive client 104 is shared with the third-party server 112. The SDK limits which information is shared based on the needs of the external resources. Each third-party server 112 provides the interactive server 124 with an HTML5 file corresponding to the web-based external resource. The interactive server 124 can add a visual representation (e.g., box design or other graphics) of the web-based external resource to the interactive client 104. Once the user selects the visual representation or instructs the interactive client 104 to access the features of the web-based external resource through the interactive client 104's GUI, the interactive client 104 obtains the HTML5 file and instantiates the resource for accessing the features of the web-based external resource.
[0081] Interactive client 104 presents a graphical user interface (GUI) for an external resource (e.g., a login page or title screen). During, before, or after presenting the login page or title screen, interactive client 104 determines whether the initiated external resource has previously been authorized to access user data of interactive client 104. In response to determining that the initiated external resource has previously been authorized to access user data of interactive client 104, interactive client 104 presents another GUI of the external resource, including its functionality and characteristics. In response to determining that the initiated external resource has not previously been authorized to access user data of interactive client 104, after displaying the login page or title screen of the external resource for a threshold time period (e.g., 3 seconds), interactive client 104 slides up a menu (e.g., animates the menu to appear from the bottom of the screen to the middle of the screen or other parts) to authorize the external resource to access user data. This menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user selection of the accept option, interactive client 104 adds the external resource to the list of authorized external resources and allows the external resource to access user data from interactive client 104. External resources are authorized by the interactive client 104 to access user data under the OAuth 2 framework.
[0082] Interactive client 104 controls the type of user data shared with external resources based on the type of authorized external resource. For example, it provides access to a first type of user data (e.g., two-dimensional avatars of users with or without different avatar characteristics) to external resources including full-scale applications (e.g., application 106). As another example, it provides access to a second type of user data (e.g., payment information, two-dimensional avatars of users, three-dimensional avatars of users, and avatars with various avatar characteristics) to external resources including smaller versions of applications (e.g., web-based versions of applications). Avatar characteristics include different ways of customizing the appearance of an avatar (e.g., different poses, facial features, clothing, etc.).
[0083] The advertising system 628 enables third parties to purchase advertisements for presentation to end users via the interactive client 104, and also handles the delivery and presentation of these advertisements.
[0084] Artificial intelligence and machine learning system 630 provides various services to different subsystems within interactive system 100. For example, AI and machine learning system 630 operates in conjunction with image processing system 602 and camera device system 604 to analyze images and extract information such as objects, text, or faces. Image processing system 602 can then use this information to enhance, filter, or manipulate images. This information can also be used to train generative AI model 208 and to provide input to generative AI model 208. AI and machine learning system 630 can be used by augmentation system 606 to generate augmented content and augmented reality experiences, such as adding virtual objects or animations to real-world images. Communication system 608 and messaging system 610 can use AI and machine learning system 630 to analyze communication patterns and provide insights into how users interact with each other, and provide intelligent message classification and tagging, such as classifying messages based on sentiment or topic. AI and machine learning system 630 can also provide chatbot functionality for messaging interactions 120 between user systems 102 and between user systems 102 and interactive server system 110. The artificial intelligence and machine learning system 630 can also work with the audio communication system 616 to provide speech recognition and natural language processing capabilities, enabling users to interact with the interactive system 100 using voice commands. Data Architecture
[0085] Figure 7 This is a schematic diagram illustrating a data structure 700 that can be stored in a database 704 of an interactive server system 110, according to certain examples. Although the contents of the database 704 are shown as including multiple tables, it should be understood that data can be stored in other types of data structures, such as object-oriented databases.
[0086] Database 704 includes message data stored in message table 706. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and a payload. See below for reference. Figure 7 Further details describe information that can be included in the message and is contained within the message data stored in message table 706.
[0087] Entity table 708 stores entity data and (for example, links to entity diagram 710 and profile data 702). Entities for which records are maintained in entity table 708 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of the entity type, any entity for which the interactive server system 110 stores data can be an identifiable entity. Each entity is assigned a unique identifier and an entity type identifier (not shown).
[0088] Entity Graph 710 stores information about relationships and associations between entities. As an example only, such relationships can be social, professional (e.g., working in a common company or organization), interest-based, or activity-based. Some relationships between entities can be one-way, such as an individual user subscribing to digital content from a business or publishing user (e.g., a newspaper or other digital media channel or brand). Other relationships can be two-way, such as the "friendship" between various users of Interactive System 100.
[0089] Certain licenses and relationships can be attached to each relationship, and also to each direction of the relationship. For example, a two-way relationship (e.g., a friendship between individual users) can include authorization for the posting of digital content items between the individual users, but certain restrictions or filters can be imposed on the posting of such digital content items (e.g., based on content characteristics, location data, or time of day data). Similarly, a subscription relationship between an individual user and a business user can impose varying degrees of restrictions on the posting of digital content from the business user to the individual user, and can significantly restrict or prevent the posting of digital content from the individual user to the business user. As an example of an entity, a specific user can (e.g., through privacy settings) record certain restrictions in the records for that entity within entity table 708. Such privacy settings can be applied to all types of relationships in the context of interaction system 100, or selectively applied to certain types of relationships.
[0090] Profile data 702 stores various types of profile data about a specific entity. Based on privacy settings specified by the specific entity, profile data 702 can be selectively used and presented to other users of the interactive system 100. In the case of an individual, profile data 702 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. A specific user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interactive system 100 and on a map interface displayed to other users by the interactive client 104. The set of avatar representations may include "status avatars," which present a graphical representation of a status or activity that the user can choose to transmit at a specific time.
[0091] In the case that the entity is a group, in addition to the group name, members and various settings for the relevant group (e.g., notifications), the profile data 702 for the group may similarly include one or more avatar representations associated with the group.
[0092] Database 704 also stores enhancement data, such as overlays or filters, in enhancement table 712. Enhancement data is associated with and applied to videos (video data is stored in video table 714) and images (image data is stored in image table 716).
[0093] In some examples, filters are displayed as overlays on images or videos during presentation to the receiving user. Filters can be of various types, including user-selected filters from a set of filters presented to the sending user by the interactive client 104 while the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, geolocation filters specific to nearby or particular locations can be presented by the interactive client 104 within the user interface based on geographic location information determined by the Global Positioning System (GPS) unit of the user system 102.
[0094] Another type of filter is a data filter, which can be selectively presented to the sending user by the interactive client 104 during the message creation process based on other inputs or information collected by the user system 102. Examples of data filters include the current temperature at a specific location, the sending user's current speed, the battery life of the user system 102, or the current time.
[0095] Other augmented data that can be stored in image table 716 includes augmented reality content items (e.g., corresponding to an application "lens" or augmented reality experience). Augmented reality content items can be real-time special effects and sounds that can be added to images or videos.
[0096] Collection table 718 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or galleries). The creation of a specific collection can be initiated by a specific user (e.g., each user for whom records are maintained in entity table 708). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcast by that user. For this purpose, the user interface of interactive client 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.
[0097] The collection can also constitute a "live story," which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic technologies. For example, a "live story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a co-located event at a specific time can be presented with the option to contribute content to a specific live story, for example, via the user interface of interactive client 104. Live stories can be identified to a user by interactive client 104 based on their location. The end result is a "live story" told from a collective perspective.
[0098] Another type of content collection is called a "location story," which allows users of user system 102 located in a specific geographic location (e.g., on a college or university campus) to contribute to a specific collection. In some examples, contributions to a location story may employ secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).
[0099] As mentioned above, video table 714 stores video data, which in some examples is associated with messages for which records are maintained within message table 706. Similarly, image table 716 stores image data associated with messages whose message data is stored in entity table 708. Entity table 708 can associate various enhancements from enhancement table 712 with various images and videos stored in image table 716 and video table 714. Data communication architecture
[0100] Figure 8 This is a schematic diagram illustrating the structure of message 800 according to some examples, generated by interactive client 104 for transmission to another interactive client 104 via interactive server 124. The content of a particular message 800 is used to populate message table 706 within database 704 accessible by interactive server 124. Similarly, the content of message 800 is stored in memory as "in-transit" or "in-flight" data for user system 102 or interactive server 124. Message 800 is shown to include the following example components: Message Identifier 802: A unique identifier that identifies message 800. Message text payload 804: The text to be generated by the user via the user interface of user system 102 and included in message 800. Message image payload 806: Image data captured by the camera device component of user system 102 or retrieved from the memory component of user system 102 and included in message 800. The image data for the sent or received message 800 can be stored in image table 716. Message video payload 808: Video data captured by the camera device component or retrieved from the memory component of the user system 102 and included in message 800. The video data for the sent or received message 800 can be stored in image table 716. Message audio payload 810: Audio data captured by the microphone or retrieved from the memory component of the user system 102 and included in message 800. Message enhancement data 812: This represents enhancement data (e.g., filters, labels, or other annotations or enhancements) to be applied to the message image payload 806, message video payload 808, or message audio payload 810 of message 800. Enhancement data for the sent or received message 800 can be stored in enhancement table 712. Message duration parameter 814: A parameter value, in seconds, indicating the amount of time that the content of the message (e.g., message image payload 806, message video payload 808, message audio payload 810) should be presented to the user via the interactive client 104 or made accessible to the user. Message geolocation parameter 816: Geographic location data (e.g., latitude and longitude coordinates) associated with the message's content payload. Multiple message geolocation parameter 816 values may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 806 or a specific video within the message video payload 808). Message Story Identifier 818: An identifier value that identifies one or more sets of content (e.g., “story” identified in set table 718) associated with a specific content item in the message image payload 806 of message 800. For example, the identifier value can be used to associate multiple images within the message image payload 806 with multiple sets of content, respectively. Message Tag 820: Each message 800 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image depicts an animal (e.g., a lion) is included in the message image payload 806, a tag value can be included within the message tag 820 indicating the relevant animal. Tag values can be manually generated based on user input, or can be automatically generated using, for example, image recognition. Message sender identifier 822: An identifier (e.g., message sending system identifier, email address, or device identifier) indicating the user of the user system 102 on which message 800 is generated and from which message 800 is sent. Message receiver identifier 824: An identifier (e.g., message sending and receiving system identifier, email address, or device identifier) indicating the user of the user system 102 to which message 800 is addressed.
[0101] The content (e.g., values) of each component of message 800 can be pointers to locations in tables where content data values are stored. For example, image values in message image payload 806 can be pointers (or addresses) to locations within image table 716. Similarly, values in message video payload 808 can point to data stored in image table 716, values in message enhancement data 812 can point to data stored in enhancement table 712, values in message story identifier 818 can point to data stored in set table 718, and values in message sender identifier 822 and message receiver identifier 824 can point to user records stored in entity table 708. Systems with head-mounted devices
[0102] Figure 9 A system 900, including a head-worn wearable device 116 with a selector input device, is shown according to some examples. Figure 9 This is a high-level functional block diagram of an example head-mounted wearable device 116 that is communicatively coupled to mobile devices 114 and various server systems 904 (e.g., interactive server system 110) via various networks 108.
[0103] The head-mounted device 116 includes one or more camera devices, each of which may be, for example, a visible light camera 906, an infrared transmitter 908, and an infrared camera 910 that are communicatively and / or operatively coupled to the secure data vault system 132.
[0104] Mobile device 114 connects to head-mounted device 116 using both low-power wireless connection 912 and high-speed wireless connection 914. Mobile device 114 also connects to server system 904 and network 916.
[0105] The head-mounted device 116 also includes two image displays of the optical component's image display 918. The two image displays 918 of the optical component include one image display associated with the left lateral side of the head-mounted device 116 and one image display associated with the right lateral side of the head-mounted device 116. The head-mounted device 116 also includes an image display driver 920, an image processor 922, a low-power circuitry system 924, and a high-speed circuitry system 926. The image displays 918 of the optical component are used to present images and videos to the user of the head-mounted device 116, including images that may include a graphical user interface.
[0106] The image display driver 920 commands and controls the image display 918 of the optical components. The image display driver 920 can directly deliver image data to the image display 918 of the optical components for presentation, or it can convert image data into a signal or data format suitable for delivery to the image display device. For example, the image data can be video data formatted according to compression formats such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., while still image data can be formatted according to compression formats such as Portable Network Group (PNG), Joint Photo Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (EXIF), etc.
[0107] The head-mounted device 116 includes a frame and poles (or temples) extending laterally from the frame. The head-mounted device 116 also includes a user input device 928 (e.g., a touch sensor or button) comprising an input surface on the head-mounted device 116. The user input device 928 (e.g., a touch sensor or button) is used to receive input selections from a user for manipulating a graphical user interface of the presented image.
[0108] Figure 9 The components shown for the head-mounted device 116 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in the modules, frame, hinges, or nose bridge of the head-mounted device 116. The left and right visible light imaging devices 906 may include digital imaging device elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data, including images of scenes with unknown objects.
[0109] The head-mounted device 116 includes a memory 902 that stores instructions for performing a subset or all of the functions described herein. The memory 902 may also include a storage device.
[0110] like Figure 9 As shown, the high-speed circuit system 926 includes a high-speed processor 930, a memory 902, and a high-speed wireless circuit system 932. In some examples, the image display driver 920 is coupled to the high-speed circuit system 926 via a secure data vault system 132 and is operated by the display and rendering system 218 to drive the left and right image displays of the image display 918 of the optical components. The high-speed processor 930 can be any processor capable of managing the high-speed communication and operation of any general-purpose computing system required by the head-mounted device 116 and is included in the display and rendering system 218. The high-speed processor 930 includes the processing resources required to manage high-speed data transmission over the high-speed wireless connection 914 to a wireless local area network (WLAN) using the high-speed wireless circuit system 932. In some examples, the high-speed processor 930 executes the operating system of the head-mounted device 116 (e.g., a LINUX operating system) or other such operating system, and the operating system is stored in the memory 902 for execution. Among other duties, the high-speed processor 930, which executes the software architecture of the head-mounted device 116, manages data transmission with the high-speed wireless circuit system 932. In some examples, the high-speed wireless circuit system 932 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi®. In some examples, the high-speed wireless circuit system 932 may implement other high-speed communication standards.
[0111] The low-power wireless circuitry system 934 and high-speed wireless circuitry system 932 of the head-mounted device 116 may include short-range transceivers (Bluetooth™) and wireless wide-area network transceivers, wireless local area network transceivers, or wide-area network transceivers (e.g., cellular or Wi-Fi®). The mobile device 114—including transceivers communicating via low-power wireless connectivity 912 and high-speed wireless connectivity 914—can be implemented using the details of the architecture of the head-mounted device 116, as can other components of the network 916.
[0112] Memory 902 includes any storage device capable of storing various data and applications, including camera data generated by the left and right visible light imaging devices 906, the infrared imaging device 910, and the image processor 922, as well as images generated for display on an image display 918 in the optical components via an image display driver 920. While memory 902 is shown as integrated with high-speed circuitry 926, in some examples, memory 902 may be a separate, independent component of the head-mounted device 116. In some such examples, electrical wiring may provide a connection from the image processor 922 or the low-power processor 936 to memory 902 via a chip including a high-speed processor 930. In some examples, the high-speed processor 930 may manage addressing of memory 902, such that the low-power processor 936 will initiate the high-speed processor 930 whenever a read or write operation involving memory 902 is required.
[0113] like Figure 9 As shown, the low-power processor 936 or high-speed processor 930 of the head-mounted device 116 may be coupled to a camera device (visible light camera 906, infrared emitter 908, or infrared camera 910), an image display driver 920, a user input device 928 (e.g., a touch sensor or button), and a memory 902. In some embodiments, both the low-power circuitry 924 and the high-speed circuitry 926 are included in the secure data vault system 132.
[0114] The head-mounted device 116 is connected to a host computer. For example, the head-mounted device 116 is paired with the mobile device 114 via a high-speed wireless connection 914 or connected to the server system 904 via a network 916. The server system 904 may be one or more computing devices as part of a service or network computing system, for example, it includes a processor, memory, and network communication interfaces to communicate with the mobile device 114 and the head-mounted device 116 via the network 916.
[0115] Mobile device 114 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via network 916, low-power wireless connection 912, or high-speed wireless connection 914. Mobile device 114 may also store at least a portion of its instructions in its memory to implement the functions described herein. In some embodiments, low-power circuitry system 924 and high-speed wireless circuitry system 932 are included in secure network service 216 and are used only for downloading data, such as update packages, to secure data vault system 132. For example, the update package may include computer instructions configured to update the sandbox system, secure network service 216, and / or display and rendering system 218 to a newer version. The update package is encrypted, and a key stored in sandbox system 210 is then used to decrypt and verify the validity of the update package. In some examples, the key used to decrypt the update package is a Protocol Good Privacy (PGP) private key. Therefore, the PGP public key shared by sandbox system 210 is used to encrypt the update package. Verification can then be performed by reading the header of the update package after decryption. The header may include, for example, a cyclic redundancy check (CRC) code to verify the integrity of the instructions included in the update package.
[0116] The output components of the head-mounted device 116 include visual components such as a display, which may be, for example, a liquid crystal display (LCD), a plasma display panel (PDP), a light-emitting diode (LED) display, a projector, or a waveguide. The image display of the optical components is driven by an image display driver 920. The output components of the head-mounted device 116 also include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (e.g., user input device 928) of the head-mounted device 116, mobile device 114, and server system 904 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide position and force for touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0117] The head-mounted device 116 may also include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-mounted device 116. For example, peripheral device elements may include any I / O components, including output components, motion components, positioning components, or any other such elements described herein.
[0118] For example, biometric components include those for detecting expressions (e.g., hand gestures, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition). Biometric components may include brain-computer interface (BMI) systems that enable communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a computer-understandable format, and then using the resulting signals to control the device or machine.
[0119] Examples of BMI technology types include: BMI based on electroencephalography (EEG) uses electrodes placed on the scalp to record electrical activity in the brain. Invasive BMI, which uses electrodes surgically implanted into the brain. Optogenetics BMI uses light to control the activity of specific nerve cells in the brain.
[0120] Any biometric data collected by the biometric component is captured and stored only with the user's approval and is deleted upon the user's request. Furthermore, such biometric data can be used for very limited purposes, such as authentication. To ensure restricted and authorized use of biometric information and other personally identifiable information (PII), access to this data is limited to authorized personnel (if applicable). Any use of biometric data can be strictly limited to authentication purposes, and biometric data will not be shared or sold to any third party without the user's explicit consent. Additionally, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.
[0121] Moving components include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components (e.g., GPS receiver components) for generating position coordinates, Wi-Fi or Bluetooth™ transceivers for generating positioning system coordinates, altitude sensor components (e.g., altimeters or barometers that detect air pressure, from which altitude can be determined), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates can also be received from the mobile device 114 via low-power wireless connection 912 and high-speed wireless connection 914 through low-power wireless connection 934 or high-speed wireless connection 932.
[0122] Figure 10This is an example process 1000 suitable for use with the secure data vault system 132, according to some embodiments. In the depicted embodiments, the process 1000 receives sensor data at block 1002 from various sensors of the AR and / or VR system (e.g., user system 102), such as sensors communicatively coupled to user system 102 (e.g., camera devices, microphones, gyroscopes, navigation system sensors, biometric sensors, etc.). In some embodiments, the sensors are only coupled to the secure data vault system 132 and therefore only send sensor signals to the secure data vault system 132.
[0123] At box 1004, processing 1000 processes sensor data within sandbox system 210. As previously mentioned, sandbox system 210 can be a hardware-based system, such as an FPGA, one or more processors, and / or one or more custom chips. In operation, programs executable via sandbox system 210 (e.g., computer instructions) are isolated, allowing programs to crash without affecting the non-sandboxed system of the AR system. For example, sandbox system 210, via execution environment 212, disallows execution of computer instructions outside a specified memory address range (e.g., memory outside sandbox system 210), and computer instructions are executed only by sandbox system 210.
[0124] Additionally, a security compatibility matrix is used to restrict program execution within the sandbox system 210. For example, the security compatibility matrix includes rows storing unique process (e.g., computer program) identifier data and columns storing instructions that can be executed by the corresponding process identified by that process identifier data. That is, the columns include functions, methods, classes, and / or object-oriented objects that can be executed by each of the processes listed in the rows. In some examples, AI models are used to detect private information. For example, images, sounds, and / or geolocation information considered private can be detected via AI models. Then, processing of the sensor data at box 1004 can blur images, mute sounds and / or add noise to sounds, remove geolocation information, etc., to protect user privacy. Processing of the sensor data at box 1004 can also notify the user via text, images, and / or voice that private data has been detected and blurred.
[0125] Then, the processed data is rendered at box 1006. For example, an image can be rendered via display and rendering system 218 to blur private data via generative AI model 208. Similarly, sound can be rendered to inject noise or be muted to protect privacy. Output for AR / VR experiences is also generated, for example, via generative AI model 208. Then, at box 1008, the rendered image and / or sound are displayed along with the surrounding real-world environment. For example, a viewed driver's license view can now be displayed as blurred or occluded, while the surrounding real-world environment, such as the table on which the driver's license is placed, is displayed as unblurred. The image also includes rendered 3D images of virtual objects, avatars, etc., that can be overlaid on the surrounding real-world environment. By applying the secure data vault system 132, the techniques described herein improve security and privacy. Machine architecture
[0126] Figure 11This is a schematic representation of machine 1100, within which instructions 1102 (e.g., software, program, application, app, or other executable code) can be executed to cause machine 1100 to perform any or more of the methods discussed herein. For example, instructions 1102 can cause machine 1100 to perform any or more of the methods described herein. Instructions 1102 transform the general, unprogrammed machine 1100 into a specific machine 1100 programmed to perform the described and illustrated functions in the described manner. Machine 1100 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1100 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1100 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1102 specifying actions to be taken by machine 1100. Furthermore, although only a single machine 1100 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 1102 to perform any or more of the methods discussed herein. Machine 1100 may, for example, include user system 102 or any of a plurality of server devices forming part of interactive server system 110. In some examples, machine 1100 may also include both client and server systems, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of said particular method or algorithm are performed on the client side.
[0127] Machine 1100 may include a processor 1104, a memory 1106, and an input / output (I / O) unit 1108 that can be configured to communicate with each other via a bus 1110. In the example, processor 1104 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processors 1112 and 1114 that execute instruction 1102. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 11 Multiple processors 1104 are shown, but machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0128] Memory 1106 includes main memory 1116, static memory 1118, and memory cell 1120, all of which are accessible by processor 1104 via bus 1110. Main memory 1106, static memory 1118, and memory cell 1120 store instructions 1102 that implement any one or more of the methods or functions described herein. Instructions 1102 may also reside wholly or partially within main memory 1116, static memory 1118, machine-readable medium 1122 within memory cell 1120, at least one processor of processor 1104 (e.g., within the processor's cache memory), or any suitable combination thereof during execution by machine 1100.
[0129] I / O component 1108 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O component 1108 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. It should be recognized that I / O component 1108 may include... Figure 11Many other components are not shown. In various examples, I / O component 1108 may include user output component 1124 and user input component 1126. User output component 1124 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 1126 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens or other haptic input components that provide the position and force of a touch or touch gesture), audio input components (e.g., microphones), etc.
[0130] In other examples, I / O component 1108 may include biometric component 1128, motion component 1130, environmental component 1132, or positioning component 1134, as well as a wide range of other components. For example, biometric component 1128 includes components for detecting expressions (e.g., hand gestures, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), and identifying a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Biometric component may include a brain-computer interface (BMI) system that allows communication between the brain and an external device or machine. This can be achieved by recording brain activity data, converting that data into a computer-understandable format, and then using the resulting signals to control the device or machine.
[0131] Examples of BMI technology types include: BMI based on electroencephalography (EEG) uses electrodes placed on the scalp to record electrical activity in the brain. Invasive BMI, which uses electrodes surgically implanted into the brain. Optogenetics BMI uses light to control the activity of specific nerve cells in the brain.
[0132] Any biometric data collected by the biometric component will only be captured and stored with the user's approval and will be deleted upon the user's request. Furthermore, such biometric data will be used for very limited purposes, such as authentication. To ensure limited and authorized use of biometric information and other personally identifiable information (PII), only authorized personnel may access the data (if necessary). Any use of biometric data may be strictly limited to authentication purposes, and the data may not be shared or sold to any third party without the user's explicit consent. In addition, appropriate technical and organizational measures have been implemented to ensure the security and confidentiality of this sensitive information.
[0133] The moving part 1130 includes an acceleration sensor part (e.g., an accelerometer), a gravity sensor part, and a rotation sensor part (e.g., a gyroscope).
[0134] The environmental component 1132 includes, for example, one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases for safety purposes or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0135] Regarding the camera device, user system 102 may have a camera device system including, for example, a front-facing camera on the front surface of user system 102 and a rear-facing camera on the rear surface of user system 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., “selfies”) of the user of user system 102, which can then be enhanced with the aforementioned enhancement data (e.g., filters). For example, the rear-facing camera may be used to capture still images and videos in a more conventional camera device mode, wherein these images are similarly enhanced with enhancement data. In addition to the front-facing and rear-facing cameras, user system 102 may also include a 360° camera for capturing 360° photos and videos.
[0136] Furthermore, the camera system of user system 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even triple, quadruple, or quintuple rear camera configurations on the front and rear sides of user system 102. For example, these multi-camera systems may include wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.
[0137] The positioning component 1134 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure and can determine altitude based on air pressure), an orientation sensor component (e.g., a magnetometer), etc.
[0138] Various technologies can be used to implement communication. I / O component 1108 also includes a communication component 1136 operable to couple machine 1100 to network 1138 or device 1140 via a corresponding coupling or connection. For example, communication component 1136 may include a network interface component or another suitable device that interfaces with network 1138. In other examples, communication component 1136 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components for providing communication via other modalities. Device 1140 may be another machine or any peripheral device from a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0139] Furthermore, communication component 1136 can detect identifiers, or includes components operable to detect identifiers. For example, communication component 1136 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, QR codes such as Quick Response (QR) codes, Aztec codes, data matrices, and data symbols). TM The device can be equipped with optical sensors for multidimensional barcodes and other optical codes, such as MaxiCode, PDF417, UltraCode, and UCC RSS-2D barcodes, or acoustic detection components (e.g., microphones for identifying audio signals of the tags). Additionally, various information can be obtained via communication component 1136, such as location obtained via Internet Protocol (IP) geolocation, location obtained via Wi-Fi® signal triangulation, or location obtained by detecting NFC beacon signals that can indicate a specific location.
[0140] Various memories (e.g., main memory 1116, static memory 1118, and the memory of processor 1104) and storage unit 1120 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1102) cause various operations to implement the disclosed examples when executed by processor 1104.
[0141] Instructions 1102 can be sent or received over network 1138 using a transmission medium via a network interface device (e.g., the network interface component included in communication component 1136) and using any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1102 can be sent or received using a transmission medium via coupling to device 1140 (e.g., peer-to-peer coupling). Software Architecture
[0142] Figure 12 This is a block diagram 1200 illustrating a software architecture 1202 that can be installed on any one or more of the devices described herein. The software architecture 1202 is supported by hardware, such as a machine 1204 including a processor 1206, memory 1208, and I / O components 1210. In this example, the software architecture 1202 can be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1202 includes layers such as an operating system 1212, libraries 1214, frameworks 1216, and applications 1218. Operationally, application 1218 activates API calls 1220 via the software stack and receives messages 1222 in response to API calls 1220.
[0143] Operating system 1212 manages hardware resources and provides public services. Operating system 1212 includes, for example, a kernel 1224, services 1226, and drivers 1228. Kernel 1224 serves as an abstraction layer between hardware and other software layers. For example, kernel 1224 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. Services 1226 can provide other public services to other software layers. Drivers 1228 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1228 may include display drivers, camera drivers, Bluetooth® or Bluetooth® Low Energy drivers, flash memory drivers, serial communication drivers (e.g., USB drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.
[0144] Library 1214 provides common low-level infrastructure used by application 1218. Library 1214 may include system library 1230 (e.g., the C standard library), which provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1214 may include API library 1232, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing capabilities), and so on. Library 1214 may also include a wide variety of other libraries 1234 to provide many other APIs to application 1218.
[0145] Framework 1216 provides common high-level infrastructure for use by application 1218. For example, framework 1216 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. Framework 1216 can provide a wide range of other APIs that can be used by application 1218, some of which may be specific to a particular operating system or platform.
[0146] In the example, application 1218 may include home application 1236, contact application 1238, browser application 1240, book reader application 1242, location application 1244, media application 1246, messaging application 1248, game application 1250, and a wide variety of other applications such as third-party application 1252. Application 1218 is a program that performs the functions defined in the program. One or more applications of application 1218 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, third-party application 1252 (e.g., an application developed by an entity other than a platform-specific vendor using the Android™ or iOS™ Software Development Kit (SDK)) may be mobile software that runs on mobile operating systems such as iOS™, Android™, Windows® Phone, or other mobile operating systems. In this example, a third-party application 1252 can activate API call 1220 provided by the operating system 1212 to facilitate the functionality described herein. Glossary
[0147] "Carrier signal" refers to any intangible medium, such as a medium capable of storing, encoding, or carrying instructions to be executed by a machine and including digital or analog communication signals, or other intangible medium that facilitates the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.
[0148] "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.
[0149] "Communication network" means, for example, one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Common Old-Style Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, coupling can enable any data transmission technology of various types, such as single-carrier radio transmission technology (1xRTT), evolved data optimization (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.
[0150] A “component” refers to a logical or physical entity having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that partition or modularize a particular processing or control function. A component can be combined with other components via its interface to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components, or part of a program that typically performs a related function. A component can constitute a software component (e.g., code implemented on a machine-readable medium) or a hardware component. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to operate to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic permanently configured to perform certain operations. Hardware components can be dedicated processors, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Hardware components can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, and no longer a general-purpose processor. It will be appreciated that a decision may be made, for cost and time considerations, whether to implement a hardware component mechanically in a dedicated and permanently configured circuit or in a temporarily configured (e.g., software-configured) circuit. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to include tangible entities, i.e., entities physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain way or perform certain operations described herein. Consider the example of hardware components being temporarily configured (e.g., programmed), without requiring each of the hardware components to be configured or instantiated at any given time. For example, in cases where the hardware components include a general-purpose processor that is configured as a dedicated processor via software, this general-purpose processor can be configured at different times as its respective dedicated processor (e.g., including different hardware components). The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times. Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled.In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In examples where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing information in a memory structure accessible to the multiple hardware components and retrieving information from the memory structure. For example, a hardware component can perform an operation and store the output of that operation in a memory device communicatively coupled to it. Another hardware component can then access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed, at least in part, by one or more processors configured, either temporarily (e.g., by software) or permanently, to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented at least in part by processors, where a particular processor or one or more processors are examples of hardware. For example, at least some of the operations of the methods can be performed by one or more processors or processor-implemented components. Furthermore, one or more processors can also operate to support the execution of related operations in a “cloud computing” environment or as a “Software as a Service” (SaaS) operation. For example, at least some of the operations can be performed by a group of computers (as an example of a machine including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed among processors, not residing within a single machine, but deployed across multiple machines. In some examples, the processor or processor-implemented component can reside in a single geographic location (e.g., in a home environment, office environment, or server cluster). In other examples, the processor or processor-implemented component can be distributed across multiple geographic locations.
[0151] "Computer-readable storage medium" refers to both, for example, machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and can be used interchangeably in this disclosure.
[0152] A "brief message" is a message that is accessible for a limited time, such as a short period of time. Brief messages can be text, images, videos, etc. The access time for a brief message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is temporary.
[0153] "Machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Therefore, this term should be considered to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium," "device storage medium," and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium."
[0154] "Non-transitory computer-readable storage medium" means, for example, a tangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine.
[0155] "Signal medium" means any intangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.
[0156] "User equipment" means, for example, a device that is accessed, controlled, or owned by a user, with which the user interacts, performs actions or interactions, including interactions with other users or computer systems.
Claims
1. A system comprising: monitor; Camera device; as well as A secure data vault system, comprising: A sandbox system operatively coupled to the camera device and configured to receive camera device data from the camera device, wherein, in operation of the sandbox system, the camera device only sends its own camera device data to the sandbox system, and wherein the sandbox system includes an execution environment configured to restrict the execution of instructions to a predefined range of memory addresses; and A display and rendering system, operatively coupled to the sandbox system, is configured to render an image based on camera device data processed via the instructions and to display the image via the display, wherein the display and rendering system is configured to blur portions of the image based on positional information derived from the image.
2. The system of claim 1, comprising an artificial intelligence (AI) model configured to derive the location from the image.
3. The system according to claim 2, wherein, The AI model is included in the secure data vault system, and the execution environment is configured to restrict the execution of the AI model to the predefined memory address range.
4. The system according to claim 3, wherein, The AI model includes a generative AI model configured to derive the location information from the image as output based on a pattern learned during training of the generative AI model, wherein the pattern represents a private location where the use of a camera device is not permitted.
5. The system according to claim 4, wherein, The generative AI model includes: a variational autoencoder (VAE) model configured to learn a probabilistic mapping between data and a latent space via a training dataset; a generative adversarial network (GAN) model including a generator and a discriminator trained in a competitive manner; a recurrent neural network (RNN) model that computes a time step that depends on a previous time step; a Transformer model for sequential data generation; an autoencoder model for data compression and generation by learning to reconstruct input data from a lower-dimensional representation via the training dataset; or a combination thereof.
6. The system of claim 1, further comprising an augmented reality (AR) system, the AR system including the display and the camera device, wherein, The image includes a real-world environment, and the display is configured to blur the real-world environment based on the location information.
7. The system of claim 1, further comprising a virtual reality (VR) system, the VR system including the display and the camera device, wherein, The VR system is configured to display the image and a blurred portion of the image within a virtual environment.
8. The system of claim 1, further comprising a hardware-based system configured to execute the sandbox system within a virtual machine.
9. The system according to claim 1, wherein, The execution environment is configured to receive an identifier for a processing of a subset of requested execution instructions and to verify, based on the identifier, that the processing is permitted to execute the subset of instructions.
10. The system according to claim 9, wherein, The execution environment is configured to verify that the processing is allowed to execute a subset of the instructions based on the identifier and the data storage included in the secure data vault system.
11. The system according to claim 10, wherein, The data storage includes a security compatibility matrix, which comprises multiple rows and multiple columns.
12. The system according to claim 11, wherein, The plurality of rows include a plurality of unique processing identifier data, and wherein the plurality of columns include a plurality of instructions that can be executed by a corresponding process identified by the processing identifier data.
13. The system according to claim 1, wherein, The system includes a microphone and a speaker, wherein, during operation of the sandbox system, the microphone sends audio data only to the sandbox system, and wherein the sandbox system is configured to detect private information in the audio data, and wherein the sandbox system is configured to prevent portions of the audio data from being sent to the speaker based on the private information.
14. The system according to claim 1, wherein, The execution environment is configured to isolate the execution of a program comprising multiple program instructions, such that the crash of the program does not affect other systems included in the AR system.
15. The system according to claim 14, wherein, The execution environment is configured to disallow the execution of any of the plurality of program instructions that access a memory address outside a specified memory address range included in the predefined memory address range.
16. The system according to claim 1, wherein, The secure data vault system includes a secure network service configured to authenticate connections to external systems and, if the connection is authenticated, download update packages for the sandbox system from the external systems.
17. The system according to claim 16, wherein, The update package includes computer instructions configured to update the sandbox system to a newer version.
18. The system according to claim 1, wherein, The display and rendering system is configured to be operatively coupled only to the sandbox system and to display images based on data received only via the sandbox system.
19. A method comprising: Data is received from a camera device via a sandbox system, wherein, during operation of the camera device, the camera device sends data only to the sandbox system, and wherein the sandbox system includes an execution environment configured to restrict the execution of instructions to a predefined range of memory addresses. Based on the camera device data processed via the instructions, an image is rendered via a display and rendering system; and The image is displayed via a monitor, wherein the display and rendering system is configured to blur portions of the image based on positional information derived from the image.
20. A non-transitory machine-readable medium storing instructions that, when executed by a computer system, cause the computer system to perform operations, the operations including: Data is received from a camera device via a sandbox system, wherein, during operation of the camera device, the camera device sends data only to the sandbox system, and wherein the sandbox system includes an execution environment configured to restrict the execution of instructions to a predefined range of memory addresses. Based on the camera device data processed via the instructions, an image is rendered via a display and rendering system; and The image is displayed via a monitor, wherein the display and rendering system is configured to blur portions of the image based on positional information derived from the image.