Generating multi-sensory content based on user state
By generating multi-sensory content computing system, using the user's physiological and emotional states, the problems of non-scalability and resource-intensiveness of the existing technology are solved, flexible training tools and attention capture are realized, and training effects are improved.
Patent Information
- Application Number
- CN202480006050.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-28
- Filing Date
- 2024-02-20
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art has problems such as non-scalability, labor-intensive and resource-intensive in training sessions, and cannot fully attract users' attention, affecting the realization of training goals.
Using computer-implemented technology, multi-sensory content is generated based on the user's physiological state and emotional state, and the component maps prompt information to the output information through pattern filling, and using machine training models such as transformer-based models and reinforcement learning, the output system is controlled to deliver guide content.
It provides a flexible tool that can serve multiple training purposes, reduces the time and resource requirements for developing and maintaining independent systems, is user-friendly and able to capture and maintain attention, and improves training success rates.
Smart Images

Figure CN120513489A_ABST
Abstract
Description
Background Art
[0001] Attempts have been made to integrate computing technologies into training sessions. In some cases, the purpose of the training phase is to achieve a specific therapeutic goal. However, the solutions proposed are sometimes not scalable. That is, the solution is developed to serve the narrow purpose of a specific training environment and cannot be easily modified to serve other training environments. Furthermore, the development, maintenance, and use of individual solutions are sometimes labor-intensive and resource-intensive. Finally, the solution may not fully capture the user's attention, which negatively impacts their ability to achieve the stated goal of the training session. Summary of the Invention
[0002] This article describes a computer-implemented technique for providing personalized multisensory content based on input information that expresses a user's current state, including information expressing the user's physiological state and / or the user's experienced emotional state. The technique generates prompt information that describes the input information and the purpose of the guidance to be delivered. The technique then uses a pattern fill component to map the prompt information to output information. The output information includes control instructions for directing an output system to deliver guidance via multisensory content.
[0003] In some implementations, the purpose expressed in the prompt information is a guided treatment goal. In some examples, the treatment goal is: (a) stress reduction; or (b) meditation; or (c) sleep induction; or (d) promoting any of attention, mindfulness, and concentration (and / or reducing the time it takes to fall asleep); or (e) controlling specified emotions or impulses; or (f) memory management; or (g) the ability to complete tasks within a specified environment; or (h) improving productivity; or (in the context of the illustrated challenge or problem) any combination thereof.
[0004] In some implementations, the user's physiological state expresses any of the following: (a) one or more vital signs; or (b) brain activity; or (c) electrodermal activity; or (d) body movement; or (e) one or more eye-related characteristics; or (f) one or more voice-related characteristics; or (f) any combination thereof. In some implementations, the technology allows a user to self-report his or her current emotional state, for example, by entering "anxiety" when the user feels anxious.
[0005] In some implementations, the pattern filling component is a machine-trained model that maps text-based input information to text-based output information. In some cases, at least some of the text-based output information includes instructions for controlling one or more sensing devices in the form of comments, commands, numerical values (associated with corresponding settings), etc., or any combination thereof. In some examples, the machine-trained model is a transformer-based model that includes an attention mechanism.
[0006] In some implementations, the output system includes: (a) an audio output system for delivering audio content; or (b) a visual output system for delivering visual content; (c) a lighting system for modifying lighting in the user's environment; or (d) a scent output system for delivering scents; or (e) a tactile output system for delivering tactile experiences; or (f) an HVAC system for controlling heating, cooling, and / or ventilation; or (g) a workflow modification system; or (h) any combination thereof. In some cases, the audio content includes a narrative generated by the pattern filling component that is intended to advance the guided therapy goals.
[0007] In some implementations, the visual output system uses another machine-trained model to synthesize visual content based on the output information generated by the pattern filling component. The environment can present the synthesized visual content together with other sensory experiences delivered by other output systems (including audio output systems, scent output systems, tactile output systems, etc.).
[0008] In some implementations, another machine-trained model also processes the input information and / or the output information. This machine-trained model is trained through reinforcement learning to promote the guided goal, for example, by promoting specific content that advances the goal and penalizing other content that hinders the goal. In other implementations, the machine-trained model operates on its own without the need for a pattern-filling component.
[0009] Overall, the technology provides a flexible tool for serving many different training purposes using a general processing model. Thus, the tool eliminates or reduces the time-intensive and resource-intensive practice previously required to develop and maintain separate systems that serve specific training purposes. Furthermore, the tool is user-friendly in that the user can easily apply it to a specific training problem. That is, the user only needs to explicitly state the training goals and his or her current emotional state to successfully set up the tool for a specific training environment. Additionally or alternatively, the multisensory content is guided by inferences drawn from other sources, including sources that provide contextual information. Contextual information includes information about the task the user is performing, information about calendar events, information from various environmental sensing devices, and the like. Furthermore, the technology provides a sensory experience that is designed to capture and hold the user's attention, which increases the chances of training success.
[0010] This summary is provided to introduce a selection of concepts in a simplified form; these concepts are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A computing system for providing multisensory content for training purposes is shown.
[0012] Figure 2 Shown by Figure 1 Examples of multisensory content delivered by a computing system.
[0013] Figure 3 shows an example of a prompt generation component that is Figure 1 A component of a computing system.
[0014] Figure 4 Shown by Figure 3 An example of seed hint information generated by the hint generation component.
[0015] Figure 5 An example of another computing system for providing multisensory content incorporating the use of reinforcement learning is shown.
[0016] Figure 6 shows an example of a machine-trained pattern filling model, which is Figure 1 and Figure 5 Another component of a computing system.
[0017] Figure 7 An example of a machine-trained image synthesis model is shown. Figure 1 and Figure 5 Another component of a computing system.
[0018] Figure 8 Shown to provide Figure 1 and Figure 5 A process that provides an overview of one mode of operation of a computing system.
[0019] Figure 9 Shown to provide Figure 5 A process that provides an overview of one mode of operation of a computing system.
[0020] Figure 10 In some implementations, the Figure 1 and Figure 5 computing devices of a computing system.
[0021] Figure 11 A computing system of an illustrative type is shown that, in some implementations, is used to implement any aspect of the features shown in the preceding figures.
[0022] Throughout this disclosure and the accompanying drawings, like numerals are used to refer to similar components and features. Figure 1 Features appearing in the 200 series are numbers that were originally Figure 2 Features appearing in the 300 series are those originally Figure 3 Features that appear in, etc. Specific implementation method
[0023] Figure 1 An illustrative computing system 102 is shown for delivering multisensory content to a user via an output system 104. As will be described in more detail below, computing system 102 is implemented by one or more computing devices. In some cases, the entirety of computing system 102 is implemented by a local computing system. In other cases, aspects of computing system 102 are distributed between local and remote computing resources.
[0024] As a term, "machine-trained model" refers to the computer-implemented logic for performing a task using machine-trained weights generated during training operations. "Weight" is shorthand for parameter values. In some contexts, terms such as "component," "module," "engine," and "tool" refer to parts of computer-based technology that perform corresponding functions. Figure 10 and Figure 11 Examples of illustrative computing devices for performing these functions are provided.
[0025] The computing system 102 receives input information 106 from various input systems. For example, the input information 106 expresses one or more aspects of the user's current physiological state, as captured by the state sensing system 108. The sensing system 108 includes one or more sensing devices (110, 112, ...). Without limitation, the user's physiological state expresses: (a) one or more vital signs (e.g., captured by sensing devices that measure heart rate, blood pressure, respiratory rate, and / or body temperature); (b) one or more brain activity states (e.g., captured by an electroencephalogram sensing device); (c) one or more eye-related states (e.g., captured by a sensing device that measures pupil dilation or determines gaze direction); (d) one or more skin response states (e.g., captured by any type of skin conductance sensing device); (e) any body movement; (f) one or more voice-related states; or any combination thereof.
[0026] Input information 106 also includes information items explicitly specified by the user via user input system 114. One such information item describes the type of environment to be depicted in the multisensory content. Examples of environments include: (a) a beach-related setting; (b) a mountain-related setting; (c) a forest-related setting; (d) an urban setting; (e) a workplace-related setting (such as a surgeon's operating room); (f) any historical setting, etc. In some cases, the historical setting corresponds to a historical setting of common interest (e.g., Brooklyn, New York in the 1970s). In other cases, the historical setting corresponds to a specific user's personal environment (e.g., a depiction of the user's own home environment in Brooklyn, New York in the 1970s, with all the accompanying sensory stimuli and reminders of people associated with that previous environment). Computing system 102 provides an interface (not shown) that allows a specific user to specify private settings, for example, by providing a textual description of the private setting, keywords, an image, a video, an audio recording, a link to a news article, etc. Alternatively or additionally, computing system 102 creates the private setting on behalf of the specific user, with the user's explicit authorization. The computing system 102 can also obtain information from one or more knowledge bases, search engines, question answering services, chatbots, etc. in a build setting.
[0027] Another information item describes the user's self-reported emotional state. Another information item describes the purpose of the training. Without limitation, some illustrative goals of the training include: (a) stress reduction; or (b) meditation; or (c) inducing sleep; or (d) inducing any of attention, mindfulness, and focus (and / or reducing the time it takes to fall asleep); or (e) controlling specified emotions or impulses; or (f) memory management; or (g) the ability to complete tasks within a specified environment; or (h) increasing productivity; or (i) any combination thereof. Different implementations of the user input system 114 solicit information items from the user in various different ways. In some examples, the user input system 114 collects information items from the user via a user interface page, for example, as free-form text entries and / or as selections within a drop-down menu, etc.
[0028] In some examples, input information 106 includes other information from any other source 116. For example, in some implementations, input information 106 includes information items specifying the user's current location obtained from any location-determining device or service. One such location-determining device is a Global Positioning System (GPS) device. The user's location can also be inferred based on the user's proximity to beacons and cellular towers with known locations. Additionally or alternatively, input information 106 specifies the current time and / or the time available to perform a task. Additionally or alternatively, input information 106 specifies the nature of the user's current physical environment, for example, by specifying whether the user is interacting with computing system 102 at home or at the user's workplace. Additionally or alternatively, input information 106 includes any data obtained from any environmental sensor or detection device, including any weather sensor, any traffic sensor, any news feed information, etc. Additionally or alternatively, input information 106 expresses any characteristics of the user, including the user's age, location, interests, media preferences, etc. Each user is given the ability to authorize (or prohibit) the collection of such information and specify the conditions for its use.
[0029] Additionally or alternatively, the input information 106 expresses any information about a task that the user is currently performing or has recently performed. Additionally or alternatively, the input information 106 includes information extracted from a personal calendar or a shared calendar (with the user's explicit authorization), for example, related to events scheduled for execution by the user or other planned matters that may affect the user. Additionally or alternatively, the input information 106 includes information from an inference mechanism that predicts the user's current physiological and / or emotional state based on behavior exhibited by the user and / or other context-based signals (such as calendar events). For example, the input information 106 may include a prediction that the user may be stressed based on the number of meetings the user has scheduled, and / or the number of email messages the user has not yet read, and / or the number of upcoming deadlines the user is facing, etc. In some implementations, the computing system 102 implements an inference mechanism that is capable of performing prediction functions using any type of machine-trained model, examples of which are described below.
[0030] The prompt generation component 118 generates prompt information 120 based on the input information 106. In some implementations, the prompt generation component 118 performs this task by assembling portions of the input information 106 and other pre-generated content into a series of text tokens. Each token represents a word or a portion of a word. For example, a portion of the prompt information 120 includes the pre-generated text segment "My current heart beats per minute is." Assume that the heart rate sensing device indicates that the user's current heart rate is 68 beats per minute. The prompt generation component 118 appends the text "68bpm" to the end of the pre-generated text segment so that the text segment now reads "My current heart beats per minute is 68bpm."
[0031] More specifically, in some examples, the computing system 102 operates in an autoregressive manner over multiple passes. The prompt generation component 118 creates seed prompt information 122 during the first pass. The seed prompt information 122 describes the primary aspects of the training to be delivered to the user, for example, by specifying the purpose of the training, the type of environment to be depicted in the multisensory content, and the user's current physiological and emotional state. The prompt generation component 118 appends added information 124 to the end of the existing prompt information 120 during each subsequent pass. The added information 124 expresses the multisensory content delivered to the user, the user's updated physiological and / or emotional state, and the like.
[0032] The pattern filling component 126 analyzes the prompt 120 and generates predictions of the text that may follow the prompt 120. For example, suppose the seed prompt 122 consists of a series of text tokens S1, S2, S3, ..., S N The pattern filling component 126 successively determines a series of one or more text tokens (A1, A2, A3, ..., A2, ...) that may follow the seed hint information 122 across multiple subsequent traversals. K ), each time adding such a token to the existing series of tokens. This autoregressive operation is represented by arrow 128.
[0033] That is, in the first traversal, the pattern filling component 126 generates the added token A1. In the second traversal, the hint generation component 126 appends the token A1 to the end of the existing string of text tokens. The pattern filling component 126 then generates the text token A2. This process continues until the pattern filling component 126 generates a stop token, which is interpreted as a request to stop generating tokens. Upon completion, the pattern filling component 126 provides output information 130 that expresses all the text tokens added by the pattern filling component 126 (here, the sequence S1, S2, ..., S2) in the manner described above. N , A1, A2, …, A K). The mode filling component 126 repeats the above process when new input information is received, the new input information including new information reporting the user's current physiological state and / or current emotional state.
[0034] The pattern filling component 126 generates each fill based on knowledge of statistical patterns expressed in many other text segments. That is, the pattern filling component 126 will determine that a token string should be added to the end of an existing token string because the pattern has been observed to exist in many other text segments.
[0035] In some implementations, the pattern filling component 126 is implemented using a machine-trained model 132. The machine-trained model 132 maps an input text token string to an output text token string. The machine-trained model 132 operates based on weights learned by a training system (not shown) during a previous training process. The training process iteratively adjusts the weights while processing a large corpus of text segments, with the goal of accurately repeating the patterns exhibited by those text segments.
[0036] In some cases, the machine training model 132 is a transformer-based model. Figure 6 Further details of this type of model are set forth. A publicly available transformer-based model for performing pattern filling is the BLOOM model available from HUGGING FACE, INC. of New York, NY, the latest version of which is version 1.3 released on July 6, 2022. Other implementations of the computing system 102 use other types of machine-trained models, including fully connected feedforward neural networks (FFNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and the like.
[0037] Output information 130 includes control instructions for controlling output system 104. For example, output information 130 specifies content to be delivered to the user, such as by specifying narrative content to be displayed and / or read, the title of a song to be played, an identifier (e.g., a name) of another sound to be played, an identifier (e.g., a name) of one or more scents to be emitted, and the like. Furthermore, output information 130 includes specific commands, values, and / or other control-related data that govern the operation of output system 104. For example, output information 130 specifies characteristics of lighting, including any of its hue, (brightness) value, saturation, and the like, on a specified corresponding scale. Additionally or alternatively, output information 130 specifies characteristics of emitted sound, including any of its volume, bass, treble, balance, and the like, on a specified corresponding scale. Additionally or alternatively, output information 130 specifies characteristics of a scent using any specification system with respect to any scale(s). All of this information is derived from predictions made by pattern-filling component 126. For example, the pattern filling component 126 predicts that a specific song title follows the initial text because many other text segments have been observed to exhibit the same pattern.
[0038] Regarding the specific output modality of a scent, some specification systems allow users to express the persistence of a scent and its pleasantness. For example, in one illustrative specification system, a "base note" refers to a scent that typically lasts for more than 6 hours, a "middle note" refers to a scent that typically lasts for 3 to 6 hours, and a "high note" refers to a scent that typically dissipates within an hour. For example, the scent of vanilla is often used as a "base note." Individual scents can be combined with individual scents that have different respective persistence. The pleasantness of a scent measures its comfort and intensity on a specified scale.
[0039] The output system 104 uses multiple output systems to deliver multi-sensory content based on the output information 130. In some implementations, the output system 104 can also directly receive and operate on portions of the input information 106 described above, such as any portion of the state information provided by the state monitoring system 108, such as Figure 1, represented by the dashed line connecting the input information 106 with the output system 104. The output system 104 includes any of the following: (a) a visual output system 134; (b) an audio output system 136; (c) an odor output system 138; (d) a lighting system 140; (d) a tactile output system 142; and (e) an HVAC output system (not shown). The visual output system 134 delivers visual content, such as narrative text, images, and video content. The audio output system 136 delivers audio content, such as spoken narration, songs, and other sounds. In some implementations, the audio output system 136 incorporates 3D audio effect technology to create an immersive soundscape. One commercially available technology for providing 3D audio effects is 360REALITY AUDIO technology from SONY GROUP CORPORATION of Tokyo, Japan. The odor output system 138 delivers specific odors to the user's environment. One manufacturer of such a system is OVRTECHNOLOGY of Burlington, Vermont. Another manufacturer is OLORAMA TECHNOLOGY LTD. of Valencia, Spain. A lighting system 140 controls the lighting in the user's environment relative to the physical space (e.g., via overhead lighting). Additionally or alternatively, head-mounted device 226 controls the lighting of the virtual scene based on the output of lighting system 140. Any other visual output system 134 can also adjust one or more control settings based on the output of lighting system 140. A haptic output system 142 delivers a tactile user experience. One manufacturer of such a device is VR ELECTRONICS LTD. of London, England. An HVAC system (not shown) controls heating, cooling, and ventilation in the physical environment to match the simulated scene, for example, by creating warm conditions to match a tropical scene. Additionally or alternatively, the HVAC system and / or some other output system uses one or more fans to direct a breeze onto the user's body (e.g., the user's face). The computing system 102 incorporates wind into the multisensory content it delivers. For example, in some cases, computing system 102 uses fan(s) to simulate the type of wind expected in a particular environment. In other cases, computing system 102 uses directional airflow to help guide the intervention. For example, during a breathing exercise, computing system 102 may activate fan(s) when the user inhales and deactivate fan(s) when the user exhales.
[0040] One visual system is an image synthesis system 144. The image synthesis system 144 uses a machine trained model 146 to synthesize an output image based on the output information 130 provided by the pattern filling component 126. For example, assume that a portion of the output information 130 includes narrative text: "Sitting on the white sand of the beach, wearing my new sunhat." The image synthesis system 144 generates an image depicting a person sitting on the white sand beach wearing a sunhat. In other cases, the image synthesis system 144 synthesizes multimodal input information, such as by autoregressively mapping a combination of text and image information into new visual content. The concept of "autoregression" is explained in more detail below. In addition, in some implementations, the image synthesis system 114 produces dynamically changing visual content. That is, as the stream of output information 130 changes, the image synthesis system 114 dynamically changes its synthesized visual content to reflect the current state of the output information 130.
[0041] In some cases, the machine-trained model 146 performing this task is a diffusion model, which is described below in conjunction with Figure 7 Describe its examples. One publicly available image synthesis engine is provided by the CompVis research group at Ludwig-Maximilians-Universität in Munich, with its latest version being model 2.1, released on December 7, 2022. Another image synthesis system is the DALL-E system provided by OpenAI in San Francisco, California, which is capable of processing streams consisting of text and / or image content. Other types of machine-trained models can achieve the same goal, including transformer-based neural networks, CNNs, RNNs, and others. One way to train these other types of models is to use generative adversarial networks (GANs).
[0042] Visual output system 134 may include one or more other types of visual systems 148. One type of other visual output system selects from a store of pre-existing images and / or videos based on output information 130.
[0043] Computing system 102 provides a flexible tool for serving many different training objectives using a common processing model. Computing system 102 thereby eliminates or reduces the time-intensive and resource-intensive practice previously required to develop and maintain separate systems to serve specific training objectives. Furthermore, computing system 102 is user-friendly because users can easily apply it to specific training problems. To do so, the user need only clearly state the training goal and their current emotional state. Furthermore, computing system 102 provides a sensory experience designed to capture and maintain the user's attention, which increases the chances of training success.
[0044] It is also important to note that the output information 130 is personalized, which further enhances the computing system's ability to capture the user's attention. This means that the computing system 102 enables the user to construct seed prompt information 122 that contains features that describe the user's unique characteristics and needs. This seed prompt information 122 enables the pattern filling component 126 to also generate personalized output information 130.
[0045] Figure 2 Shown by Figure 1 202 illustrates an example of an immersive experience generated by computing system 102. In some examples, user input system 114 enables a user to begin a therapy session by entering three items of information via user interface page 204. As a first item, user input system 114 invites the user to specify a goal for the therapy. In this example, the user specifies that the goal is to "reduce stress." As a second item, user input system 114 asks the user to self-report their emotional state. In this example, the user specifies that they are currently feeling "anxious." As a third item, user input system 114 invites the user to describe the type of environment to be depicted in the multisensory content. In this example, the user specifies that they would like the therapy to be structured within the context of "beach" and "sailing." Optionally, the user can decline to specify the type of environment, in which case mode filler component 126 will select an environment based on other cues in prompt information 120. The user can also decline to specify their current emotional state, in which case mode filler component 126 will generally infer that the user requires the requested therapy.
[0046] In other implementations, user input system 114 asks the user to specify additional parameters that will govern the therapy session. For example, user input system 114 may invite the user to specify what type of media content will be used in the therapy session. Additionally or alternatively, computing system 102 asks the user to specify a technique for structuring the multisensory content. For example, suppose the user specifies the established "guided imagery" technique. In this approach, computing system 102 presents a series of images designed to evoke mental images, memories of sounds and smells, and the like, which can reactivate sensory perceptions previously associated with those memories. Alternatively or additionally, computing system 102 allows the user to invoke various mnemonic memory strategies. One such technique is the method of loci, which, guided by a sequence of sensory cues, enhances the user's ability to remember information in a specific order. For example, suppose the user's goal is to remember events in their calendar, or facts they will later be tested on. Computing system 102 creates a story using the memory palace technique with accompanying images to weave events and / or facts into a chronological story that is conveyed alongside the multisensory content.
[0047] In some implementations, the user input system 114 allows the user to enter free-form answers without limiting which ones are specified. In other implementations, the user input system 114 allows the user to select an information item from a drop-down menu of information items, etc. Additionally or alternatively, some implementations allow the user to enter information items in explicit verbal form and / or some other form, including non-verbal audible sounds (including sighs, expressive breathing sounds, etc.), sub-vocalization information and silent voice information (e.g., lip movement information), etc., or any combination thereof. Additionally or alternatively, some embodiments allow the user to enter information via gestures, such as by providing a thumbs-up or thumbs-down signal, a facial expression, etc.
[0048] The computing system 102 also collects information about the user's current physiological state. For example, Figure 2 One or more vital sign sensing devices 206 are shown providing vital sign information (including heart rate, blood pressure, respiratory rate, etc.). Brain activity sensing device 208 collects electroencephalogram information. Skin electrical sensing device 210 collects skin response information. One or more motion sensing devices 212 collect body motion information, including any of facial expressions, body posture, motion type, motion level, etc. The motion information may indicate the user's level of anxiety, the degree of the user's physical reaction to certain stimuli, etc. Eye state sensing device 214 collects information about the user's eyes, including pupil size, gaze direction, etc. Other implementations collect additional physiological state information. Additionally or alternatively, other implementations omit Figure 2 In some cases, computing system 102 provides a workstation that integrates all sensing devices. In other cases, computing system 102 uses a collection of independent sensing devices.
[0049] Despite Figure 2 Not shown, but some implementations of computing system 102 collect additional information, such as context-based information (including location and / or current time, etc.).
[0050] Assume that the pattern-filling component 126 generates prompt information 120 based on the above information items. The pattern-filling component 126 maps the prompt information 120 to output information 130. The output system 104 generates multisensory content based on the output information 130 and delivers the content via various output devices. As an overview, the multisensory content presents beach-related images, beach-related sounds, and beach-related smells, along with a narrative designed to reduce the user's stress level. The specific nature of this experience depends on the text tokens that make up the prompt information 120. Therefore, the prompt information 120 acts as an actual request for a specific type of experience.
[0051] More specifically, the audio output system 136 generates audible content via one or more speakers. For example, the audio output system 136 presents one or more beach-related sounds, such as the sound of waves crashing and / or the sounds of seabirds. Additionally or alternatively, the audio output system 136 presents a beach-themed song and presents the song to the user. In some cases, the song is an instrumental song, and the song is found without lyrics according to the instructions in the prompt information 120. Figure 2 The songs selected in the example of include melodies in the easy listening genre, for example, with a tropical feel (eg, including the sound of a slide guitar to evoke themes from traditional Hawaiian music).
[0052] Additionally or alternatively, the audio output system 136 reads a narrative generated by the pattern filling component 126. In some examples, the narrative begins with a description of the environment depicted in the image, such as: "I am relaxing on a beach at sunset. The sun is setting. A cool breeze blows my parasol. I hear the cries of seagulls in the distance. I smell the citrus aroma in my drink, mingling with the scent of the ocean." The narrative also contains content that more directly attempts to achieve a specified goal of reducing stress, such as: "My heart rate slows to 59 beats per minute. I take a deep breath, feeling more calm and composed. The worries gradually dissipate." For repetition, the pattern filling component 126 automatically synthesizes the narrative "from scratch" based on the prompt information 120, rather than selecting the narrative from a predefined script. Thus, the narrative reflects the patterns of speech in many text segments, but there is typically no pre-existing text that matches the entire delivered narrative.
[0053] Visual output system 134 presents visual content in the form of synthesized images and / or videos. Additionally or alternatively, visual output system 134 presents pre-existing images and / or videos specified in output information 130.
[0054] The visual output system 134 presents visual content via one or more display devices. Figure 2 As shown, visual output system 134 instructs laptop computing device 216 to present image 218 on its display device 220. Image 218 shows a picture of a beach scene, including a parasol, a sailboat, and a setting sun. Additionally or alternatively, visual output system 134 presents the visual content via wall display 222. In this example, wall display 222 shows another image 224 of the beach scene. In some cases, visual output system 134 sends digital data describing the visual content to wall display 222 for display by wall display 222. Alternatively, visual output system 134 projects the visual content as an image onto wall display 222.
[0055] Additionally or alternatively, the visual output system 134 presents information using other output devices, such as a virtual or mixed reality headset 226. One commercially available headset of this type is the HOLOLENS headset produced by MICROSOFT CORPORATION of Redmond, Washington. In some cases, virtual content is rendered as if projected onto a specific physical surface in the user's environment. Additionally or alternatively, the visual output system 134 presents information via an immersive room 228 in which the user can move and explore, sometimes referred to as an immersive cave (CAVE) environment. One provider of immersive environments is VISBOX INC. of St. Joseph, Illinois. Additionally or alternatively, the visual output system 134 presents visual content via a hologram (not shown) or the like.
[0056] Scent output system 138 emits beach-related scents, such as the ocean. Lighting system 140 displays a soft yellow light. Tactile output system 142 simulates the sensation of sand touching the skin, wind on the skin, and the like. Workflow modification system 230 controls the user's applications to limit interruptions, for example, by temporarily silencing alarms, prompts, reminders, and phone ringing. Other environments include other types of stimuli and / or omit one or more of the aforementioned forms of stimulation.
[0057] In some implementations, the multisensory content informs the user of his or her progress toward the goal of reducing stress. For example, the visual output system 134 displays the user's current heart rate. Additionally or alternatively, the audio output system 136 presents a narration that informs the user of his or her progress.
[0058] In other cases, computing system 102 presents images, sounds, and narration tailored to a particular work environment. For example, computing system 102 displays an operating room environment to train surgeons to cope with the stresses of that environment. This is an example of how multisensory content need not always be soothing; here, computing system 102 challenges the user and strengthens their ability to successfully navigate the operating room environment.
[0059] In other cases, healthcare providers apply computing system 102 to the problem of reactivating or enhancing long-term memories. To this end, computing system 102 provides the sounds, images, and smells that accompanied the initial formation of memories, hoping to restore those memories. This technology can allow users to more fully retain past events as long-term memories. Alternatively, the goal of treatment can be to address the opposite problem: memories that are too vivid and intrusive, and therefore impair the user's health. Therapy helps users cope with these memories, for example, by appropriately contextualizing them or, in some cases, desensitizing them.
[0060] In other cases, computing system 102 is primarily used for entertainment purposes. For example, a user may start a conversation by inputting a target environment (e.g., "Medieval England"), a target problem ("Defeat a mythical beast"), a target goal ("Excitement after success"), and a current emotional state ("Bored"). Thereafter, computing system 102 constructs and presents a narrative based on the input information. The narrative evolves based on dynamically changing input conditions, such as the user's current physiological and emotional state.
[0061] In all examples, computing system 102 also enables the user to enter additional descriptive words during the session, such as by typing "I am now ten feet tall" during game play. This has the effect of guiding the experience to incorporate new descriptive concepts. Figure 2 During a therapy session, the user may enter "stop showing the sailboat." The computing system 102 adds this information to the prompt information 120 in its current state. The mode filling component 126 interprets the added information as a request to omit the sailboat from the multi-sensory content.
[0062] Figure 3 Further illustrative details are shown regarding the hint generation component 118. As previously described, the hint generation component 118 generates hint information 120. The hint information 120 has two components: seed hint information 122 and added information 124.
[0063] The seed prompt information 122 includes various information items. In some examples, the seed prompt information 122 includes: a description of the training goal (e.g., "reduce stress"); a description of the desired environment (e.g., "beach"); the user's current self-reported emotional state (e.g., "anxiety"); and one or more physiological states. Other implementations include additional information items in the seed prompt information 122 and / or omission of Figure 3 One or more information items shown in .
[0064] The added information 124 includes various updates to the seed prompt information 122. In some examples, the added information 124 includes a description of the selected visual content, audio content, olfactory content, lighting settings, etc. specified by the output information 130. Additionally or alternatively, the added information 124 includes updated measurements from any device that measures the user's physiological state. The computing system 102 can be programmed to provide these updates in the form of prompts periodically (e.g., every second). Alternatively or additionally, the output information 130 generated by the pattern filling component 126 includes context-specific instructions to collect updated state information from any (multiple) sensing devices. Additionally or alternatively, the added information 124 includes an updated assessment of the user's current emotional state, for example, as self-reported by the user.
[0065] Figure 4 A more detailed example of seed prompt information 122 is shown. The first portion 402 of the seed prompt information 122 provides an overview of the training to be provided. It specifies a request for an immersive experience that includes a first-person narration. The immersive experience uses modulation of light, color, scent, and audio to achieve its effects. The first portion 402 also establishes that the user experience will be guided by the user's current heart rate.
[0066] The second portion 404 establishes the type of environment (e.g., "beach") in which the training is being conducted. The second portion 404 also establishes the specific goal of the training (e.g., "reduce stress"). The user provides information in the second portion 404, for example, by interacting with the user interface page 204. The third portion 406 uses the current reading provided by the heart rate sensing device to establish the user's current heart rate. The fourth portion 408 establishes the user's current emotional state (e.g., "anxiety"), as reported by the user via the user interface page 204. The fifth portion 410 specifies the format of the control instructions to be used in the output information 130. For example, the fifth portion 410 requests the output information 130 to generate light intensity settings provided in the range of 0.0 to 1.5. The sixth portion 412 specifies constraints governing the selection of songs, scents, and narration. The seed prompt information 122 forms a long sequence of text tokens.
[0067] The seed prompt 122 guides the schema filler component 126 to construct the output experience using guiding imagery. As described above, the computing system 102 can invoke many other techniques. For example, the following seed prompt uses the aforementioned mnemonics technique to help a user enhance their memory: "Prompt" = "Create a first-person story that changes based on the user's heart rate and how they feel." + "+userScener.text+" scene." + "Your goal is" + "Use guiding imagery to make this person" +userGoal.text+". Their current heart rate is" +HRtext+" per minute. Their current feeling is" +userFeelling.text+". The seed prompt continues: “Please incorporate the following items into your story in the order they appear: “+userMemory.text+”. Use the mnemonic method of traces based on this image: “+userScenys.text;* / ” (Note that the “+” symbol is a programmatic way of specifying string concatenation. In some implementations, the actual text fed to the pattern filling component 126 does not include the “+” symbol). The prompt generation component 118 merges this seed prompt with information collected from the condition monitoring system 108, the user input system 114, and / or the other input system(s) 116. The seed prompt set forth above specifically directs the pattern filling component 126 to use the mnemonic method of traces based on information extracted from specified sources (such as “userMemory.text” and “userScenery.text”).
[0068] Referring now to the added information 124, the form and content of this text varies based on the form and content of the seed prompt information 122 fed to the pattern fill component 126. Consider an illustrative case where the seed prompt information 122 includes the following sequence of text characters: "Prompt: Create an immersive experience in Unity that matches lighting, color, scent, and audio to a first-person story based on my heart rate, my feelings, and a beach scene. Your goal is to use guided imagery to help me relax. My current heart rate is 90 beats per minute. I am currently feeling anxious. Please return this as a string with the following format: {hue HSV}; {saturation HSV}; {value HSV}; {light intensity}; {sound}; {scent}; {first-person story}, where hue HSV = 0..1; saturation HSV = 0..1; value HSV = 0..1; light intensity = 0..1.5; sound = song title with no lyrics; scent = top-middle-bottom; story = ;".
[0069] Assume that a specific implementation of the pattern fill component 126 returns the following added information 124: "0.5; 0.5; 0.5; 1; "Soft Waves Crashing" by Kevin MacLeod; lavender, bergamot, vanilla; A feeling of peace washes over you as you close your eyes. You are standing on the beach, the sun warming your skin, and the sound of waves lapping against the shore gently lulling you into a state of blissful relaxation. You feel the tension in your body melt away with each deeper, slower breath. The air is filled with the sweet aroma of lavender, bergamot, and vanilla. You stop for a moment to appreciate the beauty of the moment, allowing yourself to gradually drift into a state of tranquility." Note that the actual text consumed and produced by the pattern fill component 126 may omit one or more spaces that appear in the above description.
[0070] The text "0.5; 0.5; 0.5; 1" in the added information 124 controls the lighting system 140. The pattern fill component 126 generates this information in response to the text "{Hue HSV}; {Saturation HSV}; {Value HSV}; {Lighting Intensity}" appearing in the seed hint information 122, which describes the type of lighting information expected, and the text "Hue HSV=0..1; Saturation HSV=0..1; Value HSV=0..1; Lighting Intensity=0..1.5," which describes the format of the expected lighting information.
[0071] The text "Soft Waves Crashing" in the added information 124 identifies a specific song to be played by the audio output system 136. The pattern fill component 126 generates this text in response to the portion of the seed hint information 122 specifying "{sound}" (which indicates that audio content is desired), and the text "sound = song title without lyrics" (which indicates that a song title (without lyrics) is desired).
[0072] The text "Lavender, Bergamot, Vanilla" in added information 124 describes the "scent notes" presented by scent output system 138, which, when combined, form a single scent. Pattern filling component 126 generates this text in response to the portion of seed hint information 122 specifying "{scent}" (which indicates the desired scent content) and the text "scent = top notes, middle notes, base notes" (which describes the desired format for expressing the scent). Alternatively or additionally, seed hint information 122 may prompt pattern filling component 126 to generate a scent using the format information "scent = adjective." In this format, pattern filling component 126 is directed to specify a descriptive name for the scent.
[0073] The text beginning with "When you close your eyes..." in the added information 124 controls the narrative delivered in visual and / or audio form. In response to the portion of the seed hint information 122 specifying "{first person story}" (which indicates that a first person story is desired), and the text "story =;" (which prompts the schema fill component 126 to begin generating the stream of text tokens that make up the story), the schema fill component 126 renders the text.
[0074] Note also that all text in added information 124 responds to the contextual information provided in the initial portion of seed hint information 122. For example, the narrative in added information 124 reflects the setting set forth in seed hint information 122, as well as the selection of soundscape, lighting, and scent.
[0075] Figure 5 An illustrative computing system 502 incorporating reinforcement learning is shown. The use of reinforcement learning allows the computing system 502 to modify its behavior over multiple sessions based on the user's reaction to the delivered content. That is, the computing system 502 reinforces the selection of some stimuli when it finds that the content advances the stated therapeutic goal. The computing system 502 negatively weights the selection of these stimuli when it finds that the content hinders the achievement of the goal. Conversely, Figure 1 The pattern filling component 126 used in the computing system 102 of , when operating alone, provides output information 130 that reflects patterns in text provided by many different people; however, the pattern filling component 126 does not capture and retain any information for measuring individual user reactions to the output information 130.
[0076] Environment 504 refers to the experience delivered by output system 506. Output system 506 is used to Figure 1 The multisensory content 508 is delivered in the same manner as the output system 104 of the embodiment of the present invention. The environment 504 also includes the user's reaction to the multisensory content 508.
[0077] Information source 510 and Figure 1 The illustrated state sensing system 108, user input system 114, and other input systems 116 correspond to an aggregate. Information source 510 generates input information 512 that reflects the current state of the environment 504, including the user's current physiological and emotional state. In other words, the input information 512 is Figure 1 The corresponding part of the input information 106.
[0078] The action determination component 514 generates output information 516 based on the input information 512. In some implementations, the action determination component 514 includes any action determination logic 518. In some examples, for example, the action determination logic 518 includes any type of neural network (FFN, CNN, RNN, transformer-based network, etc.) that is used to map the input information 512 to the output information 516. In this case, the action determination logic is driven by a set of machine-trained weights 520, which collectively define a policy.
[0079] The weight update component 522 uses reinforcement learning to update the weights 520. To perform this task, the reward evaluation logic 524 determines to what extent the current state (given by the input information 512) achieves the stated training goal (also given by the input information 512), including both short-term effects and expected long-term effects. For example, assume that the goal is to reduce stress. Further assume that the user's heart rate and the user's self-reported emotional state are criteria for assessing stress. Further assume that a low-stress state is associated with a heart rate below 65 bpm and a self-reported assessment of "calm", "stress-free" or equivalent. The reward evaluation logic 524 determines the extent to which the current state deviates from the ideal state and also determines the long-term expected results of the current state in terms of the treatment goal. The weight update component 522 then updates the weights 520 based on this determination. This has the effect of reinforcing the types of stimulation that reduce stress and penalizing the types of stimulation that increase stress. In some implementations, the weight update component 522 uses stochastic gradient descent combined with backpropagation to update the weights.
[0080] exist Figure 5 In the framework shown, the action determination component 518 operates as an actor in that it selects actions that change the environment 504. The reward evaluation logic 524 operates as a critic in that it evaluates how well the changes in the environment 504 advance the stated training objectives. Figure 5 ), the reward evaluation logic 524 can be driven by its own machine-trained model with its own set of weights that are separate from the weights 520 of the actors in the action determination logic 518. Here, the weight update component 522 also updates the weights of the reward evaluation logic 524 in each iteration. One approach for updating the weights of the actors and critics in this manner is a deep deterministic policy gradient (DDPG) system.
[0081] In the above-described manner of operation, the action determination component 514 relies solely on the action determination logic 518 to generate the output information 516. The action determination logic 518 is specifically trained to output control instructions that are within the expected ranges of the components of the output system 506.
[0082] In a second implementation, the action determination component 514 relies on the action determination logic 518 in conjunction with the prompt generation component 526 and the pattern filling component 528. Figure 1 The input information 512 is converted into prompt information in the same manner as the prompt generation component 118 of the embodiment. The pattern filling component 528 performs the same Figure 1 That is, the pattern filling component 528 uses the machine training model ( Figure 5 (not shown) maps the prompt information to output information 516. The machine-trained model used by the pattern filling component 528 includes weights that reflect patterns observed in a large corpus of text segments. However, these weights do not reflect whether a specific stimulus is promoting or hindering the goal in the current situation.
[0083] Different implementations combine the analysis of the action determination logic 518 and the pattern filling component 528 in different ways. In one approach, the action determination logic 518 provides its reinforcement learning-based recommendations to the prompt generation component 526 and the pattern filling component 528. When generating the output information 516, the pattern filling component 528 treats the information as just another text item to consider.
[0084] In another approach, action determination logic 518 operates on output information 516 generated by pattern filling component 528, evaluating whether output information 516 will achieve a positive or negative effect. If the latter is the case, action determination logic 518 modifies output information 516 or requests pattern filling component 528 to generate new output information. In other words, the first approach uses action determination logic 518 as a pre-processing component, while the second approach uses action determination logic 518 as a post-processing component. Other approaches to integrating these two functions are also possible.
[0085] Figure 6 Shown in Figure 1 An implementation of a machine-trained model 132 (referred to as a “model” for brevity) used in the pattern filling module 126 of Figure 5 The model 132 maps the initial prompt information 120 to the final output information 130. The model 132 is composed in part of a pipeline of transformer components, including a first transformer component 602. Figure 6 Details are provided regarding one way to implement the first transformer component 602. Although not specifically shown, the other transformer components of the model 132 have the same architecture as the first transformer component 602 and perform the same functions as the first transformer component 602 (but governed by a separate set of weights).
[0086] The model 132 begins by receiving prompt information 120. The prompt information 120 includes a series of language tokens 604. As used herein, a "token" or "text token" refers to a unit of text having any granularity, such as a single word, a word fragment generated by byte pair encoding (BPE), a character n-gram, a word fragment identified by the WordPiece algorithm, etc. For ease of explanation, it is assumed that each token corresponds to a complete word or other unit (e.g., a measurement of a sensing device, or a formatted value).
[0087] Next, the embedding component 606 maps the token sequence 604 into corresponding embedding vectors. For example, the embedding component 606 generates a one-hot vector describing the token and then uses a machine-trained linear transformation to map the one-hot vector into an embedding vector. The embedding component 606 then adds position information to the corresponding embedding vector to generate position-supplemented embedding vectors 608. The position information added to each embedding vector describes the position of the embedding vector in the embedding vector sequence 608.
[0088] The first transformer component 602 operates on the position-supplemented input vector 608. In some implementations, the first transformer component 602 includes, in sequence, an attention component 610, a first addition and normalization component 612, a feed-forward neural network (FFN) component 614, and a second addition and normalization component 616.
[0089] The attention component 610 performs attention analysis using the following equation:
[0090] The attention component 610 calculates the query weight matrix W by multiplying the position-supplemented embedding vector 608 (or in some applications, only the last position-supplemented embedding vector associated with the last received token) by Q To generate the query information Q. Similarly, the attention component 416 multiplies the position-supplemented embedding vectors by the key weight matrix W K Sum value weighted matrix W V To generate key information K and value information V. To execute equation (1), the attention component 416 performs a dot product operation on the query information Q and the transpose of the key information K, and then divides the dot product by the scaling factor to produce a scaled result. Wherein, the symbol d represents the dimension of Q and K. The attention component 610 performs a Softmax operation (normalized exponential function) on the scaled result and then multiplies the result of the Softmax operation with the value information V to generate attention output information. More generally, the attention component 610 determines how much emphasis should be placed on parts of the input information when interpreting other parts of the input information. In some cases, the attention component 610 is referred to as performing masked attention because the attention component 610 masks output token information that has not yet been determined at any given time. Background information on the general concept of attention is provided in Vaswani et al., "Attention Is All You Need," published in the 31st Conference on Neural Information Processing Systems (NIPS2017), 2017, 11 pages.
[0091] Notice, Figure 6 The attention component 610 is shown to be composed of multiple attention heads, including a representative attention head 618. Each attention head performs the calculation specified by equation (1), but for a different representation subspace than the subspaces of the other attention heads. To accomplish this, the attention heads perform the above calculations using a different set of query, key, and value weight matrices. Although not shown, the attention component 610 concatenates the outputs of the attention components of the separate attention heads and then multiplies the concatenated result by another weight matrix W O .
[0092] The addition and normalization component 612 includes a residual connection that combines (e.g., sums) the input information fed to the attention component 610 with the output information generated by the attention component 610. The addition and normalization component 612 then normalizes the output information generated by the residual connection, for example, by normalizing the values in the output information based on the mean and standard deviation of these values. Another addition and normalization component 616 performs the same function as the first-mentioned addition and normalization component 612.
[0093] The FFN component 614 uses a feed-forward neural network with any number of layers to transform input information into output information. In some implementations, the FFN component 614 is a two-layer network that performs its function using the following equation: FNN(x)=max(0,xW fnn1 +b1)W fnn2 +b2 (2)
[0094] Symbol W fnn1 and W fnn2 are two weight matrices used by the FFN component 614, which have (d, d fnn ) and (dfnn ,d) are complementary shapes. Symbols b1 and b2 represent deviation values.
[0095] The first transformer component 602 produces an output embedding 620. A series of other transformer components (622, ..., 624) perform the same function as the first transformer component 602, each operating on the output embedding produced by the immediately preceding transformer component. Each transformer component uses its own set of machine-trained weights at a specific level. The final transformer component 624 in model 132 produces a final output embedding 626.
[0096] The post-processing component 628 performs post-processing operations on the final output embedding 626 to produce the final output information 130. In one case, for example, the post-processing component 628 performs a machine-trained linear transformation on the final output embedding 626 and processes the result of the transformation using a Softmax component (not shown).
[0097] In some implementations, the model 132 operates in an autoregressive manner. To operate in this manner, the post-processing component 628 uses a Softmax operation to predict the next token (or in some cases, a set of most likely next tokens). The model 132 then appends the next token to the end of the input token sequence 604 to provide an updated token sequence. In the next pass, the model 132 processes the updated token sequence to generate the next output token. The model 132 repeats the above process until it generates a designated stop token.
[0098] In one implementation, a training system (not shown) trains the weights of model 132 based on a large corpus of text snippets. These text snippets belong to different knowledge domains and are not targeted to any particular domain. In some cases, the training system optionally fine-tunes the weights based on a corpus of text snippets that is particularly suitable for the training category provided by computing system 102. For example, this implementation fine-tunes the weights based on a corpus of documents belonging to various applicable medical-related fields, psychological-related fields, and social-related fields.
[0099] Figure 7 Shown in Figure 1 An example of a machine-trained model 146 (referred to as a “model” for brevity) used in an image synthesis system 144. Figure 7Specifically shown is the case where model 146 is a diffusion model that maps embedding 702 and key term 704 to image 706. Assume that a machine-trained model of any encoder type (not shown) has generated embedding 702 based on text term 708, as started by a randomly generated noise information instance. Key term 704 corresponds to the randomly generated noise information instance. Text term 708 is in turn generated at least in part by pattern filling component 126. In other words, text term 708 corresponds to a portion of output information 130 (see Figure 1 ).
[0100] In some implementations, as guided by the embedding 702, the model 146 successively transforms the key item 704 (which represents a noise sample) into an image 706 using a series of image generators (710, 712, 714). The first image generator 710 produces image information having a resolution R1. The second image generator 712 produces image information having a resolution R2, where R2>R1. The third image generator 714 produces image information having a resolution R3, where R3>R2, and so on. In some implementations, the diffusion model 146 implements each image generator using a U-Net component. For example, with respect to the representative second image generator 712, the U-Net component 716 includes a series of downsampling components 718 followed by a series of upsampling components 720. Each downsampling component or upsampling component itself includes any combination of subcomponents, including any of convolutional components, feedforward components, residual connections, attention components, etc. Skip connections 722 couple downsampling and upsampling components that perform processing on the same resolution level.
[0101] More specifically, the input to the machine training model 146 can take the form of prompt information generated by the pattern filling component 126. For example, in one case, the prompt information generated by the pattern filling component 126 takes the following form: "Hint string = + timer + "An equirectangular scene based on the following story:" + movingWindowInput + ".";". The preceding information "An equirectangular scene based on the following story:" prompts the machine training model 702 to provide images in the form of spherical / skybox images rather than 2D images.
[0102] Machine-trained model 146 dynamically changes its image based on changes in the variable "movingWindowInput." The text associated with the variable "movingWindowInput" in turn changes based on changes in the prompt information 120 fed to pattern-filling component 126. Consider a scenario where a user experiences a spike in stress level for some reason. The prompt information 120 fed to pattern-filling component 126 will change to reflect the updated state information expressing the user's stress spike. This change gradually propagates to affect output information 130, which then affects the image generated by image synthesis system 144.
[0103] Figure 8 A first process 802 is shown that provides Figure 1 An overview of the computing system 102. Although Figure 1 102, but process 802 is also applicable to Figure 5 More generally, process 802 is expressed as a series of operations performed in a specific order. However, the order of these operations is merely representative, and in other implementations, the operations can vary. In addition, any two or more operations described below can be performed in parallel. In one implementation, the relevant blocks shown in the flowchart that pertain to processing-related functions are represented by a combination of Figure 10 and Figure 11 The described hardware logic circuit system is implemented by one or more processors, computer-readable storage media, etc.
[0104] In block 804, computing system 102 receives input information (e.g., input information 106), the input information expressing a physiological state of a user and / or an emotional state experienced by the user (e.g., from user input system 114), the physiological state of the user being obtained from a state sensing system (e.g., state sensing system 108). In block 806, computing system 102 generates prompt information (e.g., prompt information 120) describing the input information and the purpose of the guidance to be delivered. In block 808, computing system 102 uses a mode fill component (e.g., mode fill component 126) to map the prompt information to output information (e.g., output information 130), the output information including control instructions for controlling an output system (e.g., output system 104) to deliver the guidance via the generated content. In block 810, computing system 102 provides the output information to the output system.
[0105] Figure 9 A second process 902 is shown that provides Figure 5 An overview of the computing system 502 is provided. Figure 8 The same generalizations provided apply to Figure 9Process 902. Figure 9 The process shown is one of many possible processes.
[0106] In block 904, the computing system 502 receives input information (e.g., input information 512) representing a physiological state of a user from a state sensing system (e.g., state sensing system 108) and an emotional state experienced by the user (e.g., obtained from user input system 114). In block 906, the computing system 502 uses a machine-trained model (e.g., using action determination logic 518) to map the input information to output information (e.g., output information 516), the output information comprising control instructions for controlling an output system (e.g., output system 506) to deliver guidance via generated content (e.g., multisensory content 508). The machine-trained model has been trained through reinforcement learning to generate instances of output information that promote the identified therapeutic goal of the guidance. In block 908, the computing system provides the output information to the output system for delivery of the guidance.
[0107] Figure 10 In some implementations, the Figure 1 computing system 102 or Figure 5 10. The computing device 1002 of the computing system 502 is shown. The computing device 1002 includes a set of local devices 1004 coupled to a set of servers 1006 via a computer network 1008. Each local device corresponds to any type of computing device, including any of a desktop computing device, a laptop computing device, any type of handheld computing device (e.g., a smartphone or tablet computing device), a mixed reality device, a smart appliance, a wearable computing device (e.g., a smartwatch), an Internet of Things (IoT) device, a gaming system, an immersive "CAVE", a media device, an in-vehicle computing system, any type of robotic computing system, a computing system in a manufacturing system, etc. In some implementations, the computer network 1008 is implemented as a local area network, a wide area network (e.g., the Internet), one or more point-to-point links, or any combination thereof.
[0108] Figure 10 The dashed boxes in the figure indicate that the functionality of the computing system (102, 502) can be distributed in any manner across the local devices 1004 and / or the server 1006. For example, in some cases, each local computing device or a group of associated local computing devices implements the entire computing system (102, 502). In other implementations, one or more data processing operations are delegated to online resources provided by the server 1006. For example, in some implementations, the pattern filling component (126, 528) is implemented by the server 1006.
[0109] Figure 111 shows a computing system 1102, which, in some implementations, is used to implement any aspect of the mechanisms described in the above figures. For example, in some implementations, Figure 11 The type of computing system 1102 shown in FIG. 1 is used to implement Figure 10 100. In all cases, computing system 1102 represents a physical and tangible processing mechanism.
[0110] Computing system 1102 includes a processing system 1104, which includes one or more processors. The processor(s) include one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs) and / or one or more application-specific integrated circuits (ASICs) and / or one or more neural processing units (NPUs), etc. More generally, any processor corresponds to a general-purpose processing unit or a special-purpose processor unit.
[0111] The computing system 1102 also includes a computer-readable storage medium 1106 corresponding to one or more computer-readable medium hardware units. The computer-readable storage medium 1106 retains any kind of information 1108, such as machine-readable instructions, settings, model weights, and / or other data. In some implementations, the computer-readable storage medium 1106 includes one or more solid-state devices, one or more magnetic hard disks, one or more optical disks, magnetic tape, etc. Any instance of the computer-readable storage medium 1106 uses any technology for storing and retrieving information. In addition, any instance of the computer-readable storage medium 1106 represents a fixed or removable unit of the computing system 1102. In addition, any instance of the computer-readable storage medium 1106 provides volatile and / or non-volatile retention of information.
[0112] More generally, any storage resource or any combination of storage resources described herein will be considered a computer-readable medium. In many cases, a computer-readable medium represents some form of physical and tangible entity. The term computer-readable medium also includes propagated signals, such as those transmitted or received via a physical channel and / or air or other wireless medium. However, the specific term "computer-readable storage medium" or "storage device" explicitly excludes propagated signals themselves in transmission, while including all other forms of computer-readable media; in this regard, the computer-readable storage medium or storage device is "non-transitory."
[0113] The computing system 1102 utilizes any instance of the computer-readable storage medium 1106 in various ways. For example, in some implementations, any instance of the computer-readable storage medium 1106 represents a hardware memory unit (such as random access memory (RAM)) for storing information during execution of a program by the computing system 1102, and / or represents a hardware storage unit (e.g., a hard disk) for longer-term retention / archiving of information. In the latter case, the computing system 1102 also includes one or more drive mechanisms 1110 (e.g., a hard disk drive mechanism) for storing and retrieving information from an instance of the computer-readable storage medium 1106.
[0114] In some implementations, the computing system 1102 performs any of the functions described above when the processing system 1104 executes computer-readable instructions stored in any instance of the computer-readable storage medium 1106. For example, in some implementations, the computing system 1102 executes computer-readable instructions to perform the reference Figure 9 and Figure 10 Each box of the process is described. Figure 11 The hardware logic circuitry 1112 is generally indicated to include any combination of the processing system 1104 and the computer-readable storage media 1106 .
[0115] Additionally or alternatively, the processing system 1104 includes one or more other configurable logic units that use a collection of logic gates to perform operations. For example, in some implementations, the processing system 1104 includes a fixed configuration of hardware logic gates, e.g., that are created and set at the time of manufacture and cannot be changed thereafter. Additionally or alternatively, the processing system 1104 includes a collection of programmable hardware logic gates that are configured to perform different application-specific tasks. The latter class of devices includes programmable array logic devices (PALs), general array logic devices (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), etc. In these implementations, the processing system 1104 effectively incorporates a storage device that stores computer-readable instructions because the configurable logic unit is configured to execute the instructions and thus embodies or stores the instructions.
[0116] In some cases (e.g., where computing system 1102 represents a user computing device), computing system 1102 also includes an input / output interface 1114 for receiving various inputs (via input devices 1116) and for providing various outputs (via output devices 1118). Illustrative input devices include a keyboard device, a mouse input device, a touch screen input device, a digitizer tablet, one or more still image cameras, one or more video cameras, one or more depth camera systems, one or more microphones, a voice recognition mechanism, any location determination device (e.g., a GPS device), any motion detection mechanism (e.g., an accelerometer and / or a gyroscope), and the like. In some implementations, one specific output mechanism includes a display device 1120 and an associated graphical user interface presentation (GUI) 1122. Display device 1120 corresponds to a liquid crystal display device, a light emitting diode display (LED) device, a cathode ray tube device, a projection mechanism, and the like. Other output devices include a printer, one or more speakers, a tactile output mechanism, an archiving mechanism (for storing output information), and the like. In some implementations, computing system 1102 also includes one or more network interfaces 1124 for exchanging data with other devices via one or more communication channels 1126. One or more communication buses 1128 communicatively couple the above elements together.
[0117] Communication channel(s) 1126 can be implemented in any manner, for example, via a local area computer network, a wide area computer network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication channel(s) 1126 include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0118] Figure 11 Computing system 1102 is shown as being composed of a discrete collection of separate units. In some cases, the collection of units corresponds to independent hardware units provided in a computing device chassis having any form factor. Figure 11 In other cases, the computing system 1102 includes an integrated Figure 1 For example, in some implementations, the computing system 1102 includes a system on a chip (SoC or SOC) that is combined with a Figure 11 The functions of two or more units shown in the figure correspond to an integrated circuit.
[0119] The following overview provides a set of illustrative examples of the techniques described herein.
[0120] (A1) According to a first aspect, a computer-implemented method (e.g., 802) for providing content is described. The method includes: receiving (e.g., 804) input information (e.g., 106), the input information expressing a physiological state of a user and / or an emotional state experienced by the user, the physiological state of the user being obtained from a state sensing system; generating (e.g., 806) prompt information (e.g., 120) describing the input information and the purpose of the guidance to be delivered; mapping (e.g., 808) the prompt information to output information (e.g., 130) using a pattern fill component (e.g., 126), the output information including control instructions for controlling an output system (e.g., 104) to deliver the guidance via the generated content; and providing (810) the output information to the output system.
[0121] (A2) According to some implementations of method A1, the user's physiological state is expressed by: (a) vital signs; or (b) skin electrode activity; or (c) body movement; or (d) eye-related characteristics; or (e) voice-related characteristics; or (f) any combination thereof.
[0122] (A3) Some implementations of method A1 or A2, wherein the emotional state is self-reported by the user.
[0123] (A4) Some implementations of any individual method according to methods A1 to A3, wherein the prompt information further describes the selected environment.
[0124] (A5) According to some implementations of method A4, the therapeutic goal is: (a) stress reduction; or (b) meditation; or (c) inducing sleep; or (d) promoting any of attention, mindfulness, and concentration (and / or reducing the time it takes to fall asleep); or (e) controlling specified emotions or impulses; or (f) memory management; or (g) the ability to complete tasks within a specified environment; or (h) increasing productivity; or (i) any combination thereof.
[0125] (A6) Some implementations of any individual method according to methods A1 to A5, wherein the prompt information further describes the selected environment.
[0126] (A7) Some implementations of any individual method of methods A1 to A6, wherein the prompt information is a series of input text tokens, wherein the pattern filling component is a machine-trained model, and wherein the output information is a series of output text tokens.
[0127] (A8) According to some implementations of method A7, the machine-trained model is a transformer-based machine-trained neural network.
[0128] (A9) According to some implementations of method A7, at least some of the input text tokens describe the type of output modality to be used, and at least some of the input text tokens describe the format of control information to be provided in the output text tokens.
[0129] (A10) According to some implementations of method A7, the input information and / or the output information is also processed by another machine-trained model, which is trained through reinforcement learning to facilitate the purpose of guidance.
[0130] (A11) According to some implementations of any individual method of methods A1 to A10, the output system includes a machine-trained model for mapping the output information to visual content.
[0131] (A12) According to some implementations of any individual method of methods A1 to A11, the content is multisensory content.
[0132] (A13) According to some implementations of any individual method of Methods A1 to A12, the output system includes: (a) an audio output system for delivering audio content; or (b) a visual output system for delivering visual content; or (c) a lighting system for controlling lighting; or (d) a scent output system for delivering scent; or (e) a tactile output system for delivering a tactile experience; or (f) an HVAC system for controlling heating, cooling and / or ventilation, or (g) a workflow modification system for controlling a user's workflow; or (h) any combination thereof.
[0133] (A14) According to some implementations of any individual method of methods A1 to A13, the method also includes updating the prompt information to include aspects of the output information and updated state information to provide updated prompt information, and mapping the updated prompt information to the updated output information using a pattern fill component.
[0134] (B1) According to a second aspect, another computer-implemented method (e.g., 902) for providing content is described. The method includes: receiving (e.g., 904) input information (512) expressing a physiological state of a user and an emotional state experienced by the user, the physiological state of the user being obtained from a state sensing system; mapping (906) the input information to output information (e.g., 516) using a machine-trained model (e.g., 518), the output information comprising control instructions for controlling an output system (e.g., 506) to deliver a guidance via generated content; and providing (e.g., 908) the output information to the output system for delivering the guidance. The machine-trained model has been trained by reinforcement learning to generate instances of the output information, the instances of the output information promoting an identified therapeutic goal of the guidance.
[0135] (B2) According to some implementations of method B1, the operation further includes generating prompt information describing the input information and the guided treatment target to be delivered, and generating output information by mapping the prompt information to candidate output information using a pattern filling component.
[0136] In yet another aspect, some implementations of the techniques described herein include a computing system (e.g., computing system 1102) that includes a processing system (e.g., processing system 1104) having a processor. The computing system also includes a storage device (e.g., computer-readable storage medium 1106) for storing computer-readable instructions (e.g., information 1108). The processing system executes the computer-readable instructions to perform any method described herein (e.g., any individual method in Methods A1 to A14, Methods B1, or B2).
[0137] In yet another aspect, some implementations of the techniques described herein include a computer-readable storage medium (e.g., computer-readable storage medium 1106) for storing computer-readable instructions (e.g., information 1108). A processing system (e.g., processing system 1104) executes the computer-readable instructions to perform any of the operations described herein (e.g., the operations in any individual method of Methods A1 to A14, Methods B1, or B2).
[0138] More generally, any individual elements and steps described herein may be combined into any logically consistent arrangement or subset. Furthermore, any such combination may be embodied as a method, apparatus, system, computer-readable storage medium, data structure, article of manufacture, graphical user interface presentation, etc. The technology may also be expressed as a series of means-plus-function elements in the claims, but unless the phrase "means for..." is explicitly used in a claim, it should not be construed as invoking this format.
[0139] With respect to the terminology used in this specification, the phrase "configured to" encompasses various physical and tangible mechanisms for performing the identified operation. These mechanisms may be configured to use Figure 11 The term "logic" also encompasses various physical and tangible mechanisms for performing tasks. For example, Figure 11 Each process-related operation illustrated in the flowcharts of 1 and 12 corresponds to a logical component for performing the operation.
[0140] The description may have identified one or more features as optional. Such statements should not be interpreted as an exhaustive indication of features that may be considered optional; unless otherwise stated, any feature should be considered optional, even if not explicitly identified in the text. Furthermore, any reference to a single entity is not intended to exclude the use of multiple such entities; similarly, the description of multiple entities in the specification is not intended to exclude the use of a single entity. Thus, the statement that a device or method has feature X does not exclude the possibility of it having additional features. Furthermore, unless otherwise stated, any features described as alternative ways of implementing an identified function or an identified mechanism may be combined in any combination.
[0141] In terms of specific terms, the term "multiple" or "a plurality of" or any plural form of a term (without the explicit use of "multiple" or "a plurality of") refers to two or more items and does not necessarily imply "all" items of a particular kind, unless explicitly specified otherwise. Unless otherwise specified, the term "at least one of" refers to one or more items; in the absence of an explicit recitation of "at least one of," etc., reference to a single item is not intended to exclude the inclusion of multiple items. In addition, unless otherwise specified, the descriptors "first," "second," "third," etc. are used to distinguish different items and do not imply an ordering between items. The phrase "A and / or B" means A, or B, or A and B. The phrase "any combination thereof" refers to any combination of two or more elements in a list of elements. In addition, the terms "including," "comprising," and "having" are open terms used to identify at least a portion of a larger whole, but not necessarily all parts of the whole. A "set" includes zero members, one member, or more than one member. Finally, the term "exemplary" or "illustrative" refers to one implementation among potentially many implementations.
[0142] Finally, the functionality described herein can employ various mechanisms to ensure that any user data is processed in a manner consistent with applicable laws, social norms, and the expectations and preferences of individual users. For example, the functionality can be configured to allow users to explicitly opt-in (and then explicitly opt-out) of the functionality's provisions. The functionality can also be configured to provide appropriate security mechanisms to ensure the privacy of user data (such as data sanitization mechanisms, encryption mechanisms, and / or password protection mechanisms).
[0143] Furthermore, the description may have presented various concepts in the context of an illustrative challenge or problem. This interpretation is not intended to imply that others have recognized and / or addressed the challenge or problem in the manner specified herein. Furthermore, this interpretation is not intended to imply that the subject matter recited in the claims is limited to solving the identified challenge or problem; that is, the subject matter of the claims may be applicable in the context of challenges or problems other than those described herein.
[0144] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for providing content, comprising: receiving input information expressing a physiological state of a user and / or an emotional state experienced by the user, the physiological state of the user being obtained from a state sensing system; generating prompt information describing the input information and the purpose of the guidance to be delivered; mapping the prompt information to output information using a mode filling component, the output information including control instructions for controlling an output system to deliver the guidance via the generated content; as well as The output information is provided to the output system.
2. A method according to claim 1, wherein the physiological state of the user expresses: (a) vital signs; or (b) skin electrode activity; or (c) body movement; or (d) eye-related characteristics; or (e) voice-related characteristics; or (f) any combination thereof. The method of claim 1 , wherein the emotional state is self-reported by the user. The method according to claim 1 , wherein the purpose expressed in the prompt information is the guided treatment goal. The method according to claim 1 , wherein the prompt information further describes the selected environment.
6. The method according to claim 1, The prompt information is a series of input text tokens. wherein the pattern filling component is a machine trained model, and The output information is a series of output text tokens.
7. The method of claim 6, wherein the machine-trained model is a transformer-based machine-trained neural network.
8. A method according to claim 6, wherein at least some of the input text tokens describe the type of output modality to be used, and at least some of the input text tokens describe the format of control information to be provided in the output text tokens.
9. The method of claim 6, wherein the input information and / or the output information is further processed by another machine-trained model, which is trained by reinforcement learning to facilitate the purpose of the guidance.
10. The method of claim 1, wherein the output system comprises a machine-trained model for mapping the output information to visual content.
11. The method of claim 1 , wherein the output system comprises: (a) an audio output system for delivering audio content; or (b) a visual output system for delivering visual content; or (c) a lighting system for controlling lighting; or (d) a scent output system for delivering scents; or (e) a tactile output system for delivering a tactile experience; or (f) an HVAC system for controlling heating, cooling and / or ventilation, or (g) a workflow modification system for controlling the workflow of the user; or (h) any combination thereof.
12. The method of claim 1, further comprising updating the hint information to include aspects of the output information and updated state information to provide updated hint information, and mapping the updated hint information to updated output information using the schema fill component.
13. A processing system comprising a processor and a storage device, wherein the storage device is configured to execute the method according to any one of claims 1 to 12.
14. A computer-readable storage medium for storing computer-readable instructions, which, when executed by a processing system, perform the method according to any one of claims 1 to 12.
15. A computing system for providing content; a storage device for storing computer-readable instructions; a processing system for executing the computer-readable instructions to perform operations comprising: receiving input information expressing a physiological state of a user and / or an emotional state experienced by the user, the physiological state of the user being derived from a state sensing system; mapping the input information to output information using a machine-trained model, the output information comprising control instructions for controlling an output system to deliver the content via the generated content; and providing the output information to the output system for delivery guidance, The machine training model has been trained through reinforcement learning to generate instances of output information that further the guided identified treatment goal.