Interactive educational method and interactive companion
An AI-powered interactive companion provides personalized learning experiences and engages parents, addressing the issue of excessive screen time by promoting interactive exploration and awareness of children's activities.
Patent Information
- Application Number
- PCT/US2025/031509
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-04
AI Technical Summary
Young children spend excessive time on screens, which lacks interactive learning and discourages environmental exploration and peer interaction, while parents are often unaware of their child's activities and learning.
An AI-powered interactive companion that captures real-world images and audio, processes them to provide personalized learning experiences, and engages parents through a companion app, encouraging exploration and interaction.
Enhances learning through interactive real-world experiences, promotes environmental discovery, and informs parents about their child's activities and learning progress.
Smart Images

Figure US2025031509_04122025_PF_FP_ABST
Abstract
Description
PATENT Attorney Docket No: WELI.P2001WO / 00640412 INTERACTIVE EDUCATIONAL METHOD AND INTERACTIVE COMPANION CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 653,127, filed on 29 May 2024, the disclosure of which is incorporated herein by reference in its entirety. SUMMARY
[0002] One aspect of the present embodiments includes the realization that young children are spending an increasing amount of time looking at computer screens. Much of the screen content being viewed does not further the development of the child, and watching a screen is an individual non-interactive activity. Smartphones and tablet devices are ubiquitous and discourage environmental exploration and interaction with peers. The present embodiments solve this problem by providing an interactive companion that understands the real-world environment of the child and delivers age- appropriate learning in a way that is interesting and exciting. Advantageously, the interactive companion encourages the child to explore and learn about their environment, and to interact with other children to share their learning.
[0003] Another aspect of the present embodiments includes the realization that parents / guardians are not always fully aware of their child’s activities and learning. The present embodiments solve this problem by providing a parental engagement mechanism (e.g., via a portal and / or an app running on a smartphone or tablet device) that notifies parents / guardians about their child's discoveries using the interactive companion. The parental engagement mechanism also provides the parents / guardians with suggestions for further discussions with their child to promote co-learning experiences.
[0004] In certain embodiments, the techniques described herein relate to an interactive educational method implemented through an interactive companion, including: receiving, within a processor, an image captured by a camera of the interactive companion of an object in a real-world environment; processing the image using AI to identify the object as a discovery; generating a learning experience for a user of the interactive companion using AI based on the discovery; sending the learning experience to the interactive companion for output to the user; and storing thePATENT Attorney Docket No: WELI.P2001WO / 00640412 discovery and the learning experience in a timeline of an account associated with the user.
[0005] In certain embodiments, the techniques described herein relate to an interactive companion for a child, including: a selectable character; a wireless interface for communicating with an AI-powered backend; a camera for capturing an image from a real-world environment of the child; a microphone for capturing verbal input from the child; and a speaker for outputting a voice of the selectable character; wherein the AI- powered backend generates the voice of the selectable character to interact with the child, and to encourage exploration, discovery, and learning by the child.
[0006] In certain embodiments, the techniques described herein relate to a method for personalized educational content generation for a child, including: receiving, from a child-operable device, an image of a real-world environment; determining, based on the image, a discovery by the child; and generating, based on the discovery, a personalized learning experience that includes at least one educational component selected from a group consisting of: STEM, language, and creativity.
[0007] In certain embodiments, the techniques described herein relate to an AI- powered device for interacting with a child in a real-world environment, including: a housing having a circular part and a handle; a battery; a camera; a microphone; a speaker; a ; a communication interface for communicating with a backend; a button for triggering the camera and the microphone; a processor communicatively coupled with the camera, the microphone, the speaker, the communication interface and the button; and a memory communicatively coupled with the processor and storing machine- readable instructions that, when executed by the processor, cause the processor to: capture an image of the real-world environment using the camera in response to detecting a short-press of the button; capture audio data using the microphone in response to a long-press of the button; sends the image and the audio data to a backend via the communication interface; receive an audio response from the backend in response to the image or the audio data; and playing the audio response via the speaker. BRIEF DESCRIPTION OF THE FIGURES
[0008] FIG.1 is a schematic diagram illustrating one example AI-powered interactive real-world learning system, in embodiments.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0009] FIG.2 is a schematic diagram illustrating a backend service, an interactive companion, a computing device of learning system of FIG.1 in further example detail, in embodiments.
[0010] FIG.3 is a block diagram illustrating the backend service of FIG.1 in further example detail, in embodiments.
[0011] FIG.4A is a block diagram illustrating the interactive companion of FIG.1 in further example detail, in embodiments.
[0012] FIG.4B is a rear perspective of a main body of the interactive companion of FIGs.1 and 4A, in embodiments.
[0013] FIG.4C shows optional accessories for the main body of FIG.4B, in embodiments.
[0014] FIG.5 is a flowchart illustrating one example method for discovering with the interactive companion of FIG.1, in embodiments.
[0015] FIG.6 is a flowchart illustrating one example method for learning with the interactive companion of FIG.1, in embodiments.
[0016] FIG.7 is a flowchart illustrating one example method for providing interactive discovering to a user via the interactive companion of FIG.1, in embodiments.
[0017] FIG.8 is a flowchart illustrating one example method for providing interactive learning to a user via the interactive companion of FIG.1, in embodiments.
[0018] FIGs.9A–9D show example screenshots of the interactive companion of FIG.1, in embodiments.
[0019] FIGs.9E–9G show embodiments of the computing device of FIG.2 displaying example app views of the supervisor app running on the computing device.
[0020] FIG.10 shows an embodiment of the computing device of FIG.2 displaying an example view feed of the supervisor app running on the computing device.
[0021] FIG.11 shows an embodiment of the computing device of FIG.2 displaying an example gallery view of the supervisor app running on the computing device.
[0022] FIG.12 is a table that illustrates developmental characteristics for targeted age groups of the AI-powered interactive real-world learning system of FIG.1, in embodiments.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0023] FIG.13 is a high-level system diagram showing data flow within AI- powered interactive real-world learning system of FIG.1, in embodiments.
[0024] FIG.14 is a flowchart illustrating an embodiment method for generating a language-learning adventure, which may be implemented by the learning system of FIG. 1.
[0025] FIG.15 illustrates an example interactive story presented by an embodiment of the learning system of FIG.1.
[0026] FIG.16 is a flowchart illustrating an example interactive storytelling method , which may be implemented by embodiments of learning system of FIG.1.
[0027] FIG.17 is a flowchart illustrating an example method for generating and utilizing collectible digital flashcards, which may be implemented by the learning system of FIG.1.
[0028] FIG.18 is a flowchart illustrating an example method for generating an interactive narrative based on a real-world scene captured by the learning system of FIG.1.
[0029] FIG.19 shows app views related to example language goals functionality of embodiments of a supervisor app running on the computing device of FIG.2.
[0030] FIG.20 includes an app view illustrating an example pronunciation scoring feature of embodiments of the supervisor app running on the computing device of FIG.2.
[0031] FIG.21 is a flowchart illustrating a method for interactive discovering with the interactive companion of FIG.1; the method combines the methods of FIGs.5 and 7. DETAILED DESCRIPTION OF THE EMBODIMENTS Terminology
[0032] interactive companion: An interactive device with at least one of a camera, speaker, and LEDs that act as a companion to a user (e.g., a child).
[0033] supervisor app: An app running on a smartphone or tablet device that allows a parent / guardian to define settings for the interactive companion and displays an exploration / learning timeline provided by the interactive companion and provides suggestions for further discussions between the parent / guardian and a user of the interactive companion to promote co-learning experiences.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0034] discovery: An object identified in an image captured by the interactive companion that becomes a focus of an interactive learning experience provided to a user of the interactive companion.
[0035] child: a user of an interactive companion, regardless of age.
[0036] character: One of a plurality of characters from children’s TV and learning experiences that is selected to provide a voice for the companion provided by the interactive companion.
[0037] happiness: An indication by the interactive companion to convey successful interaction by the user. Overview
[0038] An artificial intelligence (AI) powered large language model (LLM) is used to analyze and respond to input from Dex to provide an AI-Powered Interactive Real- World Learning System for Children.
[0039] FIG.1 is a schematic diagram illustrating one example learning system 100. Learning system 100 is an example of an AI-powered interactive real-world learning system. Learning system 100 includes an interactive companion 102 that acts as a companion to a user 104 in a real-world environment 106, in embodiments. Interactive companion 102 is a small and light-weight device that is easily carried and operated by a child (e.g., a child-operable device) and acts as a companion to the child. The user is a child in certain use scenarios of learning system 100.
[0040] In the embodiment of FIG.1, interactive companion 102 is like a magnifying glass (e.g., a housing having a circular part and a handle) but may take other shapes and forms without departing from the scope hereof. For example, interactive companion 102 may take the form of an instamatic / snapshot camera, a microscope, a telescope, a pendant, and so on. Interactive companion 102 and user 104 communicate verbally 108 and interactive companion 102 includes a happiness indicator 110 that displays a happiness level of the companion. For example, interactive companion 102 receives a happiness level (e.g., see happiness level 238 of FIG.2) to simulate a satisfaction of the companion with the attention received from user 104.
[0041] Interactive companion 102 includes a camera that is triggered by a camera button 111 positioned on a handle of interactive companion 102. In certain embodiments, interactive companion 102 includes a second camera button 112PATENT Attorney Docket No: WELI.P2001WO / 00640412 positioned on a handle of interactive companion 102 for example. In certain embodiments, a central portion 114 of interactive companion 102 is a magnifying glass. In other embodiments, central portion 114 is a display screen used by interactive companion 102 to display an image of the characters and / or discoveries made by user 104 and allows user 104 to interact with interactive companion 102. In certain embodiments, central portion 114 implements both a magnifying glass and a display screen, where the display screen appears substantially transparent until activated to display information. In certain embodiments, central portion 114 is a display screen that shows a live view from the camera. When central portion is a display screen, the display screen may be a touch screen.
[0042] Interactive companion 102 may also include a talk button 113 (e.g., push- to-talk button as used in radio communication) that allows user 104 to verbally interact with interactive companion 102. For example, user 104 may press and hold talk button 113 when speaking to interactive companion 102, and may give talk button 113 a short press to request additional information from interactive companion 102. For example, a short-press on talk button 113 to request more information and / or more story. For example, after listening to an audio response 236 (FIG.2), user 104 may short-press talk button 113 to listen to a next portion of a story. Alternatively, user 104 may long-press talk button 113 to ask a question.
[0043] Learning system 100 includes a backend service 120 that may communicate wirelessly with interactive companion 102. As shown, backend service 120 is implemented as a cloud-based service (e.g., using Amazon Web Services – AWS, or other similar service platforms). As user 104 encounters objects (e.g., a dog 132, a tree 134, etc.) within real-world environment 106, user 104 may press camera button 111 to capture an image 136 of the object(s). Image 136 is sent from interactive companion 102 to backend service 120 where it is processed and initiates a conversation between interactive companion 102 and user 104.
[0044] FIG.2 is a schematic diagram illustrating learning system 100 of FIG.1 in further example detail, in embodiments. FIGS.1 and 2 are best viewed together with the following description. Backend service 120 includes a coordination server 202, a speech-to-text service 204, a text-to-speech service 206, an AI service 208, a vector database 210, and a relational database 212. Backend service 120 may include other components without departing from the scope hereof.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0045] In certain embodiments, coordination server 202 is implemented using python script. Coordination server 202 coordinates functionality of learning system 100 to receive, comprehend, and respond to communications from user 104 via interactive companion 102. For example, coordination server 202 generates an AI prompt 203 based on the communication from user 104 and sends AI prompt 203 to AI service 208. In response to AI prompt 203, AI service 208 generates an AI response text 234 that includes a learning experience, such as for a discovery in the communication from the user 104. Speech-to-text service 204 (e.g., Google, OpenAI, etc.) converts digital audio into text, and text-to-speech service 206 (e.g., Google, Gemelo, ElevenLabs, etc.) converts response text 234 into an audio response (e.g., digital audio data) for output by interactive companion 102. Particularly, text-to-speech service 206 generates audio response 236 to have the characteristics of a selected character (see character 308 of FIG.3). Coordination server 202 may also generate a happiness level 238 that is sent to interactive companion 102 for display on happiness indicator 110. Happiness level 238 is maintained by coordination server 202 and may have a plurality of levels (e.g., eight), that are represented by happiness indicator 110 on interactive companion 102. In certain embodiments, a number of happiness levels corresponds to a number of LEDs of happiness indicator 110. For each interaction by user 104 with interactive companion 102, coordination server 202 increases happiness level 238, but happiness has a limited lifetime (e.g., twenty-four hours), after which the happiness is lost. Accordingly, coordination server 202 increases happiness level 238 as user 104 interacts with interactive companion 102, and decreases happiness level 238 over time as each happiness increase ages out. Thus, happiness level 238 decreases when user 104 does not interact with interactive companion 102, and happiness level 238 increases when user 104 interacts frequently with interactive companion 102. Text-to-speech service 206 may also be controlled to generate audio response 236 to have inflection based on happiness level 238.
[0046] AI service 208 is a large language model (LLM) AI service (e.g., GPT-4V, GPT-4, etc.) that is invoked to process AI prompt 203 that includes text derived from communication by user 104 with interactive companion 102 and generate corresponding AI response text 234. AI service 208 may also process image 136, captured by the camera of interactive companion 102, to identify at least one discovery (e.g., an object) in the image, whereby AI service 208 generates AI response text 234PATENT Attorney Docket No: WELI.P2001WO / 00640412 with attributes defining the object as derived from the image. Coordination server 202 uses vector database 210 to store model selection parameters and operational directives of AI service 208 and uses relational database 212 to store account data and settings of user 104 (e.g., of interactive companion 102).
[0047] In one example of operation, when user 104 speaks to interactive companion 102, interactive companion 102 captures audio data 230 (e.g., a digital audio format) and wirelessly sends audio data 230 to backend service 120. Coordination server 202 receives audio data 230 and invokes speech-to-text service 204 to convert audio data 230 into input text 232. Coordination server 202 then generates AI prompt 203 from input text 232 and sends AI prompt 203 to AI service 208, where AI prompt 203 includes corresponding directives that instruct AI service 208 to process image 136, if included, and to generate a suitable response text 234 for user 104.
[0048] Where interactive companion 102 sends only image 136 to coordination server 202, coordination server 202 sends image 136 to AI service 208 with directives to identify any new, and / or previously found, discoveries in the image. AI service 208 generates response text 234 defining objects found in image 136. Coordination server 202 may process response text 234, which may include attributes of objects found in image 136, and invokes text-to-speech service 206 to convert response text 234 into an audio response 236 (e.g., digital audio data).
[0049] As shown in FIG.2, an adult 240 may use a computing device 242 to run a supervisor app 244 that allows adult 240 to configure parameters of an account for user 104 and / or interactive companion 102 and thereby control operation of learning system 100. Supervisor app 244 may also receive insights 246 from coordination server 202, such as suggested follow-on interactions between adult 240 and user 104 based on discoveries made by user 104 using interactive companion 102. Insights 246 may include information instructing adult 240 on how to participate further with user 104. For example, insights 246 might include "Your child saw an elephant for the first time on March second of this year! Here are three topics for further discussion at dinner... Here's a book you can read to your child... Here's a video you can watch with your child... Your child has shown a strong interest in animals (ten animal discoveries in the past thirty days), consider these ways to further stimulate their potential..." Advantageously, insights 246 provide additional details to adult 240 for understanding progress of user 104 to allow adult 240 to interact and further guide user 104 based onPATENT Attorney Docket No: WELI.P2001WO / 00640412 their discoveries and learning. Further, via supervisor app 244, adult 240 may provide feedback on any event, wherein the AI model automatically adjusts based on the parent's feedback.
[0050] FIG.3 is a block diagram illustrating backend service 120 of FIG.1 in further example detail, in embodiments. Coordination server 202 stores, in relational database 212 and / or vector database 210, account 302 for interactive companion 102 and / or user 104. For example, coordination server 202 interacts with supervisor app 244 to allow adult 240 to configure operation of interactive companion 102 for user 104. Account 302 includes an age / date-of-birth 304, learning goals 306, character 308, knowledge density 310, custom preferences 312, a timeline 314, and discoveries 316.
[0051] Age / date-of-birth 304 defines the birthday or age of user 104 and allows coordination server 202 to use appropriate directives for interacting with AI service 208. For example, date-of-birth 304 is used by coordination server 202 to select an age group directive 322, from age-appropriate discover directives 320, that is sent to AI service 208 such that discoveries and conversations are appropriate for a child of that age. Age-appropriate discover directives 320 includes a plurality of age group directives 322, such as an age group directive 322(1) for children between the ages of four and six, and an age group directive 322(2) for children between the ages of seven and ten, and so on. In one example of operation, coordination server 202 uses age group directive 322 to customize a communication style for interactive companion 102 based on age / date-of-birth 304 such that user 104 receives communications corresponding to an age-appropriate developmental stage of user 104. Particularly, age group directive 322(1) may cause AI service 208 to provide younger children with basic concepts and hands-on experiments, and age group directive 322(2) may cause AI service 208 to provide older children with more complex character traits and story themes.
[0052] Backend service 120 may also include age-appropriate learn directives 330 that instruct AI service 208 to provide educational content for existing discoveries 317. Age-appropriate learn directives 330 includes a plurality of age group directives 332, such as an age group directive 332(1) for children between the ages of four and six, an age group directive 332(2) for children between the ages of seven and ten, and so on. In one example of operation, coordination server 202 uses age group directive 332 to customize a communication style for interactive companion 102 based on age / date-of- birth 304 such that user 104 receives communications corresponding to an age-PATENT Attorney Docket No: WELI.P2001WO / 00640412 appropriate developmental stage of user 104. Particularly, age group directive 332(1) may cause AI service 208 to provide younger children with basic concepts and hands-on experiments, and age group directive 332(2) may cause AI service 208 to provide older children with more complex character traits and story themes.
[0053] Backend service 120 may also include age-appropriate challenges 360 that instruct AI service 208 to generate a challenge 248 for user 104. Age-appropriate challenges 360 includes a plurality of age group directives 362, such as an age group directive 362(1) for children between the ages of four and six, an age group directive 362(2) for children between the ages of seven and ten, and so on. In one example of operation, coordination server 202 uses age group directive 362 to generate an age- appropriate challenge, based on the age of user 104, that is designed to motivate user 104 to discover certain objects. Accordingly, where user 104 has not been interacting with interactive companion 102 and based on age-appropriate challenges 360 (and optionally other current information received from interactive companion 102, such as a current location, time of day, etc.), coordination server 202 may generate and send challenge 248 to interactive companion 102 to motivate user 104 to make discoveries to fulfill the challenge. In one example, challenge 248 may state: “Take a picture of as many square objects that are nearby.” Coordination server 202 then evaluates each image 136 received from interactive companion 102 to determine whether it includes an object related to challenge 248.
[0054] Learning goals 306 defines one or more of general knowledge, language communication, STEM, foreign language introduction / learning, social skills, creativity, and so on. Coordination server 202 selects and / or tailors AI directives and control logic within coordination server 202 based on the selected learning goals 306. For example, when adult 240 selects a foreign language introduction to Spanish, coordination server 202 selects or generates directives to increase likelihood that AI service 208 incorporates new Spanish words into picture explanations output to interactive companion 102 during the learning mode.
[0055] Character 308 defines a personality of interactive companion 102. For example, adult 240 may select one of the available (e.g., purchased) characters for interactive companion 102. Adult 240 may also deselect the currently selected character.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0056] Knowledge density 310 is selected by adult 240 from automatic, high, medium, and low to define a level of information and educational level in feedback generated by AI service 208 and sent to interactive companion 102. When automatic is selected, coordination server 202 determines an appropriate level of feedback based on responses and reactions made by user 104 to previous content.
[0057] Custom preferences 312 may be defined by adult 240 to further provide ideas and directions to AI service 208 when generating responses for user 104 via interactive companion 102. For example, adult 240 may describe preferences, ideas, and requirements for interaction of interactive companion 102 with user 104 using natural language. Coordination server 202 provides custom preferences 312 as directives to AI service 208.
[0058] Timeline 314 is generated automatically as user 104 uses interactive companion 102 and includes activities and communication performed by interactive companion 102. In certain embodiments, timeline 314 includes all communications between interactive companion 102 and user 104 for both exploration and learning modes of operation. Adult 240 may view timeline 314 via supervisor app 244 and is thereby fully apprised of the use of interactive companion 102 by user 104.
[0059] Discoveries 316 stores a discovery 317 for each identified object of interest found by user 104 in a discovery mode and within educational content provided via interactive companion 102 during a learning mode. Each discovery 317 may include images (e.g., image 136) captured of discoveries and thereby function as a photo album that may be viewed by user 104 and adult 240. Each discovery 317 is associated with at least one event 315 of timeline 314 and supervisor app 244 may present photos corresponding to each event 315 and / or each discovery 317 as memories. Each event 315 includes a communication to or from user 104 via interactive companion 102, such that timeline 314 is a transcription of communications between user 104 and interactive companion 102.
[0060] Advantageously, through supervisor app 244, adult 240 may view, participate in, and customize exploration and learning of user 104, providing feedback through supervisor app 244 that helps AI service 208 to better understand the needs of user 104.
[0061] Supervisor app 244 may allow computing device 242 to connect with interactive companion 102. In certain embodiments, interactive companion 102 mayPATENT Attorney Docket No: WELI.P2001WO / 00640412 communicate with backend service 120 via computing device 242, where other communication paths are not available to interactive companion 102.
[0062] Backend service 120 may also include an age specific black list 340 that defines words and phrases that may not be used for each age range, and / or may include an age specific white list 350 that defines words and phrases that may be used for each age range.
[0063] Although FIG.2 shows a single interactive companion 102 and user 104 and a single adult 240, supervisor app 244 may be associated with multiple interactive companions 102 (e.g., controlling multiple accounts 302), and interactive companion 102 (and its account 302) may be associated with multiple supervisor apps 244 and adult 240. Supervisor app 244 may implement security and privacy protection by requiring authentication prior to access to information of account 302. Interactive Companion
[0064] FIG.4A is a block diagram illustrating interactive companion 102 of FIG.1 in further example detail, in embodiments. FIG.4B is a rear perspective of a main body of interactive companion 102 of FIGs.1 and 4A separated from a vertical handle, in embodiments. FIG.4C shows a plurality of accessories for use with the main body of interactive companion 102 of FIG.4B, in embodiments. FIGs.4A, 4B, and 4C are best viewed together with the following description.
[0065] Interactive companion 102 includes happiness indicator 110, camera button 111, talk button 113, a camera 402, a communication interface 404, a microphone 406, a speaker 408, a processor 410, a memory 412 storing firmware 414, and a battery 432. Interactive companion 102 may also include a display 430.
[0066] Interactive companion 102 may include a backend service 401, which is an example of backend service 120. Hence, interactive companion 102 may execute any number of the functions and store data described herein as being executed or stored by backend service 120. Part of backend service 401 may be stored in memory 412 and executed by processor 410.
[0067] Camera 402 is a back-facing camera communicatively coupled with processor 410 and / or memory 412. In certain embodiments, 402 is a 5-MP back-facing camera. In other embodiments, camera 402 is a back-facing camera with a resolution of 512×512 pixels.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0068] Communication interface 404 is communicatively coupled with processor 410 and implements one or more wireless communication protocols (e.g., Bluetooth 3.0, Wi-Fi, etc.). In certain embodiments, communication interface 404 also implements one or more of near-field communication capability, UWB communication capability, and cellular communication capability. Microphone 406 is built-in to interactive companion 102 and configured to capture audio (e.g., verbal communication) from user 104. Interactive companion 102 may have multiple microphones positioned around the main body without departing from the scope hereof. Speaker 408 is built-in to interactive companion 102 and is capable of outputting audio to user 104. Interactive companion 102 may include an inertial measurement unit (IMU) 416 the detects orientation and / or movement of interactive companion 102. Interactive companion 102 may also include a real-time clock (RTC) 418 that provides interactive companion 102 with a current date and time of day.
[0069] Processor 410 is a digital processor that is communicatively coupled to memory 412 which may include non-volatile memory and volatile memory. Firmware 414 includes machine-readable instructions that are executable by processor 410 to implement functionality of interactive companion 102 as described herein. Battery 432 may be rechargeable and powers components of interactive companion 102.
[0070] Display 430 may have a relatively low resolution as compared to smart phones, etc. That is, display 430 need not be a high-resolution display as found on smart phones, etc. In certain embodiments, display 430 is a pixel screen capable of displaying pixel art. In embodiments where interactive companion 102 includes display 430, processor 410 generates a character animation (e.g., lips moving), based on character 308, as audio response 236 is played through speaker 408. Display 430 is low resolution and operates to assist the user in operation of interactive companion 102. Particularly, display 430 may not operate unless the user is actively interacting with interactive companion 102. Further, in embodiments display 430 operates only for short durations during the interaction with the user.
[0071] Happiness indicator 110 is controlled by processor 410 to indicate a happiness of the character represented by (selected for) interactive companion 102. In certain embodiments, happiness indicator 110 is an array of multi-color LEDs mounted around an edge of interactive companion 102. Where happiness level 238 is above a value displayable by happiness indicator 110, processor 410 causes happiness indicatorPATENT Attorney Docket No: WELI.P2001WO / 00640412 110 to display a rainbow animation (e.g., where the color flow circulates on the multi- color LEDs). In embodiments that include display 430, happiness indicator 110 may be displayed as graphics on display 430, whereby interactive companion 102 does not include LEDs for happiness indicator 110. In certain embodiments, interactive companion 102 and / or backend service 120 maintains a streak length 434 that is a count of consecutive days when happiness level 238 has been above a threshold level (e.g., greater than zero). Streak length 434 may be displayed on display 430 to provide an incentive for user 104 to learn using interactive companion 102.
[0072] Camera button 111 is electrically coupled with processor 410 and mounted on a front side of interactive companion 102 such that it is easily operable by user 104 when holding interactive companion 102 in one hand. Interactive companion 102 may include additional buttons and / or switches without departing from the scope hereof. Interactive companion 102 also includes volume buttons 420 that allow the user to adjust an output volume from speaker 408.
[0073] Advantageously, the main body of interactive companion 102 includes a mechanical connector that allows the main body to couple with any one of the plurality of accessories. FIG.4C shows a vertical handle 440, a table stand 450, a laptop mount 460, a portable kit 470, a horizontal handle 480, and a carrier bag 490. In certain embodiments, as shown in FIGs.1 and 2, vertical handle 440 may include an additional camera button (see camera button 112 of FIGs.1 and 2) that facilitates single-hand use of interactive companion 102.
[0074] Interactive companion 102 is implemented using the following components in various embodiments. Processor 410 is a Rockchip PX30-S Quad-core ARM Cortex-A35 CPU, memory 412 includes a DDR42G for RAM and an EMMC 32G FLASH memory. Communication interface 404 implements Wi-Fi 2.4 / 5G and BT 5.0. Camera 402 has 8MP with an image stabilizer, fast focus and is implemented on the circuit board using built-in chips. Interactive companion 102 may also include a zoom ring 422 around the outside of interactive companion 102 that allows the user to zoom a view of camera 402 in and out by turning zoom ring 422 as indicated by arrow 423, similar to a zoom of a DSLR digital camera. IMU 416 may detect 6-axes of movement and may determine an orientation of the main body of interactive companion 102. Display 430 may be an OLED 480×480 pixel touch screen with a diameter of 2.1 inches. Microphone 406 may have a sensitivity of 32dB. Happiness indicator 110 may bePATENT Attorney Docket No: WELI.P2001WO / 00640412 implemented as sixteen RGB light-emitting diodes that are controlled by processor 410. Battery 432 is a rechargeable lithium-ion battery with a capacity of 1800 mAH and is charged by providing 5V to a USB-C connector on the main body of interactive companion 102. RTC 418 is an HYM8563.
[0075] In certain embodiments, communication interface 404 also includes a 4G modem (e.g., a Quectel SC60) and Esim (e.g., Soracom eSIM) that facilitates communication over cellular networks. Adult 240 may then purchase a day or month pass to active mobile data for interactive companion 102.
[0076] FIG.5 is a flowchart illustrating one example method 500 for discovering with interactive companion 102 of FIG.1, in embodiments. Method 500 is implemented within firmware 414 of FIG.4A, for example.
[0077] In block 502, method 500 detects a short-press of a shutter button. In one example of block 502, processor 410 detects a short-press of camera button 111 by the user of interactive companion 102. In block 504, method 500 captures an image using a camera. In one example of block 504, processor 410 controls camera 402 to capture image 136.
[0078] Block 506 is a decision. If, in block 506, method 500 detects a long-press of a talk button, method 500 continues with block 508; otherwise, method 500 continues with block 510. In one example of block 506, processor 410 detects a long- press of talk button 113. In block 508, method 500 captures audio using a microphone. In one example of block 508, processor 410 captures audio data 230 from microphone 406 while talk button 113 is held down.
[0079] In block 510, method 500 sends the image and / or audio data to the backend. In one example of block 510, processor 410 sends image 136 and / or audio data 230 to backend service 120 via communication interface 404. In block 512, method 500 receives audio response and happiness level from the backend. In one example of block 512, processor 410 receives an audio response 236 and a happiness level 238 from backend service 120 via communication interface 404. In block 514, method 500 outputs the audio response using the speaker. In one example of block 514, processor 410 drives speaker 408 using audio response 236. In block 516, method 500 displays happiness. In one example of block 516, processor 410 outputs happiness level 238 on happiness indicator 110.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0080] Block 518 is a decision. If, in block 518, method 500 detects a short-press of the talk button, method 500 continues with block 520; otherwise, method 500 continues with block 522. In one example of block 518, processor 410 detects a short- press of talk button 113. In block 520, method 500 sends a “more” request to the backend. In one example of block 520, processor 410 sends a “more” request to backend service 120 via communication interface 404. Method 500 then continues with block 512.
[0081] Block 522 is a decision. If, in block 522, method 500 detects a long-press of the talk button, method 500 continues with block 524; otherwise, method 500 continues with block 502. In one example of block 522, processor 410 detects a long- press of talk button 113. In block 524, method 500 captures audio using the microphone. In one example of block 524, processor 410 captures audio data 230 from microphone 406 while talk button 113 is held down. In block 526, method 500 sends the audio data to the backend. In one example of block 526, processor 410 sends audio data 230 to backend service 120 via communication interface 404. Method 500 then continues with block 512.
[0082] Blocks 502 through 526 repeat for each push of camera button 111 as the user makes discoveries with interactive companion 102.
[0083] FIG.6 is a flowchart illustrating one example method 600 for learning with interactive companion 102 of FIG.1, in embodiments. Method 600 is implemented within firmware 414 of FIG.4A for example.
[0084] Block 602 is a decision. If, in block 602, method 600 detects a short-press of a talk button, method 600 continues with block 604; otherwise, method 600 continues with block 606. In one example of block 602, processor 410 detects a short- press of talk button 113. In block 604, method 600 sends a more request to the backend. In one example of block 604, processor 410 sends a request for more output to backend service 120 via communication interface 404. Method 600 then continues with block 612.
[0085] Block 606 is a decision. If, in block 606, method 600 detects a long-press of the talk button, method 600 continues with block 608; otherwise, method 600 continues with block 602. In one example of block 606, processor 410 detects a long push on talk button 113. In block 608, method 600 captures audio using a microphone. In one example of block 608, processor 410 captures audio data 230 from microphonePATENT Attorney Docket No: WELI.P2001WO / 00640412 406 while talk button 113 is held down. In block 610, method 600 sends audio data to the backend. In one example of block 610, processor 410 sends audio data 230 to backend service 120 via communication interface 404. Method then continue with block 612.
[0086] In block 612, method 600 receives an audio response and a happiness level from the backend. In one example of block 612, processor 410 receives an audio response 236 and a happiness level 238 from backend service 120 via communication interface 404. In block 614, method 600 outputs the audio response using the speaker. In one example of block 614, processor 410 drives speaker 408 using audio response 236. In block 616, method 600 displays happiness. In one example of block 616, processor 410 outputs happiness level 238 on happiness indicator 110.
[0087] Blocks 602 through 616 repeat for each push of talk button 113. Backend Operation
[0088] FIG.7 is a flowchart illustrating one example method 700 for providing interactive discovering to user 104 via interactive companion 102 of FIG.1, in embodiments. Method 700 may be implemented by coordination server 202 of FIG.2, by interactive companion 102, or a combination thereof
[0089] In block 702, method 700 receives an image and / or audio data from an interactive companion. In one example of block 702, coordination server 202 receives image 136 and / or audio data 230 from interactive companion 102. In block 704, method 700 transcribes the audio data to text. In one example of block 704, coordination server 202 invokes speech-to-text service 204 to convert audio data 230 into input text 232. In block 706, method 700 sends the image and the text with corresponding age level directives to the AI service. In one example of block 706, coordination server 202 sends image 136, input text 232, age group directive 322(1), and custom preferences 312 to AI service 208.
[0090] In block 708, method 700 receives response text from the AI service. In one example of block 708, coordination server 202 receives response text 234 from AI service 208. In block 710, method 700 generates an audio response from the response text based on the selected character and a current tone. In one example of block 710, coordination server 202 invokes text-to-speech service 206 to convert response text 234 to audio response 236 based on character 308 and happiness level 238, wherePATENT Attorney Docket No: WELI.P2001WO / 00640412 happiness level 238 is reflected in the voice response emotion, mimicking humans reacting to companionship. For example, happiness level 238 is factored in the prompt as parameters to influence the tone. Overall the tone will be on a spectrum of calm / neutral to excited / enchanted. In block 712, method 700 sends the audio response to the interactive companion. In one example of block 712, coordination server 202 sends audio response 236 to interactive companion 102.
[0091] Block 714 is a decision. If, in block 714, method 700 determines that a new discovery has been made, method 700 continues with block 716; otherwise, method 700 continues with block 720. For example, coordination server 202 processes response text 234 to identify a new discovery. Coordination server 202 maintains discoveries 316 (e.g., a discovery database within relational database 212) of discoveries 317 made by user 104 using interactive companion 102 (e.g., within image 136 and / or audio data 230). Coordination server 202 detects objects / concepts of educational interest in each response text 234 received from AI service 208, and determined whether the objects / concepts is already stored within discoveries 316. When the objects / concepts are new, coordination server 202 add the objects / concepts to discoveries 316 to await further processing and / or actions, such as awarding user 104 on achievements and / or recommending adult 240 takes actions based on a determined discovery pattern for user 104. In block 716, method 700 records the new discovery and increases the happiness level. In one example of block 716, coordination server 202 keeps a local record of the new discovery and increases happiness level 238 by two. In block 718, method 700 sends audio for the new discovery and happiness level to the interactive companion. In one example of block 718, coordination server 202 generates response text 234 to include praise for finding a new discovery, invokes text-to-speech service 206 to convert response text 234 into audio response 236, and sends audio response 236 and happiness level 238 to interactive companion 102. Advantageously, this additional audio response 236 and increased happiness level 238 provide a positive feedback to user 104 for the new discovery, thereby encouraging further discovery.
[0092] Block 720 is a decision. If, in block 720, method 700 determines that user 104 is providing further interaction (e.g., before a response period ends), method 700 continues with block 722; otherwise, method 700 continues with block 724.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0093] In block 722, method 700 receives an image and / or audio data from the interactive companion. In one example of block 722, coordination server 202 receives further audio data 230 from interactive companion 102. Method 700 then continues with block 704. Blocks 704 through 722 repeat as user 104 continues to interact with interactive companion 102.
[0094] In block 724, method 700 stored new discoveries in the timeline. In one example of block 724, coordination server 202 stores discovery 317(2) in discoveries 316 and adds a corresponding event 315(2) to timeline 314. In block 726, method 700 sends a thank-you message and happiness level to the interactive companion. In one example of block 726, coordination server 202 stores a message thanking the user for taking interactive companion 102 to see new and interesting things in response text 234, invokes text-to-speech service 206 to convert response text 234 into audio response 236, and sends audio response 236 and happiness level 238 to interactive companion 102.
[0095] FIG.8 is a flowchart illustrating one example method 800 for providing interactive learning to user 104 via interactive companion 102 of FIG.1, in embodiments. Method 800 is implemented in coordination server 202 of FIG.2 and invoked in response to coordination server 202 receiving a request to learn more about an existing discovery, for example.
[0096] In block 802, method 800 receives a learn request and / or audio data from an interactive companion. In one example of block 802, coordination server 202 receives an identifier of discovery 317(1) in a learn request (e.g., a message) from interactive companion 102. In another example, of block 802, coordination server 202 receives audio data 230 from 102. In block 804, method 800 transcribes the audio data to text. In one example of block 804, coordination server 202 invokes speech-to-text service 204 to convert audio data 230 into input text 232. In block 806, method 800 sends the text with corresponding age level learn directives to the AI service. In one example of block 806, coordination server 202 sends input text 232, age group directive 332(1), and custom preferences 312 to AI service 208.
[0097] In block 808, method 800 receives response text from the AI service. In one example of block 808, coordination server 202 receives response text 234 from AI service 208. In block 810, method 800 generates an audio response from the response text based on the selected character and a current tone. In one example of block 810,PATENT Attorney Docket No: WELI.P2001WO / 00640412 coordination server 202 invokes text-to-speech service 206 to convert response text 234 to audio response 236 based on character 308 and happiness level 238. Happiness level 238 may be reflected in the voice response emotion, mimicking humans reacting to companionship. For example, happiness level 238 is factored in the prompt as parameters to influence the tone of the audio response 236. Overall, the tone will be on a spectrum of calm / neutral to excited / enchanted. In block 812, method 800 sends the audio response to the interactive companion. In one example of block 812, coordination server 202 sends audio response 236 to interactive companion 102.
[0098] Block 814 is a decision. If, in block 814, method 800 determines that user 104 is providing further interaction (e.g., before a response period ends), method 800 continues with block 816; otherwise, method 800 continues with block 822.
[0099] In block 816, method 800 records new discoveries and increases the happiness level. In one example of block 816, coordination server 202 keeps a local record of the new discoveries and increases happiness level 238 by two. In block 818, method 800 sends audio for story complete and happiness level to the interactive companion. In one example of block 818, coordination server 202 stores a message indicating that the story is complete in response text 234, invokes text-to-speech service 206 to convert response text 234 into audio response 236, and sends audio response 236 and happiness level 238 to interactive companion 102. In block 820, method 800 stores new discoveries in the timeline. In one example of block 820, coordination server 202 stores discovery 317(2) in discoveries 316 and adds a corresponding event 315(2) to timeline 314. Method 800 then continues with block 826.
[0100] Block 822 is a decision. If, in block 822, method 800 determines that the user is continuing to interact, method 800 continues with block 824; otherwise, method 800 continues with block 826.
[0101] In block 824, method 800 receives further audio data from the interactive companion. In one example of block 824, coordination server 202 receives audio data 230 with new speech from interactive companion 102. Method 800 then continues with block 804. Blocks 804 through 822 repeat as user 104 continues to interact with interactive companion 102.
[0102] In block 826, method 800 sends a thank-you message and happiness level to the interactive companion. In one example of block 826, coordination server 202 stores a message thanking the user for completing the interactive story in response textPATENT Attorney Docket No: WELI.P2001WO / 00640412 234, invokes text-to-speech service 206 to convert response text 234 into audio response 236, and sends audio response 236 and happiness level 238 to interactive companion 102. Method 800 then terminates until learning mode is invoked again from interactive companion 102.
[0103] FIGs.9A–9D show example app views 902, 904, 906, and 922, respectively, of interactive companion 102 of FIG.1, in embodiments. App views 902, 904, 906, and 922 illustrate functionality of interactive companion 102 emulated on a device 942, which may be a tablet or a smartphone.
[0104] App view 902 shows simulated camera button 112 and talk button 113 and a character representation 908 of character 308, which is the currently active character for interactive companion 102. Selecting character representation 908 presents, as shown on app view 904, a character list 910 of characters available for use with interactive companion 102 for user 104. User 104 may select any available character from character list 910 for use by interactive companion 102. Character list 910 may also show, for each character, a discovery count 912 that indicate the number of new discoveries made while that character was selected, and a happy day count 914 that indicates the number of days when happiness level 238 was above a threshold value (e.g., greater than zero). App view 902 also shows a new discovery count 916 of recent new discoveries made with currently selected character 308, and a discovery button 918 that, when selected, displays a discovery list 920, as shown on app view 906, of discoveries 317 made with the selected character 308. User 104 may select one discovery 317 from discovery list 920 to initiate further learning about that discovery.
[0105] User 104 may tap on discovery 317 in 920 to display image 136 corresponding to that discovery. To initiate further interaction, user 104 may tap on, or circle, a portion of image 136 on display 430 to direct the attention of backend service 120 to that portion and thereby learn more about objects at that specific part of the image. Each image 136 may contain more than one discovery 317 and by tapping on the image, user 104 may direct backend service 120 to other discoveries within the image.
[0106] App view 922 shows user 104 controlling interactive companion 102 to capture a new discovery that is a begonia. FIG.9D shows an emulated display 930, which is an emulation of display 430. As shown emulated in FIG.9D, emulated display 930 may be circular, whereby interactive companion 102 displays a viewfinder of camera 402 that is circular. In this example, emulated display 930 includes a thumbnailPATENT Attorney Docket No: WELI.P2001WO / 00640412 image 924 of a first image (e.g., an entry point) of a current discovery album. A discovery album is a group of associated discoveries 317, such as where the discoveries are related to one another and / or occurred at a similar time.
[0107] Accordingly, new discoveries may be associated with a “current” discovery album, where thumbnail image 924 represents image 136 that was first captured for the current discovery album. Although example app view 922 shows settings (e.g., MP3 resolution, PCM resolution, AI service 208 selection) for configuring operation of interactive companion 102, these settings are provided only in the emulated embodiment and are not available (e.g., visible) to user 104 on interactive companion 102.
[0108] User 104 may have multiple adults 240, whereby each adult 240 may access and modify account 302 of user 104, and may view the exploration / learning timeline and engage in depth with the teaching of user 104.FIGs.9E–9G show example app views 952, 954, and 956, respectively, of supervisor app 244 running on computing device 242 of FIG.2 in embodiments.
[0109] Supervisor app 244 displays a dashboard, as shown in app view 952 (FIG. 9E), that allows adult 240 to select from a user list 958 of users managed by adult 240. In this example, adult 240 has four children that each use a different interactive companion 102, and the first user is selected, causing supervisor app 244 to display a discovery list 960 of discoveries made by that first user. Advantageously, adult 240 sees discoveries made by the selected user and may select a corresponding co-pilot button 962 to view and control the activity, as described below with reference to FIGs.10 and 11.
[0110] Selecting a setting option 964 for the selected user opens a setting screen, as shown in FIG.9G. Accordingly, app view 956 is also referred to herein as setting screen 956. Through setting screen 956, adult 240 controls the functionality of interactive companion 102 for that user. For example, adult 240 may use setting screen 956 to define a birthday, time allowed, and a default character of interactive companion 102 for that user. Adult 240 may also define which characters may be selected by the user.
[0111] Adult 240 may set a slider (not shown) in app view 956 that defines functionality of interactive companion 102 between usefulness and safety, and may provide example responses for certain stock photos such that adult 240 sees typicalPATENT Attorney Docket No: WELI.P2001WO / 00640412 response levels of interactive companion 102. In certain embodiments, adult 240 may receive an indication of which photographs captured by interactive companion 102 contains a minor’s face and body and adult 240 may be prompted to download such photographs to their own photo album and allow adult 240 to delete those photos from interactive companion 102.
[0112] Adult 240 may also interact with supervisor app 244 to send a challenge 248 (see FIG.2) to interactive companion 102. For example, challenge 248 may be “discover as many fruits as possible today.” In certain embodiments, challenge 248 is displayed on interactive companion 102 in a way similar to a notification on a smartphone.
[0113] Advantageously, supervisor app 244 provides adult 240 with further control over operation of interactive companion 102 for each user in user list 958. Selecting co-pilot button 962 of FIG.9F allows adult 240 to provide feedback to learning system 100 for each discovery 317. Adult 240 may correct an explanation of a photograph provided by learning system 100. For example, where discovery 317 was triggered by a picture of a horse with a somewhat stripy appearance that was identified as a zebra by learning system 100, adult 240 may provide a corrected explanation to learning system 100, which is processed by learning system 100 to provide a corrective statement to user 104, as the selected character 308, such as “Hey! Sorry I made a mistake - your parents found out this is not a zebra – just a horse that happens to have stripes…” Advantageously, the feedback appears in the storyline of discovery 317, as viewed on interactive companion 102, as if provided by character 308. In certain embodiments, the feedback is anonymized and aggregated corrections (e.g., not the actual feedback as received) are processed by learning system 100 to determine what types of objects / concepts receive the most feedback (e.g., corrections) to fine tune learning system 100 to address common reasons for errors.
[0114] Advantageously, adult 240 may encourage the exploration and learning experience of user 104 through use of supervisor app 244. Adult 240 may view and participate in discoveries 317 (e.g., the exploration / learning records) of user 104 through use of supervisor app 244. Adult 240 may provide feedback through use of supervisor app 244 to help the learning model better understand the needs of user 104.
[0115] FIG.10 shows an embodiment of computing device 242 displaying an example view feed 1002 of supervisor app 244. View feed 1002 includes a photographPATENT Attorney Docket No: WELI.P2001WO / 00640412 1004. Photograph 1004 and corresponding interactions 1006 between interactive companion 102 and user 104 for each discovery are presented as a chronological feed (e.g., interaction history). In this example, adult 240 may review corresponding interactions 1006 provided to user 104 in response to photograph 1004. Although not shown in this example, corresponding interactions 1006 may include a transcript of input (e.g., questions etc.) from user 104.
[0116] FIG.11 shows an embodiment of supervisor app 244 displaying an example gallery view 1102 of supervisor app 244. In gallery view 1102, photographs 1104 from discoveries 317 are presented, without the corresponding interactions, in chronological groups, such as groups by day, by week, or by month, similar to a camera roll of a mobile phone. Advantageously, supervisor app 244 provides an easily assimilated view of activity by each user 104 associated with adult 240. Where discoveries 317 are grouped into discovery albums, photographs 1104 may be displayed in groups corresponding to the discovery album.
[0117] FIG.12 is a table 1200 that illustrates developmental characteristics for targeted age groups of learning system 100 of FIG.1, in embodiments. learning system 100 may automatically select the appropriate level of response based on a determined age of user 104, (e.g., from birthday defined in app view 956).
[0118] FIG.13 is a high-level system diagram 1300 showing one example implementation of coordination server 202 of FIG.2, in embodiments. In FIG.13, coordination server 202 is represented as API instances 1302, where coordination server 202 is a distributed server with multiple instances. For example, coordination server 202 may include between one and one-hundred API instances 1302, is shown with API instance 1302(1) and 1302(2) for clarity of illustration. Coordination server 202 may be implemented using software such as Python. API instances 1302 may communicate using a message queue 1304. In embodiments, API instance 1302(1) controls text-to-speech service 206 and AI service 208 to generate audio response 236 (FIGs 2 and 4A) in response to image 136 received from interactive companion 102.
[0119] API instance 1302(1) may send audio response 236 to API instance 1302(2) for output to interactive companion 102. API instance 1302(1) may send audio response 236 to API instance 1302(2) via a message queue 1304. Message queue 1304 may be a real-time messaging / communication mechanism between backend services and may be implemented by AWS ElastiCache's Redis Cluster's Streaming feature.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0120] Backend service 120 may also include a load balancer 1306 that balances loads caused by servicing multiple interactive companions 102 (e.g., many different users 104) and multiple supervisor apps 244 (e.g., many different adults 240), between API instances 1302. * * *
[0121] Embodiments of learning system 100 may include one or more of the following features: 1. Camera-Activated Adventures 2. Interactive Story Library 3. Flashcard Collectibles 4. Real-World Story Triggers 5. Parent Copilot App Features The following description of the above features includes parenthetical numbers following terms recited by the method. The parenthetical number indicates that the element associated with the number in parentheses is an example of the term. For example, the description of camera-activated features recites “object (132),” which means that aforementioned dog 132 is an example of the object. 1. Camera-Activated Adventures
[0122] Feature Description. Camera-Activated Adventures allow a child to take a photograph of any object (132) and instantly launch a tailored language-learning mini- game based on that object. This feature may create an interactive game scenario with branching logic. For example, if a child snaps a picture of an apple, the system might initiate an adventure game themed around an apple character or the concept of apples. The child’s device (or a connected app) presents a series of prompts or questions about the object – such as its color, taste, or usage – all narrated in the target language. The child participates by responding verbally in that language. The game’s narrative branches in real-time according to the child’s spoken answers, creating a choose-your- own-adventure style dialogue.
[0123] This feature combines image recognition (to identify the photographed object and generate relevant content) with voice recognition (to understand the child’s responses and progress the game). It provides immediate, engaging feedback: if the child responds correctly (or in an expected way), the story or game progresses downPATENT Attorney Docket No: WELI.P2001WO / 00640412 one branch; if not, the system can gently correct the child or offer hints, keeping the experience fun and educational. This real-time loop of prompt → child’s spoken reply → system reaction leverages the AI backend to “recognize, evaluate and respond” to the child’s speech , thereby turning passive object identification into an interactive language exercise.
[0124] For example, when a child may use learning system 100 to capture an image an object (or a representation of the object such as a displayed image of the object) associated with location, learning system 100 may initiate a conversation with choices such as “how would you like to get there” by presenting two or more options: car or train (in the target language, e.g., Spanish). The options may be displayed on central portion 114. Then child may go multiple levels deeper from here to continue the adventure and learn more about the language.
[0125] Technical Workflow. When the child presses the camera button and captures an image, the device’s backend uses computer vision algorithms (e.g. a trained image classification AI) to identify the main object or concept in the photo. Based on the recognized object, the system retrieves or dynamically generates a mini-game script related to that object. This script includes a sequence of scenes or challenges – for instance, a friendly character might ask the child to choose between options (“Shall we eat the red apple or plant the green apple seed?”) in the target language. Each choice corresponds to a branch in the storyline. The child’s verbal response is captured via the microphone and sent to a speech recognition module. The speech recognition interprets what the child said (and may also evaluate pronunciation or correctness of the vocabulary).
[0126] The backend then determines which branch was chosen (e.g. the child said “red” vs “green” in the target language) and returns the next segment of the game accordingly. This loop continues, creating an interactive dialogue where the child’s voice drives the progression. Notably, the entire interaction remains in the target language to maximize immersion. The system may use natural language processing to allow some flexibility in the child’s responses – for example, recognizing synonyms or phrases, not just single-word answers, to keep the game fluid. All the while, the character’s voice (via text-to-speech or prerecorded audio) provides feedback and carries the narrative.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0127] The result is a branching conversational mini-game generated on-the-fly from the child’s environment, transforming any real-world object into a language adventure. Embodiments of learning system 100 hence combine contextual object- based content generation and verbal interactive branching logic for language learning.
[0128] Features of Camera-Activated Adventures. Camera-Activated Adventures introduce specific features. These include: 1.1. Dynamic mini-game generation centered on the identified object, as opposed to a static audio description or fact. 1.2. Branching scenarios where the child’s choices influence outcomes – essentially a real-time decision tree of content – which is beyond the one-way audio response contemplated earlier. 1.3. Voice input as the control mechanism for navigation. This mechanism spells out using speech to make structured choices in a game. These additions include an interactive pedagogical method (gameplay through speaking) that harnesses both vision and speech AI.
[0129] FIG.14 is a flowchart illustrating a method 1400 for generating a language-learning adventure based on a captured image. Method 1400 may be implemented by one or more aspects of learning system 100, e.g., by interactive companion 102, backend service 120 or a combination thereof. Method 1400 includes at least one of steps 1410, 1420, 1430, 1440, 1450, and 1460.
[0130] Herein, description of method 1400 and other methods herein (e.g., methods 1600, 1700, and 1800) may include parenthetical numbers following terms recited by the method. The parenthetical number indicates that the element associated with the number in parentheses is an example of the term. For example, the description of step 1410 below recites “capturing, by a camera (402) of a user-operated device,” which means that camera 402 of interactive companion 102 is an example of the camera of step 1410.
[0131] Step 1410 includes capturing, by a camera (402) of a child-operated device (102), an image (136) of an object (132) in the child’s environment. Step 1420 includes processing the image using an AI-based vision model (208) to identify the object and determine one or more attributes or categories associated with the object (the “discovery”).PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0132] Step 1430 includes generating a scripted interactive mini-game tailored to the identified object. The mini-game includes a sequence of narrative segments and branching decision points related to the object’s attributes. The interactive mini-game may involve educational queries about the object’s characteristics – including at least one of color, shape, size, quantity, taste, or function. The child’s correct verbal answers in the target language cause the story to progress along a positive reinforcement path, while incorrect or different answers trigger corrective feedback or an alternate branch of the story (providing additional hints or practice).
[0133] In embodiments, a script of the mini-game is generated or adapted by an AI narrative engine (208) based on the identified object, such that different objects yield distinct game storylines. The branching logic (branching decision points) may be dynamically adjustable based on the child’s age or language proficiency (e.g., offering simpler binary choices for beginners versus open-ended questions for advanced learners).
[0134] Step 1440 includes outputting a first prompt or question to the child at a first decision point. The prompt may require the child to respond to the prompt, e.g., to make a choice or input by speaking aloud. The first prompt or question may be in a target language, in which case responding to the prompt may require the child to speak aloud in the target language.
[0135] Step 1450 includes capturing the child’s response. When the response is a vocal response (230), step 1450 may include interpreting the response using a speech recognition engine to determine the child’s choice or answer. The response may be captured by a microphone (406). Said determining may include evaluating whether the spoken input matches an expected word or phrase in the target language. The response may be a verbal response (e.g., vocal or audio), a gesture detected by a camera of learning system 100, or a press of a button of learning system 100.
[0136] Step 1460 includes selecting a next narrative branch of the mini-game based on the interpreted response. Each branch may correspond to a different outcome or path in the mini-game. Step 1470 includes outputting, to an output component of the device, a next segment of the narrative in the target language according to the selected branch, thereby advancing the interactive adventure in real time based on the child’s spoken inputs. The output component may be a speaker (408) or display (430) of the device.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0137] In method 1400, learning system 100 may monitor the child’s pronunciation and word usage during the adventure. In such embodiments, learning system 100 may provide immediate feedback on pronunciation quality or correctness as part of the game dialogue (e.g., requiring the child to repeat a word correctly to advance), thereby integrating real-time language coaching into the gameplay. 2. Interactive Story Library
[0138] Embodiments of learning system 100 that include an Interactive Story Library include at least one of the following features: 2.1. a menu-driven story selection interface for children 2.2. the integration of voice-driven branching narrative as a primary mode of interaction (versus the camera-initiated paradigm), and 2.3. using speech recognition to advance a plotted story. As such, these embodiments of learning system 100 result in its functioning as an interactive story narrator that includes a library of multilingual interactive stories for language practice.
[0139] Feature Description. The Interactive Story Library is a user experience that provides a collection of re-imagined stories (such as fairy tales, adventures, or original narratives), which a child can select and engage with through voice-interactive reading. Unlike on-demand, object-triggered content, this feature offers a curated library of stories accessible at any time – effectively a digital bookshelf of interactive, language-learning stories.
[0140] The child (with or without a parent’s help) browses the library via the device or a companion app and chooses a story. Once a story is launched, the system narrates it (using the target language) in segments. At key plot points, the narration pauses and the child is prompted to make a decision or answer a question by speaking, which is an example of audio response 236.
[0141] In a first example, learning system 100 presents an interactive Mandarin Chinese story, the narrator might ask, “The little dragon comes to a river. Should it fly over or go around? You decide – say it in Chinese!” The child’s spoken response (“^过 去” for fly over, or “绕道” for go around, for instance) is recognized by the voice recognition system. FIG.15 illustrates a second example of an interactive story, in whichPATENT Attorney Docket No: WELI.P2001WO / 00640412 learning system 100 presents the child with three choices displayed on central portion 114. The choices prompt the child to speak in a target language, in this case Spanish.
[0142] The story then branches accordingly and continues with the next segment in response to the child’s choice. This turns passive story time into an active language exercise: the child practices speaking words or phrases in context to influence the story. Importantly, the voice interaction may be natural and open-ended – not limited to pressing buttons or choosing from a screen, but actually speaking the choice, which is a more engaging and educational experience. The technology ensures that even young learners can participate; the system can be tuned to listen for specific expected phrases (possibly displaying visual hints or options on a screen if available), and robustly handle variations or mispronunciations by using confidence scoring.
[0143] When the child’s response isn’t recognized or is incorrect, learning system 100 may gently reprompt or provide the correct phrase, then allow the child to try again, thus creating a supportive learning environment. This feature essentially merges interactive storytelling with speech-driven interaction to reinforce language skills. Embodiments of learning system 100 hence provide a library of selectable stories and a dedicated voice-driven UX for story progression, independent of taking a new photo.
[0144] Technical flow. Embodiments of the Interactive Story Library involve both front-end UX and back-end AI components (120). On the front end, the child may be presented with a menu or library interface (likely icon-based for pre-readers, or text titles for older kids) to choose a story. Once a story is selected, learning system 100 retrieves the story data (which could be a predefined script with branches, or dynamically generated content following a known theme). The story is structured into segments separated by decision points. Learning system 100 plays or displays the story’s first segment. When a decision point is reached, the system uses text-to-speech (206) (or pre-recorded audio) to ask the user a question or present choices in the target language.
[0145] On the back end, a speech recognition module listens for the child’s spoken answer (236). The recognition system may be configured to listen for particular keywords / phrases corresponding to the known choices (e.g., expecting the word for “fly” vs “go around”), or even do open-ended parsing if the interaction is more free- form. Natural Language Understanding (NLU) logic then maps the recognized utterancePATENT Attorney Docket No: WELI.P2001WO / 00640412 to the next branch of the storyline. Once the branch is determined, the next segment of the story is delivered. This continues until the story concludes.
[0146] Throughout the process, learning system 100 may log the child’s responses (what they said and whether it was correct or intelligible) for learning assessment purposes. The interactive stories are likely authored or generated to provide multiple outcomes or paths, enhancing replay value – a child could re-read the same story and make different choices to practice different vocabulary. Additionally, because the stories are re-imagined for interaction, they might incorporate educational elements (such as teaching new words in context or cultural notes) seamlessly in the narrative.
[0147] Technically, this feature leverages learning system 100’s ability to output high-quality audio (for narration) and microphone 406 and / or ASR (Automatic Speech Recognition) to enable voice input. It can be thought of as a story-based skill of the interactive companion, distinct from the camera discovery mode. Notably, this mode might not require using camera 402 during the story (except perhaps if a story prompt asks the child to find something to show, but that blends into other features). It’s a self- contained interactive reading experience guided by the child’s voice. Backend service 120 may also function to mass generate / transform those stories and localize to all languages and personalize to all levels.
[0148] FIG.16 is a flowchart illustrating an interactive storytelling method 1600, which may be implemented by one or more aspects of learning system 100, e.g., by interactive companion 102, backend service 120 or a combination thereof. Method 1600 includes at least one of steps 1610, 1620, 1630, 1640, 1650, 1660, and 1670.
[0149] Step 1610 includes providing a library of interactive story content accessible via a user interface on a child-operated device (102), wherein each story in the library has a narrative with one or more decision points at which user input is required. Step 1620 includes receiving a user selection of a particular story from the library and initiating playback of that story on the device. Step 1630 includes outputting a first segment of the story’s narrative to the user in a target language. Said outputting may be via audio through a speaker (408) and / or text on a display (114).
[0150] Step 1640 includes, at a predefined decision point in the story, presenting the user with a prompt in the target language that elicits a spoken choice or response. The prompt may offer two or more narrative options that the user can select from byPATENT Attorney Docket No: WELI.P2001WO / 00640412 speaking. The prompt at the decision point may require the user to speak a word or phrase in the target language that corresponds to a choice, e.g., such as speaking the name of a desired action or item. The system may use natural language processing to match the user’s utterance to an expected choice even if the utterance is not an exact predefined phrase (allowing for minor pronunciation errors or synonyms), thereby creating a robust voice-driven selection mechanism for story branching.
[0151] When the user’s spoken input is not recognized or is incorrect, method 1600 may include providing real-time feedback or guidance. Examples of such feedback and guidance include repeating the prompt, providing the correct phrase in the target language, or encouraging the user to try again – such that the user receives pronunciation support and additional chances to respond, turning potential errors into learning opportunities.
[0152] Step 1650 includes capturing the user’s spoken response via a microphone (406) and processing the response with a speech recognition module (204) to interpret the user’s words.
[0153] Step 1660 includes determining, based on the recognized response, a next storyline segment from the available branches. For example, step 1660 may include mapping the spoken word or phrase to one of the offered choices in the narrative. Step 1670 includes continuing the story by outputting the next segment of the narrative corresponding to the determined branch. Step 1670 hence includes advancing the story in accordance with the user’s verbal input, until a story endpoint is reached.
[0154] The interactive story of method 1600 may be used as a language learning exercise by embedding vocabulary and phrases in the narrative and requiring the user to use those words appropriately to progress. Method 1600 may include tracking the user’s spoken responses and generating a performance report or summary at the end of the story (e.g. which words the child pronounced correctly or struggled with, how many branching choices were taken), which can be stored in the user’s learning timeline for review by a parent or educator. 3. Flashcard Collectibles
[0155] Key aspect of Flashcard Collectibles are as follows: 3.1. Creating a persistent, gamified vocabulary collection from the child’s real-world interactions, andPATENT Attorney Docket No: WELI.P2001WO / 00640412 3.2. Using those collected items in further interactive exercises to reinforce learning. This is different from a simple log or list of past items – it adds game design elements and a concrete representation (cards) for each learned word. It encourages active recall practice as part of play. Embodiments of learning system 100 that include flashcard collectables may auto-generate these collectable cards from real-world camera input and integrate them into an AI-powered interactive system. This functionality bridges the gap between augmented reality (capturing real objects) and a virtual collection of knowledge.
[0156] Feature Description. Flashcard Collectibles is a gamified vocabulary system present in embodiments of learning system 100. In essence, when a child takes a photo of an object (a “discovery”), interactive companion 102 not only engages in an immediate learning interaction (like describing the object or playing a mini-game), but now also generates a collectible digital flashcard for that object.
[0157] These flashcards may be presented in a fun, collectible-card-style format – meaning they are visually engaging, potentially with a character or graphic representing the word, and can be “collected” and reviewed over time. Each flashcard may include at least one of the object’s image or an illustration, the word in the target language (with possibly phonetic help), and an audio clip of pronunciation. The child builds a personal collection of these vocabulary cards as they explore new things with camera 402. Over time, the collection becomes a tangible record of words the child has learned, turning language acquisition into a collecting game.
[0158] In embodiments, the flashcards are not static; the system leverages them in spaced-repetition practice and mini-games. For example, the app might have a flashcard quiz game where the child is shown a card from their collection and asked to pronounce the word or answer a question about it. Interactive companion 102 may implement a memory matching game that uses the cards.
[0159] Interactive companion 102 may include a user interface that allows the child or parent to browse the full collection of vocabulary flashcards, select cards for review, and see summary statistics (e.g. “Words collected: 50”, or individual card mastery levels). This user interface may ties into the parent’s dashboard as well, e.g., via supervisor app 244, giving the user of computing device 242 insight into which words have been learned through the child’s explorations.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0160] Technical Workflow. When an image is captured and backend service 120 identifies the object (say the child takes a photo of a “dog”), learning system 100 may generate a vocabulary flashcard entry for “dog” in the chosen target language. This involves retrieving the appropriate translation of the word (if the child is learning Spanish, “perro” would be on the card), and possibly an illustrative image or using the actual photo taken as the card image. The card might also include or store metadata like the date and context of discovery, example sentences, or an associated character. The card is then added to the child’s digital collection, which could appear as a gallery in the app or device. The UX might show a notification or animation (“New Card Collected!”) to reward the child.
[0161] Backend service 120 may store these cards in the user’s profile and may schedule them for review (e.g., using spaced repetition algorithms to prompt the child to revisit “dog” after a few days). The Parent Copilot app (discussed later) may also display the collection or highlight new cards to parents. Additionally, mini-games may leverage the flashcards: for instance, a pronunciation game might cycle through a few cards and ask the child to say each word, using the speech recognition to score how well they pronounced it.
[0162] Another game could be a memory game where the child matches a spoken word to one of the cards. Technically, this feature may access a database of vocabulary linked to object recognition (to know what word / card corresponds to a detected object) and the ability to generate the card media (which might involve an image library or even AI-generated cartoon images for each object). It also requires tracking user progression (which cards collected, how many, etc.) and integrating with the learning games logic. Importantly, since this is aimed at children, the design of the cards and games is kept playful and encouraging rather than test-like. For instance, each card could have levels or “evolution” (in a Pokémon sense) that upgrades as the child masters the pronunciation or usage of the word in different contexts – an idea that further gamifies the learning process.
[0163] FIG.17 is a flowchart illustrating a method 1700 for generating and utilizing collectible digital flashcards in an interactive learning system. Method 1700 may be implemented by one or more aspects of learning system 100, e.g., by interactive companion 102, backend service 120 or a combination thereof. Method 1700 includes at least one of steps 1710, 1720, 1730, 1740, 1750, and 1760.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0164] Step 1710 includes identifying an object (132) in an image. Step 1710 may include capturing the image of the object via a learning device’s camera (402) and identifying the object as a new “discovery” using an image recognition process. The image recognition process may employ AI.
[0165] Step 1720 includes creating a digital vocabulary card for the identified object. The vocabulary card includes at least one of: the object’s name in the target learning language (and optionally in the child’s native language), an image or illustration of the object, and metadata for learning (including pronunciation guidance or an audio recording of the word).
[0166] Method 1700 may include rendering the digital vocabulary card, e.g., on central portion 114. The rendered vocabulary card may have game-like attributes and collectible features, including one or more of: a point value or rarity level assigned to the card (to reward discovery of uncommon objects), a character or avatar linked to the word that can “level up” as the child masters the word, and categories or series of cards that encourage the child to find related items (e.g. a set of “animal” cards, “food” cards, etc., to collect complete sets).
[0167] In a use scenario, the captured image captured image corresponds to a concept for which a flashcard already exists in the user’s collection. Method 1700 may include determining whether the captured image corresponds to an existing flashcard. In such cases, instead of or in addition to step 1720, said determination may trigger an alternative interaction – such as a brief quiz or an advanced fact about that object – instead of creating a duplicate card, thereby reinforcing existing knowledge and keeping the collection unique. (In this way, the feature handles repeat discoveries intelligently by switching into practice mode for known vocabulary rather than simply repeating initial content.)
[0168] Step 1730 includes storing the digital vocabulary card in a user-specific collection accessible within the interactive learning system (100). Step 1730 may include or result in updating the user’s profile to include the newly learned item. Step 1740 includes presenting a notification or visual indicator that a new collectible flashcard has been obtained. The notification may be presented by display (430) of microphone (406) of the interactive learning system.
[0169] Step 1750 includes integrating the stored vocabulary cards into subsequent learning activities. Step 1750 may include, as part of said integrating,PATENT Attorney Docket No: WELI.P2001WO / 00640412 retrieving one or more cards from the user’s collection and engaging the child in an interactive mini-game or quiz involving those cards. Examples of said engagement include prompting the child to pronounce a word from a card, or to recall the meaning of a word when shown the card, and using speech recognition to evaluate the child’s response.
[0170] Step 1760 includes tracking the child’s performance with each flashcard (such as accuracy of pronunciation or recall). Step 1760 may also include updating the card’s status or the child’s progress metrics accordingly, which can inform spaced repetition scheduling or difficulty tuning for future games. 4. Real-World Story Triggers
[0171] Feature Description. This feature expands on the aforementioned interactive storytelling of embodiments of learning system 100 by using real-world context as the springboard for narrative adventures. Real-World Story Triggers may be viewed as a hybrid of the camera-based discovery and the interactive story concepts: a child takes a photo of not just a single object, but of a broader scene or environment (e.g. their backyard, a park, their kitchen), and the system responds by launching a choose-your-own-adventure style narrative (e.g., a branching narrative) that is grounded in that real-world setting. In effect, the child’s immediate surroundings become the stage and props for a story.
[0172] For example, a child at the park might snap a photo of the playground. The system could interpret the scene – recognizing key elements like trees, grass, maybe a slide or swing in the image – and initiate a story such as “Adventure in the Enchanted Park”, where those recognized elements become part of the plot (e.g., a talking tree character or a magical swing ride). The story unfolds as a series of scenarios where the child must make choices or perform actions, again using voice input in the target language to advance the plot. The story may have a branching narrative.
[0173] The difference from Camera-Activated Adventures (feature 1) is scale and context: instead of focusing on a single object’s attributes, Real-World Story Triggers incorporate multiple elements and the overall setting to generate a richer narrative experience. It’s like the system is telling a story about the place the child is in or the scene they captured. This could include incorporating the location – if location services arePATENT Attorney Docket No: WELI.P2001WO / 00640412 available, the system might know the child is “at the park” or “in the kitchen” and tailor the story theme accordingly (outdoor adventure vs. cooking quest, for instance).
[0174] The interactive voice-based branching is similar to features 1 and 2: the child listens to part of the story and then responds to prompts to influence what happens next. Here, the content is highly contextualized to the child’s current real- world environment, making the experience feel personalized and immersive. It blurs the line between reality and imagination – the child might hear the device narrate, “The slide you see in front of you turns into a rainbow slide to a magical land – do you climb up or stay on the ground? Say your choice!” and the child’s choice (spoken in the target language) will determine what happens. This feature leverages advanced image recognition (to detect multiple objects and scene type) and possibly location data, combined with AI content generation to craft a narrative on the fly. It uses augmented reality principles (using the real world as content) in a voice-interactive story format.
[0175] Technical Workflow. A real-world story process may begin when the child takes a photo intended as a “story trigger.” learning system 100’s computer vision backend 120 analyzes the image (136) for both specific objects and overall scene context. It might use object detection algorithms to list out key items (e.g., “tree”, “bench”, “dog”, “ball”) and scene recognition to classify the environment (“park” or “outdoor nature”, or “kitchen” etc.).
[0176] Learning system 100, e.g., backend service 120, then either selects a pre- existing story template that matches the scene or dynamically generates a storyline incorporating these elements. For example, if the recognized context is “kitchen”, the system could choose a cooking adventure template and populate it with the specific items detected (spoon, apple, etc. become items the story characters use). The story is then delivered segment by segment as with interactive stories: narrated via audio with points where the child must respond. Learning system 100 may employ voice recognition handles the child’s inputs, and the narrative branches accordingly.
[0177] One technical challenge is ensuring the story logic can gracefully handle varying inputs – one scene might have a dog and a tree, another might have no living beings but a lake and mountains, etc. Learning system 100 may therefore employ an AI storytelling engine or a set of flexible story frameworks, e.g., as part of backend service 120. This could involve using machine learning models like large language models toPATENT Attorney Docket No: WELI.P2001WO / 00640412 improvise a story with the given elements, or a rule-based engine with modular story pieces. The outcome is an on-demand personalized story.
[0178] Voice interactions between the child and interactive companion 102 may include not just choices but also possibly having the child role-play speaking lines for a character or making sound effects in the target language, to further enrich the experience. Learning system 100 may capture one or more of such interactions evaluate them (for language practice), similar to prior features. Because these are longer narratives, the system may also introduce a “happiness” or engagement meter to adjust the story difficulty or length. For example, if the child is very engaged (speaking a lot and excited), the story might offer more branches, whereas if engagement drops, the story might conclude sooner to maintain interest. At the end, as usual, the session may be logged to the timeline with a summary (e.g., “Story: Enchanted Park – Completed” along with any new words learned).
[0179] Key features of the Real-World Story Triggers include the following. 4.1. Multi-object recognition in a scene to inform content (versus a single object discovery) – the system synthesizes a narrative from a constellation of recognized inputs. 4.2. Location-aware storytelling – optionally using GPS or context to choose story themes. 4.3. Dynamic narrative generation – potentially employing AI storytelling engines to create a unique story instance for the user’s scene, which goes beyond predefined content. 4.4. Maintaining the choose-your-own-adventure style (e.g., branching narrative) voice interaction in this generated story.
[0180] FIG.18 is a flowchart illustrating a method 1800 for generating an interactive narrative based on a real-world scene captured by a user. Method 1800 may be implemented by one or more aspects of learning system 100, e.g., by interactive companion 102, backend service 120 or a combination thereof. Method 1800 includes at least one of steps 1810, 1820, 1830, 1840, 1850, 1860, 1870, 1880, and 1890.
[0181] Step 1810 includes receiving an image (136) of a scene captured by a child-operated learning device (102). The image may include multiple elements such as objects, scenery, or context indicators. Step 1810 may also include capturing the image.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0182] Step 1820 includes analyzing the image using one or more AI models to identify a set of elements (132, 134) present in the scene. The analysis may include object recognition to detect distinct items and scene classification to determine an environmental context of the image. For example, learning system 100 may recognizing objects such as “tree, slide, dog” and determining the scene is “outdoor park.”
[0183] Step 1830 includes outputting a narrative storyline based on the analyzed image. Herein, the narrative storyline is also referred to as the story. Outputting the narrative storyline may include selecting and / or generating a narrative storyline. All or part of the narrative storyline may be selected from previously-generated, existing, or “stock” storylines stored in a memory accessible to learning system 100. Examples of the memory include memory 412 and a memory accessible by backend service 120. The narrative storyline may be thematically tailored to the identified environmental context and that incorporates at least a subset of the recognized objects as characters or props in the narrative story.
[0184] In step 1830, the production of the narrative storyline may be performed by an AI storytelling engine (208) that uses the set of identified scene elements as inputs to craft a personalized story. The resulting story may be unique to the user’s captured scene and may vary even for the same location depending on which objects (132, 134) are present or emphasized. For example, two different images of a park might yield different adventures: a first image that focuses on a playground slide yields a story about a magical slide, while a second image that focuses on a pond yields a story about aquatic creatures – each using the child’s actual surroundings as plot elements.
[0185] In step 1830, generating the narrative storyline may include utilizing additional context data in selecting the narrative, including at least one of: the geolocation of the device (to identify the place, if known, and match it with a relevant story template, such as “farm”, “zoo”, “supermarket”), the time of day or weather (for context, e.g., a story that mentions nighttime if it is evening), and recent user learning history (for example, favoring a story that reinforces a vocabulary word the child learned recently).
[0186] In embodiments, step 1830 includes incorporating interactive learning moments into the story. For example, if a stop sign is detected in the image, the story’s character might encounter a stop sign and ask the child to say “stop” in the target language to proceed, thereby teaching that vocabulary in context. In embodiments,PATENT Attorney Docket No: WELI.P2001WO / 00640412 method 1800 includes requiring receiving the child’s correct verbal response for before the story continues, this linking the real-world object, the narrative, and the language exercise together.
[0187] Step 1840 includes initiating an interactive story that is based on the selected / generated narrative. The interactive story may have a branching-narrative. Said initiating may include outputting an introductory scene to the child via audio (in the target language) that sets up a scenario involving the real-world scene or objects. Step 1840 may include referencing elements identified in step 1820 as part of the story setting.
[0188] Step 1850 includes, at one or more junctures in the interactive story, presenting the child with a decision point or interactive prompt (in the target language) that requires the child to respond, e.g., verbally, to influence the story’s direction. The decision point or prompt may require the child to indicate what the character should do with a recognized object, or how to overcome an obstacle in the environment.
[0189] Step 1860 includes capturing the child’s response of step 1850. The response may be a spoken input. Capturing may include capturing via the device’s microphone and utilizing speech recognition to interpret the input and determine the child’s choice or instruction.
[0190] Step 1870 includes adjusting, based on the child’s input, a narrative path of the interactive story to yield an updated path. Step 1870 may include selecting the next story segment or outcome branch that corresponds to the child’s choice. The outcome branch may involve different real-world elements among those identified by the child. For example, if the child indicates to learning system 100 to interact with a dog in the image rather than to climb a tree in the image, the subsequent story segment focuses on the indicated element (the dog). The adjusting of step 1870 may be dynamic, e.g., in response to multiple interactions of step 1860 that generate multiple responses from the child.
[0191] Step 1880 includes continuing to deliver the interactive story with the updated path. Said continuing may incorporate additional real-world details from the image as the story progresses. One or more of steps 1820–1880 may be iterated until a narrative conclusion is reached.
[0192] Step 1890 includes, upon conclusion of the story, storing a record of the story session in the user’s timeline or log. The session may include data on which real-PATENT Attorney Docket No: WELI.P2001WO / 00640412 world elements were used and which choices or words the child spoke, enabling review of real-world-triggered learning interactions by parents or educators via the companion app. 5. Parent Copilot App Features
[0193] Embodiments of supervisor app 244 of computing device 242 may include one or more of the following features: personalized language goals and profiles, an interaction timeline with detailed context, pronunciation scoring feedback, and smart activity suggestions.
[0194] Personalized Language Goals. In embodiments, supervisor app 244 enables users (e.g., parents and educators) to set up a tailored language learning profile for the child. This includes selecting target language(s), setting proficiency level or age level, and defining specific goals (e.g., “learn 100 new words by end of the month” or “practice speaking for 10 minutes daily”). Supervisor app 244 then tracks a child’s progress toward these goals, using data from the child’s interactions. For example, if the goal is daily speaking practice, the supervisor app 244 checks how many minutes the child converses with interactive companion 102 each day. If the goal is vocabulary acquisition, it looks at how many unique flashcards have been collected or words successfully used. The language learning profile might also allow customization of content preferences (like if a parent wants more STEM-related content or story content).
[0195] FIG.19 shows app views 1910 and 1920 related to the language goals personalization described above. One or both of app views 1910 and 1920 may be part of embodiments of supervisor app 244.
[0196] Timeline of Interactions (Enhanced). While the original system logged discoveries and interactions in a timeline accessible to parents, the enhanced timeline is much more detailed and user-friendly. In the Parent Copilot app, parents can see a chronological feed of the child’s activity – each entry might show a thumbnail of the photo taken, the word learned or game played, the date / time, and any outcomes (such as “pronunciation score: 4 / 5 stars” or “story completed: The Little Red Riding Hood in Mandarin”). It’s akin to an interactive diary of learning sessions. Parents can tap on an entry to get more details – for instance, tapping a “flashcard collected” entry could show the card (word, picture, etc.), or tapping a story session could show which choices thePATENT Attorney Docket No: WELI.P2001WO / 00640412 child made or even a transcript of the child’s spoken responses (given the system’s speech recognition, transcripts of what the child said can be saved). This timeline not only keeps parents informed but also celebrates the child’s progress in a shareable way. Parents and children might review the timeline together, reinforcing learning (“Tell me about this picture you took – oh, it was an apple and you played a game about colors in Spanish!”). The timeline hence turns raw data logs into a human-readable format.
[0197] Pronunciation Scoring. Embodiments of computing device 242 may include pronunciation evaluation. Using the speech recognition data (e.g., of speech-to- text service 204), the backend service 120 analyzes how accurately the child pronounced words in the target language during their interactions. Supervisor app 244 may give a score or rating for pronunciation in various tasks. For example, if the child was repeating words in a flashcard game, learning system 100 might give each attempt a score (perhaps behind the scenes using a technique like Goodness of Pronunciation (GOP) scoring or comparing against native speaker models). Supervisor app 244 may display something like “Pronunciation of ‘apple’ – 80% (Good, could improve the ‘l’ sound)”.
[0198] To keep it child-friendly, interactive companion 102 might use stars or playful feedback, whereas the parent might see more detailed metrics or tips. Over time, supervisor app 244 may identify problem areas (e.g., certain vowels the child consistently mispronounces) and suggest practice.
[0199] Suggested Activities (Camera Challenges). Embodiments of supervisor app 244 generate and display specific camera-based missions or prompts for the child. For example, the app might suggest: “Language Challenge: Snap five fruits today and learn their names!” or “Try finding something blue and ask Dex about it in French.” These suggestions may be personalized based on the child’s recent activity and goals. For example, if learning system 100 notices the child hasn’t done any food-related words, it might encourage finding fruits; if the child has been inactive, it may prompt an outdoor exploration. Adult 240 can see these prompts and encourage the child to undertake them, or even do it together as a fun task. Some suggestions may be time-based (“this week’s challenge”) or tied to events (e.g., “It’s spring – find a flower and see what Dex says”).PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0200] In supervisor app 244, once a challenge is issued, supervisor app 244 can track its completion (for instance, it knows when five fruit photos have been taken) and then reward the child in-app (maybe a special badge or just a congratulatory message).
[0201] Technical Workflow. For the Parent Copilot feature, functionality may occur in backend service 120 and / or computing device 242. The personalized goals feature may include a settings interface for adult 240 to input or select goals and a profile setup (which could be done at onboarding or anytime). Learning system 100 then maps these goals to measurable metrics in the child’s data (like number of interactions, vocabulary count, streak of daily usage, etc.). On each sync or refresh, supervisor app 244 queries backend service 120 for progress (backend service 120 may compile stats from the child’s timeline). Graphs or progress bars may be shown for each goal. The timeline feed uses the logged data of interactions – the backend likely stores each interaction with type, timestamp, details (object recognized, content delivered, etc.). Supervisor app 244 may pull this and present it nicely. Images or icons for each entry may require storing thumbnails of the child’s photos or representative icons for activities. Privacy and security are considered (the data is only visible to authorized parents).
[0202] The pronunciation scoring involves analyzing audio clips. Likely, when the child’s speech is processed, learning system 100 may calculate a score immediately (e.g., using phoneme confidence scores from the speech recognizer or a specialized pronunciation model). These scores are saved per word or per session. Supervisor app 244 may then display either an aggregate (like an average score today) or itemized feedback. It might highlight words that were difficult. This might also tie into goals (e.g., a goal to reach a certain pronunciation score).
[0203] The suggested activities may be generated by rules or AI: learning system 100 checks the child’s recent learning history and goal state, then formulates a suggestion. For example, if a goal is incomplete (“5 fruits this week” and only 2 done), it reminds about that. Or it notices the child hasn’t done any new discovery in a few days, so it suggests something fresh. These suggestions are delivered to supervisor app 244 as notifications or a list of “Today’s Suggestions.” Adult 240 may then encourage the child, effectively making the parent an active participant. Technically, some of these suggestions might also be delivered directly to the child through interactive companion 102’s character voice (“Dex”) – for instance, Dex might say, “Let’s go find some fruits toPATENT Attorney Docket No: WELI.P2001WO / 00640412 take pictures of!” – but supervisor app 244 may be where the logic can be managed and monitored.
[0204] Embodiments of supervisor app 244 executed by computing device 242 may include one or more of the following features: (a) Personalized Goal Setting, (b) Enhanced Timeline Feed, (c) Pronunciation Scoring and Feedback, (d) Activity Suggestions and Camera Challenges. Each of these is described below.
[0205] (a) Personalized Goal Setting: An interface defines individual learning goals and profile preferences for the child, including target language selection, skill level, and quantitative goals (vocabulary count, daily speaking duration, etc.). The system tracks progress toward said goals using data from the child’s interactions (automatically updating metrics like number of new words learned or minutes of conversation recorded).
[0206] (b) Enhanced Timeline Feed: A timeline view in supervisor app 244 displays a chronological log of the child’s learning interactions with context. Each entry includes information such as the image captured or activity type, the content or words learned, and the date / time, allowing the parent to review what the child has been doing (with options to drill down into details like transcripts or scores for each interaction).
[0207] Embodiments of the timeline feed are interactive, allowing adult 240 to click on an entry corresponding to a “camera discovery + flashcard” event to see the actual flashcard that was generated (image and word), or on a “story adventure” event to see the plot summary and which choices the child made, or on a “pronunciation practice” event to listen to a recording of the child’s attempt if desired. This provides transparency and the ability for adult 240 to reinforce learning offline (e.g., the parent can ask the child about those story choices or practice the flashcard words together), leveraging the recorded data to facilitate parent-child discussion and co-learning.
[0208] (c) Pronunciation Scoring and Feedback: Supervisor app 244 may include an analysis module that evaluates the child’s spoken responses captured by the device to determine pronunciation accuracy or fluency, generating a score or rating for the child’s pronunciation of words or phrases , and presenting this information through supervisor app 244 (along with qualitative feedback or tips for improvement), thereby giving adult 240 insight into the child’s speech development (for example, highlighting words that the child mispronounces and tracking improvement over time). FIG.20PATENT Attorney Docket No: WELI.P2001WO / 00640412 includes an app view 2010 illustrating example output of the pronunciation scoring feature of an embodiment of supervisor app 244.
[0209] The pronunciation scoring may utilize machine learning models trained on children’s speech in the target language to ensure accuracy in evaluation. The scoring result may be visualized in the app (running on interactive companion 102) using child- friendly graphics (such as star ratings or color-coded performance) to allow adult 240 to easily gauge the child’s speaking proficiency at a glance and share simplified feedback with the child (motivating them without overwhelming with technical details).
[0210] (d) Activity Suggestions and Camera Challenges. An automated suggestion engine that recommends language-learning activities to encourage continued engagement, wherein the suggestions are displayed to the parent (and / or directly to the child via the device) and are tailored to the child’s context and progress – including prompts tied to using the camera (e.g., “Treasure Hunt: find and snap pictures of five different animals to learn their names” or “Practice Colors: take a photo of something blue and ask Dex about it”). The system may monitor the completion of these suggested tasks and provides feedback or rewards upon completion.
[0211] Parent application (244) may also allow adult 240 to input feedback or preferences that influence the child’s learning experience, including: selecting content themes (for instance, if the parent wants more stories or more science facts, the system will adjust the content mix), adjusting difficulty levels, or scheduling “focus days” (e.g., a day to focus on a particular category of vocabulary), with the system then adapting the challenges and content delivered to the child accordingly.
[0212] FIG.21 is a flowchart illustrating a method 2100 for interactive discovering with interactive companion 102. Method 2100 is a combination of methods 500 and 700 introduced in FIGs.5 and 7, respectively, and illustrates example interaction between interactive companion 102 and coordination server 202 of FIG.2.
[0213] Features described above, as well as those claimed below, may be combined in various ways without departing from the scope hereof. The following enumerated examples illustrate some possible, non-limiting combinations.
[0214] Embodiment 1. An interactive educational method implemented through an interactive companion, includes: receiving, within a processor, an image captured by a camera of the interactive companion of an object in a real-world environment;PATENT Attorney Docket No: WELI.P2001WO / 00640412 processing the image to identify the object as a discovery; generating a learning experience for a user of the interactive companion using AI based on the discovery; sending the learning experience to the interactive companion for output to the user; and storing the discovery and the learning experience in a timeline of an account associated with the user.
[0215] Embodiment 2. The interactive educational method of embodiment 1, the generating the learning experience including generating the learning experience as an audio response.
[0216] Embodiment 3. The interactive educational method of either one of embodiments 1 and 2 further includes: generating a happiness level indicative of attention the user gives to the interactive companion over time; and sending the happiness level to the interactive companion for display to the user.
[0217] Embodiment 4. The interactive educational method of any one of embodiments 1–3 further includes: generating an insight from the timeline; and sending the insight to a supervisor app running on a computing device.
[0218] Embodiment 5. The interactive educational method of any one of embodiments 1–4 further includes: generating a subsequent learning experience based on the discovery; and sending the subsequent learning experience to the interactive companion for output to the user.
[0219] Embodiment 6. The interactive educational method of embodiment 1 further includes: generating a challenge defining a discovery goal; and sending the challenge to the interactive companion.
[0220] Embodiment 7. An interactive companion for a child, including: a selectable character; a wireless interface for communicating with an AI-powered backend; a camera for capturing an image from a real-world environment of the child; a microphone for capturing verbal input from the child; and a speaker for outputting a voice of the selectable character; wherein the AI-powered backend generates the voice of the selectable character to interact with the child, and to encourage exploration, discovery, and learning by the child.
[0221] Embodiment 8. The interactive companion of embodiment 7, wherein the interactive companion has no high-resolution display.
[0222] Embodiment 9. A method for personalized educational content generation for a child, including: receiving, from a child-operable device, an image of a real-worldPATENT Attorney Docket No: WELI.P2001WO / 00640412 environment; determining, based on the image, a discovery by the child; and generating, based on the discovery, a personalized learning experiences that includes at least one educational component selected from the group consisting of: STEM, language, and creativity.
[0223] Embodiment 10. The method of embodiment 9, wherein the personalized learning experience comprises an interactive story related to the discovery.
[0224] Embodiment 11. The method of either one of embodiments 9 or 10, wherein the personalized learning experience comprises a verbal description of the discovery.
[0225] Embodiment 12. The method of any one of embodiments 9–11 further includes outputting an audio response to the child.
[0226] Embodiment 13. The method of embodiment 12 further includes generating content of the audio response based on an age of the child.
[0227] Embodiment 14. The method of either one of embodiment 12 or 13, said outputting including generating a voice of a character selected by the child.
[0228] Embodiment 15. An AI-powered device for interacting with a child in a real-world environment, including: a housing having a first portion and a graspable portion; a battery; a camera; a microphone; a speaker; a communication interface for communicating with a backend; a button for triggering the camera and the microphone; a processor communicatively coupled with the camera, the microphone, the speaker, , the communication interface and the button; and a memory communicatively coupled with the processor and storing machine-readable instructions that, when executed by the processor, cause the processor to: capture an image of the real-world environment using the camera in response to detecting a short-press of the button; capture audio data using the microphone in response to a long-press of the button; send the image and the audio data to a backend via the communication interface; receive an audio response from the backend in response to the image or the audio data; and output the audio response via the speaker.
[0229] Embodiment 16. The device of embodiment 15 further includes a display, the processor being communicatively coupled thereto and further storing machine- readable instructions that, when executed by the processor, cause the processor to display the image on the display.PATENT Attorney Docket No: WELI.P2001WO / 00640412
[0230] Embodiment 17. The device of either one of embodiments 15 or 16 further includes LEDs positioned around an edge of the first portion of the housing, the processor being communicatively coupled to the LEDs and further storing machine- readable instructions that, when executed by the processor, cause the processor to control illumination of the LEDs in response to a frequency of presses of the button.
[0231] Embodiment 18. The device of any one of embodiments 15–17, the first portion having a shape that is circular or a polygon with five or more sides.
[0232] Embodiment 19. A method for generating a voice-driven branching narrative, includes: receiving, via a microphone of a device, spoken input in a target language at a narrative decision point; selecting, by a processor of the device or communicatively coupled to the device, a next branch of an interactive story responsive to the spoken input; and outputting, by a speaker of the device, a subsequent story segment.
[0233] Embodiment 20. A method for generating and utilizing collectible digital flashcards in an interactive learning system, includes: identifying an object in an image; creating a digital vocabulary card for the identified object; storing the digital vocabulary card in a user-specific collection accessible within the interactive learning system; and presenting a notification or visual indicator that a new collectible flashcard has been obtained.
[0234] Changes may be made in the above methods and systems without departing from the scope of the present embodiments. It should thus be noted that the matter contained in the above description or shown in the accompanying drawings should be interpreted as illustrative and not in a limiting sense. Herein, and unless otherwise indicated the phrase “in embodiments” is equivalent to the phrase “in certain embodiments,” and does not refer to all embodiments.
[0235] As used in this specification, any appendices thereto, and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise. It should also be noted that the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise. Regarding instances of the terms “and / or” and “at least one of,” for example, in the cases of “A and / or B,” “at least one of A and B,” and “at least one of A or B,” such phrasing encompasses the selection of (i) A only, or (ii) B only, or (iii) both A and B. In the cases of “A, B, and / or C, ” “at least one of A, B, and C,” and “at least one of A, B, or C,” suchPATENT Attorney Docket No: WELI.P2001WO / 00640412 phrasing encompasses the selection of (i) A only, or (ii) B only, or (iii) C only, or (iv) A and B only, or (v) A and C only, or (vi) B and C only, or (vii) each of A and B and C. This may be extended for as many items as are listed.
[0236] The following claims are intended to cover all generic and specific features described herein, as well as all statements of the scope of the present method and system, which, as a matter of language, might be said to fall therebetween.
Claims
AMENDED CLAIMS received by the International Bureau on 08 November 2025 (08.11.2025)What is claimed is:
1. An interactive educational method implemented through an interactive companion, comprising: receiving, within a processor, an image of an object in a real- world environment captured by a camera of the interactive companion when triggered by a user of the interactive companion ; processing the image to identify the object as a discovery; generating, using an artificial intelligence (Al) model, a learning experience comprising educational content that is selected based on (i) the discovery and (ii) a learning history of the user, the learning history comprising per-concept proficiency scores derived from automated evaluation of prior user interactions stored in a timeline; wherein selecting the educational content includes applying a spaced-repetition scheduler that schedules content for review for at least one concept related to the discovery; wherein generating the learning experience includes applying an age-group directive that constrains reading level and concept complexity and applies age-specific word usage; sending the learning experience to the interactive companion for output to the user; and storing the discovery and the learning experience in the timeline of an account associated with the user.
2. The interactive educational method of claim 1, the generating the learning experience comprising generating the learning experience as an audio response.
3. The interactive educational method of claim 1, further comprising: generating a happiness level indicative of attention the user gives to the interactive companion over time; and sending the happiness level to the interactive companion for display to the user.
4. The interactive educational method of claim 1, further comprising: generating an insight from the timeline; and sending the insight to a supervisor app running on a computing device.
5. The interactive educational method of claim 1, further comprising: generating a subsequent learning experience based on the discovery; and sending the subsequent learning experience to the interactive companion for output to the user.
6. The interactive educational method of claim 1, further comprising: generating a challenge defining a discovery goal; and sending the challenge to the interactive companion.
7. An interactive companion for a child, comprising: a selectable character; a wireless interface for communicating with an Al-powered backend; a camera for capturing an image from a real- world environment of the child; a microphone for capturing verbal input from the child; and a speaker for outputting a voice of the selectable character; wherein the Al-powered backend generates the voice of the selectable character to interact with the child, and to encourage exploration, discovery, and learning by the child.
8. The interactive companion of claim 7, wherein the interactive companion has no high- resolution display.
9. (Currently Amended) A method for personalized educational content generation for a child, comprising: receiving, from a child-operable device, an image of a real-world environment; determining, based on the image, a discovery by the child; maintaining, in a timeline associated with the child, per-concept proficiency scores derived from automated evaluation of prior interactions;generating, based on the discovery and the per-concept proficiency scores, a personalized learning experiences that includes at least one educational component selected from the group consisting of: STEM, language, and creativity; and selecting portions of the personalized learning experience for review using a spaced-repetition scheduler.
10. The method of claim 9, wherein the personalized learning experience comprises an interactive story related to the discovery.
11. The method of claim 9, wherein the personalized learning experience comprises a verbal description of the discovery.
12. The method of claim 9, further comprising outputting an audio response to the child.
13. (Currently Amended) The method of claim 12, wherein the content and vocabulary of the audio response are generated by applying an age-group directive that constrains reading level and concept complexity and enforces age-specific word usage .
14. The method of claim 9, wherein the automated evaluation comprises pronunciation scoring of spoken responses from the child and updates the per-concept proficiency scores .
15. An artificial intelligence (Al)-powered device for interacting with a child in a real- world environment, comprising: a housing having a first portion and a graspable portion; a battery; a camera; a microphone; a speaker; a communication interface for communicating with a backend; a button for triggering the camera and the microphone; a processor communicatively coupled with the camera, the microphone, the speaker, the communication interface and the button; anda memory communicatively coupled with the processor and storing machine-readable instructions that, when executed by the processor, cause the processor to: capture an image of the real-world environment using the camera in response to detecting a short-press of the button; capture audio data using the microphone in response to a long-press of the button; send the image and the audio data to a backend via the communication interface; receive an audio response from the backend in response to the image or the audio data; and output the audio response via the speaker.
16. The device of claim 15, further comprising a display, the processor being communicatively coupled thereto and further storing machine-readable instructions that, when executed by the processor, cause the processor to display the image on the display.
17. The device of claim 15, further comprising LEDs positioned around an edge of the first portion of the housing, the processor being communicatively coupled to the LEDs and further storing machine-readable instructions that, when executed by the processor, cause the processor to control illumination of the LEDs in response to a frequency of presses of the button.
18. The device of claim 15, the first portion having a shape that is circular or a polygon with five or more sides.
19. A method for generating a voice-driven branching narrative, comprising: receiving, via a microphone of a device, spoken input in a target language at a narrative decision point; selecting, by a processor of the device or communicatively coupled to the device, a next branch of an interactive story responsive to the spoken input; and outputting, by a speaker of the device, a subsequent story segment.
20. A method for generating and utilizing collectible digital flashcards in an interactive learning system, comprising:identifying an object in an image; creating a digital vocabulary card for the identified object; storing the digital vocabulary card in a user-specific collection accessible within the interactive learning system; and presenting a notification or visual indicator that a new collectible flashcard has been obtained.
Citation Information
Patent Citations
Systems and methods to interactively control delivery of serial content using a handheld device
US11941185B1
Adaptive, immersive, and emotion-driven interactive media system
US20160005326A1
Immersive story creation
US20160225187A1
Wearable multimedia device and cloud computing platform with laser projection system
US20210117680A1
Virtual learning environment for children
US6517351B2