An AI-powered device for interactive storytelling and image generation

AIKO addresses the limitations of existing storytelling devices by integrating speech recognition, text generation, and print functions into a single device, enhancing engagement and accessibility with graspable results, ensuring privacy and security through real-time processing.

DE202025100803U1Active Publication Date: 2025-07-10CHANDRANA FENIL RAJKOT +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE202025100803
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-10
Estimated Expiration
2035-02-28

AI Technical Summary

Technical Problem

Existing AI-based storytelling devices primarily focus on digital content generation without providing tangible results, are complex and require extensive user interaction, and lack integration of speech recognition, AI-controlled image generation, and print functions, limiting accessibility and collaborative experiences.

Method used

An AI-based interactive device, AIKO, integrates speech recognition, text generation, and print functions into a single user-friendly platform, using a cat-shaped housing with embedded microphones, a Raspberry Pi microcontroller, and an integrated printer to convert verbal inputs to printed visual representations.

Benefits of technology

Enhances user engagement and creativity by providing graspable results that can be shared and preserved, promoting a collaborative and accessible experience for all age groups, while ensuring privacy and security through real-time processing and cloud-based APIs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An AI-powered interactive device for storytelling and image generation, consisting of: a cat-shaped main body that houses a variety of internal components of the device; one or more speakers positioned in speaker holes in the main body, the speakers being embedded into the main body through the speaker holes to provide clear audio output for both speech recognition feedback and text-to-speech functions; one or more microphones embedded in speaker holes in the main body, the microphones being configured to ensure optimal audio recording quality and to capture the user's verbal instructions; a three-button arrangement that facilitates user input and enables the device to perform certain functions, consisting of: a left button assembly for starting the narration of a story, comprising a first top cap, a first button, and a first recording hole; a central key assembly configured to confirm transcription accuracy and comprising a second top cap, a second key, and a second receiving hole; and a next button assembly configured to initiate image formation and printing and comprising a third top cap, a third button, and a third receiving hole; wherein the receiving holes are provided on the main body of the device and buttons are integrated therein; a core processing unit with a Raspberry Pi microcontroller configured as follows: Process audio inputs from the dual microphones; Manage real-time speech-to-text conversion using the speech-to-text mechanism. Control text-to-speech playback using the text-to-speech mechanism. Facilitating image generation using an artificial intelligence-based image generation model; and Controlling printing processes; a power system configured to switch on all electrical components of the device, the power system comprising: a rechargeable battery to power the components; and a USB Type-C charging port that facilitates battery charging; an embedded printer system operatively connected to the core processing unit and configured to facilitate printing of the generated images, the embedded printer system comprising: a thermal printer mechanism; a base plate for supporting the printer; and an output slot for dispensing printing materials; a plurality of structural support components integrated to ensure the stability and integrity of the internal components of the device, the structural support components including: a plurality of assembly support systems for component alignment, wherein the assembly support systems are strategically placed to maintain the alignment and stability of critical components of the device; a gasket between the main body and a lid assembly to create an airtight seal and thus ensure optimum functionality and safety during operation of the device; and a safety valve integrated into the lid assembly for regulating the internal pressure of the device through a vent pipe, the vent pipe allowing the release of steam and thus preventing the build-up of excess pressure; and a memory within the core processing unit that is configured to store instructions that, when executed by the Raspberry Pi microcontroller, cause the device to do the following: Capture verbal narrative input when activating the left key layout; convert captured verbal input into text without local audio storage; Playback of converted text via the dual speakers after activating the middle button arrangement; Using the AI-based image generation model, generate corresponding images when the next key arrangement is activated. Manage real-time printing of generated images; and Maintaining operational safety through continuous monitoring of internal conditions.
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF THE INVENTION

[0001] The present disclosure relates to an AI-powered interactive storytelling and image generation device. More specifically, the present invention relates to an interactive storytelling device called AIKO, which significantly advances the field by integrating speech recognition, text generation, image generation, and printing functions into a single, user-friendly device. AIKO uses a multi-component system that includes embedded microphones, an intuitive button interface, a Raspberry Pi microcontroller, and an integrated printer. These components work together to convert verbal narrative prompts into text, generate corresponding images using AI models, and create printed visual representations of the stories. BACKGROUND OF THE INVENTION

[0002] Storytelling is a fundamental method of communication and education that fosters imagination, creativity, and cognitive development in all ages. Traditional narrative media such as oral narratives and written texts effectively convey stories, but often lack interactive and visual components that can enhance user engagement and understanding. While these methods are invaluable, they may fall short of providing the dynamic and immersive experiences that modern educational tools are designed to provide.

[0003] In recent years, technological advances have given rise to interactive storytelling tools that integrate digital elements such as images, sounds, and animations to enrich the storytelling experience. Devices such as tablets, computers, and smartphones allow users to engage with stories in a more dynamic and immersive way. Furthermore, AI-powered applications increasingly support the generation of visual content based on text or voice input, thus enhancing the storytelling process. However, many of these solutions remain limited to digital platforms and require extensive user interaction with screens and interfaces, which can be challenging for young children and limit collaborative storytelling between parents and children.

[0004] Furthermore, existing AI-powered storytelling devices predominantly focus on generating digital content such as animated images or digital illustrations without providing tangible results. This digital-centric approach can limit the tactile and hands-on learning experiences critical for early childhood development. The lack of physical artifacts such as printed images or storybooks can also hinder the ability to share and preserve stories in a tangible format. Furthermore, many current solutions involve complex setups and multiple components, making them less accessible and user-friendly for everyday use.The need for separate devices or peripherals to record audio, generate images, and create printed materials can complicate the storytelling process and hinder the creative and collaborative aspects essential to an engaging experience.

[0005] There is a growing demand for a comprehensive, all-in-one storytelling device that seamlessly integrates speech recognition, AI-driven image generation, and printing capabilities into a single, intuitive platform. Such a device should appeal to users of all ages, especially children and their parents, by providing a user-friendly interface that facilitates collaborative storytelling and creates both digital and physical representations of their narratives. Integrating voice-activated input and instant physical output can significantly improve the accessibility and appeal of storytelling tools, making them more versatile and effective in educational and recreational settings.

[0006] The present invention, AIKO, overcomes these limitations by providing an innovative, interactive storytelling tool that combines voice-guided input, real-time AI processing, and integrated printing capabilities. AIKO enables users to narrate their stories orally, convert those narratives into text using advanced speech recognition, generate corresponding images using AI models such as OpenAI's DALL-E, and create printed visual representations of the stories. This comprehensive approach not only increases user engagement and creativity but also delivers tangible results that can be shared, preserved, and cherished.By bridging the gap between digital storytelling and physical artifacts, AIKO represents a significant advancement in educational and creative technologies, fostering an enriching and collaborative storytelling environment for families and educators alike.

[0007] Additionally, AIKO features user-friendly design elements tailored to both children and adults, ensuring the device is accessible and intuitive for a wide range of users. The integration of cloud-based APIs for speech-to-text and text-to-speech processing ensures accurate and efficient processing of user input without the need for extensive local storage, thus improving user privacy and data security. The integration of a built-in printer allows for instant physical storybook creation, enabling a seamless transition from digital creativity to tangible results. By consolidating multiple functions into a single device, AIKO simplifies the storytelling process, making it more engaging, interactive, and fun for all users. SUMMARY OF THE INVENTION

[0008] This disclosure relates to an AI-powered interactive storytelling and image generation device called AIKO. AIKO is an interactive AI-powered storytelling device housed in a cat-shaped case that transforms spoken narratives into printed storybooks. The device uses speech recognition, text processing, and AI image generation to create visual representations of spoken stories, providing an engaging and tangible storytelling experience for users of all ages.

[0009] The present disclosure is intended to provide an AI-powered interactive device for storytelling and image generation. The device comprises: a cat-shaped main housing configured to house several internal components of the device; one or more speakers positioned in speaker holes in the main housing, wherein the speakers are embedded into the main housing through the speaker holes to provide clear audio output for both speech recognition feedback and text-to-speech functions; one or more microphones embedded in speaker holes in the main housing, wherein the microphones are configured to ensure optimal audio recording quality and to capture the user's verbal prompts;a three-button arrangement that facilitates user input and controls the device to perform certain functions, comprising a left button arrangement configured to start narration of a story and including a first top cap, a first button, and a first recording hole, a middle button arrangement configured to confirm transcription accuracy and including a second top cap, a second button, and a second recording hole, and a next button arrangement configured to start image generation and printing and including a third top cap, a third button, and a third recording hole, wherein the recording holes are provided on the main body of the device and buttons are integrated therein;a core processing unit including a Raspberry Pi microcontroller configured to process audio input from the dual microphones, manage real-time speech-to-text conversion using a speech-to-text mechanism, control text-to-speech playback using a text-to-speech mechanism, enable image generation using an artificial intelligence-based image generation model, and control printing operations; a power supply system configured to power all electrical components of the device, the power supply system comprising: a rechargeable battery for powering the components; and a USB Type-C charging port that facilitates charging the battery;an embedded printer system operatively connected to a core processing unit configured to enable printing of the generated images, the embedded printer system comprising: a thermal printer mechanism; a base plate for supporting the printer; and an output slot for outputting printing materials;a plurality of structural support components integrated to ensure the stability and integrity of the internal components of the device, the structural support components comprising: a plurality of mounting support systems for component alignment, the mounting support systems strategically placed to maintain the alignment and stability of critical components of the device, a gasket between the main body and a lid assembly to create an airtight seal ensuring optimal functionality and safety during operation of the device, and a safety valve integrated into the lid assembly for regulating the internal pressure levels of the device through a vent tube, the vent tube allowing the release of steam, thus preventing overpressurization;and a memory within the core processing unit configured to store instructions that, when executed by the Raspberry Pi microcontroller, cause the device to capture verbal narrative input upon activation of the left key array, convert captured verbal input into text without local audio storage, play converted text through the dual speakers upon activation of the middle key array, generate corresponding images using the AI-based image generation model upon activation of the next key array, manage real-time printing of generated images, and maintain operational reliability by continuously monitoring internal conditions.;

[0010] One objective of the present disclosure is to provide an AI-powered interactive device for storytelling and image generation.

[0011] Another objective of the present disclosure is to provide an intuitive and interactive platform that transforms verbal storytelling into physical, illustrated narratives using advanced AI technologies.

[0012] Another objective of this disclosure is to create an engaging educational tool that encourages creativity and imagination while delivering tangible results from user stories.

[0013] Another objective of the present disclosure is to implement a secure, user-friendly system that preserves privacy while delivering high-quality story visualization through integrated speech recognition and image generation capabilities.

[0014] To further clarify the advantages and features of the present disclosure, a more detailed description of the invention will be given with reference to specific embodiments thereof illustrated in the accompanying drawings. It should be noted that these drawings represent only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained in additional detail and in greater detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE CHARACTERS

[0015] These and other features, aspects, and advantages of the present disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings, in which like characters represent like parts throughout. Fig. 1 shows a block diagram of an AI-based interactive storytelling and image generation device according to an embodiment of the present disclosure; Fig. 2 A, Fig. 2B, Fig. 2C show various views of the proposed device according to an embodiment of the present disclosure; and Fig. 3 illustrates a diagram showing a side cross-sectional view with labeling of the proposed device according to an embodiment of the present disclosure.

[0016] Furthermore, those skilled in the art will appreciate that elements in the drawings are shown for convenience and may not necessarily be drawn to scale. For example, the flowcharts illustrate the method by key steps to enhance understanding of aspects of the present disclosure. Furthermore, with respect to device construction, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only the specific details relevant to understanding embodiments of the present disclosure in order not to clutter the drawings with details that would be readily apparent to those skilled in the art who would benefit from the description herein. DETAILED DESCRIPTION:

[0017] To promote an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and described in specific language. It is to be understood, however, that no limitation upon the scope of the invention is intended thereby, since such changes and further modifications of the illustrated system, and such further applications of the principles of the invention as illustrated therein, are contemplated as would normally occur to one skilled in the art to which the invention pertains.

[0018] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not intended to be limiting thereof.

[0019] References in this specification to "one aspect," "another aspect," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, the occurrences of the phrase "in one embodiment," "in another embodiment," and similar language throughout this specification may or may not all refer to the same embodiment.

[0020] The terms "comprises," "comprising," or other variations thereof are intended to cover non-exclusive inclusion, such that a process or method comprising a list of steps not only includes those steps, but may also include other steps not expressly listed or inherent in such process or method. Likewise, one or more devices or subsystems or elements or structures or components preceded by "comprises...a" does not preclude, without further limitation, the existence of other devices or other subsystems or other elements or other structures or other components or additional devices or additional subsystems or additional elements or additional structures or additional components.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The system, methods, and examples provided herein are for illustrative purposes only and should not be considered limiting.

[0022] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0023] Fig. 1 shows a block diagram of an AI-based device for interactive storytelling and image generation according to an embodiment of the present disclosure.

[0024] With reference to Fig. 1, the device (100) comprises: a cat-shaped main housing (102) configured to house a plurality of internal components of the device (100); one or more speakers (104) positioned in speaker holes in the main housing (102), the speakers (104) being embedded into the main housing (102) through the speaker holes to provide clear audio output for both speech recognition feedback and text-to-speech functions; one or more microphones (106) embedded in speaker holes in the main housing (102), the microphones (106) being configured to provide optimal audio recording quality and to capture the user's verbal prompts;three key assemblies (108) that enable user input and instruct the device to perform certain functions, comprising a left key assembly (108a) configured to start telling a story and comprising a first top cap, a first key, and a first receiving hole, a middle key assembly (108b) configured to confirm transcription accuracy and comprising a second top cap, a second key, and a second receiving hole, and a next key assembly (108c) configured to start image generation and printing and comprising a third top cap, a third key, and a third receiving hole, wherein the receiving holes are provided on the main body of the device and have keys integrated therein;a core processing unit (110) with a Raspberry Pi microcontroller (112) configured to process audio input from the dual microphones, manage real-time speech-to-text conversion using a speech-to-text mechanism, control text-to-speech playback using a text-to-speech mechanism, facilitate image generation using an artificial intelligence-based image generation model, and control printing operations; a power system (114) configured to power all electrical components of the device, the power system comprising: a rechargeable battery (114a) for powering the components; and a USB Type-C charging port (114b) for facilitating charging of the battery;an embedded printer system (116) operatively connected to the core processing unit (110) and configured to facilitate printing of the generated images, the embedded printer system (116) comprising: a thermal printer mechanism; a base plate for supporting the printer; and an output slot for dispensing print materials;a plurality of integrated structural support components (118) to ensure the stability and integrity of the internal components of the device (100), the structural support components comprising a plurality of mounting support systems (118a) for component alignment, the mounting support systems (118a) being strategically placed to maintain the alignment and stability of critical components of the device (100), a gasket (118b) between the main body (102) and a lid assembly (118c) to create an airtight seal that ensures optimal functionality and safety during operation of the device (100), and a safety valve (118d) integrated into the lid assembly (118c) for regulating the internal pressure levels of the device (100) through a vent tube (118e), the vent tube (118e) allowing the release of steam and thereby preventing overpressurization;and a memory (110a) within the core processing unit (110) configured to store instructions that, when executed by the Raspberry Pi microcontroller (112), cause the device (100) to capture verbal narrative input upon activation of the left key array, convert captured verbal input into text without local audio storage, play converted text through the dual speakers upon activation of the middle key array, generate corresponding images using the AI-based image generation model upon activation of the next key array, manage real-time printing of generated images, and maintain operational reliability by continuously monitoring internal conditions.;

[0025] In one embodiment, the microphones and speakers (106 and 104) are configured to: operate simultaneously for real-time audio feedback, capture stereo audio input for improved speech recognition accuracy, and provide balanced audio output during text-to-speech playback, wherein pressing the left button array (108a) triggers microphones (106) to capture the user's prompt, and pressing the middle button array (108b) triggers speakers (104) to play the captured prompt.

[0026] In one embodiment, the left button assembly (108a) is configured to activate audio recording only when the left button remains pressed, to automatically stop audio recording when the left button is released, to provide tactile feedback during activation, and to prevent accidental activation through the design of the top cap; wherein the middle button assembly (108b) is configured to confirm the prompt captured via the left button assembly, to play the converted text upon pressing the middle button, and to allow the user to verify the accuracy of the transcription; and wherein the next button assembly (108c) is configured to initiate the conversion of the validated text into images upon pressing the next button and to trigger the embedded printer system (116) to print the generated image, which is then output via the output slot of the printer system (116).

[0027] In one embodiment, the Raspberry Pi microcontroller (112) is further configured to perform real-time error checking during speech-to-text conversion, maintain a secure connection and communicate with cloud-based APIs, manage memory allocation for image processing tasks, and optimize system performance based on battery status.

[0028] In one embodiment, the core processing unit (110) is further configured to perform all speech data processing in real time using text-to-speech and speech-to-text functions, whereby verbal audio data is not stored locally, and wherein the core processing unit (110) is operatively connected to speakers (104), microphones (106), and three key arrays (108).

[0029] In one embodiment, the embedded printer system (116) uses the Common UNIX Printing System (CUPS) to manage print jobs and generate high-quality print images based on user reports.

[0030] In one embodiment, the safety valve (118d) and the vent pipe (118e) act together to: automatically engage at predetermined pressure thresholds, regulate the internal pressure during prolonged operation, prevent overpressure buildup during printing processes, and maintain optimal operating conditions for electronic components.

[0031] In one embodiment, the power supply system (114) with rechargeable battery (114a) provides portability, and the USB Type-C port (114b) provided on the main body (102) facilitates rapid charging of the battery (114a) and enables seamless integration with various chargers.

[0032] In one embodiment, the structural support components (118) are configured to: provide shock absorption for sensitive electronics; facilitate pressure regulation during operation; enable replacement of modular components; and maintain precise alignment of the printing mechanisms, wherein the mounting support components enable straightforward troubleshooting and component replacement and allow the user easy access to internal components by opening the designated portion of the main body (102) for maintenance and troubleshooting.

[0033] In one embodiment, the lid assembly (118c) comprises the safety valve (118d) connected to the vent pipe (118e), wherein the safety valve (118d) and the vent pipe (118e) together enable pressure regulation so that steam can be released from the vent pipe (118e) by actuating the safety valve (118d), thereby preventing overpressure within the device (100) and ensuring the safety of the user and the device.

[0034] The present invention relates to an AI-powered interactive storytelling and image generation device called AIKO. The AIKO device represents a significant advance in interactive storytelling technology, seamlessly integrating several cutting-edge components into a user-friendly, cat-shaped device. At its core, the system utilizes a Raspberry Pi microcontroller that orchestrates the interaction between various hardware and software components, enabling a smooth transition from verbal input to printed output.

[0035] The device begins operation when a user presses and holds the left button array, which activates the dual microphones embedded in the device's speaker openings. These microphones capture the user's verbal narration, which is immediately processed via the Speech-to-Text API, converting the spoken words to text without storing any audio data locally. This real-time conversion ensures both user privacy and efficient processing. Once the verbal input has been converted to text, the user can verify the accuracy of the transcription by pressing the middle button array. This triggers the device's text-to-speech functionality, using the text-to-speech service to play the transcribed narration through the dual speakers.This verification step ensures that the story has been accurately captured before proceeding with image generation. After confirming the accuracy of the transcription, the user presses the next set of buttons to initiate the image generation process. The device sends the verified text to an artificial intelligence-based image generation model, which generates appropriate visual representations of the narrative. The generated images are then processed and sent to the embedded printer, which produces high-quality physical prints via the output slot, creating a tangible storybook that brings the user's narrative to life. The entire device is powered by a rechargeable battery with USB Type-C charging, ensuring portability and convenience.Multiple safety features, including a pressure regulation system with a safety valve and vent pipe, protect both the user and the device during operation. The structural design, with various support systems and a seal, ensures component alignment and ensures reliable operation over extended periods.

[0036] Fig. illustrate different views of the proposed device according to an embodiment of the present disclosure.

[0037] In Fig. 2A, (a), (b) and (c) respectively represent the top view, the bottom view and the perspective view of the proposed device. In Fig. 2B, (d) and (e) respectively represent the left and right side views of the proposed device. In Fig. 2C, (f) and (g) represent the front and back of the proposed device, respectively.

[0038] In Fig. A novel interactive storytelling device called AIKO is presented, significantly advancing the field by integrating speech recognition, text generation, image generation, and printing capabilities into a single, user-friendly device. AIKO uses a multi-component system that includes embedded microphones, an intuitive button interface, a Raspberry Pi microcontroller, and an integrated printer. These components work together to convert verbal story prompts into text, generate corresponding images using AI models, and create printed visual representations of the stories. This comprehensive approach not only promotes a more immersive and engaging storytelling experience but also provides users with a tangible way to visualize and retain their narratives.The AIKO device's unique combination of voice-activated input, real-time processing, and instant output through printed images represents a significant advance in educational technology, promoting interactive learning and creative expression.

[0039] Fig. 3 illustrates a diagram showing a side cross-sectional view with labeling of the proposed device according to an embodiment of the present disclosure.

[0040] According to Fig. 3, the device comprises a variety of main and secondary components, each of which and its functionality are described in detail below.

[0041] Main Body: AIKO's main structure is a cat-shaped main body designated [1]. This ergonomic and aesthetically pleasing design is tailored for easy handling by children and parents, making the device accessible and inviting. The main body [1] houses all essential components, ensuring a compact and uncluttered layout that allows for smooth operation and maintenance.

[0042] Speakers and speaker holes: The main body incorporates two speakers, labeled [1a1] and [1a2]. These speakers are housed in dedicated speaker holes, labeled [1b1] and [1b2], respectively. The strategic placement of these speakers ensures clear audio output for both voice recognition feedback and text-to-speech functions, thus enhancing user interaction and device responsiveness.

[0043] Button Layouts: AIKO features a sophisticated button layout system with three distinct buttons, each consisting of a top cap, a button, and a corresponding hole in the main body. These button layouts facilitate user interaction by controlling various functions of the device. 1. Left button assembly: a. Top cap ([1c1]): Serves as a protective and aesthetic cover for the left button. It is designed to fit securely over the button, preventing accidental presses and enhancing the overall appearance of the device. b. Button ([1c2]): Acts as the primary input interface for starting narration. When pressed, the device's microphones are activated to record verbal prompts. c. Receiving hole ([1c3]): Designed to precisely accommodate the top cap ([1c1]) and knob ([1c2]). This design ensures that the knob assembly stays securely in place while allowing for smooth and responsive operation.

[0044] Left button layout operation: Users begin narrating a story by holding down the left button layout ([1c1], [1c2], [1c3]) while speaking their narrative. This action activates the embedded microphones ([1a1], [1a2]) to record the user's verbal input. The spoken words are then processed in real time using Google's Speech-to-Text API, which converts the audio input to text without storing the verbal data locally, thus ensuring user privacy. 2. Mounting the middle button: a. Top cap ([1d1]): Covers the center button, providing protection and maintaining the aesthetic consistency of the device. b. Button ([1d2]): Serves as a confirmation interface. When pressed, this button verifies the accuracy of the transcribed text. c. Mounting hole ([1d3]): Secures the top cap ([1d1]) and knob ([1d2]) to the main body, ensuring stable and reliable operation.

[0045] Middle button operation: After the user has finished narrating the story, they press the middle button ([1d1], [1d2], [1d3]) to confirm the captured prompt. This action triggers the device's speakers ([1a1], [1a2]) to play the converted text using Google Text-to-Speech (gTTS) technology. This playback allows users to verify the accuracy of the transcription and ensure that the AI has correctly interpreted their narrative. 3. Next button assembly: a. Top cap ([1e1]): Protects the next button, prevents accidental activation, and contributes to the overall design of the device. b. Button ([1e2]): Starts image generation and printing based on the validated text. c. Retaining hole ([1e3]): Holds the top cap ([1e1]) and knob ([1e2]) in place in the main body, enabling reliable and consistent functionality.

[0046] Next button operation: If the prompt is confirmed as correct, users press the Next button ([1e1], [1e2], [1e3]) to initiate the conversion of the validated text into images. This process leverages OpenAI's DALL-E model via an open API to generate corresponding images based on the transcribed story. The generated images are then printed by the embedded printer (labeled [5]) and output through an output slot ([1f]), creating a tangible visual representation of the narrated story.

[0047] Microphones and audio recording: The microphones ([1a1], [1a2]) are strategically placed in the speaker openings ([1b1], [1b2]) to ensure optimal audio recording quality. These microphones are responsible for capturing the user's verbal prompts, which are then transmitted to the Raspberry Pi microcontroller ([3]) for processing. The dual-microphone configuration improves the device's ability to accurately record audio in diverse environments and minimize interference from background noise.

[0048] Core processing unit and power supply: At the heart of AIKO is the Raspberry Pi microcontroller ([3]), which manages all data processing tasks. This includes converting speech input to text, generating images from text, and controlling printing operations. The Raspberry Pi is a versatile and powerful component that ensures efficient and reliable execution of the device's functions. The device is powered by a rechargeable battery ([4]), ensuring portability and convenience. The battery can be conveniently charged via a USB Type-C port ([2a]), allowing users to easily charge AIKO with standard charging cables. The inclusion of a USB Type-C port ensures compatibility with a wide range of chargers and improves the overall user experience.

[0049] Printer and printing system: AIKO is equipped with an integrated printer (labeled [5]) that facilitates printing the generated images. Located within the main body, the printer receives image data from the Raspberry Pi microcontroller ([3]) and creates physical copies that are output through an output slot ([1f]). This integrated printing system allows users to obtain tangible visual representations of their narrated stories and enhances the overall storytelling experience by combining digital creativity with physical output.

[0050] Structural components: To ensure the stability and integrity of the internal components, AIKO integrates several structural elements: 1. Base plate ([2b]): Serves as a base plate for the printer and ensures that it remains securely positioned in the main body. 2. Mounting support systems ([2c], [3a], [4b], [4c]): These mounts are strategically placed to maintain the alignment and stability of critical components such as the Raspberry Pi microcontroller ([3]) and the battery ([4]). The mounting support systems facilitate maintenance and component replacement, ensuring that the device remains functional and reliable over extended periods.

[0051] In one embodiment, the device operates with user input. AIKO's intuitive interface and integrated features provide a seamless and engaging storytelling experience. The streamlined workflow involves three main steps. First, to tell the story, the user presses and holds the left button array ([1c1], [1c2], [1c3]) to activate the microphones ([1a1], [1a2]). The user then narrates the story verbally, and the embedded microphones capture the audio, which is converted to text in real time using the Speech-to-Text API. Next, for instant confirmation, the user presses the middle button array ([1d1], [1d2], [1d3]) to play the transcribed text through the speakers ([1a1], [1a2]). The device uses text-to-speech technology to read the converted text aloud, allowing the user to verify its accuracy.Finally, the user presses the next key combination ([1e1], [1e2], [1e3]) to generate and print images. The validated text is sent to an artificial intelligence-based image generation model to create corresponding images. These images are then printed by the embedded printer ([5]) and output through the output slot ([1f]), creating a tangible visual representation of the story being told.

[0052] The proposed device incorporates a variety of security and privacy features, with the device being designed with user security and data privacy in mind. The device does not store any verbal audio data locally; all speech data is processed in real time via Google's Speech-to-Text API and Text-to-Speech (gTTS) services. This approach ensures that user data remains secure and private, as no audio recordings are stored on the device. Furthermore, the integrated safety valve ([6]) and vent tube ([2e]) mechanisms prevent accidental overpressurization, thus improving the overall security profile of the device.

[0053] The design of the proposed AIKO device incorporates easily accessible components to facilitate maintenance and troubleshooting. These components include: a USB Type-C port ([2a]), located on the main body. This port allows for easy charging of the battery ([4]) using standard cables; an output slot ([1f]), positioned for easy retrieval of printed images, allowing users to conveniently and easily retrieve their printed stories; and assembly support systems ([2c], [3a], [4b], [4c]), which support supports for straightforward troubleshooting and component replacement, ensuring that AIKO remains functional and reliable over extended periods of time. Users can easily access the internal components for maintenance or troubleshooting purposes by opening the designated parts of the main body, marked [2].

[0054] The present invention presents an innovative interactive storytelling device called AIKO, which represents a significant advance in the field by seamlessly integrating speech recognition, text generation, image generation, and printing capabilities into a single, user-friendly device. AIKO operates via a multi-component system consisting of embedded microphones, an intuitive button interface, a Raspberry Pi microcontroller, and an integrated printer. These components work together to convert verbal storytelling prompts into text, generate relevant images using AI models, and create printed visual representations of the narratives. This holistic approach enhances the storytelling experience by making it more immersive and engaging, while also providing users with a tangible way to visualize and retain their narratives.AIKO's unique combination of voice-activated input, real-time processing, and instant output through printed images represents a major advance in educational technology, promoting interactive learning and creative expression.

[0055] The device features a cat-shaped casing that is easily accessible and user-friendly for people of all ages. Microphones are embedded within the casing to effectively capture users' verbal input. AIKO also features a button system with three distinct buttons. The left button assembly, consisting of a top cap, a button, and a recording hole, is designed to initiate the narration of a story when pressed. The middle button assembly, which has similar components, is used to confirm the accuracy of the transcribed text. The next button assembly, constructed in the same way, facilitates the initiation of image generation and printing processes. Furthermore, speakers housed within the casing enable audio playback of the transcribed text and utilize text-to-speech technology for an enhanced user experience.

[0056] At the heart of AIKO is a Raspberry Pi microcontroller, which is responsible for processing verbal input, converting speech to text, generating images from the text, and managing printing operations. The device is powered by a rechargeable battery paired with a USB Type-C port, providing both power and convenient charging. AIKO features an embedded printer within its casing, capable of generating and printing images, which are then output through an output slot for easy user access. Structural support components, including a baseplate and mounting support systems, ensure the stability and alignment of all internal components, enhancing the device's durability and reliability.

[0057] AIKO's main body is specially designed with speaker holes to accommodate the speakers, ensuring optimal audio output for clear user interaction. The left button array acts as an activation mechanism for the embedded microphones and captures verbal narration prompts when the button is pressed and held. The middle button array plays a critical role in the user verification process by triggering the playback of transcribed text through the speakers using Google Text-to-Speech technology. The next button array facilitates the conversion of validated text into images via an AI-based image generation model accessed via an open API, before the generated images are printed using the embedded printer.The Raspberry Pi microcontroller is carefully programmed to handle real-time speech-to-text conversion, text-to-speech playback, image generation, and printing tasks, effectively coordinating all core functions. The embedded printer operates with the Common UNIX Printing System (CUPS), which manages print jobs and delivers high-quality printed images based on the user's narrative input.

[0058] Safety and ease of use are key considerations in AIKO's design. To improve the unit's operational efficiency and safety, a gasket is installed between the main body and a lid assembly, ensuring an airtight seal. The unit features a vent tube integrated into the lid assembly, allowing for controlled steam release, preventing overpressure and ensuring operational safety. Additionally, a safety valve in the lid autonomously regulates internal pressure levels to further enhance the safety profile. The button assembly system features protective top caps and securely mounted receiving holes to prevent accidental pressing, making AIKO safe for use by children and adults alike.AIKO's structural support components play a critical role in maintaining the stability and alignment of the Raspberry Pi microcontroller, rechargeable battery, and embedded printer, ensuring consistent and reliable performance. The device follows a user interaction workflow that includes pressing and holding the left button to capture verbal prompts, pressing the middle button to review the transcribed text through audio playback, and finally pressing the next button to generate and print images, providing a seamless and engaging storytelling experience.

[0059] The rechargeable battery allows for mobile use, enabling storytelling activities without the need for a fixed power source. The integrated USB Type-C port ensures fast charging capabilities and seamless integration with various chargers, improving both accessibility and convenience. AIKO also features mounting support systems that enable easy maintenance and component replacement, thus extending the long-term functionality and reliability of the device.

[0060] With the ability to print high-resolution images, AIKO provides users with clear and detailed visual representations of their narrated stories. Its speech recognition and processing capabilities rely on cloud-based APIs, specifically Google's Speech-to-Text API and Google Text-to-Speech Services. These ensure accurate and efficient conversion of verbal input while preserving user privacy and data security, as no local data storage is required. Image generation is performed using OpenAI's DALL-E model via an open API, producing customized, high-quality images that accurately reflect the content of the narrated stories. The printed images are output via a specially designed output slot, allowing users to easily retrieve their visual representations.This feature enhances the entire storytelling experience by delivering instant and tangible output, fostering creativity and engagement. AIKO's innovative combination of voice input, AI-powered text and image generation, and instant printing makes it a groundbreaking educational and creative tool that redefines interactive storytelling.

[0061] The device's ability to create printed storybooks enhances its educational value, allowing users to visualize and preserve their stories in a tangible format. This feature not only promotes memory and cognitive development but also provides families and educators with an opportunity to share and celebrate creative storytelling. By transforming verbal narratives into printed images, AIKO supports diverse learning styles and encourages interactive and hands-on educational activities.

[0062] The drawings and the foregoing description provide examples of embodiments. Those skilled in the art will recognize that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be separated into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of the processes described herein may be changed and are not limited to the manner described herein. Furthermore, the actions of any flowchart need not be implemented in the order shown; nor do all actions necessarily need to be performed. Also, those actions that are not dependent on other actions may be performed in parallel with the other actions. The scope of the embodiments is in no way limited by these specific examples.Numerous variations, whether explicitly stated in the specification or not, such as differences in structure, dimensions, and use of materials, are possible. The scope of the embodiments is at least as broad as indicated in the following claims.

[0063] Advantages, other benefits, and solutions to problems have been described above with respect to specific embodiments. However, the advantages, benefits, solutions to problems, and any components that may cause an advantage or solution to occur or become more apparent are not to be construed as a critical, required, or essential feature or component of any or all of the claims. REFERENCES 100 The device comprises: A cat-shaped main housing. 102 Main housing 104 One or more speakers 106 One or more microphones 108 three-button assemblies 108a Left button assembly 108b Middle key unit 108c Next button arrangement 110 Core Processing Unit 110a storage 112 Raspberry Pi microcontrollers 114 Power grid 114a Rechargeable Battery 114b USB Type-C charging port 116 Embedded printer system 118 Variety of structural support components 118a Variety of assembly support systems 118b Seal 118c cover assembly 118d safety valve 118e vent pipe

Claims

[1] An AI-powered interactive device for storytelling and image generation, consisting of: a cat-shaped main body that houses a variety of internal components of the device; one or more speakers positioned in speaker holes in the main body, the speakers being embedded into the main body through the speaker holes to provide clear audio output for both speech recognition feedback and text-to-speech functions; one or more microphones embedded in speaker holes in the main body, the microphones being configured to ensure optimal audio recording quality and to capture the user's verbal instructions; a three-button arrangement that facilitates user input and enables the device to perform certain functions, consisting of: a left button assembly for starting the narration of a story, comprising a first top cap, a first button, and a first recording hole; a central key assembly configured to confirm transcription accuracy and comprising a second top cap, a second key, and a second receiving hole; and a next button assembly configured to initiate image formation and printing and comprising a third top cap, a third button, and a third receiving hole; wherein the receiving holes are provided on the main body of the device and buttons are integrated therein; a core processing unit with a Raspberry Pi microcontroller configured as follows: Process audio inputs from the dual microphones; Manage real-time speech-to-text conversion using the speech-to-text mechanism. Control text-to-speech playback using the text-to-speech mechanism. Facilitating image generation using an artificial intelligence-based image generation model; and Controlling printing processes; a power system configured to switch on all electrical components of the device, the power system comprising: a rechargeable battery to power the components; and a USB Type-C charging port that facilitates battery charging; an embedded printer system operatively connected to the core processing unit and configured to facilitate printing of the generated images, the embedded printer system comprising: a thermal printer mechanism; a base plate for supporting the printer; and an output slot for dispensing printing materials; a plurality of structural support components integrated to ensure the stability and integrity of the internal components of the device, the structural support components including: a plurality of assembly support systems for component alignment, wherein the assembly support systems are strategically placed to maintain the alignment and stability of critical components of the device; a gasket between the main body and a lid assembly to create an airtight seal and thus ensure optimum functionality and safety during operation of the device; and a safety valve integrated into the lid assembly for regulating the internal pressure of the device through a vent pipe, the vent pipe allowing steam to escape and thus preventing overpressure build-up; and a memory within the core processing unit that is configured to store instructions that, when executed by the Raspberry Pi microcontroller, cause the device to do the following: Capture verbal narrative input when activating the left key layout; convert captured verbal input into text without local audio storage; Playback of converted text via the dual speakers after activating the middle button arrangement; Using the AI-based image generation model, generate corresponding images when the next key arrangement is activated. Manage real-time printing of generated images; and Maintaining operational safety through continuous monitoring of internal conditions. [2] The device of claim 1, wherein the microphones and speakers are configured to: operate simultaneously for real-time audio feedback, capture stereo audio input for improved speech recognition accuracy, and provide balanced audio output during text-to-speech playback, wherein pressing the left button array triggers microphones to capture the user's prompt, and pressing the middle button array triggers speakers to play the captured prompt. [3] The device of claim 1, wherein the left key assembly is configured to activate audio recording only when the left key remains pressed, to automatically stop audio recording when the left key is released, to provide tactile feedback during activation, and to prevent accidental activation through the design of the top cap; wherein the middle key assembly is configured to validate the prompt captured via the left key assembly, to play the converted text upon pressing the middle key, and to allow the user to verify the accuracy of the transcription; and wherein the next key assembly is configured to initiate the conversion of the validated text to images upon pressing the next key and to trigger the embedded printer system to print the generated image, which is then output via the output slot of the printer system. [4] The device of claim 1, wherein the Raspberry Pi microcontroller is additionally configured to perform real-time error checking during speech-to-text conversion, maintain a secure connection and communicate with cloud-based APIs, manage memory allocation for image processing tasks, and optimize system performance based on battery status. [5] The device of claim 1, wherein the core processing unit is further configured to perform all speech data processing in real time using text-to-speech and speech-to-text functions, whereby verbal audio data is not stored locally, and wherein the core processing unit is operatively connected to speakers, microphones, and a three-button assembly. [6] The apparatus of claim 1, wherein the embedded printer system uses the Common UNIX Printing System (CUPS) to manage print jobs and generate high-quality print images based on user reports. [7] The device according to claim 1, wherein the safety valve and the vent pipe together have the following effect: they are automatically activated at predetermined pressure thresholds, regulate the internal pressure during prolonged operation, prevent overpressure build-up during printing operations and maintain optimal operating conditions for electronic components. [8] The device according to claim 1, wherein the rechargeable battery power system provides portability, and the USB-C port provided on the main body facilitates fast battery charging and enables seamless integration with various chargers. [9] The device of claim 1, wherein the structural support components are configured to provide shock absorption for sensitive electronics; facilitate pressure regulation during operation; enable modular component replacement; and maintain precise alignment of the printing mechanisms, wherein the mounting support components enable straightforward troubleshooting and component replacement and allow the user easy access to the internal components by opening the designated portion of the main body for maintenance and troubleshooting. [10] The appliance of claim 1, wherein the lid assembly includes the safety valve communicating with the vent pipe, the safety valve and the vent pipe together enabling pressure regulation and, upon actuation of the safety valve, allowing steam to be released from the vent pipe, thereby preventing overpressure in the appliance and ensuring the safety of the user and the appliance.