Audio reading device and method introducing learning mode engine and supporting printed picture attachment
By generating printable images and audio content through a mobile application, and combining a learning mode engine and a safety control module, this technology solves the problems of high hardware complexity and insufficient security in existing audiobook technology, achieving flexible multi-mode interaction and safe and controllable use for children.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-03-31
AI Technical Summary
Existing audiobook technology for early childhood education suffers from problems such as high hardware complexity, increased costs, lack of flexibility, difficulty in supporting user-defined content updates, and lack of security controls.
Learning content can be generated or selected through a mobile application, automatically generating printable images and audio content. Multi-mode interaction is achieved through a matrix key logic layer, combined with a learning mode engine and a safety control module to ensure the safety and controllability of children's use.
It enables synchronized updates of printed and audio content, supports multi-mode interaction, reduces hardware complexity, and enhances safety and controllability during children's use through safety controls.
Smart Images

Figure CN121768253A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of early childhood education, audiobooks, human-computer interaction, and artificial intelligence voice technology. Specifically, it relates to an audiobook device and method combining a printed book and an electronic voice generator, and more particularly to an audiobook device and method incorporating a learning mode engine. This device achieves multi-mode interaction through a configurable matrix button logic layer, and introduces a microphone and artificial intelligence voice interaction in a dialogue learning mode. Simultaneously, safety controls and parental control mechanisms ensure the safety and controllability of children's use. The safety controls incorporate permission judgments and policy controls during voice acquisition, voice interaction triggering, and audio output to technically constrain children's interactive behavior. Background Technology
[0002] Existing audiobook technology has proposed multiple implementation methods. For example: Camera recognition: Captures paper pages or images with a camera, identifies features, and matches them with corresponding audio content; Voice-based interactive software: This method uses voice recognition to answer user questions or commands and then plays corresponding explanations. However, this method relies on user voice interaction, and in the early stages of childhood education, children often lack complete language expression skills and cannot achieve effective operation through dialogue. NFC / RFID: NFC tags or other RFID components are embedded in paper books to identify pages and play corresponding audio by reading the tags; Page number recognition: Determines the current page by detecting page turning actions or reading printed codes, and matches the audio playback accordingly; User-initiated input type: Users enter page numbers or content keywords in the application or terminal to retrieve and play the corresponding audio files.
[0003] While the above solutions have achieved a certain degree of connection between printed content and voice playback, they still have the following shortcomings: they rely on cameras or sensors, resulting in high hardware complexity and increased costs; users can only use the device's preset pages or recognition methods, lacking flexibility; it is difficult to support users to customize printed content and update the corresponding voice content at the same time; and when voice interaction is introduced, there is a lack of safety controls and parental supervision mechanisms for children's use scenarios.
[0004] With the widespread adoption of AI-powered speech synthesis, mobile applications, and home printing devices, users desire the ability to flexibly generate or select printed content based on their learning and entertainment needs, automatically generating corresponding audio content, while ensuring child safety and providing a richer, more controllable interactive learning experience. Therefore, current technology does not yet offer an audiobook solution that maintains a simple hardware structure while supporting flexible content updates, multiple learning modes, and safe, controllable voice interaction. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention provides an audiobook device and method. The overall technical solution involves using a mobile terminal application as a unified entry point to generate or select learning content, automatically generating printable image content and audio content, and establishing a direct correspondence between the printed content and matrix buttons through user printing and pasting. Based on this, a learning mode engine and a configurable matrix button logic layer are introduced, enabling the same set of physical buttons to have different interactive semantics in different learning modes. In the dialogue learning mode, a microphone and artificial intelligence voice interaction are also introduced, and safety controls and parental management mechanisms are used to constrain voice acquisition, dialogue content, and usage permissions, thereby ensuring the safety and controllability of children's use. The safety control achieves proactive control of interactive behavior by performing policy judgments before learning mode switching, voice acquisition initiation, and voice output.
[0006] Specifically, the mobile application serves as the core entry point for content management and strategy configuration. Users select or generate learning content through this application. After user interaction, the application automatically generates two parts of information: first, a printable image for paper output, which includes alignment markers, page numbers, page switching button markers, and content button markers; second, audio content corresponding to the printable image. After printing the image, the user affixes it to the designated printable image placement position on the paper book or cover page as prompted by the application, thus establishing a spatial correspondence between the printable image and the matrix buttons.
[0007] The audio content is transmitted to and stored in the audiobook device via wireless communication. When a child presses the corresponding matrix button according to the printed image prompt during reading, the audiobook device converts the physical button trigger signal into a logical identifier through a logic mapping module. The learning mode engine then parses the logical identifier according to the currently active learning mode to trigger the corresponding audio playback or interactive behavior. Before triggering the audio playback or interactive behavior, the security control module verifies the current mode and permission status.
[0008] In non-dialogue learning mode, button presses are used to recall and play pre-stored voice content. In dialogue learning mode, button presses act as voice acquisition commands, activating the microphone to acquire the user's voice when permitted by the safety control module. The acquired voice data is then uploaded to the mobile terminal application, where the application generates a voice response based on artificial intelligence and plays it back. The safety control module controls the start and stop of voice acquisition, dialogue permissions, and content filtering. These policies can be configured by parents through the mobile terminal application. The safety control module performs voice acquisition start and stop control on the device side and content review and policy configuration on the mobile terminal side, working together to achieve safe management of voice interaction.
[0009] Compared with existing technologies, the present invention has at least the following beneficial effects: it can realize the synchronous updating of paper content and audio content; it supports the same physical button to realize different interactive semantics in different learning modes; it can realize multi-mode interaction and dialogue learning without relying on complex recognition hardware; and through safety control and parental management mechanisms, it can effectively improve the safety and controllability of children's use. Attached Figure Description
[0010] Figure 1 is a block diagram of the overall structure of the audio reading device of the present invention.
[0011] Figure 2 is a schematic diagram of the structure of the sound-generating device of the present invention.
[0012] Figure 3 is a flowchart of the audiobook content update method and dialogue of the present invention.
[0013] Figure 4 is a schematic diagram of the physical structure of the paper book or book cover, the sound-generating device, the matrix buttons, and the location where the printed image is attached according to the present invention. Detailed Implementation
[0014] As shown in Figure 1, the audio reading device of the present invention includes: a paper book or book cover 140, a printed image attachment area 141 disposed on the paper book or book cover, a sound-generating device 130 embedded inside the book cover, and a mobile terminal application 110 running on a mobile terminal.
[0015] The sound-generating device 130 includes a communication module 131, a storage module 132, a logic mapping module 133, a learning mode engine 134, a security control module 135, a matrix keypad 136, a speaker 137, a microphone module 138, and a power module 139.
[0016] The sound-generating device 130 is connected to a mobile terminal application 110 via a communication module 131 and a wireless communication 120 to transmit and update voice content and configuration data. The mobile terminal application 110 is used to select or generate printed and voice content and includes a security control module 112 for checking the printed and voice content. The security control module 112 in the mobile terminal application is used to configure security policies and distribute them to the sound-generating device. The security control module 135 in the sound-generating device performs local control according to the policies.
[0017] The printed image attachment area 141 on the paper book or book cover 140 is used to attach image content generated and printed by the mobile terminal application 110. The printed images correspond one-to-one with the matrix buttons 136 in spatial position. When the user presses the matrix buttons 136, the sound device 130 triggers the playback of the corresponding voice content according to the correspondence.
[0018] As shown in Figure 2, the sound-generating device 130 includes a communication module 131, a storage module 132, a logic mapping module 133, a learning mode engine 134, a security control module 135, a matrix keypad 136, a speaker 137, a microphone module 138, and a power module 139. The communication module 131 receives voice content and keypad logic configuration data from the mobile terminal application 110; the storage module 132 stores the voice content and configuration data; the logic mapping module 133 maps the physical trigger signals of the matrix keypad 136 to logical identifiers; the learning mode engine 134 parses the logical identifiers according to the current learning mode; and the security control module 135 performs permission judgment and control before triggering audio playback or voice acquisition based on the parsing results.
[0019] As shown in Figure 3, the method of the present invention includes the following steps: S101: Users can select learning content, generate custom content, or have the learning content automatically generated by the application's AI module in the mobile terminal application; S102: The mobile terminal application generates printable image files and audio files corresponding to the learning content; S103: The mobile terminal application sends the voice file and its mapping relationship with the matrix keypad to the sound-generating device; S104: The sound-generating device receives and stores voice files and mapping relationships; S105: The user attaches the printed image to the designated location on the book or book cover, and aligns it with the attachment mark on the printed image and the attachment area on the book cover to establish the correct content and key mapping relationship. S106: When the user presses the matrix button, the sound device will play the corresponding voice content or enter the dialogue learning mode after the safety control module verifies the information. S107: In dialogue learning mode, the voice-generating device collects the user's voice when permitted by the safety control module, and generates a response through the AI module of the mobile terminal application; S108: The voice-generating device plays AI-generated voice responses to complete voice interaction learning.
[0020] Through the above embodiments, this invention, while maintaining the advantages of the original printed image-attached audiobook structure, introduces a learning mode engine, multi-mode interaction, and safe and controllable artificial intelligence voice interaction capabilities, making it suitable for applications such as early childhood education, audiobooks, and interactive learning.
Claims
1. An audiobook device, characterized in that, include: A communication module for wireless communication with mobile terminal applications; The storage module is used to store voice content and key logic configuration data; Matrix buttons are used to receive key presses from users. A logic mapping module is used to map the physical trigger signals of the matrix keys into logical identifiers; The learning mode engine is used to parse the logical identifier according to the currently active learning mode in order to generate the corresponding interactive behavior request. The security control module is located between the learning mode engine and the execution module, and is used to perform permission verification and policy judgment on the interactive behavior request before the interactive behavior request triggers audio playback or voice acquisition. A speaker, used to output voice content when permitted by the safety control module; A microphone module for capturing user voice when permitted by the security control module.
2. An audiobook system, characterized in that, include: Mobile terminal applications, and audiobook devices connected to them. The mobile terminal application includes a first security control module, which is used to configure voice acquisition permissions, dialogue mode permissions, or content filtering policies. The audio reading device includes a second security control module, which is used to locally control the start and stop of voice acquisition and audio output according to the security policy issued by the mobile terminal application.
3. An audiobook device, characterized in that, Including a security control module, The security control module is configured as follows: Before any voice playback or voice acquisition action is executed, the current device status, learning mode status, or user permission status is verified, and the corresponding action is allowed to be executed if the verification is successful.
4. An audiobook device, characterized in that, The device includes a learning mode engine and a security control module. The learning mode engine is used to parse key logic identifiers to generate interaction behavior requests. The security control module is used to restrict or allow the interactive behavior request based on the current learning mode before the interactive behavior request is executed.
5. An interactive control method for audiobooks, characterized in that, include: Obtain the user's matrix key press trigger operations; Map the key-triggered operation to a logical identifier; Based on the current learning mode, the logical identifier is parsed to generate an interactive behavior request; Before executing the interactive behavior request, the security control module performs permission verification on the interactive behavior request; If the verification passes, perform voice playback or voice capture.
6. A voice interaction control method, characterized in that, When entering the dialogue learning mode, the microphone is activated to collect user voice only if the security control module allows it, and voice collection stops if the security control module prohibits it.
7. An audiobook system, characterized in that, The mobile application is used to generate printable image content and corresponding audio content, and to configure security policies; The audio reading device is used to control the playback of the audio content or the audio acquisition behavior according to the security policy.
8. An audiobook device, characterized in that, The security control module is configured to intercept and control voice acquisition requests or voice playback requests locally, without relying on a real-time network connection.
9. An audiobook device suitable for children's use, characterized in that, The device includes a security control module for restricting the activation conditions of the dialogue learning mode to prevent voice interaction from starting when preset conditions are not met.
10. An audiobook device, characterized in that, The device includes an execution module for performing user interaction actions, and a security control module for making control judgments before the execution module takes action.