Picture Book Projector Story Machine Q&A Interactive Method and Device

By acquiring audio and page data from the picture book projector story machine, performing speech recognition and age determination, matching it with an age-appropriate developmental framework, and generating personalized responses, the problem of monotonous interactive design in the picture book projector story machine is solved, enhancing the scientific growth and educational value for children.

CN122489744APending Publication Date: 2026-07-31ANHUI SHENGYUN INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI SHENGYUN INTELLIGENT TECH CO LTD
Filing Date
2026-07-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing picture book projection story machines have simplistic interactive designs that fail to meet children's needs for exploratory and engaging learning. Their question-and-answer response mechanisms are outdated, they cannot update content in real time, and they lack an overall educational plan.

Method used

By acquiring audio data of questions and picture book page data, speech recognition and age range determination are performed to match an age-appropriate developmental framework. AI is then used to generate personalized response content, and facial expression recognition and auxiliary images are combined to enhance the interactive experience.

Benefits of technology

It enables personalized interaction of the picture book projection story machine, enhances children's scientific growth and educational value, and strengthens the interactive experience and fun.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489744A_ABST
    Figure CN122489744A_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of picture book projection technology, specifically a question-and-answer interaction method and device for a picture book projection story machine. The data processing steps include: acquiring question audio data and picture book page data; performing speech recognition on the question audio data to obtain question text; determining a preset age range for the question audio data to obtain a first age range to which the questioner belongs; matching an age-appropriate growth framework from a preset growth framework library based on the first age range; calling AI to match a developmental direction from the age-appropriate growth framework based on the picture book page data and the question text to obtain a target developmental direction; and calling AI to generate a first response content that conforms to the first age range and the target developmental direction based on the picture book page data and the question text. This solution can enhance the interactive experience between children and the picture book projection story machine and promote children's scientific development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of picture book projection technology, and in particular to a question-and-answer interactive method and device for a picture book projection story machine. Background Technology

[0002] This section is intended to provide background or context for the embodiments disclosed herein. The description herein is not intended to imply that it is prior art simply because it is included in this section.

[0003] As a mainstream device for children's early education, the picture book projection story machine is based on the principles of diffuse reflection imaging and infrared safety monitoring. It brings static picture books to life through film projection and audio synchronization, emphasizing eye protection and creating an immersive story experience for children through visual images.

[0004] However, current picture book projector story machines have significant shortcomings in their interactive design, failing to meet the core needs of children's exploratory learning, fun learning, and all-round development. The main reasons include: 1. The interaction methods are singular and fixed, generally only supporting basic operations such as "play / pause / switch pages"; 2. The question-and-answer response mechanism is lagging behind, relying heavily on the device's local preset knowledge base and failing to update the question-and-answer content in real time; 3. The product design lacks an overall educational plan, failing to fully realize the enlightenment and educational value of picture book projector story machines. Summary of the Invention

[0005] Therefore, it is necessary to provide a question-and-answer interactive method and device for a picture book projection story machine that can enhance children's interactive experience and promote children's scientific growth, in order to address one of the above-mentioned technical problems.

[0006] Firstly, this disclosure provides a question-and-answer interactive method for a picture book projection story machine. The method includes:

[0007] Obtain audio data of the questions and page data of the picture book;

[0008] The audio data of the question is subjected to speech recognition to obtain the text of the question;

[0009] The question audio data is subjected to a preset age range determination to obtain the first age range to which the questioner belongs;

[0010] Based on the first age range, an age-appropriate growth framework is matched from a preset growth framework library; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction.

[0011] The AI ​​is invoked to match the development direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target development direction.

[0012] The AI ​​is invoked to generate a first response that conforms to the first age range and the target educational direction, based on the picture book page data and the question text.

[0013] Optionally, the picture book page data includes the current picture book page data and the picture book context data; the preset age range includes 0-3 years, 3-6 years, and 6 years and above; the appropriate growth framework for 0-3 years includes at least the following development directions: basic emotional perception ability, daily scene description ability, and the ability to observe and perceive things; the appropriate growth framework for 3-6 years includes at least the following development directions: emotional regulation awareness, love and companionship awareness, proactive expression ability, healthy living habits, safety awareness, imagination, and memory; the appropriate growth framework for 6 years and above includes at least the following development directions: frustration and emotional regulation awareness, care and understanding awareness, complex expression ability, time management awareness, self-learning awareness, and analytical and creative ability.

[0014] Optionally, before acquiring the audio data of the question, the method further includes:

[0015] Receive a trigger signal for question-and-answer mode; the trigger signal includes a key signal or a voice signal;

[0016] Based on the trigger signal, a preset question-and-answer prompt will be played.

[0017] Optionally, the method further includes:

[0018] Based on the content of the first response, match or generate an auxiliary screen;

[0019] The first response is played aloud, and the auxiliary screen is projected simultaneously.

[0020] Optionally, before generating the first response content, the method further includes:

[0021] The AI ​​was used to analyze the text of the question to obtain a second age range.

[0022] The intersection of the first age range and the second age range is used to obtain the third age range;

[0023] Update the first age range using the third age range.

[0024] Optionally, the method further includes:

[0025] Obtain the questioner's facial image;

[0026] The facial image is used to perform expression recognition to obtain the questioner's expression;

[0027] Determine whether the questioner's expression belongs to the preset expressions that need to be repaired;

[0028] When the questioner's expression is a preset expression to be repaired, the AI ​​is invoked to add expression repair content to the first response content to obtain the second response content;

[0029] Update the first response content using the second response content.

[0030] Optionally, the generation of the first response content includes:

[0031] The AI ​​is invoked to clarify the question content based on the picture book page data and the question text, and to obtain the question recognition text;

[0032] The AI ​​is invoked to generate the first response content based on the picture book page data, the question recognition text, the first age range, and the target training direction.

[0033] Optionally, the AI ​​is an intelligent agent, and the method further includes:

[0034] The first response will be played back via audio.

[0035] If no new question audio data is received within a preset time after the first response content has finished playing, the question-and-answer mode will be exited and the picture book projection will continue to play.

[0036] By summarizing the question text, the picture book page data, and the first response content, historical interaction data is obtained.

[0037] The intelligent agent is controlled to optimize its ability to match training directions and generate response content based on the historical interaction data.

[0038] Secondly, this disclosure also provides a picture book projection story machine question-and-answer interactive device. The device includes:

[0039] The data acquisition module is used to acquire audio data of the question and picture book page data; the picture book page data includes the current picture book page data and picture book context data.

[0040] The speech recognition module is used to perform speech recognition on the question audio data to obtain the question text;

[0041] An age range determination module is used to determine the age range of the question audio data to obtain the first age range to which the questioner belongs.

[0042] The framework matching module is used to match an age-appropriate growth framework from a preset growth framework library based on the first age range; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction;

[0043] The direction matching module is used to call the AI ​​to match the cultivation direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target cultivation direction.

[0044] The content generation module is used to call AI to generate the first response content based on the picture book page data, the first age range, the question text, and the target training direction.

[0045] Optionally, the device further includes a picture book projection module, an interactive wake-up module, an audio output module, a storage module, a power module, and an intelligent agent module; the intelligent agent module is used to call an intelligent agent for data processing, and the intelligent agent is deployed in the storage module or the cloud; the interactive wake-up module includes a button wake-up unit and a voice wake-up unit, used to receive the user's wake-up command and trigger the device to enter a question-and-answer mode; the picture book projection module is used to project picture books or auxiliary images; the audio output module is used to output voice; the power module is used to supply power to the device; the data acquisition module includes an image acquisition unit and a voice acquisition unit.

[0046] The aforementioned question-and-answer interaction method and device for the picture book projection story machine, when a questioner asks a question about the content projected by the picture book projection story machine, firstly processes the collected question audio data in two ways: firstly, it performs speech recognition to obtain the question text; secondly, it determines the questioner's age range based on the question audio data and further matches it with an age-appropriate developmental framework. Compared to existing technologies that only perform question recognition, this achieves more full utilization of the question audio data. Combined with the design of an age-appropriate developmental framework, it enables the picture book projection story machine to provide personalized interaction based on the questioner's age characteristics and educational goals. Then, AI is invoked twice. The first time, AI uses the question scenario composed of the picture book page content and the question to intelligently match a suitable developmental direction from the age-appropriate developmental framework to help achieve context-specific educational guidance. The second time, AI comprehensively considers four aspects of information—picture book page data, question text, the first age range, and the target developmental direction—to intelligently generate the first response content, thereby improving the questioner's interactive experience at the response content level and promoting the questioner's scientific growth. Attached Figure Description

[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0048] Figure 1 This is a flowchart illustrating a question-and-answer interaction method for a picture book projection story machine in one embodiment;

[0049] Figure 2This is a structural block diagram of a picture book projection story machine question-and-answer interactive device in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description, in conjunction with the accompanying drawings and embodiments, provides a more comprehensive understanding of this disclosure. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure. In the following detailed description of this application, certain specific details are described in detail. Those skilled in the art can fully understand this application even without these detailed descriptions. To avoid obscuring the essence of this application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0051] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0052] The question-and-answer interaction method for a picture book projection story machine provided in this disclosure can be applied to various application environments, such as in the usage scenarios of a picture book projection story machine. In one usage scenario, the picture book projection story machine projects picture book story images while simultaneously playing introductory audio for the corresponding story images. Children watch the projected story images and listen to the introductory audio. When children have questions about the projected content, they can ask questions. The audio acquisition module of the picture book projection story machine can collect the child's question audio and retrieve relevant data from the currently projected picture book page. Then, it processes the data using the method of this disclosure to obtain a first response. The picture book projection story machine can then verbally broadcast the first response through an audio device as an answer to the child's question. Specific application scenarios can also be found in the descriptions in subsequent device embodiments.

[0053] In one embodiment, such as Figure 1 As shown, a question-and-answer interaction method for a picture book projection story machine is provided. Taking the application of this method in the aforementioned application scenario as an example, the method includes the following steps:

[0054] Step 102: Obtain the audio data of the question and the page data of the picture book.

[0055] Specifically, the audio data for asking questions can refer to recorded audio data containing the questioner's question information and acoustic characteristic information. The picture book page data can refer to data containing one or more pages of content projected onto a picture book projector / story machine.

[0056] Specifically, in the application scenario of picture book projection story machines, the audio data for questions is generally a recording of children asking questions. The picture book page data includes at least information about the picture book page currently being projected.

[0057] Step 104: Perform speech recognition on the audio data of the question to obtain the text of the question.

[0058] The question text can refer to the text describing the question being asked.

[0059] Specifically, mature speech recognition technologies already exist that can convert audio recordings into text. Technicians can choose appropriate speech recognition technologies to process the audio data of questions, such as commercial speech recognition algorithms or artificial intelligence, and obtain the text of the question after recognition.

[0060] Step 106: Determine the age range of the question audio data to obtain the first age range to which the questioner belongs.

[0061] The preset age range can refer to the pre-defined age segments of the questioners.

[0062] Specifically, questioners of different ages have different acoustic characteristics, meaning that a person's acoustic characteristics contain age information. This is especially true for children, whose acoustic characteristics include many age-related features such as fundamental frequency (also known as the "F0 value"), formants, speech rate, and articulation. Using these acoustic characteristics, the approximate age range of a child can be identified relatively accurately. The system can preset the age range division for questioners, dividing their possible ages into multiple segments, such as 0-3 years, 5-8 years, etc., and setting reference ranges for various acoustic characteristics for each age group based on statistical data, such as acoustic characteristic data from children's acoustic studies. Therefore, for the acquired question audio data, acoustic characteristics can be extracted as needed. Then, based on the questioner's acoustic characteristic data and the preset reference ranges for each acoustic characteristic within the preset age range in the system, an age matching judgment is performed. The preset age range with the highest matching degree is taken as the questioner's age range, i.e., the first age range. The division of the preset age range should consider the intellectual characteristics of children in each age group, as well as the distinguishability of their acoustic features, to achieve a scientific division. This embodiment does not limit the specific division method. Regarding the specific matching method, based on the preset age range and the corresponding acoustic feature reference range for each age range, a specific algorithm can be used, or artificial intelligence can be directly invoked. Existing technologies already include software capable of voice age identification, and there are many publicly available patented technologies and test data. The solution in this disclosure does not involve innovations related to the accuracy of audio age recognition, and therefore will not be elaborated upon.

[0063] Step 108: Based on the first age range, match an age-appropriate growth framework from the preset growth framework library; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction.

[0064] Here, a pre-defined growth framework can refer to a pre-set ability development plan adapted to a pre-defined age range. A pre-defined growth framework library can refer to a database composed of multiple pre-defined growth frameworks. An age-appropriate growth framework can refer to a pre-defined growth framework applicable to a specific age range.

[0065] Specifically, for each preset age range, such as different age groups of children, an age-appropriate ability development plan, i.e., a preset growth framework, can be established. Each preset growth framework includes at least three developmental dimensions: emotion, language, and thinking. Each developmental dimension includes at least one developmental direction. For example, the developmental direction of the emotion dimension can include perceiving emotions, expressing emotions, or regulating emotions, etc. For example, the developmental direction of the language dimension can include language imitation, polite speech, proactive expression, or accurate expression, etc. For example, the developmental direction of the thinking dimension can include multi-sensory coordination, content memory, logical understanding, or innovative imagination, etc. Existing technologies have mature theoretical systems and plans for children's ability development at different ages. This embodiment does not limit the specific content of the developmental directions in the preset growth framework. Technicians can make scientific designs based on existing research results in this field. Therefore, a preset growth framework corresponding to the first age range can be matched from the preset growth framework library to obtain an age-appropriate growth framework. It should be noted that step 104 does not necessarily need to be completed before step 106; it only needs to be completed before step 110. For example, it can also be completed after step 108, or simultaneously with step 106 or step 108.

[0066] Step 110: The AI ​​is invoked to match the cultivation direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target cultivation direction.

[0067] AI stands for Artificial Intelligence, which can be deployed locally or in the cloud, and can be a general-purpose large model or intelligent agent.

[0068] Specifically, after obtaining the age-appropriate development framework, since it contains more than one development direction, it is also necessary to match a suitable development direction. The report page data and the text of the question can be sent to the AI, requesting it to select the most suitable development direction from the multiple development directions within the age-appropriate development framework based on the provided content. This most suitable development direction is called the target development direction.

[0069] Step 112: Call the AI ​​to generate a first response that conforms to the first age range and the target training direction based on the picture book page data and the question text.

[0070] The first response is a response generated based on the data from the picture book pages, which is within the first age range and the target training direction, and is able to answer the questions asked.

[0071] Specifically, the system can send picture book page data, the first age range, the question text, and the target educational direction to the AI. The AI ​​is then tasked with generating a response to the question text based on the picture book page content described in the picture book page data. The AI ​​then adjusts the language style and expands the content of the response to ensure the language style matches the first age range, the content is appropriate for the questioner's knowledge level within that age range, and aligns with the educational guidance in the target educational direction. This results in the first response. After generating the first response, the picture book projector can read it aloud, enabling voice interaction with the questioner. In other words, using the first response to answer a question is based on the projected content of the picture book projector, improving the accuracy of the answer; the language style matches the questioner's age, enhancing the interactive experience; and the inclusion of expanded content related to the appropriate educational direction enriches the response and provides guidance to the questioner on a suitable educational path, increasing the fun and educational value of the interaction.

[0072] In this embodiment, when a questioner asks a question about the content projected by the picture book projection story machine, the collected audio data of the question is processed in two ways: first, speech recognition is performed to obtain the text of the question; second, the questioner's age range is determined based on the audio data, and an age-appropriate developmental framework is matched accordingly. Compared to existing technologies that only perform question recognition, this approach achieves more full utilization of the audio data. Combined with the design of the age-appropriate developmental framework, the picture book projection story machine can provide personalized interaction based on the questioner's age characteristics and developmental goals. Then, AI is invoked twice. The first time, AI intelligently matches a suitable developmental direction from the age-appropriate developmental framework based on the question scenario composed of the picture book page content and the question, helping to achieve context-specific educational guidance. The second time, AI comprehensively considers four aspects of information—picture book page data, question text, the first age range, and the target developmental direction—to intelligently generate the first response content, thereby improving the questioner's interactive experience at the response level and promoting the questioner's scientific development.

[0073] In one embodiment, the picture book page data includes the current picture book page data and picture book context data. There are three preset age ranges: 0-3 years, 3-6 years, and 6 years and above. The appropriate developmental framework for 0-3 years includes at least the following developmental directions: basic emotional perception ability, daily scene description ability, and observation and perception of things. The appropriate developmental framework for 3-6 years includes at least the following developmental directions: emotional regulation awareness, awareness of love and companionship, proactive expression ability, healthy living habits, safety awareness, imagination, and memory. The appropriate developmental framework for 6 years and above includes at least the following developmental directions: awareness of frustration and emotional regulation, awareness of care and understanding, complex expression ability, time management awareness, self-learning awareness, and analytical and creative abilities.

[0074] The current picture book page data can refer to data containing the page content currently being projected by the picture book projector. The picture book context data can refer to data containing the content of projected pages that have been or will be projected by the picture book projector. In this embodiment, referring to the research results of existing child development systems, three age ranges and age-appropriate development frameworks for each age range are designed, which helps to make the setting of the preset development framework library more scientific and the matching of target development directions more convenient and accurate.

[0075] In one embodiment, before acquiring the audio data of the question, the method further includes:

[0076] Receive a trigger signal for question-and-answer mode; the trigger signal includes a key signal or a voice signal;

[0077] Based on the trigger signal, a preset question-and-answer prompt will be played.

[0078] Here, "key signal" refers to a trigger signal emitted via a key. "Voice signal" refers to a trigger signal emitted via voice. "Preset question guidance voice" refers to a pre-set voice used to guide the questioner in asking questions.

[0079] Specifically, the picture book projector story machine can be set to projection mode and question-and-answer mode. When turned on, the picture book projector story machine defaults to projection mode. Upon receiving a trigger signal, it switches from projection mode to trigger mode, during which the projected page remains fixed until it returns to projection mode. The trigger signal can be a button signal or a voice signal. The button signal can be emitted through the buttons on the picture book projector story machine. The voice signal can be a preset voice command, such as "Hello, Xiaoying" or "Xiaoying, Xiaoying." When the system recognizes the preset voice command, it controls the device to enter question-and-answer mode. After entering question-and-answer mode, a preset question-and-answer guidance voice can be played immediately, such as "Hello, what question do you have?" or "Baby, what do you want to ask?". The design in this embodiment facilitates intelligent control of the picture book projector story machine and guides the questioner to ask questions.

[0080] In one embodiment, the method further includes:

[0081] Based on the content of the first response, match or generate an auxiliary screen;

[0082] The first response is played aloud, and the auxiliary screen is projected simultaneously.

[0083] In this context, "auxiliary screen" can refer to a projected image used to help the questioner better understand the content of the first response.

[0084] Specifically, the system can pre-set auxiliary screens to help answer the questioner's questions, along with an auxiliary screen matching mechanism. For example, each auxiliary screen could be assigned several matching keywords, allowing for keyword-based matching. After generating the first response, a suitable auxiliary screen is matched. Alternatively, AI can be used to directly generate auxiliary screens based on the first response. When the first response is played aloud by the storytelling projector, the corresponding auxiliary screens can be projected simultaneously, providing the questioner with a synchronized audiovisual interactive experience.

[0085] In one embodiment, before generating the first response content, the method further includes:

[0086] The AI ​​was used to analyze the text of the question to obtain a second age range.

[0087] The intersection of the first age range and the second age range is used to obtain the third age range;

[0088] Update the first age range using the third age range.

[0089] Specifically, questions asked by children of different ages generally differ in length and depth. Therefore, AI can be used to estimate the questioner's age based on the content of the question text, thus obtaining a second age range. Then, it is determined whether the first and second age ranges overlap. If they do, the overlapping range is designated as the third age range; otherwise, no processing is performed. After obtaining the third age range, it can be used as a new first age range for other steps, such as generating the first response. In this embodiment, narrowing the first age range using the question text helps obtain more accurate questioner information, thereby improving the relevance of the generated first response and providing a better user experience.

[0090] In one embodiment, the method further includes:

[0091] Obtain the questioner's facial image;

[0092] The facial image is used to perform expression recognition to obtain the questioner's expression;

[0093] Determine whether the questioner's expression belongs to the preset expressions that need to be repaired;

[0094] When the questioner's expression is a preset expression to be repaired, the AI ​​is invoked to add expression repair content to the first response content to obtain the second response content;

[0095] Update the first response content using the second response content.

[0096] Among these, the preset expressions to be repaired can refer to pre-defined expressions that require intervention. The expression repair content can refer to content designed to encourage the questioner to improve their expression.

[0097] Specifically, a camera can be added to a picture book projection story machine to capture the questioner's facial image with the user's permission. After acquiring the questioner's facial image, an expression recognition algorithm or AI can be used to identify the questioner's expression. Then, it is determined whether the questioner's expression belongs to a preset expression to be repaired pre-stored in the system. If it does not belong, no processing is performed; if it does belong, AI is invoked to add expression repair content to the first response content, resulting in an updated first response content. The method added in this embodiment can add expression repair content to the first response content when the questioner's state is detected to be poor through expression recognition, thereby adding an expression adjustment function to the question response.

[0098] In one embodiment, generating the first response content includes:

[0099] The AI ​​is invoked to clarify the question content based on the picture book page data and the question text, and to obtain the question recognition text;

[0100] The AI ​​is invoked to generate the first response content based on the picture book page data, the question recognition text, the first age range, and the target training direction.

[0101] Specifically, children's questions are often incomplete, and failing to identify or misidentifying the question can lead to a poor user experience. Therefore, AI can be used to first identify the question based on picture book page data and the question text. By leveraging AI's content analysis and expansion capabilities, the question can be clearly defined, resulting in a identified question text. This identified question text, rather than the original question text, is then used to generate the first response. In this embodiment, using AI to clarify the question beforehand helps improve the relevance of the response and optimizes the user experience.

[0102] In one embodiment, the method further includes:

[0103] After the first response is played, if no new question audio data is received within a preset time after the playback ends, the question-and-answer mode will exit and the picture book projection will continue. Technicians can set the preset time according to the actual situation; a setting of 3-5 seconds is recommended, such as 3 seconds, 4 seconds, or 5 seconds.

[0104] In one embodiment, the AI ​​is an intelligent agent, and the method further includes:

[0105] By summarizing the question text, the picture book page data, and the first response content, historical interaction data is obtained.

[0106] The intelligent agent is controlled to optimize its ability to match training directions and generate response content based on the historical interaction data.

[0107] Specifically, an intelligent agent can be deployed within a picture book projection story machine to handle tasks requiring AI. At this point, the question text, picture book page data, and the first response can be aggregated to obtain historical interaction data. This historical interaction data can be used as training data, allowing the intelligent agent to learn and evolve, optimizing its ability to match training directions and generate response content. This helps improve the accuracy of training direction matching and the quality of response generation, providing a better interactive experience for the questioner and enhancing the overall interaction level.

[0108] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0109] Based on the same inventive concept, this disclosure also provides a picture book projection story machine question-and-answer interaction device for implementing the above-mentioned question-and-answer interaction method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more picture book projection story machine question-and-answer interaction device embodiments provided below can be found in the limitations of the picture book projection story machine question-and-answer interaction method above, and will not be repeated here.

[0110] In one embodiment, such as Figure 2 As shown, a picture book projection story machine question-and-answer interactive device 01 is provided, including: a data acquisition module 10, a speech recognition module 20, an age range judgment module 30, a frame matching module 40, a direction matching module 50, and a content generation module 60, wherein:

[0111] The data acquisition module 10 is used to acquire audio data of the question and picture book page data; the picture book page data includes the current picture book page data and picture book context data.

[0112] The speech recognition module 20 is used to perform speech recognition on the audio data of the question to obtain the text of the question.

[0113] The age range determination module 30 is used to determine the age range of the question audio data to obtain the first age range to which the questioner belongs.

[0114] The framework matching module 40 is used to match an age-appropriate growth framework from a preset growth framework library according to the first age range; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction.

[0115] The direction matching module 50 is used to call the AI ​​to match the cultivation direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target cultivation direction.

[0116] The content generation module 60 is used to call the AI ​​to generate the first response content based on the picture book page data, the first age range, the question text, and the target training direction.

[0117] In one embodiment, the picture book projection story machine Q&A interactive device includes a data storage module. The data storage module stores picture book page data, including the current picture book page data and picture book context data. The preset age ranges stored in the data storage module include 0-3 years old, 3-6 years old, and 6 years old and above. The data storage module stores age-appropriate growth frameworks for each age group, including: for 0-3 years old, the age-appropriate growth framework should include at least the following development directions: basic emotional perception ability, daily scene description ability, and the ability to observe and perceive things; for 3-6 years old, the age-appropriate growth framework should include at least the following development directions: emotional regulation awareness, love and companionship awareness, proactive expression ability, healthy living habits, safety awareness, imagination, and memory; for 6 years old and above, the age-appropriate growth framework should include at least the following development directions: frustration and emotional regulation awareness, care and understanding awareness, complex expression ability, time management awareness, self-learning awareness, and analytical and creative ability.

[0118] In one embodiment, the device further includes a voice playback module, and the data acquisition module is further configured to receive a trigger signal for a question-and-answer mode; the trigger signal includes a button signal or a voice signal. The voice playback module is configured to play a preset question-and-answer guidance voice according to the trigger signal.

[0119] In one embodiment, the device further includes an auxiliary image module and a projection module. The auxiliary image module is used to match or generate an auxiliary image based on the first response content. The projection module is used to synchronously project the auxiliary image while the voice playback module plays the first response content.

[0120] In one embodiment, before generating the first response content, the device further includes an age analysis module, a comprehensive evaluation module, and an age update module. The age analysis module is used to invoke AI to analyze the question text to obtain a second age range. The comprehensive evaluation module is used to take the intersection of the first age range and the second age range to obtain a third age range. The age update module is used to update the first age range using the third age range.

[0121] In one embodiment, the device further includes an expression recognition module, an expression judgment module, a content optimization module, and a content update module. The data acquisition module is further configured to acquire a facial image of the questioner. The expression recognition module is configured to perform expression recognition on the facial image to obtain the questioner's expression. The expression judgment module is configured to determine whether the questioner's expression belongs to a preset expression to be repaired. The content optimization module is configured to, when the questioner's expression belongs to a preset expression to be repaired, invoke AI to add expression repair content to the first response content to obtain a second response content. The content update module is configured to update the first response content using the second response content.

[0122] In one embodiment, the content generation module includes a semantic understanding module and an intelligent processing module. The semantic understanding module is used to invoke AI to clarify the question content based on the picture book page data and the question text, obtaining question recognition text. The intelligent processing module is used to invoke AI to generate a first response content based on the picture book page data, the question recognition text, the first age range, and the target educational direction.

[0123] In one embodiment, the AI ​​is an intelligent agent, and the device further includes:

[0124] The voice playback module is used to play the first response content via voice.

[0125] The mode adjustment module is used to exit the question-and-answer mode and continue playing the picture book projection if no new question audio data is received within a preset time after the first response content has finished playing by voice.

[0126] The data collection module is used to summarize the question text, the picture book page data, and the first response content to obtain historical interaction data;

[0127] The self-training module is used to control the agent to optimize its direction matching ability and response content generation ability based on the historical interaction data.

[0128] In one embodiment, the device further includes a picture book projection module, an interactive wake-up module, an audio output module, a storage module, a power module, and an intelligent agent module. The intelligent agent module is used to invoke an intelligent agent for data processing; the intelligent agent is deployed in the storage module or in the cloud. The interactive wake-up module includes a button wake-up unit and a voice wake-up unit, used to receive user wake-up commands and trigger the device to enter a question-and-answer mode. The picture book projection module is used to project picture books or auxiliary images. The audio output module is used to output voice. The power module is used to supply power to the device. The data acquisition module includes an image acquisition unit and a voice acquisition unit.

[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0130] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (RRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0131] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.

Claims

1. A question-and-answer interactive method for a picture book projection story machine, characterized in that, The method includes: Obtain audio data of the questions and page data of the picture book; The audio data of the question is subjected to speech recognition to obtain the text of the question; The question audio data is subjected to a preset age range determination to obtain the first age range to which the questioner belongs; Based on the first age range, an age-appropriate growth framework is matched from a preset growth framework library; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction. The AI ​​is invoked to match the development direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target development direction. The AI ​​is invoked to generate a first response that conforms to the first age range and the target educational direction, based on the picture book page data and the question text.

2. The method according to claim 1, characterized in that, The picture book page data includes the current picture book page data and the picture book context data; the preset age range includes 0-3 years, 3-6 years, and 6 years and above; the appropriate growth framework for 0-3 years includes at least the following development directions: basic emotional perception ability, daily scene description ability, and the ability to observe and perceive things; the appropriate growth framework for 3-6 years includes at least the following development directions: emotional regulation awareness, love and companionship awareness, proactive expression ability, healthy living habits, safety awareness, imagination, and memory; the appropriate growth framework for 6 years and above includes at least the following development directions: frustration and emotional regulation awareness, care and understanding awareness, complex expression ability, time management awareness, self-learning awareness, and analytical and creative ability.

3. The method according to claim 1, characterized in that, Before acquiring the audio data of the question, the method further includes: Receive a trigger signal for question-and-answer mode; the trigger signal includes a key signal or a voice signal; Based on the trigger signal, a preset question-and-answer prompt will be played.

4. The method according to claim 1, characterized in that, The method further includes: Based on the content of the first response, match or generate an auxiliary screen; The first response is played aloud, and the auxiliary screen is projected simultaneously.

5. The method according to claim 1, characterized in that, Before generating the first response content, the method further includes: The AI ​​was used to analyze the text of the question to obtain a second age range. The intersection of the first age range and the second age range is used to obtain the third age range; Update the first age range using the third age range.

6. The method according to claim 1, characterized in that, The method further includes: Obtain the questioner's facial image; The facial image is used to perform expression recognition to obtain the questioner's expression; Determine whether the questioner's expression belongs to the preset expressions that need to be repaired; When the questioner's expression is a preset expression to be repaired, the AI ​​is invoked to add expression repair content to the first response content to obtain the second response content; Update the first response content using the second response content.

7. The method according to claim 2, characterized in that, The content of the generated first response includes: The AI ​​is invoked to clarify the question content based on the picture book page data and the question text, and to obtain the question recognition text; The AI ​​is invoked to generate the first response content based on the picture book page data, the question recognition text, the first age range, and the target training direction.

8. The method according to claim 3, characterized in that, The AI ​​is an intelligent agent, and the method further includes: The first response will be played back via audio. If no new question audio data is received within a preset time after the first response content has finished playing, the question-and-answer mode will be exited and the picture book projection will continue to play. By summarizing the question text, the picture book page data, and the first response content, historical interaction data is obtained. The intelligent agent is controlled to optimize its ability to match training directions and generate response content based on the historical interaction data.

9. A question-and-answer interactive device for a picture book projection story machine, characterized in that, The device includes: The data acquisition module is used to acquire audio data of the question and picture book page data; the picture book page data includes the current picture book page data and picture book context data. The speech recognition module is used to perform speech recognition on the question audio data to obtain the question text; An age range determination module is used to determine the age range of the question audio data to obtain the first age range to which the questioner belongs. The framework matching module is used to match an age-appropriate growth framework from a preset growth framework library based on the first age range; the age-appropriate growth framework includes at least three developmental dimensions: emotion, language, and thinking, and each developmental dimension includes at least one developmental direction; The direction matching module is used to call the AI ​​to match the cultivation direction from the age-appropriate growth framework based on the picture book page data and the question text, and obtain the target cultivation direction. The content generation module is used to call AI to generate the first response content based on the picture book page data, the first age range, the question text, and the target training direction.

10. The apparatus according to claim 9, characterized in that, The device further includes a picture book projection module, an interactive wake-up module, an audio output module, a storage module, a power module, and an intelligent agent module. The intelligent agent module is used to call upon an intelligent agent for data processing, and the intelligent agent is deployed in the storage module or the cloud. The interactive wake-up module includes a button wake-up unit and a voice wake-up unit, used to receive the user's wake-up command and trigger the device to enter a question-and-answer mode. The picture book projection module is used to project picture books or auxiliary images. The audio output module is used to output voice. The power module is used to supply power to the device. The data acquisition module includes an image acquisition unit and a voice acquisition unit.