Large model-based child dialogue method and device, electronic equipment and storage medium

By acquiring and analyzing children's audio data, combining it with parental feedback to generate learning audio and provide scoring, the problems of lack of personalization and parental participation in existing systems are solved, and the interactive experience and learning effect of children's dialogue systems are improved.

CN119152738BActive Publication Date: 2025-10-21广州词元信息科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411197916.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-10-21
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

Existing dialogue systems and educational robots have obvious shortcomings in meeting the diverse needs of children and parents. They lack personalized optimization, parent participation and control mechanisms, are unable to meet educational needs, and the interactive experience is stiff and lacks specificity.

Method used

By obtaining children's audio data and parsing it into text data, parent feedback data is used to determine the target learning section, learning audio is generated and output, learning progress and scoring data are collected, and personalized optimization and parent participation are combined with large models to provide visual learning progress and scoring.

Benefits of technology

It improves the interactive experience of the children's dialogue system, enhances parental participation and the targetedness of the dialogue system, meets the diverse needs of children and parents, and realizes personalized learning guidance and scoring feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152738B_ABST
    Figure CN119152738B_ABST
Patent Text Reader

Abstract

The application provides a large model-based child dialogue method and device, electronic equipment and storage medium, and relates to the technical field of computers. The method comprises the following steps: obtaining child audio data corresponding to a child, analyzing and converting the child audio data into corresponding child text data, sending the child text data to a parent device, obtaining parent feedback data corresponding to the child text data based on the parent, determining a target learning block based on the parent feedback data, generating learning audio based on the target learning block by a large model, outputting the learning audio to the child, and sending the learning audio to the parent device; after the dialogue ends, collecting learning progress of the target learning block, obtaining score data of the child learning the target learning block, and feeding back the learning progress and the score data to the parent device. The application solves the problem of obvious deficiency in meeting the diversified needs of children and parents in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a method, device, electronic device, and storage medium for children's dialogue based on a large model. Background Art

[0002] With the rapid development of modern society and the increasingly accelerated pace of life, parents are generally faced with time constraints, resulting in a significant reduction in the time they spend with their children. This situation not only affects the establishment and maintenance of parent-child relationships but also limits parents' effective participation in their children's education. Furthermore, due to the rapid pace of knowledge change and the limitations of parents' own abilities and energy, they often struggle to provide timely and reasonable explanations to children's questions, making it difficult to satisfy their children's growing curiosity and thirst for knowledge.

[0003] While current conversational systems on the market offer users a convenient way to access information, they are primarily designed for adults and are often presented via web or mobile platforms. Answers are presented in a monotonous format, relying primarily on text output, and lack an intuitive and efficient voice interaction experience. This interaction approach not only limits the efficiency of information delivery but also fails to fully consider the specific needs of children, such as the naturalness and engaging nature of voice communication. More importantly, existing conversational systems generally lack the ability to analyze long-term chat logs, making them unable to tailor their interactions to children's habits and interests. They are often based on pre-set, fixed responses, making it difficult to dynamically adjust to a child's specific situation, resulting in a stilted and unfocused interaction experience. Furthermore, these systems lack effective parental engagement and control mechanisms, failing to provide necessary intervention and guidance, limiting their active role in their children's education. Furthermore, current conversational systems lack goal-oriented or task-oriented approaches, focusing primarily on casual conversation and failing to meet parents' educational needs.

[0004] In summary, existing dialogue systems and educational robots have obvious shortcomings in meeting the diverse needs of children and parents. Summary of the Invention

[0005] This application provides a method, device, electronic device, and storage medium for children's dialogue based on a large model, which can solve the problem that existing dialogue systems and educational robots in the related art are obviously insufficient in meeting the diverse needs of children and parents. The technical solution is as follows:

[0006] According to one aspect of the present application, a child conversation method based on a large model includes: obtaining child audio data corresponding to the child, parsing the child audio data and converting it into corresponding child text data, and delivering the child text data to a parent device, wherein the child text data includes the knowledge section involved in the child audio; obtaining parent feedback data corresponding to the child text data based on the parent feedback data, determining the corresponding target learning section based on the parent feedback data, generating corresponding learning audio based on the target learning section with a large model, outputting the learning audio to the child, and sending the learning audio to the parent device in text form; after the conversation ends, collecting the learning progress for the target learning section, and obtaining the scoring data of the child's learning of the target learning section, and feeding back the learning progress and the scoring data to the parent device.

[0007] In an exemplary embodiment, after obtaining the child audio data corresponding to the child, the method further includes:

[0008] Preprocessing the children's audio data and extracting corresponding children's audio features from the children's audio data;

[0009] Retrieve a pre-trained wake-up word detection model, wherein the corresponding wake-up word dataset and non-wake-up word dataset are obtained, the wake-up word dataset and non-wake-up word dataset are input into the training model for training until the training model converges, and the corresponding wake-up word detection model is determined;

[0010] The child's audio features are input into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

[0011] In an exemplary embodiment, when outputting learning audio to a child, it is necessary to select a timbre of the output audio, and the method further includes:

[0012] Retrieve a pre-recorded sound data set;

[0013] Playing the corresponding basic speech in the timbre data set to the child, and determining the corresponding target timbre data based on the child's selection;

[0014] The process of recording timbre data is as follows:

[0015] Acquire recorded audio data of the character's timbre, automatically segment the recorded audio data into several audio segments based on a speech segmentation model, and store them in audio files;

[0016] A timbre conversion model is used to extract features of each audio file to obtain corresponding audio timbre features, the audio timbre features can be stored locally, and the audio timbre features are associated with the characters.

[0017] In an exemplary embodiment, the method further comprises:

[0018] After the parent logs in for the first time, the corresponding age-appropriate learning section is determined based on the child's age input;

[0019] Sending the planned schedule corresponding to the age-appropriate learning section to the parent's device, and determining the corresponding target schedule based on the selection information of the parent's device;

[0020] After the child's conversation ends, the learning progress corresponding to the target learning section will be determined, and visual data will be provided based on the learning progress and the target progress chart and sent to the parent's device.

[0021] In an exemplary embodiment, when a long-term goal is selected, after the child initiates the conversation, the method further includes:

[0022] If there is a long-term goal, the large model generates corresponding related content for playback based on the long-term goal, where the related content is the conversation content close to the long-term goal, and guides the conversation direction and content according to the long-term goal;

[0023] If there are multiple long-term goals, play the content of each long-term goal to the child, and determine the corresponding long-term goal based on the child's feedback;

[0024] The content to be played is selected based on the long-term goal.

[0025] In an exemplary embodiment, after obtaining parent feedback data corresponding to the child's text data, the method further includes:

[0026] Determining corresponding target filtering words based on the parent feedback data;

[0027] Set corresponding blocked words and restricted words based on the target filtering words;

[0028] During the conversation, the number of times the restrictive word appears is obtained. If the number of times the restrictive word appears is greater than a restrictive word threshold, the restrictive word is converted into a blocked word.

[0029] According to one aspect of the present application, a large model-based children's dialogue device includes:

[0030] A child text data acquisition module, which acquires the child's audio data corresponding to the child, parses the child's audio data and converts it into corresponding child text data, and transmits the child text data to the parent's device, wherein the child text data includes the knowledge section related to the child's audio;

[0031] A target learning section determination module obtains parent feedback data corresponding to the child's text data, determines the corresponding target learning section based on the parent feedback data, generates corresponding learning audio based on the target learning section, outputs the learning audio to the child, and sends the learning audio to the parent's device in text form;

[0032] The parent feedback module collects the learning progress of the target learning section and obtains the scoring data of the child's learning target learning section after the conversation ends, and is used to feed back the learning progress and the scoring data to the parent device.

[0033] In an exemplary embodiment, the apparatus includes, but is not limited to:

[0034] a child audio feature extraction module, which preprocesses the child audio data and extracts corresponding child audio features from the child audio data;

[0035] A wake-up word detection model retrieval module is used to retrieve a pre-trained wake-up word detection model, wherein the corresponding wake-up word dataset and non-wake-up word dataset are obtained, and the wake-up word dataset and non-wake-up word dataset are input into the training model for training until the training model converges, thereby determining the corresponding wake-up word detection model;

[0036] The wake-up word judgment module inputs the child's audio features into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

[0037] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the large model-based children's dialogue method as described above.

[0038] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, wherein the computer-readable instructions are executed by one or more processors to implement the large model-based children's dialogue method as described above.

[0039] According to one aspect of the present application, a computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the large model-based children's dialogue method as described above.

[0040] The beneficial effects of the technical solution provided by this application are:

[0041] In the above technical solution, the system first obtains the children's audio data corresponding to the children, and then parses the children's audio data and converts it into corresponding children's text data. The children's text data contains the knowledge sections involved in the children's audio. Then, through the existing common intention recognition model, the knowledge sections that the children may design when they just talk can be parsed, and then the children's text data can be transmitted to the parent's device, so that the parents can be notified of the knowledge sections that the children may be involved in at the first time; after the parents receive the corresponding information, they can make a selection on the software on the parent's device, and can determine whether to conduct knowledge learning of the relevant sections this time, and after obtaining the corresponding parent feedback data of the parents. After that, the corresponding target learning section can be determined based on the parent feedback data. The target learning section here can be the previously scheduled learning goals or the learning section currently arranged by the parents. The system can then retrieve the learning audio related to the target learning section and output the learning audio to the child, thereby guiding the child to learn accordingly. In addition, after each conversation with the child, the system can collect the learning progress of the target learning section and score it according to the conversation process, that is, it can obtain the scoring data of the child's learning target learning section, and feed back the learning progress and the scoring data to the parent device. The parent can view the child's relative learning situation from the parent device. Through the above process, the system increases the parent's participation in the conversation process with the child. Parents can adjust the knowledge that needs to be learned according to the child's actual situation. Parents not only participate in it, but also can know the child's learning situation at the first time after each conversation, thereby effectively solving the problem that the relevant technology is obviously insufficient in meeting the diverse needs of children and parents. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.

[0043] Figure 1 It is a schematic diagram of the implementation environment involved in this application;

[0044] Figure 2 is a flowchart of a children's dialogue method based on a large model according to an exemplary embodiment;

[0045] Figure 3 is a flowchart from S001 to S003 of a children's dialogue method based on a large model according to an exemplary embodiment;

[0046] Figure 4is a flowchart from S004 to S005 in a children's dialogue method based on a large model according to an exemplary embodiment;

[0047] Figure 5 is a flowchart from S101 to S103 of a children's dialogue method based on a large model according to an exemplary embodiment;

[0048] Figure 6 is a flowchart from S111 to S113 of a children's dialogue method based on a large model according to an exemplary embodiment;

[0049] Figure 7 The figure is a structural block diagram of a children's dialogue device based on a large model according to an exemplary embodiment. DETAILED DESCRIPTION

[0050] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0051] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0052] As mentioned above, relevant technologies have obvious shortcomings in meeting the diverse needs of children and parents.

[0053] To this end, the large-model-based children's conversation method provided in this application can effectively improve the accuracy of large-model-based children's conversation. Accordingly, the large-model-based children's conversation method is suitable for a large-model-based children's conversation device, which can be deployed on an electronic device.

[0054] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0055] Figure 1 The following is a schematic diagram of the implementation environment involved in a large-scale model-based children's dialogue method. The implementation environment includes a collection end and a service end.

[0056] See also Figure 2 The embodiment of the present application provides a method for children's dialogue based on a large model, which is applicable to electronic devices, which can be Figure 1 The server side in the implementation environment is shown; in the following method embodiment, for the convenience of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0057] like Figure 2 As shown, the method may include the following steps:

[0058] S100, obtaining the child's corresponding audio data, parsing the child's audio data and converting it into corresponding child's text data, and transmitting the child's text data to the parent's device.

[0059] Among them, children's text data includes the knowledge sections involved in children's audio;

[0060] It should be pointed out here that after the child issues a voice command, the voice data is collected, that is, the collected child audio data. At this time, the child audio data needs to be parsed. The system can understand the content of the child audio data and convert the corresponding content from audio into corresponding child text data. At the same time, the child text data is sent to the parents as soon as possible, and the parents can understand it on the device.

[0061] S110, obtain parent feedback data corresponding to the child's text data, determine the corresponding target learning section based on the parent feedback data, the large model generates corresponding learning audio based on the target learning section, outputs the learning audio to the child, and sends the learning audio to the parent's device in text form.

[0062] Among them, every time the parent receives the child's text data sent by the system on the parent device, the parent can make a selection on the software page of the parent device. At this time, the system can receive the data feedback from the software, that is, the learning section selected by the parent on the software page. The system can then determine the learning section that the child needs to involve in this conversation.

[0063] The large model generates corresponding learning audio based on the target learning section and outputs the learning audio to the child. For each set learning section, the system speaker will start the session with "Hello! [child's name], would you like to study today?" or guide the child through the target learning section through dialogue. At the same time, the system will send the learning audio generated by the large model to the parent's device in text form, so that the parent can view the content of the child's upcoming learning in real time.

[0064] In addition, it should be pointed out here that in addition to being used to determine the learning section, parent feedback data will also participate in the result generation strategy of each round of conversation (such as encouraging expression and not doing dangerous things) and selecting specific answers (such as giving two answers for parents to choose from, and after selecting one of them, the system will output it to the child side).

[0065] Parent feedback plays different roles in data at different stages, such as:

[0066] Learning Area Determination: The system displays a panel listing all available learning areas and allows parents to select those they believe are most important for their child. Based on their understanding of their child's interests, ability level, and learning needs, parents can select specific learning areas or courses. For example, if parents believe their child needs to strengthen math skills, they can select "Math" as a priority learning area.

[0067] Dialogue strategy selection: The system can suggest several ways to communicate with children, such as encouragement, praise, or gentle criticism. Parents can choose the most suitable dialogue strategy based on their understanding of their children's personality. When a child completes a task, the system provides two types of feedback: "Great job!" and "Good this time, but try a more difficult task next time." Parents can choose the former if they want to boost their children's self-confidence.

[0068] Conversation outcome selection: The system may provide several preset answers, allowing parents to choose the most appropriate response to their children. This helps ensure that the content of the conversation is consistent with the parents' values ​​and educational philosophy. When a child asks about safety issues, the system may give two answer options: "You can play outside as long as you don't leave the yard." and "If you want to go out to play, make sure you are accompanied by an adult." Parents can choose the latter if they want their children to be supervised by adults at all times.

[0069] During the communication between the child and the system, the system will generate the corresponding communication process through a large model and send the content of the communication process to the parent's device. Therefore, the parent's feedback data will be different based on the different communication content.

[0070] S120, after the conversation ends, collect the learning progress of the target learning section, obtain the scoring data of the child's learning target learning section, and feed back the learning progress and scoring data to the parent's device.

[0071] Among them, after each conversation with the child, the system will determine the learning progress of the target learning section based on the actual conversation situation, and record the child's score for this learning through questioning, so as to obtain the corresponding score data, and finally send the learning progress and score data of each conversation to the parent's device.

[0072] It should also be noted that after the system finishes the conversation with the child, the large model will summarize the entire content of the conversation and send the child's response measures and the parent's response measures to the parent client. For example, the content sent to the parent's device is:

[0073] System: Hi, Xiao Ming! How was your day at school today?

[0074] Xiao Ming: Hi, I’m very happy today because we had an interesting science class!

[0075] System: Wow, so what did you learn?

[0076] Xiao Ming: We learned how plants grow, and the teacher also taught us how to grow bean sprouts.

[0077] System: That sounds great! Do you like growing bean sprouts?

[0078] Xiao Ming: I love it! I think it's magical to watch the seeds slowly turn into bean sprouts.

[0079] System: It's truly amazing! Did you know that many plants, not just bean sprouts, grow from tiny seeds? Have you ever thought about growing some plants at home?

[0080] Xiao Ming: I’ve thought about it, but I don’t know how to start.

[0081] System: That's okay, I can give you some tips. First you need some soil and seeds, and then find a sunny place.

[0082] Summary of the conversation generated by the large model: In this conversation, Xiao Ming expressed his interest in school science classes, particularly the part about growing bean sprouts. He mentioned that he found the process of plant growth fascinating and expressed interest in trying to grow plants at home. The system encouraged Xiao Ming's interest and provided a simple guide to growing plants at home. Overall, Xiao Ming appeared active and curious during the conversation, demonstrating his passion for natural science. Parents are advised to further support Xiao Ming's interests, such as preparing planting activities together. This will not only strengthen the parent-child relationship but also help Xiao Ming learn more about nature.

[0083] After each conversation between the system and the child, the system will record and summarize the entire conversation process. In addition, after each conversation, the model will score the importance of each conversation and give reasons. For example, the system can pre-set the scoring mechanism, and parents can see the scoring status on the interface. For more important conversations, the software will send notifications to parents.

[0084] The system has a built-in list of keywords related to key areas such as children's safety, health, emotional changes, and learning needs. When these keywords appear in a conversation, the conversation score will be improved; using the sentiment analysis module in natural language processing (NLP) technology, the system can identify positive, negative or neutral emotions in the conversation. Strong emotional expressions (whether positive or negative) may increase the importance score of the conversation; by analyzing the context, length and complexity of the conversation, the system can evaluate whether the conversation involves in-depth discussion of topics, which may also improve the score.

[0085] In the embodiment of the present application, parents can set their children's learning goals in advance, which can be long-term goals or short-term goals, such as Figure 3 The specific process is as follows:

[0086] S001, after the parent logs in for the first time, the corresponding age-appropriate learning section is determined based on the child’s input age.

[0087] Among them, at this time, the system will display all learning content suitable for the age of the child to the parents according to the age of the child. At this time, parents can choose the appropriate age-appropriate learning section in the system according to actual needs; in addition, the content of the age-appropriate learning section can be used as the child’s short-term goal and long-term goal. The short-term goal allows the child to complete the learning in one conversation, while the long-term goal requires the child to learn through multiple conversations to complete the entire progress of the long-term goal.

[0088] S002, sending the planned schedule corresponding to the age-appropriate learning section to the parent device, and determining the corresponding target schedule based on the selection information of the parent device.

[0089] It should be pointed out here that parents can view the entire plan schedule on the parent device, and can study on the parent device based on the content of the section that the child is currently interested in or currently studying, and the system will record the corresponding target schedule.

[0090] S003: After the child's conversation ends, the learning progress corresponding to the target learning section will be determined, and visual data will be provided based on the learning progress and target progress chart and sent to the parent's device.

[0091] After each child finishes a conversation, the system will determine the learning progress corresponding to the current target learning section based on the currently ended conversation; and after each conversation ends, the corresponding learning progress will be sent to the parents.

[0092] It should be pointed out here that when children start to have a conversation, the system can let the children learn the corresponding section according to the settings of the parents. In addition, during the learning stage, the system can guide the children to learn the corresponding learning section, such as the above-mentioned "Hello! [child's name], would you like to learn today?", and determine whether to proceed with subsequent learning based on the child's reply.

[0093] At the same time, after determining the target learning section, the system can check whether the child has learned the content of the target learning section, which is the long-term goal mentioned above. At this time, it is necessary to determine the progress point to be learned this time based on the learning progress of the last conversation, and at the end of each conversation, send the data in the form of a table to the parents to visualize the results.

[0094] During a certain period of time, parents may set one or more long-term goals for their children. The system will handle these different situations differently. That is, after each conversation is initiated, the method also includes:

[0095] If there is a long-term goal, the big model generates corresponding related content for playback based on the long-term goal. It should be pointed out here that the related content is the conversation content close to the long-term goal. According to the long-term goal, the direction and content of the conversation are guided. For the case of a long-term goal, after the conversation is started, it is processed according to the normal processing flow.

[0096] If there are multiple long-term goals, it is necessary to first determine with the child which long-term goal needs to be studied in the current conversation. At this time, the system can play the content of each long-term goal to the child, for example, "Hello! [child's name], would you like to learn [Tang poetry] today?" to ask the child. If the child is sure, it can be confirmed. Otherwise, the content of different long-term goals will be switched for inquiry, and then the corresponding long-term goal will be determined based on the child's feedback. After determining the long-term goal of the current conversation, the system can select the content to be played based on the long-term goal.

[0097] Through the above process, the system can guide children to learn accordingly during each conversation. In addition, the system can use various guidance methods during the conversation with children, such as the following:

[0098] Repeat and clarify: If children do not fully understand, patiently explain it again and try to express the same idea in different ways; Provide examples: Use concrete examples to help children understand abstract concepts, which can help them connect new information with what they already know.

[0099] Vivid Explanations: Use metaphors and stories to explain complex concepts, making the information more vivid and memorable; Provide Choices: Give children several different viewpoints or options, encouraging them to think and make choices, which helps develop their decision-making skills.

[0100] Encourage expression: Encourage children to share their thoughts and feelings, which helps them build confidence and learn to express themselves; Questioning skills: Provide several questions for children to choose from, which can stimulate their curiosity and thinking ability.

[0101] Encourage questioning: Encourage children to raise their own questions, which helps them to gain a deeper understanding and develop critical thinking; Try and practice: Encourage children to try and practice the new knowledge they have learned personally. Practice is an important way to consolidate learning.

[0102] Safety warnings: Clearly tell children which behaviors are dangerous and explain why they should not be done to ensure their safety.

[0103] Family communication: Remind children to communicate with their parents in a timely manner when they encounter problems or have important matters. This helps build trust and support within the family.

[0104] Knowledge testing: Simple questions or quizzes to check if children have grasped the content of the communication help ensure understanding and retention.

[0105] For different short-term or long-term goals, the system will use the above different methods to communicate with children, allowing children to learn and master different communication methods through a variety of dialogue methods.

[0106] When the system is communicating with children, it can select a more appropriate timbre based on the children's past habits. The timbre selected here can be the system's initial preset or a subsequent input. Figure 4 As shown in the figure, the process of recording a sound includes:

[0107] S004, obtaining recorded audio data of the character's timbre, automatically segmenting the recorded audio data into several audio segments based on a speech segmentation model, and storing each segment as an audio file;

[0108] S005: A timbre conversion model is used to extract features from each audio file to obtain corresponding audio timbre features. The audio timbre features can be stored locally and associated with the character. This involves using a corresponding timbre synthesis model, which can be a spectral envelope model. The spectral envelope model is a commonly used timbre synthesis method that synthesizes timbre by detailed simulation of the timbre's spectral characteristics. In the spectral envelope method, timbre is typically represented as the product of a spectral characteristic function and an envelope function.

[0109] First, the system receives and obtains the original audio data of the character. The original audio data needs to be associated with the character. Then, it uses an advanced speech segmentation model to automatically and accurately segment these continuous audio streams into multiple independent audio clips. Each clip contains a clear and independent voice part and is saved as an independent audio file. Next, the timbre conversion model is applied to each segmented audio file for in-depth feature extraction to capture the unique timbre characteristics of each audio clip. These extracted audio timbre features are not only securely stored in the local database, but also establish a correspondence with the corresponding character information through an intelligent matching mechanism to ensure that the timbre characteristics of each character can be accurately recorded and distinguished. This process not only simplifies the tedious process of traditional timbre entry and management, but also significantly improves the processing efficiency and quality of timbre data.

[0110] In the process of the system selecting the corresponding timbre, the system first retrieves the pre-recorded timbre data set and plays the corresponding basic voice in the timbre data set to the child. It should be pointed out here that the system can play different timbre data in turn, and determine the corresponding target timbre data based on the child's feedback and the child's choice. Then, during this conversation, the timbre corresponding to the target timbre data can be played; in addition, if the child uses a certain timbre for conversation for a long time and the child does not resist, the system can continue to use the timbre used for a long time for conversation; in addition, it should be pointed out here that the above mentioned the selection of timbre, but for the system, in addition to timbre, it also includes character setting, that is, by inputting a certain character's classic quotes and deeds, the large model completes the establishment of the character setting. Subsequently, the large model will generate an answer that conforms to the character positioning based on the character setting, and generate the final audio answer through its timbre. In other words, each timbre is associated with a certain character, and the timbre selection can be achieved directly by selecting the corresponding character.

[0111] Each time a child has a conversation, the system needs to be awakened. After the system is awakened, subsequent conversations can be carried out, such as Figure 5 As shown, the specific process includes:

[0112] S101, preprocessing the children's audio data and extracting corresponding children's audio features from the children's audio data;

[0113] S102: Retrieve a pre-trained wake-up word detection model, wherein a corresponding wake-up word dataset and a non-wake-up word dataset are obtained, the wake-up word dataset and the non-wake-up word dataset are input into the training model for training until the training model converges, and a corresponding wake-up word detection model is determined;

[0114] S103: Input the child's audio features into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

[0115] First, preprocess the children's audio data to extract key children's audio features. Then, collect and digitize the audio signals, converting analog audio signals into digital signals for computer processing. Noise reduction is performed on the digital audio signals to eliminate or reduce interference factors such as background noise and electronic noise, improving the quality of the audio signals. Then, audio segmentation is performed to divide the continuous audio stream into smaller segments based on time or content to facilitate subsequent feature extraction and processing. On this basis, audio enhancement is performed, which may include volume adjustment, equalization, and other processing to optimize the auditory effect of the audio signal. Finally, for specific application requirements, such as speech recognition or audio classification, feature extraction is performed to extract useful information from the audio signal, such as spectral features and Mel-Frequency Cepstral Coefficients (MFCCs), providing a basis for subsequent processing and analysis.

[0116] Subsequently, a dedicated wake-up word detection model is constructed by using the pre-collected wake-up word dataset and non-wake-up word dataset and training the model until it converges.

[0117] Finally, the extracted children's audio features are input into the wake-up word detection model. The system detects and determines in real time whether the preset wake-up word is recognized. Once the wake-up word is detected, the conversation mode is automatically turned on and prepares for subsequent voice interaction.

[0118] Through the above process, the response speed and user experience of children's voice interaction devices can be improved, ensuring that the device can be quickly activated and enter a conversation state when the child issues a command.

[0119] Finally, during the conversation between the system and the child, some words involved in the conversation need to be shielded, that is, after obtaining the parent feedback data corresponding to the child's text data, such as Figure 6 As shown, the method further includes:

[0120] S111, determining corresponding target filtering words based on parent feedback data.

[0121] It should be pointed out here that the system generally has some pre-set filters, and after parents log in to the system, they can also add some new filters to the system and mark the important levels of each filter.

[0122] S112: Setting corresponding blocked words and restricted words based on the target filtering words.

[0123] In the embodiment of the present application, after the parent makes the mark, the system determines the corresponding blocked words and restricted words according to the importance level. Blocked words are prohibited from being issued during the conversation process, and restricted words can be issued, but the number of times they are issued is limited accordingly.

[0124] S113, during the conversation, obtaining the number of times the restrictive word appears, and if the number of times the restrictive word appears is greater than the restrictive word threshold, converting the restrictive word into a blocked word.

[0125] Among them, a restricted word threshold can be set. When the number of restricted words exceeds the restricted word threshold, it means that the system has sent the corresponding word multiple times during the conversation and has been prohibited from sending the word again. In addition, parent feedback data is used to determine the target filter words, and based on these filter words, blocked words and restricted words are set. The usage rules of these words are dynamically adjusted during the conversation. This can mainly play the following roles:

[0126] Enhanced family monitoring capabilities: Allows parents to participate more directly in the safety management of their children's online communications. By customizing filtering words, parents can set appropriate blocked and restricted words based on their children's specific circumstances and age groups, thereby more effectively controlling the information content their children are exposed to and protecting their children from harmful information.

[0127] Improve the accuracy of information filtering: The system's preset filtering words may not cover all the content that parents are concerned about, but parent-defined filtering words can make up for this deficiency, making information filtering more in line with the specific needs of each family and improving the accuracy and effectiveness of filtering.

[0128] Dynamically adjust protection strategies: During a conversation, the status of restricted words is dynamically adjusted (converted from restricted words to blocked words) based on the number of times they are used. This flexible adjustment mechanism ensures that the system can respond quickly to new situations or challenges, further strengthening protection for children.

[0129] The following is an embodiment of the device of the present application, which can be used to implement the large model-based children's conversation method involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the large model-based children's conversation method involved in the present application.

[0130] See also Figure 7 In the embodiment of the present application, a large model-based children's dialogue device is provided, including but not limited to:

[0131] The child text data acquisition module 200 acquires the child audio data corresponding to the child, parses the child audio data and converts it into corresponding child text data, and transmits the child text data to the parent's device. The child text data includes the knowledge section related to the child audio.

[0132] The target learning section determination module 210 obtains parent feedback data corresponding to the child's text data, determines the corresponding target learning section based on the parent feedback data, generates corresponding learning audio based on the target learning section, outputs the learning audio to the child, and sends the learning audio to the parent's device in text form;

[0133] After the conversation is over, the parent feedback module 220 collects the learning progress of the target learning section and obtains the scoring data of the child's learning target learning section, and is used to feed back the learning progress and scoring data to the parent device.

[0134] In an exemplary embodiment, the apparatus further includes, but is not limited to:

[0135] A children's audio feature extraction module preprocesses the children's audio data and extracts corresponding children's audio features from the children's audio data;

[0136] A wake-up word detection model retrieval module is used to retrieve a pre-trained wake-up word detection model, wherein the corresponding wake-up word dataset and non-wake-up word dataset are obtained, and the wake-up word dataset and non-wake-up word dataset are input into the training model for training until the training model converges, thereby determining the corresponding wake-up word detection model;

[0137] The wake-up word judgment module inputs the child's audio features into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

[0138] In an exemplary embodiment, the apparatus further includes, but is not limited to:

[0139] A timbre data retrieval module, used to retrieve a pre-recorded timbre data set;

[0140] a target timbre data determination module that plays the corresponding basic speech in the timbre data set to the child and determines the corresponding target timbre data based on the child's selection;

[0141] The process of recording timbre data is as follows:

[0142] An input audio data acquisition module is used to obtain the input audio data of the character's timbre, automatically segment the input audio data into several audio segments based on the speech segmentation model, and store them into audio files respectively;

[0143] The audio timbre feature storage module uses the timbre conversion model to extract features from each audio file to obtain the corresponding audio timbre features, which can be stored locally and correspond the audio timbre features to the characters.

[0144] In an exemplary embodiment, the apparatus further includes, but is not limited to:

[0145] The module for determining the appropriate learning section for a parent is used to determine the appropriate learning section based on the child's age input after the parent logs in for the first time.

[0146] The target schedule determination module sends the plan schedule corresponding to the age-appropriate learning section to the parent device end, and determines the corresponding target schedule based on the selection information of the parent device end;

[0147] The visual data sending module will determine the learning progress corresponding to the target learning section after the child's conversation ends, and provide visual data based on the learning progress and target progress table for sending to the parent's device.

[0148] In an exemplary embodiment, the apparatus further includes, but is not limited to:

[0149] The associated content playback module, if there is a long-term goal, generates corresponding associated content for playback based on the long-term goal. The associated content is the conversation content close to the long-term goal. The conversation direction and content are guided according to the long-term goal.

[0150] The current long-term goal determination module, if there are multiple long-term goals, plays the content of each long-term goal to the child and determines the corresponding long-term goal based on the child's feedback;

[0151] The playback content selection module is used to select playback content based on the long-term goal.

[0152] In an exemplary embodiment, the apparatus further includes, but is not limited to:

[0153] A target filter word determination module is used to determine corresponding target filter words based on parent feedback data;

[0154] The vocabulary setting module is used to set corresponding blocked words and restricted words based on the target filter words;

[0155] The vocabulary conversion module obtains the number of restricted words that appear during the conversation. If the number of restricted words is greater than the restricted word threshold, it is used to convert the restricted words into blocked words.

[0156] It should be noted that the large-model-based children's dialogue device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when conducting a large-model-based children's dialogue. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the large-model-based children's dialogue device will be divided into different functional modules to complete all or part of the functions described above.

[0157] In addition, the large-model-based children's dialogue device and the large-model-based children's dialogue method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs operations has been described in detail in the method embodiment and will not be repeated here.

[0158] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the large model-based children's dialogue method as described above.

[0159] In addition, an embodiment of the present application provides a storage medium on which computer-readable instructions are stored. The computer-readable instructions are executed by one or more processors to implement the large model-based children's dialogue method as described above.

[0160] A computer program product is provided in an embodiment of the present application. The computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the large model-based children's dialogue method as described above.

[0161] Compared with the related technology, in the above technical solution, the system first obtains the corresponding children's audio data of the children, and then parses the children's audio data and converts it into corresponding children's text data. The children's text data contains the knowledge sections involved in the children's audio. It is through the existing common intention recognition model that the knowledge sections that the children may design when they just talk can be parsed, and then the children's text data can be transmitted to the parent device end, so that the knowledge sections that the children may involve can be notified to the parents as soon as possible; after the parents receive the corresponding information, they can make a selection on the software on the parent device end, and can determine whether to conduct knowledge learning of the relevant sections this time. After obtaining the corresponding parent feedback data of the parents, the corresponding target learning section can be determined according to the parent feedback data. The target learning section here can be a previously scheduled learning goal or a learning section currently arranged by the parents. Then the system can call up the learning audio related to the target learning section and output the learning audio to the children, thereby guiding the children to conduct corresponding learning.

[0162] In addition, after each conversation with the child, the system can collect the learning progress of the target learning section and score it based on the conversation process. That is, it can obtain the scoring data of the child's target learning section, and feed back the learning progress and the scoring data to the parent device. Parents can check the child's relative learning situation from the parent device.

[0163] Through the above process, the system increases the participation of parents in the dialogue process with children. Parents can adjust the knowledge that needs to be learned according to the actual situation of the children. Parents not only participate in it, but also participate in the result generation strategy of each round of conversation during the dialogue between the system and the children. After each conversation, they can understand the learning situation of the children at each time at the first time, thereby meeting the diverse needs of parents.

[0164] In addition, by setting restricted words and blocked words, parents can monitor their children's online communication content in a personalized and precise manner, enhance the flexibility and effectiveness of information filtering, promote positive and healthy communication habits, and deepen parents' trust in the system and the system's own adaptability.

[0165] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0166] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A children's dialogue method based on a large model, characterized in that: include: Acquire child audio data corresponding to the child, parse the child audio data and convert it into corresponding child text data, and transmit the child text data to the parent's device; After the parent logs in for the first time, the corresponding age-appropriate learning section is determined based on the child's age input; Sending the planned schedule corresponding to the age-appropriate learning section to the parent's device, and determining the corresponding target schedule based on the selection information of the parent's device; After the child's conversation is over, the learning progress corresponding to the target learning section will be determined, and visual data based on the learning progress and the target progress chart will be provided and sent to the parent's device; Obtaining parent feedback data corresponding to the child's text data, determining a corresponding target learning section based on the parent feedback data, generating corresponding learning audio based on the target learning section by the large model, outputting the learning audio to the child, and sending the learning audio to the parent's device in text form; When a long-term goal is selected, after the child initiates a conversation, the following steps are performed: if there is a long-term goal, the large model generates corresponding related content based on the long-term goal for playback, wherein the related content is the conversation content close to the long-term goal, and the conversation direction and content are guided according to the long-term goal; if there are multiple long-term goals, the content of each long-term goal is played to the child, and the corresponding long-term goal is determined based on the child's feedback; and the playback content is selected based on the long-term goal; After the conversation is over, the learning progress of the target learning section is collected, and the scoring data of the child's learning target learning section is obtained, and the learning progress and the scoring data are fed back to the parent's device.

2. The method according to claim 1, wherein After obtaining the child audio data corresponding to the child, the method further includes: Preprocessing the children's audio data and extracting corresponding children's audio features from the children's audio data; Retrieve a pre-trained wake-up word detection model, wherein the corresponding wake-up word dataset and non-wake-up word dataset are obtained, the wake-up word dataset and non-wake-up word dataset are input into the training model for training until the training model converges, and the corresponding wake-up word detection model is determined; The child's audio features are input into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

3. The method according to claim 2, wherein When outputting learning audio to children, it is necessary to select the timbre of the output audio. The method further includes: Retrieve a pre-recorded sound data set; Playing the corresponding basic speech in the timbre data set to the child, and determining the corresponding target timbre data based on the child's selection; The process of recording timbre data is as follows: Acquire recorded audio data of the character's timbre, automatically segment the recorded audio data into several audio segments based on a speech segmentation model, and store them in audio files; A timbre conversion model is used to perform feature extraction on each of the audio files to obtain corresponding audio timbre features, the audio timbre features are stored locally, and the audio timbre features are associated with the characters.

4. The method according to claim 1, wherein After obtaining parent feedback data corresponding to the child's text data, the method further includes: Determining corresponding target filtering words based on the parent feedback data; Set corresponding blocked words and restricted words based on the target filtering words; During the conversation, the number of times the restrictive word appears is obtained. If the number of times the restrictive word appears is greater than a restrictive word threshold, the restrictive word is converted into a blocked word.

5. A device for executing the large model-based children's dialogue method according to any one of claims 1 to 4, characterized in that: include: A child text data acquisition module, which acquires the child's audio data corresponding to the child, parses the child's audio data and converts it into corresponding child text data, and transmits the child text data to the parent's device, wherein the child text data includes the knowledge section related to the child's audio; A target learning section determination module obtains parent feedback data corresponding to the child's text data, determines the corresponding target learning section based on the parent feedback data, generates corresponding learning audio based on the target learning section, outputs the learning audio to the child, and sends the learning audio to the parent's device in text form; The parent feedback module collects the learning progress of the target learning section and obtains the scoring data of the child's learning target learning section after the conversation ends, and is used to feed back the learning progress and the scoring data to the parent device.

6. The device according to claim 5, characterized in that The device further comprises: a child audio feature extraction module, which preprocesses the child audio data and extracts corresponding child audio features from the child audio data; A wake-up word detection model retrieval module is used to retrieve a pre-trained wake-up word detection model, wherein the corresponding wake-up word dataset and non-wake-up word dataset are obtained, and the wake-up word dataset and non-wake-up word dataset are input into the training model for training until the training model converges, thereby determining the corresponding wake-up word detection model; The wake-up word judgment module inputs the child's audio features into the wake-up word detection model to determine whether the wake-up word is detected. If detected, the conversation mode is turned on.

7. An electronic device, characterized in that: include: at least one processor and at least one memory, wherein: The memory has computer-readable instructions stored thereon; The computer-readable instructions are executed by one or more processors to enable the electronic device to implement the large model-based children's dialogue method according to any one of claims 1 to 4.

8. A storage medium having computer-readable instructions stored thereon, characterized in that: The computer-readable instructions are executed by one or more processors to implement the large model-based children's dialogue method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Guided scene dialogue method and system

    CN110399471A

  • Dialogue generation method, apparatus and device, and computer readable medium

    CN117972043A

  • Programmable language teaching system and method

    US20180122266A1