Speech teaching method and device, storage medium and terminal

By acquiring and analyzing target speech topics through a large-scale speech teaching model, personalized speech drafts and examples are generated, which solves the problem of lack of systematicness and personalization in traditional teaching methods and improves children's speech skills and expressiveness.

CN120808645APending Publication Date: 2025-10-17SHENZHEN SANLIJIE INTELLIGENT MANUFACTURING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510913062.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing speech teaching methods lack systematic and personalized guidance, which limits the improvement of children's speech skills. Traditional classroom teaching and parental guidance lack objectivity and accuracy.

Method used

The system uses a large-scale speech teaching model to identify target speech topics, conduct topic analysis, generate personalized speech drafts and speech examples, provide multimodal display and evaluation optimization functions, and support voice and body language guidance.

Benefits of technology

It provides a personalized speech teaching experience, enhances children's speech skills and personal expression, reduces writing pressure, improves the efficiency and quality of speech writing, and provides rich learning resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808645A_ABST
    Figure CN120808645A_ABST
Patent Text Reader

Abstract

The invention discloses a speech teaching method and device, a storage medium and a terminal, and the method comprises the steps: obtaining a target speech theme, carrying out the theme analysis of the target speech theme through a speech teaching large model, and obtaining the theme type of the target speech theme; generating a speech draft corresponding to the target speech theme according to the theme type, and generating a speech example corresponding to the target speech theme based on the theme type; and displaying the speech and the speech example to the user. A target speech theme is deeply analyzed through a speech teaching large model, the theme type can be accurately positioned, personalized speech draft writing guidance and speech skill teaching are further obtained and visually displayed accordingly, a user can more clearly understand how to build a speech framework, organize languages and apply speech skill, and the speech teaching efficiency is improved. Therefore, speech skills and personal expressive force are effectively improved, and personalized speech teaching experience is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a speech teaching method and device, a storage medium, and a terminal. BACKGROUND

[0002] In the current educational environment, the cultivation of children's speech ability mainly depends on traditional classroom teaching and real-time guidance by parents. This traditional speech teaching method often lacks systematic and personalized guidance, and the limitations of teaching resources and teaching content also limit the improvement and development of children's speech thinking and expression skills to some extent. SUMMARY

[0003] The present application provides a speech teaching method, device, storage medium, and terminal to solve the technical problems of insufficient flexibility and strong subjectivity of existing speech teaching methods.

[0004] In a first aspect, an embodiment of the present application provides a speech teaching method, which comprises:

[0005] obtaining a target speech topic, performing topic analysis on the target speech topic through a speech teaching large model to obtain a topic type of the target speech topic;

[0006] generating a speech script corresponding to the target speech topic according to the topic type, and generating a speech example corresponding to the target speech topic based on the topic type;

[0007] displaying the speech script and the speech example to a user.

[0008] In a second aspect, an embodiment of the present application provides a speech teaching device, which comprises:

[0009] a topic analysis module configured to obtain a target speech topic, perform topic analysis on the target speech topic through a speech teaching large model, and obtain a topic type of the target speech topic;

[0010] a content generation module configured to generate a speech script corresponding to the target speech topic according to the topic type, and generate a speech example corresponding to the target speech topic based on the topic type;

[0011] a content display module configured to display the speech script and the speech example to a user.

[0012] In a third aspect, an embodiment of the present application provides a computer storage medium, which stores a plurality of instructions, and the instructions are adapted to be loaded by a processor and execute the steps of the above method.

[0013] In a fourth aspect, an embodiment of the present application provides a terminal, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the computer program is adapted to be loaded by the processor and execute the steps of the method described above.

[0014] The technical solutions provided by some embodiments of the present application have at least the following beneficial effects:

[0015] The present application provides a speech teaching method, obtaining a target speech topic, performing topic analysis on the target speech topic through a speech teaching large model to obtain a topic type of the target speech topic; generating a speech script corresponding to the target speech topic according to the topic type, and generating a speech example corresponding to the target speech topic based on the topic type; and displaying the speech script and the speech example to a user. First, the target speech topic of the user is obtained, and in-depth analysis is performed on the target speech topic, which can accurately identify and classify the core connotation and characteristics of the target speech topic, and clearly define the specific content or direction of the speech, thereby providing more targeted reference basis for subsequent writing of the speech script and generation of the speech example; next, a speech script with clear structure and coherent logic is automatically generated according to the topic type of the target speech topic, so that the content of the speech script not only meets the theme requirements, but also has certain attractiveness and appeal, which not only relieves the pressure of the user in writing the speech script, but also facilitates the user to quickly grasp the core points of the speech by referring to the speech script, thereby improving the efficiency and quality of writing the speech script; at the same time, the speech example matched with the target speech topic is provided based on the topic type, which can enable the user to quickly and accurately master the correct speech method and skill through intuitive imitation and learning, thereby improving the expressiveness and appeal of the user's speech; finally, the generated speech script and speech example are intuitively displayed to the user, which provides rich learning resources for the user to refer to and learn at any time, so as to help them continuously improve and perfect their speech skills. Through the speech teaching method of the present application, the user can more clearly understand how to organize language and use speech skills, thereby effectively improving the speech skills and personal expressiveness, and realizing personalized speech teaching experience. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 An exemplary system architecture diagram of a speech teaching method provided by an embodiment of the present application is shown in the figure.

[0018] Figure 2A flowchart of a speech teaching method provided by an embodiment of the present application is shown in FIG. 1.

[0019] Figure 3 A flowchart of a speech teaching method provided by an embodiment of the present application is shown in FIG. 1.

[0020] Figure 4 A flowchart of a speech teaching method provided by an embodiment of the present application is shown in FIG. 1.

[0021] Figure 5 A flowchart of a model training method of a speech teaching large model provided by an embodiment of the present application is shown in FIG. 1.

[0022] Figure 6 A structural block diagram of a speech teaching device provided by an embodiment of the present application is shown in FIG. 1.

[0023] Figure 7 A structural diagram of a terminal provided by an embodiment of the present application is shown in FIG. 1. DETAILED DESCRIPTION

[0024] In order to make the features and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0025] The following description refers to the accompanying drawings. Unless otherwise indicated, same or similar elements in different drawings are denoted by the same reference numerals. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0026] The speech teaching in traditional classrooms often adopts a "one-size-fits-all" teaching mode, which is difficult to provide targeted teaching according to the differences of each child in speech skills, content organization, voice tone, etc. Moreover, due to the limitation of teaching resources, teachers may not be able to provide students with the whole process from topic mining, content construction to speech skill training, which means that children may only have fragmented knowledge, but not a complete speech skill system, thereby limiting the development of children's speech skills. Although parents can provide real-time guidance to some extent, their guidance is often based on personal subjective feelings and lacks objectivity and accuracy due to the lack of professional speech teaching knowledge and experience, which makes it difficult to help children improve their overall speech ability.

[0027] Therefore, the embodiment of the present application provides a speech teaching method to solve the technical problems of insufficient flexibility and strong subjectivity of the existing speech teaching method.

[0028] Please refer to Figure 1 , Figure 1 An exemplary system architecture diagram of a speech teaching method provided by the embodiment of the present application is shown.

[0029] As Figure 1 shown, the system architecture can include a terminal 101, a network 102 and a server 103. The network 102 is used to provide a communication link medium between the terminal 101 and the server 103. The network 102 can include various types of wired communication links or wireless communication links, for example, the wired communication links include optical fiber, twisted pair or coaxial cable, and the wireless communication links include Bluetooth communication link, Wireless-Fidelity (Wi-Fi) communication link or microwave communication link, etc.

[0030] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103, or the terminal 101 can interact with the server 103 through the network 102 to receive messages or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various electronic devices, including but not limited to smart watches, smart phones, tablet computers, laptop computers and desktop computers, etc. When the terminal 101 is software, it can be installed in the above-mentioned electronic devices, which can be implemented as multiple software or software modules (for example, used to provide distributed services) or a single software or software module, which is not specifically limited here.

[0031] In the embodiment of the present application, the terminal 101 first acquires a target speech topic, performs topic analysis on the target speech topic through a speech teaching large model to obtain a topic type of the target speech topic; then the terminal 101 generates a speech script corresponding to the target speech topic according to the topic type, and generates a speech example corresponding to the target speech topic based on the topic type; finally, the terminal 101 displays the speech script and the speech example to the user.

[0032] The server 103 can be a service server providing various services. It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or as a single software or software module, which is not specifically limited here.

[0033] Or, the system architecture can also not include the server 103, in other words, the server 103 can be an optional device in the embodiments of the present specification, that is, the method provided in the embodiments of the present specification can be applied to a system structure including only the terminal 101, and the embodiments of the present application do not limit this.

[0034] It should be understood that Figure 1 The number of terminals, networks and servers in the system architecture is only illustrative, and can be any number of terminals, networks and servers according to the needs of implementation.

[0035] Please refer to Figure 2 , Figure 2 A flowchart of a speech teaching method provided by the embodiments of the present application. The execution subject of the embodiments of the present application can be a terminal executing speech teaching, or a processor in the terminal executing the speech teaching method, or a speech teaching service in the terminal executing the speech teaching method. For convenience of description, the specific execution process of the speech teaching method will be introduced below taking the execution subject as the processor in the terminal.

[0036] As Figure 2 indicated, the speech teaching method can at least include:

[0037] S202, obtaining a target speech topic, performing topic analysis on the target speech topic by a speech teaching large model to obtain a topic type of the target speech topic.

[0038] Optionally, in the scenario of speech teaching, different topic types (such as personal experience, knowledge popularization, emotional expression, etc.) are suitable for different content structures and expression methods. For example, the personal experience type topic of "My pet" may pay more attention to emotional expression and personal story, and usually includes structures such as background introduction, event description and emotional summary; while the knowledge popularization type topic such as "The mystery of the universe" emphasizes the accuracy and logicality of information, and needs clear definition, explanation and supporting materials. Therefore, before the speech teaching of the target speech topic, the target speech topic needs to be analyzed first to determine the core content and key points of this speech, and to provide more targeted teaching suggestions for subsequent speech writing and skill teaching.

[0039] Optionally, a simple and easy-to-use input box is set on the terminal interface (such as a child watch interface) or related application interface running the method of the present application, for users (hereinafter referred to as "children" as an example of users) to input their target speech topics of interest through voice or text. Some suggestive icons or words can be attached next to the input box to help children understand which types of topics can be input, such as "my pet", "unforgettable travel experience", etc. After the children input the topic in the input box, a preliminary verification is first performed to ensure that the input content meets the basic requirements (such as non-empty, moderate length, no sensitive words, etc.), and if the input does not meet the basic requirements, a prompt box will pop up to guide the children to re-input. Further, data cleaning operation and standardization processing are performed on the input data, such as unifying case, correcting spelling errors, removing irrelevant characters and noise, etc., to reduce the interference on subsequent topic analysis.

[0040] Optionally, the terminal built-in or bound with the speech teaching large model running the method of the present application, which has been pre-trained with a large number of speech topics and speech texts, can accurately understand and analyze the input target speech topic based on deep learning algorithm. Once the target speech topic input by the children is verified, the method immediately calls the speech teaching large model to perform multi-dimensional analysis on the input target speech topic to identify the topic type of the target speech topic. These dimensions can include but are not limited to the following aspects: emotional tendency (such as positive, negative, neutral), content type (such as personal experience, knowledge sharing, emotional expression, etc.), audience group (such as children, teenagers, adults, etc.). For example, if the topic input by the children is "my pet", the speech teaching large model will first identify that it is a positive topic biased towards personal experience and emotional expression, and further determine its general direction as describing the appearance, personality and interesting stories of the pet, etc. to more accurately guide the subsequent speech writing and example generation. Through topic analysis, the model can accurately identify the topic type of the target speech topic, and can choose to display the identified topic type to the children in a clear and understandable way, which can include textual description, icon identification or simple animation effect, to help children quickly understand the topic type.

[0041] S204, generating a speech script corresponding to the target speech topic according to the topic type, and generating a speech example corresponding to the target speech topic based on the topic type.

[0042] Optionally, different target speech topics have different connotations and emphases, and children may not be able to closely focus on the topic in the process of writing a speech, resulting in scattered content or even deviation from the target speech topic. Therefore, in order to ensure that the speech content closely revolves around the target speech topic and avoid deviation from the topic or empty content, the speech teaching large model can further generate a corresponding speech according to the determined topic type, helping children to clearly define the core content and direction of the speech, and making the speech more targeted. Specifically, when generating a speech corresponding to the target speech topic, the structure framework of the speech can be first sorted for the children according to the topic type, for example, the structure can include introduction, expansion, summary and other parts; then relevant material suggestions are provided for the structure framework to help children enrich the content of the speech; finally, the structure framework and the corresponding material suggestions are integrated into a complete speech.

[0043] Optionally, different topic types of speeches have different language styles and expression requirements. For example, a personal experience type of speech such as "My pet" requires warm and kind language, while a science popularization type of speech such as "The mystery of the universe" requires strict and professional expression. Different children have different ways of understanding and expressing the same speech, and if the specific expression of the speech is not intuitively displayed, children may have difficulty grasping the expression elements such as speech tone, speed, rhythm and emotional connotation, affecting the final speech effect. Based on this, after generating a speech corresponding to the target speech topic, high-quality speech examples can be generated based on the topic type using speech synthesis technology (for example, Text-to-Speech, TTS). These speech examples have clear pronunciation, natural tone, and appropriate speed, providing clear and accurate pronunciation and tone demonstrations for children, so that they can learn and imitate the speech examples more quickly to master the speech skills and improve their speech ability. In addition to speech examples, the body language during the speech can also be displayed in the form of animation or video, and children can follow the animation or video for imitation practice, and the system can capture the children's body movements through the camera for real-time evaluation and guidance.

[0044] S206, display the speech and speech examples to the user.

[0045] Optionally, a concise and intuitive multi-modal display interface is set up on the terminal running the method of the present application, which interface contains at least a speech text area, a speech example playing area and related operation buttons (such as play, pause, repeat play, etc.). Specifically, in the speech text area, multiple display modes are provided, such as normal text mode, segmented display mode, etc., to adapt to children of different age groups and reading habits; meanwhile, a display form of pictures and texts is also supported, and pictures or animations are inserted in key parts to help children better understand and remember the speech content, for example, a cute pet picture can be inserted when describing the appearance of a pet; in addition, scrolling, selecting and copying, voice playing, editing, etc. operations are also supported for children. In the speech example playing area, functions such as volume adjustment, progress bar dragging, speed play, etc. are provided to meet the user's needs in different scenarios; and subtitles can be displayed to synchronize the display of the text content of the speech example, helping children better understand the content and emotion expressed by voice.

[0046] Optionally, in order to help children understand their learning progress and effect, the method of the present application will also automatically save the learning records of children, including the target speech topics learned, the number of times of browsing the speech, the number of times of playing the speech example, etc. Further, learning feedback reports can be generated regularly according to the learning records of children, and the report content can include the score situation of children in different dimensions (such as content integrity, voice performance, body language, etc.), progress trend and targeted improvement suggestions, etc. And when displaying the speech and the speech example, similar topics or styles of speech content can also be recommended in combination with the historical learning data and interest preferences of children, to increase the participation and interest of children. In addition, buttons such as collection, sharing and downloading are also set up on the terminal interface to facilitate children to save their favorite speeches or share the content with parents or teachers to get more feedback.

[0047] In the embodiment of the present application, a speech teaching method is provided. A target speech topic is obtained, the target speech topic is analyzed by a speech teaching large model to obtain a topic type of the target speech topic, a speech script corresponding to the target speech topic is generated according to the topic type, and a speech example corresponding to the target speech topic is generated based on the topic type. The speech script and the speech example are displayed to the user. First, the target speech topic of the user is obtained, and the target speech topic is analyzed in depth. The core connotation and characteristics of the target speech topic can be accurately identified and classified, and the specific content or direction of the speech is clear. Therefore, more targeted reference basis is provided for subsequent writing of the speech script and generation of the speech example. Next, the speech script with clear structure and logical coherence is automatically generated according to the topic type of the target speech topic. The content of the speech script not only meets the theme requirements, but also has certain attraction and appeal. This not only reduces the pressure of the user to write the speech script, but also facilitates the user to quickly grasp the core points of the speech by referring to the speech script, thereby improving the efficiency and quality of writing the speech script. At the same time, the speech example matched with the target speech topic is provided based on the topic type. The user can quickly and accurately master the correct speech method and skill through intuitive imitation and learning, thereby improving the expressiveness and appeal of the user's speech. Finally, the generated speech script and speech example are intuitively displayed to the user. This provides rich learning resources for the user to check and learn at any time, so as to help them continuously improve and perfect their speech skills. Through the speech teaching method of the present application, the user can more clearly understand how to organize language and use speech skills, thereby effectively improving the speech skills and personal expressiveness, and realizing personalized speech teaching experience.

[0048] Please refer to Figure 3 , Figure 3 A flowchart of a speech teaching method provided in the embodiment of the present application is shown.

[0049] As Figure 3 shown, the speech teaching method can at least include:

[0050] In S302, a target speech topic is obtained, and a speech teaching large model is controlled to analyze the target speech topic from a first preset dimension to obtain a topic type of the target speech topic. The first preset dimension includes at least one of semantic understanding, keyword extraction, and topic classification.

[0051] Optionally, after obtaining the target speech topic input by the child, the speech teaching large model uses advanced natural language processing technology (NLP) to conduct in-depth topic analysis on the target speech topic, such as semantic understanding, keyword extraction, topic classification, etc. Then the speech teaching large model fuses the results of semantic understanding, keyword extraction and topic classification to form a comprehensive topic analysis result, providing decision support for subsequent speech script generation and speech example making.

[0052] Specifically, first, semantic understanding is performed on the target speech topic. By analyzing the words, phrases and their mutual relationships in the target speech topic, the model can accurately grasp the core meaning and potential implications of the topic, avoiding the generated speech script and speech examples deviating from the topic in the subsequent process. In this process, the model not only considers the literal expression of the target speech topic itself, but also combines contextual information (such as the child's previous speech history, interests and hobbies, etc.) to more comprehensively grasp the theme intention and ensure that the speech script and speech examples are in line with the actual needs of the child.

[0053] Optionally, in order to determine the focus and structure of the speech script, the model also performs keyword extraction. Specifically, the model automatically extracts key information points such as theme words, core concepts, important events, etc. from the target speech topic. In addition, the model can assign different weights to each keyword according to its importance and frequency of appearance in the target speech topic, and place the key content in a more prominent position according to their respective weights in the subsequent speech script generation and speech example process. For example, keywords with high weights can be used as main chapter or paragraph titles in the speech script, while keywords with lower weights can be used as auxiliary content or details.

[0054] Optionally, when writing the speech script, target speech topics with similar theme types may use the same or similar materials and cases, so in order to facilitate the selection of material content closely related to the target speech topic, the speech teaching large model also performs topic classification on the target speech topic when performing topic analysis. Specifically, the model uses a multi-dimensional classification system to classify the target speech topic into corresponding categories, which may include content types (such as personal experiences, social phenomena, technological development, etc.), emotional tendencies (such as positive, negative, neutral, etc.), audience groups (such as children, teenagers, adults, etc.), etc. These classification results will be used for subsequent teaching strategy development. For example, for different age groups of the audience, the model can adjust the language style and content depth of the speech script; for different emotional tendency topics, the model can recommend appropriate speech skills and expression methods.

[0055] S304. According to the subject type, the structure framework of the speech script corresponding to the target speech subject is combed; the material suggestion corresponding to the structure framework is generated based on the subject type; and the speech script corresponding to the target speech subject is generated according to the structure framework and the material suggestion.

[0056] Optionally, in the process of generating the speech script corresponding to the target speech subject according to the subject type, the speech teaching large model first combs the personalized speech script structure framework for the child according to the analyzed subject type. Specifically, the model has multiple types of structure framework templates built-in, which can automatically match and adjust the structure framework template according to the subject type, ensuring that the structure is reasonable and consistent with the subject characteristics. For example, for personal experience type subjects, the framework may include an introduction (introducing through specific events), a main body (detailed description of experiences, feelings, growth, etc.), and a conclusion (summarizing the enlightenment or insight brought by the experience); for science and technology development type subjects, it may focus more on the current situation introduction, development trend analysis, personal opinions, etc.

[0057] Illustratively, taking the target speech subject of "My pet" as an example, the model can guide the child how to start the speech by catching the audience's attention with interesting pet introduction, how to expand in the middle part from the appearance, personality, interesting stories of the pet, how to summarize the feelings for the pet and sublimate the theme in the end, etc. And an example structure framework is displayed in the interface, such as: beginning - everyone, today I want to share a story about my cat; middle - first, let me describe its appearance; end - in short, my cat is not only my friend, but also an indispensable part of my life.

[0058] Optionally, based on the subject type and the structure framework, the speech teaching large model will provide detailed material suggestions closely related to the content of the speech script, which include but are not limited to stories, cases, data, famous quotes, etc., aiming to help children enrich the content of the speech script. Specifically, the model relies on a child-specific corpus, quickly filters out content materials that match the subject type and structure framework through intelligent search and matching technology. At the same time, considering the age characteristics and interest preferences of children, the materials are appropriately adjusted and optimized to ensure that the materials have both educational significance and can stimulate the interest of children.

[0059] Illustratively, when describing the appearance of the pet, the model can prompt the child to start with details such as hair color, eye shape, etc.; when telling the pet story, the model suggests that the child describe the cause, process and result of the event, etc. And some specific demonstration material suggestions are displayed in the interface for the child to refer to, such as: "My cat has a pair of blue eyes, like two shining gems. It also has soft white fur, which feels very comfortable to touch."

[0060] Optionally, the speech teaching model will automatically generate a complete speech text corresponding to the target speech topic based on the selected structural framework and the provided material suggestions, so that children can further edit and improve it based on the generated text.

[0061] S306: Determine the speech features corresponding to the speech manuscript, and generate a speech voice example of the speech manuscript based on the speech features, where the speech features include at least one of voice style, pronunciation, intonation, and speech speed.

[0062] Optionally, the speech teaching model determines the speech features corresponding to the target speech topic, the content structure of the speech, and the expected emotional expression, and generates speech examples based on these features. These features include, but are not limited to, voice style, pronunciation accuracy, intonation, and speaking speed.

[0063] Specifically, first, the model has multiple built-in voice style templates (such as lively and cheerful, calm and professional, and humorous and witty). It automatically matches the appropriate voice style based on the topic type or allows children to manually select the appropriate voice style, ensuring that the speech examples meet the specific topic requirements. For example, for a personal experience topic like "My Pet," the model would select a lively and cheerful voice style; while for a popular science topic like "The Mysteries of the Universe," it would choose a calm and professional voice style. Second, the model also uses high-precision speech synthesis technology to ensure that the generated speech examples have clear and accurate pronunciation, providing a good model for children to imitate. For example, when describing a pet's appearance, the character "rong" in "furry little ball" is pronounced accurately and clearly. Third, the model can adjust the tone of voice appropriately based on the topic content, enhancing the emotional expression and making the speech more appealing. For example, when telling a story about a pet's mischief, the tone of voice can be more lively and cheerful to convey the fun of the story. Fourth, the speaking speed is adjusted appropriately according to the characteristics of the topic and the needs of the audience to ensure the appropriate rhythm of the speech. For example, for younger children, generate examples at a slower speech rate so that they can more easily understand and imitate; for older children, you can choose a normal speech rate.

[0064] Optionally, after determining the speech features, the model uses text-to-speech (TTS) technology to convert the speech script into a speech sample with the aforementioned speech features. The generated results can optionally be post-processed, such as by adding background music and sound effects, to enhance the appeal and expressiveness of the speech sample. Furthermore, after the speech sample is generated, it undergoes quality assessment and optimization. Through automated evaluation or manual review, the clarity, fluency, and emotional accuracy of the speech sample are checked, and necessary adjustments are made based on the feedback.

[0065] S308, display the speech script to the user, and play the speech voice sample to the user according to the first playing rule.

[0066] Optionally, the generated speech script is displayed to the child, and during the display process, the speech voice sample can be played to the child according to the preset first playing rule, such as loop playing, sentence-by-sentence playing, segment-by-segment playing, etc., and the child is allowed to control the playing of the voice sample at any time through simple operations (such as clicking a button, sliding a progress bar, etc.), such as pausing, replaying, or adjusting the playing speed, volume, etc., so as to better learn and imitate.

[0067] In the embodiments of the present application, a speech teaching method is provided, which can more accurately grasp the core connotation and type of the theme by controlling the speech teaching large model to deeply analyze the target speech theme from the first preset dimensions such as semantic understanding, keyword extraction, and theme classification, helping children to clearly understand the direction and key points of the speech, and providing a solid foundation for subsequent speech script writing and sample generation; further, the speech script structure framework is sorted according to the theme type, corresponding material suggestions are generated, and finally a complete speech script is generated, providing systematic and personalized speech content support for children, so that children can quickly build high-quality speech content and further modify and improve it on this basis; by determining the corresponding voice characteristics of the speech script and generating speech voice samples based on these characteristics, and then displaying them to the child according to the preset rules, a high-quality imitation object is provided for the child, which can help them master the correct pronunciation, intonation, and speech speed and other key speech skills, and improve the voice expressiveness and speech appeal.

[0068] Please refer to Figure 4 , Figure 4 A flowchart of a speech teaching method provided in the embodiments of the present application.

[0069] As Figure 4 shown, the speech teaching method can at least include:

[0070] S402, obtaining a target speech theme, and performing theme analysis on the target speech theme by a speech teaching large model to obtain a theme type of the target speech theme.

[0071] Optionally, for step S402, please refer to the detailed description in step S202, which will not be repeated here.

[0072] S404, in response to an evaluation optimization request for an initial speech script corresponding to the target speech theme, evaluating the initial speech script from a second preset dimension to obtain an evaluation result of the initial speech script, the second preset dimension including at least one of grammatical correctness, vocabulary richness, and semantic coherence.

[0073] Optionally, the child can write an initial speech script on the target speech topic by himself / herself, at this time the method of the embodiment of the application provides an evaluation optimization function for the initial speech script. A special button or menu item is set on the smart terminal (such as a child watch), and the child or his / her parents, teachers can click and submit the initial speech script, select the "evaluation optimization" function to enter the evaluation and optimization process. In this way, the speech teaching large model will comprehensively evaluate the initial speech script according to the preset second preset dimension, which includes but is not limited to grammatical correctness, vocabulary richness, semantic coherence, etc.

[0074] Specifically, first, the speech teaching large model will use NLP technology, combined with a large number of grammar rule libraries, to perform syntax analysis on the initial speech script, check whether there are syntax errors such as tense errors, subject-verb inconsistencies, etc., and ensure that the expression is accurate and correct. Second, the model will also evaluate the use of vocabulary in the initial speech script according to the vocabulary library and context relationship, analyze whether there are repetitive or monotonous words in the initial speech script, and encourage the use of more vivid and specific words to enrich language expression. For example, for the repeated "lovely", it is suggested to replace it with synonyms such as "fascinating" and "charming" to increase vocabulary diversity; it can also prompt the child to change "my cat is fat" to "my cat is like a fluffy little ball, with a chubby appearance that is very lovable". Third, the model will check the overall structure of the initial speech script and the logical coherence between paragraphs to ensure that the content is clear and logical. For example, if it is found that there are no transition sentences between paragraphs, it can be suggested to add a connecting sentence such as "Next, let's take a look at its mischievous side!"

[0075] Optionally, the speech teaching large model will display the above evaluation results to the child in the form of a visual report, pointing out the problems and deficiencies in the initial speech script and providing corresponding improvement suggestions to facilitate the child's understanding and modification. In addition, the child can further edit based on the optimization suggestions provided by the model, and the model will continuously update the evaluation results according to the child's modifications and allow the child to view the real-time optimization suggestions given by the model at any time. For example, the child deletes a paragraph during editing, and the model will re-evaluate the coherence of the remaining part and give new evaluation results and optimization suggestions.

[0076] S406, optimizing the initial speech script based on the evaluation results, and displaying the optimized initial speech script to the user.

[0077] Optionally, in order to improve the convenience and flexibility of the speech teaching method, the speech teaching large model can automatically modify the initial speech script according to the evaluation results using language generation technology. The optimization content includes correcting grammatical errors, replacing more rich vocabulary, adjusting sentence structure to enhance semantic coherence, etc. Then the optimized initial speech script is displayed to the children for review and confirmation. Children can continue to make evaluation optimization requests as needed, and perform multiple iterations of optimization until a satisfactory result is achieved. This method also displays the optimized speech script through the user interface, providing functions such as saving, exporting, printing, etc. to facilitate subsequent use by children.

[0078] S408, generating a speech script corresponding to the target speech theme according to the theme type, and generating a speech example corresponding to the target speech theme based on the theme type.

[0079] Optionally, for step S408, please refer to the detailed description in step S204, which will not be repeated here.

[0080] S410, determining the body characteristics corresponding to the speech script, and generating a speech animation example of the speech script based on the body characteristics, the body characteristics including at least one of gestures, body, eye contact; playing the speech animation example to the user according to the second playing rule.

[0081] Optionally, in the speech teaching method, in order to further improve the speech expressiveness and appeal of children, the method of the present application embodiment also provides a speech animation example generation and display function based on body characteristics. The core of this function is to analyze the content of the speech script through intelligent algorithm, determine the matching body characteristics including gestures, body posture, eye contact, etc., and then generate vivid and vivid speech animation examples to help children better understand and imitate professional speech body language.

[0082] Optionally, the speech teaching large model uses NLP technology to deeply analyze the generated speech script, identifies key paragraphs, emotional expression points, and information points that need to be emphasized, and then marks the speech segments that need to be accompanied by specific body language according to the analysis results. Further, combined with the experience in the field of speech and the cognitive characteristics of children, suitable body feature templates are matched for different types of speech content.

[0083] Specifically, the model features a built-in library of gestures, automatically recommending appropriate gestures based on the emotional expression of the speech. For example, when describing a pet's adorable movements, it recommends a gentle petting gesture; when recounting a funny event, it recommends incorporating lively gestures. Secondly, the model also recommends appropriate standing postures and movements based on the emotion and style of the speech, ensuring the speaker's posture is natural and expressive, enhancing the speaker's appeal. For example, in formal science presentations, it's recommended to maintain an upright posture; while in casual sharing of personal experiences, it's allowed to move about to enhance approachability. Furthermore, the model can guide speakers on how to connect with the audience through eye contact, providing specific advice on how to allocate gaze and when to establish eye contact, thereby enhancing the interactivity and appeal of the speech. For example, during the opening remarks, it recommends alternating between short eye contact with different sections of the audience to capture attention; while maintaining steady eye contact when delivering key points enhances persuasiveness.

[0084] Furthermore, 3D animation technology is used to generate a corresponding speech animation example based on the aforementioned body features. The characters in the animation example have a child-friendly appearance, natural and smooth body movements, and are closely aligned with the content of the speech. A second playback rule is then set to play the speech animation example, such as automatically playing the corresponding animation example according to the paragraph division of the speech manuscript, or manually playing it when the child requests it. In order to enhance the children's learning effect of the animation example, functions such as slow motion, pause, and replay of the speech animation example are also provided to facilitate children to carefully observe and learn each body movement. While playing the animation example, voice commentary is provided to explain the meaning and function of each body movement to help children better understand. The model will also encourage children to imitate and practice after watching the animation example, and capture the children's body movements through the camera, compare them with the speech animation example, and provide real-time feedback and suggestions.

[0085] S412: Show the speech manuscript and speech examples to the user.

[0086] Optionally, regarding step S412, please refer to the detailed description in step S206, which will not be repeated here.

[0087] In the embodiment of the present application, a speech teaching method is provided, which evaluates the initial speech script from the dimensions of grammatical correctness, vocabulary richness, and semantic coherence, and optimizes based on the evaluation results, so as to improve the language quality and expression effect of the initial speech script, and make the speech content of children more accurate, vivid and logically clear; by determining the corresponding body features of the speech script and generating a speech animation example, the gestures, body posture and eye contact during the speech can be displayed in an intuitive way, which helps children better understand and imitate professional speaking skills, enhances the appeal and expressiveness of the speech, and improves the overall speech effect.

[0088] Please refer to Figure 5 , Figure 5 The flowchart of the model training method of the speech teaching large model provided in the embodiment of the present application is shown.

[0089] As Figure 5 shown, the model training method of the speech teaching large model can at least include:

[0090] S502, constructing an initial speech teaching large model for the speech teaching scene based on a basic multi-modal large model.

[0091] Optionally, since the speech teaching large model needs to be able to process speech information data from different sources and different types (such as the target speech theme input by children, the initial speech script, the speech voice content, etc.), a multi-modal fusion architecture needs to be used to construct the initial speech teaching large model, so that the model can more comprehensively understand the speech information and improve the accuracy of speech teaching. Based on this, an existing and verified basic multi-modal large model is selected as a starting point, which has the ability to output prediction results for prediction objects based on different types of features of the prediction objects, and an initial speech teaching large model for the speech teaching scene is constructed accordingly.

[0092] Optionally, in order to enhance the understanding ability and response speed of the model, preset technical tools are integrated in the initial speech teaching large model, such as large models with natural language processing technology, speech synthesis technology or intelligent evaluation technology and other technical capabilities. These large models are used to process and analyze speech teaching related tasks through their powerful text understanding and generation capabilities, for example, a speech evaluation model based on intelligent evaluation technology can comprehensively evaluate a child's speech in different dimensions, pointing out the performance in content integrity, logical clarity, pronunciation accuracy, speech speed, and giving corresponding scores and improvement suggestions. Based on this, the basic multi-modal large model is fused with the preset technical tools, and in the fusion process, the interface compatibility between the two models needs to be ensured, and efficient collaboration under a unified framework is required, which involves adjusting the model architecture, sharing intermediate layer feature representation, etc. Then according to the specific speech teaching needs, the parameters of the initial speech teaching large model are customized to make it more focused on the speech teaching of the target speech theme through theme analysis. In this way, not only the advantages of each model are retained, but also the overall performance of the initial speech teaching large model is improved through collaborative work.

[0093] S504, acquire a plurality of sample theme data, and the plurality of sample theme data are sample data with standard speech content labels.

[0094] Optionally, when the basic multi-modal large model is directly applied in a specific scenario, the unadjusted large model is difficult to adapt to the new scenario, and therefore, after constructing the initial speech teaching large model for the speech teaching scenario, the initial speech teaching large model needs to be trained.

[0095] Optionally, first, according to the project requirements and goals, the types and quantities of sample data required are determined, and diversified sample theme data are collected. For example, we need to cover common speech teaching event types, including but not limited to personal experiences, knowledge popularization, emotional expression, etc. These sample theme data cover different speech teaching types and scenarios, providing sufficient and accurate training materials for the model. And use data enhancement techniques such as feature transformation, data synthesis, etc. to expand the original sample theme data, increase the generalization ability of the model, and reduce the risk of overfitting.

[0096] Specifically, the acquisition approach of sample theme data includes but is not limited to the following aspects: historical speech material collection, that is, representative speech cases are selected from past speech competition videos, school speech activity records, public speech platforms, etc. These cases cover different types of speech themes, styles and skill applications, providing rich practical materials for the model; expert written samples, inviting experts or experienced teachers in the field of speech teaching to write or provide speech themes and speech samples. These samples not only have high content quality, but also contain professional speech skills and strategies, which can help the model learn more advanced speech teaching methods; public speech dataset utilization, that is, learning from domestic and foreign public speech datasets such as TED speeches and academic conference speeches. These datasets usually contain a large amount of speech videos, audios and text content, which can be used for pre-training and generalization ability improvement of the model; user feedback and interaction, which can be set up in the speech teaching platform or application to encourage children and their parents, teachers to upload speech works and collect their evaluations and suggestions on speech content and skill application. At the same time, through user interaction (such as likes and comments), popular speech themes and expression methods can be mined as supplementary sample theme data; in addition, a dynamic updating mechanism for sample theme data can be established to regularly follow new speech trends, popular topics and the latest research achievements in the education field, and incorporate these contents into the sample dataset in a timely manner to ensure that the model can keep pace with the times and adapt to the changing needs of speech teaching.

[0097] Optionally, the collected sample theme data is then labeled with standard speech content labels to ensure that each sample theme data has clear speech scripts and speech examples for reference. After the standard speech content label labeling is completed, a part of the sample theme data is randomly selected for review to ensure the accuracy and consistency of the labeling; statistical methods can also be used to evaluate the overall quality of the sample theme data set, such as whether the proportion of each type of theme is reasonable and whether there is obvious deviation in the standard speech content label, etc. If problems are found, affected data is corrected and re-labeled in a timely manner.

[0098] S506, inputting the plurality of sample theme data into the initial speech teaching large model, and training the initial speech teaching large model.

[0099] Optionally, the preprocessed sample theme data is input into the initial speech teaching large model in an appropriate format, and the initial speech teaching large model is controlled to automatically extract key information according to the characteristics of the input data and perform speech teaching.

[0100] S508, in the training process of the initial speech teaching large model, the initial speech teaching large model is controlled to output predicted speech content labels for the plurality of sample theme data, and the parameters of the initial speech teaching large model are adjusted according to the predicted speech content labels and the standard speech content labels of the plurality of sample theme data until the initial speech teaching large model converges, obtaining the trained speech teaching large model.

[0101] Optionally, in the training process of the initial speech teaching large model, the initial speech teaching large model outputs predicted speech content labels according to the input sample theme data, which are the speech teaching results of the initial speech teaching large model for the plurality of sample theme data, including speech scripts and speech examples. Then compare these predicted speech content labels with the standard speech content labels in the sample theme data. The gap between the predicted speech content labels and the standard speech content labels is the gap between the current state of the initial speech teaching large model and the expected performance.

[0102] Further, the loss function is calculated according to the gap between the predicted speech content labels and the standard speech content labels, and the parameters of the initial speech teaching large model are adjusted according to the loss value until the initial speech teaching large model converges to obtain the trained speech teaching large model. For example, the learning rate is dynamically adjusted according to the change of the loss function of the model. When the loss function decreases slowly, the learning rate is appropriately reduced to avoid the model falling into local optimum; when the loss function decreases rapidly, the learning rate is appropriately increased to accelerate convergence.

[0103] Optionally, the model training process is iterated multiple times, and each iteration uses all or part of the sample theme data for training, and adjusts the model architecture, increases the training data, optimizes the training strategy, etc., thereby gradually improving the teaching ability of the model for the speech theme. In addition, distributed training technology can be used to distribute the model training task to multiple computing nodes for parallel execution, which not only shortens the training time, but also uses more computing resources to process larger training data sets.

[0104] In the embodiments of the present application, a speech teaching method is provided, which includes a training method of a speech teaching large model. By combining a basic multi-modal large model with a preset technical tool to construct an initial speech teaching large model, combining the multi-modal data processing capability and the powerful natural language understanding capability, not only the speech teaching performance of the system is enhanced, but also the robustness and adaptability of the model are improved; further, a large amount of sample theme data with standard speech content labels is used to train the large model, so that the model can automatically learn and accurately identify various speech themes, thereby realizing the targeted teaching for the target speech theme.

[0105] Please refer to Figure 6 ,Figure 6 A structural block diagram of a speech teaching device is provided for an embodiment of the present application.

[0106] As shown in Figure 6 The speech teaching device 600 includes:

[0107] A theme analysis module 610 is configured to obtain a target speech theme, perform theme analysis on the target speech theme by using a speech teaching large model, and obtain a theme type of the target speech theme.

[0108] A content generation module 620 is configured to generate a speech script corresponding to the target speech theme according to the theme type, and generate a speech example corresponding to the target speech theme based on the theme type.

[0109] A content display module 630 is configured to display the speech script and the speech example to a user.

[0110] In some possible embodiments, the theme analysis module 610 is further configured to control the speech teaching large model to perform theme analysis on the target speech theme from a first preset dimension to obtain the theme type of the target speech theme, and the first preset dimension includes at least one of semantic understanding, keyword extraction, and theme classification.

[0111] In some possible embodiments, the content generation module 620 is further configured to sort a structural framework of the speech script corresponding to the target speech theme according to the theme type, generate material suggestions corresponding to the structural framework based on the theme type, and generate the speech script corresponding to the target speech theme according to the structural framework and the material suggestions.

[0112] In some possible embodiments, the speech teaching device 600 further includes an evaluation optimization module configured to, in response to an evaluation optimization request for an initial speech script corresponding to the target speech theme, perform evaluation on the initial speech script from a second preset dimension to obtain an evaluation result of the initial speech script, the second preset dimension includes at least one of grammatical correctness, vocabulary richness, and semantic coherence, optimize the initial speech script based on the evaluation result, and display the optimized initial speech script to the user.

[0113] In some possible embodiments, the content generation module 620 is further configured to determine a speech feature corresponding to the speech script, and generate a speech voice example of the speech script based on the speech feature, the speech feature includes at least one of a speech style, pronunciation, intonation, and speech speed, and the content display module 630 is further configured to display the speech script to the user, and play the speech voice example to the user according to a first playing rule.

[0114] In some possible embodiments, the speech teaching apparatus 600 further includes an animation example module configured to determine a body feature corresponding to the speech script, and generate a speech animation example of the speech script based on the body feature, the body feature including at least one of a hand gesture, a body, and an eye contact; and play the speech animation example to the user according to a second playing rule.

[0115] In some possible embodiments, the speech teaching apparatus 600 further includes a model training module configured to construct an initial speech teaching large model for a speech teaching scene based on a basic multi-modal large model; obtain a plurality of sample theme data, the plurality of sample theme data each being sample data with a standard speech content label; input the plurality of sample theme data into the initial speech teaching large model, and train the initial speech teaching large model; in a training process of the initial speech teaching large model, control the initial speech teaching large model to output a predicted speech content label for the plurality of sample theme data, and adjust parameters of the initial speech teaching large model according to the predicted speech content label and the standard speech content label of the plurality of sample theme data until the initial speech teaching large model converges, to obtain a trained speech teaching large model.

[0116] In the embodiment of the present application, a speech teaching device is provided, wherein a theme analysis module is configured to obtain a target speech theme, perform theme analysis on the target speech theme by using a speech teaching large model, and obtain a theme type of the target speech theme; a content generation module is configured to generate a speech script corresponding to the target speech theme according to the theme type, and generate a speech example corresponding to the target speech theme based on the theme type; and a content display module is configured to display the speech script and the speech example to a user. First, the target speech theme of the user is obtained by using the theme analysis module, and in-depth analysis is performed on the target speech theme, so that the core connotation and characteristics of the target speech theme can be accurately identified and classified, and the specific content or direction of the speech is clear, thereby providing more targeted reference basis for subsequent writing of the speech script and generation of the speech example. Next, the content generation module automatically generates a speech script with clear structure and coherent logic according to the theme type of the target speech theme, so that the content of the speech script not only meets the theme requirement, but also has certain attractiveness and appeal, thereby not only reducing the pressure of the user in writing the speech script, but also facilitating the user to quickly master the core points of the speech by referring to the speech script, and improving the efficiency and quality of writing the speech script. Meanwhile, the content generation module provides a speech example matched with the target speech theme based on the theme type, so that the user can quickly and accurately master the correct speech method and skill through intuitive imitation and learning, and improve the expressiveness and appeal of the speech. Finally, the content display module intuitively displays the generated speech script and speech example to the user, which provides rich learning resources for the user to refer to and learn at any time, so as to help them continuously improve and perfect their speech skills. Through the speech teaching method of the present application, the user can more clearly understand how to organize language and use speech skills, thereby effectively improving the speech skills and personal expressiveness, and realizing the personalized speech teaching experience.

[0117] The embodiment of the present application also provides a computer storage medium, which can store a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to perform the steps of the method in any one of the above embodiments.

[0118] Please refer to Figure 7 , Figure 7 A structural schematic diagram of a terminal is provided in the embodiment of the present application. As shown in Figure 7 , the terminal 700 can include at least one terminal processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.

[0119] The communication bus 702 is configured to realize the connection communication between the components.

[0120] The user interface 703 can include a display screen (Display) and a camera (Camera), and the optional user interface 703 can further include a standard wired interface and a wireless interface.

[0121] The network interface 704 can optionally include a standard wired interface, a wireless interface (e.g., a WI-FI interface).

[0122] The terminal processor 701 can include one or more processing cores. The terminal processor 701 connects various parts within the terminal 700 through various interfaces and lines, executes various functions of the terminal 700 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 705, and calling data stored in the memory 705. Optionally, the terminal processor 701 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The terminal processor 701 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU is mainly responsible for processing an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing content to be displayed on a display screen; and the modem is responsible for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the terminal processor 701, but can be implemented by a separate chip.

[0123] The memory 705 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory 705 includes a non-transitory computer-readable storage medium. The memory 705 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 705 can include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 705 can optionally be at least one storage device located away from the terminal processor 701. For example, Figure 7As shown, the memory 705 as a computer storage medium can include an operating system, a network communication module, a user interface module, and a speech teaching program.

[0124] In Figure 7 In the terminal 700 as shown, the user interface 703 is mainly used to provide an interface for user input and obtain user input data; and the terminal processor 701 can be used to call the speech teaching program stored in the memory 705 and specifically perform the following operations:

[0125] In some possible embodiments, when the terminal processor 701 performs the subject analysis on the target speech topic by the speech teaching large model to obtain the subject type of the target speech topic, the terminal processor 701 specifically performs the following steps: controls the speech teaching large model to perform subject analysis on the target speech topic from a first preset dimension to obtain the subject type of the target speech topic, and the first preset dimension includes at least one of semantic understanding, keyword extraction, and subject classification.

[0126] In some possible embodiments, when the terminal processor 701 performs the generation of the speech script corresponding to the target speech topic according to the subject type, the terminal processor 701 specifically performs the following steps: combs a structure framework of the speech script corresponding to the target speech topic according to the subject type; generates material suggestions corresponding to the structure framework based on the subject type; and generates the speech script corresponding to the target speech topic according to the structure framework and the material suggestions.

[0127] In some possible embodiments, the terminal processor 701 further specifically performs the following steps: in response to an evaluation optimization request for the initial speech script corresponding to the target speech topic, performs evaluation on the initial speech script from a second preset dimension to obtain an evaluation result of the initial speech script, and the second preset dimension includes at least one of grammatical correctness, vocabulary richness, and semantic coherence; optimizes the initial speech script based on the evaluation result, and displays the optimized initial speech script to the user.

[0128] In some possible embodiments, when the terminal processor 701 performs the generation of the speech example corresponding to the target speech topic based on the subject type, the terminal processor 701 specifically performs the following steps: determines a voice feature of the speech script, and generates a speech voice example of the speech script based on the voice feature, and the voice feature includes at least one of voice style, pronunciation, intonation, and speech speed; and when the terminal processor 701 performs the display of the speech script and the speech example to the user, the terminal processor 701 specifically performs the following steps: displays the speech script to the user, and plays the speech voice example to the user according to a first playing rule.

[0129] In some possible embodiments, the terminal processor 701 further specifically performs the following steps: determining a body feature corresponding to the speech script, and generating a speech animation example of the speech script based on the body feature, the body feature including at least one of a hand gesture, a body, and an eye contact; and playing the speech animation example to the user according to a second playing rule.

[0130] In some possible embodiments, the terminal processor 701 further specifically performs the following steps: constructing an initial speech teaching large model for a speech teaching scene based on the basic multi-modal large model; obtaining a plurality of sample theme data, the plurality of sample theme data each being sample data with a standard speech content label; inputting the plurality of sample theme data into the initial speech teaching large model, and training the initial speech teaching large model; in the training process of the initial speech teaching large model, controlling the initial speech teaching large model to output a predicted speech content label for the plurality of sample theme data, and adjusting parameters of the initial speech teaching large model according to the predicted speech content label and the standard speech content label of the plurality of sample theme data until the initial speech teaching large model converges, to obtain a trained speech teaching large model.

[0131] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the apparatus embodiments described above are merely schematic; the division of the modules is merely a logical function division; an actual implementation can be another division manner, for example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different modules can be indirect couplings or communication connections through some interfaces, devices or modules, and can be electrical, mechanical or in other forms.

[0132] The modules illustrated as separated components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0133] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The above computer program product includes one or more computer instructions. When the above computer program instructions are loaded and executed on a computer, all or part of the processes or functions described above according to the embodiments of the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted by the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.

[0134] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0135] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0136] The above is the description of the speech teaching method, device, storage medium and terminal provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in specific implementation and application range. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A speech teaching method, characterized in that: The method comprises: Obtaining a target speech topic, performing topic analysis on the target speech topic using a speech teaching model, and obtaining a topic type of the target speech topic; Generating a speech manuscript corresponding to the target speech topic according to the topic type, and generating a speech example corresponding to the target speech topic based on the topic type; The speech manuscript and the speech example are presented to the user.

2. The method according to claim 1, characterized in that The subject analysis of the target speech topic by the speech teaching model to obtain the subject type of the target speech topic includes: The speech teaching model is controlled to perform topic analysis on the target speech topic from a first preset dimension to obtain a topic type of the target speech topic, wherein the first preset dimension includes at least one of semantic understanding, keyword extraction, and topic classification.

3. The method according to claim 1, characterized in that Generating a speech manuscript corresponding to the target speech topic according to the topic type includes: Sorting out the structural framework of the speech manuscript corresponding to the target speech topic according to the topic type; Generating material suggestions corresponding to the structural framework based on the theme type; Generate a speech manuscript corresponding to the target speech topic based on the structural framework and the material suggestions.

4. The method according to claim 1, wherein The method further comprises: In response to an evaluation optimization request for an initial speech draft corresponding to the target speech topic, evaluating the initial speech draft based on a second preset dimension to obtain an evaluation result of the initial speech draft, wherein the second preset dimension includes at least one of grammatical correctness, lexical richness, and semantic coherence; The initial speech draft is optimized based on the evaluation result, and the optimized initial speech draft is presented to the user.

5. The method according to claim 1, wherein Generating a speech example corresponding to the target speech topic based on the topic type includes: Determining speech features corresponding to the speech draft, and generating a speech example of the speech draft based on the speech features, wherein the speech features include at least one of speech style, pronunciation, intonation, and speech speed; Presenting the speech manuscript and the speech example to the user includes: The speech manuscript is presented to the user, and the speech voice example is played to the user according to the first playing rule.

6. The method according to claim 1, characterized in that The method further comprises: Determining body features corresponding to the speech draft, and generating a speech animation example of the speech draft based on the body features, wherein the body features include at least one of gestures, body, and eye expressions; The speech animation example is played to the user according to a second playing rule.

7. The method according to claim 1, characterized in that The method further comprises: Construct an initial speech teaching model for speech teaching scenarios based on the basic multimodal model; Acquire a plurality of sample topic data, wherein the plurality of sample topic data are all sample data with standard speech content labels; Inputting the plurality of sample topic data into the initial speech teaching model to train the initial speech teaching model; During the training process of the initial speech teaching model, the initial speech teaching model is controlled to output predicted speech content labels for the multiple sample topic data, and the parameters of the initial speech teaching model are adjusted according to the predicted speech content labels and the standard speech content labels of the multiple sample topic data until the initial speech teaching model converges, thereby obtaining the trained speech teaching model.

8. A lecture teaching device, characterized in that: The device comprises: A topic analysis module is used to obtain a target speech topic, perform topic analysis on the target speech topic using a speech teaching model, and obtain a topic type of the target speech topic; A content generation module, configured to generate a speech manuscript corresponding to the target speech topic according to the topic type, and to generate a speech example corresponding to the target speech topic based on the topic type; The content display module is used to display the speech manuscript and the speech example to the user.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 7.

10. A terminal, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the program.