Interactive expression training method and device, equipment and medium
By acquiring the speech data of the speaker, performing speech recognition and emotional feature extraction, and using a multimodal large model to generate multi-dimensional expressive feedback, the problem of single evaluation dimension in existing technologies is solved, thereby improving the effectiveness and efficiency of expression training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 特赞(上海)信息科技有限公司
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies using "speech scoring" products have a single evaluation dimension, resulting in poor training effects for expression skills.
By acquiring the speech data of the speaker, performing speech recognition, extracting emotional expression features, and using a multimodal large model to generate expressive feedback information, combined with text data and emotional expression features, multi-dimensional expressive feedback is generated.
It enriches the dimensions of feedback for expression training, improves the effectiveness and systematic nature of expression training, and provides specific directions for improvement and overall optimization suggestions.
Smart Images

Figure CN121963743A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence and human-computer interaction technology, and more specifically, to an interactive expression training method, apparatus, device, and medium. Background Technology
[0002] With the development of technologies such as Large Language Model (LLM), Automatic Speech Recognition (ASR), and Text-to-Speech (TTS), several "speech scoring" and "speech evaluation" products have emerged on the market. However, these products typically only use ASR to convert speech into text and then score it according to rules, resulting in a single evaluation dimension and poor training effect. Summary of the Invention
[0003] The main purpose of this disclosure is to provide interactive expression training methods, devices, equipment, and media to solve the technical problem of poor expression training effects in the prior art, and to achieve the technical effect of improving the effectiveness of expression training.
[0004] To achieve the above objectives, a first aspect of this disclosure proposes an interactive expression training method, comprising: Acquire the speech data of the speaker; Speech recognition is performed on the speech data to obtain the corresponding text data; Extracting emotional expression features from speech data; By using a multimodal large model, based on text data and emotional expression features, expressive feedback information of the subject is generated.
[0005] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Determine the audience information of the speaker; Determine the cognitive decision-making characteristics of the audience group represented by the audience information; By using a multimodal large model, based on text data, emotional expression features, and cognitive decision-making features, expressive feedback information of the subject is generated.
[0006] In some possible implementations, determining the cognitive decision-making characteristics of the audience object represented by the audience information includes: Determine the cognitive decision-making characteristics of each of the multiple audience groups representing the audience information; and Through a multimodal large model, based on text data, sentiment expression features, and cognitive decision-making features, expressive feedback information of the subject is generated, including: By using a multimodal large model, based on text data, emotional expression features, and the cognitive decision-making features of each audience member, the expressive feedback information of each audience member is generated for the expressive subject. The feedback information from each audience member to the speaker is aggregated to obtain the aggregated feedback information from the speaker.
[0007] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain the demonstration content corresponding to the voice data; By using a multimodal large model, based on the presentation content, text data, and emotional expression features, the expressive feedback information of the subject is generated.
[0008] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: By using a multimodal large model, based on text data and emotional expression features, fragmented feedback information and whole-segment feedback information of the expressing subject are generated; Fragment feedback information and whole-segment feedback information are identified as the expressive feedback information of the subject.
[0009] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain historical feedback information from the subject of expression; By using a multimodal large model, based on text data, sentiment expression features, and historical expression feedback information, the current expression feedback information of the subject is generated.
[0010] In some possible implementations, a multimodal large model is used to generate the current expression feedback information of the expressing subject based on text data, sentiment expression features, and historical expression feedback information, including: In response to any expression feedback information of the generated expression subject, determine the feedback type of the expression feedback information, store the expression feedback information according to the feedback type, and use the stored expression feedback information as historical expression feedback information; Based on voice data, the target historical expression feedback information is determined from the stored historical expression feedback information of each feedback type. By using a multimodal large model, based on text data, sentiment expression features, and target historical expression feedback information, the current expression feedback information of the subject is generated.
[0011] In some possible implementations, the above method further includes: The historical expression feedback information belonging to the feedback type is aggregated to obtain the aggregated historical expression feedback information. Update the historical feedback information of each feedback type before aggregation to the historical feedback information after aggregation.
[0012] In some possible implementations, the method further includes, before acquiring the speech data of the speaker: To obtain information about the expressive tendencies of the subject; Based on the information about the expression tendency, expressive suggestions are generated for the subject of the expression.
[0013] Secondly, embodiments of this disclosure provide an interactive expression training device, comprising: The first acquisition unit is configured to acquire the speech data of the subject of expression; The recognition unit is configured to: perform speech recognition on speech data to obtain text data corresponding to the speech data; The extraction unit is configured to extract emotional expression features from speech data; The first generation unit is configured to generate the expressive feedback information of the subject through a multimodal large model, based on text data and sentiment expression features.
[0014] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Determine the audience information of the speaker; Determine the cognitive decision-making characteristics of the audience group represented by the audience information; By using a multimodal large model, based on text data, emotional expression features, and cognitive decision-making features, expressive feedback information of the subject is generated.
[0015] In some possible implementations, determining the cognitive decision-making characteristics of the audience object represented by the audience information includes: Determine the cognitive decision-making characteristics of each of the multiple audience groups representing the audience information; and Through a multimodal large model, based on text data, sentiment expression features, and cognitive decision-making features, expressive feedback information of the subject is generated, including: By using a multimodal large model, based on text data, emotional expression features, and the cognitive decision-making features of each audience member, the expressive feedback information of each audience member is generated for the expressive subject. The feedback information from each audience member to the speaker is aggregated to obtain the aggregated feedback information from the speaker.
[0016] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain the demonstration content corresponding to the voice data; By using a multimodal large model, based on the presentation content, text data, and emotional expression features, the expressive feedback information of the subject is generated.
[0017] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: By using a multimodal large model, based on text data and emotional expression features, fragmented feedback information and whole-segment feedback information of the expressing subject are generated; Fragment feedback information and whole-segment feedback information are identified as the expressive feedback information of the subject.
[0018] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain historical feedback information from the subject of expression; By using a multimodal large model, based on text data, sentiment expression features, and historical expression feedback information, the current expression feedback information of the subject is generated.
[0019] In some possible implementations, a multimodal large model is used to generate the current expression feedback information of the expressing subject based on text data, sentiment expression features, and historical expression feedback information, including: In response to any expression feedback information of the generated expression subject, determine the feedback type of the expression feedback information, store the expression feedback information according to the feedback type, and use the stored expression feedback information as historical expression feedback information; Based on voice data, the target historical expression feedback information is determined from the stored historical expression feedback information of each feedback type. By using a multimodal large model, based on text data, sentiment expression features, and target historical expression feedback information, the current expression feedback information of the subject is generated.
[0020] In some possible implementations, the above-described apparatus further includes: The aggregation unit is configured to aggregate historical expression feedback information belonging to the feedback type to obtain aggregated historical expression feedback information. The update unit is configured to update the historical expression feedback information of each feedback type before aggregation to the aggregated historical expression feedback information.
[0021] In some possible implementations, the device further includes: The second acquisition unit is configured to: acquire information on the expression tendency of the expression subject; The second generation unit is configured to generate expression suggestion information for the expression subject based on expression tendency information.
[0022] Thirdly, embodiments of this disclosure provide an electronic device, including: Memory, used to store computer programs; A processor is configured to execute a computer program stored in the memory, wherein, when the computer program is executed, it implements the method of any embodiment of the interactive expression training method of the first aspect of this disclosure.
[0023] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any embodiment of the interactive expression training method of the first aspect described above.
[0024] Fifthly, embodiments of this disclosure provide a computer program including computer-readable code, wherein when the computer program instructions are executed by a processor, they implement the method of any embodiment of the interactive expression training method of the first aspect described above.
[0025] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: In this embodiment of the disclosure, the speaker's speech data is acquired, then speech recognition is performed on the speech data to obtain the corresponding text data. Then, the emotional expression features of the speech data are extracted. Subsequently, a multimodal large model is used to generate the speaker's expression feedback information based on the text data and emotional expression features. The speaker's expression feedback information can be generated by combining the text data and emotional expression features, which can enrich the dimensions of expression training feedback and thus improve the effectiveness of expression training. Attached Figure Description
[0026] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of the disclosure and to make other features, objects, and advantages of the disclosure more apparent. The illustrative embodiments of the disclosure, along with their descriptions, are used to explain the disclosure and do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart of an interactive expression training method provided in this embodiment of the disclosure; Figure 2 A flowchart of another interactive expression training method provided in this disclosure embodiment; Figure 3A flowchart of yet another interactive expression training method provided in this disclosure embodiment; Figure 4 This is a schematic diagram of an interactive expression training device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] In this disclosure, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. These terms are primarily for the purpose of better describing this disclosure and its embodiments, and are not intended to limit the indicated devices, elements, or components to having a specific orientation, or to be constructed and operated in a specific orientation.
[0030] Furthermore, in addition to indicating location or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in certain circumstances to indicate a dependency or connection. Those skilled in the art can understand the specific meaning of these terms in this disclosure according to the specific circumstances.
[0031] Furthermore, the terms "installation," "setup," "equipped with," "connection," "linked," and "socketing" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0032] Figure 1 This is a flowchart illustrating an interactive expression training method provided in an embodiment of this disclosure. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0033] like Figure 1 As shown, the method specifically includes: Step 101: Obtain the speech data of the speaker.
[0034] In this embodiment, the subject of expression can be an object undergoing interactive expression training. As an example, the subject of expression can include, but is not limited to, at least one of the following: salesperson, speaker, interviewer, etc.
[0035] Voice data can be the speech generated by the subject during the training process. As an example, voice data could be the audio stream generated by a user "simulating a phone call" during telephone sales training, or the audio stream generated by a job seeker answering questions from an interviewer during video or in-person interview training.
[0036] In some alternative implementations, speech data generated by the speaker can be acquired using an audio acquisition device (such as a microphone). This audio acquisition device can be a device with audio acquisition capabilities, such as a microphone, smartphone, or voice recorder.
[0037] Step 102: Perform speech recognition on the speech data to obtain the corresponding text data.
[0038] In this embodiment, Automatic Speech Recognition (ASR) can be used to convert speech data into text data.
[0039] Text data can be the text obtained by performing speech recognition on speech data. As an example, if the speech data is the statement "I have been responsible for the technical development of three projects", the text data obtained through speech recognition can be the text "I have been responsible for the technical development of three projects".
[0040] In some alternative implementations, speech recognition technology can be used to perform speech recognition on the speech data, and the recognition result can be used as the corresponding text data.
[0041] Step 103: Extract the emotional expression features of the speech data.
[0042] In this embodiment, emotional expression features can be features in speech data that reflect the speaker's emotions, rhythm, tone, etc. As an example, emotional expression features may include, but are not limited to, at least one of the following: speech rate, fundamental frequency (pitch), energy, pitch variation amplitude, pause duration, and number of pauses.
[0043] In some optional implementations, the speech data can be processed frame by frame (e.g., emotion recognition feature extraction, tone feature extraction) to obtain emotional expression features; or, the speech data can be processed as a whole segment (e.g., emotion recognition feature extraction, tone feature extraction) to obtain emotional expression features; or, if the speech data is a data stream, the speech data can be processed according to a preset duration (e.g., emotion recognition feature extraction, tone feature extraction) to obtain emotional expression features.
[0044] Step 104: Using a multimodal large model, based on text data and emotional expression features, generate the expression feedback information of the subject.
[0045] In this embodiment, the multimodal big data model can be used to process various modalities of data, such as the emotional expression features of text and voice data, thereby generating expression feedback information tailored to the speaker. As an example, the multimodal big data model can evaluate expression based on the semantic features of text data and the emotional expression features of voice data, thereby obtaining scores and suggestions encompassing the following dimensions: comprehensibility (whether the audience can understand), persuasiveness (whether it addresses the audience's pain points or interests), structure (whether it is logically organized), expression style (appropriate speaking speed, natural tone, and emotional control), and business goal achievement (whether the sale / persuasion / self-presentation was successfully completed). The scores and suggestions for the above dimensions can be determined by inputting prompt words into the multimodal big data model.
[0046] Feedback information can be evaluations and suggestions generated regarding the content and performance of the speaker. For example, feedback could be an evaluation in telephone sales training stating, "The pain points are not sufficiently addressed; we need to highlight the cost-saving points for the customer," or a suggestion in speech training stating, "The speaking speed is too slow; we need to improve the rhythm of key parts."
[0047] In some optional implementations of this embodiment, the following approach can be adopted: using a multimodal large model, based on text data and sentiment expression features, to generate the expressive feedback information of the subject: The first step is to generate fragmented feedback information and full-length feedback information of the subject of expression based on text data and sentiment expression features through a multimodal large model.
[0048] In this context, segment feedback information can be feedback on specific paragraphs, sections, or segments in the speaker's presentation. For example, segment feedback information could be, in telephone sales training, feedback to a user regarding their product introduction segment: "The speaking speed is too fast; you need to slow down."
[0049] The complete feedback message can be a summary of the speaker's overall delivery. For example, the complete feedback message could be the conclusion after a sales call training session: "The predicted likelihood of a sale is 60%, and the next round of training should focus on improving the ability to uncover customer needs."
[0050] In some alternative implementations, a multimodal large model can be used to analyze fragments of text data and corresponding fragments of sentiment expression features, thereby outputting fragment feedback information of the expressing subject; or a multimodal large model can be used to analyze the overall text data and the overall sentiment expression features, thereby outputting fragment feedback information of the expressing subject.
[0051] The second step is to identify the fragment feedback information and the whole paragraph feedback information as the expression feedback information of the subject.
[0052] Here, fragment feedback information and whole-segment feedback information can be integrated into the final expression feedback information.
[0053] It is understandable that among the above-mentioned optional implementation methods, generating fragment feedback information through a multimodal large model allows the speaker to more accurately understand the problems and advantages of each specific segment in the expression process, facilitating targeted improvements to specific areas and avoiding the inability to pinpoint specific improvement points while only knowing the overall problems. Generating whole-segment feedback information allows the speaker to understand their overall performance, global strengths, and core weaknesses, forming a holistic understanding of the expression effect. Integrating fragment feedback information and whole-segment feedback information into the speaker's expression feedback information achieves a combination of local detail guidance and global overall evaluation, providing the speaker with both directions for improvement of specific segments and overall optimization suggestions, thereby enhancing the systematicness and efficiency of expression training.
[0054] Optionally, when there are multiple pieces of expressive feedback information from the subject, the highest priority expressive feedback information can be determined from the multiple expressive feedback information and displayed.
[0055] Each feedback message can correspond to a priority level. The priority of this feedback message can be determined as follows: First, determine the expression scenario of the target audience, such as business negotiation, teaching demonstration, or customer service communication. Then, configure multiple evaluation dimensions for the expression scenario. For example, for the business negotiation scenario, the following evaluation dimensions can be configured: logical rigor, persuasiveness of language, adaptability, and goal achievement. Different expression scenarios can be pre-associated with multiple different evaluation dimensions, and each evaluation dimension can be pre-associated with a certain level of importance. This pre-associated importance can be determined by the audience. Then, sort and display these multiple evaluation dimensions according to their importance. For example, they can be sorted and displayed in the order of goal achievement, logical rigor, adaptability, and persuasiveness of language. Afterwards, the sorting operation on these evaluation dimensions can be detected, resulting in a reordered set of evaluation dimensions. For example, the reordered evaluation dimensions might be logical rigor, goal achievement, persuasiveness of language, and adaptability. Then, based on the reordered order of the evaluation dimensions, a first weight coefficient is determined for each dimension, and a second weight coefficient is determined based on the importance of each dimension. For example, evaluation dimensions ranked higher have higher first weight coefficients, and evaluation dimensions with higher importance have higher second weight coefficients. Subsequently, based on the first and second weight coefficients of each evaluation dimension and the scores of the expressed feedback information on each dimension, a comprehensive score for the expressed feedback information is calculated. The expressed feedback information with the highest priority is the one with the highest comprehensive score. For example, the average of the first and second weight coefficients can be weighted and summed with the scores of the expressed feedback information on each evaluation dimension to obtain the comprehensive score of the expressed feedback information. Thus, the priority of expressed feedback information can be adaptively adjusted based on different expression scenarios, different audiences, and different user subjective preferences, making the expressed feedback information more in line with actual needs.
[0056] In some optional implementations of this embodiment, the following approach can be adopted: using a multimodal large model, based on text data and sentiment expression features, to generate the expressive feedback information of the subject: The first step is to obtain historical feedback information from the subject of the expression.
[0057] Historical feedback information can be the feedback information generated and stored by the speaker during previous interactive expression training. As an example, historical feedback information may include feedback such as "insufficient demand identification," "speaking too fast," and "lack of logical product explanation" obtained from the user's first 5 telephone sales training sessions.
[0058] In some optional implementations, past feedback information of the expressing subject can be retrieved from stored feedback data. As an example, Retrieval-Augmented Generation (RAG) can be used to retrieve relevant historical feedback information of the user. Specifically, RAG can be used to retrieve historical feedback information related to the current expressing subject from a database.
[0059] The second step involves using a multimodal large model to generate the current expression feedback information of the subject based on text data, sentiment expression features, and historical expression feedback information.
[0060] The current expression feedback information can be the feedback information generated most recently for the current expression training. As an example, the current expression feedback information can be the expression feedback information generated in the user's last telephone sales training, combined with the expression feedback information of the previous 5 historical expressions, which says "The ability to discover needs has been improved, and the logic of product explanation needs to be optimized."
[0061] In some alternative implementations, the multimodal large model can combine text data, sentiment expression features, and historical expression feedback information from the current expression training process to generate current expression feedback information for the subject. As an example, the current expression feedback information can be a new expression feedback information generated by the multimodal large model by comparing the weaknesses in the user's current expression (e.g., text data, sentiment expression features) with historical expression feedback information, evaluating improvements, and generating new expression feedback information.
[0062] It is understandable that obtaining historical expression feedback information from the subject allows the multimodal large model to understand the subject's past performance, historical weaknesses, and areas for improvement, providing historical reference for generating current expression feedback information. By generating current expression feedback information based on text data, sentiment expression features, and historical expression feedback information through the multimodal large model, the current expression feedback information can reflect the training changes of the subject, achieving continuity and progression of feedback, and helping the subject to improve their expression.
[0063] In some application scenarios of the above-mentioned optional implementation methods, the following approach can be adopted to generate the current expression feedback information of the expressing subject based on text data, sentiment expression features, and historical expression feedback information through a multimodal large model: The first step is to respond to any expression feedback information of the generated expression subject, determine the feedback type of the expression feedback information, store the expression feedback information according to the feedback type, and use the stored expression feedback information as historical expression feedback information.
[0064] The feedback type can be a category based on the content of the feedback information. For example, feedback types may include, but are not limited to, at least one of the following: speech rate related, needs assessment related, structure related, persuasiveness related, emotional expression related, business logic related, and language standardization related.
[0065] Here, after each expressive feedback message is generated, its feedback type can be determined, and the expressive feedback message can be stored according to its feedback type. For example, expressive feedback messages of the same feedback type can be stored in one storage unit, and expressive feedback messages of different feedback types can be stored in different storage units. Furthermore, the stored expressive feedback messages can be used as historical expressive feedback messages.
[0066] The second step is to determine the target historical expression feedback information from the stored historical expression feedback information of each feedback type, based on the voice data.
[0067] Specifically, the target historical expression feedback information can be historical expression feedback information related to voice data that is filtered from stored historical expression feedback information. As an example, the target historical expression feedback information can be feedback information related to the logic of the product explanation retrieved from historical expression feedback information.
[0068] In some alternative implementations, target historical expression feedback information can be filtered from historical expression feedback information based on the content and context of the current speech data. As an example, the target historical expression feedback information can be obtained by retrieving historical expression feedback information that is most relevant (e.g., has the highest similarity) to the content and context of the current speech data using the RAG method.
[0069] The third step involves using a multimodal large model to generate the current expression feedback information of the subject based on text data, sentiment expression features, and target historical expression feedback information.
[0070] In some optional implementations, a multimodal large model can combine current text data, sentiment expression features, and historical expression feedback information of the target to analyze and output current expression feedback information. As an example, a multimodal large model can, based on records of "insufficient demand mining" in the historical expression feedback information of the target, evaluate the demand mining situation of the current expression based on the current text data and sentiment expression features, and generate feedback information, thereby obtaining the current expression feedback information.
[0071] Understandably, in the above application scenarios, historical expression feedback information can be structured and managed according to feedback type, facilitating rapid retrieval and use later. Determining target historical expression feedback information from historical expression feedback information of various feedback types based on voice data allows for the filtering of key historical expression feedback information relevant to the current expression (i.e., text data and emotional expression features), avoiding interference from irrelevant historical expression feedback information in the generation of current expression feedback information, and improving the accuracy of generating current expression feedback information. Generating current expression feedback information based on text data, emotional expression features, and target historical expression feedback information using a multimodal large model enables the current expression feedback information to focus more on the historical weaknesses of the expressing subject and the relevance of the current expression, improving the targeting of the expression feedback.
[0072] In some of the above application scenarios, the following steps can also be performed: The first step is to aggregate the historical expression feedback information that belongs to the feedback type to obtain the aggregated historical expression feedback information.
[0073] Among them, the aggregated historical expression feedback information can be the information obtained by aggregating multiple historical expression feedback information under the same feedback type.
[0074] In some optional implementations, the aggregated historical expression feedback information can be aggregating 10 historical expression feedback information related to speech speed into one message: "The overall speech speed is too fast, and it needs to be deliberately slowed down during the product explanation."
[0075] The second step is to update the historical expression feedback information that belongs to the feedback type before aggregation with the historical expression feedback information after aggregation.
[0076] In some alternative implementations, aggregated historical feedback information can be used to replace multiple original historical feedback information of the same type, thereby achieving context compression and preventing the context from becoming too bloated.
[0077] Understandably, in the above scenario, aggregating historical expression feedback information belonging to the same feedback type can combine multiple duplicate or similar historical expression feedback information, reduce the amount of historical expression feedback information data, avoid wasting storage resources, and extract key improvement directions (i.e., the aggregated historical expression feedback information) for easy subsequent retrieval and use. Updating the historical expression feedback information before aggregation to the aggregated historical expression feedback information can optimize the storage structure of historical expression feedback information, avoid multimodal large models processing a large amount of redundant data when calling historical expression feedback information, reduce context length pressure, and improve the efficiency of expression feedback information generation.
[0078] In some optional implementations of this embodiment, the following steps may be performed before acquiring the speech data of the speaker: The first step is to obtain information about the expressive tendencies of the subject.
[0079] Among these, the information expressing the speaker's intention can include the goals they want to achieve, the key points they want to emphasize, and the core direction of their expression during the training. For example, this intention could be a user's response during telephone sales training: "I want to achieve the goal of recommending the product within 3 minutes," or "I want to highlight the product's advantages in increasing table turnover."
[0080] In some optional implementation methods, information such as the speaker's expression goals and key points can be collected. For example, users can be interviewed in a 10-second quick interview and their answers recorded to obtain information about their expression tendencies.
[0081] The second step is to generate expression suggestions for the subject of the expression based on the information on expression tendency.
[0082] Among these, suggestive information can be used to guide the speaker on how to express themselves. For example, suggestive information could be something like, "It is recommended to first clearly explain the product positioning in one sentence, prioritizing solutions that can help increase table turnover."
[0083] In some alternative implementations, a large model can be used to generate expression suggestions tailored to the speaker based on their expression preferences. For example, by combining the user's expression preferences with uploaded script documents, expression suggestions on "what to say and what not to say" can be generated.
[0084] It is understandable that among the above-mentioned optional implementation methods, obtaining the expression tendency information of the expression subject can clarify the expression goal and core focus of the expression subject, providing a clear basis for the generation of expression suggestion information, making the expression suggestion information more in line with the needs of the expression subject; generating expression suggestion information for the expression subject based on the expression tendency information can provide targeted guidance before the expression subject begins formal expression training, helping the expression subject to clarify the expression direction in advance, focus on the core focus, avoid expression problems that may deviate from the goal, improve the initial effect and efficiency of expression training, allow the expression subject to obtain effective feedback more quickly, and enhance the enthusiasm for training.
[0085] It should be noted that, where there is no conflict, the technical features described in different alternative implementations can be included in the same embodiment. For the sake of brevity, they will not be elaborated here.
[0086] Based on the embodiments of this disclosure, by acquiring the speech data of the speaker, then performing speech recognition on the speech data to obtain the corresponding text data, then extracting the emotional expression features of the speech data, and subsequently generating the speaker's expression feedback information based on the text data and emotional expression features through a multimodal large model. The speaker's expression feedback information can be generated by combining the text data and emotional expression features, thus enriching the dimensions of expression training feedback and improving the effectiveness of expression training.
[0087] Figure 2 This is a flowchart illustrating another interactive expression training method provided in an embodiment of this disclosure. Figure 2 As shown, the method specifically includes: Step 201: Obtain the speech data of the speaker.
[0088] In this embodiment, step 201 and Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.
[0089] Step 202: Perform speech recognition on the speech data to obtain the corresponding text data.
[0090] In this embodiment, step 202 and Figure 1 Step 102 in the corresponding embodiment is basically the same, and will not be repeated here.
[0091] Step 203: Extract the emotional expression features of the speech data.
[0092] In this embodiment, step 203 and Figure 1 Step 103 in the corresponding embodiment is basically the same, and will not be repeated here.
[0093] Step 204: Determine the audience information of the speaker.
[0094] In this embodiment, audience information can be various descriptive information related to the audience. As an example, audience information may include, but is not limited to: industry, job title, cognitive level, and points of interest in the topic.
[0095] In some alternative implementations, users (such as the speaker) can input target customer group information through the interface and use that target customer group information as the speaker's audience information.
[0096] Step 205: Determine the cognitive decision-making characteristics of the audience in the audience information representation.
[0097] In this embodiment, the audience can be the individuals or groups to whom the speaker's message is addressed. For example, the audience could be a regional manager of a chain restaurant brand in telephone sales training, or a technical supervisor in interview training. The audience can be represented using an AI (Artificial Intelligence) model.
[0098] Cognitive decision-making characteristics may include, but are not limited to, the audience's knowledge boundaries, comprehension preferences, decision-making logic, and focus.
[0099] Step 206: Using a multimodal large model, based on text data, emotional expression features, and cognitive decision-making features, generate the expression feedback information of the expressing subject.
[0100] In some optional implementations, multimodal big data models can combine textual data, sentiment expression features, and cognitive decision-making features to generate expressive feedback information from the expressive subject. As an example, a multimodal big data model can start from the cognitive decision-making features of the chain restaurant manager (the audience) to assess whether the expressive subject's expression addresses their concern about table turnover rate, thereby generating questions or feedback and obtaining the expressive subject's expressive feedback information.
[0101] In some optional implementations of this embodiment, the cognitive decision characteristics of the audience object represented by the audience information can be determined in the following ways: Determine the cognitive decision-making characteristics of each of the multiple audience groups in the audience information representation.
[0102] Based on this, the following approach can be used to generate the expressive feedback information of the expressing subject through a multimodal large model, based on text data, sentiment expression features, and cognitive decision-making features: The first step is to generate feedback information for each audience member on the subject of expression by using a multimodal large model based on text data, emotional expression features, and cognitive decision-making features of each audience member.
[0103] In some alternative implementations, the cognitive decision-making characteristics of each audience member can be analyzed separately based on audience information. Then, using a multimodal large model, based on text data, sentiment expression characteristics, and the cognitive decision-making characteristics of each audience member, expressive feedback information for each audience member is generated for the expressing subject.
[0104] In some cases, the target audience may include: human resources managers, department heads, company executives, etc.
[0105] The second step is to aggregate the feedback information from each audience member to the speaker, resulting in aggregated feedback information from the speaker.
[0106] Among them, multiple audience objects can be two or more different individual audiences or groups corresponding to the audience information.
[0107] In some optional implementations, common problems and personalized suggestions can be extracted from the expressive feedback information of multiple audiences according to dimensions such as comprehension and persuasiveness, and then aggregated to obtain the aggregated expressive feedback information of the expressive subjects.
[0108] It is understandable that by identifying the cognitive decision-making characteristics of multiple audience members, the differences in perspectives among different audiences can be covered, avoiding the one-sidedness of evaluation caused by a single audience perspective; by generating feedback information from each audience member regarding the speaker's expression, multi-dimensional and differentiated evaluation opinions can be provided to the speaker, allowing the speaker to understand the views of different types of audiences on their expression; by aggregating and processing the feedback information from each audience member, the core content of multi-perspective feedback can be integrated, enabling the speaker to grasp the core direction for improvement more quickly and improve the effectiveness of expression training.
[0109] It should be noted that, in addition to the contents described above, this embodiment may also include... Figure 1 The corresponding technical features described in the corresponding embodiments, thereby achieving Figure 1 For details on the technical effects of the interactive expression training method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0110] Based on the embodiments of this disclosure, by determining the cognitive decision-making characteristics of the audience through audience information, the multimodal large model can more accurately grasp the audience's understanding preferences and decision-making logic, avoiding evaluations that are detached from the audience's actual needs; by generating expressive feedback information based on text data, emotional expression characteristics and cognitive decision-making characteristics through the multimodal large model, targeted feedback from the audience's perspective is achieved.
[0111] Figure 3 This is a flowchart illustrating yet another interactive expression training method provided in an embodiment of this disclosure. Figure 3 As shown, the method specifically includes: Step 301: Obtain the speech data of the speaker.
[0112] In this embodiment, step 301 and Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.
[0113] Step 302: Perform speech recognition on the speech data to obtain the corresponding text data.
[0114] In this embodiment, step 302 and Figure 1 Step 102 in the corresponding embodiment is basically the same, and will not be repeated here.
[0115] Step 303: Extract the emotional expression features of the speech data.
[0116] In this embodiment, step 303 and Figure 1 Step 103 in the corresponding embodiment is basically the same, and will not be repeated here.
[0117] Step 304: Obtain the demonstration content corresponding to the voice data.
[0118] In this embodiment, the demonstration content can be relevant content displayed by the presenter during the presentation. As an example, the demonstration content may include demonstration documents uploaded by users during roadshow training.
[0119] In some alternative implementations, the presentation content presented by the subject can be captured (e.g., via a camera) or received (e.g., via a network protocol) as the subject expresses itself, such as presentation slides, charts, or videos.
[0120] Step 305: Using a multimodal large model, based on the presentation content, text data, and emotional expression features, generate the expression feedback information of the subject.
[0121] In this embodiment, the multimodal large model can combine presentation content, text data, and emotional expression features to generate expression feedback information of the expressing subject, such as the information "The content of page 3 of the document is not mentioned in the expression".
[0122] It should be noted that, in addition to the contents described above, this embodiment may also include... Figure 1 The corresponding technical features described in the corresponding embodiments, thereby achieving Figure 1 For details on the technical effects of the interactive expression training method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0123] Based on the embodiments of this disclosure, by acquiring the demonstration content corresponding to the voice data, the feedback of the multimodal large model can be made more consistent with the actual expression scenario (such as speeches, roadshows, teaching demonstrations, and other scenarios that require accompanying demonstration content). Through the multimodal large model, expression feedback information is generated based on the demonstration content, text data, and emotional expression features. This can provide feedback on the matching degree, relevance, and logic between the expression and the demonstration content, avoiding the one-sidedness of feedback caused by expression feedback based solely on voice data. This makes the expression feedback information more comprehensive and helps the expression subject optimize the synergistic effect between the expression and the demonstration content.
[0124] The following describes the embodiments of this disclosure by way of example. However, it should be noted that the following content is only used to understand the technical solutions of the embodiments of this disclosure and does not constitute a limitation on the protection scope of the embodiments of this disclosure.
[0125] With the development of large-scale models and technologies such as speech recognition and speech synthesis, several "speech scoring" and "speech rating" products have emerged on the market.
[0126] In developing this disclosure, the inventors discovered that such solutions typically only utilize ASR to convert speech data into text data, and then score it using rules or small models. This makes it difficult to comprehensively consider speech features such as speaking speed, pauses, tone of voice, and stress, and even more so, it fails to realistically simulate the comprehension process of different types of listeners. Furthermore, existing solutions often use "standard answer matching" and "keyword coverage" as core indicators, ignoring the impact of listener differences in real-world business scenarios (e.g., professional background, knowledge level, and focus), and cannot provide differentiated feedback to "target customers / interviewers / judges." Some products only provide an overall score and comments after the user has finished speaking a long passage (e.g., an entire speech or phone call), resulting in delayed feedback. Users have to wait a long time for feedback, making it difficult to form a high-frequency, high-engagement training habit. Existing solutions often simply save historical training results without structurally archiving and compressing multi-round training records and feedback content. They lack a closed-loop mechanism of "long-term memory + dynamic optimization," and cannot adaptively adjust the training focus and difficulty according to the user's growth.
[0127] Therefore, the inventors discovered that there is an urgent need for an interactive expression training system that can simultaneously utilize multimodal signals (speech, text, tone features, etc.), combine a multi-role audience model, and have long-term memory and feedback optimization mechanisms, in order to improve the relevance and efficiency of expression training.
[0128] This solution aims to address the following issues in existing expression training schemes: relying solely on ASR text fails to fully utilize speech features such as speech rate, volume, and intonation, resulting in a single evaluation dimension; lacking an understanding and evaluation mechanism tailored to specific audiences makes it difficult to assess whether the message is clear to a particular group of people; feedback is delayed during training, failing to provide effective prompts within a short timeframe (e.g., 10 seconds), leading to a poor user experience; and historical training feedback from multiple rounds is not structured, compressed, or utilized, making it impossible to achieve adaptive optimization of the training strategy.
[0129] In view of this, the inventor proposes the following solution: The multimodal acquisition and feature extraction module collects users' voice data and optional video / screen content (i.e. presentation content) in scenarios such as telephone sales, interviews, or speeches. The transcribed text is obtained through ASR to obtain text data. The voice data is then processed to extract tone features such as speech rate, pitch, energy, and pause duration (belonging to emotional expression features) to form a multimodal feature vector.
[0130] The understanding and simulation module is built by constructing an audience model. Based on the speech documents uploaded by users and the target audience information (such as industry, job position, cognitive level, and topic of interest), one or more AI audience models (i.e. audience objects) are constructed. The knowledge boundary, understanding preference, and decision logic (i.e. cognitive decision features) of the audience model are generated using a large language model. During the training process, the audience's expressed content (i.e. voice data) is simulated to understand and ask questions (i.e. express feedback information) from the perspective of the audience.
[0131] Through the expression effectiveness evaluation and feedback generation module, the text, multimodal tone feature vectors, and audience model persona settings (i.e., cognitive decision features) are input into the multimodal big model to evaluate the speech segment-level (i.e., segment feedback information) and overall (i.e., whole-segment feedback information). The output includes, but is not limited to, the following dimensions of scores and suggestions (i.e. expression feedback information): comprehensibility (whether the audience can understand), persuasiveness (whether it hits the audience's pain points and benefits), structure (whether it is clear and organized), expression style (appropriate speaking speed, natural tone, emotional control), and business goal achievement (whether the sales / persuasion / self-presentation were successfully completed).
[0132] The user's training history is managed through a feedback memory and dynamic compression module. Specifically, a retrieval-enhanced generation approach can be used to manage the user's training history. For example, after each training session, feedback information (such as "speech speed too fast," "lack of customer pain point identification") can be categorized and stored in a storage space. When the user starts the next round of training, the most relevant content (i.e., the target historical feedback information) can be retrieved from the past feedback (i.e., historical feedback information), rather than all of it. This allows the multimodal large model to recall the user's past improvements more quickly. Furthermore, these feedbacks (i.e., historical feedback information) can be automatically compressed and aggregated periodically (e.g., aggregating 10 feedback information related to "speech speed" into one core suggestion (i.e., the aggregated historical feedback information)) to prevent the context from becoming too bloated. Thus, from the user's perspective, when the user receives feedback information "insufficient demand identification" during the first training session, and "I noticed you have improved by 30% in demand identification recently; the current issue is the logic of the product explanation" during the fifth training session, the system can effectively address these issues. This continuous feedback makes the user feel that the system is like a "coach with memory," rather than starting from scratch each time. This avoids context "explosion," recalling only the most relevant historical information and preventing the large multimodal model from processing too many irrelevant historical adaptive training paths. The system can adjust the focus of the next round of representation training based on "which dimensions have improved and which are still weak," and provide rapid feedback. Users don't need to wait for the large model to re-understand the entire history; they only need to generate feedback quickly based on the relevant context.
[0133] Through the feedback optimization and adaptive training strategy module, the training strategy can be automatically adjusted in subsequent expression training based on the feedback distribution in long-term memory and user performance trends. For example, training tasks can be generated by prioritizing user weaknesses (such as "unclear business logic" or "speaking too fast"); at the beginning of expression training, users can be given prompts on "what to say and what not to say" through 10-second quick interviews, allowing users to perceive the value of the system in advance; real-time micro-feedback can be provided in sections during the presentation, and an overall debriefing report can be provided after the presentation.
[0134] As a first example, in a telephone sales training scenario, the user (the speaker) first selects the "Telephone Sales Training" scenario on the interface and uploads sales script documents or product introduction documents. Then, they fill in the target customer group information (e.g., regional manager of a chain restaurant brand, store manager of a front-line store, etc.) to generate an initial audience model. Afterwards, based on the uploaded sales script documents and configured business information, a large language model can be used to construct the following fields: Industry background: e.g., "Operations manager of a chain fast-food brand in East China"; Key Performance Indicators (KPIs): average order value, table turnover rate, raw material cost, customer repurchase rate, etc.; Past experience: whether they have worked with similar condiment / curry sauce brands; Risk preference: whether they value stable supply or innovative collaborations; Language preference: preferring honesty and disliking excessive packaging. Before the user officially begins their practice, a 10-second interview can be conducted, such as asking: "What goal do you hope to achieve in this 3-minute call?"; "What do you think is the one question the other party is most likely to be worried about?". Based on user responses and uploaded sales script documents, an instant prompt is generated, such as: suggest clearly stating the product positioning in one sentence; avoid continuously discussing formula details; prioritize solutions that can help the other party "increase table turnover rate." These prompts can be highlighted on the interface to provide users with initial guidance. Subsequently, multimodal data acquisition and feature extraction are performed: the user begins "simulating a phone call" through the interface; the audio stream (i.e., voice data) is sent in real-time to the backend ASR service to obtain text data; simultaneously, the backend performs frame-level processing on the audio, extracting features such as speech rate, pitch, energy, and pause length to construct a time-series feature vector. Next, multimodal fusion and audience understanding simulation are performed: for example, the following can be assembled into multimodal input: the transcribed text of the current section; the aggregated vector of tone features for the corresponding time period (e.g., average speaking speed, number of pauses, pitch variation); the structured description of the audience model; and the following inference tasks are performed through the multimodal large model: comprehension judgment: whether the expression here is "comprehensible" to the audience represented by the current audience model; persuasiveness analysis: whether it hits the KPIs that the audience cares about; expression style analysis: whether the speaking speed is too fast, whether the tone is monotonous, and whether the pauses are appropriate. Then, segment-by-segment feedback and overall report generation are performed: For each logical segment, the system generates structured feedback objects, such as: comprehension score, pain point relevance, and expression rhythm suggestions; 2-3 actionable suggestions are given in natural language (such as "Here you can first ask the other party about the average order value of the store, and then introduce your product"); After the entire call ends, a summary report for each segment can be generated, namely the segment feedback information: prediction of the likelihood of closing the deal (based on the matching degree between the script and the audience); replay of key turning points (where the customer was "lost"); the focus of the next expression training (such as "reduce the introduction of product history and ask more about the other party's pain points").Next, feedback memory and dynamic compression can be implemented: each training record is broken down into: scene tags, audience information, score vectors for each segment, text summaries, etc.; multi-round training data are aggregated and long-term summaries are generated according to "skill dimensions" (e.g., needs mining, product explanation, sales ability); in subsequent expression training, only the user's key summaries and weakness tags can be injected into the larger model context, thereby maintaining continuous awareness of the user's history within a limited context. Finally, strategy optimization and task recommendation can be performed: the progress curves of each dimension are calculated periodically based on historical training records; when a certain skill is consistently weak, the system will prioritize recommending corresponding templates (e.g., training scenarios that only practice needs mining opening remarks), forming an adaptive training path.
[0135] As a second example, in a phone / video interview training scenario, building upon the first example, the audience model can be replaced with an "interviewer model," such as a "technical manager" or "director" model. Evaluation metrics can include: consistency between the resume and the content expressed; depth of understanding of job responsibilities; and completeness of expression in behavioral interview questions. During the quick interview phase, questions can be asked: What position are you applying for? What are the 1-2 key experiences you most want to emphasize? This generates "interview focus tips," helping users focus on the narrative before formal training.
[0136] As a third example, in the context of speech / roadshow training, based on the first example, the following aspects can be appropriately expanded: support uploading complete roadshow materials; in the quick interview stage, automatically identify chapters based on the material structure and provide suggestions on "how many sentences to say on each page and not to read the original text"; and place greater emphasis on "rhythm changes" and "emotional fluctuations" in the evaluation of tone characteristics to help speakers enhance their persuasiveness.
[0137] It should be noted that this solution is not limited to a specific ASR service and can use any engine that supports streaming recognition. Speech features can be extracted using traditional DSP (Digital Signal Processing) methods, or an end-to-end emotion recognition model can output "emotion vectors," which are essentially a type of multimodal feature. In simulations involving multi-role audiences and group discussions, it can be extended to a "jury mode" (such as a roadshow judge) with multiple audience members participating simultaneously. The large model produces feedback from different perspectives, which are then aggregated. Regarding the storage medium and structure of the memory module, long-term memory based on relational and vector databases can be used, or object storage and local indexing can be employed; essentially, it is a combination of structured storage and vector retrieval.
[0138] It should be noted that, in addition to the contents described above, this embodiment may also include the technical features described in the above embodiments, thereby achieving the technical effect of the interactive expression training method shown above. Please refer to the above description for details. For the sake of brevity, it will not be elaborated here.
[0139] The above technical solution enables an upgrade from "LLM that only scores" to "expression coaching based on the audience's perspective, with long-term memory, and continuous optimization." Compared with existing technologies, this solution can introduce intonation features such as speech rate, energy, and pauses on top of traditional ASR and text analysis. These features, along with text semantics, are input into a multimodal large model, enabling a more accurate assessment of emotional impact and expression rhythm. Furthermore, this solution does not simply judge "whether it's good or bad," but evaluates from the perspective of "whether a specific audience (such as a certain type of customer or interviewer) can understand and be moved," which is closer to real-world business scenarios. After users upload their script documents and audience information, suggestions on "what to say and what not to say" can be provided through short interviews (e.g., Q&A within 10 seconds), greatly improving user conversion rates and willingness to continue training. Multi-round training records are not simply accumulated but compressed and stored through hierarchical summarization and structured vector indexing. This ensures information fidelity while controlling context length, making subsequent training and feedback more continuous and targeted. Furthermore, it can automatically adjust the training focus based on the user's historical performance in different dimensions, achieving a personalized and progressive path to improve expressive abilities, rather than a uniform template.
[0140] Figure 4 This is a schematic diagram of an interactive expression training device provided in an embodiment of the present disclosure. The interactive expression training device includes: The first acquisition unit 401 is configured to acquire the speech data of the subject of expression; The recognition unit 402 is configured to: perform speech recognition on speech data to obtain text data corresponding to the speech data; Extraction unit 403 is configured to: extract emotional expression features from speech data; The first generation unit 404 is configured to generate the expression feedback information of the subject through a multimodal large model based on text data and sentiment expression features.
[0141] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Determine the audience information of the speaker; Determine the cognitive decision-making characteristics of the audience group represented by the audience information; By using a multimodal large model, based on text data, emotional expression features, and cognitive decision-making features, expressive feedback information of the subject is generated.
[0142] In some possible implementations, determining the cognitive decision-making characteristics of the audience object represented by the audience information includes: Determine the cognitive decision-making characteristics of each of the multiple audience groups representing the audience information; and Through a multimodal large model, based on text data, sentiment expression features, and cognitive decision-making features, expressive feedback information of the subject is generated, including: By using a multimodal large model, based on text data, emotional expression features, and the cognitive decision-making features of each audience member, the expressive feedback information of each audience member is generated for the expressive subject. The feedback information from each audience member to the speaker is aggregated to obtain the aggregated feedback information from the speaker.
[0143] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain the demonstration content corresponding to the voice data; By using a multimodal large model, based on the presentation content, text data, and emotional expression features, the expressive feedback information of the subject is generated.
[0144] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: By using a multimodal large model, based on text data and emotional expression features, fragmented feedback information and whole-segment feedback information of the expressing subject are generated; Fragment feedback information and whole-segment feedback information are identified as the expressive feedback information of the subject.
[0145] In some possible implementations, a multimodal large model is used to generate expressive feedback information of the subject based on text data and sentiment expression features, including: Obtain historical feedback information from the subject of expression; By using a multimodal large model, based on text data, sentiment expression features, and historical expression feedback information, the current expression feedback information of the subject is generated.
[0146] In some possible implementations, a multimodal large model is used to generate the current expression feedback information of the expressing subject based on text data, sentiment expression features, and historical expression feedback information, including: In response to any expression feedback information of the generated expression subject, determine the feedback type of the expression feedback information, store the expression feedback information according to the feedback type, and use the stored expression feedback information as historical expression feedback information; Based on voice data, the target historical expression feedback information is determined from the stored historical expression feedback information of each feedback type. By using a multimodal large model, based on text data, sentiment expression features, and target historical expression feedback information, the current expression feedback information of the subject is generated.
[0147] In some possible implementations, the above-described apparatus further includes: The aggregation unit (not shown in the figure) is configured to aggregate historical expression feedback information belonging to the feedback type to obtain aggregated historical expression feedback information. The update unit (not shown in the figure) is configured to update the historical expression feedback information of each feedback type before aggregation to the aggregated historical expression feedback information.
[0148] In some possible implementations, the device further includes: The second acquisition unit (not shown in the figure) is configured to acquire information on the expression tendency of the subject. The second generation unit (not shown in the figure) is configured to generate expression suggestion information for the expression subject based on expression tendency information.
[0149] The interactive expression training device provided in this embodiment can execute the corresponding steps of the interactive expression training methods described above, thereby achieving the technical effects of the interactive expression training methods described above. The interactive expression training device and the interactive expression training methods can refer to and cite each other in terms of specific implementation and technical effects. For the sake of brevity, they will not be elaborated here.
[0150] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Figure 5 The illustrated electronic device 500 includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the electronic device 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to implement communication between these components. In addition to a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 5 The general designated all buses as Bus System 505.
[0151] The user interface 503 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).
[0152] It is understood that the memory 502 in this embodiment of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0153] In some implementations, memory 502 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 5021 and application program 5022.
[0154] The operating system 5021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 5022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 5022.
[0155] In this embodiment, by calling the program or instructions stored in memory 502, specifically the program or instructions stored in application program 5022, processor 501 executes the method steps provided in each method embodiment, including, for example: Acquire the speech data of the speaker; Speech recognition is performed on the speech data to obtain the corresponding text data; Extracting emotional expression features from speech data; By using a multimodal large model, based on text data and emotional expression features, expressive feedback information of the subject is generated.
[0156] The methods disclosed in the above embodiments of this disclosure can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in processor 501. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 502. Processor 501 reads the information in memory 502 and, in conjunction with its hardware, completes the steps of the above method.
[0157] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described above, or combinations thereof.
[0158] For software implementation, the techniques described herein can be implemented by units that perform the functions described above. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or external to the processor.
[0159] The electronic device provided in this embodiment may be as follows: Figure 5 The electronic device shown can execute all the steps of the interactive expression training methods described above, thereby achieving the technical effects of the interactive expression training methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, further details are not provided here.
[0160] This disclosure also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include combinations of the above types of memory.
[0161] When one or more programs in the storage medium can be executed by one or more processors to achieve the interactive expression training method described above that is executed on the electronic device side.
[0162] The processor described above is used to execute the interactive representation training program stored in memory to implement the following steps of the interactive representation training method executed on the electronic device side: Acquire the speech data of the speaker; Speech recognition is performed on the speech data to obtain the corresponding text data; Extracting emotional expression features from speech data; By using a multimodal large model, based on text data and emotional expression features, expressive feedback information of the subject is generated.
[0163] Furthermore, the computer program product provided in this disclosure embodiment may include computer-readable code that, when executed on a device, causes a processor in the device to implement the steps of the interactive representation training method executed on the electronic device side: Acquire the speech data of the speaker; Speech recognition is performed on the speech data to obtain the corresponding text data; Extracting emotional expression features from speech data; By using a multimodal large model, based on text data and emotional expression features, expressive feedback information of the subject is generated.
[0164] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0165] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0166] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0167] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An interactive expression training method, characterized in that, The method includes: Acquire the speech data of the speaker; Perform speech recognition on the speech data to obtain the text data corresponding to the speech data; Extract the emotional expression features of the speech data; Using a multimodal large model, based on the text data and the emotional expression features, the expressive feedback information of the subject is generated.
2. The method according to claim 1, characterized in that, The step of generating the expressive feedback information of the subject through a multimodal large model, based on the text data and the emotional expression features, includes: Determine the audience information of the speaker; Determine the cognitive decision-making characteristics of the audience represented by the audience information; Using a multimodal large model, based on the text data, the emotional expression features, and the cognitive decision features, the expressive feedback information of the subject is generated.
3. The method according to claim 2, characterized in that, The determination of the cognitive decision-making characteristics of the audience objects represented by the audience information includes: Determine the cognitive decision-making characteristics of each of the multiple audience objects represented by the audience information; and The step of generating the expressive feedback information of the subject through a multimodal large model, based on the text data, the emotional expression features, and the cognitive decision features, includes: Using a multimodal large model, based on the text data, the emotional expression features, and the cognitive decision features of each audience member, expressive feedback information for each audience member is generated for the expressing subject. The expression feedback information of each audience member to the expression subject is aggregated to obtain the aggregated expression feedback information of the expression subject.
4. The method according to claim 1, characterized in that, The step of generating the expressive feedback information of the subject through a multimodal large model, based on the text data and the emotional expression features, includes: Obtain the demonstration content corresponding to the voice data; Using a multimodal large model, based on the demonstration content, the text data, and the emotional expression features, the expressive feedback information of the subject is generated.
5. The method according to claim 1, characterized in that, The step of generating the expressive feedback information of the subject through a multimodal large model, based on the text data and the emotional expression features, includes: Using a multimodal large model, based on the text data and the emotional expression features, fragment feedback information and whole-segment feedback information of the expressing subject are generated; The fragment feedback information and the whole segment feedback information are determined as the expression feedback information of the expression subject.
6. The method according to any one of claims 1-5, characterized in that, The step of generating the expressive feedback information of the subject through a multimodal large model, based on the text data and the emotional expression features, includes: Obtain the historical expression feedback information of the expressing subject; Using a multimodal large model, based on the text data, the emotional expression features, and the historical expression feedback information, the current expression feedback information of the expression subject is generated.
7. The method according to claim 6, characterized in that, The step of generating the current expression feedback information of the subject through a multimodal large model, based on the text data, the sentiment expression features, and the historical expression feedback information, includes: In response to any expression feedback information of the expressed subject that has been generated, the feedback type of the expression feedback information is determined, and the expression feedback information is stored according to the feedback type, and the stored expression feedback information is used as historical expression feedback information. Based on the voice data, the target historical expression feedback information is determined from the stored historical expression feedback information of each feedback type; Using a multimodal large model, based on the text data, the emotional expression features, and the target historical expression feedback information, the current expression feedback information of the expression subject is generated.
8. An interactive expression training device, characterized in that, The device includes: The first acquisition unit is configured to acquire the speech data of the subject of expression; The recognition unit is configured to: perform speech recognition on the speech data to obtain text data corresponding to the speech data; The extraction unit is configured to: extract the emotional expression features of the speech data; The first generation unit is configured to generate the expression feedback information of the expression subject based on the text data and the emotional expression features through a multimodal large model.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.