Eloquence training method, device and medium based on data retrieval and fusion

By calculating the matching degree between user query words and preset resources and evaluating eloquence ability, and combining learning habits to generate personalized feedback suggestions, the accuracy and personalization problems of eloquence training methods in the existing technology are solved, and the effect of eloquence training is improved.

CN118410139BActive Publication Date: 2025-09-16CHINA NEW LINE EDUCATION TECH CO LTD

Patent Information

Application Number
CN202410429472.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-09-16
Estimated Expiration
2044-04-10

AI Technical Summary

Technical Problem

The existing eloquence training methods have low accuracy in eloquence assessment, making it difficult to fully capture all dimensions and details of a speech, and are unable to generate personalized customized training plans based on the user's actual usage needs.

Method used

By obtaining the user's query terms and eloquence information, calculating the matching degree between the query terms and preset resources, combining the eloquence ability assessment results and learning habits, generating personalized feedback suggestions, and evaluating according to the preset eloquence indicators to generate eloquence training methods.

Benefits of technology

It improves the accuracy and personalization of eloquence training, can provide targeted training for users' weaknesses, and improve learning efficiency and motivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118410139B_ABST
    Figure CN118410139B_ABST
Patent Text Reader

Abstract

The present invention discloses an eloquence training method, device, and medium based on data retrieval and fusion. The method comprises: sorting the resources by calculating the matching degree between the query word and the resource to obtain a first resource; obtaining the user's eloquence ability evaluation result, and sorting the first resource according to the result to obtain knowledge feedback suggestions; filtering the resources according to learning habits to generate personalized feedback suggestions; evaluating eloquence information according to eloquence indicators, and generating evaluation feedback suggestions according to the evaluation results; and generating an eloquence training method based on the above feedback suggestions to train the user. The present invention proposes an eloquence training method, device, and medium based on data retrieval and fusion. By considering multiple different dimensions, corresponding feedback suggestions are generated, and a comprehensive eloquence training method is constructed to train the user. This can solve the problem of difficulty in generating personalized and comprehensive eloquence training programs based on the user's actual usage needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech training, and in particular to an eloquence training method, device and medium based on data retrieval and fusion. Background Art

[0002] Eloquence, as the art of expressing information, communicating, and influencing others, is an indispensable skill in learning and life, especially when speaking and expressing oneself in public. Eloquence training can help users build self-confidence, which not only impacts their academic performance but also extends to all aspects of life, helping them better cope with challenges and difficulties. Existing eloquence training methods use deep learning technology to extract features from multimodal data such as speech, semantics, and body language. These features are used to evaluate eloquence, and then, based on the eloquence evaluation results, eloquence training methods are generated to provide users with specialized eloquence training.

[0003] However, the eloquence assessment accuracy of existing eloquence training methods is not high, and it is difficult to fully capture all dimensions and details of the speech, resulting in poor training results; moreover, the current eloquence training program unilaterally provides training content to users, and it is difficult to generate corresponding personalized customization plans based on the user's actual usage needs, and it is impossible to provide accurate feedback based on the individual differences of users. Summary of the Invention

[0004] The present invention provides an eloquence training method, device and medium based on data retrieval and fusion, so as to solve the problem that it is difficult to generate a personalized and comprehensive eloquence training program according to the actual usage needs of users.

[0005] In order to solve the above problems, the present invention provides an eloquence training method based on data retrieval and fusion, comprising:

[0006] Obtain the user's query terms and eloquence information;

[0007] By calculating the matching degree between the query word and the preset resources, the preset resources are sorted to obtain a first resource;

[0008] Obtaining an evaluation result of the user's eloquence ability from the eloquence information, and sorting the first resources according to the evaluation result of the eloquence ability to obtain knowledge feedback suggestions;

[0009] Filtering the preset resources according to the user's learning habits to generate personalized feedback suggestions;

[0010] Evaluate the eloquence information according to the preset eloquence indicators, and generate evaluation feedback suggestions based on the evaluation results;

[0011] An eloquence training method is generated according to the knowledge feedback suggestion, the personalized feedback suggestion and the evaluation feedback suggestion, and the user is trained according to the eloquence training method.

[0012] The present invention sorts the preset resources according to the matching degree between the query words and the preset resources, which can make the content presented by the obtained first resource highly relevant to the query words. Since the query words reflect the actual eloquence training needs of the user, it helps the subsequently generated eloquence training method to provide substantial help to the user; the first resources are sorted according to the eloquence ability evaluation results to obtain knowledge feedback suggestions, which can be used to conduct targeted training on the user's weak points in eloquence. Personalized feedback suggestions are generated according to the user's learning habits, which can help the user to absorb and master new knowledge faster when encountering new concepts similar to existing eloquence knowledge, thereby improving learning efficiency. The eloquence information is evaluated according to the preset eloquence indicators, which accurately reflects the user's real eloquence status from an objective perspective, so that the generated evaluation feedback suggestions can clarify the user's training focus and direction.

[0013] Compared with the existing technology, the present invention constructs an eloquence training method to train users based on feedback suggestions in multiple dimensions, wherein the preset resources are sorted according to the matching degree between the query words and the preset resources, which helps to meet the user's actual eloquence training needs; personalized feedback suggestions are generated according to learning habits, which can improve the user's learning efficiency and learning motivation; evaluation feedback suggestions are generated according to the user's eloquence evaluation results, which can comprehensively and objectively carry out targeted training on the user's eloquence weaknesses, thereby solving the problem of difficulty in generating personalized and comprehensive eloquence training plans based on the user's actual usage needs.

[0014] As a preferred solution, obtaining the user's eloquence ability evaluation result from the eloquence information is specifically as follows:

[0015] According to a preset value range, the eloquence information is scored from the modal dimensions of voice, text, and video to obtain a speech performance value;

[0016] The dependency value and dynamic change value between the user's knowledge level and the speech performance value at different times are captured through a probabilistic network to obtain the user's oral ability evaluation result.

[0017] This preferred solution scores the eloquence information, and the resulting speech performance value can fully reflect the user's performance in the modal dimensions of voice, text, and video; it captures the dependency value and dynamic change value between the user's knowledge level and speech performance value at different times, and can promptly discover the user's progress and regression, so that the eloquence ability assessment results can objectively reflect the user's eloquence knowledge level and eloquence performance.

[0018] As a preferred solution, the preset resources are filtered according to the user's learning habits to generate personalized feedback suggestions, specifically:

[0019] Filtering the preset resources according to the historical resources that the user has learned, and obtaining similar resources that have a high similarity with the historical resources;

[0020] Obtaining preferred resources of other trainees based on the similarity of training resources and similarity of ratings between the user and other trainees;

[0021] Filtering the preset resources using a linear weighted method to obtain mixed resources;

[0022] The personalized feedback suggestion is generated according to the similar resources, the preferred resources and the mixed resources.

[0023] This preferred solution obtains similar resources based on the historical resources that the user has learned. This helps users to use their existing cognitive framework to absorb and master new knowledge more quickly when encountering new concepts that are similar to their existing knowledge. In addition, learning similar content helps promote communication and resonance, helping to improve the user's learning efficiency.

[0024] In addition, by using the similarity of training resources and ratings between users and other trainers as criteria, obtaining the preferred resources of other trainers can identify the user's eloquence level at the first time. Since the preferred resources have usually been screened and sorted, they can help users quickly locate high-quality eloquence learning materials, thus avoiding getting lost in massive amounts of information and saving time on information screening. By learning from the preferred resources that others have learned, users can be exposed to knowledge and viewpoints in different fields, which helps to broaden their own eloquence skills.

[0025] As a preferred solution, the preset resources are filtered using a linear weighted method to obtain mixed resources, specifically:

[0026] A linear weighting method is used to assign corresponding weights to the similar resources and the preferred resources, and the degree of integration of the similar resources and the preferred resources is adjusted by a preset coefficient to obtain the mixed resources.

[0027] This preferred solution fuses resources through multimodal data fusion technology, which can further extract feature data from similar resources and preferred resources to improve the effectiveness of information.

[0028] As a preferred solution, the eloquence information is evaluated according to the preset eloquence indicators, and evaluation feedback suggestions are generated based on the evaluation results, specifically:

[0029] Evaluate the eloquence information from the perspectives of speech expression ability, body language ability, and situation control ability, respectively, to obtain a first evaluation result, a second evaluation result, and a third evaluation result;

[0030] According to the target speech scenario and target eloquence performance value of the user, the weights of the first evaluation result, the second evaluation result and the third evaluation result are adjusted, and the evaluation feedback suggestion is generated according to the adjustment results.

[0031] This preferred solution evaluates the eloquence information using pre-set eloquence indicators, so that the obtained first evaluation result, second evaluation result, and third evaluation result can accurately reflect the user's voice expression ability, body language ability, and situation control ability respectively;

[0032] Moreover, since the ability to convey ideas clearly and accurately through language helps users communicate effectively with others, appropriate body language helps build trust and consensus, and the ability to control situations helps users adjust their communication strategies according to the occasion and the object in different social and professional environments, so as to be able to handle them with ease, the evaluation feedback suggestions generated based on the evaluation results can improve the user's overall eloquence.

[0033] As a preferred solution, the eloquence information is evaluated from the perspective of speech expression ability to obtain a first evaluation result, specifically:

[0034] Using a preset formula set, the first evaluation result is obtained by measuring the pitch value and the speech speed value of the eloquence information;

[0035] The formula set is established by considering the weights of different dimensions and the influence of time factors.

[0036] As a preferred solution, the preset resources are specifically:

[0037] Acquire initial resources; wherein the initial resources refer to modal data including voice, text, and video;

[0038] Calculating the similarity between different data in the initial resource to obtain a resource similarity value;

[0039] The data in the initial resources are sorted according to the resource similarity values ​​to obtain the preset resources.

[0040] As a preferred solution, after the user is trained according to the eloquence training method, the method further includes:

[0041] Based on the similarity between the preset source domain and target domain features, the migration weight of the feature is calculated;

[0042] Using the migration weights to migrate the features of the source domain to the target domain, and generating a migration learning strategy based on the migration results;

[0043] Formulate predefined strategies based on the learning characteristics and verbal ability of the user;

[0044] Optimizing the predefined strategy according to the user's learning behavior and social data to obtain an individual feature learning strategy;

[0045] The user is trained in eloquence according to the transfer learning strategy and the individual feature learning strategy.

[0046] This preferred solution uses transfer weights to transfer features from the source domain to the target domain, and generates a transfer learning strategy based on the transfer results. This allows the knowledge from the source domain to be directly applied to the target domain, thereby reducing the data requirements of the target domain and accelerating the learning process. In addition, the transfer learning approach can help data better adapt to the characteristics of the target domain, making the transfer learning strategy more adaptable to speech learning tasks in different domains.

[0047] Moreover, by formulating predefined strategies based on the user's learning characteristics and the strength of their eloquence, we can focus on improving the user's weak points in eloquence and enhance the quality of their expression.

[0048] As a preferred solution, after the user is trained according to the eloquence training method, the method further includes:

[0049] By measuring the overall eloquence level of the user, a comprehensive score is obtained;

[0050] By evaluating the user's ability performance in different eloquence dimensions, a sub-dimension score is obtained;

[0051] Generate personalized feedback based on the comprehensive score and the sub-dimension scores, and display the personalized feedback.

[0052] In this preferred option, improving eloquence is a process of continuous learning and progress. By displaying personalized feedback based on comprehensive scores and dimensional scores, users can understand their own progress speed and direction, thereby adjusting their learning strategies and methods more specifically and clarifying their own training focus and direction.

[0053] As a preferred solution, after obtaining the first resource, the method further includes:

[0054] The first resources are sorted according to the user's knowledge level target and resource preference to obtain sorted first resources.

[0055] This preferred solution sorts the first resources according to the user's knowledge level goal and resource preference, so that the data of the first resources can better meet the user's actual usage needs.

[0056] The present invention also provides an eloquence training device based on data retrieval and fusion, comprising an information module, a retrieval module, an evaluation module, a habit module, an eloquence module and a comprehensive module;

[0057] Wherein, the information module is used to obtain the user's query words and eloquence information;

[0058] The retrieval module is configured to calculate a matching degree between the query term and the preset resources, sort the preset resources, and obtain a first resource;

[0059] The evaluation module is configured to obtain an evaluation result of the user's eloquence ability from the eloquence information, and sort the first resources according to the evaluation result of the eloquence ability to obtain knowledge feedback suggestions;

[0060] The habit module is used to filter the preset resources according to the user's learning habits and generate personalized feedback suggestions;

[0061] The eloquence module is used to evaluate the eloquence information according to the preset eloquence indicators and generate evaluation feedback suggestions based on the evaluation results;

[0062] The comprehensive module is used to generate an eloquence training method based on the knowledge feedback suggestion, the personalized feedback suggestion and the evaluation feedback suggestion, and train the user according to the eloquence training method.

[0063] As a preferred solution, the evaluation module includes a performance unit and a network unit;

[0064] The performance unit is configured to score the eloquence information based on the modal dimensions of voice, text, and video according to a preset value range to obtain a speech performance value;

[0065] The network unit is used to capture the dependency value and dynamic change value between the user's knowledge level at different times and the speech performance value through a probabilistic network, and obtain the user's oral ability evaluation result.

[0066] As a preferred solution, the habit module includes a similarity unit, a preference unit, a mixing unit and a feedback unit;

[0067] The similarity unit is configured to filter the preset resources according to the historical resources that the user has learned, and obtain similar resources that have a high degree of similarity with the historical resources;

[0068] The optimization unit is configured to obtain the preferred resources of other trainees based on the training resource similarity and rating similarity between the user and other trainees;

[0069] The mixing unit is configured to filter the preset resources in a linear weighted manner to obtain mixed resources;

[0070] The feedback unit is configured to generate the personalized feedback suggestion based on the similar resources, the preferred resources, and the mixed resources.

[0071] As a preferred solution, the mixing unit is specifically:

[0072] A linear weighting method is used to assign corresponding weights to the similar resources and the preferred resources, and the degree of integration of the similar resources and the preferred resources is adjusted by a preset coefficient to obtain the mixed resources.

[0073] As a preferred solution, the eloquence module includes an ability unit and a weight unit;

[0074] The ability unit is used to evaluate the eloquence information from the perspectives of speech expression ability, body language ability, and situation control ability, respectively, to obtain a first evaluation result, a second evaluation result, and a third evaluation result;

[0075] The weighting unit is used to adjust the weights of the first evaluation result, the second evaluation result and the third evaluation result according to the target speech scenario and the target eloquence performance value of the user, and generate the evaluation feedback suggestion according to the adjustment result.

[0076] As a preferred solution, the capability unit is specifically:

[0077] Using a preset formula set, the first evaluation result is obtained by measuring the pitch value and the speech speed value of the eloquence information;

[0078] The formula set is established by considering the weights of different dimensions and the influence of time factors.

[0079] As a preferred solution, the preset resources are specifically:

[0080] Acquire initial resources; wherein the initial resources refer to modal data including voice, text, and video;

[0081] Calculating the similarity between different data in the initial resource to obtain a resource similarity value;

[0082] The data in the initial resources are sorted according to the resource similarity values ​​to obtain the preset resources.

[0083] As a preferred solution, after the comprehensive module, it also includes a migration unit, a strategy unit, a prefabrication unit, an optimization unit and a training unit;

[0084] The migration unit is used to calculate the migration weight of the feature according to the similarity between the preset source domain and target domain features;

[0085] The strategy unit is configured to migrate the features of the source domain to the target domain using the migration weights, and generate a migration learning strategy based on the migration results;

[0086] The prefabrication unit is used to formulate a predefined strategy based on the learning characteristics and oral ability of the user;

[0087] The optimization unit is configured to optimize the predefined strategy according to the user's learning behavior and social data to obtain an individual feature learning strategy;

[0088] The training unit is used to perform eloquence training on the user according to the transfer learning strategy and the individual feature learning strategy.

[0089] As a preferred solution, after the user is trained according to the eloquence training method, it further includes an overall unit, a dimension unit and a display unit;

[0090] The overall unit is used to obtain a comprehensive score by measuring the overall eloquence level of the user;

[0091] The dimension unit is used to evaluate the user's ability performance in different eloquence dimensions to obtain sub-dimension scores;

[0092] The display unit is used to generate personalized feedback according to the comprehensive score and the sub-dimensional scores, and to display the personalized feedback.

[0093] As a preferred solution, after obtaining the first resource, an adjustment unit is further included;

[0094] The adjustment unit is configured to sort the first resources according to the user's knowledge level target and resource preference to obtain sorted first resources.

[0095] The present invention also provides a storage medium having a computer program stored thereon. The computer program is called and executed by a computer to implement the above-mentioned eloquence training method based on data retrieval and fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1This is a flow chart of an eloquence training method based on data retrieval and fusion provided by an embodiment of the present invention;

[0097] Figure 2 It is a structural diagram of an eloquence training device based on data retrieval and fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0098] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0099] In the description of this application, it should be understood that the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first," "second," and "third" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, "several" means two or more.

[0100] An eloquence training method based on data retrieval and fusion provided by an embodiment of the present invention is mainly used in situations where it is difficult to generate a personalized and comprehensive eloquence training program based on the actual usage needs of the user to train the user's eloquence.

[0101] Example 1:

[0102] See also Figure 1 The embodiment of the present invention provides an eloquence training method based on data retrieval and fusion, including S1 to S6. The specific implementation steps are as follows:

[0103] S1. Obtain the user's query terms and eloquence information.

[0104] In step S1 of the embodiment of the present invention, S1 includes S1.1 to S1.2; S1.1 is the process of obtaining query words and eloquence information; S1.2 is the process of processing the data, specifically:

[0105] S1.1. Obtaining the user's query term and eloquence information; wherein the eloquence information includes the user's speech video and user data;

[0106] In addition, user data includes information such as learning goals, knowledge level, learning style, oral ability, and user historical learning behavior data; user historical learning behavior data includes learned resources, learning time and learning results, etc.

[0107] S1.2. Extract features such as learning goals, knowledge level, learning style, and oral ability from user data, and extract features such as learning preferences, learning time, and learning effects from user historical learning behavior data for subsequent calculations.

[0108] S2. Calculate the matching degree between the query term and the preset resources, sort the preset resources, and obtain a first resource.

[0109] In step S2 of the embodiment of the present invention, S2 includes S2.1 to S2.5; wherein S2.1 is the process of acquiring initial resources, S2.2 is the process of defining eloquence dimension indicators, S2.3 is the process of performing similarity calculation, S2.4 is the process of performing matching calculation, and S2.5 is the process of performing resource preference calculation, specifically:

[0110] S2.1. Acquire initial resources; initial resources refer to modal data including voice, text, and video;

[0111] Initial resources are stored in a knowledge resource library, which is updated regularly by adding new learning resources. The knowledge resource library stores a vast amount of speech knowledge and learning resources, including:

[0112] Speech theory: covers the basic principles, techniques and methods of speech;

[0113] Speech Cases: Provide excellent speech cases for users to refer to and learn from;

[0114] Speech materials: Provide rich materials for users to practice speeches;

[0115] Eloquence training course: Provide systematic eloquence training courses to help users improve their speaking skills.

[0116] S2.2. Define eloquence dimension indicators and update them regularly;

[0117] Among them, the eloquence dimension indicators include:

[0118] Speech dimensions: voice intonation, volume, speaking speed, pauses, and clarity;

[0119] Text dimensions: vocabulary, sentence structure, grammar, logic and semantics;

[0120] Video dimensions: body language, facial expressions, eye contact, and demeanor;

[0121] Knowledge dimension: the depth, breadth and accuracy of the knowledge in the speech;

[0122] Emotional dimension: emotional expression and appeal of the speech, etc.

[0123] S2.3. Calculate the similarity between different data in the initial resource using the improved multimodal fusion weighted cosine similarity according to the first formula to obtain a resource similarity value;

[0124] Sort the data in the initial resources according to the resource similarity value to obtain the preset resources;

[0125] Among them, the first formula is:

[0126]

[0127] Among them, x and y are vector representations of two resource data, n is the number of eloquence dimension indicators, and w i is the weight of the i-th eloquence dimension indicator, α i , β i and γ i They are the weights of the three modes of speech, text, and video. and are the values ​​of x in speech, text, and video modes at time t on the i-th eloquence dimension index, and are the values ​​of y in voice, text, and video modes at time t on the i-th eloquence dimension indicator; f(x) is the user preference function, which can be learned based on the user's historical behavior; t is time.

[0128] S2.4. Calculate the matching degree between the query term and the preset resources according to the second formula in the search method, and sort the preset resources according to the matching degree to obtain the first resource;

[0129] Among them, the second formula is:

[0130] R u ={r i |Sim(q u ,r i )>θ∧F i (K i )>δ∧P u (r i )>γ}

[0131] Among them, R u is the retrieved resource collection, q u is the query vector of user u, Sim(q u ,r i ) is the query vector q u With resource vector r i similarity, θ is the similarity threshold, F i (K i ) is the knowledge level function of resource i, δ is the knowledge level threshold, P u (r i) is the preference degree of user u for resource i, and γ is the preference degree threshold.

[0132] S2.5. Calculate the user's knowledge level goal and resource preference according to the third formula, and sort the first resources according to the calculation results to obtain sorted first resources;

[0133] Among them, the third formula is:

[0134] R u (t) = {r i |Sim(r i ,C u (t),u,t)>θ∧L i =L u ∧P u (r i )>δ}

[0135] Among them, R u (t) is the resource set recommended to user u at time t, r i is resource i, Sim(r i ,C u (t),u,t) is the similarity between resource i and user u’s eloquence at time t, θ is the similarity threshold, L i is the knowledge level of resource i, L u is the target knowledge level of user u, P u (r i ) is the preference degree of user u for resource i, and δ is the preference degree threshold.

[0136] In this embodiment, S2.5 sorts the first resource according to the user's knowledge level goal and resource preference, so that the data of the first resource can better meet the user's actual usage needs.

[0137] S3. Obtain the user's eloquence ability evaluation result from the eloquence information, and sort the first resource according to the eloquence ability evaluation result to obtain knowledge feedback suggestions.

[0138] Step S3 of the embodiment of the present invention is specifically as follows:

[0139] According to the preset value range, the eloquence information is scored from the modal dimensions of voice, text and video to obtain the speech performance value;

[0140] Using the fourth formula, the probability network is used to capture the dependency and dynamic change between the user's knowledge level and speech performance at different times, and the user's oral ability evaluation result is obtained. The first resource is then sorted according to the oral ability evaluation result to obtain knowledge feedback suggestions.

[0141] Among them, the fourth formula is:

[0142]

[0143] Among them, C t represents the user's speaking ability at time t, represents the user's speech performance in the i-th modality at time t; i∈{v,a,t} represents the voice, text, and video modalities respectively, and the speech performance of each modality is a continuous variable that can take a value between 0 and 1, representing the quality of the modality performance; B t represents the user's knowledge level at time t; C t-1 represents the user's speaking ability at time t-1, which is a discrete variable and is related to C t same;

[0144] represents the probability of a user's eloquence at time t, given the user's speech performance in voice, text, and video modalities, knowledge level at time t, and eloquence at time t-1; P(B) represents the probability of a user’s speech performance in the i-th mode at time t, given the user’s eloquence at time t. t |C t ) represents the probability of the user’s knowledge level at time t given the user’s oral ability at time t; P(C t |C t-1 ) represents the dependency of the user’s eloquence ability at time t on his eloquence ability at time t-1; represents the joint probability of user’s speech performance in voice, text and video modes at time t; P(B t ) represents the probability of the user’s knowledge level at time t.

[0145] It should be noted that in the above formula, the specific meaning of each operation formula is:

[0146] Sim(x,y): the similarity between resources x and y;

[0147] f(x): user preference function;

[0148] R u (t): The set of resources recommended to user u at time t;

[0149] The user's speaking ability at time t;

[0150] R u : The retrieved resource collection;

[0151] Sim(q u ,r i ):Query vector q uWith resource vector r i similarity;

[0152] F i (K i ): knowledge level function of resource i;

[0153] P u (r i ): User u’s preference for resource i.

[0154] In this embodiment, S3 scores the eloquence information, and the obtained speech performance value can fully reflect the user's performance in the modal dimensions of voice, text and video; it captures the dependency value and dynamic change value between the user's knowledge level and speech performance value at different times, and can timely discover the user's progress and regression, so that the eloquence ability evaluation result can objectively reflect the user's eloquence knowledge level and eloquence performance.

[0155] S4. Filter preset resources based on the user's learning habits and generate personalized feedback suggestions.

[0156] In step S4 of the embodiment of the present invention, S4 includes S4.1 to S4.4. S4.1 is a process of calculating based on historical resources to obtain similar resources, S4.2 is a process of calculating based on the resources of other trainers to obtain preferred resources, S4.3 is a process of constructing mixed resources, and S4.4 is a process of generating personalized feedback suggestions, specifically:

[0157] S4.1. Using the first resource formula, filter the preset resources based on the historical resources that the user has learned, and obtain similar resources that have a high degree of similarity with the historical resources;

[0158] Among them, the first resource formula is:

[0159]

[0160] Where u is the user, i is the resource, t is the time; m is the modality type, m = 1 represents voice, m = 2 represents text, and m = 3 represents video; α m is the weight of the mth mode, is the multimodal attention mechanism function, is the feature vector of the advanced eloquence index of the mth mode of user u at time t, is the characteristic vector of the advanced eloquence index of the mth mode of resource i at time t, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; β m is the mth modal fusion coefficient, which is used to adjust the weights of different modalities.

[0161] S4.2. Using the second resource formula, based on the training resource similarity and rating similarity between the user and other trainers, obtain the preferred resources of other trainers;

[0162] Among them, the second resource formula is:

[0163]

[0164] Among them, u and v are two users, t is time, n is the number of resources, w i is the weight of the i-th resource, R ui (t) is the rating of resource i by user u at time t, R u (t) is the average rating of user u at time t, R vi (t) is the rating of resource i by user v at time t, R v (t) is the average rating of user v at time t, λ is the time decay factor, f(P u (i)) is the preference function of user u for resource i.

[0165] S4.3. Using a linear weighting method, assign corresponding weights to similar resources and preferred resources according to the third resource formula, and adjust the degree of integration of similar resources and preferred resources using a preset coefficient to obtain a mixed resource;

[0166] Among them, the third resource formula is:

[0167] R(u,i,t)=α(t)·Sim(u,i,t)+β(t)·Sim(u,v i (t),t)·f(P u (i))·e -λt γ

[0168] Among them, u is the user, i is the resource, and t is the time; α(t) and β(t) are two time-related coefficients used to adjust the weights of the two recommendation methods; Sim(u,i,t) is the content-based recommendation similarity, Sim(u,v i (t), t) is the recommendation similarity based on collaborative filtering, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; γ is the fusion coefficient, which is used to adjust the fusion degree of the two recommendation methods.

[0169] In this embodiment, S4.3 uses multimodal data fusion technology to fuse resources, which can further extract feature data from similar resources and preferred resources, thereby improving the effectiveness of information.

[0170] Moreover, the mathematical formula used can effectively fuse different modal data and extract more representative features, further improving the accuracy of the evaluation.

[0171] S4.4. Generate personalized feedback suggestions based on similar resources, preferred resources, and mixed resources.

[0172] Overall, in this embodiment, S4 obtains similar resources with high similarity based on the historical resources that the user has learned. This can help the user, when encountering new concepts similar to existing eloquence knowledge, use the existing cognitive framework to absorb and master new knowledge more quickly. In addition, learning similar content helps promote communication and resonance, helping to improve the user's learning efficiency.

[0173] In addition, by using the similarity of training resources and ratings between users and other trainers as criteria, obtaining the preferred resources of other trainers can identify the user's eloquence level at the first time. Since the preferred resources have usually been screened and sorted, they can help users quickly locate high-quality eloquence learning materials, thus avoiding getting lost in massive amounts of information and saving time on information screening. By learning from the preferred resources that others have learned, users can be exposed to knowledge and viewpoints in different fields, which helps to broaden their own eloquence skills.

[0174] S5. Evaluate the eloquence information according to the preset eloquence indicators, and generate evaluation feedback suggestions based on the evaluation results.

[0175] In step S5 of the embodiment of the present invention, S5 includes S5.1 to S5.4; wherein S5.1 is a process of calculating a first evaluation result based on pitch and speech rate, S5.2 is a process of calculating a second evaluation result based on body language, S5.3 is a process of calculating a third evaluation result based on situational control ability, and S5.4 is a process of generating evaluation feedback suggestions, specifically:

[0176] S5.1. Using a preset formula set, obtain a first evaluation result by measuring the pitch value and speech rate value of the eloquence information; wherein the first evaluation result includes a pitch variation amplitude evaluation result, a pitch variation flexibility evaluation result, a pitch control skill evaluation result, a speech rate variation amplitude evaluation result, and a speech rate variation rhythm evaluation result;

[0177] The formula set is established by considering the weights of different dimensions and the influence of time factors, including the formula for pitch change amplitude, pitch change flexibility, pitch control skills, speaking speed change amplitude, and speaking speed change rhythm:

[0178] ① Pitch variation:

[0179]

[0180] Among them, F p is the pitch change amplitude, N is the number of sampling points of the speech segment, p i is the pitch value of the i-th sample; is the weight of the pitch variation, which is adjusted by experts according to different speech scenarios and goals; p is the attenuation coefficient of the pitch change amplitude, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0181] ②Flexibility of pitch changes:

[0182]

[0183] Among them, F v is the pitch change flexibility, N is the number of sampling points of the speech segment, and p i is the pitch value of the i-th sample; is the weight of the flexibility of pitch variation, which is adjusted by experts according to different speech scenarios and goals; v is the attenuation coefficient of the flexibility of pitch change, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0184] ③ Tone control skills:

[0185]

[0186] Among them, F t is the pitch control technique, N is the number of sampling points of the speech segment, p i is the pitch value of the i-th sample; p target is the target pitch value, which is adjusted by experts according to different speech scenarios and goals; is the weight of the pitch control skill, which is adjusted by experts according to different speech scenarios and goals; t is the attenuation coefficient of the pitch control technique, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0187] ④Speech speed variation:

[0188]

[0189] Among them, V p is the amplitude of speech rate change, N is the number of sampling points of speech segment, v i is the speech rate value of the i-th sample; is the weight of the speech speed variation, which is adjusted by experts according to different speech scenarios and goals; p is the attenuation coefficient of the speech rate change amplitude, which is used to reflect the influence of time factors; ti is the timestamp of the i-th sample.

[0190] ⑤ Rhythm of speech speed changes:

[0191]

[0192] Among them, V r is the rhythm of speech rate change, N is the number of sampling points of the speech segment, v i is the speech rate value of the i-th sample; is the weight of the rhythm of speech speed change, which is adjusted by experts according to different speech scenarios and goals; r is the attenuation coefficient of the speech rate change rhythm.

[0193] S5.2. Evaluate the eloquence information from the perspective of body language ability to obtain a second evaluation result;

[0194] Specifically:

[0195] First, human posture estimation technology is used to extract body language features such as posture and expression;

[0196] Then, the results of the posture assessment are calculated: indicators such as the naturalness, coordination, and expressiveness of body language are calculated. A time decay factor and user preference function are added to adjust the assessment results.

[0197] Secondly, calculate the results of expression evaluation: calculate indicators such as the richness, accuracy and appeal of expressions; add time decay factors and user preference functions to adjust the evaluation results.

[0198] Finally, the results of posture evaluation and expression evaluation are integrated to obtain the final result of body language evaluation, that is, the second evaluation result.

[0199] S5.3. Evaluate the eloquence information from the perspective of situational control ability to obtain the third evaluation result;

[0200] Specifically: First, use speech recognition technology to extract features such as the speaker's interaction with the audience and control of the stage.

[0201] Then, the results of the audience interaction evaluation are calculated: indicators such as the frequency, effectiveness, and attractiveness of the audience interaction are calculated; the time decay factor and user preference function are added to adjust the evaluation results.

[0202] Secondly, the results of the stage control evaluation are calculated: indicators such as the stability, fluency and appeal of the stage control are calculated; the time decay factor and user preference function are added to adjust the evaluation results.

[0203] Finally, the results of the audience interaction evaluation and the stage control evaluation are integrated to obtain the final result of the situation control evaluation, which is the third evaluation result.

[0204] S5.4. Based on the user's target speech scenario and target eloquence performance value, the weights of the first evaluation result, the second evaluation result, and the third evaluation result are adjusted, and evaluation feedback suggestions are generated based on the adjustment results.

[0205] It should be noted that the first evaluation result, the second evaluation result, and the third evaluation result in S5 of this embodiment are obtained by evaluating according to the speech evaluation indicators;

[0206] Among them, speech evaluation indicators include:

[0207] Language expression: pronunciation, speaking speed, intonation, vocabulary and grammar, etc.;

[0208] Logical thinking: speech structure, logic and argumentation;

[0209] Body language: eye contact, facial expressions, movements and postures, etc.;

[0210] Situational control: managing nervousness, controlling the stage, and interacting with the audience, etc.

[0211] Overall, in this embodiment S5, eloquence information is evaluated using pre-set eloquence indicators, so that the obtained first evaluation result, second evaluation result, and third evaluation result can accurately reflect the user's speech expression ability, body language ability, and situation control ability, respectively;

[0212] Moreover, since the ability to convey ideas clearly and accurately through language helps users communicate effectively with others, appropriate body language helps build trust and consensus, and the ability to control situations helps users adjust their communication strategies according to the occasion and the object in different social and professional environments, so as to be able to handle them with ease, the evaluation feedback suggestions generated based on the evaluation results can improve the user's overall eloquence.

[0213] S6. Generate an eloquence training method based on the knowledge feedback suggestions, personalized feedback suggestions, and evaluation feedback suggestions, and train the user based on the eloquence training method.

[0214] Step S6 of the embodiment of the present invention is specifically as follows:

[0215] An eloquence training method is generated based on the knowledge feedback suggestions, personalized feedback suggestions and evaluation feedback suggestions, and users are trained based on the eloquence training method.

[0216] The first embodiment also includes S7, using meta-learning transfer and adaptive learning strategies to train the user's eloquence.

[0217] In step S7 of the embodiment of the present invention, S7 includes S7.1 to S7.4; wherein S7.1 is the process of preparing data, S7.2 is the process of building a transfer model, S7.3 is the process of generating a transfer learning strategy, and S7.4 is the process of generating an individual feature learning strategy and training the user, specifically:

[0218] S7.1. Prepare data, including:

[0219] Collect speech data from source and target domains;

[0220] Extract the eloquence dimension features of the source and target domains, including voice expression, body language, and situational control;

[0221] The Kullback-Leibler divergence of the distribution of eloquence dimension indicators of the source and target domain tasks was calculated to measure the degree of difference between the eloquence ability assessment results of the two domains.

[0222] S7.2. First, perform model training:

[0223] Initialize the parameters of the preset model, including meta-learning algorithm parameters and transfer learning parameters;

[0224] For each task: ① Sample a batch of data, including speech videos, texts, and eloquence dimension indicators; ② Calculate the eloquence dimension indicators and adjust the indicator weights based on the Kullback-Leibler divergence; ③ Use an improved meta-learning algorithm, such as MAML++, Reptile++, or Prototypical Networks++, to update the model parameters.

[0225] Next, perform model evaluation:

[0226] Evaluate the model's performance using data from the target domain, including evaluation metrics for speech expression, body language, and contextual control.

[0227] The performance of the model under different eloquence dimension indicators is analyzed, and the model is tuned according to the evaluation results to obtain the migration model.

[0228] S7.3. Calculate the similarity between the preset source domain and target domain features using a similarity measurement method; wherein the method may be, for example, cosine similarity or Euclidean distance;

[0229] Calculate the transfer weight of each feature based on the similarity to obtain the feature's transfer weight, and adjust the transfer weight based on the weight of the eloquence dimension indicator to make the source domain and target domain features more compatible;

[0230] Use transfer weights to transfer features from the source domain to the target domain, and integrate features from the target domain to train the transfer model;

[0231] The transfer model is enhanced through instance selection, and the source and target domain features are more closely matched through instance migration. The generalization ability of the transfer model is improved through instance weighting, resulting in an optimized transfer model. A transfer learning strategy is then generated based on the optimized transfer model.

[0232] The specific instance migration methods are as follows:

[0233] First, similarity between instances in the source and target domains is calculated using a similarity metric, such as KNN or spectral clustering.

[0234] Then, instances related to the target domain are selected based on similarity, taking into account the differences in eloquence dimension indicators;

[0235] Finally, the model of the target domain is trained using the selected instances and the knowledge of the source domain is combined for model enhancement.

[0236] The specific instance migration methods are as follows:

[0237] First, feature extraction techniques are used to extract features of instances in the source and target domains;

[0238] Then, the distance between the source and target domain features is calculated using a similarity measure;

[0239] Finally, the features of the source domain instances are adjusted according to the distance to make them more closely match the target domain features.

[0240] The specific method of instance weighting is as follows:

[0241] First, the differences in the eloquence dimension indicators between source and target domain instances are calculated;

[0242] Then, the transfer weight of each instance is calculated based on the difference;

[0243] Finally, the transferred weights are used to train the model in the target domain to improve the generalization ability of the model.

[0244] It should be noted that the source domain refers to the domain where knowledge and data already exist, and the target domain refers to the new domain where knowledge needs to be learned and applied;

[0245] In this embodiment of the present invention, the source domain may be:

[0246] ① A database with rich speech data and eloquence assessment results, such as public speaking competitions and eloquence training courses;

[0247] ② Experts or speakers with specific speaking styles or eloquence skills, such as speakers, teachers, and actors.

[0248] Target areas could be:

[0249] ① Areas where there is a lack of speech data or eloquence assessment results, such as emerging industries, specific groups of people, etc.

[0250] ② Individuals who need to learn specific speaking styles or eloquence skills, such as students, job seekers, and business people.

[0251] S7.4. Develop predefined strategies based on the user’s learning characteristics and verbal skills;

[0252] Optimize predefined strategies based on users' learning behaviors and social data to obtain individual feature learning strategies. Continuously optimize and adjust individual feature learning strategies through trial and error and reward mechanisms.

[0253] Users are trained in eloquence based on transfer learning strategies and individual feature learning strategies.

[0254] Overall, in this embodiment S7, the features of the source domain are transferred to the target domain using transfer weights, and a transfer learning strategy is generated based on the transfer results. This allows the knowledge of the source domain to be directly applied to the target domain, thereby reducing the data requirements of the target domain and accelerating the learning process. In addition, the transfer learning approach can help the data better adapt to the characteristics of the target domain, making the transfer learning strategy more suitable for speech learning tasks in different domains.

[0255] Furthermore, by developing predefined strategies based on a user's learning characteristics and speaking ability, we can focus on improving weaknesses in their speaking and enhance the quality of their presentation. Furthermore, this embodiment leverages meta-learning principles to extract common patterns across speech learning tasks. Combined with transfer learning methods based on eloquence dimensional indicators, this enables the transfer model to rapidly transfer learned speaking strategies to new domains or effectively adapt to the individual characteristics of different users.

[0256] The first embodiment further includes S8, displaying the user's performance through intelligent speech evaluation and feedback.

[0257] In step S8 of the embodiment of the present invention, S8 includes S8.1 to S8.4; wherein S8.1 is the process of generating personalized feedback for intelligent speech evaluation, S8.2 is the process of calculating the comprehensive score, S8.3 is the process of calculating the sub-dimension scores, and S8.4 is the process of generating and displaying personalized feedback, specifically:

[0258] S8.1. Utilize technologies such as speech recognition, semantic analysis, and natural language processing, combined with eloquence dimension indicators, to conduct a multi-dimensional and multi-level evaluation of the user's speech performance, and generate personalized feedback on intelligent speech evaluation based on the evaluation results.

[0259] S8.2. According to the comprehensive formula, the user's overall eloquence level is measured to obtain a comprehensive score;

[0260] Among them, the comprehensive formula is:

[0261]

[0262] Among them, S i is the comprehensive score of the ith eloquence dimension indicator, is the weight of the jth feature under the i-th eloquence dimension indicator, is the value of the jth feature on the i-th sample under the i-th eloquence dimension indicator, λ is the weight of the eloquence dimension indicator, is the evaluation result of the kth evaluation indicator under the i-th eloquence dimension indicator, μ is the personal trait weight, is the evaluation result of the lth personal trait indicator under the i-th eloquence dimension indicator, ν is the speech situation weight, is the evaluation result of the qth speech situation indicator under the i-th eloquence dimension indicator, is the cognitive ability weight, is the evaluation result of the tth cognitive ability indicator under the i-th eloquence dimension indicator, ψ is the weight of social and cultural factors, It is the evaluation result of the u-th social and cultural factor indicator under the i-th eloquence dimension indicator.

[0263] S8.3. Based on the dimensionality formula, evaluate the user's performance in different eloquence dimensions to obtain dimensionality scores;

[0264] The dimension formula is:

[0265]

[0266] in, is the score of the ith eloquence dimension indicator under the dth dimension, is the weight of the jth feature under the dth dimension under the i-th eloquence dimension indicator, is the value of the jth feature in the dth dimension under the i-th eloquence dimension indicator on the i-th sample, λ (d) is the weight of the eloquence dimension indicator under the dth dimension, is the evaluation result of the kth evaluation indicator under the dth dimension under the i-th eloquence dimension indicator, ω (d) is the weight of personal traits under the d-th dimension, It is the evaluation result of the lth personal trait indicator under the dth dimension under the i-th eloquence dimension indicator.

[0267] S8.4. Generate personalized feedback based on the overall score and the sub-dimension scores, taking into account both strengths and weaknesses, and present the personalized feedback in the form of text, voice, and video.

[0268] Among them, the specific ways to consider the advantages and disadvantages are:

[0269] For strengths: provide encouragement and maintenance suggestions.

[0270] Targeting deficiencies: Generate personalized improvement suggestions based on factors such as personal traits, speech context, cognitive ability, socio-cultural factors, speech goals, audience perception, and speech topic.

[0271] Personalized feedback includes:

[0272] ① Specific description of the behavior, such as "You are speaking too fast, please slow down."

[0273] ② Recommended practice tasks, such as "It is recommended to practice abdominal breathing to improve voice control."

[0274] ③Recommended learning resources, such as "Recommended article "How to Improve Body Language in Speech"".

[0275] Overall, in this embodiment S8, the improvement of eloquence is a process of continuous learning and progress. By displaying personalized feedback based on the comprehensive score and dimensional scores, users can understand their own progress speed and direction, thereby adjusting their learning strategies and methods more specifically and clarifying their own training focus and direction.

[0276] Overall, this embodiment also has the following beneficial effects:

[0277] The present invention sorts the preset resources according to the matching degree between the query words and the preset resources, which can make the content presented by the obtained first resource highly relevant to the query words. Since the query words reflect the actual eloquence training needs of the user, it helps the subsequently generated eloquence training method to provide substantial help to the user; the first resources are sorted according to the eloquence ability evaluation results to obtain knowledge feedback suggestions, which can be used to conduct targeted training on the user's weak points in eloquence. Personalized feedback suggestions are generated according to the user's learning habits, which can help the user to absorb and master new knowledge faster when encountering new concepts similar to existing eloquence knowledge, thereby improving learning efficiency. The eloquence information is evaluated according to the preset eloquence indicators, which accurately reflects the user's real eloquence status from an objective perspective, so that the generated evaluation feedback suggestions can clarify the user's training focus and direction.

[0278] In addition, this embodiment adopts multimodal data fusion technology, which comprehensively considers multiple aspects of information such as voice, semantics and body language, and can fully capture all dimensions and details of the speech, thereby improving the accuracy of the evaluation; based on factors such as the user's personal characteristics, speech context and cognitive ability, it generates targeted improvement suggestions, which can help users improve their eloquence more effectively.

[0279] Example 2:

[0280] See also Figure 2 , an embodiment of the present invention provides an eloquence training device based on data retrieval and fusion, comprising an information module 10, a retrieval module 20, an evaluation module 30, a habit module 40, an eloquence module 50 and a comprehensive module 60;

[0281] Among them, the information module 10 is used to obtain the user's query words and eloquence information;

[0282] A search module 20 is configured to sort the preset resources by calculating the matching degree between the query term and the preset resources to obtain a first resource;

[0283] An evaluation module 30 is configured to obtain an evaluation result of the user's eloquence ability from the eloquence information, and sort the first resources according to the evaluation result of the eloquence ability to obtain knowledge feedback suggestions;

[0284] Habit module 40, used to filter preset resources according to the user's learning habits and generate personalized feedback suggestions;

[0285] The eloquence module 50 is used to evaluate the eloquence information according to the preset eloquence indicators and generate evaluation feedback suggestions based on the evaluation results;

[0286] The comprehensive module 60 is used to generate an eloquence training method based on the knowledge feedback suggestions, the personalized feedback suggestions and the evaluation feedback suggestions, and train the user according to the eloquence training method.

[0287] In one embodiment, the information module 10 includes an information unit and a processing unit, wherein the information unit is a process of obtaining query words and eloquence information; and the processing unit is a process of processing data, specifically:

[0288] An information unit, used to obtain the user's query words and eloquence information; wherein the eloquence information includes the user's speech video and user data;

[0289] In addition, user data includes information such as learning goals, knowledge level, learning style, oral ability, and user historical learning behavior data; user historical learning behavior data includes learned resources, learning time and learning results, etc.

[0290] The processing unit is used to extract features such as learning goals, knowledge level, learning style and oral ability from user data, and to extract features such as learning preferences, learning time and learning effect from user historical learning behavior data for subsequent calculations.

[0291] In one embodiment, the retrieval module 20 includes a resource unit, a dimension unit, a similarity unit, a matching unit, and an adjustment unit. The resource unit is a process for acquiring initial resources, the dimension unit is a process for defining eloquence dimension indicators, the similarity unit is a process for performing similarity calculations, the matching unit is a process for performing matching calculations, and the adjustment unit is a process for performing resource preference calculations. Specifically,

[0292] A resource unit, used to obtain initial resources; wherein the initial resources refer to modal data including voice, text, and video;

[0293] The resource unit is also used to store initial resources in the knowledge resource library and update the knowledge resource library by regularly adding new learning resources. The knowledge resource library stores a large amount of speech knowledge and learning resources, including:

[0294] Speech theory: covers the basic principles, techniques and methods of speech;

[0295] Speech Cases: Provide excellent speech cases for users to refer to and learn from;

[0296] Speech materials: Provide rich materials for users to practice speeches;

[0297] Eloquence training course: Provide systematic eloquence training courses to help users improve their speaking skills.

[0298] Dimension unit, used to define eloquence dimension indicators and regularly update them;

[0299] Among them, the eloquence dimension indicators include:

[0300] Speech dimensions: voice intonation, volume, speaking speed, pauses, and clarity;

[0301] Text dimensions: vocabulary, sentence structure, grammar, logic and semantics;

[0302] Video dimensions: body language, facial expressions, eye contact, and demeanor;

[0303] Knowledge dimension: the depth, breadth and accuracy of the knowledge in the speech;

[0304] Emotional dimension: emotional expression and appeal of the speech, etc.

[0305] A similarity unit is used to calculate the similarity between different data in the initial resource using the improved multimodal fusion weighted cosine similarity according to the first formula to obtain a resource similarity value;

[0306] The similarity unit is further used to sort the data in the initial resources according to the resource similarity value to obtain the preset resources;

[0307] Among them, the first formula is:

[0308]

[0309] Among them, x and y are vector representations of two resource data, n is the number of eloquence dimension indicators, and w i is the weight of the i-th eloquence dimension indicator, α i , β i and γ i They are the weights of the three modes of speech, text, and video. and are the values ​​of x in speech, text, and video modes at time t on the i-th eloquence dimension index, and are the values ​​of y in voice, text, and video modes at time t on the i-th eloquence dimension indicator; f(x) is the user preference function, which can be learned based on the user's historical behavior; t is time.

[0310] a matching degree unit, configured to calculate the matching degree between the query term and the preset resources according to the second formula in a retrieval manner, and sort the preset resources according to the matching degree to obtain the first resource;

[0311] Among them, the second formula is:

[0312] R u ={r i |Sim(q u ,r i )>θ∧F i (K i )>δ∧P u (r i )>γ}

[0313] Among them, R u is the retrieved resource collection, q u is the query vector of user u, Sim(q u ,r i ) is the query vector q u With resource vector r i similarity, θ is the similarity threshold, F i (K i ) is the knowledge level function of resource i, δ is the knowledge level threshold, P u(r i ) is the preference degree of user u for resource i, and γ is the preference degree threshold.

[0314] an adjusting unit, configured to calculate the user's knowledge level target and resource preference according to a third formula, and sort the first resources according to the calculation result to obtain sorted first resources;

[0315] Among them, the third formula is:

[0316] R u (t) = {r i |Sim(r i ,C u (t),u,t)>θ∧L i =L u ∧P u (r i )>δ}

[0317] Among them, R u (t) is the resource set recommended to user u at time t, r i is resource i, Sim(r i ,C u (t),u,t) is the similarity between resource i and user u’s eloquence at time t, θ is the similarity threshold, L i is the knowledge level of resource i, L u is the target knowledge level of user u, P u (r i ) is the preference degree of user u for resource i, and δ is the preference degree threshold.

[0318] In this embodiment, the adjustment unit sorts the first resources according to the user's knowledge level target and resource preference, so that the data of the first resources can better meet the user's actual usage needs.

[0319] In one embodiment, the evaluation module 30 includes a presentation unit and a network unit;

[0320] The performance unit is used to score the eloquence information from the modal dimensions of voice, text and video according to a preset value range to obtain a speech performance value;

[0321] a network unit configured to use the fourth formula to capture the dependency relationship and dynamic change value between the user's knowledge level and speech performance value at different times through a probabilistic network, obtain an evaluation result of the user's speaking ability, and sort the first resource according to the evaluation result of the speaking ability to obtain knowledge feedback suggestions;

[0322] Among them, the fourth formula is:

[0323]

[0324] Among them, C t represents the user's speaking ability at time t, represents the user's speech performance in the i-th modality at time t; i∈{v,a,t} represents the voice, text, and video modalities respectively, and the speech performance of each modality is a continuous variable that can take a value between 0 and 1, representing the quality of the modality performance; B t represents the user's knowledge level at time t; C t-1 represents the user's speaking ability at time t-1, which is a discrete variable and is related to C t same;

[0325] represents the probability of a user's eloquence at time t, given the user's speech performance in voice, text, and video modalities, knowledge level at time t, and eloquence at time t-1; P(B) represents the probability of a user’s speech performance in the i-th mode at time t, given the user’s eloquence at time t. t |C t ) represents the probability of the user’s knowledge level at time t given the user’s oral ability at time t; P(C t |C t-1 ) represents the dependency of the user’s eloquence ability at time t on his eloquence ability at time t-1; represents the joint probability of user’s speech performance in voice, text and video modes at time t; P(B t ) represents the probability of the user’s knowledge level at time t.

[0326] It should be noted that in the above formula, the specific meaning of each operation formula is:

[0327] Sim(x,y): the similarity between resources x and y;

[0328] f(x): user preference function;

[0329] R u (t): The set of resources recommended to user u at time t;

[0330] The user's speaking ability at time t;

[0331] R u : The retrieved resource collection;

[0332] Sim(q u ,r i ):Query vector q u With resource vector r i similarity;

[0333] F i (K i ): knowledge level function of resource i;

[0334] P u (r i ): User u’s preference for resource i.

[0335] The evaluation module 30 of this embodiment scores the eloquence information, and the obtained speech performance value can fully reflect the user's performance in the modal dimensions of voice, text and video; it captures the dependency value and dynamic change value between the user's knowledge level and speech performance value at different times, and can timely discover the user's progress and regression, so that the eloquence ability evaluation result can objectively reflect the user's eloquence knowledge level and eloquence performance.

[0336] In one embodiment, the habit module 40 includes a similarity unit, a preferred unit, a mixed unit, and a feedback unit. The similarity unit is a process of calculating based on historical resources to obtain similar resources. The preferred unit is a process of calculating based on the resources of other trainees to obtain preferred resources. The mixed unit is a process of constructing mixed resources. The feedback unit is a process of generating personalized feedback suggestions. Specifically,

[0337] A similarity unit is configured to use the first resource formula to filter preset resources according to historical resources that the user has learned, and obtain similar resources that have a high degree of similarity with the historical resources;

[0338] Among them, the first resource formula is:

[0339]

[0340] Where u is the user, i is the resource, t is the time; m is the modality type, m = 1 represents voice, m = 2 represents text, and m = 3 represents video; α m is the weight of the mth mode, is the multimodal attention mechanism function, is the feature vector of the advanced eloquence index of the mth mode of user u at time t, is the characteristic vector of the advanced eloquence index of the mth mode of resource i at time t, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; β m is the mth modal fusion coefficient, which is used to adjust the weights of different modalities.

[0341] A preference unit, configured to use a second resource formula to obtain preferred resources of other trainers based on the training resource similarity and rating similarity between the user and other trainers;

[0342] Among them, the second resource formula is:

[0343]

[0344] Among them, u and v are two users, t is time, n is the number of resources, w i is the weight of the i-th resource, R ui (t) is the rating of resource i by user u at time t, R u (t) is the average rating of user u at time t, R vi (t) is the rating of resource i by user v at time t, R v (t) is the average rating of user v at time t, λ is the time decay factor, f(P u (i)) is the preference function of user u for resource i.

[0345] a mixing unit, configured to assign corresponding weights to similar resources and preferred resources respectively using a linear weighting method according to the third resource formula, and adjust the degree of integration of similar resources and preferred resources using a preset coefficient to obtain mixed resources;

[0346] Among them, the third resource formula is:

[0347] R(u,i,t)=α(t)·Sim(u,i,t)+β(t)·Sim(u,v i (t),t)·f(P u (i))·e -λt γ

[0348] Among them, u is the user, i is the resource, and t is the time; α(t) and β(t) are two time-related coefficients used to adjust the weights of the two recommendation methods; Sim(u,i,t) is the content-based recommendation similarity, Sim(u,v i (t), t) is the recommendation similarity based on collaborative filtering, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; γ is the fusion coefficient, which is used to adjust the fusion degree of the two recommendation methods.

[0349] The hybrid unit of this embodiment fuses resources through multimodal data fusion technology, which can further extract feature data from similar resources and preferred resources, thereby improving the effectiveness of information.

[0350] Moreover, the mathematical formula used can effectively fuse different modal data and extract more representative features, further improving the accuracy of the evaluation.

[0351] The feedback unit is used to generate personalized feedback suggestions based on similar resources, preferred resources and mixed resources.

[0352] Overall, the habit module 40 of this embodiment obtains similar resources with high similarity based on the historical resources that the user has learned. This can help the user to use the existing cognitive framework to absorb and master new knowledge more quickly when encountering new concepts similar to existing eloquence knowledge. In addition, learning similar content helps promote communication and resonance, helping to improve the user's learning efficiency.

[0353] In addition, by using the similarity of training resources and ratings between users and other trainers as criteria, obtaining the preferred resources of other trainers can identify the user's eloquence level at the first time. Since the preferred resources have usually been screened and sorted, they can help users quickly locate high-quality eloquence learning materials, thus avoiding getting lost in massive amounts of information and saving time on information screening. By learning from the preferred resources that others have learned, users can be exposed to knowledge and viewpoints in different fields, which helps to broaden their own eloquence skills.

[0354] In one embodiment, the eloquence module 50 includes a first capability unit, a second capability unit, a third capability unit, and a weighting unit. The first capability unit is a process for calculating a first evaluation result based on pitch and speech rate, the second capability unit is a process for calculating a second evaluation result based on body language, and the third capability unit is a process for calculating a third evaluation result based on situational control ability. The weighting unit is a process for generating evaluation feedback suggestions, specifically:

[0355] A first capability unit is configured to use a preset formula set to measure the pitch value and speech rate value of the eloquence information to obtain a first evaluation result; wherein the first evaluation result includes a pitch variation amplitude evaluation result, a pitch variation flexibility evaluation result, a pitch control skill evaluation result, a speech rate variation amplitude evaluation result, and a speech rate variation rhythm evaluation result;

[0356] The formula set is established by considering the weights of different dimensions and the influence of time factors, including the formula for pitch change amplitude, pitch change flexibility, pitch control skills, speaking speed change amplitude, and speaking speed change rhythm:

[0357] ① Pitch variation:

[0358]

[0359] Among them, F p is the pitch change amplitude, N is the number of sampling points of the speech segment, p i is the pitch value of the i-th sample; is the weight of the pitch variation, which is adjusted by experts according to different speech scenarios and goals; p is the attenuation coefficient of the pitch change amplitude, which is used to reflect the influence of time factors; ti is the timestamp of the i-th sample.

[0360] ②Flexibility of pitch changes:

[0361]

[0362] Among them, F v is the pitch change flexibility, N is the number of sampling points of the speech segment, and p i is the pitch value of the i-th sample; is the weight of the flexibility of pitch variation, which is adjusted by experts according to different speech scenarios and goals; v is the attenuation coefficient of the flexibility of pitch change, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0363] ③ Tone control skills:

[0364]

[0365] Among them, F t is the pitch control technique, N is the number of sampling points of the speech segment, p i is the pitch value of the i-th sample; p target is the target pitch value, which is adjusted by experts according to different speech scenarios and goals; is the weight of the pitch control skill, which is adjusted by experts according to different speech scenarios and goals; t is the attenuation coefficient of the pitch control technique, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0366] ④Speech speed variation:

[0367]

[0368] Among them, V p is the amplitude of speech rate change, N is the number of sampling points of speech segment, v i is the speech rate value of the i-th sample; is the weight of the speech speed variation, which is adjusted by experts according to different speech scenarios and goals; p is the attenuation coefficient of the speech rate change amplitude, which is used to reflect the influence of time factors; t i is the timestamp of the i-th sample.

[0369] ⑤ Rhythm of speech speed changes:

[0370]

[0371] Among them, V ris the rhythm of speech rate change, N is the number of sampling points of the speech segment, v i is the speech rate value of the i-th sample; is the weight of the rhythm of speech speed change, which is adjusted by experts according to different speech scenarios and goals; r is the attenuation coefficient of the speech rate change rhythm.

[0372] The second ability unit is used to evaluate the eloquence information from the perspective of body language ability to obtain a second evaluation result;

[0373] Specifically:

[0374] First, human posture estimation technology is used to extract body language features such as posture and expression;

[0375] Then, the results of the posture assessment are calculated: indicators such as the naturalness, coordination, and expressiveness of body language are calculated. A time decay factor and user preference function are added to adjust the assessment results.

[0376] Secondly, calculate the results of expression evaluation: calculate indicators such as the richness, accuracy and appeal of expressions; add time decay factors and user preference functions to adjust the evaluation results.

[0377] Finally, the results of posture evaluation and expression evaluation are integrated to obtain the final result of body language evaluation, that is, the second evaluation result.

[0378] The third ability unit is used to evaluate the eloquence information from the perspective of situation control ability and obtain the third evaluation result;

[0379] Specifically: First, use speech recognition technology to extract features such as the speaker's interaction with the audience and control of the stage.

[0380] Then, the results of the audience interaction evaluation are calculated: indicators such as the frequency, effectiveness, and attractiveness of the audience interaction are calculated; the time decay factor and user preference function are added to adjust the evaluation results.

[0381] Secondly, the results of the stage control evaluation are calculated: indicators such as the stability, fluency and appeal of the stage control are calculated; the time decay factor and user preference function are added to adjust the evaluation results.

[0382] Finally, the results of the audience interaction evaluation and stage control evaluation are integrated to obtain the final result of the situation control evaluation, which is the third evaluation result.

[0383] The weighting unit is used to adjust the weights of the first evaluation result, the second evaluation result and the third evaluation result according to the user's target speech scenario and target eloquence performance value, and generate evaluation feedback suggestions based on the adjustment results.

[0384] It should be noted that the first evaluation result, the second evaluation result, and the third evaluation result of the eloquence module 50 of this embodiment are obtained by evaluating according to the speech evaluation index;

[0385] Among them, speech evaluation indicators include:

[0386] Language expression: pronunciation, speaking speed, intonation, vocabulary and grammar, etc.;

[0387] Logical thinking: speech structure, logic and argumentation;

[0388] Body language: eye contact, facial expressions, movements and postures, etc.;

[0389] Situational control: managing nervousness, controlling the stage, and interacting with the audience, etc.

[0390] Overall, the eloquence module 50 of this embodiment evaluates the eloquence information using pre-set eloquence indicators, so that the obtained first evaluation result, second evaluation result, and third evaluation result can accurately reflect the user's speech expression ability, body language ability, and situation control ability, respectively;

[0391] Moreover, since the ability to convey ideas clearly and accurately through language helps users communicate effectively with others, appropriate body language helps build trust and consensus, and the ability to control situations helps users adjust their communication strategies according to the occasion and the object in different social and professional environments, so as to be able to handle them with ease, the evaluation feedback suggestions generated based on the evaluation results can improve the user's overall eloquence.

[0392] In one embodiment, the integration module 60 is specifically:

[0393] An eloquence training method is generated based on the knowledge feedback suggestions, personalized feedback suggestions and evaluation feedback suggestions, and users are trained based on the eloquence training method.

[0394] The second embodiment further includes a migration module 70;

[0395] Among them, the transfer module 70 is used to use meta-learning transfer and adaptive learning strategies to train users' eloquence.

[0396] In one embodiment, the migration module 70 includes a domain unit, a model unit, a migration unit, a strategy unit, a prefabrication unit, an optimization unit, and a training unit. The domain unit is the process of data preparation, the model unit is the process of building a migration model, the migration unit and the strategy unit are the processes of generating a migration learning strategy, and the prefabrication unit, the optimization unit, and the training unit are the processes of generating an individual feature learning strategy and training users. Specifically,

[0397] Domain units, used for data preparation, include:

[0398] Collect speech data from source and target domains;

[0399] Extract the eloquence dimension features of the source and target domains, including voice expression, body language, and situational control;

[0400] The Kullback-Leibler divergence of the distribution of eloquence dimension indicators of the source and target domain tasks was calculated to measure the degree of difference between the eloquence ability assessment results of the two domains.

[0401] The model unit is used to first perform model training:

[0402] Initialize the parameters of the preset model, including meta-learning algorithm parameters and transfer learning parameters;

[0403] For each task: ① Sample a batch of data, including speech videos, texts, and eloquence dimension indicators; ② Calculate the eloquence dimension indicators and adjust the indicator weights based on the Kullback-Leibler divergence; ③ Use an improved meta-learning algorithm, such as MAML++, Reptile++, or Prototypical Networks++, to update the model parameters.

[0404] Next, perform model evaluation:

[0405] Evaluate the model's performance using data from the target domain, including evaluation metrics for speech expression, body language, and contextual control.

[0406] The performance of the model under different eloquence dimension indicators is analyzed, and the model is tuned according to the evaluation results to obtain the migration model.

[0407] A migration unit, configured to calculate the similarity between the features of the source domain and the target domain using a similarity measurement method, such as cosine similarity or Euclidean distance;

[0408] The transfer unit is also used to calculate the transfer weight of each feature based on the similarity, obtain the transfer weight of the feature, and adjust the transfer weight in combination with the weight of the eloquence dimension indicator to make the source domain and target domain features more compatible;

[0409] The strategy unit is used to transfer the features of the source domain to the target domain using the transfer weights, and to integrate the features of the target domain for transfer model training;

[0410] The strategy unit is also used to enhance the transfer model through instance selection, make the source domain and target domain features more closely matched through instance migration, and improve the generalization ability of the transfer model through instance weighting. This results in an optimized transfer model and generates a transfer learning strategy based on the optimized transfer model.

[0411] The specific instance migration methods are as follows:

[0412] First, similarity between instances in the source and target domains is calculated using a similarity metric, such as KNN or spectral clustering.

[0413] Then, instances related to the target domain are selected based on similarity, taking into account the differences in eloquence dimension indicators;

[0414] Finally, the model of the target domain is trained using the selected instances and the knowledge of the source domain is combined for model enhancement.

[0415] The specific instance migration methods are as follows:

[0416] First, feature extraction techniques are used to extract features of instances in the source and target domains;

[0417] Then, the distance between the source and target domain features is calculated using a similarity measure;

[0418] Finally, the features of the source domain instances are adjusted according to the distance to make them more closely match the target domain features.

[0419] The specific method of instance weighting is as follows:

[0420] First, the differences in the eloquence dimension indicators between source and target domain instances are calculated;

[0421] Then, the transfer weight of each instance is calculated based on the difference;

[0422] Finally, the transferred weights are used to train the model in the target domain to improve the generalization ability of the model.

[0423] It should be noted that the source domain refers to the domain where knowledge and data already exist, and the target domain refers to the new domain where knowledge needs to be learned and applied;

[0424] In this embodiment of the present invention, the source domain may be:

[0425] ① A database with rich speech data and eloquence assessment results, such as public speaking competitions and eloquence training courses;

[0426] ② Experts or speakers with specific speaking styles or eloquence skills, such as speakers, teachers, and actors.

[0427] Target areas could be:

[0428] ① Areas where there is a lack of speech data or eloquence assessment results, such as emerging industries, specific groups of people, etc.

[0429] ② Individuals who need to learn specific speaking styles or eloquence skills, such as students, job seekers, and business people.

[0430] Prefabricated units for developing predefined strategies based on the user's learning characteristics and verbal skills;

[0431] The optimization unit is used to optimize the predefined strategy based on the user's learning behavior and social data to obtain the individual feature learning strategy, and continuously optimize and adjust the individual feature learning strategy through trial and error and reward mechanisms;

[0432] The training unit is used to train users' eloquence based on the transfer learning strategy and the individual feature learning strategy.

[0433] Overall, the transfer module 70 of this embodiment uses transfer weights to transfer features from the source domain to the target domain and generates a transfer learning strategy based on the transfer results. This allows the knowledge from the source domain to be directly applied to the target domain, thereby reducing the data requirements of the target domain and accelerating the learning process. Furthermore, the transfer learning approach can help data better adapt to the characteristics of the target domain, making the transfer learning strategy more adaptable to speech learning tasks in different domains.

[0434] Furthermore, by developing predefined strategies based on a user's learning characteristics and speaking ability, we can focus on improving weaknesses in their speaking and enhance the quality of their presentation. Furthermore, this embodiment leverages meta-learning principles to extract common patterns across speech learning tasks. Combined with transfer learning methods based on eloquence dimensional indicators, this enables the transfer model to rapidly transfer learned speaking strategies to new domains or effectively adapt to the individual characteristics of different users.

[0435] The second embodiment further includes a display module 80;

[0436] Among them, the display module 80 is used to display the user's performance through intelligent speech evaluation and feedback.

[0437] In one embodiment, the presentation module 80 includes an intelligent evaluation unit, an overall unit, a dimension unit, and a presentation unit. The intelligent evaluation unit generates personalized feedback on intelligent speech evaluation, the overall unit calculates the overall score, the dimension unit calculates the sub-dimension scores, and the presentation unit generates and presents personalized feedback. Specifically,

[0438] The intelligent evaluation unit is used to use technologies such as speech recognition, semantic analysis and natural language processing, combined with eloquence dimension indicators, to conduct multi-dimensional and multi-level evaluation of the user's speech performance, and generate personalized feedback on intelligent speech evaluation based on the evaluation results.

[0439] The overall unit is used to obtain a comprehensive score by measuring the user's overall eloquence level according to a comprehensive formula;

[0440] Among them, the comprehensive formula is:

[0441]

[0442] Among them, S i is the comprehensive score of the ith eloquence dimension indicator, is the weight of the jth feature under the i-th eloquence dimension indicator, is the value of the jth feature on the i-th sample under the i-th eloquence dimension indicator, λ is the weight of the eloquence dimension indicator, is the evaluation result of the kth evaluation indicator under the i-th eloquence dimension indicator, μ is the personal trait weight, is the evaluation result of the lth personal trait indicator under the i-th eloquence dimension indicator, ν is the speech situation weight, is the evaluation result of the qth speech situation indicator under the i-th eloquence dimension indicator, is the cognitive ability weight, is the evaluation result of the tth cognitive ability indicator under the i-th eloquence dimension indicator, ψ is the weight of social and cultural factors, It is the evaluation result of the u-th social and cultural factor indicator under the i-th eloquence dimension indicator.

[0443] Dimension units are used to evaluate the user's performance in different eloquence dimensions according to the dimension formula to obtain dimension scores;

[0444] The dimension formula is:

[0445]

[0446] in, is the score of the ith eloquence dimension indicator under the dth dimension, is the weight of the jth feature under the dth dimension under the i-th eloquence dimension indicator, is the value of the jth feature in the dth dimension under the i-th eloquence dimension indicator on the i-th sample, λ (d) is the weight of the eloquence dimension indicator under the dth dimension, is the evaluation result of the kth evaluation indicator under the dth dimension under the i-th eloquence dimension indicator, ω (d) is the weight of personal traits under the d-th dimension, It is the evaluation result of the lth personal trait indicator under the dth dimension under the i-th eloquence dimension indicator.

[0447] A display unit is used to generate personalized feedback based on the comprehensive score and the sub-dimensional scores by considering the strengths and weaknesses of the dimensions, and to display the personalized feedback in the form of text, voice, and video;

[0448] Among them, the specific ways to consider the advantages and disadvantages are:

[0449] For strengths: provide encouragement and maintenance suggestions.

[0450] Targeting deficiencies: Generate personalized improvement suggestions based on factors such as personal traits, speech context, cognitive ability, socio-cultural factors, speech goals, audience perception, and speech topic.

[0451] Personalized feedback includes:

[0452] ① Specific description of the behavior, such as "You are speaking too fast, please slow down."

[0453] ② Recommended practice tasks, such as "It is recommended to practice abdominal breathing to improve voice control."

[0454] ③Recommended learning resources, such as "Recommended article "How to Improve Body Language in Speech"".

[0455] Overall, in the display module 80 of this embodiment, the improvement of eloquence is a process of continuous learning and progress. By displaying personalized feedback generated based on the comprehensive score and dimensional scores, users can understand their own progress speed and direction, thereby adjusting their learning strategies and methods more specifically and clarifying their own training focus and direction.

[0456] Overall, this embodiment also has the following beneficial effects:

[0457] The present invention sorts the preset resources according to the matching degree between the query words and the preset resources, which can make the content presented by the obtained first resource highly relevant to the query words. Since the query words reflect the actual eloquence training needs of the user, it helps the subsequently generated eloquence training method to provide substantial help to the user; the first resources are sorted according to the eloquence ability evaluation results to obtain knowledge feedback suggestions, which can be used to conduct targeted training on the user's weak points in eloquence. Personalized feedback suggestions are generated according to the user's learning habits, which can help the user to absorb and master new knowledge faster when encountering new concepts similar to existing eloquence knowledge, thereby improving learning efficiency. The eloquence information is evaluated according to the preset eloquence indicators, which accurately reflects the user's real eloquence status from an objective perspective, so that the generated evaluation feedback suggestions can clarify the user's training focus and direction.

[0458] In addition, this embodiment adopts multimodal data fusion technology, which comprehensively considers multiple aspects of information such as voice, semantics and body language, and can fully capture all dimensions and details of the speech, thereby improving the accuracy of the evaluation; based on factors such as the user's personal characteristics, speech context and cognitive ability, it generates targeted improvement suggestions, which can help users improve their eloquence more effectively.

[0459] Example 3:

[0460] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the eloquence training method based on data retrieval and fusion;

[0461] Wherein, the eloquence training method based on data retrieval and fusion, if implemented in the form of a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0462] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for training eloquence based on data retrieval and fusion, characterized in that: include: Obtain the user's query terms and eloquence information; By calculating the matching degree between the query word and the preset resources, the preset resources are sorted to obtain a first resource; Obtaining an evaluation result of the user's eloquence ability from the eloquence information, and sorting the first resources according to the evaluation result of the eloquence ability to obtain knowledge feedback suggestions; Filtering the preset resources according to the user's learning habits to generate personalized feedback suggestions; Specifically, the first resource formula is used to filter the preset resources according to the historical resources that the user has learned, and similar resources with a high similarity to the historical resources are obtained; Based on the training resource similarity and rating similarity between the user and other trainees, the preferred resources of the other trainees are obtained; using a linear weighting method, corresponding weights are assigned to the similar resources and the preferred resources, and the degree of integration of the similar resources and the preferred resources is adjusted by a preset coefficient to obtain a mixed resource; and the personalized feedback suggestions are generated based on the similar resources, the preferred resources, and the mixed resources; Evaluate the eloquence information according to the preset eloquence indicators, and generate evaluation feedback suggestions based on the evaluation results; generating an eloquence training method according to the knowledge feedback suggestion, the personalized feedback suggestion, and the evaluation feedback suggestion, and training the user according to the eloquence training method; Among them, the first resource formula is: Where u is the user, i is the resource, t is the time; m is the modality type, m = 1 represents voice, m = 2 represents text, and m = 3 represents video; α m is the weight of the mth mode, is the multimodal attention mechanism function, is the feature vector of the advanced eloquence index of the mth mode of user u at time t, is the characteristic vector of the advanced eloquence index of the mth mode of resource i at time t, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; β m is the mth modal fusion coefficient, which is used to adjust the weights of different modalities.

2. The method for eloquence training based on data retrieval and fusion according to claim 1, characterized in that: Obtaining the user's eloquence ability evaluation result from the eloquence information, specifically: According to a preset value range, the eloquence information is scored from the modal dimensions of voice, text, and video to obtain a speech performance value; The dependency value and dynamic change value between the user's knowledge level and the speech performance value at different times are captured through a probabilistic network to obtain the user's oral ability evaluation result.

3. The method for training eloquence based on data retrieval and fusion according to claim 1, characterized in that: Evaluate the eloquence information according to the preset eloquence indicators, and generate evaluation feedback suggestions based on the evaluation results, specifically: Evaluate the eloquence information from the perspectives of speech expression ability, body language ability, and situation control ability, respectively, to obtain a first evaluation result, a second evaluation result, and a third evaluation result; According to the target speech scenario and target eloquence performance value of the user, the weights of the first evaluation result, the second evaluation result and the third evaluation result are adjusted, and the evaluation feedback suggestion is generated according to the adjustment results.

4. The method for training eloquence based on data retrieval and fusion according to claim 3, characterized in that: The eloquence information is evaluated from the perspective of speech expression ability to obtain a first evaluation result, specifically: Using a preset formula set, the first evaluation result is obtained by measuring the pitch value and the speech speed value of the eloquence information; The formula set is established by considering the weights of different dimensions and the influence of time factors.

5. The method for training eloquence based on data retrieval and fusion according to claim 1, characterized in that: The method for obtaining the preset resources is specifically as follows: Acquire initial resources; wherein the initial resources refer to modal data including voice, text, and video; Calculating the similarity between different data in the initial resource to obtain a resource similarity value; The data in the initial resources are sorted according to the resource similarity values ​​to obtain the preset resources.

6. The method for training eloquence based on data retrieval and fusion according to claim 1, characterized in that: After the user is trained according to the eloquence training method, the method further includes: Based on the similarity between the preset source domain and target domain features, the migration weight of the feature is calculated; Using the migration weights to migrate the features of the source domain to the target domain, and generating a migration learning strategy based on the migration results; Formulate predefined strategies based on the learning characteristics and verbal ability of the user; Optimizing the predefined strategy according to the user's learning behavior and social data to obtain an individual feature learning strategy; The user is again trained in eloquence according to the transfer learning strategy and the individual feature learning strategy.

7. The method for training eloquence based on data retrieval and fusion according to claim 1, characterized in that: After the user is trained according to the eloquence training method, the method further includes: By measuring the overall eloquence level of the user, a comprehensive score is obtained; By evaluating the user's ability performance in different eloquence dimensions, a sub-dimension score is obtained; Generate personalized feedback based on the comprehensive score and the sub-dimension scores, and display the personalized feedback.

8. The method for training eloquence based on data retrieval and fusion according to claim 1, characterized in that: After obtaining the first resource, the method further includes: The first resources are sorted according to the user's knowledge level target and resource preference to obtain sorted first resources.

9. An eloquence training device based on data retrieval and fusion, characterized in that: It includes information module, retrieval module, evaluation module, habit module, eloquence module and comprehensive module; Wherein, the information module is used to obtain the user's query words and eloquence information; The retrieval module is configured to calculate a matching degree between the query term and the preset resources, sort the preset resources, and obtain a first resource; The evaluation module is configured to obtain an evaluation result of the user's eloquence ability from the eloquence information, and sort the first resources according to the evaluation result of the eloquence ability to obtain knowledge feedback suggestions; The habit module is used to filter the preset resources according to the user's learning habits and generate personalized feedback suggestions; specifically, the first resource formula is used to filter the preset resources according to the historical resources that the user has learned, and similar resources with high similarity to the historical resources are obtained; the training resource similarity and score similarity between the user and other trainees are used as standards to obtain the preferred resources of the other trainees; a linear weighting method is used to assign corresponding weights to the similar resources and the preferred resources, and the degree of integration of the similar resources and the preferred resources is adjusted by a preset coefficient to obtain a mixed resource; the personalized feedback suggestions are generated based on the similar resources, the preferred resources, and the mixed resources. The eloquence module is used to evaluate the eloquence information according to the preset eloquence indicators and generate evaluation feedback suggestions based on the evaluation results; The comprehensive module is used to generate an eloquence training method based on the knowledge feedback suggestion, the personalized feedback suggestion and the evaluation feedback suggestion, and train the user according to the eloquence training method; Among them, the first resource formula is: Where u is the user, i is the resource, t is the time; m is the modality type, m = 1 represents voice, m = 2 represents text, and m = 3 represents video; α m is the weight of the mth mode, is the multimodal attention mechanism function, is the feature vector of the advanced eloquence index of the mth mode of user u at time t, is the characteristic vector of the advanced eloquence index of the mth mode of resource i at time t, f(P u (i)) is the preference function of user u for resource i, λ is the time decay factor; β m is the mth modal fusion coefficient, which is used to adjust the weights of different modalities.

10. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: The evaluation module includes a performance unit and a network unit; The performance unit is configured to score the eloquence information based on the modal dimensions of voice, text, and video according to a preset value range to obtain a speech performance value; The network unit is used to capture the dependency value and dynamic change value between the user's knowledge level at different times and the speech performance value through a probabilistic network, and obtain the user's oral ability evaluation result.

11. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: The eloquence module includes an ability unit and a weight unit; The ability unit is used to evaluate the eloquence information from the perspectives of speech expression ability, body language ability, and situation control ability, respectively, to obtain a first evaluation result, a second evaluation result, and a third evaluation result; The weighting unit is used to adjust the weights of the first evaluation result, the second evaluation result and the third evaluation result according to the target speech scenario and the target eloquence performance value of the user, and generate the evaluation feedback suggestion according to the adjustment result.

12. The eloquence training device based on data retrieval and fusion according to claim 11, characterized in that: The capability units are specifically: Using a preset formula set, the first evaluation result is obtained by measuring the pitch value and the speech speed value of the eloquence information; The formula set is established by considering the weights of different dimensions and the influence of time factors.

13. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: The method for obtaining the preset resources is specifically as follows: Acquire initial resources; wherein the initial resources refer to modal data including voice, text, and video; Calculating the similarity between different data in the initial resource to obtain a resource similarity value; The data in the initial resources are sorted according to the resource similarity values ​​to obtain the preset resources.

14. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: It also includes migration units, strategy units, prefabrication units, optimization units, and training units; The migration unit is used to calculate the migration weight of the feature according to the similarity between the preset source domain and target domain features; The strategy unit is configured to migrate the features of the source domain to the target domain using the migration weights, and generate a migration learning strategy based on the migration results; The prefabrication unit is used to formulate a predefined strategy based on the learning characteristics and oral ability of the user; The optimization unit is configured to optimize the predefined strategy according to the user's learning behavior and social data to obtain an individual feature learning strategy; The training unit is used to perform eloquence training on the user again according to the transfer learning strategy and the individual feature learning strategy.

15. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: It also includes overall units, dimensional units and display units; The overall unit is configured to obtain a comprehensive score by measuring the overall eloquence level of the user after the user is trained according to the eloquence training method; The dimension unit is used to evaluate the user's ability performance in different eloquence dimensions to obtain sub-dimension scores; The display unit is used to generate personalized feedback according to the comprehensive score and the sub-dimensional scores, and to display the personalized feedback.

16. The eloquence training device based on data retrieval and fusion according to claim 9, characterized in that: Also included is an adjustment unit; The adjustment unit is configured to, after obtaining the first resource, sort the first resource according to the user's knowledge level target and resource preference to obtain the sorted first resource.

17. A storage medium, characterized in that: The storage medium stores a computer program, which is called and executed by a computer to implement any one of the eloquence training methods based on data retrieval and fusion as claimed in claims 1 to 8.

Citation Information

Patent Citations

  • Speech effect evaluation method and device, evaluation equipment and readable storage medium

    CN112884340A

  • Interactive virtual reality oral talent expression training method, device and equipment and medium

    CN117541444A

  • Multi-modal feedback method, device and equipment for talent training and storage medium

    CN117788239A

Cited By

  • Preparation method and application of recombinant I-type human collagen filler

    CN120733120A

  • A preparation method and application of a recombinant type I human collagen filler

    CN120733120B