Speech evaluation system and method based on multimodal feature analysis
Through multimodal feature analysis, combined with voice signals and environmental data, a comprehensive evaluation model for speech effects was constructed, which solved the problem of inaccurate evaluation of speech effects in the existing technology, and achieved accurate evaluation and effective feedback on speech effects.
Patent Information
- Application Number
- CN202510896692.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The existing technology is difficult to evaluate the performance of speech objectively and comprehensively, and ignores the impact of environmental noise and speech recognition errors on speech clarity, resulting in loss and distortion of speech information transmission, and cannot provide accurate feedback on speech effect.
Through multimodal feature analysis, combined with speaker voice signals, environmental data and text data, a model is built to transmit speech information accuracy and content recognition accuracy, comprehensively evaluate speech effects, and provide comprehensive feedback.
It realizes an accurate evaluation of the speech effect, ensures the effective transmission of speech content information, and provides objective and comprehensive feedback on the speech effect.
Smart Images

Figure CN120408099B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data analysis, and in particular to a speech evaluation system and method based on multimodal feature analysis. Background Art
[0002] In information dissemination and communication scenarios, speeches serve as a core vehicle for knowledge transfer and opinion expression. The evaluation of their effectiveness is crucial for the effective transfer of knowledge and the improvement of speakers' abilities. The effectiveness of a speech depends not only on the accuracy of the speaker's own information transmission but is also deeply intertwined with multiple factors, such as environmental interference and speech recognition. For example, if a speaker's clear and accurate speech output is interfered with by complex environmental noise, or if the audience experiences recognition bias due to voice quality or environmental factors, this can lead to loss and distortion of the speech's information transmission, thus affecting the effective communication of knowledge and ideas. Traditional speech evaluation methods often rely on subjective manual evaluation. These methods rely solely on subjective judgment of speech effectiveness through human listening, which is significantly influenced by individual experience and cognitive differences. The results lack stability and objectivity, making it difficult to accurately capture the true loss of information transmission during a speech, and unable to objectively and comprehensively evaluate the effectiveness of a speech. Consequently, they struggle to meet the current demand for accurate evaluation in scenarios such as knowledge dissemination and speech training.
[0003] At the same time, existing technologies only evaluate the clarity of speech during a speech based on the short-term energy of the speech signal, ignoring the impact of the complex interaction between the human ear's auditory characteristics and environmental noise on the clarity of speech during a speech, or only quantifying the speech effect from the degree of matching of text semantics, without considering the impact of errors introduced by the environment and recognition during speech transmission on the speech effect. Therefore, there is a lack of collaborative analysis of multimodal characteristic factors in the entire speech process, making it difficult for existing technologies to accurately capture the actual loss of content information transmission in a speech, and thus unable to objectively and comprehensively evaluate the speech effect.
[0004] In order to solve these problems, this application designs a speech evaluation system and method based on multimodal feature analysis. Summary of the Invention
[0005] The purpose of the present invention is to provide a speech evaluation system and method based on multimodal feature analysis, which comprehensively analyzes the accuracy of the speaker's speech information transmission and the accuracy of the speech content recognition; and then evaluates the speaker's speech effect; and judges the speaker's speech effect based on the speech effect evaluation results; it can provide the speaker with accurate and comprehensive speech effect feedback, while ensuring the effective transmission of speech content information.
[0006] The present invention is achieved in that:
[0007] In a first aspect, the present invention provides a speech evaluation method based on multimodal feature analysis, comprising the following steps:
[0008] S1. Acquire the speaker's reading content text data, speaker's voice signal data, and speech environment data during the speaker's speech;
[0009] S2. Importing the speaker's reading content text data and the speaker's voice signal data into a speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy;
[0010] S3, importing the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content;
[0011] S4. Construct a comprehensive evaluation model for speech effectiveness, import the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy into the comprehensive evaluation model to evaluate the speaker's speech effectiveness;
[0012] S5. Judge the speaker’s speech effectiveness based on the speech effectiveness evaluation results.
[0013] In the preferred technical solution of this embodiment, step S2 analyzes the accuracy of the speaker's information transmission, specifically including:
[0014] S21, extracting the speaker's reading content text data and the speaker's voice signal data;
[0015] S22. Construct a speech information transmission accuracy analysis model, import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model, analyze the speaker's speech information transmission accuracy, and obtain the speaker's speech information transmission accuracy analysis results.
[0016] In the preferred technical solution of this embodiment, the process of constructing the speech information transmission accuracy analysis model in step S22 specifically includes:
[0017] S221, analyzing the speaker's speech information vocabulary matching accuracy based on the speaker's reading content text data and the speaker's voice signal data, to obtain a speaker's speech information vocabulary matching accuracy analysis result;
[0018] S222, analyzing the semantic consistency of the speaker's speech information based on the speaker's reading content text data and the speaker's voice signal data, to obtain a semantic consistency analysis result of the speaker's speech information;
[0019] S223, analyzing the accuracy of the speaker's speech information transmission based on the analysis results of the speaker's speech information vocabulary matching accuracy and the speech information semantic consistency analysis results;
[0020] The formula for calculating the accuracy of speech information transmission is:
[0021] ;
[0022] Where CD is the accuracy of the speaker's speech information transmission, Cp is the analysis result of the accuracy of speech information vocabulary matching, and Cy is the analysis result of the semantic consistency of speech information.
[0023] In the preferred technical solution of this embodiment, step S3 analyzes the recognition accuracy of the speech content, which specifically includes the following steps:
[0024] S31, extracting simultaneously the speaker's speech signal data and speech environment data during the speaker's speech;
[0025] S32. Construct a speech content recognition accuracy analysis model, import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model, analyze the recognition accuracy of the speech content, and obtain the speech content recognition accuracy analysis result.
[0026] In the preferred technical solution of this embodiment, the process of constructing the speech content recognition accuracy analysis model in step S32 includes the following specific steps:
[0027] S321. Analyze the speech clarity of the speech content based on the speaker's voice signal data and the speech environment data to obtain a speech clarity analysis result of the speech content;
[0028] S322: Analyze the degree of environmental interference to the speech content based on the speaker's voice signal data and the speech environment data, and obtain an analysis result of the degree of environmental interference to the speech content;
[0029] S323, analyzing the recognition accuracy of the speech content based on the analysis results of the speech clarity and the analysis results of the degree of environmental interference of the speech content;
[0030] The calculation formula for recognition accuracy is:
[0031] ;
[0032] Where SQ is the recognition accuracy of the speech content, Yq is the speech clarity analysis result of the speech content, and Hg is the analysis result of the environmental interference degree of the speech content.
[0033] In the preferred technical solution of this embodiment, the construction of a comprehensive evaluation model for speech effectiveness in step S4 includes the following specific steps:
[0034] S41, obtaining the analyzed results of the accuracy of the speaker's speech information transmission and the accuracy of the speech content recognition analysis;
[0035] S42. Evaluate the speaker's speech effectiveness based on the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy;
[0036] The formula for calculating the speaker's speech effect is:
[0037] ;
[0038] Where XG is the speaker's speech effect, a and b are the influence weights of information transmission accuracy and recognition accuracy, respectively.
[0039] In the preferred technical solution of this embodiment, step S5 judges the speaker's speech effect based on the speech effect evaluation result, specifically including:
[0040] S51. Obtaining the evaluation result of the speaker's speech effect;
[0041] S52. Preset a speech effect threshold. When the evaluation result of the speaker's speech effect is greater than the speech effect threshold, the speaker's speech effect is determined to be qualified; when the evaluation result of the speaker's speech effect is less than or equal to the speech effect threshold, the speaker's speech effect is determined to be unqualified.
[0042] In a second aspect, the present invention provides a speech evaluation system based on multimodal feature analysis, comprising:
[0043] A data acquisition module is used to acquire the speaker's reading content text data, speaker voice signal data and speech environment data during the speaker's speech;
[0044] The speech information transmission accuracy analysis module is used to import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy;
[0045] A speech content recognition accuracy analysis module is used to import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content;
[0046] The speech effect comprehensive evaluation module is used to build a speech effect comprehensive evaluation model, import the results of the speaker's speech information transmission accuracy analysis and the results of the speech content recognition accuracy analysis into the speech effect comprehensive evaluation model, and evaluate the speaker's speech effect;
[0047] A speech effect judgment module is used to judge the speaker's speech effect based on the speech effect evaluation results;
[0048] The control module is used to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, and the speech effect judgment module.
[0049] In a third aspect, the present invention provides an electronic device comprising: a processor and a memory, wherein the memory stores a computer program that can be called by the processor, and the processor executes a speech evaluation method based on multimodal feature analysis by calling the computer program stored in the memory.
[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] The present invention imports the speaker's reading content text data and the speaker's voice signal data into a speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy; imports the speaker's voice signal data and speech environment data into a speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; constructs a comprehensive evaluation model for speech effect, imports the speaker's speech information transmission accuracy analysis results and the speech content recognition accuracy analysis results into the comprehensive evaluation model for speech effect, and evaluates the speaker's speech effect; judges the speaker's speech effect based on the speech effect evaluation results; and can provide the speaker with accurate and comprehensive speech effect feedback while ensuring the effective transmission of speech content information. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0053] Figure 1 Schematic diagram of the overall process of the speech evaluation method based on multimodal feature analysis of the present invention;
[0054] Figure 2 Schematic diagram of the structure of the speech evaluation system based on multimodal feature analysis of the present invention;
[0055] Figure 3 This is an analysis flow chart of step S2 of the speech evaluation method based on multimodal feature analysis of the present invention;
[0056] Figure 4 This is an analysis flow chart of step S3 of the speech evaluation method based on multimodal feature analysis of the present invention. DETAILED DESCRIPTION
[0057] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0058] Example 1
[0059] like Figure 1 As shown, this embodiment provides a speech evaluation method based on multimodal feature analysis, which specifically includes the following steps:
[0060] S1. Acquire the speaker's reading content text data, speaker's voice signal data, and speech environment data during the speaker's speech;
[0061] S2. Importing the speaker's reading content text data and the speaker's voice signal data into a speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy;
[0062] S3, importing the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content;
[0063] S4. Construct a comprehensive evaluation model for speech effectiveness, import the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy into the comprehensive evaluation model to evaluate the speaker's speech effectiveness;
[0064] S5. Judge the speaker’s speech effectiveness based on the speech effectiveness evaluation results.
[0065] In this embodiment, if Figure 3 As shown, in step S2, the accuracy of the speaker's information transmission is analyzed, specifically including:
[0066] S21, extracting the speaker's reading content text data and the speaker's voice signal data;
[0067] S22. Construct a speech information transmission accuracy analysis model, import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model, analyze the speaker's speech information transmission accuracy, and obtain the speaker's speech information transmission accuracy analysis results.
[0068] In this embodiment, the process of constructing the speech information transmission accuracy analysis model in step S22 specifically includes:
[0069] S221, analyzing the speaker's speech information vocabulary matching accuracy based on the speaker's reading content text data and the speaker's voice signal data, to obtain a speaker's speech information vocabulary matching accuracy analysis result;
[0070] The calculation formula for the accuracy of speech information vocabulary matching is:
[0071] ;
[0072] Where Cp is the accuracy of speech information vocabulary matching, At is the set of reading content texts in the speaker's reading content text data, and Ay is the set of speech texts after ASR speech transcription in the speaker's speech signal data. is the Damerau-Levenshtein distance between the reading content text set At and the speech text set Ay, is the symmetric difference between the reading content text set At and the voice text set Ay, and len() is the length of the set in the brackets;
[0073] For example, during a speech, a speaker may often make slips of the tongue, such as reversing adjacent words. The traditional edit distance calculation method is to determine the distance between the reading content text set and the speech text set by calculating the minimum number of basic operations such as insertion, deletion, and replacement. However, when analyzing the accuracy of speech information vocabulary matching, this method cannot analyze the situation where the speaker often reverses adjacent words. Therefore, this embodiment takes the situation of reversal of adjacent words into account during the speech, and calculates the Damerau-Levenshtein distance between the reading content text set and the speech text set to quantify the impact of the exchange of adjacent characters in the reading content text set and the speech text set on the accuracy of speech information vocabulary matching, thereby more accurately analyzing the vocabulary deviation in the speech process. Further, the symmetric difference set of the reading content text set and the speech text set is the set of all elements in At and Ay that do not belong to the intersection of At and Ay. This embodiment calculates the length of the symmetric difference set of the reading content text set and the speech text set through len(), which can quantify the impact of the increase or decrease of vocabulary during the speech on the accuracy of speech information vocabulary matching. Further, this embodiment calculates the length of the symmetric difference set of the reading content text set and the speech text set through len(), which can quantify the impact of the increase or decrease of vocabulary during the speech on the accuracy of speech information vocabulary matching. To balance the influence of the exchange of adjacent characters in the reading content text set and the speech text set and the increase or decrease of vocabulary in the speech process on the accuracy of speech information vocabulary matching; further, this embodiment normalizes the numerator calculation result by taking the total length of the reading content text set and the speech text set as the denominator, and , the differences in adjacent character exchanges between the reading content text set and the voice text set, as well as the differences in vocabulary addition or reduction that occur during the speaker's speech, are converted into the accuracy of speech information vocabulary matching; this embodiment avoids single difference analysis, such as analyzing only the vocabulary deviation caused by the number of vocabulary replacements during the speech, by integrating the differences in adjacent character exchanges and the differences in vocabulary addition or reduction, and can more comprehensively reflect the speaker's transmission accuracy at the speech vocabulary level when conveying speech information to the audience.
[0074] S222, analyzing the semantic consistency of the speaker's speech information based on the speaker's reading content text data and the speaker's voice signal data, to obtain a semantic consistency analysis result of the speaker's speech information;
[0075] The calculation formula for the semantic consistency of speech information is:
[0076] ;
[0077] Where Cy is the semantic consistency of speech information, n is the number of common words in the reading content text set and the speech text set, Vti is the BERT word vector of the i-th common word wti in the reading content text set, and Vyi is the BERT word vector of the i-th common word wyi in the speech text set. is the shortest path distance between the i-th common word wti in the reading content text set and the i-th common word wyi in the speech text set in the WordNet semantic tree, is the cosine similarity between the BERT word vector of the i-th common word wti in the reading content text set and the BERT word vector of the i-th common word wyi in the speech text set;
[0078] For example, in this embodiment, when common words appear in the speaker's reading content text set and voice text set, the common words themselves may have different meanings. For example, when the common word is "apple", "apple" can refer to both fruit and a company. At the same time, when the common words have the same meaning in both the reading content text set and the voice text set, the semantics of the common words may differ due to the different contexts in the reading content text set and the voice text set. In this embodiment, the WordNet semantic tree can map the different meanings of the same word form to different nodes in the tree by constructing a semantic network structure of the words. When common words with the same word form appear in the reading content text set and the voice text set, if the two have different meanings, the shortest path distance in the semantic tree will be larger; conversely, if the meanings are similar, the distance will be smaller. By quantifying the shortest path distance between common words in the reading content text set and the speech text set in the WordNet semantic tree, different meanings under the same word form can be effectively identified, that is, the larger the shortest path distance, the more significant the difference in the meaning of the common words; therefore, this embodiment quantifies the difference in meaning of common words with the same word form through the shortest path distance between the i-th common word wti in the reading content text set and the i-th common word wyi in the speech text set in the WordNet semantic tree, which can avoid the occurrence of polysemy; further, in the process of human perception of semantic differences, the lexical differences within the same semantic branch are less perceived by humans; while the lexical differences across semantic branches are perceived significantly more differently by humans. Therefore, compared to the linear attenuation function, the exponential function used in this embodiment can better highlight the nonlinear attenuation characteristic that the contribution to the semantic consistency of speech information decreases sharply as the shortest path distance increases; specifically, the exponential function It has a nonlinear attenuation characteristic of first rapid and then slow. When smaller, will drop rapidly, which enables the contribution of close words in the semantic tree to the semantic consistency of speech information to be quickly corrected, avoiding misjudgment due to subtle semantic differences; When it is larger, The downward trend slows down and gradually approaches 0, which can avoid the formula used in this embodiment from excessively punishing words with large semantic differences, and at the same time provide space for calculating the cosine similarity of the BERT vector to capture the potential semantic association of the context; further, this embodiment uses cosine similarity Quantify the semantic proximity of word vector space, The closer it is to 1, the closer the BERT word vector directions of the common words in the reading content text set and the voice text set are, and the more similar the semantics represented by the common words in the reading content text set and the voice text set in the semantic space are; through cosine similarity The specific technical process of quantifying the semantic proximity of word vector space belongs to the existing technology and will not be described here. and Complementary advantages. When the common words themselves have small differences in meaning and the semantic differences in the context are small, the multiplication result becomes larger, and the degree of semantic consistency of the speech information is higher; if one of the two is more different, the degree of semantic consistency of the speech information becomes lower. This embodiment adopts multiplication operation to avoid the one-sidedness of single semantic coding. It solves the problem of semantic differences in dynamic contexts and It solves the problem of static word meaning ambiguity and normalizes the calculation results by dividing them by the number of common words. It can accurately quantify the quality of semantic transmission of speech information and meet the needs of actual speech evaluation.
[0079] S223, analyzing the accuracy of the speaker's speech information transmission based on the analysis results of the speaker's speech information vocabulary matching accuracy and the speech information semantic consistency analysis results;
[0080] The formula for calculating the accuracy of speech information transmission is:
[0081] ;
[0082] Where CD is the accuracy of the speaker's speech information transmission, Cp is the analysis result of the accuracy of speech information vocabulary matching, and Cy is the analysis result of the semantic consistency of speech information.
[0083] For example, this embodiment is calculated based on the classical calculation formula structure of the traditional harmonic mean, that is, through the product term To quantify the collaborative contribution of the accuracy of speech information vocabulary matching and the semantic consistency of speech information; the denominator is Perform addition operation by The sum of the direct contributions of the accuracy of speech information vocabulary matching and the semantic consistency of speech information is reflected by the absolute value It measures the deviation between the accuracy of the speech information vocabulary matching and the semantic consistency of the speech information. When the two are added together, When the difference is large, the denominator increases significantly, thereby lowering the calculation result to quantify the imbalance between the accuracy of vocabulary matching and the degree of semantic consistency during the speaker's speech. The classic calculation formula structure of the traditional harmonic mean is improved to ensure that the accuracy of the speaker's speech information transmission will approach the ideal value only when the accuracy of vocabulary matching and the degree of semantic consistency are balanced during the speaker's speech.
[0084] In this embodiment, if Figure 4 As shown, step S3 analyzes the recognition accuracy of the speech content, which specifically includes the following steps:
[0085] S31, extracting simultaneously the speaker's speech signal data and speech environment data during the speaker's speech;
[0086] S32. Construct a speech content recognition accuracy analysis model, import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model, analyze the recognition accuracy of the speech content, and obtain the speech content recognition accuracy analysis result.
[0087] In this embodiment, the process of constructing the speech content recognition accuracy analysis model in step S32 includes the following specific steps:
[0088] S321. Analyze the speech clarity of the speech content based on the speaker's voice signal data and the speech environment data to obtain a speech clarity analysis result of the speech content;
[0089] The formula for calculating the speech clarity of speech content is:
[0090] ;
[0091] Where Yq is the speech clarity of the speech content, Sy is the speaker's speech time-domain audio sequence extracted from the speaker's speech signal data, and Se is the ambient noise time-domain audio sequence extracted from the speech environment data. represents the time-frequency matrix obtained by performing short-time Fourier transform on the j-th frame of the speaker's speech in the speaker's speech time-domain audio sequence Sy, It represents the time-frequency matrix output after the j-th frame of the environmental noise in the environmental noise time domain audio sequence Se is processed by the Gammatone filter bank. is the Hadamard product, m is the total number of frames in the time domain audio sequence, and j is any item from 1 to m.
[0092] For example, this embodiment simulates human hearing through a Gammatone filter bank and performs speech signal analysis through short-time Fourier transform. The speaker's speech time-domain audio sequence Sy and the ambient noise time-domain audio sequence Se during the speech are divided into m short-time frames according to a fixed frame length and frame shift. The j-th frame of the speaker's speech is subjected to short-time Fourier transform to obtain a time-frequency matrix , the j-th frame environmental noise is processed by the Gammatone filter bank, and the time-frequency matrix simulating the frequency response of the human auditory system is output; then the Hadamard product is used to and By element-by-element multiplication, the frequency domain features of a single frame of the speaker's speech after being disturbed by ambient noise are obtained, thereby reflecting the interference of ambient noise on the audience's reception and understanding of the speaker's speech. Finally, the sum of the frequency domain features of m frames is normalized by dividing the sum by the total number of frames m to obtain the speech clarity of the speech content. The higher the speech clarity of the speech content, the more speech content information is retained after the frequency domain superposition of the speech and noise, and the clearer the speaker's speech. In this embodiment, short-time Fourier transform and gammatone filter bank processing are both existing signal processing and analysis technologies and will not be described in detail here.
[0093] S322: Analyze the degree of environmental interference to the speech content based on the speaker's voice signal data and the speech environment data, and obtain an analysis result of the degree of environmental interference to the speech content;
[0094] The formula for calculating the degree of environmental interference to the speech content is:
[0095] ;
[0096] Where Hg is the degree of environmental interference to the speech content, represents the time-frequency matrix obtained by performing short-time Fourier transform on the j-th frame of the ambient noise in the ambient noise time-domain audio sequence Se. is the Frobenius norm, which is used to calculate the energy of the time-frequency matrix. The calculation process of the Frobenius norm is an existing technology and will not be repeated here.
[0097] For example, during a speaker's speech, the frequency and energy of the speech change dynamically with the content, and the ambient noise may also exhibit time-varying characteristics due to interference sources. This embodiment uses short-time Fourier transform to convert long-term non-stationary signals into short-term approximately stationary time-frequency segments, which can accurately quantify the frequency energy distribution of the speech and ambient noise in each frame; further, this embodiment uses the Frobenius norm to calculate the energy of the time-frequency matrix, which can objectively reflect the total energy intensity of the time-frequency matrix in the two-dimensional space of time and frequency. The complex time-frequency matrix energy is converted into a scalar value through the Frobenius norm, so that the energy of speech and noise are directly comparable. In this embodiment, Used to calculate the proportion of speech energy to total energy in a single frame. According to the acoustic masking effect, the human ear's perception of sound depends on energy competition; when the energy of the ambient noise is dominant, the speaker's speech is easily masked, which increases the degree of environmental interference to the speech content; when the speaker's speech energy is dominant, the degree of environmental interference to the speech content is relatively weak. This embodiment converts the energy competition at the physical level into a quantifiable proportional relationship by calculating the degree of environmental interference to the speech content, which can directly reflect the impact of environmental interference on the speech content; and by subtracting 1 , converting the speech energy ratio into the degree of environmental interference; further, during the speech process, the environmental noise will fluctuate over time. Analyzing the degree of environmental interference for only a single frame of speech cannot reflect the environmental interference during the entire speech. Therefore, this embodiment solves the dynamic time-varying problem of environmental interference during the speech by dividing by m to obtain the average value, and can objectively reflect the degree of environmental interference during the entire speech.
[0098] S323, analyzing the recognition accuracy of the speech content based on the analysis results of the speech clarity and the analysis results of the degree of environmental interference of the speech content;
[0099] The calculation formula for recognition accuracy is:
[0100] ;
[0101] Where SQ is the recognition accuracy of the speech content, Yq is the speech clarity analysis result of the speech content, and Hg is the analysis result of the environmental interference degree of the speech content.
[0102] Illustratively, this embodiment uses multiplication operation and 1-Hg to quantify the suppression of speech clarity caused by environmental interference, thereby reflecting the influence of speech clarity under environmental interference on the recognition accuracy of speech content.
[0103] In this embodiment, the construction of the comprehensive evaluation model for speech effectiveness in step S4 includes the following specific steps:
[0104] S41, obtaining the analyzed results of the accuracy of the speaker's speech information transmission and the accuracy of the speech content recognition analysis;
[0105] S42. Evaluate the speaker's speech effectiveness based on the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy;
[0106] The formula for calculating the speaker's speech effect is:
[0107] ;
[0108] Where XG is the speaker's speech effect, a and b are the influence weights of information transmission accuracy and recognition accuracy, respectively.
[0109] In this embodiment, step S5 judges the speaker's speech effectiveness based on the speech effectiveness evaluation result, specifically including:
[0110] S51. Obtaining the evaluation result of the speaker's speech effect;
[0111] S52: A speech effect threshold is preset. When the evaluation result of the speaker's speech effect is greater than the speech effect threshold, the speaker's speech effect is determined to be qualified; when the evaluation result of the speaker's speech effect is less than or equal to the speech effect threshold, the speaker's speech effect is determined to be unqualified. The setting parameters (e.g., weights and thresholds) in this embodiment are obtained by experiments conducted by those skilled in the art. The specific experimental method is as follows: obtaining text data of speaker reading content, speaker voice signal data, and speech environment data from multiple historical speakers' speeches; substituting the speaker reading content text data, speaker voice signal data, and speech environment data into each step of this embodiment to obtain speech effect evaluation results of the multiple historical speakers; obtaining judgment results on whether the speech effects of the multiple historical speakers are qualified; importing the speech effect evaluation results of the multiple historical speakers obtained by substituting the obtained results into each step of this embodiment and the corresponding judgment results on whether the speech effects of the multiple historical speakers are qualified into fitting software, and outputting the setting parameters (e.g., weights and thresholds) that meet the highest speech effect judgment accuracy.
[0112] Example 2
[0113] like Figure 2 As shown, this embodiment provides a speech evaluation system based on multimodal feature analysis, including:
[0114] A data acquisition module is used to acquire the speaker's reading content text data, speaker voice signal data and speech environment data during the speaker's speech;
[0115] The speech information transmission accuracy analysis module is used to import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy;
[0116] A speech content recognition accuracy analysis module is used to import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content;
[0117] The speech effect comprehensive evaluation module is used to build a speech effect comprehensive evaluation model, import the results of the speaker's speech information transmission accuracy analysis and the results of the speech content recognition accuracy analysis into the speech effect comprehensive evaluation model, and evaluate the speaker's speech effect;
[0118] A speech effect judgment module is used to judge the speaker's speech effect based on the speech effect evaluation results;
[0119] The control module is used to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, and the speech effect judgment module.
[0120] For the above-mentioned parameters and steps for each unit module to realize corresponding functions in the speech evaluation system based on multimodal feature analysis of the present invention, reference can be made to the parameters and steps in the embodiment of the speech evaluation method based on multimodal feature analysis above, and no further details will be given here.
[0121] Example 3
[0122] An electronic device according to an embodiment of the present invention includes: a processor and a memory, wherein the memory stores a computer program that can be called by the processor, and the processor executes a speech evaluation method based on multimodal feature analysis by calling the computer program stored in the memory. It should be noted that all computer programs of the speech evaluation method based on multimodal feature analysis are implemented using the C language, wherein the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, the speech effect judgment module, and the control module are all controlled by a remote server.
[0123] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0124] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A speech evaluation method based on multimodal feature analysis is characterized by: The steps include: S1. Acquire the speaker's reading content text data, speaker's voice signal data, and speech environment data during the speaker's speech; S2. Importing the speaker's reading content text data and the speaker's voice signal data into a speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy; S3, importing the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; S4. Construct a comprehensive evaluation model for speech effectiveness, import the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy into the comprehensive evaluation model to evaluate the speaker's speech effectiveness; S5. Judging the speaker’s speech effectiveness based on the speech effectiveness evaluation results; The construction process of the speech information transmission accuracy analysis model specifically includes: S221, analyzing the speaker's speech information vocabulary matching accuracy based on the speaker's reading content text data and the speaker's voice signal data, to obtain a speaker's speech information vocabulary matching accuracy analysis result; S222, analyzing the semantic consistency of the speaker's speech information based on the speaker's reading content text data and the speaker's voice signal data, to obtain a semantic consistency analysis result of the speaker's speech information; S223, analyzing the accuracy of the speaker's speech information transmission based on the analysis results of the speaker's speech information vocabulary matching accuracy and the speech information semantic consistency analysis results; The formula for calculating the accuracy of speech information transmission is: ; Wherein, CD is the accuracy of the speaker's speech information transmission, Cp is the analysis result of the accuracy of the speech information vocabulary matching, and Cy is the analysis result of the semantic consistency of the speech information; The construction process of the speech content recognition accuracy analysis model includes the following specific steps: S321. Analyze the speech clarity of the speech content based on the speaker's voice signal data and the speech environment data to obtain a speech clarity analysis result of the speech content; S322: Analyze the degree of environmental interference to the speech content based on the speaker's voice signal data and the speech environment data, and obtain an analysis result of the degree of environmental interference to the speech content; S323, analyzing the recognition accuracy of the speech content based on the analysis results of the speech clarity and the analysis results of the degree of environmental interference of the speech content; The calculation formula for recognition accuracy is: ; Where SQ is the recognition accuracy of the speech content, Yq is the analysis result of the speech clarity, and Hg is the analysis result of the degree of environmental interference to the speech content; Constructing a comprehensive evaluation model for speech effectiveness includes the following specific steps: S41, obtaining the analyzed results of the accuracy of the speaker's speech information transmission and the accuracy of the speech content recognition analysis; S42. Evaluate the speaker's speech effectiveness based on the analysis results of the speaker's speech information transmission accuracy and the analysis results of the speech content recognition accuracy; The formula for calculating the speaker's speech effect is: ; Where XG is the speaker's speech effect, a and b are the influence weights of information transmission accuracy and recognition accuracy, respectively.
2. The speech evaluation method based on multimodal feature analysis according to claim 1 is characterized in that: The analysis of the speaker's information delivery accuracy in step S2 specifically includes: S21, extracting the speaker's reading content text data and the speaker's voice signal data; S22. Construct a speech information transmission accuracy analysis model, import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model, analyze the speaker's speech information transmission accuracy, and obtain the speaker's speech information transmission accuracy analysis results.
3. The speech evaluation method based on multimodal feature analysis according to claim 2 is characterized in that: The analysis of the recognition accuracy of the speech content in step S3 specifically includes the following steps: S31, extracting simultaneously the speaker's speech signal data and speech environment data during the speaker's speech; S32. Construct a speech content recognition accuracy analysis model, import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model, analyze the recognition accuracy of the speech content, and obtain the speech content recognition accuracy analysis result.
4. The speech evaluation method based on multimodal feature analysis according to claim 3 is characterized in that: In step S5, judging the speaker's speech effect according to the speech effect evaluation result specifically includes: S51. Obtaining the evaluation result of the speaker's speech effect; S52. Preset a speech effect threshold. When the evaluation result of the speaker's speech effect is greater than the speech effect threshold, the speaker's speech effect is determined to be qualified; when the evaluation result of the speaker's speech effect is less than or equal to the speech effect threshold, the speaker's speech effect is determined to be unqualified.
5. A speech evaluation system based on multimodal feature analysis, used to implement the speech evaluation method based on multimodal feature analysis according to any one of claims 1 to 4, characterized in that: The system comprises: A data acquisition module is used to acquire the speaker's reading content text data, speaker voice signal data and speech environment data during the speaker's speech; The speech information transmission accuracy analysis module is used to import the speaker's reading content text data and the speaker's voice signal data into the speech information transmission accuracy analysis model to analyze the speaker's speech information transmission accuracy; A speech content recognition accuracy analysis module is used to import the speaker's voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; The speech effect comprehensive evaluation module is used to build a speech effect comprehensive evaluation model, import the results of the speaker's speech information transmission accuracy analysis and the results of the speech content recognition accuracy analysis into the speech effect comprehensive evaluation model, and evaluate the speaker's speech effect; A speech effect judgment module is used to judge the speaker's speech effect based on the speech effect evaluation results; The control module is used to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, and the speech effect judgment module.
6. An electronic device comprising: A processor and a memory, wherein the memory stores a computer program that can be called by the processor; it is characterized in that the processor executes the speech evaluation method based on multimodal feature analysis as described in any one of claims 1 to 4 by calling the computer program stored in the memory.
Citation Information
Patent Citations
Text matching method and device, computer equipment and storage medium
CN113486659A
Real-time spoken language assessment system and method on mobile devices
US20160253923A1