Speech evaluation system and method based on multi-modal feature analysis
Through a speech evaluation system with multimodal feature analysis, combined with voice signals and environmental data, a comprehensive evaluation model is built, which solves the shortcomings of speech effect evaluation in the existing technology and achieves accurate evaluation and feedback on speech effect.
Patent Information
- Application Number
- CN202510896692.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
The existing technology is difficult to evaluate the speech effect objectively and comprehensively, and ignores the impact of environmental noise and audience identification errors on speech clarity, resulting in loss and distortion of speech information transmission, and cannot provide accurate speech effect feedback.
Through multimodal feature analysis, combining speaker voice signals, reading content text data and environmental data, an analysis model is constructed to comprehensively evaluate the speech effect and provide accurate speech effect feedback.
It achieves an accurate and comprehensive evaluation of the speech effect, ensures the effective transmission of speech content information, provides objective feedback on speech effect, and improves the accuracy of speaker ability training.
Smart Images

Figure CN120408099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data analysis, and particularly to a speech evaluation system and method based on multimodal feature analysis. Background Art
[0002] In the scenarios of information dissemination and communication, as a core carrier for knowledge transfer and opinion expression, the evaluation of speech effects is crucial for the effective transfer of knowledge and the improvement of speakers' abilities. The speech effect not only depends on the accuracy of information transfer by the speaker himself, but is also deeply related to multiple factors such as environmental interference and speech recognition. For example, if the clear and accurate speech output of the speaker is affected by complex environmental noise interference, or if the audience has recognition deviations due to speech quality and environmental factors, the transmission of speech information will be lost and distorted, affecting the effective communication of knowledge and opinions. Traditional speech evaluation methods often rely on subjective manual evaluation, only subjectively judging the speech effect by manual listening. Greatly affected by personal experience and cognitive differences, the results lack stability and objectivity, are difficult to accurately capture the real loss of information transfer in the speech, and cannot objectively and comprehensively evaluate the speech effect, making it difficult to meet the requirements of accurate evaluation in current scenarios such as knowledge dissemination and speech ability training.
[0003] At the same time, the existing technology only evaluates the clarity of speech during a speech based on the short-time energy of the speech signal, ignoring the complex interaction between the human ear's auditory characteristics and environmental noise on the clarity of speech during a speech, or only quantifies the speech effect from the degree of text semantic matching, without considering the influence of errors introduced by the environment and recognition during the speech transmission on the speech effect. Therefore, the lack of collaborative analysis of multimodal feature factors in the entire speech process makes it difficult for the existing technology to accurately capture the real loss of content information transfer in the speech, and thus unable to objectively and comprehensively evaluate the speech effect.
[0004] To solve these problems, the present application designs a speech evaluation system and method based on multimodal feature analysis. Summary of the Invention
[0005] The purpose of the present invention is to provide a speech evaluation system and method based on multimodal feature analysis. By comprehensively analyzing the accuracy of speech information transfer by the speaker and the accuracy of speech content recognition; then evaluating the speech effect of the speaker; and judging the speech effect of the speaker according to the speech effect evaluation result; it can provide accurate and comprehensive speech effect feedback for the speaker, and at the same time ensure the effective transfer of speech content information.
[0006] The present invention is implemented as follows: In the first aspect, the present invention provides a speech evaluation method based on multimodal feature analysis, including the following steps: S1. Obtain the speaker's reading content text data, speaker voice signal data, and speech environment data during the speech; S2. Import the speaker's reading content text data and speaker voice signal data into the speech information transmission accuracy analysis model to analyze the accuracy of the speaker's speech information transmission; S3. Import the speaker voice signal data and speech environment data into the speech content recognition accuracy analysis model to analyze the accuracy of speech content recognition; S4. Construct a comprehensive speech effect evaluation model, import the analysis results of the speaker's speech information transmission accuracy and the speech content recognition accuracy analysis results into the comprehensive speech effect evaluation model to evaluate the speaker's speech effect; S5. Judge the speaker's speech effect according to the speech effect evaluation result.
[0007] In the preferred technical solution of this embodiment, the analysis of the accuracy of the speaker's information transmission in step S2 specifically includes: S21. Extract the speaker's reading content text data and speaker voice signal data; S22. Construct a speech information transmission accuracy analysis model, import the speaker's reading content text data and speaker voice signal data into the speech information transmission accuracy analysis model, analyze the accuracy of the speaker's speech information transmission, and obtain the analysis result of the speaker's speech information transmission accuracy.
[0008] In the preferred technical solution of this embodiment, the construction process of the speech information transmission accuracy analysis model in step S22 specifically includes: S221. Based on the speaker's reading content text data and speaker voice signal data, analyze the accuracy of the speech information vocabulary matching of the speaker to obtain the analysis result of the speech information vocabulary matching accuracy of the speaker; S222. Based on the speaker's reading content text data and speaker voice signal data, analyze the semantic consistency of the speaker's speech information to obtain the analysis result of the semantic consistency of the speaker's speech information; S223. According to the analysis results of the speech information vocabulary matching accuracy and the semantic consistency of the speech information of the speaker, analyze the accuracy of the speaker's speech information transmission; The calculation formula for the accuracy of speech information transmission is: ; In the formula, CD is the accuracy of the speaker's speech information transmission, Cp is the analysis result of the speech information vocabulary matching accuracy, and Cy is the analysis result of the semantic consistency of the speech information.
[0009] In the preferred technical solution of this embodiment, in step S3, the accuracy of the recognition of the speech content is analyzed, which specifically includes the following steps: S31. Simultaneously extract the speaker's speech signal data and the speech environment data during the speaker's speech; S32. Construct an analysis model for the accuracy of speech content recognition, import the speaker's speech signal data and the speech environment data into the analysis model for the accuracy of speech content recognition, analyze the accuracy of the recognition of the speech content, and obtain the analysis result of the accuracy of the recognition of the speech content.
[0010] In the preferred technical solution of this embodiment, the construction process of the analysis model for the accuracy of speech content recognition in step S32 includes the following specific steps: S321. Based on the speaker's speech signal data and the speech environment data, analyze the speech clarity of the speech content to obtain the analysis result of the speech clarity of the speech content; S322. Based on the speaker's speech signal data and the speech environment data, analyze the degree of environmental interference suffered by the speech content to obtain the analysis result of the degree of environmental interference suffered by the speech content; S323. According to the analysis result of the speech clarity of the speech content and the analysis result of the degree of environmental interference suffered by the speech content, analyze the accuracy of the recognition of the speech content; The calculation formula for the accuracy of recognition is: ; In the formula, SQ is the accuracy of the recognition of the speech content, Yq is the analysis result of the speech clarity of the speech content, and Hg is the analysis result of the degree of environmental interference suffered by the speech content.
[0011] In the preferred technical solution of this embodiment, in step S4, constructing a comprehensive evaluation model for the speech effect includes the following specific steps: S41. Obtain the analysis result of the accuracy of the speech information transmission of the speaker and the analysis result of the accuracy of the recognition of the speech content obtained through analysis; S42. Evaluate the speech effect of the speaker according to the analysis result of the accuracy of the speech information transmission of the speaker and the analysis result of the accuracy of the recognition of the speech content; The calculation formula for the speech effect of the speaker is: ; In the formula, XG is the speech effect of the speaker, and a and b are the influence weights of information transmission accuracy and recognition accuracy respectively.
[0012] In the preferred technical solution of this embodiment, in step S5, according to the speech effect evaluation result, the speech effect of the speaker is judged, which specifically includes: S51. Obtain the speech effect evaluation result of the speaker obtained by evaluation; S52. Preset a speech effect threshold. When the speech effect evaluation result of the speaker obtained by evaluation is greater than the speech effect threshold, it is determined that the speech effect of the speaker is qualified; when the speech effect evaluation result of the speaker obtained by evaluation is less than or equal to the speech effect threshold, it is determined that the speech effect of the speaker is unqualified.
[0013] In a second aspect, the present invention provides a speech evaluation system based on multimodal feature analysis, including: A data acquisition module, configured to acquire the text data of the content read by the speaker, the voice signal data of the speaker, and the speech environment data during the speech of the speaker; A speech information transmission accuracy analysis module, configured to import the text data of the content read by the speaker and the voice signal data of the speaker into a speech information transmission accuracy analysis model to analyze the speech information transmission accuracy of the speaker; A speech content recognition accuracy analysis module, configured to import the voice signal data of the speaker and the speech environment data into a speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; A speech effect comprehensive evaluation module, configured to construct a speech effect comprehensive evaluation model, and import the speech information transmission accuracy analysis result of the speaker and the speech content recognition accuracy analysis result into the speech effect comprehensive evaluation model to evaluate the speech effect of the speaker; A speech effect judgment module, configured to judge the speech effect of the speaker according to the speech effect evaluation result; A control module, configured to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, and the speech effect judgment module.
[0014] In a third aspect, the present invention provides an electronic device, including: a processor and a memory, wherein a computer program callable by the processor is stored in the memory, and the processor executes a speech evaluation method based on multimodal feature analysis by calling the computer program stored in the memory.
[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects: The present invention imports the text data of the content read by the speaker and the speaker's voice signal data into the analysis model for the accuracy of speech information transmission to analyze the accuracy of the speaker's speech information transmission; imports the speaker's voice signal data and the speech environment data into the analysis model for the accuracy of speech content recognition to analyze the accuracy of speech content recognition; constructs a comprehensive speech effect evaluation model, and imports the analysis results of the accuracy of the speaker's speech information transmission and the analysis results of the accuracy of speech content recognition into the comprehensive speech effect evaluation model to evaluate the speaker's speech effect; judges the speaker's speech effect according to the speech effect evaluation results; can provide accurate and comprehensive speech effect feedback for the speaker, and at the same time ensure the effective transmission of speech content information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments read with reference to the accompanying drawings: Figure 1 It is a schematic diagram of the overall process of the speech evaluation method based on multi-modal feature analysis of the present invention; Figure 2 It is a schematic diagram of the structure of the speech evaluation system based on multi-modal feature analysis of the present invention; Figure 3 It is an analysis flowchart of step S2 of the speech evaluation method based on multi-modal feature analysis of the present invention; Figure 4 It is an analysis flowchart of step S3 of the speech evaluation method based on multi-modal feature analysis of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0018] Embodiment 1
[0019] As Figure 1 shown, this embodiment provides a speech evaluation method based on multi-modal feature analysis, which specifically includes the following steps: S1. Obtain the text data of the content read by the speaker, the speaker's voice signal data, and the speech environment data during the speaker's speech; S2. Import the text data of the content read by the speaker and the speaker's voice signal data into the analysis model for the accuracy of speech information transmission to analyze the accuracy of the speaker's speech information transmission; S3. Import the speaker's speech signal data and the speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; S4. Construct a comprehensive speech effect evaluation model, import the analysis results of the accuracy of the speaker's speech information transmission and the analysis results of the recognition accuracy of the speech content into the comprehensive speech effect evaluation model to evaluate the speaker's speech effect; S5. Judge the speaker's speech effect according to the speech effect evaluation results.
[0020] In this embodiment, as Figure 3 shown, the analysis of the accuracy of the speaker's information transmission in step S2 specifically includes: S21. Extract the speaker's reading content text data and the speaker's speech signal data; S22. Construct a speech information transmission accuracy analysis model, import the speaker's reading content text data and the speaker's speech signal data into the speech information transmission accuracy analysis model, analyze the accuracy of the speaker's speech information transmission, and obtain the analysis results of the accuracy of the speaker's speech information transmission.
[0021] In this embodiment, the construction process of the speech information transmission accuracy analysis model in step S22 specifically includes: S221. Analyze the accuracy of the speech information vocabulary matching of the speaker based on the speaker's reading content text data and the speaker's speech signal data to obtain the analysis results of the accuracy of the speech information vocabulary matching of the speaker; The calculation formula for the accuracy of the speech information vocabulary matching is: ; In the formula, Cp is the accuracy of the speech information vocabulary matching, At is the reading content text set in the speaker's reading content text data, Ay is the speech text set after ASR speech transcription in the speaker's speech signal data, is the Damerau-Levenshtein distance between the reading content text set At and the speech text set Ay, is the symmetric difference set of the reading content text set At and the speech text set Ay, and len() is the length of the set in the parentheses; Exemplarily, during a speech, the speaker often makes slips of the tongue such as reversing adjacent words; the traditional edit distance calculation method determines the distance between the reading content text set and the speech text set by calculating the minimum number of basic operations such as insertion, deletion, and replacement. However, when analyzing the accuracy of word matching in speech information, this method cannot analyze the situation where the speaker often reverses adjacent words; therefore, this embodiment takes into account the situation where adjacent words are reversed during the speech, and quantifies the impact of the adjacent character swapping situation existing in the reading content text set and the speech text set on the accuracy of word matching in speech information by calculating the Damerau-Levenshtein distance between the reading content text set and the speech text set, so as to more accurately analyze the word deviation existing during the speech; further, the symmetric difference set of the reading content text set and the speech text set is the set of all elements that do not belong to the intersection of At and Ay in At and Ay. This embodiment calculates the length of the symmetric difference set of the reading content text set and the speech text set through len(), which can quantify the impact of the situation of word addition or reduction by the speaker during the speech on the accuracy of word matching in speech information; further, this embodiment balances the impact of the adjacent character swapping situation existing in the reading content text set and the speech text set and the situation of word addition or reduction by the speaker during the speech on the accuracy of word matching in speech information; further, this embodiment normalizes the calculation result of the numerator by using the total length of the reading content text set and the speech text set as the denominator, and through , converts the adjacent character swapping difference and the word addition or reduction difference existing in the reading content text set and the speech text set into the accuracy of word matching in speech information; this embodiment integrates the adjacent character swapping difference and the word addition or reduction difference, avoiding single difference analysis, such as the word deviation caused by only analyzing the number of word replacements during the speech, and can more comprehensively reflect the transmission accuracy at the word level of the speaker when transmitting speech information to the audience.
[0022] S222. Analyze the semantic consistency degree of the speaker's speech information based on the speaker's reading content text data and the speaker's speech signal data, and obtain the analysis result of the semantic consistency degree of the speaker's speech information; The calculation formula for the semantic consistency degree of speech information is: ; Where Cy is the semantic consistency degree of the speech information, n is the number of common words in the reading content text set and the speech text set, Vti is the BERT word vector of the i-th common word wti in the reading content text set, and Vyi is the BERT word vector of the i-th common word wyi in the speech text set. Is the shortest path distance between the i-th common word wti in the reading content text set and the i-th common word wyi in the speech text set in the WordNet semantic tree. Is the cosine similarity between the BERT word vector of the i-th common word wti in the reading content text set and the BERT word vector of the i-th common word wyi in the speech text set. Exemplarily, in this embodiment, when there are common words in the reading content text set and the speech text set of the speaker, there will be differences in the meanings of these common words themselves. For example, when the common word is "apple", it can refer to a fruit or a company. At the same time, when the meanings of the common words in the reading content text set and the speech text set are the same, there will still be differences in the semantics of the common words due to the different contexts in the reading content text set and the speech text set. In this embodiment, the WordNet semantic tree can map different meanings of the same word form to different nodes in the tree by constructing a semantic network structure of the vocabulary. When there are common words with the same word form in the reading content text set and the speech text set, if their meanings are different, the shortest path distance in the semantic tree will be larger; conversely, if the meanings are similar, the distance will be smaller. By quantifying the shortest path distance between the common words in the reading content text set and the speech text set in the WordNet semantic tree, it is possible to effectively identify different meanings under the same word form, that is, the larger the shortest path distance, the more significant the difference in the meanings of the common words. Therefore, in this embodiment, the shortest path distance between the i-th common word wti in the reading content text set and the i-th common word wyi in the speech text set in the WordNet semantic tree is used to quantify the difference in the meanings of the common words with the same word form, which can avoid the situation of polysemy. Further, in the process of human perception of semantic differences, the difference in vocabulary within the same semantic branch is perceived by humans with less change; while the difference in vocabulary across semantic branches will cause a significant jump in human perception. Therefore, compared with the linear attenuation function, the exponential function adopted in this embodiment can more prominently show the non-linear attenuation characteristic that the contribution of the speech information semantic consistency degree decreases sharply as the shortest path distance increases. Specifically, the exponential function Has a non-linear attenuation characteristic of being steep first and then gentle. When Is relatively small, Will decrease rapidly, which enables the contribution of the vocabulary with a short distance in the semantic tree to the speech information semantic consistency degree to be quickly corrected, avoiding misjudgment due to subtle semantic differences. And when When it is relatively large, the downward trend slows down and gradually approaches zero, which can avoid the formula adopted in this embodiment from over-punishing words with large semantic differences, and at the same time give room for calculating the cosine similarity of BERT vectors to capture potential semantic associations in the context; further, this embodiment adopts the cosine similarity to quantify the semantic proximity of the word vector space, the closer it is to 1, indicating that the directions of the BERT word vectors of the common words in the reading content text set and the speech text set are more similar, and the semantics represented by the common words in the reading content text set and the speech text set in the semantic space are also more similar; through the cosine similarity the specific technical processes for quantifying the semantic proximity of the word vector space all belong to the prior art and will not be elaborated here. Further, the multiplication operation has the characteristics of mutual enhancement or inhibition, which can make and complement each other's advantages. When the difference in the meaning of the common word itself is small and the semantic difference in the context is small, the multiplication result becomes larger, and the semantic consistency degree of the speech information is higher; if one of them has a large difference, the semantic consistency degree of the speech information becomes lower. This embodiment adopts the multiplication operation to avoid the one-sidedness of single semantic encoding, and not only deals with the problem of semantic differences in dynamic contexts, but also solves the problem of static word meaning ambiguity, and at the same time normalizes the calculation result by dividing by the number of common words, which can accurately quantify the semantic transmission quality of speech information and meet the needs of actual speech evaluation.
[0023] S223. Analyze the accuracy of the speech information transmission of the speaker according to the analysis results of the accuracy of the speech information word matching and the analysis results of the semantic consistency degree of the speech information obtained; The calculation formula for the accuracy of the speech information transmission is: ; In the formula, CD is the accuracy of the speech information transmission of the speaker, Cp is the analysis result of the accuracy of the speech information word matching, and Cy is the analysis result of the semantic consistency degree of the speech information.
[0024] Exemplarily, this embodiment calculates based on the classical calculation formula structure of the traditional harmonic mean, and not only uses the product term to quantify the collaborative contribution of the accuracy of the speech information word matching and the semantic consistency degree of the speech information; the denominator uses for addition operation, and through Measure the degree of deviation between the accuracy of vocabulary matching and the semantic consistency of the speech information. After adding the two, when the difference is large, the denominator increases significantly, thereby dragging down the calculation result to quantify the unbalanced state of the accuracy of vocabulary matching and the degree of semantic consistency during the information transmission process of the speaker's speech. This embodiment adopts to improve the classical calculation formula structure of the traditional harmonic mean to ensure that only when the accuracy of vocabulary matching and the degree of semantic consistency are balanced during the information transmission process of the speaker's speech, the accuracy of the speaker's speech information transmission tends to the ideal value.
[0025] In this embodiment, as Figure 4 shown, in step S3, analyze the accuracy of speech content recognition, which specifically includes the following steps: S31. Extract the speaker's speech signal data and speech environment data during the speaker's speech at the same time; S32. Construct an analysis model for the accuracy of speech content recognition, import the speaker's speech signal data and speech environment data into the analysis model for the accuracy of speech content recognition, analyze the accuracy of speech content recognition, and obtain the analysis result of the accuracy of speech content recognition.
[0026] In this embodiment, the construction process of the analysis model for the accuracy of speech content recognition in step S32 includes the following specific steps: S321. Based on the speaker's speech signal data and speech environment data, analyze the speech clarity of the speech content to obtain the analysis result of the speech clarity of the speech content; The calculation formula for the speech clarity of the speech content is: ; In the formula, Yq is the speech clarity of the speech content, Sy is the speaker's speech time-domain audio sequence extracted from the speaker's speech signal data, Se is the environmental noise time-domain audio sequence extracted from the speech environment data, represents the time-frequency matrix obtained by performing short-time Fourier transform on the j-th frame of the speaker's speech in the speaker's speech time-domain audio sequence Sy, represents the time-frequency matrix output after processing the j-th frame of environmental noise in the environmental noise time-domain audio sequence Se through a Gammatone filter bank, is the Hadamard product, m is the total number of frames of the time-domain audio sequence, and j is any item from 1 to m.
[0027] Exemplarily, in this embodiment, the human ear hearing is simulated through a Gammatone filter bank, and the speech signal is analyzed through a short-time Fourier transform. The speech time-domain audio sequence Sy and the environmental noise time-domain audio sequence Se during the speech process are segmented into m short-time frames according to a fixed frame length and frame shift. The short-time Fourier transform is performed on the speech of the speaker in the j-th frame to obtain a time-frequency matrix , the j-th frame of environmental noise is processed by a Gammatone filter bank, and a time-frequency matrix outputting the frequency response of the simulated human ear auditory system is obtained; then through the Hadamard product for and element-by-element multiplication, the frequency-domain characteristics of the single-frame speaker speech interfered by environmental noise are obtained, so as to be able to reflect the interference of environmental noise on the audience's reception and understanding of the speaker's speech when the audience is watching; finally, the sum of the superposition results of the frequency-domain characteristics of m frames is summed and divided by the total number of frames m for normalization processing, and the speech clarity of the speech content is obtained. When the speech clarity of the speech content is higher, it indicates that more speech content information is retained after the superposition of the speech and noise frequency domains, and the speaker's speech is clearer. In this embodiment, the short-time Fourier transform and the Gammatone filter bank processing both belong to existing signal processing and analysis technologies, which will not be elaborated here
[0028] S322. Based on the speaker speech signal data and the speech environment data, analyze the degree of environmental interference suffered by the speech content, and obtain the analysis result of the degree of environmental interference suffered by the speech content The calculation formula for the degree of environmental interference suffered by the speech content is ; In the formula, Hg is the degree of environmental interference suffered by the speech content represents the time-frequency matrix obtained by performing a short-time Fourier transform on the j-th frame of environmental noise in the environmental noise time-domain audio sequence Se is the Frobenius norm, which is used to calculate the energy of the time-frequency matrix. The calculation process of the Frobenius norm is an existing technology and will not be elaborated here
[0029] Exemplarily, during the speaker's speech, the frequency and energy of the speech vary dynamically with the content, and the environmental noise may also show time-varying characteristics due to interference sources. In this embodiment, through the short-time Fourier transform, the long-time non-stationary signal is transformed into short-time approximately stationary time-frequency segments, which can accurately quantify the frequency and energy distribution of the speech and environmental noise within each frame; further, in this embodiment, the Frobenius norm is used to calculate the energy of the time-frequency matrix, which can objectively reflect the total energy intensity of the time-frequency matrix in the two-dimensional space of time and frequency. By using the Frobenius norm to transform the energy of the complex time-frequency matrix into a scalar value, the energy of the speech and noise is directly comparable. In this embodiment To calculate the proportion of the speech energy of a single frame in the total energy. According to the acoustic masking effect, the human ear's perception of sound depends on energy competition; when the environmental noise energy dominates, the speech of the speaker is easily masked, enhancing the degree of environmental interference on the speech content; when the speech energy of the speaker dominates, the degree of environmental interference on the speech content is weaker. In this embodiment, by calculating the degree of environmental interference on the speech content, the energy competition at the physical level is converted into a quantifiable proportional relationship, which can directly reflect the impact of environmental interference on the speech content; and by subtracting 1 from , the proportion of speech energy is converted into the degree of environmental interference; further, during the speech process, the environmental noise will fluctuate over time. Analyzing the degree of environmental interference only for a single-frame speech cannot reflect the environmental interference throughout the speech. Therefore, in this embodiment, by taking the average value after dividing by m, the dynamic time-varying problem of environmental interference during the speech process is solved, and the degree of environmental interference throughout the speech can be objectively reflected.
[0030] S323. Analyze the recognition accuracy of the speech content based on the analysis results of the speech clarity degree of the speech content and the analysis results of the degree of environmental interference on the speech content; The calculation formula for the recognition accuracy is: ; In the formula, SQ is the recognition accuracy of the speech content, Yq is the analysis result of the speech clarity degree of the speech content, and Hg is the analysis result of the degree of environmental interference on the speech content.
[0031] Exemplarily, this embodiment uses multiplication operation and quantifies the suppression of environmental interference on the speech clarity degree by 1 - Hg, so as to reflect the impact of the speech clarity degree under environmental interference on the recognition accuracy of the speech content.
[0032] In this embodiment, in step S4, constructing a comprehensive evaluation model for the speech effect includes the following specific steps: S41. Obtain the analysis results of the accurate transmission degree of the speech information of the speaker and the analysis results of the recognition accuracy of the speech content obtained through analysis; S42. Evaluate the speech effect of the speaker based on the analysis results of the accurate transmission degree of the speech information of the speaker and the analysis results of the recognition accuracy of the speech content; The calculation formula for the speech effect of the speaker is: ; In the formula, XG is the speech effect of the speaker, and a and b are the influence weights of information transmission accuracy and recognition accuracy respectively.
[0033] In this embodiment, in step S5, according to the speech effect evaluation result, the speech effect of the speaker is judged, specifically including: S51. Obtain the speech effect evaluation result of the speaker obtained by evaluation; S52. Preset a speech effect threshold. When the speech effect evaluation result of the speaker obtained by evaluation is greater than the speech effect threshold, it is determined that the speech effect of the speaker is qualified; when the speech effect evaluation result of the speaker obtained by evaluation is less than or equal to the speech effect threshold, it is determined that the speech effect of the speaker is unqualified. Among them, the acquisition method of the set parameters (such as weights and thresholds) in this embodiment is obtained through experiments by those skilled in the art. The specific experimental method is as follows: Obtain the text data of the speaker's reading content, the speaker's voice signal data, and the speech environment data during the speech of multiple historical speakers; substitute the text data of the speaker's reading content, the speaker's voice signal data, and the speech environment data into each step of this embodiment to obtain the speech effect evaluation results of multiple historical speakers; obtain the judgment results of whether the speech effects of multiple historical speakers are qualified, and import the speech effect evaluation results of multiple historical speakers obtained by substituting into each step of this embodiment and the corresponding judgment results of whether the speech effects of multiple historical speakers are qualified into the fitting software, and output the values of the set parameters (such as weights and thresholds) that meet the highest speech effect judgment accuracy.
[0034] Embodiment 2
[0035] As Figure 2 shown, this embodiment provides a speech evaluation system based on multi-modal feature analysis, including: A data acquisition module, which is used to obtain the text data of the speaker's reading content, the speaker's voice signal data, and the speech environment data during the speech of the speaker; A speech information transmission accuracy analysis module, which is used to import the text data of the speaker's reading content and the speaker's voice signal data into the speech information transmission accuracy analysis model to analyze the speech information transmission accuracy of the speaker; A speech content recognition accuracy analysis module, which is used to import the speaker's voice signal data and the speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; A speech effect comprehensive evaluation module, which is used to construct a speech effect comprehensive evaluation model, and import the speech information transmission accuracy analysis result of the speaker and the speech content recognition accuracy analysis result into the speech effect comprehensive evaluation model to evaluate the speech effect of the speaker; A speech effect judgment module, which is used to judge the speech effect of the speaker according to the speech effect evaluation result; A control module, configured to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, and the speech effect judgment module.
[0036] For the parameters and the steps for each unit module in the above speech evaluation system based on multi-modal feature analysis of the present invention to implement corresponding functions, reference can be made to the parameters and steps in the embodiments of the speech evaluation method based on multi-modal feature analysis in the foregoing text, which will not be elaborated herein.
[0037] Embodiment 3
[0038] An electronic device according to an embodiment of the present invention includes: a processor and a memory. Among them, a computer program that can be called by the processor is stored in the memory, and the processor executes a speech evaluation method based on multi-modal feature analysis by calling the computer program stored in the memory. It should be noted that: all computer programs of the speech evaluation method based on multi-modal feature analysis are implemented using the C language. Among them, the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the speech effect comprehensive evaluation module, the speech effect judgment module, and the control module are all controlled by a remote server.
[0039] In the description of this specification, the descriptions referring to terms such as "an embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0040] The above-disclosed preferred embodiments of the present invention are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the relevant technical fields can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A speech evaluation method based on multimodal feature analysis, characterized in that, Including the following steps: S1. Obtain the text data of the speaker's reading content, the speaker's voice signal data, and the speech environment data during the speech; S2. Import the text data of the speaker's reading content and the speaker's voice signal data into the speech information transmission accuracy analysis model to analyze the accuracy of the speaker's speech information transmission; S3. Import the speaker's voice signal data and the speech environment data into the speech content recognition accuracy analysis model to analyze the accuracy of the speech content recognition; S4. Build a comprehensive speech effect evaluation model, import the analysis results of the speaker's speech information transmission accuracy and the speech content recognition accuracy into the comprehensive speech effect evaluation model to evaluate the speaker's speech effect; S5. Judge the speaker's speech effect according to the speech effect evaluation results.
2. The speech evaluation method based on multi-modal feature analysis according to claim 1, wherein The analysis of the speaker's information transmission accuracy in step S2 specifically includes: S21. Extract the text data of the speaker's reading content and the speaker's voice signal data; S22. Build a speech information transmission accuracy analysis model, import the text data of the speaker's reading content and the speaker's voice signal data into the speech information transmission accuracy analysis model, analyze the accuracy of the speaker's speech information transmission, and obtain the analysis result of the speaker's speech information transmission accuracy.
3. The speech evaluation method based on multi-modal feature analysis according to claim 2, wherein The construction process of the speech information transmission accuracy analysis model in step S22 specifically includes: S221. Analyze the accuracy of the speech information vocabulary matching of the speaker based on the text data of the speaker's reading content and the speaker's voice signal data to obtain the analysis result of the speech information vocabulary matching accuracy of the speaker; S222. Analyze the semantic consistency of the speaker's speech information based on the text data of the speaker's reading content and the speaker's voice signal data to obtain the analysis result of the semantic consistency of the speaker's speech information; S223. Analyze the accuracy of the speaker's speech information transmission according to the analysis results of the speech information vocabulary matching accuracy and the semantic consistency of the speaker's speech information; The calculation formula for the accuracy of speech information transmission is: ; In the formula, CD is the accuracy of the speaker's speech information transmission, Cp is the analysis result of the speech information vocabulary matching accuracy, and Cy is the analysis result of the semantic consistency of the speaker's speech information.
4. The speech evaluation method based on multi-modal feature analysis according to claim 3, wherein The analysis of the accuracy of the speech content recognition in step S3 specifically includes the following steps: S31. Extract the speaker's voice signal data and the speech environment data during the speaker's speech at the same time; S32. Build a speech content recognition accuracy analysis model, import the speaker's voice signal data and the speech environment data into the speech content recognition accuracy analysis model, analyze the accuracy of the speech content recognition, and obtain the analysis result of the speech content recognition accuracy.
5. The speech evaluation method based on multi-modal feature analysis according to claim 4, wherein The construction process of the speech content recognition accuracy analysis model in step S32 includes the following specific steps: S321. Analyze the speech clarity of the speech content based on the speaker's speech signal data and the speech environment data to obtain the analysis result of the speech clarity of the speech content; S322. Analyze the degree of environmental interference suffered by the speech content based on the speaker's speech signal data and the speech environment data to obtain the analysis result of the degree of environmental interference suffered by the speech content; S323. Analyze the recognition accuracy of the speech content according to the analysis results of the speech clarity of the speech content and the analysis results of the degree of environmental interference suffered by the speech content; The calculation formula for the recognition accuracy is: ; In the formula, SQ is the recognition accuracy of the speech content, Yq is the analysis result of the speech clarity of the speech content, and Hg is the analysis result of the degree of environmental interference suffered by the speech content.
6. The speech evaluation method based on multi-modal feature analysis according to claim 5, characterized in that In step S4, constructing a comprehensive speech effect evaluation model includes the following specific steps: S41. Obtain the analysis results of the accuracy of the speaker's speech information transmission and the recognition accuracy of the speech content obtained by analysis; S42. Evaluate the speaker's speech effect according to the analysis results of the accuracy of the speaker's speech information transmission and the recognition accuracy of the speech content; The calculation formula for the speaker's speech effect is: ; In the formula, XG is the speaker's speech effect, and a and b are the influence weights of information transmission accuracy and recognition accuracy respectively.
7. The speech evaluation method based on multi-modal feature analysis according to claim 6, wherein, In step S5, judging the speaker's speech effect according to the speech effect evaluation result specifically includes: S51. Obtain the speech effect evaluation result of the speaker obtained by evaluation; S52. Preset a speech effect threshold. When the speech effect evaluation result of the speaker obtained by evaluation is greater than the speech effect threshold, it is determined that the speaker's speech effect is qualified; when the speech effect evaluation result of the speaker obtained by evaluation is less than or equal to the speech effect threshold, it is determined that the speaker's speech effect is unqualified.
8. A speech evaluation system based on multimodal feature analysis, which is used to implement the speech evaluation method based on multimodal feature analysis described in any one of claims 1-7, and is characterized in that, The system includes: A data acquisition module for acquiring the speaker's reading content text data, the speaker's speech signal data, and the speech environment data during the speech; A speech information transmission accuracy analysis module for importing the speaker's reading content text data and the speaker's speech signal data into the speech information transmission accuracy analysis model to analyze the accuracy of the speaker's speech information transmission; A speech content recognition accuracy analysis module for importing the speaker's speech signal data and the speech environment data into the speech content recognition accuracy analysis model to analyze the recognition accuracy of the speech content; A comprehensive speech effect evaluation module for constructing a comprehensive speech effect evaluation model, importing the analysis results of the accuracy of the speaker's speech information transmission and the recognition accuracy of the speech content into the comprehensive speech effect evaluation model to evaluate the speaker's speech effect; A speech effect judgment module for judging the speaker's speech effect according to the speech effect evaluation result; A control module, configured to control the operation of the data acquisition module, the speech information transmission accuracy analysis module, the speech content recognition accuracy analysis module, the comprehensive speech effect evaluation module, and the speech effect judgment module.
9. An electronic device, comprising: A processor and a memory, wherein the memory stores a computer program that can be called by the processor; characterized in that the processor executes the speech evaluation method based on multi-modal feature analysis according to any one of claims 1-7 by calling the computer program stored in the memory.
Citation Information
Patent Citations
Text matching method and device, computer equipment and storage medium
CN113486659A
Real-time spoken language assessment system and method on mobile devices
US20160253923A1
System, plug-in, and method for improving text composition by modifying character prominence according to assigned character information measures
US8306356B1