A method and system for speech rate analysis

By extracting and analyzing the total number of syllables in the pronunciation data, the accuracy of speech speed analysis in teaching resources is solved, and an effective evaluation of teachers' teaching activities is achieved.

CN115171724BActive Publication Date: 2025-07-01DMAI (GUANGZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110359348.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-01
Publication Date
2025-07-01
Estimated Expiration
2041-04-01

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve accurate analysis of speech speed in teaching resources, which affects the evaluation of teachers' teaching activities.

Method used

By obtaining the speech data to be analyzed and its total duration, the total number of syllables is extracted, and the speech speed is determined based on the total number of syllables and the total duration. The method includes dividing audio clips, extracting sound features, inputting a preset syllable regression model to obtain the syllable number, and finally calculating the speech speed.

Benefits of technology

It realizes accurate analysis of speech speed in teaching resources, and can analyze speech speed data for teachers' teaching dialogues in online teaching platforms, providing data support for teaching evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171724B_ABST
    Figure CN115171724B_ABST
Patent Text Reader

Abstract

The present invention provides a speech rate analysis method and system. The method includes: obtaining the speech data to be analyzed and its corresponding total duration; extracting the total number of syllables included in the speech data to be analyzed; and determining the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration. Thus, by extracting the total number of syllables included in the speech data to be analyzed, the speech rate is analyzed, and accurate analysis of the speech rate in teaching resources is achieved. Thereby, the teaching conversations of teachers in Chinese or English teaching scenarios on the online teaching platform can be analyzed to obtain their speech rate data, providing data support for teaching analysis, which is of great significance for the evaluation of the overall teaching activities of teachers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and particularly to a speech rate analysis method and system. Background Art

[0002] With the rapid development of the mobile Internet, the application of communication software has become more and more extensive. For example, more and more teachers use instant messaging software to provide online teaching guidance to students, replacing the traditional face-to-face teaching method. Compared with the traditional offline education mode, online education has the advantage of space, with flexible teaching locations, which to a certain extent promotes the spread of high-quality educational resources.

[0003] Online education usually completes teaching through recorded audio and video. Since the speech rate of teachers during teaching affects the listening effect of students, the speech rate is usually one of the important evaluation indicators for evaluating teachers' teaching activities. Therefore, how to accurately analyze the speech rate in teaching resources is of great significance for evaluating teachers' overall teaching activities. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a speech rate analysis method and system to overcome the problem that it is difficult to accurately analyze the speech rate in teaching resources in the prior art.

[0005] Embodiments of the present invention provide a speech rate analysis method, including:

[0006] Obtaining the speech data to be analyzed and its corresponding total duration;

[0007] Extracting the total number of syllables included in the speech data to be analyzed;

[0008] Determining the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration.

[0009] Optionally, the extracting the total number of syllables included in the speech data to be analyzed includes:

[0010] Dividing the speech data to be analyzed into multiple audio segments based on the total duration of the speech data to be analyzed and a preset audio duration;

[0011] Extracting the voice features of each audio segment;

[0012] Inputting the voice features corresponding to each audio segment into a preset syllable number regression model to obtain the number of syllables corresponding to each audio segment;

[0013] Summing up the number of syllables corresponding to all audio segments to obtain the total number of syllables.

[0014] Optionally, the extracting the voice features of each audio segment includes:

[0015] Convert the current audio segment into a magnitude spectrum;

[0016] Extract the depth features included in the current audio segment based on the magnitude spectrum;

[0017] Aggregate all the depth features to obtain the voice feature corresponding to the current audio segment.

[0018] Optionally, the method further includes:

[0019] Determine whether there is an audio segment whose duration is less than the preset audio duration;

[0020] When there is an audio segment whose duration is less than the preset audio duration, pad the audio segment to meet the preset audio duration.

[0021] Optionally, calculate the speech rate through the following formula:

[0022]

[0023] where v represents the speech rate, n represents the total number of audio segments into which the speech data to be analyzed is divided, l represents the total duration of the speech data to be analyzed, and ρ(x i ) represents the number of syllables output by the i-th audio segment input into the preset syllable number regression model ρ.

[0024] Optionally, the preset syllable number regression model is trained through the following method:

[0025] Construct a training data set, where the training data set includes: the voice features of audio samples and the actual number of syllables included in the text corresponding to the audio samples;

[0026] Input the voice features of each audio sample in the training data set into the initial syllable number regression model to obtain the predicted number of syllables corresponding to each audio sample;

[0027] Adjust the model parameters of the initial syllable number regression model based on the relationship between the predicted number of syllables and the actual number of syllables of each audio sample until the preset training requirements of the model are met, and obtain the preset syllable number regression model.

[0028] Optionally, the preset syllable number regression model is a neural network model.

[0029] An embodiment of the present invention further provides a speech rate analysis system, including:

[0030] An acquisition module, configured to acquire the speech data to be analyzed and its corresponding total duration;

[0031] A first processing module, configured to extract the total number of syllables included in the speech data to be analyzed;

[0032] A second processing module, configured to determine the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration.

[0033] An embodiment of the present invention further provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the speech rate analysis method provided by the embodiment of the present invention.

[0034] An embodiment of the present invention further provides a computer-readable storage medium storing computer instructions for causing a computer to execute the speech rate analysis method provided by the embodiment of the present invention.

[0035] The technical solution of the present invention has the following advantages:

[0036] An embodiment of the present invention provides a speech rate analysis method and system. By obtaining speech data to be analyzed and its corresponding total duration; extracting the total number of syllables included in the speech data to be analyzed; and determining the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration. Thus, the speech rate is analyzed by extracting the total number of syllables included in the speech data to be analyzed, realizing accurate analysis of the speech rate in teaching resources, so that the teaching conversations of teachers in Chinese or English teaching scenarios on an online teaching platform can be analyzed to obtain their speech rate data, providing data support for teaching analysis, which is of great significance for the evaluation of the overall teaching activities of teachers. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a flowchart of the speech rate analysis method in the embodiment of the present invention;

[0039] Figure 2 It is a schematic structural diagram of the speech rate analysis system in the embodiment of the present invention;

[0040] Figure 3 It is a schematic structural diagram of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] The technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0043] With the rapid development of the mobile Internet, the application of communication software is becoming more and more extensive. For example, more and more teachers use instant messaging software to conduct online teaching and tutoring for students to replace the traditional face-to-face teaching method. Compared with the traditional offline education mode, online education has advantages in terms of space, with flexible teaching locations, which to a certain extent also promotes the spread of high-quality educational resources.

[0044] Online education usually completes teaching through recorded audio and video. Since the speaking speed of teachers during teaching will affect the listening effect of students, the speaking speed is usually one of the important evaluation indicators for evaluating teachers' teaching activities. Therefore, how to accurately analyze the speaking speed in teaching resources is of great significance for evaluating the overall teaching activities of teachers.

[0045] Based on the above problems, the embodiments of the present invention provide a speaking speed analysis method, which can be applied to the speaking speed analysis of teaching resources in an online teaching platform, such as Figure 1 shown, the speaking speed analysis method mainly includes the following steps:

[0046] Step S101: Obtain the voice data to be analyzed and its corresponding total duration.

[0047] Specifically, the voice data to be analyzed is audio data containing human voices. For example, teaching audio recorded on an online teaching platform or audio data extracted from a teaching video containing voice data, etc. The acquisition method of the voice data to be analyzed can be directly downloading the audio data or extracting it from a preset voice database to be analyzed, etc. The present invention is not limited thereto.

[0048] Step S102: Extract the total number of syllables contained in the voice data to be analyzed.

[0049] Specifically, the way to divide syllables in speech data is related to the language of the speech data. For example, in Chinese, one character generally corresponds to one syllable. In special cases, there are cases where erhua sounds and individual Chinese characters represent two syllables. Targeted text normalization can be performed according to the data set situation before counting the syllables. If it is English, the number of syllables can be counted by segmenting the words in the sentence and determining how many syllables the word has. The number of syllables corresponding to the sentence can also be obtained by calculating the vowels and louder consonants in the phonetic symbols. Other languages ​​need to extract the number of syllables based on their grammatical characteristics, but the present invention is not limited to this.

[0050] Step S103: Determine the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration.

[0051] Among them, the more total syllables contained in the audio data of a fixed length, the faster the speaking speed in the audio data, and vice versa. Therefore, the speaking speed can be evaluated by the total length of the voice data and the total number of syllables contained therein.

[0052] Through the above steps S101 to S103, the speech rate analysis method provided by the embodiment of the present invention analyzes the speech rate by extracting the total number of syllables contained in the voice data to be analyzed, thereby realizing accurate analysis of the speech rate in the teaching resources, thereby analyzing the teacher's teaching dialogue in the Chinese or English teaching scene on the online teaching platform, and obtaining its speech rate data, providing data support for teaching analysis, which is of great significance to the evaluation of the teacher's overall teaching activities.

[0053] Specifically, in one embodiment, the above step S102 specifically includes the following steps:

[0054] Step S201: based on the total duration of the voice data to be analyzed and a preset audio duration, the voice data to be analyzed is divided into a plurality of audio segments.

[0055] Among them, the preset audio duration is a duration set according to the sound feature extraction method and actual needs. By dividing the voice data to be analyzed with a longer total duration into several audio segments, and then processing each audio segment in parallel, the processing speed of the overall speech rate analysis is improved, and it is conducive to realizing real-time speech rate analysis.

[0056] Specifically, in order to facilitate the processing of audio segments, the segmented audio segments need to be normalized. For the last segmented audio segment of the speech data to be analyzed, it is determined whether there is an audio segment whose duration is less than the preset audio duration; if there is an audio segment whose duration is less than the preset audio duration, the audio segment is padded to meet the preset audio duration. This ensures that the duration of all audio segments is the same, which is convenient for subsequent data processing.

[0057] Step S202: Extract the acoustic features of each audio segment.

[0058] Specifically, the acoustic features are extracted through the following process:

[0059] Convert the current audio segment into a magnitude spectrum. Specifically, the main steps include performing short-time Fourier transform on the audio segment, calculating the magnitude spectrum, and performing data preprocessing operations such as normalization to convert the audio signal into a two-dimensional normalized magnitude spectrum.

[0060] Based on the magnitude spectrum, extract the depth features contained in the current audio segment.

[0061] Aggregate all the depth features to obtain the acoustic features corresponding to the current audio segment.

[0062] In the embodiment of the present invention, the audio segment is input into a trained deep neural network model to obtain the acoustic features. Among them, mobilenet-v2 is used as the backbone network to obtain several depth features of the audio, and then these depth features are aggregated to obtain the dense features of the audio, and finally sent to a preset syllable number regression model for syllable number prediction. The backbone network mobilenet-v2 uses depthwise separable convolution instead of traditional convolution, with faster inference speed, which has been widely used in the industry and will not be introduced in depth here; in the feature aggregation stage, a more effective feature aggregation method NetVLADPooling is adopted. Assume that the depth features obtained by the backbone network are {x1, x2, …, x T}. The intermediate output of NetVLAD Pooling is a K×D matrix V, where K represents the predefined number of clusters, and D represents the dimension size of each cluster center. Then each row of the matrix V is obtained through the following formula:

[0063]

[0064] where {w k}, {b k}, {c k} are training parameters and are trained together with the classification model. After performing L2 regularization on the matrix V and splicing them together, they are the features aggregated by NetVLAD Pooling. Finally, they are sent to a preset syllable number regression model for syllable number regression to obtain the number of syllables in the input audio segment. The entire model uses the mean square loss function as the loss function to adjust and train the model.

[0065] Step S203: Input the acoustic features corresponding to each audio segment into a preset syllable number regression model to obtain the syllable number corresponding to each audio segment.

[0066] Among them, in the embodiment of the present invention, the preset syllable number regression model is taken as an example of a neural network model to improve data processing efficiency and facilitate real-time analysis of speech rate. It should be noted that in actual applications, different models can also be selected according to different speech rate analysis requirements.

[0067] The preset syllable number regression model is trained in the following way:

[0068] Construct a training data set, which includes: the sound features of the audio sample and the actual number of syllables contained in the text corresponding to the audio sample. Among them, the original data required to construct the training data set needs to include the audio and its corresponding text, and then the number of syllables corresponding to the audio annotation text is obtained through the syllable calculation method of the corresponding language. This step can be performed by manual annotation or scripting according to the general grammar. For example, in Chinese, one character generally corresponds to one syllable. In special cases, there are erhua sounds and individual Chinese characters representing two syllables. Targeted text normalization can be performed according to the data set situation and then the syllables can be counted. If it is English, the number of syllables can be counted by segmenting the words in the sentence and determining how many syllables the word has. The number of syllables corresponding to the sentence can also be calculated by calculating the vowels and louder consonants in the phonetic symbols. Other languages ​​need to calculate the number of syllables according to their grammatical characteristics to complete the data set construction.

[0069] Input the sound features of each audio sample in the training data set into the initial syllable number regression model to obtain the predicted number of syllables corresponding to each audio sample. First, input the normalized amplitude spectrum obtained in the data pre-processing into the above-mentioned deep neural network, and then adjust the various parameters of the initialized deep neural network model according to the data label. Use all training data to perform the above training operation in a loop until the model converges and the training is completed, and then input the sound features corresponding to each audio sample output after the training is completed into the initial syllable number regression model.

[0070] Based on the relationship between the predicted number of syllables and the actual number of syllables of each audio sample, the model parameters of the initial syllable number regression model are adjusted until the preset training requirements of the model are met, thereby obtaining a preset syllable number regression model.

[0071] Step S204: summing up the numbers of syllables corresponding to all audio clips to obtain the total number of syllables.

[0072] The speaking rate is calculated using the following formula:

[0073]

[0074] Where v represents the speaking speed, n represents the total number of audio segments into which the speech data to be analyzed is divided, l represents the total duration of the speech data to be analyzed, and ρ(x i) represents the number of syllables output by the i-th audio segment input into the preset syllable number regression model ρ.

[0075] Next, the speech rate analysis method provided by the embodiments of the present invention will be described in detail in combination with specific application examples.

[0076] During the first run, load the trained deep neural network model, the preset syllable number regression model, and their respective model parameters;

[0077] Then, perform preprocessing on the incoming audio, including operations such as audio segmentation, padding, conversion to amplitude spectrum, etc., and then perform inference using the deep neural network model and the preset syllable number regression model to obtain the regression results of the number of syllables in each audio sub-segment. In practical applications, rounding operations can be performed according to the required accuracy, such as only retaining the integer number of syllables, etc.

[0078] Then, after obtaining the regression results of the number of pronunciation syllables of all sub-segments after the current audio segmentation, perform a summation calculation on them to obtain the total number of pronunciation syllables.

[0079] Finally, calculate the average speech rate of the current audio based on the total number of pronunciation syllables and the actual duration of the audio and return the result.

[0080] By performing the above steps, the speech rate analysis method provided by the embodiments of the present invention analyzes the speech rate by extracting the total number of syllables contained in the speech data to be analyzed, realizes the accurate analysis of the speech rate in teaching resources, and thus can analyze the teaching dialogues of teachers in Chinese or English teaching scenarios on the online teaching platform to obtain their speech rate data, providing data support for teaching analysis, which is of great significance for the evaluation of the overall teaching activities of teachers.

[0081] The embodiments of the present invention also provide a speech rate analysis system, as Figure 2 shown, this speech rate analysis system includes:

[0082] An acquisition module 101, configured to acquire the speech data to be analyzed and its corresponding total duration. For detailed content, refer to the relevant description of step S101 in the above method embodiment, and details will not be elaborated here.

[0083] A first processing module 102, configured to extract the total number of syllables contained in the speech data to be analyzed. For detailed content, refer to the relevant description of step S102 in the above method embodiment, and details will not be elaborated here.

[0084] A second processing module 103, configured to determine the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration. For detailed content, refer to the relevant description of step S103 in the above method embodiment, and details will not be elaborated here.

[0085] Through the collaborative cooperation of the above-mentioned various components, the speech rate analysis system provided by the embodiments of the present invention analyzes the speech rate by extracting the total number of syllables contained in the speech data to be analyzed, realizes the accurate analysis of the speech rate in teaching resources, and thus can analyze the teaching conversations of teachers in Chinese or English teaching scenarios on the online teaching platform to obtain their speech rate data, providing data support for teaching analysis and being of great significance for the evaluation of the overall teaching activities of teachers.

[0086] According to the embodiments of the present invention, there is also provided an electronic device, as Figure 3 shown. The electronic device may include a processor 901 and a memory 902. The processor 901 and the memory 902 may be connected through a bus or other means. Figure 3 Taking the connection through the bus as an example.

[0087] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.

[0088] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the method embodiments of the present invention. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 902, that is, to implement the methods in the above method embodiments.

[0089] The memory 902 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely provided with respect to the processor 901, and these remote memories may be connected to the processor 901 through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0090] One or more modules are stored in the memory 902 and, when executed by the processor 901, perform the methods in the above method embodiments.

[0091] For the specific details of the above electronic device, reference may be made to the corresponding relevant descriptions and effects in the above method embodiments for understanding, and details are not described herein again.

[0092] Those skilled in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0093] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A speech rate analysis method, characterized in that, including: obtaining the speech data to be analyzed and its corresponding total duration; extracting the total number of syllables included in the speech data to be analyzed; determining the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration; wherein, the extracting the total number of syllables included in the speech data to be analyzed includes: dividing the speech data to be analyzed into multiple audio segments based on the total duration of the speech data to be analyzed and a preset audio duration; extracting the acoustic features of each audio segment; inputting the acoustic features corresponding to each audio segment into a preset syllable number regression model to obtain the number of syllables corresponding to each audio segment; and summing up the number of syllables corresponding to all audio segments to obtain the total number of syllables; extracting the acoustic features of each audio segment includes: inputting the audio segment into a trained deep neural network model to obtain acoustic features, wherein the backbone network of the deep neural network model is used to output several deep features of the audio segment, aggregating the deep features to obtain the dense features of the audio segment, and using the dense features of the audio segment as the acoustic features.

2. The method according to claim 1, wherein The extracting the acoustic features of each audio segment includes: converting the current audio segment into a magnitude spectrum; extracting the deep features included in the current audio segment based on the magnitude spectrum; performing feature aggregation on all the deep features to obtain the acoustic features corresponding to the current audio segment.

3. The method according to claim 1, wherein further including: judging whether there is an audio segment with a duration less than the preset audio duration; when there is an audio segment with a duration less than the preset audio duration, padding the audio segment to meet the preset audio duration.

4. The method according to claim 3, wherein calculating the speech rate through the following formula: Among them, v represents the speech rate, n represents the total number of audio segments into which the speech data to be analyzed is divided, l represents the total duration of the speech data to be analyzed, and ρ(x i ) represents the number of syllables output by the i-th audio segment input into the preset syllable number regression model ρ.

5. The method according to claim 1, characterized in that, The preset syllable number regression model is trained in the following manner: constructing a training data set, which includes: the acoustic features of audio samples and the actual number of syllables included in the text corresponding to the audio samples; inputting the acoustic features of each audio sample in the training data set into an initial syllable number regression model to obtain the predicted number of syllables corresponding to each audio sample; adjusting the model parameters of the initial syllable number regression model based on the relationship between the predicted number of syllables and the actual number of syllables of each audio sample until meeting the preset training requirements of the model to obtain the preset syllable number regression model.

6. The method according to claim 1, characterized in that, The preset syllable number regression model is a neural network model.

7. A speech rate analysis system, characterized in that, including: an obtaining module, configured to obtain the speech data to be analyzed and its corresponding total duration; a first processing module, configured to extract the total number of syllables included in the speech data to be analyzed; a second processing module, configured to determine the speech rate of the speech data to be analyzed based on the total number of syllables and the total duration; wherein, the first processing module includes: a dividing unit, configured to divide the speech data to be analyzed into multiple audio segments based on the total duration of the speech data to be analyzed and a preset audio duration; an extracting unit, configured to extract the acoustic features of each audio segment; an inputting unit, configured to input the acoustic features corresponding to each audio segment into a preset syllable number regression model to obtain the number of syllables corresponding to each audio segment; a summing unit, configured to sum up the number of syllables corresponding to all audio segments to obtain the total number of syllables; Among them, an extraction unit is configured to input the audio clip into a trained deep neural network model to obtain voice features. Among them, the backbone network of the deep neural network model is used to output several deep features of the audio clip, aggregate the deep features to obtain a dense feature of the audio clip, and use the dense feature of the audio clip as the voice feature.

8. An electronic device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. Computer instructions are stored in the memory, and the processor executes the computer instructions to execute the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method for estimating speech speed of multiple speakers based on segmentation and clustering of speakers

    CN102543063A