Intelligent sentence segmentation method based on voice spectrogram, computer device and storage medium

Through the spectral pattern-based pronunciation sentence classification method, the convolutional neural network is used to identify pause categories in the speech data, and the problem of insufficient accuracy of pronunciation sentence classification in the prior art is solved, especially the recognition effect of European Portuguese, and efficient pronunciation sentence classification is achieved.

CN114512118BActive Publication Date: 2025-07-11MACAO POLYTECHNIC INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210005950.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-04
Publication Date
2025-07-11
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

The existing pronunciation and sentence classification methods have insufficient accuracy, especially for non-common languages such as European Portuguese, which are not ideal for speech recognition. The existing methods often require multiple steps of speech recognition and sentence classification processing, resulting in inefficiency.

Method used

By obtaining the spectrum map of speech data, identifying the spectrum mute segment, and using a preset classification model, combining the pre- and post-spectrogram features, a convolutional neural network is used for pause category recognition to realize intelligent sentence segmentation of speech files.

Benefits of technology

It improves the accuracy of pronunciation and sentence classification process, and is suitable for multiple languages, especially for non-common languages such as European Portuguese.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512118B_ABST
    Figure CN114512118B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent sentence segmentation method, a computer device and a storage medium based on a voice spectrogram. The method includes: obtaining voice data to be segmented, and converting the voice data to be segmented into spectrogram data to be segmented; identifying spectrogram silent segments according to the spectrogram data to be segmented; obtaining a pre-spectrum of a first preset duration before the spectrogram silent segment and a post-spectrum of a second preset duration after the spectrogram silent segment, and combining the pre-spectrum and the post-spectrum into a spectrogram to be recognized; using a preset classification model to recognize the spectrogram to be recognized, and confirming the pause category of the spectrogram silent segment; and performing sentence segmentation on the voice file according to the pause category. By applying the intelligent sentence segmentation method based on the voice spectrogram of the present invention, the accuracy of voice sentence segmentation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition. Specifically, it relates to an intelligent sentence segmentation method based on a voice spectrogram, a computer device applying the intelligent sentence segmentation method based on the voice spectrogram, and a computer-readable storage medium applying the intelligent sentence segmentation method based on the voice spectrogram. Background Art

[0002] Language is an important tool for people to convey and obtain information. When sound is transmitted to a person's ear through a medium, the brain processes the speech and forms its own understanding, and then responds with language or actions. Making a computer understand human language depends on an important technology for human-computer interaction - speech recognition technology. Computer speech recognition, that is, speech-to-text (STT) or automatic speech recognition (ASR), is the process by which a computer recognizes spoken language and translates it into text.

[0003] One of the core factors for the accuracy of automatic speech recognition (ASR) is the collection of a parallel corpus (speech and text) and ensuring its quality. Although the preprocessing work of the corpus is cumbersome and time-consuming, it is still a crucial step. The preprocessing work includes data scraping, sentence segmentation, corpus alignment, noise removal, speech normalization, etc.

[0004] There are many ways to automatically split a piece of sound into smaller files by a computer, and there are also some methods in the field of natural language processing to automatically divide human speech into sentences (speech sentence segmentation algorithms). Currently, there are several common methods for splitting speech: the first is to split by a fixed duration, the second is to split by a fixed file size, the third is to automatically detect the silent segments in the speech and use the silent segments for splitting, and the fourth is to first recognize the speech into text and then use the corresponding language model to split the text into sentences.

[0005] These existing methods all have certain problems: 1. Since the content of a sentence is not of a fixed length, the first and second methods are not suitable for most cases; 2. The third method can only split a piece of speech into many small clauses, but in most cases, they are short sentences or phrases and cannot be intelligently split into complete sentences; 3. The fourth case involves speech recognition and sentence segmentation technologies, that is, it requires two steps of conversion and processing, and the effect of sentence segmentation depends on the result of speech recognition. Although the current accuracy of speech recognition is very good, it is only limited to common languages such as English and Mandarin, and the speech recognition accuracy of other languages such as European Portuguese is still not very satisfactory. Therefore, a more accurate speech sentence segmentation method is needed. Summary of the Invention

[0006] The first object of the present invention is to provide an intelligent sentence segmentation method based on a voice spectrogram that can effectively improve the accuracy of speech sentence segmentation.

[0007] The second object of the present invention is to provide a computer device that can effectively improve the accuracy of speech clause segmentation.

[0008] The third object of the present invention is to provide a computer-readable storage medium that can effectively improve the accuracy of speech clause segmentation.

[0009] To achieve the above first object, the intelligent clause segmentation method based on the voice spectrogram provided by the present invention includes: obtaining the speech data to be clause-segmented, and converting the speech data to be clause-segmented into spectrogram data to be clause-segmented; identifying the spectral silent segments according to the spectrogram data to be clause-segmented; obtaining the pre-spectrum of the first preset duration before the spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment, and combining the pre-spectrum and the post-spectrum into a spectrogram to be recognized; using a preset classification model to recognize the spectrogram to be recognized, and confirming the pause category of the spectral silent segment; and performing sentence segmentation on the speech file according to the pause category.

[0010] As can be seen from the above solution, when preprocessing the speech data to be clause-segmented in the intelligent clause segmentation method based on the voice spectrogram of the present invention, by obtaining the pre-spectrum of the first preset duration before the spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment, and combining the pre-spectrum and the post-spectrum into a spectrogram to be recognized, the preset classification model analyzes and recognizes the spectrogram to be recognized, and uses the voice spectrogram features before and after the speech pause, so as to determine the pause category of the spectral silent segment, and further the speech data to be clause-segmented can be sentence-segmented, thereby effectively improving the accuracy of speech clause segmentation.

[0011] In a further solution, the step of combining the pre-spectrum and the post-spectrum into a spectrogram to be recognized includes: adding a silent audio spectrum of the third preset duration between the pre-spectrum and the post-spectrum to obtain the spectrogram to be recognized.

[0012] Thus, when obtaining the spectrogram to be recognized, by adding a silent audio spectrum of the third preset duration between the pre-spectrum and the post-spectrum, it is convenient to recognize the audio before and after the pause, and the spectrogram to be recognized can meet the recognition standard, which is convenient for model analysis.

[0013] In a further solution, the value range of the third preset duration is 1 / 5 to 1 / 4 of the total spectral duration in the spectrogram to be recognized.

[0014] Thus, the value range of the third preset duration is one-fourth or one-fifth of the total spectral duration in the spectrogram to be recognized, which can ensure the accuracy and recognition rate of recognition.

[0015] In a further solution, the second preset duration is three times the first preset duration.

[0016] It can be seen that since the audio at the beginning of a sentence is more representative than the audio at the end of the sentence, increasing the proportion of the audio at the beginning of the sentence can improve the recognition accuracy.

[0017] In a further solution, the steps of identifying the spectral silent segment according to the spectrogram data to be segmented include: when the frequency amplitude in the spectrogram data to be segmented is less than a preset value and lasts for a preset duration, then this spectral segment is considered as the spectral silent segment.

[0018] It can be seen that when identifying the spectral silent segment, by identifying the frequency amplitude of the audio, when the frequency amplitude is less than the preset value and lasts for the preset duration, then it can be considered that the spectral silent segment appears.

[0019] In a further solution, the preset classification model is obtained by learning with a convolutional neural network.

[0020] It can be seen that using a convolutional neural network to learn and obtain the preset classification model can improve the recognition accuracy of the model.

[0021] In a further solution, the steps of learning with a convolutional neural network include: obtaining the spectrogram data corresponding to the training speech data; performing pause category annotation on all the spectral silent segments in the spectrogram data; obtaining the pre-spectrum of the first preset duration before each spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment in the spectrogram data to form the training spectrogram; using the convolutional neural network algorithm to perform model training on the training spectrogram to obtain the preset classification model.

[0022] It can be seen that when learning with a convolutional neural network to obtain the preset classification model, by performing pause category annotation on all the spectral silent segments in the spectrogram data, it is convenient for recognition and classification judgment. At the same time, obtaining the pre-spectrum of the first preset duration before each spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment in the spectrogram data to form the training spectrogram can reduce the amount of data analysis and improve the model training efficiency.

[0023] In a further solution, after the step of segmenting the speech data to be segmented according to the pause category, it further includes: storing the segmented sentences in a preset format.

[0024] It can be seen that storing the segmented sentences in a preset format is convenient for subsequent application processing.

[0025] To achieve the second object of the present invention, the present invention provides a computer device including a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the above-mentioned intelligent sentence segmentation method based on the voice spectrogram are implemented.

[0026] To achieve the third object of the present invention, the computer-readable storage medium provided by the present invention stores a computer program, and when the computer program is executed by a controller, the steps of the above-mentioned intelligent sentence segmentation method based on a voice spectrogram are implemented. Description of the Drawings

[0027] Figure 1 is a flowchart of an embodiment of the intelligent sentence segmentation method based on a voice spectrogram of the present invention.

[0028] Figure 2 is a schematic diagram of a spectrogram to be recognized in an embodiment of the intelligent sentence segmentation method based on a voice spectrogram of the present invention.

[0029] Figure 3 is a flowchart of the convolutional neural network learning step in an embodiment of the intelligent sentence segmentation method based on a voice spectrogram of the present invention.

[0030] The present invention will be further described below with reference to the drawings and embodiments. Detailed Embodiments

[0031] Embodiment of the intelligent sentence segmentation method based on a voice spectrogram:

[0032] The intelligent sentence segmentation method based on a voice spectrogram in this embodiment is an application program applied to a computer and is used for performing sentence segmentation operations on voice data.

[0033] As Figure 1 shown, in this embodiment, when the intelligent sentence segmentation method based on a voice spectrogram is working, first, step S1 is executed to obtain the voice data to be segmented and convert the voice data to be segmented into spectrogram data to be segmented. Before performing sentence segmentation on the voice data, the voice data to be segmented is first obtained. The voice data to be segmented can be real-time obtained voice data or pre-recorded voice data. After obtaining the voice data to be segmented, in order to facilitate the analysis of the voice data, it is necessary to convert the voice data into spectrogram data. Converting voice data into spectrogram data is a well-known technology to those skilled in the art and will not be elaborated here.

[0034] After converting the voice data to be segmented into spectrogram data to be segmented, step S2 is executed to identify the spectral silent segments according to the spectrogram data to be segmented. In natural language, due to people's speaking habits or language rules, when a complete sentence ends, there will be a pause for a period of time before the next sentence. Therefore, there are usually pauses in voice data, that is, spectral silent segments will appear in the audio, and the end of a sentence can be judged by the spectral silent segments. However, spectral silent segments will also appear when there are non-sentence-ending pauses such as commas, semicolons or insertions. Therefore, it is necessary to identify sentence-ending pauses and non-sentence-ending pauses, but first, it is necessary to identify the silent segments in the audio.

[0035] In this embodiment, the steps of identifying the spectral silent segment according to the spectral diagram data to be segmented include: when the frequency amplitude in the spectral diagram data to be segmented is less than the preset value and lasts for the preset duration, it is considered that this spectral segment is the spectral silent segment. When a pause occurs, the frequency amplitude in the spectral diagram will last for a period of lower amplitude. By identifying the frequency amplitude of the audio, when the frequency amplitude is less than the preset value and lasts for the preset duration, it can be considered that the spectral silent segment appears.

[0036] After identifying the spectral silent segment, perform step S3 to obtain the pre-spectrum of the first preset duration before the spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment, and combine the pre-spectrum and the post-spectrum into the spectral diagram to be recognized. Through research, in some languages, the words and phrases commonly used at the beginning and end of sentences. For example, in Portuguese, among the most commonly used words at the beginning of sentences, the frequencies of "O" and "A" are 7.88% and 5.15% respectively, and the frequencies are 0.54% and 0.47% respectively. The words and phrases commonly used at the beginning and end will have corresponding spectra. Therefore, using the spectral diagram features of the speech before and after the speech pause can facilitate the identification of whether the spectral silent segment is the end pause of the sentence. The first preset duration and the second preset duration are set in advance according to the experimental data. Since the audio at the beginning of the sentence is more representative than the audio at the end of the sentence, increasing the proportion of the audio at the beginning of the sentence can improve the accuracy of recognition. In this embodiment, the second preset duration is three times the first preset duration. Preferably, the first preset duration is 100 ms and the second preset duration is 300 ms.

[0037] In this embodiment, the step of combining the pre-spectrum and the post-spectrum into the spectral diagram to be recognized includes: adding a silent audio spectrum of the third preset duration between the pre-spectrum and the post-spectrum to obtain the spectral diagram to be recognized. As Figure 2 shown, combine the pre-spectrum 1, the silent audio spectrum 2 and the post-spectrum 3 into the spectral diagram to be recognized. In order to facilitate the identification of the pre-spectrum and the post-spectrum, a section of silent audio spectrum is set between the pre-spectrum and the post-spectrum to facilitate the identification of the audio before and after the pause, and setting the third preset duration can make the spectral diagram to be recognized meet the model recognition standard and facilitate the model analysis. In order to ensure the accuracy and recognition rate of the recognition, the third preset duration can be set in advance according to the experimental data. In this embodiment, the value range of the third preset duration is 1 / 5 to 1 / 4 of the total spectral duration in the spectral diagram to be recognized.

[0038] After obtaining the spectral diagram to be recognized, perform step S4 to use the preset classification model to recognize the spectral diagram to be recognized and confirm the pause category of the spectral silent segment. Among them, the pause category includes end pause and non-end pause. In this embodiment, the preset classification model is obtained by learning with a convolutional neural network. Using a convolutional neural network to learn to obtain the preset classification model can improve the accuracy of model recognition.

[0039] See Figure 3 In this embodiment, when performing convolutional neural network learning, step S41 is first executed to obtain spectrogram data corresponding to training speech data. In order to enable the training model to accurately classify, a large amount of speech data needs to be learned. In the field of deep learning, the application of image classification emerged earlier and is more mature than others. The training set can also train an accurate model with a smaller number of contents, and the efficiency is much higher than that of training speech recognition applications. Therefore, it is necessary to obtain training speech data, process it to obtain corresponding spectrogram data, and use the spectrogram data for model training.

[0040] After obtaining the spectrogram data corresponding to the training speech data, step S42 is executed to label the pause categories for all spectrogram silent segments in the spectrogram data. When performing the pause category labeling, each spectrogram silent segment is manually labeled. When the spectrogram silent segment corresponds to the end of a speech sentence, the number "1" is labeled to indicate an end-of-sentence pause. When the end of the speech is not the end of a sentence, that is, when there is a pause caused by a comma, semicolon, or parenthesis, or it is judged from the tone and intonation that the sentence is not completed, the number "0" is labeled to indicate a non-end-of-sentence pause.

[0041] After performing the pause category labeling, step S43 is executed to obtain a training spectrogram composed of a pre-spectrum of a first preset duration before each spectrogram silent segment and a post-spectrum of a second preset duration after the spectrogram silent segment in the spectrogram data. In order to reduce the amount of data analysis and improve the model training speed, a training spectrogram is obtained by combining the pre-spectrum of a first preset duration before each spectrogram silent segment and the post-spectrum of a second preset duration after the spectrogram silent segment in the spectrogram data. The training spectrogram has the same flat spectrum structure as the spectrogram to be recognized, that is, a silent audio spectrum of a third preset duration is set between the pre-spectrum and the post-spectrum. The obtained training spectrogram is stored in the training spectrogram library for training use.

[0042] After obtaining the training spectrogram, step S44 is executed to perform model training on the training spectrogram using the convolutional neural network algorithm to obtain a preset classification model. Using the convolutional neural network algorithm for classification model training is a well-known technology to those skilled in the art and will not be elaborated here. In this embodiment, a preset classification model that can recognize whether the spectrogram silent segment is an end-of-sentence pause or a non-end-of-sentence pause is obtained using the convolutional neural network algorithm.

[0043] After confirming the pause category of the spectrum silent segment, step S5 is executed to perform sentence segmentation on the speech file according to the pause category. After confirming that the pause category of the spectrum silent segment is an end-of-sentence pause or a non-end-of-sentence pause, sentence segmentation can be performed on the speech file according to the pause category, thereby completing the clause division of the sentence. For example, if the pause category of the current spectrum silent segment is an end-of-sentence pause, it indicates that the sentence is complete, and the current spectrum silent segment is used as a segmentation point for segmentation; if the pause category of the current spectrum silent segment is a non-end-of-sentence pause, it indicates that the sentence is not yet complete, and the next spectrum silent segment needs to be continuously judged until the pause category of the spectrum silent segment is an end-of-sentence pause for segmentation. By performing sentence segmentation on the speech file, all sentences in the speech file can be identified.

[0044] Of course, if there is no spectrum silent segment in a section of speech, it means that there is only one sentence in this section of speech and no segmentation is required.

[0045] After sentence segmentation, step S6 is executed to store the segmented sentences in a preset format. The preset format can be set according to application requirements. In this embodiment, the preset format is the MP3 format. Storing the segmented sentences in a preset format can facilitate subsequent applications, such as text conversion, translation, editing, etc.

[0046] As can be seen from the above, when the intelligent clause division method based on the voice spectrogram of the present invention preprocesses the voice data to be divided into clauses, by obtaining the pre-spectrum of the first preset duration before the spectrum silent segment and the post-spectrum of the second preset duration after the spectrum silent segment, combining the pre-spectrum and the post-spectrum into a spectrum to be recognized, and analyzing and recognizing the spectrum to be recognized by a preset classification model, using the voice spectrogram features before and after the voice pause, thereby determining the pause category of the spectrum silent segment, and further the voice data to be divided into clauses can be subjected to sentence segmentation, thereby effectively improving the accuracy of voice clause division.

[0047] Embodiment of the computer device:

[0048] The computer device of this embodiment includes a controller, and when the controller executes a computer program, it implements the steps in the above-mentioned embodiment of the intelligent clause division method based on the voice spectrogram.

[0049] For example, the computer program can be divided into one or more modules, and one or more modules are stored in the memory and executed by the controller to complete the present invention. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.

[0050] The computer device may include, but is not limited to, a controller and a memory. Those skilled in the art can understand that the computer device may include more or fewer components, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0051] For example, the controller may be a Central Processing Unit (CPU), or may also be other general controllers, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general controller may be a microcontroller or the controller may also be any conventional controller, etc. The controller is the control center of the computer device, and connects various parts of the entire computer device using various interfaces and lines.

[0052] The memory can be used to store computer programs and / or modules. The controller realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. For example, the memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, and so on. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0053] Examples of computer-readable storage media:

[0054] If the modules integrated in the computer device of the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiments of the intelligent clause-splitting method based on voice spectrogram, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a controller, the steps of the above embodiments of the intelligent clause-splitting method based on voice spectrogram can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0055] It should be noted that the above are only the preferred embodiments of the present invention, but the design concept of the invention is not limited thereto. Any non-substantive modifications made to the present invention using this concept also fall within the protection scope of the present invention.

Claims

1. An intelligent sentence splitting method based on voice spectrogram, characterized in that: Including: Obtain the speech data to be segmented, and convert the speech data to be segmented into spectrogram data to be segmented; Identify the spectral silent segments according to the spectrogram data to be segmented; Obtain the pre-spectrum of the first preset duration before the spectral silent segment and the post-spectrum of the second preset duration after the spectral silent segment, and combine the pre-spectrum and the post-spectrum into a spectrogram to be recognized; Use a preset classification model to recognize the spectrogram to be recognized, and confirm the pause category of the spectral silent segment; Segment the speech data to be segmented according to the pause category; Among them, the step of combining the pre-spectrum and the post-spectrum into a spectrogram to be recognized includes: adding a static spectrogram of a third preset duration between the pre-spectrum and the post-spectrum to obtain the spectrogram to be recognized.

2. The intelligent sentence segmentation method based on voice spectrogram according to claim 1, wherein: The value range of the third preset duration is 1 / 5 to 1 / 4 of the total spectral duration in the spectrogram to be recognized.

3. The intelligent sentence segmentation method based on voice spectrogram according to claim 2, wherein: The second preset duration is three times the first preset duration.

4. The intelligent sentence segmentation method based on voice spectrogram according to any one of claims 1 to 3, wherein: The step of identifying the spectral silent segment according to the spectrogram data to be segmented includes: When the frequency amplitude in the spectrogram data to be segmented is less than a preset value and lasts for a preset duration, it is considered that this spectral segment is a spectral silent segment.

5. The intelligent sentence segmentation method based on voice spectrogram according to any one of claims 1 to 3, wherein: The preset classification model is obtained by learning with a convolutional neural network.

6. The intelligent sentence segmentation method based on voice spectrogram according to claim 5, wherein: The steps of learning with the convolutional neural network include: Obtain the spectrogram data corresponding to the training speech data; Perform pause category annotation on all the spectral silent segments in the spectrogram data; Obtain the pre-spectrum of the first preset duration before each spectral silent segment in the spectrogram data and the post-spectrum of the second preset duration after the spectral silent segment to form a training spectrogram; Use the convolutional neural network algorithm to perform model training on the training spectrogram to obtain the preset classification model.

7. The intelligent sentence segmentation method based on voice spectrogram according to any one of claims 1 to 3, wherein: After the step of segmenting the speech data to be segmented according to the pause category, it further includes: Store the segmented sentences in a preset format.

8. A computer device, comprising a processor and a memory, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the steps of the intelligent sentence segmentation method based on voice spectrogram according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the controller, the steps of the intelligent sentence segmentation method based on voice spectrogram according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Audio processing method and device, electronic equipment and storage medium

    CN112669822A

  • Voice detection method and device, computer equipment and storage medium

    CN112802498A

  • Voice interaction system, related method, device and equipment

    CN113160854A