Chinese dialect speech automatic segmentation method and device

Through voice collection and template construction based on dialect sound and rhyme coordination rules, combined with dynamic time alignment technology and general speech detection technology, automatic segmentation of Chinese dialect pronunciation is realized, solving the problems of poor applicability and limited data of automatic speech segmentation in Chinese dialects in the existing technology, and improving segmentation efficiency and accuracy.

CN120108382APending Publication Date: 2025-06-06BEIJING LANGUAGE AND CULTURE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212038.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively apply to automatic speech segmentation in Chinese dialects, especially when the annotation processing voice data is limited, it is difficult to train a robust unit model.

Method used

By using dialect voice and rhyme coordination rules to collect and construct dialect voice templates, combined with dynamic time alignment technology and mean calculation method, dialect voice templates are constructed, and the sound-like area positioning is used to use general speech detection technology to filter phoneme segments through frame-by-frame matching and matching distance methods to realize automatic segmentation of Chinese dialect voice.

Benefits of technology

On the premise of saving training costs and time, the efficiency and accuracy of automatic speech segmentation are improved, and are suitable for automatic speech segmentation tasks in any Chinese dialect, expanding the application field and scope of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108382A_ABST
    Figure CN120108382A_ABST
Patent Text Reader

Abstract

The invention provides a Chinese dialect speech automatic segmentation method and device, and relates to the technical field of speech signal processing. The method comprises the following steps: constructing a dialect voice sample based on a dialect initial and final matching rule; constructing a dialect voice template according to the dialect voice sample based on a dynamic time warping technology and a mean value calculation method; obtaining a to-be-segmented dialect voice; according to the to-be-segmented dialect voice, voice type area positioning is carried out by using a general voice detection technology, and a sound blocking area voice, a vowel area voice and a non-vowel area voice are obtained; based on the dialect voice template, frame-by-frame matching is carried out according to the voice of the sound blocking area, the voice of the vowel area and the voice of the non-vowel area, and a potential phoneme fragment set is obtained; and based on a dialect initial and final matching rule, according to the potential phoneme fragment set, carrying out phoneme fragment screening by using a matching distance method to obtain a dialect voice segmentation result. The invention relates to a Chinese dialect speech automatic segmentation method based on speech data processed by a small amount of annotations, aiming at Chinese dialect characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and in particular to a method and device for automatic segmentation of Chinese dialect speech. Background Art

[0002] In the speech resource construction project, an important basic work is to divide the continuous speech signal into many speech segments. Each speech segment corresponds to the basic unit of language, such as phoneme, syllable, or word. This process is usually called speech segmentation. In actual operation, the basic unit to be segmented is often determined according to the language characteristics and application requirements. Usually, the speech signals of Mandarin Chinese and dialects are segmented into initial consonants, finals, and syllables.

[0003] With the help of speech analysis software such as Praat, operators can manually determine the boundaries of speech segments. However, manual segmentation requires skilled domain experts to complete. Currently, there are mainly two methods for automatic speech segmentation: unit recognition-based methods and boundary detection-based methods. The unit recognition-based method uses the built unit model to identify the speech signal to be segmented into the most likely unit sequence, and uses dynamic programming to output the start and end time of each unit. The boundary detection-based method determines the potential unit boundaries based on the changes in speech attributes in the feature space.

[0004] Both of the above automatic speech segmentation methods have certain limitations. The unit recognition-based method is usually language-related, that is, the trained unit model is difficult to apply to other languages. In order to obtain accurate segmentation results, a large amount of accurately labeled training data and a complex network with many parameters are required to build a robust unit model, which runs counter to the goal of automatic speech segmentation. The boundary detection-based method is theoretically not restricted by language, but it needs to adjust parameters to meet different types of speech signals. The segmentation accuracy is not easy to control and it is difficult to meet and promote to application environments with large amounts of data.

[0005] As representatives of monosyllabic languages, Mandarin Chinese and its dialects have the same phonetic structure, that is, one Chinese character is one syllable. Each syllable consists of three parts: initial consonant, final vowel and tone. However, Mandarin Chinese and its dialects differ in the number of initial consonants, final vowels and tones, pronunciation realization and collocation relationship. In addition, the annotated speech data is relatively limited, which is particularly prominent in Chinese dialects. Therefore, it is difficult to extend the current automatic speech segmentation method suitable for Mandarin Chinese to Chinese dialects.

[0006] In the prior art, there is a lack of an automatic Chinese dialect speech segmentation method based on a small amount of annotated and processed speech data that is tailored to the characteristics of Chinese dialects. Summary of the invention

[0007] In order to solve the technical problems in the prior art that the applicability of Chinese dialects is poor and the speech data processed by Chinese dialect annotation is relatively limited and cannot train a robust unit model, the embodiment of the present invention provides a method and device for automatic segmentation of Chinese dialect speech. The technical solution is as follows:

[0008] On the one hand, a method for automatic segmentation of Chinese dialect speech is provided, which is implemented by a Chinese dialect speech automatic segmentation device, and the method comprises:

[0009] Based on the dialect consonant-rhyme coordination rules, use language acquisition equipment to collect dialect speech and construct dialect speech samples;

[0010] Based on dynamic time warping technology and mean calculation method, a dialect speech template is constructed according to the dialect speech samples;

[0011] Acquire the dialect speech to be segmented; use the general speech detection technology to locate the sound area according to the dialect speech to be segmented, and obtain the speech in the obstructed sound area, the speech in the vowel area, and the speech in the non-vowel area;

[0012] Based on the dialect speech template, frame-by-frame matching is performed according to the speech in the obstruent area, the speech in the vowel area, and the speech in the non-vowel area to obtain a set of potential phoneme segments;

[0013] Based on the dialect consonant-rhyme coordination rules and the potential phoneme segment set, the matching distance method is used to screen the phoneme segments to obtain the dialect speech segmentation results.

[0014] On the other hand, a device for automatic segmentation of Chinese dialect speech is provided, which is applied to the method for automatic segmentation of Chinese dialect speech, and the device comprises:

[0015] A dialect sample construction module is used to collect dialect speech based on the dialect consonant-rhythm coordination rules using a language collection device and to construct a dialect speech sample;

[0016] A dialect template construction module is used to construct a dialect speech template according to a dialect speech sample based on a dynamic time warping technology and a mean calculation method;

[0017] The sound category area positioning module is used to obtain the dialect speech to be segmented; according to the dialect speech to be segmented, the general speech detection technology is used to locate the sound category area to obtain the speech in the obscuration area, the speech in the vowel area and the speech in the non-vowel area;

[0018] The first dialect segmentation module is used to perform frame-by-frame matching based on the dialect speech template and the speech in the obstruent area, the vowel area and the non-vowel area to obtain a set of potential phoneme segments;

[0019] The second dialect segmentation module is used to screen the phoneme segments based on the dialect consonant-rhythm coordination rules and the potential phoneme segment set using the matching distance method to obtain the dialect speech segmentation result.

[0020] On the other hand, a device for automatic segmentation of Chinese dialect speech is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned methods for automatic segmentation of Chinese dialect speech is implemented.

[0021] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned methods for automatic segmentation of Chinese dialect speech.

[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0023] The present invention proposes a method for automatic segmentation of Chinese dialect speech. The initial and final templates are constructed by a small amount of annotated and processed speech data, combined with the detection results of the general sound category, and the continuous Chinese dialect speech signal is automatically segmented into speech segments of initials, finals and syllables by a template matching method. The efficiency and accuracy of automatic speech segmentation are improved while saving training costs and time costs. The method is suitable for promotion to the automatic speech segmentation task of any Chinese dialect, and the field and scope of application of the method are improved. The present invention is a method for automatic segmentation of Chinese dialect speech based on a small amount of annotated and processed speech data, which is targeted at the characteristics of Chinese dialects. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 This is a flow chart of a method for automatic segmentation of Chinese dialect speech provided by an embodiment of the present invention;

[0026] Figure 2 It is a schematic diagram of dynamic time warping matching of speech samples with different frame lengths provided by an embodiment of the present invention;

[0027] Figure 3 A schematic diagram of matching initial consonants or finals appearing in consecutive frames provided by an embodiment of the present invention;

[0028] Figure 4It is a schematic diagram of automatically segmenting a speech signal to be segmented into syllables, initial consonants and finals provided by an embodiment of the present invention;

[0029] Figure 5 It is a block diagram of a device for automatic segmentation of Chinese dialect speech provided by an embodiment of the present invention;

[0030] Figure 6 It is a structural schematic diagram of a Chinese dialect speech automatic segmentation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0032] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0033] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0034] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.

[0035] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0036] The embodiment of the present invention provides a method for automatic segmentation of Chinese dialect speech, which can be implemented by a Chinese dialect speech automatic segmentation device, and the Chinese dialect speech automatic segmentation device can be a terminal or a server. Figure 1 The flowchart of the method for automatic segmentation of Chinese dialect speech is shown in the figure. The processing flow of the method may include the following steps:

[0037] S1. Based on the dialect phonetic and rhyme coordination rules, use language acquisition equipment to collect dialect speech and construct dialect speech samples.

[0038] Optionally, based on the dialect consonant-rhyme coordination rules, a language collection device is used to collect dialect speech and construct a dialect speech sample, including:

[0039] Through the language collection equipment, the voice of common words is collected to obtain the dialect voice subset;

[0040] Based on the dialect consonant-rhyme coordination rules and according to the dialect speech subset, language analysis tools are used to annotate and obtain dialect speech samples.

[0041] In a feasible implementation mode, the present invention uses 3700 commonly used words in the "Dialect Survey Word List" as the collection object of voice samples. The 3700 commonly used words are designed and distributed into multiple word subsets, each subset contains 400-600 words, as long as they can cover all syllables composed of initials and finals, and overlapping words are allowed between each subset.

[0042] Combining the knowledge of dialect survey and experimental linguistics, multiple speakers of the same Chinese dialect were selected to collect speech samples of initials and finals, so as to cover people of different ages and genders. The age of 18-70 years old was divided into 3-5 age groups, and at least one male speaker and one female speaker were required for each age group. Each speaker was randomly assigned a single word subset to collect speech samples. Before recording, the speaker was required to be familiar with the content of the single word subset to ensure accurate and natural pronunciation.

[0043] Record the speech sample using Praat speech analysis software. Set the sampling frequency to 44100Hz and save the speech sample in .wav format. Manually annotate the speech sample using Praat speech analysis software to obtain the start and end boundaries of the initials and finals. The annotation process requires combining a lot of information to make decisions, such as spectrogram changes, waveform changes, and human perception, so that the annotation results have higher accuracy and rationality. Save the annotation results in TXT text.

[0044] Among them, the dialect phonetic and rhyme coordination rules include dialect initials, dialect finals, dialect tones and dialect syllables.

[0045] In a feasible implementation, linguistics or Chinese dialectology professionals conduct field surveys to obtain the phonetic system of the Chinese dialect under investigation, clarify the dialect consonant-vowel coordination rules of the dialect phonetic system, including the number and content of initials, finals and tones, and record them using international phonetic symbols. From the single-word phonetic table obtained from the survey, summarize the coordination rules of initials and finals, and clarify which initials and finals can be combined into single-word syllables.

[0046] S2. Based on dynamic time warping technology and mean calculation method, a dialect speech template is constructed according to the dialect speech samples.

[0047] Optionally, based on the dynamic time warping technology and the mean calculation method, a dialect speech template is constructed according to the dialect speech sample, including:

[0048] Performing frame processing on the dialect speech sample to obtain a framed speech sample;

[0049] Based on the short-time average energy and Mel frequency cepstral coefficients, the voice sample features are obtained by calculating the mean of the voices with the same content in the framed voice samples;

[0050] Based on dynamic time warping technology, the matching distance method is used to screen samples according to the framed speech samples to obtain the optimal dialect samples;

[0051] A dialect speech template is constructed according to the speech sample features and the frame number of the optimal dialect sample.

[0052] In a feasible implementation manner, the present invention takes into account that the duration of initial consonants relative to finals is relatively short. In order to obtain a better description, the frame length in the framing processing parameters is set to 10ms and the frame offset is set to 5ms. The frame feature vector is composed of the short-time average energy of the time domain features and the Mel-frequency cepstral coefficients of the frequency domain features. The Mel-frequency cepstral coefficients of the 12-dimensional MFCC coefficients and the short-time average energy of 1 dimension are extracted for each frame. They are normalized and combined with their first-order differences and second-order differences to form a total of 39-dimensional frame feature vectors. The implementation method of framing, Mel-frequency cepstral coefficients and short-time average energy is widely used in the field of natural language processing technology and will not be repeated here.

[0053] Select a framed speech sample as a temporary template, and use dynamic time warping technology to match other speech samples with the temporary template, such as Figure 2 As shown. In order to calculate the matching distance between a speech sample and a temporary template, the distance between their corresponding frames should be calculated, and the Euclidean distance is usually used. The matching distance between each speech sample and the temporary template is accumulated to obtain the sum. The speech samples are selected as temporary templates in turn, and the above operation is repeated. The use of dynamic time warping technology for sample matching and matching distance calculation is a common technical means in this field and will not be repeated here.

[0054] Select the speech sample with the smallest sum to construct the template of the initial or final. The number of frames of the framed speech sample is used as the number of frames of the template, and the average value of the feature vectors of the aligned frames is used as the frame feature vector of the template. Finally, the template of the initial or final is represented as the frame feature vector Number of frames.

[0055] Constructing the template of initials or finals is to determine its representation method and specific parameters. The present invention represents the template of initials or finals as time series data, so it is necessary to determine the number of frames and the feature vector of each frame. Since dynamic time warping has matched speech samples with different frame lengths. The number of frames of the speech sample with the smallest matching distance is selected as the number of frames of the template, and the average value of the feature vectors of the aligned frames can be used as the frame feature vector of the template.

[0056] S3, obtaining the dialect speech to be segmented; according to the dialect speech to be segmented, using the general speech detection technology to locate the sound area, and obtaining the speech in the obstructed sound area, the speech in the vowel area and the speech in the non-vowel area.

[0057] Optionally, according to the dialect speech to be segmented, a general speech detection technology is used to locate the sound category area to obtain the speech in the obstructed sound area, the speech in the vowel area and the speech in the non-vowel area, including:

[0058] According to the dialect speech to be segmented, the speech activity detection technology is used to locate the effective sound area and obtain the speech with sound area;

[0059] According to the voice in the sound area, the sonorous sound area detection technology is used to segment the sound area to obtain the speech in the sonorous sound area and the speech in the obstructive sound area;

[0060] According to the sonorous area speech, the vowel area detection technology is used to segment the sound area to obtain the vowel area speech and the non-vowel area speech.

[0061] In a feasible implementation, in order to avoid the large amount of calculation and low accuracy caused by directly using template matching on the speech signal to be segmented, the present invention uses a general sound class detection technology to preferentially find the effective speech area, thereby reducing and clarifying the scope of template matching. This operation is due to the fact that the initial consonants in Chinese dialects mainly correspond to the consonant sound class in speech, and the finals are mainly composed of the sonorants in vowels and consonants.

[0062] Phonetic research divides human speech into two categories: sonorants and obstruents. Sonorants include nasals, approximants and linguistic consonants in vowels and consonants, and obstruents include fricatives, stops and affricates in consonants. Speech signals are usually composed of background noise, pauses and silences, and effective speech that transmits information. Voice activity detection technology can locate the sound area in which information is transmitted, and regard background noise and pauses and silence as silent areas. The operation of implementing the location of the sound area can refer to the common method of existing voice activity detection, and the present invention will not be repeated here.

[0063] Sonorous sounds are an important part of speech, and humans tend to pronounce and perceive most sonorous sounds clearly. Localizing the sonorous sound area can improve the efficiency of many downstream automatic speech processing. The operation of implementing the sonorous sound area localization can refer to the common method of existing sonorous sound area detection, and the present invention will not be described in detail here.

[0064] Vowels play an important role in constructing syllables and words in language, and can serve as the core segment of a syllable. They are also characterized by long duration, stable pronunciation, and quasi-periodicity, making them easy to detect in speech signals. Vowel area localization has been proven to improve the efficiency and accuracy of many downstream automatic speech processing. The operation of implementing vowel area localization can refer to the commonly used methods of existing vowel area detection, and the present invention will not elaborate on them here.

[0065] S4. Based on the dialect speech template, frame-by-frame matching is performed according to the speech in the obscurant area, the speech in the vowel area and the speech in the non-vowel area to obtain a set of potential phoneme segments.

[0066] Optionally, based on the dialect speech template, frame-by-frame matching is performed according to the speech in the obstructed region, the speech in the vowel region, and the speech in the non-vowel region to obtain a set of potential phoneme segments, including:

[0067] Based on the dialect speech template, frame processing is performed according to the speech in the obstruent area, the speech in the vowel area and the speech in the non-vowel area, and a frame feature vector matrix is ​​constructed;

[0068] Based on the preset number of frames, according to the frame feature vector matrix and the dialect speech template, the distance matching calculation is performed frame by frame through the dynamic time warping technology to obtain a frame-by-frame matching result set;

[0069] Perform statistical calculations based on the frame-by-frame matching result set to obtain a matching distance threshold;

[0070] Based on the matching distance threshold, the phoneme boundary area is judged for the speech in the obscurant area, the speech in the vowel area and the speech in the non-vowel area to obtain a set of potential phoneme segments.

[0071] In a feasible implementation, a region is selected from the speech signal and represented as a frame feature vector The number of frames is then compared with the initial or final template to calculate the matching distance. If the matching distance is lower than the matching distance threshold, it is determined that the initial or final exists in the region.

[0072] The same frame processing technology and feature vector description method as used in constructing the dialect speech template are used. Here, the frame processing technology and feature vector description method used in template construction must be the same.

[0073] Select a frame as the starting point in the frame feature vector matrix, select an area with a set frame length, and calculate the matching distance between the selected area and the initial or final template through dynamic time warping. The simplest way to set the frame length is to be the same as the frame length of the initial or final. In order to cope with the changes in the duration of the pronunciation of the same initial or final by different people or the same person, multiple frames of different lengths can be set. For example, the frame length is set to 0.5 times, 0.75 times, 1.0 times, 1.25 times and 1.5 times the template frame length, etc. Select a frame as the starting point, and select multiple areas with a set frame length. Calculate the matching distance between each selected area and the initial or final template. Only the frame length with the smallest matching distance and its distance value are retained as the matching result of the frame. The operation of implementing distance calculation also uses dynamic time warping technology.

[0074] Repeat the above operation frame by frame to obtain the matching distance of each frame and statistically obtain the mode and standard deviation. Use the value of the mode minus the standard deviation as the distance threshold. For each frame, obtain the frame length with the minimum matching distance and its distance value. Use statistical tools to obtain the mode and standard deviation from these distance values.

[0075] If the matching distance of several consecutive frames is lower than the matching distance threshold, it is determined that the initial consonant or final appears in the selected area corresponding to these frames. The selected area with the smallest matching distance is used as the starting and ending boundary of the initial consonant or final; the number of consecutive frames is set to 3 to 5 frames, that is, when 3 consecutive frames, 4 frames or 5 frames are used as the starting point, the distance value between the selected area and the initial consonant or final template is less than the distance threshold, then the initial consonant or final is considered to appear in the selected area corresponding to these frames. After matching several consecutive frames, the situation of the initial consonant or final appearing is as follows. Figure 3 As shown. There are many ways to determine the boundaries of initials or finals from these selected areas. The present invention selects the simplest way, and uses the selected area with the smallest matching distance as the start and end boundaries of the initials or finals. After dividing the phoneme boundaries of the initials or finals, a plurality of potential phoneme fragment sets composed of initials and finals are finally obtained.

[0076] S5. Based on the dialect consonant-rhyme coordination rules and the potential phoneme segment set, the matching distance method is used to screen the phoneme segments to obtain the dialect speech segmentation result.

[0077] Optionally, based on the dialect consonant-rhyme matching rules and according to the potential phoneme segment set, the phoneme segment is screened using a matching distance method to obtain the dialect speech segmentation result, including:

[0078] Based on the dialect phonetic and rhyme coordination rules, the dialect syllable rules are screened according to the potential phoneme segment set to obtain the segment set that conforms to the syllable rule and the corresponding segment syllable set;

[0079] Based on the dialect phoneme matching rules, according to the set of segments that conform to the syllable rules, the matching distance method is used to screen the dialect phoneme rules to obtain the optimal speech segmentation segment;

[0080] Obtaining a syllable segmentation result according to the optimal speech segmentation segment and the corresponding segment syllable set;

[0081] Based on the dialect consonant-resonance coordination rules, the phoneme region matching is performed according to the optimal speech segmentation segment to obtain the initial consonant segmentation result and the final consonant segmentation result;

[0082] According to the syllable segmentation results, the initial consonant segmentation results and the final vowel segmentation results, the dialect speech segmentation results are obtained.

[0083] In a feasible implementation manner, in the matching process of the present invention, prior knowledge and language information such as context, phonetic grammar, etc. are not used as guidance. Therefore, after template matching, a plurality of sequences consisting of initial consonants and final vowels are usually generated. In order to reduce the workload brought by building a high-level language knowledge model, the embodiment of the present invention chooses to use the initial and final vowel matching rules and the principle of the minimum sum of matching distances to screen and determine the final initial consonant, final vowel and syllable segmentation results.

[0084] For each potential phoneme segment sequence composed of initials and finals generated. Check the first speech segment of the sequence, which can be an initial consonant or a final. If it is a final, it must be a final that can form a syllable. Check the last speech segment of the sequence, which must be a final. Check the sequence from front to back, the initial consonant must be followed by a final, and the combination must comply with the rules of initial-final coordination. The final can be followed by an initial consonant or a final. If it is a final, it must be a final that can form a syllable.

[0085] When there are multiple potential phoneme segments of initials and finals that match the syllable structure, the sequence with the smallest sum of matching distances is selected as the final segmentation result. Correspondingly, the start and end points of the syllable correspond to the starting point of the initial consonant and the end point of the final.

[0086] After screening, there may be multiple syllable-compliant segments of initials and finals that conform to the syllable structure. These sequences may be partially identical and overlapped, or they may be completely different. The present invention does not introduce a language model to generate or determine which of these sequences is more grammatical. The sum of the matching distances of the initials and finals in the sequence is used as the basis for selection, and the sequence with the smallest sum of matching distances is finally retained as the optimal speech segmentation segment. The initials and finals form the start and end boundaries of the syllable. The speech signal to be segmented is automatically segmented into multiple syllables, and each syllable is segmented into initials and finals, such as Figure 4 shown.

[0087] The present invention proposes a method for automatic segmentation of Chinese dialect speech. The initial and final templates are constructed by a small amount of annotated and processed speech data, combined with the detection results of the general sound category, and the continuous Chinese dialect speech signal is automatically segmented into speech segments of initials, finals and syllables by a template matching method. The efficiency and accuracy of automatic speech segmentation are improved while saving training costs and time costs. The method is suitable for promotion to the automatic speech segmentation task of any Chinese dialect, and the field and scope of application of the method are improved. The present invention is a method for automatic segmentation of Chinese dialect speech based on a small amount of annotated and processed speech data, which is targeted at the characteristics of Chinese dialects.

[0088] Figure 5 1 is a block diagram of a device for automatic segmentation of Chinese dialect speech according to an exemplary embodiment, wherein the device is used in a method for automatic segmentation of Chinese dialect speech. Figure 5 The device includes a dialect sample construction module 510, a dialect template construction module 520, a sound category area positioning module 530, a first dialect segmentation module 540 and a second dialect segmentation module 550. Among them:

[0089] The dialect sample construction module 510 is used to collect dialect speech based on the dialect consonant and rhyme coordination rules using a language collection device and to construct a dialect speech sample;

[0090] A dialect template construction module 520 is used to construct a dialect speech template according to the dialect speech sample based on the dynamic time warping technology and the mean calculation method;

[0091] The sound category region positioning module 530 is used to obtain the dialect speech to be segmented; according to the dialect speech to be segmented, the general speech detection technology is used to locate the sound category region to obtain the speech in the obscuration region, the speech in the vowel region and the speech in the non-vowel region;

[0092] The first dialect segmentation module 540 is used to perform frame-by-frame matching based on the dialect speech template and the speech in the obstruent area, the vowel area and the non-vowel area to obtain a set of potential phoneme segments;

[0093] The second dialect segmentation module 550 is used to screen the phoneme segments based on the dialect consonant-rhythm matching rules and the potential phoneme segment set using the matching distance method to obtain the dialect speech segmentation result.

[0094] Optionally, the dialect sample construction module 510 is further used to:

[0095] Through the language collection equipment, the voice of common words is collected to obtain the dialect voice subset;

[0096] Based on the dialect consonant-rhyme coordination rules and according to the dialect speech subset, language analysis tools are used to annotate and obtain dialect speech samples.

[0097] Among them, the dialect phonetic and rhyme coordination rules include dialect initials, dialect finals, dialect tones and dialect syllables.

[0098] Optionally, the dialect template construction module 520 is further used to:

[0099] Performing frame processing on the dialect speech sample to obtain a framed speech sample;

[0100] Based on the short-time average energy and Mel frequency cepstral coefficients, the voice sample features are obtained by calculating the mean of the voices with the same content in the framed voice samples;

[0101] Based on dynamic time warping technology, the matching distance method is used to screen samples according to the framed speech samples to obtain the optimal dialect samples;

[0102] A dialect speech template is constructed according to the speech sample features and the frame number of the optimal dialect sample.

[0103] Optionally, the sound category region positioning module 530 is further configured to:

[0104] According to the dialect speech to be segmented, the speech activity detection technology is used to locate the effective sound area and obtain the speech with sound area;

[0105] According to the voice in the sound area, the sonorous sound area detection technology is used to segment the sound area to obtain the speech in the sonorous sound area and the speech in the obstructive sound area;

[0106] According to the sonorous area speech, the vowel area detection technology is used to segment the sound area to obtain the vowel area speech and the non-vowel area speech.

[0107] Optionally, the first dialect segmentation module 540 is further configured to:

[0108] Based on the dialect speech template, frame processing is performed according to the speech in the obstruent area, the speech in the vowel area and the speech in the non-vowel area, and a frame feature vector matrix is ​​constructed;

[0109] Based on the preset number of frames, according to the frame feature vector matrix and the dialect speech template, the distance matching calculation is performed frame by frame through the dynamic time warping technology to obtain a frame-by-frame matching result set;

[0110] Perform statistical calculations based on the frame-by-frame matching result set to obtain a matching distance threshold;

[0111] Based on the matching distance threshold, the phoneme boundary area is judged for the speech in the obscurant area, the speech in the vowel area and the speech in the non-vowel area to obtain a set of potential phoneme segments.

[0112] Optionally, the second dialect segmentation module 550 is further configured to:

[0113] Based on the dialect phonetic and rhyme coordination rules, the dialect syllable rules are screened according to the potential phoneme segment set to obtain the segment set that conforms to the syllable rule and the corresponding segment syllable set;

[0114] Based on the dialect phoneme matching rules, according to the set of segments that conform to the syllable rules, the matching distance method is used to screen the dialect phoneme rules to obtain the optimal speech segmentation segment;

[0115] Obtaining a syllable segmentation result according to the optimal speech segmentation segment and the corresponding segment syllable set;

[0116] Based on the dialect consonant-resonance coordination rules, the phoneme region matching is performed according to the optimal speech segmentation segment to obtain the initial consonant segmentation result and the final consonant segmentation result;

[0117] According to the syllable segmentation results, the initial consonant segmentation results and the final vowel segmentation results, the dialect speech segmentation results are obtained.

[0118] The present invention proposes a method for automatic segmentation of Chinese dialect speech. The initial and final templates are constructed by a small amount of annotated and processed speech data, combined with the detection results of the general sound category, and the continuous Chinese dialect speech signal is automatically segmented into speech segments of initials, finals and syllables by a template matching method. The efficiency and accuracy of automatic speech segmentation are improved while saving training costs and time costs. The method is suitable for promotion to the automatic speech segmentation task of any Chinese dialect, and the field and scope of application of the method are improved. The present invention is a method for automatic segmentation of Chinese dialect speech based on a small amount of annotated and processed speech data, which is targeted at the characteristics of Chinese dialects.

[0119] Figure 6 is a schematic diagram of the structure of a Chinese dialect speech automatic segmentation device provided by an embodiment of the present invention, such as Figure 6 As shown, the Chinese dialect speech automatic segmentation device may include the above Figure 5 Optionally, the Chinese dialect speech automatic segmentation device 610 may include a first processor 2001 .

[0120] Optionally, the Chinese dialect speech automatic segmentation device 610 may further include a memory 2002 and a transceiver 2003 .

[0121] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0122] Combine the following Figure 6 The various components of the Chinese dialect speech automatic segmentation device 610 are specifically introduced:

[0123] The first processor 2001 is the control center of the Chinese dialect speech automatic segmentation device 610, which can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).

[0124] Optionally, the first processor 2001 can perform various functions of the Chinese dialect speech automatic segmentation device 610 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0125] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 are shown in FIG.

[0126] In a specific implementation, as an embodiment, the Chinese dialect speech automatic segmentation device 610 may also include multiple processors, such as Figure 6 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0127] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0128] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0129] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0130] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0131] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0132] It should be noted that Figure 6 The structure of the Chinese dialect speech automatic segmentation device 610 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0133] In addition, the technical effects of the Chinese dialect speech automatic segmentation device 610 can refer to the technical effects of the Chinese dialect speech automatic segmentation method described in the above method embodiment, and will not be repeated here.

[0134] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0135] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0136] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0137] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0138] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0139] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0140] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0141] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0142] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0143] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0145] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0146] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A method for automatic segmentation of Chinese dialect speech, characterized in that: The method comprises: Based on the dialect consonant-rhyme coordination rules, use language acquisition equipment to collect dialect speech and construct dialect speech samples; Based on dynamic time warping technology and mean calculation method, a dialect speech template is constructed according to the dialect speech samples; Acquire the dialect speech to be segmented; use the general speech detection technology to locate the sound area according to the dialect speech to be segmented, and obtain the speech in the obstructed sound area, the speech in the vowel area, and the speech in the non-vowel area; Based on the dialect speech template, frame-by-frame matching is performed according to the speech in the obstruent area, the speech in the vowel area, and the speech in the non-vowel area to obtain a set of potential phoneme segments; Based on the dialect consonant-rhyme coordination rules and the potential phoneme segment set, the matching distance method is used to screen the phoneme segments to obtain the dialect speech segmentation results.

2. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The method of collecting dialect speech based on the dialect consonant-rhyme coordination rules using a language collection device and constructing a dialect speech sample includes: Through the language collection equipment, the voice of common words is collected to obtain the dialect voice subset; Based on the dialect consonant-rhyme coordination rules and according to the dialect speech subset, language analysis tools are used to annotate and obtain dialect speech samples.

3. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The dialect phonetic-phonological coordination rules include dialect initials, dialect finals, dialect tones and dialect syllables.

4. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The method of constructing a dialect speech template based on a dialect speech sample based on a dynamic time warping technology and a mean calculation method includes: Performing frame processing on the dialect speech sample to obtain a framed speech sample; Based on the short-time average energy and Mel frequency cepstral coefficients, the average value of the speech with the same content in the framed speech sample is calculated to obtain the speech sample features; Based on dynamic time warping technology, the matching distance method is used to screen samples according to the framed speech samples to obtain the optimal dialect samples; A dialect speech template is constructed according to the speech sample features and the frame number of the optimal dialect sample.

5. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The method uses a general speech detection technology to locate the sound area according to the dialect speech to be segmented, and obtains the speech in the obstructed sound area, the speech in the vowel area, and the speech in the non-vowel area, including: According to the dialect speech to be segmented, the speech activity detection technology is used to locate the effective sound area and obtain the speech with sound area; According to the voice in the sound area, the sonorous sound area detection technology is used to segment the sound area to obtain the voice in the sonorous sound area and the voice in the obstructive sound area; According to the sonorous area speech, the vowel area detection technology is used to segment the sound area to obtain the vowel area speech and the non-vowel area speech.

6. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The method is based on the dialect speech template, and performs frame-by-frame matching according to the speech in the obstructed area, the speech in the vowel area, and the speech in the non-vowel area to obtain a set of potential phoneme segments, including: Based on the dialect speech template, frame processing is performed according to the speech in the obstruent area, the speech in the vowel area and the speech in the non-vowel area, and a frame feature vector matrix is ​​constructed; Based on the preset number of frames, according to the frame feature vector matrix and the dialect speech template, the distance matching calculation is performed frame by frame through the dynamic time warping technology to obtain a frame-by-frame matching result set; Perform statistical calculations based on the frame-by-frame matching result set to obtain a matching distance threshold; Based on the matching distance threshold, the phoneme boundary area is judged for the speech in the obscurant area, the speech in the vowel area and the speech in the non-vowel area to obtain a set of potential phoneme segments.

7. The method for automatic segmentation of Chinese dialect speech according to claim 1, characterized in that: The method of screening the phoneme segments based on the dialect consonant-rhyme matching rules and the potential phoneme segment set by using the matching distance method to obtain the dialect speech segmentation result includes: Based on the dialect phonetic and rhyme coordination rules, the dialect syllable rules are screened according to the potential phoneme segment set to obtain the segment set that conforms to the syllable rule and the corresponding segment syllable set; Based on the dialect phoneme matching rules, according to the set of segments that conform to the syllable rules, the matching distance method is used to screen the dialect phoneme rules to obtain the optimal speech segmentation segment; Obtaining a syllable segmentation result according to the optimal speech segmentation segment and the corresponding segment syllable set; Based on the dialect consonant-resonance coordination rules, the phoneme region matching is performed according to the optimal speech segmentation segment to obtain the initial consonant segmentation result and the final consonant segmentation result; According to the syllable segmentation results, the initial consonant segmentation results and the final vowel segmentation results, the dialect speech segmentation results are obtained.

8. A device for automatically segmenting Chinese dialect speech, the device for automatically segmenting Chinese dialect speech is used to implement the method for automatically segmenting Chinese dialect speech as claimed in any one of claims 1 to 7, characterized in that: The device comprises: A dialect sample construction module is used to collect dialect speech based on the dialect consonant-rhythm coordination rules using a language collection device and to construct a dialect speech sample; A dialect template construction module is used to construct a dialect speech template according to a dialect speech sample based on a dynamic time warping technology and a mean calculation method; The sound category area positioning module is used to obtain the dialect speech to be segmented; according to the dialect speech to be segmented, the general speech detection technology is used to locate the sound category area to obtain the speech in the obscuration area, the speech in the vowel area and the speech in the non-vowel area; The first dialect segmentation module is used to perform frame-by-frame matching based on the dialect speech template and the speech in the obstruent area, the vowel area and the non-vowel area to obtain a set of potential phoneme segments; The second dialect segmentation module is used to screen the phoneme segments based on the dialect consonant-rhythm coordination rules and the potential phoneme segment set using the matching distance method to obtain the dialect speech segmentation result.

9. A device for automatic segmentation of Chinese dialect speech, characterized in that: The Chinese dialect speech automatic segmentation device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.