Miao language speech adaptive segmentation method and system based on time domain features

Through the Miao language pronunciation adaptive segmentation method based on time domain features, the adaptive evaluation function is constructed using short-time energy and zero-crossing rate, and the boundary search is optimized by combining elite strategies and genetic algorithms, the Miao language pronunciation syllable boundary fuzzy problem is solved, and more accurate speech segmentation is achieved.

CN120279889APending Publication Date: 2025-07-08GUIZHOU MINZU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311122252.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有技术难以有效解决苗语语音音节边界模糊的问题,导致语音分割不准确。

Method used

Adaptive segmentation method of Miao language speech based on time domain features is adopted to extract speech frames through short-time energy and short-time zero-crossing rate, build a fitness evaluation function model, perform adaptive segmentation, and optimize boundary search using elite strategies and genetic algorithms.

Benefits of technology

The adaptive boundary search capability of phonological syllables is significantly improved, the precise boundary between phonological syllables and silent segments is extracted, the syllable boundary blur problem is solved, and the accuracy and recall rate of segmentation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279889A_ABST
    Figure CN120279889A_ABST
Patent Text Reader

Abstract

The invention provides a Miao language speech adaptive segmentation method and system based on time domain features, and relates to the technical field of speech segmentation. The method comprises the following steps: pre-recording voice audio of a single channel, and preprocessing the voice audio to obtain short-time energy and a short-time zero-crossing rate; the syllables are extracted through the short-time energy and the short-time zero-crossing rate, and voice syllables are obtained; calculating the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculating the short-time zero-crossing rate to obtain a minimum zero-crossing rate; pre-segmenting the voice frame through the maximum short-time energy, the minimum short-time energy and the minimum zero-crossing rate to obtain a plurality of pre-segmented voice segments; and constructing a fitness evaluation function model, and performing adaptive segmentation on the plurality of pre-segmented voice segments through the fitness evaluation function model to obtain a plurality of optimal voice segments. A syllable boundary fuzzy problem is converted into an actual relation problem between a syllable real boundary and a predicted boundary, so that the speech syllable adaptive boundary search capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of speech segmentation, and particularly relates to a method and system for adaptive segmentation of Miao language speech based on time-domain features. Background Art

[0002] Speech segmentation is a speech signal preprocessing technology for extracting speech units (such as syllables or phonemes) by removing silent segments (such as silence or noise). At present, certain research results have been achieved in Chinese, English, and some ethnic minority languages (such as Tibetan or Uyghur). However, as a low-resource and non-written ethnic minority language in the southwestern region of China, Miao language has problems such as an imperfect speech corpus, regional differences, and the lack of a written language, which pose great challenges to Miao language speech recognition. Although there is a certain foundation for the current research on Miao language speech, it mainly focuses on the recognition of isolated Miao language words. Relatively speaking, there is less research on the key issue of fuzzy processing of syllable boundaries in Miao language speech segmentation. With the increasing demand for this technology in the inheritance and protection of ethnic cultures such as Miao language, the research on speech segmentation methods has become a hot and difficult point in ethnic speech research in recent years, attracting the attention of scholars at home and abroad.

[0003] In summary, in terms of speech segmentation based on boundary detection, although some auxiliary methods are used to find the speech boundary and better speech units can be segmented, it is difficult to find the optimal segmentation point using the characteristics of the speech itself in the face of the problem of fuzzy syllable boundaries in Miao language speech. Therefore, a boundary fuzzy automatic search method based on the optimization idea is designed to solve the problem of Miao language speech syllable segmentation. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for adaptive segmentation of Miao language speech based on time-domain features in view of the deficiencies of the prior art.

[0005] The technical solution of the present invention for solving the above technical problem is as follows:

[0006] A method for adaptive segmentation of Miao language speech based on time-domain features includes the following steps:

[0007] Pre-record a single-channel speech audio, extract the speech frame length from the speech audio according to a preset length to obtain speech frames, and preprocess the speech frames to obtain short-time energy and short-time zero-crossing rate;

[0008] Extract syllables of the speech frames through the short-time energy and the short-time zero-crossing rate to obtain speech syllables;

[0009] Calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate;

[0010] Pre-segment the speech frame by the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain a plurality of pre-segmented speech segments;

[0011] Construct a fitness evaluation function model, and adaptively segment the plurality of pre-segmented speech segments through the fitness evaluation function model to obtain a plurality of optimal speech segments.

[0012] The beneficial effects of the present invention are as follows: Aiming at the problem of speech syllable segmentation, the initial speech boundary is obtained by time-domain feature segmentation. By constructing a fitness evaluation function model for the speech syllable boundary, the problem of blurred boundary between syllables and silent segments is transformed into the actual relationship problem between the true boundary and the predicted boundary of syllables, and the accurate boundary between speech syllables and silent segments is extracted, significantly improving the ability of adaptive boundary search for speech syllables.

[0013] Further, the preprocessing of the speech frame to obtain the short-time energy and the short-time zero-crossing rate is specifically as follows:

[0014] Perform pre-emphasis processing on the speech frame through a high-pass digital filter to obtain pre-emphasized speech,

[0015] The pre-emphasis processing is:

[0016] H(Z) = 1 - αZ -1 ,

[0017] where H(Z) is the transfer function of the high-pass digital filter, Z is the time interval of signal sampling, α is the pre-emphasis coefficient, and 0.9 < α < 1.0;

[0018] Perform frame windowing processing on the pre-emphasized speech to obtain the short-time energy and the short-time zero-crossing rate;

[0019] The short-time energy is:

[0020]

[0021] where E m is the short-time energy, x(n) is the pre-emphasized speech, and is the discrete speech audio signal sampling value at time n, ω(m) is the window function with length m, m is the time position of the short-time energy on the speech frame, and m = 1, 2, 3,..., M, where M is the length of the speech frame;

[0022] The short-time zero-crossing rate is:

[0023]

[0024] where Z mLet \(ZCR\) be the short-time zero-crossing rate, \(x(n)\) be the pre-emphasized speech, which is the sampled value of the discrete speech audio signal at time \(n\), \(\omega(m)\) be the window function of length \(m\), \(m\) be the time position of the short-time zero-crossing rate on the speech frame, \(m = 1, 2, 3,\cdots,M\), where \(M\) is the length of the speech frame, and \(sgn(x(m))\) be the sign function:

[0025]

[0026] The so-called short-time energy (STE) refers to the energy magnitude of the audio signal within a certain period of time.

[0027] Performing frame windowing on the pre-emphasized speech also results in a stationary speech time series;

[0028] If adjacent samples in the stationary speech time series have different algebraic signs, then a zero-crossing occurs. Therefore, the number of zero-crossings is calculated, and the number of zero-crossings per unit time is called the zero-crossing rate.

[0029] The so-called short-time zero-crossing rate (STZCR) refers to the number of times the signal passes through the zero value within each frame.

[0030] The beneficial effects of adopting the above further scheme are as follows: The preprocessing is to perform pre-emphasis processing on the speech, enhance the high-frequency part of the speech, remove the influence of lip radiation, increase the high-frequency resolution of the speech, and make the speech frequency smoother; while the speech is a non-stationary sequence. In order to obtain stationary speech features, the non-stationary Miao speech is frame windowed to obtain a stationary Miao speech time series, and the short-time energy and short-time zero-crossing rate generated during preprocessing and frame windowing are obtained as features for speech segmentation.

[0031] Further, extracting the syllables of the speech frame through the short-time energy and the short-time zero-crossing rate to obtain speech syllables specifically includes:

[0032] Extracting the vowels and consonants of the syllables of the speech frame through the short-time energy to obtain the vowels and the consonants; extracting the voiceless and voiced sounds of the syllables of the speech frame through the short-time zero-crossing rate to obtain the voiceless and voiced sounds;

[0033] Inputting the vowels, the consonants, the voiceless sounds, and the voiced sounds into a preset syllable list to obtain speech syllables.

[0034] The syllable list is used to save the position nodes of the pre-segmented positions of the speech frame.

[0035] The beneficial effects of adopting the above further scheme are as follows: Miao language speech is monosyllabic, with high energy in vowel syllables, low energy in consonant syllables, and it is easy to distinguish between high and low energy in Miao language syllables. Short-time energy can distinguish between consonants and vowels in speech; in Miao language speech, the number of zero-crossing rates of voiced sounds is large, and the number of zero-crossing rates of voiceless sounds is small. The short-time zero-crossing rate can well distinguish between voiced and voiceless sounds.

[0036] Further, calculating the short-time energy to obtain the maximum short-time energy and the minimum short-time energy specifically includes:

[0037] Calculating the short-time energy to obtain the average short-time energy;

[0038] Calculating the maximum short-time energy through the average short-time energy, and the maximum short-time energy is:

[0039]

[0040] where M H is the maximum short-time energy, is the average short-time energy, μ is a parameter, and 0 < μ ≤ 1;

[0041] Calculating the minimum short-time energy through the maximum short-time energy and the short-time energy value, and the minimum short-time energy is:

[0042]

[0043] where M L is the minimum short-time energy, is the short-time energy value of the first 5 frames, θ is a parameter, and 0 < θ ≤ 1.

[0044] The beneficial effects of adopting the above further scheme are as follows: obtaining the optimal value through multiple calculations Calculating the minimum short-time energy and obtaining the minimum short-time energy is convenient for speech pre-segmentation.

[0045] Further, calculating the short-time zero-crossing rate to obtain the minimum zero-crossing rate specifically includes:

[0046] Calculating the short-time zero-crossing rate to obtain the average short-time zero-crossing rate;

[0047] Calculating the minimum zero-crossing rate through the average short-time zero-crossing rate, and the minimum zero-crossing rate is:

[0048]

[0049] where Z S is the minimum zero-crossing rate, is the average short-time zero-crossing rate, β is a parameter, and 0 < β ≤ 1.

[0050] The beneficial effect of adopting the above further scheme is that calculating the minimum zero-crossing rate facilitates voice pre-segmentation.

[0051] Further, the voice frames are pre-segmented by the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain a plurality of pre-segmented voice segments. Specifically:

[0052] Set the noise and the minimum unvoiced interval according to the voice syllables to obtain the number of noise frames and the number of minimum unvoiced intervals;

[0053] Pre-segment the voice frames by the number of noise frames, the number of minimum unvoiced intervals, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain a plurality of pre-segmented voice segments.

[0054] The beneficial effect of adopting the above further scheme is that pre-segmenting the voice in the time-domain features facilitates subsequent adaptive segmentation of the pre-segmented voice segments.

[0055] Further, the construction of the fitness evaluation function model is specifically as follows:

[0056] Let S be the true boundary segment and P be the predicted boundary segment. The fitness evaluation function model is as follows:

[0057]

[0058] s.t.S i ∈ R, P i ∈ R,

[0059] i = 1, 2,..., M,

[0060] where f(x) is the objective function of the true segmentation and the predicted segmentation, S i is the true boundary point of the i-th frame, P i is the predicted boundary point of the i-th frame, M is the length of the voice frame, i is the number of boundary frames, R is the set of real numbers, and s.t. is the abbreviation of subject to, indicating the constraint condition.

[0061] The beneficial effect of adopting the above further scheme is that two parameters are added to limit the noise and judge the interval of unvoiced sounds; since the time-domain feature segmentation has limited ability to search for syllables, a relationship model between the true boundary and the predicted boundary of the voice syllables is constructed, and this relationship model is used as the fitness evaluation function for adaptive segmentation to solve the problem of fuzzy boundaries of voice syllables, thereby realizing the adaptive segmentation of voice syllables.

[0062] Further, adaptively segmenting the multiple pre-segmented speech segments through the fitness evaluation function model to obtain multiple optimal speech segments specifically includes:

[0063] Randomly generating an initial population for the multiple pre-segmented speech segments through binary coding, where the initial population includes multiple individuals; and performing fitness calculation, selection operation, elite strategy, crossover operation, mutation operation, and evaluation operation on the initial population;

[0064] The fitness calculation includes: performing the fitness calculation on multiple individuals in the initial population through the fitness evaluation function model to obtain fitness values;

[0065] The selection operation includes: selecting the multiple individuals according to the fitness values, retaining multiple excellent individuals through the elite strategy, and generating a first-generation population from multiple inferior individuals;

[0066] The crossover operation includes: performing two-point crossover on multiple inferior individuals in the next-generation population to obtain a diverse population;

[0067] The mutation operation includes: mutating multiple inferior individuals in the diverse population with a preset probability to obtain a mutated population;

[0068] Combining multiple excellent individuals in the elite strategy and multiple inferior individuals in the mutated population to generate a new population;

[0069] The evaluation operation includes: performing target judgment on multiple individuals in the new population through the fitness evaluation function model,

[0070] When the minimum objective function is not obtained, repeat the fitness calculation, selection operation, elite strategy, crossover operation, mutation operation, and evaluation operation on multiple inferior individuals in the new population;

[0071] When the minimum objective function is obtained, output the new population, obtain two optimization parameters through the new population, set noise and the minimum voiceless interval through the speech syllables and optimization parameters, obtain the optimized noise frame number and the minimum optimized voiceless interval number; perform adaptive segmentation on the speech audio through the optimized noise frame number, the minimum optimized voiceless interval number, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple optimal speech segments.

[0072] The adaptive model includes three steps: the crossover operation, the mutation operation, and the elite strategy. New populations are generated through the crossover and mutation operations. The individuals of the new populations are used to evaluate the fuzzy boundaries of speech syllables, and the individuals of the populations that meet the evaluation optimization function are selected. The elite strategy is used to retain the two best individuals without performing the crossover and mutation operations again, so as to find the population individuals with the best speech syllable segmentation boundaries, thereby realizing the adaptive segmentation of Miao speech syllables until the iteration stops.

[0073] The beneficial effects of adopting the above further scheme are as follows: Through continuous iteration, the predicted segmentation boundary can be closer to the real segmentation boundary. Using the elite strategy as the population retention strategy, the best speech syllable segmentation boundary is obtained.

[0074] Further, after obtaining multiple speech segments, it also includes the step of calculating the accuracy rate, recall rate, and harmonic mean of the adaptive segmentation of the speech, specifically:

[0075] Calculating the accuracy rate of the adaptive segmentation of the speech includes:

[0076] The accuracy rate is calculated by the number of correctly segmented speech points and the total number of segmented speech points, and the accuracy rate is:

[0077]

[0078] Among them, PRC is the accuracy rate, d is the number of correctly segmented speech points, l is the total number of segmented speech points, and · represents multiplication;

[0079] Calculating the recall rate of the adaptive segmentation of the speech includes:

[0080] The recall rate is calculated by the number of correctly segmented speech points and the number of segmented speech points, and the recall rate is:

[0081]

[0082] Among them, RCL is the recall rate, d is the number of correctly segmented speech points, h is the number of segmented speech points, and · represents multiplication;

[0083] Calculating the harmonic mean of the adaptive segmentation of the speech includes:

[0084] The harmonic mean is calculated by the accuracy rate and the recall rate, and the harmonic mean is:

[0085]

[0086] Among them, F is the harmonic mean, PRC is the accuracy rate, RCL is the recall rate, and · represents multiplication.

[0087] The beneficial effect of adopting the above further solution is: calculate the harmonic mean to test the feasibility of this solution.

[0088] Another technical solution for the present invention to solve the above technical problems is as follows:

[0089] A Miao language speech adaptive segmentation system based on time-domain features, comprising: a preprocessing module, a time-domain feature calculation module, a speech pre-segmentation module, and an adaptive speech segmentation module;

[0090] The preprocessing module is used to pre-record a single-channel speech audio, extract the speech frame length from the speech audio according to a preset length to obtain speech frames, and preprocess the speech frames to obtain short-time energy and short-time zero-crossing rate;

[0091] The time-domain feature calculation module is used to extract syllables of the speech frames through the short-time energy and the short-time zero-crossing rate to obtain speech syllables; calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate;

[0092] The speech pre-segmentation module is used to pre-segment the speech frames through the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain a plurality of pre-segmented speech segments;

[0093] The adaptive speech segmentation module is used to construct a fitness evaluation function model, and adaptively segment the plurality of pre-segmented speech segments through the fitness evaluation function model to obtain a plurality of optimal speech segments.

[0094] The beneficial effect of the present invention is: aiming at the problem of speech syllable segmentation, obtaining an initial speech boundary by time-domain feature segmentation, and by constructing a fitness evaluation function model for the speech syllable boundary, transforming the problem of blurred boundary between syllables and silent segments into the actual relationship problem between the true boundary and the predicted boundary of the syllables, extracting the precise boundary between speech syllables and silent segments, and significantly improving the adaptive boundary search ability of speech syllables. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 It is a flowchart of a Miao language speech adaptive segmentation method based on time-domain features provided by an embodiment of the present invention;

[0096] Figure 2 It is a module block diagram of a Miao language speech adaptive segmentation system based on time-domain features provided by an embodiment of the present invention;

[0097] Figure 3 It is a structural diagram of a Miao language speech example diagram including silent segments provided by an embodiment of the present invention;

[0098] Figure 4 It is the structural diagram of the 60ms Miao language speech time-domain feature segmentation map provided by the embodiment of the present invention;

[0099] Figure 5 It is the structural diagram of the 60ms Miao language speech adaptive segmentation map provided by the embodiment of the present invention. Specific implementation manner

[0100] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0101] The following embodiments take the Miao language as the research object.

[0102] Speech segmentation is to find the boundaries of speech syllables and solve the problem of the boundary positions between speech syllables and silent segments. However, since the syllable part may be misjudged as the silent segment part or the silent segment part is misjudged as the syllable part, over-segmentation and missed segmentation of syllables occur, making it difficult to determine the syllable boundaries. As Figure 3 shown, the part framed by the square is the speech segment, the long side line of the square is the speech syllable boundary, and the part pointed by the arrow is the silent segment (i.e., silence or noise).

[0103] As Figure 1 shown, a Miao language speech adaptive segmentation method based on time-domain features includes the following steps:

[0104] Pre-record a single-channel speech audio, extract the speech frame length from the speech audio according to a preset length to obtain speech frames, and preprocess the speech frames to obtain short-time energy and short-time zero-crossing rate;

[0105] Extract the syllables of the speech frames through the short-time energy and the short-time zero-crossing rate to obtain speech syllables;

[0106] Calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate;

[0107] Pre-segment the speech frames through the maximum short-time energy, the minimum short-time energy and the minimum zero-crossing rate to obtain a plurality of pre-segmented speech segments;

[0108] Construct a fitness evaluation function model, and adaptively segment the plurality of pre-segmented speech segments through the fitness evaluation function model to obtain a plurality of optimal speech segments.

[0109] Specifically, manually record Miao language speech. The recorded speech format is single-channel, the sampling frequency is 44100HZ, and it is saved as a lossless speech in WAV format.

[0110] The beneficial effects of the present invention are as follows: Regarding the problem of speech syllable segmentation, the initial speech boundaries are obtained by segmenting time-domain features. By constructing a fitness evaluation function model for speech syllable boundaries, the problem of blurred boundaries between syllables and silent segments is transformed into the actual relationship problem between the true boundaries and predicted boundaries of syllables, and the precise boundaries between speech syllables and silent segments are extracted, significantly improving the ability to adaptively search for speech syllable boundaries.

[0111] Preferably, the preprocessing of the speech frame to obtain the short-time energy and short-time zero-crossing rate is specifically as follows:

[0112] The speech frame is pre-emphasized by a high-pass digital filter to obtain pre-emphasized speech.

[0113] The pre-emphasis process is as follows:

[0114] H(Z) = 1 - αZ -1 ,

[0115] where H(Z) is the transfer function of the high-pass digital filter, Z is the time interval of signal sampling, α is the pre-emphasis coefficient, and 0.9 < α < 1.0;

[0116] The pre-emphasized speech is framed and windowed to obtain the short-time energy and short-time zero-crossing rate;

[0117] The short-time energy is as follows:

[0118]

[0119] where E m is the short-time energy, x(n) is the pre-emphasized speech, ω(m) is the window function with length m, m is the time position of the short-time energy on the speech frame, and m = 1, 2, 3, …, M, where M is the length of the speech frame;

[0120] The short-time zero-crossing rate is as follows:

[0121]

[0122] where Z m is the short-time zero-crossing rate, x(n) is the pre-emphasized speech, ω(m) is the window function with length m, m is the time position of the short-time zero-crossing rate on the speech frame, m = 1, 2, 3, …, M, where M is the length of the speech frame, and sgn(x(m)) is the sign function:

[0123]

[0124] Specifically, the speech is pre-emphasized by a high-pass digital filter with a first-order transfer function; a Hamming window is used for windowing, and the frame length is set to 25 ms and the frame shift is set to 10 ms.

[0125] The short-time energy (STE) refers to the energy magnitude of an audio signal within a certain period of time.

[0126] For the stationary speech time series, if adjacent samples have different algebraic signs, it is called a zero crossing. Therefore, the number of zero crossings is calculated, and the number of zero crossings per unit time is called the zero-crossing rate.

[0127] The short-time zero-crossing rate (STZCR) refers to the number of times the signal passes through the zero value within each frame.

[0128] In the above embodiments, the preprocessing is to perform pre-emphasis processing on the speech, emphasize the high-frequency part of the speech, remove the influence of lip radiation, increase the high-frequency resolution of the speech, and make the speech frequency smoother; and the speech is a non-stationary sequence. In order to obtain stationary speech features, the non-stationary Miao language speech is framed and windowed to obtain a stationary Miao language speech time series, and the short-time energy and short-time zero-crossing rate generated during preprocessing and framing and windowing are obtained as features for speech segmentation.

[0129] Preferably, the syllables of the speech frame are extracted through the short-time energy and the short-time zero-crossing rate to obtain speech syllables, specifically:

[0130] The vowels and consonants of the syllables of the speech frame are extracted through the short-time energy to obtain the vowels and the consonants; the voiceless and voiced sounds of the syllables of the speech frame are extracted through the short-time zero-crossing rate to obtain the voiceless and voiced sounds;

[0131] The vowels, the consonants, the voiceless and the voiced sounds are input into a preset syllable list to obtain speech syllables.

[0132] In the above embodiments, the Miao language speech is monosyllabic, the vowel syllables have high energy, the consonant syllables have low energy, and it is easy to distinguish between high and low energy of Miao language syllables, and the short-time energy can distinguish the consonants and vowels of the speech; in the Miao language speech, the number of zero-crossing rates of voiced sounds is large, the number of zero-crossing rates of voiceless sounds is small, and the short-time zero-crossing rate can well distinguish between voiced and voiceless sounds.

[0133] Preferably, the short-time energy is calculated to obtain the maximum short-time energy and the minimum short-time energy, specifically:

[0134] The short-time energy is calculated to obtain the average short-time energy;

[0135] The maximum short-time energy is calculated through the average short-time energy, and the maximum short-time energy is:

[0136]

[0137] Among them, M H is the maximum short-time energy, is the average short-time energy, μ is a parameter, and 0 < μ ≤ 1;

[0138] The minimum short-time energy is calculated through the maximum short-time energy and the short-time energy value, and the minimum short-time energy is:

[0139]

[0140] Among them, M L is the minimum short-time energy, is the short-time energy value of the first 5 frames, θ is a parameter, and 0 < θ ≤ 1.

[0141] In the above embodiments, the optimal value is obtained through multiple calculations The minimum short-time energy is calculated, and obtaining the minimum short-time energy facilitates voice pre-segmentation.

[0142] Preferably, the short-time zero-crossing rate is calculated to obtain the minimum zero-crossing rate, specifically:

[0143] The short-time zero-crossing rate is calculated to obtain the average short-time zero-crossing rate;

[0144] The minimum zero-crossing rate is calculated through the average short-time zero-crossing rate, and the minimum zero-crossing rate is:

[0145]

[0146] Among them, Z S is the minimum zero-crossing rate, is the average short-time zero-crossing rate, β is a parameter, and 0 < β ≤ 1.

[0147] In the above embodiments, obtaining the minimum zero-crossing rate through calculation facilitates voice pre-segmentation.

[0148] Preferably, the speech frames are pre-segmented through the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple pre-segmented speech segments, specifically:

[0149] According to the speech syllables, noise and the minimum unvoiced interval are set to obtain the number of noise frames and the number of minimum unvoiced intervals;

[0150] The speech frames are pre-segmented through the number of noise frames, the number of minimum unvoiced intervals, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple pre-segmented speech segments.

[0151] Specifically, the speech is initially segmented to obtain d1, d2, … d p pre-segmented speech segments.

[0152] In the above embodiments, the speech is pre-segmented in the time-domain features to facilitate subsequent adaptive segmentation of the pre-segmented speech segments.

[0153] Preferably, the construction of the fitness evaluation function model is specifically as follows:

[0154] The construction of the fitness evaluation function model is specifically as follows:

[0155] Let S be the true boundary segment and P be the predicted boundary segment. The fitness evaluation function model is as follows:

[0156]

[0157] s.t. S i ∈R, P i ∈R,

[0158] i = 1, 2, …, M,

[0159] where f(x) is the objective function of the true segmentation and the predicted segmentation, S i is the true boundary point of the i-th frame, P i is the predicted boundary point of the i-th frame, M is the length of the speech frame, i is the number of boundary frames, R is the set of real numbers; s.t. is the abbreviation of subject to, indicating the constraint condition.

[0160] In the above embodiments, in order to limit noise and judge the minimum unvoiced interval, two parameters σ and δ are respectively added in the time-domain feature segmentation to limit noise and judge the interval of unvoiced sounds. Since the time-domain feature segmentation has limited ability to search for syllables, a relationship model between the true boundary and the predicted boundary of the speech syllables is constructed, and this relationship model is used as the fitness evaluation function for adaptive segmentation to solve the problem of fuzzy boundaries of speech syllables, thereby realizing the adaptive segmentation of speech syllables.

[0161] Preferably, the adaptive segmentation of the multiple pre-segmented speech segments through the fitness evaluation function model to obtain multiple optimal speech segments is specifically as follows:

[0162] An initial population is randomly generated for the multiple pre-segmented speech segments through binary coding. The initial population includes multiple individuals; and fitness calculation, selection operation, elite strategy, crossover operation, mutation operation and evaluation operation are performed on the initial population;

[0163] The fitness calculation includes: performing the fitness calculation on multiple individuals in the initial population through the fitness evaluation function model to obtain fitness values;

[0164] The selection operation includes: selecting the multiple individuals according to the fitness value, retaining multiple excellent individuals through the elite strategy, and generating the first-generation population from multiple inferior individuals;

[0165] The crossover operation includes: performing two-point crossover on multiple inferior individuals in the next-generation population to obtain a diverse population;

[0166] The mutation operation includes: mutating multiple inferior individuals in the diverse population with a preset probability to obtain a mutated population;

[0167] Combining multiple excellent individuals in the elite strategy and multiple inferior individuals in the mutated population to generate a new population;

[0168] The evaluation operation includes: making a target judgment on multiple individuals in the new population through the fitness evaluation function model,

[0169] When the minimum objective function is not obtained, repeat the fitness calculation, the selection operation, the elite strategy, the crossover operation, the mutation operation, and the evaluation operation for multiple inferior individuals in the new population;

[0170] When the minimum objective function is obtained, output the new population, obtain two optimization parameters through the new population, set the noise and the minimum voiceless interval through the speech syllables and the optimization parameters to obtain the optimized noise frames and the minimum optimized voiceless interval number; adaptively segment the speech audio through the optimized noise frames, the minimum optimized voiceless interval number, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple optimal speech segments.

[0171] Specifically, randomly generate an initial population x for the multiple pre-segmented speech segments through binary coding t ={x1, x2,..., x N}, t = 0;

[0172] Selection operation, calculate the fitness value for the individuals in x t through the accuracy formula, perform the selection operation according to the fitness value, and generate the next-generation population x t+1 ; if f1(x t ) < f1(x t '), then x t+1 = x t , otherwise, x t+1 = x t ';

[0173] Crossover operation, for x tIndividuals in [[]] perform two-point crossover to increase population diversity;

[0174] Mutation operation, mutate individuals in x t at the set probability;

[0175] Elitist strategy, directly copy individuals with very good evaluations in x t to the next-generation population without performing crossover and mutation operations;

[0176] Evaluate individuals in the offspring population x t+1 ;

[0177] t = t + 1.

[0178] The elitist strategy is to better solve the problem of blurred boundaries between syllables and silent segments in speech segmentation, and keep the size of the optimal population unchanged. Suppose that by the t-th generation, x t in the population is the optimal individual. Also suppose that x t+1 is the new generation population. If there is no individual in x t+1 better than x t , then add x t to x t+1 as the k-th individual in x t+1 , where k is the sequence of the population. If the elite individual is added to the new generation population, then the individual with the smallest fitness value in the new generation population can be eliminated.

[0179] The adaptive model ASA includes the above three steps of crossover operation, mutation operation and elitist strategy. Generate a new population through crossover and mutation operations, evaluate the fuzzy boundaries of speech syllables for the individuals in the new population, and select the population individuals that meet the evaluation optimization function. Use the elitist strategy to retain the best 2 individuals without performing crossover and mutation operations anymore, so as to find the population individuals with the best speech syllable segmentation boundary, thereby realizing adaptive Miao speech syllable segmentation until the iteration stops.

[0180] In the above embodiment, through continuous iteration, the predicted segmentation boundary can be closer to the true segmentation boundary. Using the elitist strategy as the population retention strategy, the best speech syllable segmentation boundary is obtained.

[0181] Preferably, after obtaining multiple speech segments, it further includes the step of calculating the accuracy rate, recall rate and harmonic mean of the speech adaptive segmentation, specifically:

[0182] Calculating the accuracy rate of the speech adaptive segmentation includes:

[0183] Calculate the accuracy rate through the number of correctly segmented speech points and the total number of speech segmentation points, and obtain the accuracy rate, where the accuracy rate is:

[0184]

[0185] Among them, PRC is the accuracy rate, d is the number of correctly segmented voice points, l is the total number of segmented voice points, and · represents multiplication;

[0186] Calculating the recall rate of the adaptive segmentation of the voice includes:

[0187] Calculating the recall rate through the number of correctly segmented voice points and the number of segmented voice points to obtain the recall rate, and the recall rate is:

[0188]

[0189] Among them, RCL is the recall rate, d is the number of correctly segmented voice points, h is the number of segmented voice points, and · represents multiplication;

[0190] Calculating the harmonic mean of the adaptive segmentation of the voice includes:

[0191] Calculating the harmonic mean through the accuracy rate and the recall rate to obtain the harmonic mean, and the harmonic mean is:

[0192]

[0193] Among them, F is the harmonic mean, PRC is the accuracy rate, RCL is the recall rate, and · represents multiplication.

[0194] In the above embodiment, calculating the harmonic mean verifies the feasibility of this solution.

[0195] Preferably, the calculation steps of the computational complexity of this solution are:

[0196] Assume that the total time required for time-domain feature segmentation is n, then the time for the first time is n - 1. Combining the elite strategy evolutionary optimization boundary step is mainly the evaluation function optimization time. Each time the time-domain feature is segmented, it needs to go through m evaluations. Then the time required for evaluation function optimization is That is, the total time T is Therefore, the computational complexity is O(nm). Since the optimization process is the evolutionary process of the population, the corresponding computational complexity is At the same time, constants do not affect the computational complexity, that is, the final computational complexity is O((nm) 2 )

[0197] As Figure 2 shown, a Miao language voice adaptive segmentation system based on time-domain features includes: a preprocessing module, a time-domain feature calculation module, a voice pre-segmentation module, and an adaptive voice segmentation module;

[0198] The preprocessing module is used to pre-record a single-channel voice audio, extract the voice frame length from the voice audio according to a preset length to obtain voice frames, and preprocess the voice frames to obtain short-time energy and short-time zero-crossing rate;

[0199] The time-domain feature calculation module is used to extract syllables of the voice frames through the short-time energy and the short-time zero-crossing rate to obtain voice syllables; calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate;

[0200] The voice pre-segmentation module is used to pre-segment the voice frames through the maximum short-time energy, the minimum short-time energy and the minimum zero-crossing rate to obtain multiple pre-segmented voice segments;

[0201] The adaptive voice segmentation module is used to construct a fitness evaluation function model and adaptively segment the multiple pre-segmented voice segments through the fitness evaluation function model to obtain multiple optimal voice segments.

[0202] This solution aims at the problem of voice syllable segmentation, obtains the initial voice boundaries through time-domain feature segmentation, constructs a fitness evaluation function model for the voice syllable boundaries, transforms the boundary fuzzy problem between syllables and silent segments into the actual relationship problem between the true boundaries and the predicted boundaries of syllables, extracts the precise boundaries between voice syllables and silent segments, and significantly improves the adaptive boundary search ability of voice syllables.

[0203] Verify the effectiveness of this solution through experiments:

[0204] To verify the effectiveness of the voice segmentation of this solution, this solution conducts experiments on Miao language data and TIMIT test data sets respectively. The TIMIT test set includes 24 speakers using standard American English, each with 8 sentences, a total of 192 sentences. It covers different voice situations and scenarios. To verify the effectiveness of the evolution combined with the elite strategy, three groups of experiments are conducted to verify the experimental results under the errors of 60ms, 40ms and 20ms respectively. In the time-domain feature segmentation, the pre-emphasis coefficient is α = 0.98, the noise frame number parameter σ = 2, and the minimum voiced interval frame number δ = 15. In order to retain the originality of the time-domain feature segmentation, the parameters μ, θ, and β in this solution are all set to 1. The time-domain feature segmentation results are shown in Table 1.

[0205] Miao language data Precision Recall F-value 60ms 39.92 35.26 36.50 40ms 21.30 18.54 19.18 20ms 14.62 11.86 12.77

[0206] Table 1 Time-domain feature segmentation results

[0207] As can be seen from Table 1, the F value of the error time-domain feature at 20 ms is 12.77%, the F value of the segmented error time-domain feature at 40 ms is 19.18%, and the F value of the segmented error time-domain feature at 60 ms is 36.50%. The time-domain feature segmentation did not achieve good results in the Miao language speech segmentation. It could not well solve the problem of the blurred boundary of Miao language speech, verifying that the time-domain feature segmentation could not adaptively segment the blurred boundary of Miao language speech syllables. To better solve the problem of the blurred boundary between Miao language speech syllables and silent segments, an experiment will be conducted on the adaptive model ASA for Miao language speech. The population number of the experiment is 6, the crossover probability is 0.4, the mutation probability is 0.1, the number of individuals in the previous generation retained by the elitist strategy is 2, and the maximum number of iterations is 20. The experimental results are shown in Table 2.

[0208] Miao language data Precision Recall F-value 60ms 91.48 93.54 92.12 40ms 76.24 81.61 78.27 20ms 62.66 68.77 65.05

[0209] Table 2 Adaptive Model Segmentation Results

[0210] As can be seen from Table 2, for the segmentation of Miao language speech syllables, the error precision at 60 ms is 91.48%, the recall rate is 93.54%, and the F value is 92.12%. This fully shows that the segmentation achieved good results, verifying that the adaptive model ASA can better find Miao language speech syllables, solved the problem that the time-domain feature segmentation could find the blurred boundary problem of Miao language speech syllables, and achieved the adaptive segmentation of Miao language speech. It shows that under the error of 60 ms, the adaptive model ASA can be used for the segmentation of Miao language speech syllables. Moreover, experiments with larger errors were also conducted on the time-domain feature segmentation. It was found that when the error was 600 ms, the F value of the time-domain feature segmentation could only be close to the segmentation result of the adaptive model ASA. However, when the error was 600 ms, the segmentation was significantly ineffective and would cause speech distortion, indicating that the adaptive model ASA can better find the speech segmentation boundary while keeping the speech unchanged. However, as the segmentation error decreases, the experimental effect also decreases. To better understand the problems existing in the model in syllable segmentation, a visualization result display was conducted to better analyze the problems still existing in the segmentation of Miao language speech syllables. The visualization results are as Figure 4 and Figure 5 shown.

[0211] It was found in the experiment that when the pronunciations of syllables are similar, the proposed scheme will misjudge two or three syllables as one syllable under this condition, resulting in misjudgment in the speech syllable segmentation. As the speech segmentation error decreases, there is a small frame number gap between the optimized segmentation point and the actual marked segmentation point, resulting in a decline in the segmentation performance.

[0212] TIMIT test set Precision Recall F-value 60ms 83.74 93.53 88.14 40ms 78.26 90.62 83.98 20ms 72.60 88.50 79.76

[0213] Table 3 TIMIT Test Set Segmentation Results

[0214] As can be seen from Table 3, at an error of 60 ms, the accuracy of the adaptive model ASA is 83.74%, the recall rate is 93.53%, and the F-value is 88.14%. The results show that the adaptive model ASA can also segment English speech. Through experiments, it is found that the proposed adaptive model ASA still maintains the speech segmentation performance when the error decreases. This is for English speech segmentation, and there is no problem of similar Miao language speech syllables. Compared with Miao language speech syllable segmentation, it can better find speech syllables and has a higher segmentation accuracy rate.

[0215] TIMIT test set Precision Recall F-value <![CDATA[SCPC 1 > 68.75 79.24 74.98 <![CDATA[SCPC 2 > 68.92 78.43 74.55 ASA 72.60 88.50 79.76

[0216] Table 4 Comparison Results of TIMIT Test Set

[0217] To verify the effectiveness of the model on other TIMIT data sets, the adaptive model ASA is used to segment English speech words on the TIMIT test set. The F-value is 88.14% at an error of 60 ms, 83.98% at an error of 40 ms, and 79.76% at an error of 20 ms, which proves that the model can effectively segment English speech. To verify the quality of the adaptive model ASA and the current popular segmentation models, under the condition of an error of 20 ms, this scheme is compared with the existing SCPC method on the TIMIT test set. The F-value of the adaptive model ASA is significantly improved by about 5% compared with the SCPC method, which proves that the model has certain advantages in speech segmentation.

[0218] Under the experimental error of 60 ms, it can effectively segment Miao language speech syllables, verifying that the adaptive model ASA can better find the speech segmentation boundary. Moreover, the proposed adaptive model ASA of this scheme is also verified for speech segmentation on the TIMIT test set. The results show that this scheme has a significant improvement compared with the existing speech segmentation methods.

[0219] For the above Miao language speech adaptive segmentation system based on time-domain features, reference can be made to the implementation content and its beneficial effects specifically described for a Miao language speech adaptive segmentation method based on time-domain features as above, which will not be elaborated here.

[0220] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0221] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adaptive segmentation method for Miao language speech based on time-domain features, characterized in that It includes the following steps: Pre-record a single-channel voice audio, extract the voice frame length from the voice audio according to a preset length to obtain voice frames, and preprocess the voice frames to obtain short-time energy and short-time zero-crossing rate; Extract syllables of the voice frames through the short-time energy and the short-time zero-crossing rate to obtain voice syllables; Calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate; Pre-segment the voice frames through the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple pre-segmented voice segments; Construct a fitness evaluation function model, and adaptively segment the multiple pre-segmented voice segments through the fitness evaluation function model to obtain multiple optimal voice segments.

2. The method for adaptively segmenting Miao language speech based on time domain features according to claim 1, wherein, The preprocessing of the voice frames to obtain short-time energy and short-time zero-crossing rate is specifically as follows: Perform pre-emphasis processing on the voice frames through a high-pass digital filter to obtain pre-emphasized voice, The pre-emphasis processing is: H(Z) = 1 - αZ -1 , Where H(Z) is the transfer function of the high-pass digital filter, Z is the time interval of signal sampling, α is the pre-emphasis coefficient, and 0.9 < α < 1.0; Perform frame-by-frame windowing processing on the pre-emphasized voice to obtain short-time energy and short-time zero-crossing rate; The short-time energy is: where E m is the short-time energy, x(n) is the pre-emphasized speech, ω(m) is a window function of length m, m is the time position of the short-time energy on the speech frame, and m = 1, 2, 3, …, M, where M is the length of the speech frame; The short-time zero-crossing rate is: where Z m is the short-time zero-crossing rate, x(n) is the pre-emphasized speech, ω(m) is a window function of length m, m is the time position of the short-time zero-crossing rate on the speech frame, m = 1, 2, 3, …, M, M is the length of the speech frame, and sgn(x(m)) is the sign function:

3. The method for adaptively segmenting Miao language speech based on time domain features according to claim 1, wherein The extraction of syllables of the voice frames through the short-time energy and the short-time zero-crossing rate to obtain voice syllables is specifically as follows: Extract vowels and consonants of the syllables of the voice frames through the short-time energy to obtain the vowels and the consonants; extract voiceless and voiced sounds of the syllables of the voice frames through the short-time zero-crossing rate to obtain the voiceless and the voiced sounds; Input the vowels, the consonants, the voiceless, and the voiced sounds into a preset syllable list to obtain voice syllables.

4. A Miao language speech adaptive segmentation method based on time-domain features according to claim 1, characterized in that, The calculation of the short-time energy to obtain the maximum short-time energy and the minimum short-time energy is specifically as follows: Calculate the short-time energy to obtain the average short-time energy; Calculate the maximum short-time energy through the average short-time energy, and the maximum short-time energy is: Among them, M H is the maximum short-time energy, is the average short-time energy, μ is a parameter, and 0 < μ ≤ 1; Calculate the minimum short-time energy through the maximum short-time energy and the short-time energy value, and the minimum short-time energy is: Among them, M L is the minimum short-time energy, is the short-time energy value of the first 5 frames, θ is a parameter, and 0 < θ ≤ 1.

5. A method for adaptive segmentation of Miao language speech based on time-domain features according to claim 1, characterized in that The calculation of the short-time zero-crossing rate to obtain the minimum zero-crossing rate is specifically as follows: Calculate the short-time zero-crossing rate to obtain the average short-time zero-crossing rate; Calculate the minimum zero-crossing rate through the average short-time zero-crossing rate, and the minimum zero-crossing rate is: Among them, Z S is the minimum zero-crossing rate, is the average short-time zero-crossing rate, β is a parameter, and 0 < β ≤ 1.

6. The method for adaptively segmenting Miao language speech based on time-domain features according to claim 1, wherein The pre-segmentation of the voice frames through the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple pre-segmented voice segments is specifically as follows: Set noise and the minimum voiced interval according to the voice syllables to obtain the number of noise frames and the number of minimum voiced intervals; Pre-segment the voice frames through the number of noise frames, the number of minimum voiced intervals, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple pre-segmented voice segments.

7. The method for adaptively segmenting Miao language speech based on time-domain features according to claim 6, wherein The construction of the fitness evaluation function model is specifically as follows: Let S be the true boundary segment and P be the predicted boundary segment. The fitness evaluation function model is as follows: s.t.S i ∈R,P i ∈R, i = 1, 2,..., M, where f(x) is the objective function of the true segmentation and the predicted segmentation, S i is the true boundary point of the i-th frame, P i is the predicted boundary point of the i-th frame, M is the length of the speech frame, i is the number of boundary frames, R is the set of real numbers, and s.t. is the abbreviation of subject to, indicating the constraint condition.

8. An adaptive segmentation method for Miao language speech based on time domain features according to claim 7, characterized in that, The adaptive segmentation of the multiple pre-segmented speech segments through the fitness evaluation function model to obtain multiple optimal speech segments is specifically as follows: Randomly generate an initial population for the multiple pre-segmented speech segments through binary coding. The initial population includes multiple individuals; and perform fitness calculation, selection operation, elite strategy, crossover operation, mutation operation, and evaluation operation on the initial population; The fitness calculation includes: performing the fitness calculation on multiple individuals in the initial population through the fitness evaluation function model to obtain fitness values; The selection operation includes: selecting the multiple individuals according to the fitness values, retaining multiple excellent individuals through the elite strategy, and generating the first-generation population from multiple inferior individuals; The crossover operation includes: performing two-point crossover on multiple inferior individuals in the next-generation population to obtain a diverse population; The mutation operation includes: mutating multiple inferior individuals in the diverse population with a preset probability to obtain a mutated population; Combine multiple excellent individuals in the elite strategy and multiple inferior individuals in the mutated population to generate a new population; The evaluation operation includes: performing target judgment on multiple individuals in the new population through the fitness evaluation function model, When the minimum objective function is not obtained, repeat the fitness calculation, the selection operation, the elite strategy, the crossover operation, the mutation operation, and the evaluation operation on multiple inferior individuals in the new population; When the minimum objective function is obtained, output the new population, and obtain two optimization parameters through the new population. Set the noise and the minimum unvoiced interval through the speech syllables and the optimization parameters to obtain the optimized noise frames and the minimum optimized unvoiced interval number; perform adaptive segmentation on the speech audio through the optimized noise frames, the minimum optimized unvoiced interval number, the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain multiple optimal speech segments.

9. The method for adaptively segmenting Miao language speech based on time domain features according to claim 1, wherein After obtaining the multiple speech segments, it further includes the step of calculating the accuracy rate, recall rate, and harmonic mean of the speech adaptive segmentation, specifically as follows: The calculation of the accuracy rate of the speech adaptive segmentation includes: Calculating the accuracy rate through the number of correctly segmented speech points and the total number of segmented speech points. The accuracy rate is: Where PRC is the accuracy rate, d is the number of correctly segmented speech points, l is the total number of segmented speech points, and · represents multiplication; The calculation of the recall rate of the speech adaptive segmentation includes: Calculating the recall rate through the number of correctly segmented speech points and the number of segmented speech points. The recall rate is: Where RCL is the recall rate, d is the number of correctly segmented speech points, h is the number of segmented speech points, and · represents multiplication; The calculation of the harmonic mean of the speech adaptive segmentation includes: The harmonic mean is calculated using the accuracy rate and the recall rate, and the harmonic mean is: where F is the harmonic mean, PRC is the accuracy rate, RCL is the recall rate, and · represents multiplication.

10. An adaptive segmentation system for Miao language speech based on time-domain features, characterized in that, It includes: a preprocessing module, a time-domain feature calculation module, a voice pre-segmentation module, and an adaptive voice segmentation module; The preprocessing module is used to pre-record a single-channel voice audio, extract the voice frame length from the voice audio according to a preset length to obtain voice frames, and preprocess the voice frames to obtain short-time energy and short-time zero-crossing rate; The time-domain feature calculation module is used to extract the syllables of the voice frames through the short-time energy and the short-time zero-crossing rate to obtain voice syllables; calculate the short-time energy to obtain the maximum short-time energy and the minimum short-time energy; calculate the short-time zero-crossing rate to obtain the minimum zero-crossing rate; The voice pre-segmentation module is used to pre-segment the voice frames through the maximum short-time energy, the minimum short-time energy, and the minimum zero-crossing rate to obtain a plurality of pre-segmented voice segments; The adaptive voice segmentation module is used to construct a fitness evaluation function model and adaptively segment the plurality of pre-segmented voice segments through the fitness evaluation function model to obtain a plurality of optimal voice segments.