Chinese dialects-oriented voice and text alignment archive method

CN122470730BActive Publication Date: 2026-09-29SHAOXING UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610957476.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0004]但现有的面向中文方言保护的语音与文本对齐存档方法在实际应用中,多依赖通用强制对齐工具或通用语音识别框架,缺乏针对方言合音、弱读、文白异读等特殊音变现象的专属检测与修正机制,忽略方言音变片段中时间边界非线性压缩、音素合并及标签替换等复杂对齐难题,没有建立可量化、可迭代的方言音系规则库与自适应反馈学习体系,无法覆盖不同方言片区、不同发音人及不同语体下的复杂音变对齐需求逻辑,导致现有对齐方法存在重通用性轻方言适配性的局限性

Benefits of technology

1、本发明通过构建集强制对齐模型构建、方言音系规则库定向修正、置信度阈值比对反馈、多维度关联存档为一体的面向中文方言保护的语音与文本对齐存档方法,结合连接时序分类算法与方言专用微调策略,实现方言语音数据与分层文本的对齐,同时通过预设的合音合并规则、弱读时长规则及文白异读映射规则,对初始时间边界进行针对性的定向修正,解决通用对齐工具在处理方言特殊音变现象时边界漂移、合并错位及标签错误的问题,提升对齐结果的准确性与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470730B_ABST
    Figure CN122470730B_ABST
Patent Text Reader

Abstract

The application discloses a speech and text alignment archiving method for Chinese dialect protection, and relates to the technical field of text archiving, and comprises the following steps: S1, acquiring data and marking segments; S2, based on a forced alignment model, inputting the data into the forced alignment model, and generating initial time boundaries of phonemes; S3, presetting a rule base and a division threshold, and performing directional correction on the initial time boundaries of the phonemes; S4, comparing the confidence in the initial time boundaries of the phonemes after directional correction with the boundary division threshold, and feeding back and updating; S5, associating the result with the data, and generating an archiving unit to import a dialect digital archive library. The application integrates the forced alignment model construction, the directional correction of the dialect phonological rule base and the multi-dimensional associated archiving into the speech and text alignment archiving method for Chinese dialect protection, combines the connection time sequence classification algorithm with the dialect special fine-tuning strategy, and realizes the alignment of dialect speech data and layered texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text archiving technology, and more particularly to a method for aligning speech and text archiving for the protection of Chinese dialects. Background Technology

[0002] The method of aligning and archiving speech and text for the protection of Chinese dialects serves as a fundamental supporting technology for the digital protection of dialects, the construction of dialect corpora, and the research of dialect speech. Its alignment accuracy, degree of automation, rule iterability, and archiving standardization level directly determine the quality of dialect corpus data, the credibility of subsequent research, and the efficiency and sustainability of dialect protection work.

[0003] The phenomena of phonological combination, weak pronunciation, and literary / colloquial pronunciation that are common in dialects are the core constraints affecting the high-precision alignment of speech and text. These phenomena are dynamically influenced by multiple factors, such as differences in the phonological systems of different dialect regions, differences in the speaker's speech speed and pronunciation habits, differences in language style, and differences in recording environment. Meanwhile, the dedicated alignment optimization technology for special dialect sound change phenomena is a key means to solve the pain points of low accuracy and high dependence on manual methods in traditional alignment methods.

[0004] However, existing methods for aligning and archiving speech and text for the protection of Chinese dialects often rely on general forced alignment tools or general speech recognition frameworks in practical applications. They lack specific detection and correction mechanisms for special sound change phenomena such as dialectal phonological changes, weak pronunciations, and literary and colloquial pronunciations. They ignore complex alignment problems such as nonlinear compression of time boundaries, phoneme merging, and label replacement in dialectal sound change segments. They have not established a quantifiable and iterative dialectal phonological rule base and an adaptive feedback learning system. They cannot cover the complex sound change alignment requirements of different dialect regions, different speakers, and different language styles. As a result, existing alignment methods have the limitation of emphasizing generality while neglecting dialect adaptability.

[0005] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0006] In response to the problems in related technologies, this invention proposes a method for aligning and archiving speech and text for the protection of Chinese dialects, in order to overcome the aforementioned technical problems existing in the existing related technologies.

[0007] To achieve the above objectives, the specific technical solution adopted by the present invention is as follows: A method for aligning speech and text for the preservation of Chinese dialects includes the following steps: S1. Obtain dialect speech data, speaker metadata, and layered text containing the international phoneme layer, and mark chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments in the international phoneme layer; S2. Construct a forced alignment model based on the connection time-series classification algorithm, and force-align dialect speech data with hierarchical text input to generate initial time boundaries for phonemes; As a preferred embodiment, the step of constructing a forced alignment model based on a connection-time classification algorithm and forcibly aligning dialect speech data with hierarchical text input to generate initial time boundaries for phonemes includes the following steps: S21. Perform acoustic preprocessing on the dialect speech data to obtain a standardized acoustic feature sequence, extract the international phoneme sequence from the layered text, and construct a phoneme label mapping table. S22. Construct a basic forced alignment network based on the connection time-series classification algorithm. Input the standardized acoustic feature sequence and the international phoneme sequence into the basic forced alignment network for supervised training. Iteratively adjust the model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model. As a preferred embodiment, the method of constructing a basic forced alignment network based on a connection-time classification algorithm, inputting standardized acoustic feature sequences and international phoneme sequences into the basic forced alignment network for supervised training, and iteratively adjusting the model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model includes the following steps: S221. Construct a basic forced alignment network structure based on the connection-based temporal classification algorithm, which includes a convolutional feature extraction layer, a bidirectional long short-term memory network layer, and a connection-based temporal classification output layer, and initialize the weights and bias parameters of each layer. S222. Divide the standardized acoustic feature sequences and international phoneme sequences into training sets and independent validation sets; S223. The connection-time classification loss function is used as the target loss function, and the adaptive moment estimation optimizer is used to supervise the training of the basic forced alignment network based on the training set, and the model weights and bias parameters are iteratively adjusted. S224. After each round of training, the phoneme boundary alignment accuracy of the model is calculated using an independent validation set. When the alignment accuracy no longer improves for several consecutive rounds, training is stopped, and the optimal model weights are saved to obtain a dialect-specific forced alignment model.

[0008] S23. Convert the dialect speech data into the corresponding standardized acoustic feature sequence, input it into the dialect-specific forced alignment model for temporal reasoning, generate the start timestamp, end timestamp and alignment confidence score for each phoneme, and form the initial time boundary of the phoneme.

[0009] S3. Preset the dialect phonological rule base and boundary division threshold, and perform targeted correction on the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak reading segments and literary and colloquial reading segments; As a preferred embodiment, the preset dialect phonological rule base and boundary segmentation threshold, and the targeted correction of the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments, include the following steps: S31. A dialect phonological rule library is preset, which includes rules for merging consonants, rules for the duration of weak pronunciations, and rules for mapping literary and colloquial pronunciations. Boundary division thresholds are set, which include duration deviation thresholds, boundary offset thresholds, and confidence thresholds. S32. Based on the international phoneme layer in the layered text, locate the time intervals and phoneme sequences corresponding to all chord segments, weak pronunciation segments, and literary / colloquial pronunciation segments from the initial time boundary of the phonemes. S33. Based on the merging rules in the dialect phonology rule base, calculate the duration deviation value of adjacent phonemes in the merging segment, compare the deviation with the duration deviation threshold, and then merge the time boundaries of adjacent phonemes according to the deviation comparison result, while adjusting the overall duration to the standard duration range of the merging segment. As a preferred embodiment, the step of calculating the duration deviation value of adjacent phonemes in a chord segment based on the chord merging rules in the dialect phonology rule base, comparing the deviation with the duration deviation threshold, merging the time boundaries of adjacent phonemes according to the deviation comparison result, and adjusting the overall duration to the standard chord duration range includes the following steps: S331. Extract the adjacent phoneme sequences corresponding to the chord segment and their respective start timestamps and end timestamps, retrieve the standard duration range of the chord corresponding to the chord from the dialect phonology rule base, calculate the sum of the individual durations of the adjacent phonemes, obtain the actual total duration of the chord segment, and calculate the difference between the actual total duration and the median of the standard duration range of the chord to obtain the duration deviation value. S332. Compare the duration deviation value with the duration deviation threshold to generate a deviation comparison result. Based on the deviation comparison result, merge the time boundaries of adjacent phonemes, set the start timestamp of the chorus segment to the start timestamp of the first phoneme, set the end timestamp to the end timestamp of the last phoneme, and then adjust the overall duration to within the standard duration range of the chorus.

[0010] S34. Based on the weak reading duration rules in the dialect phonology rule base, calculate the offset of the phoneme boundary in the weak reading segment, compare it with the boundary offset threshold, compress the phoneme duration to the weak reading standard duration range according to the offset comparison result, and correct the boundary offset at the same time. S35. Based on the literary and colloquial pronunciation mapping rules in the dialect phonological rule base, replace the corresponding phonemes in the literary and colloquial pronunciation segments and recalculate the alignment confidence. Then, compare the confidence with the confidence threshold and redivide the time boundary according to the confidence comparison result. S36. Integrate all corrected phoneme time boundaries with uncorrected phoneme time boundaries to generate directional corrected initial phoneme time boundaries.

[0011] S4. Compare the confidence level in the initial time boundary of the phonemes after directional correction with the boundary division threshold, and feed back the boundary division comparison result to update the dialect phonology rule base. As a preferred embodiment, the step of comparing the confidence level in the initial time boundary of the directionally corrected phonemes with the boundary segmentation threshold, and feeding back the boundary segmentation comparison result to update the dialect phonological rule base includes the following steps: S41. Extract the alignment confidence, duration deviation and boundary offset of each phoneme in the initial time boundary of the directionally corrected phonemes. S42. The alignment confidence, duration deviation value and boundary offset are compared with the confidence threshold, duration deviation threshold and boundary offset threshold respectively, and the boundary division comparison results containing qualified segments and unqualified segments are generated. S43. For the unqualified segments in the boundary segmentation comparison results, identify the corresponding speech phenomenon type, and extract the average duration, average boundary offset and average confidence statistics of the unqualified segments of this type. S44. Based on average confidence statistics, the weighted coefficients of the rules for merging consonants, the rules for weak pronunciation duration, and the rules for mapping literary and colloquial pronunciations are adjusted to automatically iterate and update the dialect phonological rule base.

[0012] As a preferred embodiment, the automatic iterative update of the dialect phonological rule base, based on average confidence statistics, involves adjusting the weighted coefficients of the merging rules, weak pronunciation duration rules, and literary / colloquial pronunciation mapping rules, and includes the following steps: S441. Based on the type of speech phenomenon corresponding to the unqualified segments, the statistical data of average duration, average boundary offset and average confidence are divided into statistical data of chords, statistical data of weak readings and statistical data of literary and colloquial pronunciations. S442. Based on the average duration and average confidence of the statistical data of phonological combinations, weak pronunciations, and literary / colloquial pronunciations, the gradient descent method is used to calculate the weighting coefficient adjustment of the phonological combination rules, weak pronunciation duration rules, and literary / colloquial pronunciation mapping rules. The weighting coefficients of the corresponding rules are updated using each adjustment, and the dialect phonological rule base is automatically iteratively updated.

[0013] As a preferred embodiment, the formula for calculating the weighted coefficient adjustment amount of the rules for merging consonants, weak pronunciation duration, and literary / colloquial pronunciation mapping using the gradient descent method is as follows:

[0014] In the formula, This is the adjustment amount for the weighting coefficients in the current iteration round t; For the current iteration round t, the weighted coefficient values ​​for each group corresponding to the merging rule of consonant sounds, the weak reading duration rule, and the literary and colloquial reading mapping rule are: The learning rate; For the loss function L with the current weighting coefficients gradient at; The momentum coefficient; The weighting coefficients are those from the previous iteration round t-1; And loss function The formula is:

[0015] In the formula, N is the total number of unqualified segments in the boundary segmentation comparison results; For the i-th non-compliant segment, the current weighting coefficient Alignment confidence; The duration error penalty coefficient is M; M is the total number of harmonic segments and weak reading segments involved in the directional correction. Let j be the actual duration of the j-th segment under the current rule; This refers to the standard duration of the harmonic or the standard duration of the weak pronunciation corresponding to this segment.

[0016] S5. Associate the boundary division comparison results with the speaker's metadata and generate standardized archival units to import into the dialect digital archive.

[0017] As a preferred embodiment, the step of associating the boundary segmentation comparison results with the speaker's metadata and generating standardized archival units for import into the dialect digital archive includes the following steps: S51. Obtain the dialect speech data, layered text, initial time boundaries of the phonemes after directional correction and boundary division comparison results during this processing, and retrieve the corresponding speaker metadata. S52. Using the unique identifier of the speech segment as an index, the dialect speech data, layered text, the initial time boundary of the corrected phonemes, the boundary division comparison results and the speaker's metadata are correlated to construct a multi-dimensional associated dataset. S53. In accordance with the unified format specifications of the dialect digital archive, the multi-dimensional associated dataset is encapsulated into a standardized archiving unit containing speech files, hierarchical text files, phoneme boundary files, quality assessment files and metadata files, and a globally unique archive number is generated. As a preferred embodiment, the process of encapsulating multi-dimensional associated datasets into standardized archival units containing speech files, hierarchical text files, phoneme boundary files, quality assessment files, and metadata files, and generating globally unique archive numbers, according to the unified format specifications of the dialect digital archive, includes the following steps: S531. According to the unified format specifications of the dialect digital archive, the dialect speech data in the multi-dimensional associated dataset is converted into standard digital audio format to generate speech files, the hierarchical text is converted into structured markup language format to generate hierarchical text files, and the initial time boundaries of the directionally corrected phonemes are converted into key-value pair format to generate phoneme boundary files. S532. Organize the boundary division comparison results into a quality assessment file containing a list of qualified segments and a list of unqualified segments, and integrate the speaker metadata with the timestamp of this processing, processing device information, and alignment model version information to generate a metadata file; S533. A globally unique file number is generated using a globally unique identifier generation algorithm, and the speech file, layered text file, phoneme boundary file, quality assessment file and metadata file are packaged into a standardized archive unit.

[0018] S54. Verify the data integrity and format compliance of the standardized archiving unit. After verification, import it into the dialect digital archive and establish a search index based on the archive number, dialect area, speaker information and speech phenomenon tags.

[0019] The beneficial effects of this invention are as follows: 1. This invention constructs a speech and text alignment and archiving method for the protection of Chinese dialects, integrating forced alignment model construction, dialect phonological rule base-oriented correction, confidence threshold comparison feedback, and multi-dimensional association archiving. Combining a connection-time classification algorithm with a dialect-specific fine-tuning strategy, it achieves alignment of dialect speech data with layered text. At the same time, through preset rules for merging consonants, rules for weak pronunciation duration, and rules for mapping literary and colloquial pronunciations, it performs targeted correction of the initial time boundary, solving the problems of boundary drift, merging misalignment, and labeling errors when general alignment tools deal with special dialect sound change phenomena, thus improving the accuracy and robustness of the alignment results.

[0020] 2. This invention generates statistical data on unqualified segments by comparing the directionally corrected confidence level with the boundary segmentation threshold. It then automatically adjusts the weighting coefficients in the rule base using a gradient descent momentum update formula, forming a closed-loop feedback learning mechanism. This enables adaptive iterative optimization of the dialect phonological rule base, reducing reliance on large-scale manual refining and minimizing the workload and professional requirements of manual annotation. Furthermore, by associating the final alignment result with speaker metadata and encapsulating it into standardized archiving units, it supports multi-dimensional retrieval by dialect region, speaker information, and speech phenomenon tags. This ensures the consistency and reusability of dialect corpus construction, avoiding the problems of high manual dependence, fixed and unupdable rules, inconsistent archiving formats, and difficult retrieval in dialect annotation methods. Ultimately, this invention achieves standardization, intelligence, and sustainability in the protection of Chinese dialects. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a method for aligning and archiving speech and text for the protection of Chinese dialects according to an embodiment of the present invention. Detailed Implementation

[0023] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0024] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0025] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, the speech and text alignment archiving method for Chinese dialect protection according to an embodiment of the present invention includes the following steps: S1. Obtain dialect speech data, speaker metadata, and layered text containing the international phoneme layer, and mark chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments in the international phoneme layer; Specifically, in a standard silent recording studio environment, a professional recording device with a 16-bit mono 16kHz sampling rate is used to collect continuous dialect speech data. The collected raw speech is then preprocessed by noise reduction, de-silence segment removal, and endpoint detection to obtain effective dialect speech segments. Collect basic information, language background information, and pronunciation habit information of the speakers, and organize them into standardized speaker metadata; based on the collected dialect speech data, create a three-layer structured text containing Chinese character layer, Pinyin layer, and International Phonetic Alphabet layer, in which the International Phonetic Alphabet layer uses IPA for phoneme-by-phoneme transcription; according to the preset speech phenomenon marking standard, uniformly mark all chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments in the International Phonetic Alphabet layer.

[0026] It should be explained that the criteria for collecting speaker metadata include: Basic information: name, gender, age, place of origin, education level; Language background information: dialect area, level of native language proficiency, second language use, daily language use environment; Pronunciation habit information: speech speed range, commonly used spoken expressions, whether there are accent variations and special pronunciation habits.

[0027] It should be noted that the structural specifications for layered text are as follows: Chinese character layer: Simplified Chinese standard characters are used, the standard spelling of dialect-specific vocabulary is preserved, and there are no typos; Pinyin layer: Pinyin is used for annotation, and the tone is marked with numbers (1-4 tones and neutral tone 0); International phoneme layer: The International Phonetic Alphabet (IPA) 2015 standard is used for transcription, and each phoneme is marked separately, including consonant, vowel and tone information.

[0028] It should be explained that the International Phoneme Layer's phonetic phenomenon labeling standards include: Consonant fragments: Mark single syllable fragments formed by the fusion of two or more adjacent phonemes due to connected speech, labeled with "[consonant]", for example, the "bu" in "buzhi" in Mandarin merges with "zhi" to form the sound "bu"; Weak pronunciation fragments: Mark phoneme or syllable fragments whose pronunciation is weakened, shortened in duration, and reduced in intensity due to sound changes in connected speech, labeled with "[weak pronunciation]", for example, the "zi" in "zhuozi" in Mandarin; Literary and colloquial pronunciation fragments: Mark fragments of the same Chinese character with different pronunciations in written and spoken language, labeled with "[literary pronunciation]" or "[colloquial pronunciation]", for example, the literary pronunciation of "jia" in Wu dialect is / ka / , and the colloquial pronunciation is / ko / .

[0029] S2. Construct a forced alignment model based on the connection time-series classification algorithm, and force-align dialect speech data with hierarchical text input to generate initial time boundaries for phonemes; In this embodiment of the invention, the step of constructing a forced alignment model based on a connection-time classification algorithm and forcibly aligning dialect speech data with hierarchical text input to generate initial time boundaries for phonemes includes the following steps: S21. Perform acoustic preprocessing on the dialect speech data to obtain a standardized acoustic feature sequence, extract the international phoneme sequence from the layered text, and construct a phoneme label mapping table. Specifically, effective dialect speech segments are sequentially pre-emphasized, framed, and windowed to suppress high-frequency attenuation of the speech signal and enhance the formant features of the vocal tract. The Mel frequency cepstral coefficients and their first-order and second-order difference features of each frame of speech signal are extracted and spliced ​​to form an initial acoustic feature vector. The initial acoustic feature vector is normalized to eliminate feature shifts caused by differences in recording equipment and speakers, resulting in a standardized acoustic feature sequence. Continuous international phoneme sequences are extracted line by line from the international phoneme layer of the layered text, and a one-to-one correspondence between international phonemes and integer labels is established to generate a phoneme label mapping table.

[0030] It should be explained that the core parameter standards for acoustic preprocessing are as follows: Preemphasis coefficient: 0.97; framing parameters: frame length 25ms, frame shift 10ms; window function: Hamming window; number of Mel filter banks: 40; Mel frequency cepstral coefficient dimensions: 13-dimensional static features + 13-dimensional first-order difference features + 13-dimensional second-order difference features, for a total of 39 dimensions; normalization method: global mean-variance normalization, making the mean of all feature dimensions 0 and the variance 1.

[0031] It should be noted that the construction specifications for the phoneme tag mapping table are as follows: The mapping table adopts a key-value pair structure, where the key is an international phoneme symbol and the value is a unique non-negative integer label. The label numbers are assigned consecutively starting from 0, without skipping or repeating. Special labels are reserved: 0 corresponds to the silence phoneme, 1 corresponds to the pause phoneme between sentences, and 2 corresponds to the unknown phoneme. It includes all independently pronounced consonants, vowels, and tone phonemes in this dialect, covering all phoneme types appearing in the international phoneme layer.

[0032] S22. Construct a basic forced alignment network based on the connection time-series classification algorithm. Input the standardized acoustic feature sequence and the international phoneme sequence into the basic forced alignment network for supervised training. Iteratively adjust the model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model. In this embodiment of the invention, the construction of a basic forced alignment network based on a connection-time classification algorithm, the input of standardized acoustic feature sequences and international phoneme sequences into the basic forced alignment network for supervised training, and the iterative adjustment of model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model includes the following steps: S221. Construct a basic forced alignment network structure based on the connection-based temporal classification algorithm, which includes a convolutional feature extraction layer, a bidirectional long short-term memory network layer, and a connection-based temporal classification output layer, and initialize the weights and bias parameters of each layer. Specifically, a convolutional feature extraction layer, a bidirectional long short-term memory network layer, and a connected temporal classification output layer are stacked sequentially to construct an end-to-end sequence labeling-based forced alignment network architecture. The input and output dimensions, convolutional kernel parameters, number of hidden units, and activation functions of each layer are set to determine the complete topology of the network. A hierarchical differentiated initialization strategy is adopted to independently initialize the weights and bias parameters of different types of network layers to ensure the stability of the initial state of the network.

[0033] It should be explained that the core parameter configurations for each layer of the basic forced alignment network are as follows: Convolutional Feature Extraction Layer: Contains two 1D convolutional neural networks with a kernel size of 3, a stride of 1, same padding, ReLU activation function, and 64 output channels. Each layer is followed by a batch normalization layer. Bidirectional Long Short-Term Memory Network Layer: Contains two bidirectional long short-term memory networks with 128 unidirectional hidden units per layer. Dropout is used to prevent overfitting, with an input dropout rate of 0.2 and a cyclic dropout rate of 0.1. Connected Temporal Classification Output Layer: Employs fully connected layers with Softmax activation function. The output dimension equals the total number of labels in the phoneme label mapping table, and the output result is the probability distribution of each phoneme label at each time step.

[0034] S222. Divide the standardized acoustic feature sequences and international phoneme sequences into training sets and independent validation sets; Specifically, the standardized acoustic feature sequences are paired and verified one by one with the corresponding international phoneme sequences, and invalid sample pairs with incorrect or mismatched labels are removed; the valid sample pairs are randomly divided into training sets and independent validation sets according to a preset ratio; the sequence length of all the divided samples is aligned to unify the input dimension within the batch; and the mapping relationship between the dataset partitioning results and sample labels is saved to ensure that the experimental process is reproducible.

[0035] It should be explained that the core specifications for dataset partitioning are as follows: The training set and the independent validation set are configured with a 7:3 ratio of samples. The training set is used for iterative updates of model parameters, while the independent validation set is used only for evaluating model generalization ability and early stopping detection, and does not participate in any parameter training process. Pairing verification rules: Each standardized acoustic feature sequence must correspond one-to-one with a unique international phoneme sequence, and the actual duration of the speech segment must deviate from the theoretical duration of the phoneme sequence by no more than 5%. Sequence length handling: Short sequences are padded with zero padding at the end, and long sequences exceeding the maximum length are truncated at the end. The maximum sequence length is uniformly set to 1000 time steps. Independence guarantee: All samples from the same speaker can only belong to either the training set or the independent validation set. Samples from the same speaker are prohibited from being distributed across sets to avoid data leakage leading to inflated model evaluation results. Data storage specifications: Feature sequence files, label sequence files, and sample index files for the training set and the independent validation set are stored separately in binary format to improve reading efficiency.

[0036] S223. The connection-time classification loss function is used as the target loss function, and the adaptive moment estimation optimizer is used to supervise the training of the basic forced alignment network based on the training set, and the model weights and bias parameters are iteratively adjusted. Specifically, the basic hyperparameters and termination conditions for model training are pre-set; the state parameters of the adaptive moment estimation optimizer are initialized and bound to the trainable parameters of the network; the standardized acoustic feature sequences and corresponding phoneme label sequences in the training set are loaded in batches iteratively; the feature sequences are input into the basic forced alignment network for forward propagation, and the connection temporal classification loss function value is calculated; the backpropagation algorithm is executed based on the loss value to solve the gradient of the weights and bias parameters of each layer; the adaptive moment estimation optimizer is used to iteratively update the model parameters according to the gradient information, and gradient clipping is applied to prevent gradient explosion; after each round of full training, the learning rate and state parameters of the optimizer are updated synchronously.

[0037] It should be explained that the core configuration for model training is as follows: Connection-time classification loss function: Adopts standard CTC loss calculation method, automatically handles the length difference between input and output label sequences, introduces blank labels to solve sequence alignment problems, and uses reserved labels in the phoneme label mapping table to ignore the loss calculation of inter-sentence pause labels; Adaptive moment estimation optimizer: Sets reasonable initial learning rate and moment estimation coefficients, introduces a weight decay mechanism to prevent overfitting, and uses a phased learning rate decay strategy to gradually reduce the learning rate and improve model convergence stability; Training control parameters: Sets reasonable batch size and maximum number of training epochs, and adopts an early stopping mechanism. When the phoneme boundary alignment accuracy on the independent validation set no longer improves for several consecutive epochs, training is terminated early and the current optimal model parameters are saved; Gradient constraint parameters: Employs gradient pruning technology. When the L2 norm of the gradient exceeds a preset threshold, the gradient is normalized to prevent gradient explosion during training that could lead to model divergence.

[0038] S224. After each round of training, the phoneme boundary alignment accuracy of the model is calculated using an independent validation set. When the alignment accuracy no longer improves for several consecutive rounds, training is stopped, and the optimal model weights are saved to obtain a dialect-specific forced alignment model.

[0039] Specifically, after each round of full training, all independent validation set samples are loaded; the standardized acoustic feature sequences from the validation set are input into the current training base forced alignment network to obtain phoneme prediction sequences at each time step; the prediction sequences are aligned and matched with the real international phoneme sequences to calculate the phoneme boundary alignment accuracy of this round; the historical best accuracy is compared, and if the current accuracy is higher, the best accuracy record is updated and the current model weights are saved; if the accuracy of multiple consecutive training rounds does not exceed the historical best value, the early stopping mechanism is triggered to terminate training, and the saved best model weights are used as the final dialect-specific forced alignment model.

[0040] It should be explained that the core specifications for model validation and early stopping are as follows: Validation set evaluation process: Forward propagation logic, identical to that used in the training process, is employed for prediction without any parameter updates; only the model's generalization ability is evaluated. Phoneme boundary alignment accuracy calculation: The degree of matching between the predicted boundary and the true boundary is statistically analyzed on a phoneme-by-phoneme basis. A correct match is determined only when both the start and end positions of the phoneme are within the allowable error range. Early stopping trigger rule: A reasonable threshold for consecutive rounds without improvement is set. When the validation set accuracy fails to improve for several consecutive rounds, the model is considered to have converged, and training is terminated early to prevent overfitting. Optimal model preservation strategy: Model weights are preserved only when the validation set accuracy exceeds the historical best value. The globally optimal model state is always maintained, and the final exported dialect-specific forced alignment model is the model corresponding to this optimal state.

[0041] S23. Convert the dialect speech data into the corresponding standardized acoustic feature sequence, input it into the dialect-specific forced alignment model for temporal reasoning, generate the start timestamp, end timestamp and alignment confidence score for each phoneme, and form the initial time boundary of the phoneme.

[0042] Specifically, the dialect speech data to be processed undergoes an acoustic preprocessing procedure identical to that in the training phase to obtain the corresponding standardized acoustic feature sequence. The standardized acoustic feature sequence is then input into a dialect-specific forced alignment model, and pure forward propagation temporal inference is performed to obtain the probability distribution of each phoneme label at each time step. The probability distribution is decoded using a sequence decoding method matching the training process to generate a continuous phoneme sequence and calculate the alignment confidence score for each phoneme. Based on the decoding results, the time step range corresponding to each phoneme is determined, mapped and converted into actual start and end timestamps, and integrated to form a complete initial time boundary for the phoneme.

[0043] It should be explained that the core specifications for model inference and initial boundary generation are as follows: Inference Preprocessing Specifications: Strictly reuse the acoustic preprocessing parameters and processes from the training phase to ensure that the feature distribution in the inference phase is completely consistent with that in the training phase, avoiding a decrease in alignment accuracy caused by feature shift; Model Inference Specifications: The inference process only performs forward propagation calculations, without any parameter updates or gradient calculations, and adopts batch inference to improve processing efficiency; Phoneme Decoding and Confidence Specifications: Adopt a decoding algorithm that matches the connection-time classification loss function, and calculate the alignment confidence score based on the average probability value of the corresponding time step of the phoneme, reflecting the reliability of the phoneme boundary; Time Boundary Mapping Specifications: Based on the correspondence between the time steps in the preprocessing phase and the actual duration, accurately convert the decoded phoneme time step range into start and end timestamps in seconds, forming structured initial time boundary data for the phonemes.

[0044] S3. Preset the dialect phonological rule base and boundary division threshold, and perform targeted correction on the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak reading segments and literary and colloquial reading segments; As a preferred embodiment, the preset dialect phonological rule base and boundary segmentation threshold, and the targeted correction of the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments, include the following steps: S31. A dialect phonological rule library is preset, which includes rules for merging consonants, rules for the duration of weak pronunciations, and rules for mapping literary and colloquial pronunciations. Boundary division thresholds are set, which include duration deviation thresholds, boundary offset thresholds, and confidence thresholds. Specifically, the system systematically analyzes the phonological features and speech change rules of the target dialect, collects typical examples of merging, weak pronunciation, and literary / colloquial pronunciation in the dialect, extracts rules for merging merging, weak pronunciation duration, and literary / colloquial pronunciation mapping, and constructs a structured and scalable dialect phonological rule base. Combining the phonetic characteristics and alignment accuracy requirements of the dialect, duration deviation thresholds, boundary offset thresholds, and confidence thresholds are set to form a unified boundary division threshold system.

[0045] It should be explained that the construction specifications for the dialect phonology rule base are as follows: Phonetic merging rules: Clarify the criteria, merging methods, and resulting phonemes for merging adjacent phonemes in the dialect, covering all common phonetic merging phenomena in daily spoken language; Weak pronunciation duration rules: Specify the duration criteria, pronunciation weakening characteristics, and correction principles for weak pronunciation phonemes, and distinguish the differences in the degree of weak pronunciation in different contexts; Literary and colloquial pronunciation mapping rules: Establish a one-to-one correspondence between the written and spoken pronunciations of the same Chinese character, and clarify the applicable scenarios and usage conditions for different pronunciations.

[0046] It should be explained that the setting specifications for the boundary division threshold are as follows: Duration deviation threshold: used to determine whether the actual duration of chord segments and weak reading segments conforms to the phonetic rules of the dialect. Segments exceeding the threshold are marked as needing correction. Boundary offset threshold: used to determine whether the positional error of the initial boundary of phonemes is within the allowable range. Boundaries with offsets exceeding the threshold are marked as needing adjustment. Confidence threshold: used to filter reliable alignment results obtained from model inference. Segments with confidence scores below the threshold are marked as unqualified and require further correction.

[0047] S32. Based on the international phoneme layer in the layered text, locate the time intervals and phoneme sequences corresponding to all chord segments, weak pronunciation segments, and literary / colloquial pronunciation segments from the initial time boundary of the phonemes. Specifically, the process involves loading the pre-labeled speech phenomenon tags from the international phoneme layer in the layered text; traversing all phoneme entries at the initial time boundary of the phonemes and matching them position by position with the tags from the international phoneme layer; identifying all speech segments labeled as consonants, weak pronunciations, and literary / colloquial pronunciations, and extracting the corresponding continuous time intervals and complete phoneme sequences for each segment; and classifying and organizing the extracted results according to the speech phenomenon type to form a structured set of segments to be corrected.

[0048] It should be explained that the core specifications for speech segment localization are as follows: Tag matching specifications: Matching and positioning are strictly based on the exclusive tags for chords, weak pronunciations, and literary / colloquial pronunciations pre-added in the international phoneme layer to ensure that no labeled speech phenomena are missed or mismatched; Time interval extraction specifications: The time interval of each speech segment is generated by merging the start and end timestamps of all phonemes contained in the segment to ensure the continuity and integrity of the time interval; Phoneme sequence extraction specifications: The complete international phoneme sequence corresponding to the speech segment is extracted, and the original phoneme order and tone information are preserved as the basis for subsequent targeted correction; Classification and organization specifications: The positioning results are stored separately according to three categories: chord segments, weak pronunciation segments, and literary / colloquial pronunciation segments, which facilitates the subsequent targeted application of different correction rules.

[0049] S33. Based on the merging rules in the dialect phonology rule base, calculate the duration deviation value of adjacent phonemes in the merging segment, compare the deviation with the duration deviation threshold, and then merge the time boundaries of adjacent phonemes according to the deviation comparison result, while adjusting the overall duration to the standard duration range of the merging segment. As a preferred embodiment, the step of calculating the duration deviation value of adjacent phonemes in a chord segment based on the chord merging rules in the dialect phonology rule base, comparing the deviation with the duration deviation threshold, merging the time boundaries of adjacent phonemes according to the deviation comparison result, and adjusting the overall duration to the standard chord duration range includes the following steps: S331. Extract the adjacent phoneme sequences corresponding to the chord segment and their respective start timestamps and end timestamps, retrieve the standard duration range of the chord corresponding to the chord from the dialect phonology rule base, calculate the sum of the individual durations of the adjacent phonemes, obtain the actual total duration of the chord segment, and calculate the difference between the actual total duration and the median of the standard duration range of the chord to obtain the duration deviation value. Specifically, the process iterates through all the chord segments to be corrected, extracts the complete adjacent phoneme sequence corresponding to each chord segment, and the start and end timestamps of each phoneme in the sequence; based on the extracted adjacent phoneme sequence, the standard duration range of the chord corresponding to the chord is retrieved from the dialect phonology rule base; the individual durations of all adjacent phonemes in the sequence are accumulated to obtain the actual total duration of the chord segment; the median of the standard duration range of the chord is determined, and the difference between the actual total duration and the median is calculated to obtain the duration deviation value of the chord segment.

[0050] It should be explained that the core specifications for calculating the duration deviation of harmony segments are as follows: Harmony Information Extraction Standard: Completely extract all continuous phonemes and their time information contained in the harmony segment to ensure the integrity of the phoneme sequence and the accuracy of the timestamps; Standard Duration Retrieval Standard: Based on the phoneme combination characteristics of the harmony, perform precise matching to retrieve the general standard duration range of this type of harmony in the corresponding dialect; Actual Duration Calculation Standard: Obtain the duration of a single phoneme by subtracting the start timestamp from the end timestamp of the phoneme, and sum the durations of all phonemes to obtain the actual total duration of the harmony segment; Deviation Value Calculation Standard: Calculate the deviation based on the median value of the standard duration range of the harmony, with the positive and negative values ​​of the deviation value indicating whether the actual duration is longer or shorter than the standard duration, respectively.

[0051] S332. Compare the duration deviation value with the duration deviation threshold to generate a deviation comparison result. Based on the deviation comparison result, merge the time boundaries of adjacent phonemes, set the start timestamp of the chorus segment to the start timestamp of the first phoneme, set the end timestamp to the end timestamp of the last phoneme, and then adjust the overall duration to within the standard duration range of the chorus.

[0052] Specifically, a preset duration deviation threshold is loaded, and the duration deviation value of each chord segment is compared with the threshold one by one to generate the corresponding deviation comparison result. For chord segments whose deviation value exceeds the threshold range, a time boundary merging operation is performed, and the start timestamp of the chord segment is uniformly set to the start timestamp of the first phoneme it contains, and the end timestamp is uniformly set to the end timestamp of the last phoneme it contains. According to the standard duration range corresponding to the chord in the dialect phonology rule library, the overall duration of the merged chord segment is reasonably adjusted to ensure that the adjusted duration falls within the standard duration range.

[0053] It should be explained that the core specifications for merging harmony segment boundaries and adjusting duration are as follows: Deviation Comparison Specifications: Using a preset duration deviation threshold as the judgment standard, chorus segments with deviation values ​​within the threshold range are judged as normal and do not require correction; those exceeding the threshold range are judged as needing correction, and subsequent boundary merging and duration adjustment operations are performed; Boundary Merging Specifications: The independent time boundaries of all adjacent phonemes contained in a chorus segment are merged into a continuous overall time boundary, eliminating the problem of internal boundary segmentation of the chorus caused by model alignment; Duration Adjustment Specifications: Based on the standard duration range of chorus in the dialect phonology rule base, the overall duration of the merged segment is fine-tuned, prioritizing the preservation of the overall time position of the chorus segment, and only adjusting the end timestamp to adapt to the standard duration; Result Recording Specifications: The original time boundary, deviation value, corrected time boundary, and adjustment basis of each chorus segment are recorded to facilitate subsequent quality checks and rule base optimization.

[0054] S34. Based on the weak reading duration rules in the dialect phonology rule base, calculate the offset of the phoneme boundary in the weak reading segment, compare it with the boundary offset threshold, compress the phoneme duration to the weak reading standard duration range according to the offset comparison result, and correct the boundary offset at the same time. Specifically, the process iterates through all weak reading segments to be corrected, extracting the phoneme sequence and its original time boundary information for each weak reading segment; retrieves the standard duration range of weak reading segments corresponding to this type of weak reading segment from the dialect phonological rule base; calculates the offset of the phoneme boundary of the weak reading segment based on the original time boundary and the standard duration range; compares the calculated offset with the preset boundary offset threshold to generate the offset comparison result; for weak reading segments whose offset exceeds the threshold range, the overall duration of the phoneme is compressed to the standard duration range according to the weak reading duration rule, while simultaneously correcting the start and end time boundaries of the phoneme to eliminate the boundary offset.

[0055] It should be explained that the core specifications for weak read segment duration compression and boundary correction are as follows: Weak pronunciation rule calling specifications: Based on the phoneme type and contextual features of weak pronunciation segments, accurately match the corresponding weak pronunciation duration rules and standard duration ranges from the dialect phonological rule base; Offset calculation specifications: Using the median value of the standard duration range of weak pronunciations as a benchmark, calculate the positional deviation between the original phoneme boundary and the theoretical boundary to obtain the boundary offset; Offset comparison specifications: Using a preset boundary offset threshold as the judgment standard, weak pronunciation segments with offsets within the threshold range are judged as normal and do not require correction; those exceeding the threshold range are judged as needing correction, and subsequent duration compression and boundary correction operations are performed; Duration compression specifications: Adjust the duration of weak pronunciation segments using an overall proportional compression method, prioritizing ensuring that the center time position of the weak pronunciation segment remains unchanged, and only synchronously adjusting the start and end timestamps; Boundary correction specifications: After duration compression is completed, synchronously update the phoneme boundary information of the weak pronunciation segment to ensure that the corrected boundary is continuous and conforms to the weak pronunciation rules of the dialect.

[0056] S35. Based on the literary and colloquial pronunciation mapping rules in the dialect phonological rule base, replace the corresponding phonemes in the literary and colloquial pronunciation segments and recalculate the alignment confidence. Then, compare the confidence with the confidence threshold and redivide the time boundary according to the confidence comparison result. Specifically, the process iterates through all the literary and colloquial pronunciation segments to be corrected, extracting the original phoneme sequence, temporal boundaries, and contextual information for each segment; retrieves the corresponding literary and colloquial pronunciation mapping rules from the dialect phonology rule base; replaces the corresponding phonemes in the literary and colloquial pronunciation segments according to the mapping rules, generating a corrected phoneme sequence; recalculates the alignment confidence of the segment based on the matching degree between the corrected phoneme sequence and speech features; compares the recalculated alignment confidence with a preset confidence threshold to generate a confidence comparison result; for segments with a confidence level that meets the requirements, re-divides the temporal boundaries based on the corrected phoneme sequence, updating the phoneme and boundary information of the literary and colloquial pronunciation segments.

[0057] It should be explained that the core specifications for correcting literary and colloquial pronunciation fragments and redefining boundaries are as follows: Mapping rule calling specifications: Based on the Chinese character information and contextual features of the literary and colloquial pronunciation segments, the corresponding literary and colloquial pronunciation mapping rules are accurately matched from the dialect phonology rule base to distinguish the different pronunciation requirements of written and spoken language scenarios; Phoneme replacement specifications: Phoneme replacement is performed strictly according to the mapping rules, replacing only the phoneme parts corresponding to the literary and colloquial pronunciations, and preserving the original information and order of other phonemes in the segment; Confidence recalculation specifications: Based on the acoustic feature matching degree between the replaced phoneme sequence and the corresponding speech segment, the overall alignment confidence of the literary and colloquial pronunciation segment is recalculated to reflect the reliability of the corrected boundary; Confidence comparison specifications: Using the preset confidence threshold as the judgment standard, the correction result with a confidence level higher than the threshold is judged as valid and subsequent boundary re-division operation is performed; the result with a confidence level lower than the threshold is retained as the original result and marked as pending manual verification; Boundary re-division specifications: Based on the correspondence between the corrected phoneme sequence and acoustic features, the time boundary of the literary and colloquial pronunciation segment is re-divided to ensure that the corrected boundary is consistent with the actual pronunciation position.

[0058] S36. Integrate all corrected phoneme time boundaries with uncorrected phoneme time boundaries to generate directional corrected initial phoneme time boundaries.

[0059] Specifically, the process involves collecting all phoneme time boundaries after processing such as merging, weak pronunciation correction, and literary / colloquial pronunciation replacement, as well as the original phoneme time boundaries that do not require correction; sorting all phoneme boundaries in chronological order; checking the continuity of adjacent phoneme boundaries to eliminate boundary overlaps and time gaps; verifying the correspondence between all phoneme boundaries and their corresponding phoneme sequences; and integrating all processed phoneme boundaries into a complete continuous sequence to generate the initial time boundaries of the phonemes after directional correction.

[0060] It should be explained that the core specifications for phoneme time boundary integration are as follows: Boundary Priority Specification: The corrected phoneme time boundaries have higher priority than the original uncorrected boundaries, and all segments that have undergone targeted correction use the corrected boundary data; Time Order Specification: Strictly sort the phonemes according to their actual pronunciation times to ensure that the integrated boundary sequence is completely consistent with the time axis of the speech signal; Continuity Processing Specification: For overlaps or gaps between adjacent phoneme boundaries, fine-tuning is performed based on the corrected boundaries to ensure that the boundaries are continuous and without overlap or omission; Integrity Verification Specification: After integration, it is verified that all phonemes are included in the final boundary sequence and correspond one-to-one with the international phoneme sequence in the layered text; Result Output Specification: Output structured phoneme time boundary data, including the international phonetic symbol, start timestamp, end timestamp, and alignment confidence information for each phoneme.

[0061] S4. Compare the confidence level in the initial time boundary of the phonemes after directional correction with the boundary division threshold, and feed back the boundary division comparison result to update the dialect phonology rule base. As a preferred embodiment, the step of comparing the confidence level in the initial time boundary of the directionally corrected phonemes with the boundary segmentation threshold, and feeding back the boundary segmentation comparison result to update the dialect phonological rule base includes the following steps: S41. Extract the alignment confidence, duration deviation and boundary offset of each phoneme in the initial time boundary of the directionally corrected phonemes. Specifically, the process involves iterating through all phoneme entries in the initial time boundary after the orientation correction; extracting the alignment confidence data for each phoneme from the model inference record and correction log; calculating the difference between the actual duration and the corresponding standard duration for each phoneme based on the standard duration in the dialect phonology rule base to obtain the duration deviation value; comparing the phoneme boundary positions before and after correction to calculate the positional deviation between the start and end boundaries of each phoneme to obtain the boundary offset; and associating and storing the above three types of parameters with the unique identifier of the corresponding phoneme to form a complete phoneme quality assessment dataset.

[0062] It should be explained that the core specifications for phoneme quality parameter extraction are as follows: Alignment confidence extraction specifications: For ordinary phonemes that have not undergone directional correction, the original alignment confidence obtained from model inference is used directly; for phonemes that have undergone correction for chords, weak forms, or literary / colloquial pronunciations, the alignment confidence recalculated during the correction process is used; Duration deviation extraction specifications: For chord and weak form segments, the difference between the corrected actual duration and the median of the corresponding standard duration range is used; for ordinary phonemes, the difference between the actual duration and the average standard duration of similar phonemes in the dialect is used; Boundary offset extraction specifications: For phonemes that have undergone boundary correction, the position difference between the corrected boundary and the original initial boundary is used; for uncorrected phonemes, the position difference between the predicted boundary and the theoretical standard boundary is used; Data association specifications: All extracted parameters are bound to the unique identifier of the corresponding phoneme, the type of speech phenomenon, and the correction record to ensure that the source of each parameter is traceable and the calculation process is reproducible.

[0063] S42. The alignment confidence, duration deviation value and boundary offset are compared with the confidence threshold, duration deviation threshold and boundary offset threshold respectively, and the boundary division comparison results containing qualified segments and unqualified segments are generated. Specifically, the process iterates through all phoneme entries in the phoneme quality assessment dataset; it compares the alignment confidence of each phoneme with a preset confidence threshold, compares the duration deviation with a preset duration deviation threshold, and compares the boundary offset with a preset boundary offset threshold; based on the combined results of these three comparisons, it determines whether the segment corresponding to each phoneme is a qualified or unqualified segment; it categorizes all phonemes according to segment type, records the specific reasons for the unqualified segment, and generates a complete boundary division comparison result.

[0064] It should be explained that the core specifications for boundary delimitation and comparison are as follows: Itemized Comparison Standards: Each parameter is compared strictly according to its corresponding threshold, ensuring consistent evaluation standards for each parameter and conforming to the dialect's phonetic characteristics. Pass / Fail Judgment Standards: When the alignment confidence, duration deviation, and boundary offset of a phoneme all meet the corresponding threshold requirements, the segment corresponding to that phoneme is deemed a pass / fail segment. Failure Classification Standards: Based on the type of parameter that does not meet the requirements, unqualified segments are classified into categories with insufficient confidence, excessive duration deviation, and excessive boundary offset, with specific reasons for failure recorded for each. Result Recording Standards: The boundary classification comparison results include a unique identifier for each phoneme, the values ​​of the three parameters, the comparison result, and the reason for failure, facilitating subsequent rule base updates and manual review.

[0065] S43. For the unqualified segments in the boundary segmentation comparison results, identify the corresponding speech phenomenon type, and extract the average duration, average boundary offset and average confidence statistics of the unqualified segments of this type. Specifically, the process iterates through all non-compliant segments in the boundary segmentation comparison results, and identifies the speech phenomenon type corresponding to each non-compliant segment by combining the speech phenomenon markers of the segments with historical correction records. The non-compliant segments are then classified and grouped according to the identified speech phenomenon types, and all segments of the same type are grouped into the same statistical unit. The average duration, average boundary offset, and average confidence statistics of all non-compliant segments in each statistical unit are calculated. The statistical results are then associated and stored with the corresponding speech phenomenon types to form a structured dataset for statistical analysis of non-compliant segments.

[0066] It should be explained that the core specifications for the identification and statistics of non-compliant segments are as follows: Speech Phenomenon Type Recognition Standard: Based on the original labeling of segments at the international phoneme layer and the processing records during the directional correction process, accurately determine the corresponding speech phenomenon type, covering four major categories: chord segments, weak pronunciation segments, literary and colloquial pronunciation segments, and ordinary phoneme segments; Grouping and Statistical Standard: Speech phenomenon types are strictly grouped and statistically analyzed independently, and segments of different types are not mixed in the calculation to ensure that the statistical results can accurately reflect the alignment problem characteristics of various speech phenomena; Statistical Indicator Calculation Standard: When calculating each statistical indicator, obviously abnormal outliers are automatically removed to ensure the objectivity and representativeness of the statistical results, and the statistical results retain reasonable accuracy; Result Association Standard: Statistical results are bound and stored with the corresponding speech phenomenon type, the type of non-compliance reason, and the number of samples, fully recording the key information of the statistical process, and providing data support for the subsequent iterative optimization of the dialect phonology rule base.

[0067] S44. Based on average confidence statistics, the weighted coefficients of the rules for merging consonants, the rules for weak pronunciation duration, and the rules for mapping literary and colloquial pronunciations are adjusted to automatically iterate and update the dialect phonological rule base.

[0068] As a preferred embodiment, the automatic iterative update of the dialect phonological rule base, based on average confidence statistics, involves adjusting the weighted coefficients of the merging rules, weak pronunciation duration rules, and literary / colloquial pronunciation mapping rules, and includes the following steps: S441. Based on the type of speech phenomenon corresponding to the unqualified segments, the statistical data of average duration, average boundary offset and average confidence are divided into statistical data of chords, statistical data of weak readings and statistical data of literary and colloquial pronunciations. Specifically, the statistical analysis dataset of unqualified segments is loaded, and a mapping relationship is established between speech phenomenon types and statistical data categories. Statistical data corresponding to chord phenomena are grouped into chord category statistical data, statistical data corresponding to weak pronunciation phenomena are grouped into weak pronunciation category statistical data, and statistical data corresponding to literary and colloquial pronunciation differences are grouped into literary and colloquial pronunciation difference statistical data. Unqualified data of ordinary phonemes that do not belong to the above three categories of speech phenomena are stored separately in isolation. A unique type identifier is added to each type of statistical data to unify the field structure of various types of data. The statistical data set after classification is saved to provide a data foundation for subsequent targeted adjustment of rule parameters.

[0069] It should be explained that the core specifications for statistical data classification are as follows: Classification Mapping Standards: Strictly adhere to the unique identifiers of speech phenomenon types for accurate mapping, ensuring that all statistical data of the same type are classified into their corresponding categories without misclassification or omission; Unified Data Structure Standards: Maintain completely consistent field structures for the three types of statistical data, including average duration, average boundary offset, average confidence level, sample size, and distribution of reasons for non-compliance, facilitating subsequent unified processing; Abnormal Data Isolation Standards: Store non-compliant data of ordinary phonemes separately, excluding them from rule parameter adjustments for specific speech phenomena, avoiding interference with targeted optimization effects; Type Identification Standards: Add clear and unambiguous type identifiers to each type of statistical data, facilitating rapid identification and retrieval of statistical data of the corresponding category.

[0070] S442. Based on the average duration and average confidence of the statistical data of phonological combinations, weak pronunciations, and literary / colloquial pronunciations, the gradient descent method is used to calculate the weighting coefficient adjustment of the phonological combination rules, weak pronunciation duration rules, and literary / colloquial pronunciation mapping rules. The weighting coefficients of the corresponding rules are updated using each adjustment, and the dialect phonological rule base is automatically iteratively updated.

[0071] As a preferred embodiment, the formula for calculating the weighted coefficient adjustment amount of the rules for merging consonants, weak pronunciation duration, and literary / colloquial pronunciation mapping using the gradient descent method is as follows:

[0072] In the formula, This is the adjustment amount for the weighting coefficients in the current iteration round t; For the current iteration round t, the weighted coefficient values ​​for each group corresponding to the merging rule of consonant sounds, the weak reading duration rule, and the literary and colloquial reading mapping rule are: The learning rate; For the loss function L with the current weighting coefficients gradient at; The momentum coefficient; The weighting coefficients are those from the previous iteration round t-1; And loss function The formula is:

[0073] In the formula, N is the total number of unqualified segments in the boundary segmentation comparison results; For the i-th non-compliant segment, the current weighting coefficient Alignment confidence; The duration error penalty coefficient is M; M is the total number of harmonic segments and weak reading segments involved in the directional correction. Let j be the actual duration of the j-th segment under the current rule; This refers to the standard duration of the harmonic or the standard duration of the weak pronunciation corresponding to this segment.

[0074] S5. Associate the boundary division comparison results with the speaker's metadata and generate standardized archival units to import into the dialect digital archive.

[0075] As a preferred embodiment, the step of associating the boundary segmentation comparison results with the speaker's metadata and generating standardized archival units for import into the dialect digital archive includes the following steps: S51. Obtain the dialect speech data, layered text, initial time boundaries of the phonemes after directional correction and boundary division comparison results during this processing, and retrieve the corresponding speaker metadata. Specifically, the system sequentially extracts the original dialect speech data, annotated layered text data, directionally corrected initial time boundary data of phonemes, and boundary division comparison results from the system's temporary storage module; retrieves the speaker metadata corresponding to the speech data from the system's speaker information management module; encapsulates all the acquired data in a unified manner, establishes a unique association between each data, and forms a complete dataset of the processing results; and performs integrity verification on the encapsulated dataset to ensure that all key data has been acquired and is in a consistent format.

[0076] It should be noted that the core specifications for data acquisition and encapsulation in this process are as follows: Raw Data Extraction Standards: Strictly extract the raw data generated during this processing, without any modification or secondary processing, ensuring data authenticity and traceability; Metadata Retrieval Standards: Accurately retrieve the corresponding speaker's metadata based on the unique identifier of the voice data. The metadata includes key information such as the speaker's dialect background, recording environment, and equipment information; Data Association Standards: Establish associations between all data using a unified unique identifier, ensuring that all relevant data can be quickly located through any data entry; Integrity Verification Standards: Verify that all required data is complete, check if the data format meets system requirements, and verify the validity of data associations, ensuring the packaged dataset is complete and usable.

[0077] S52. Using the unique identifier of the speech segment as an index, the dialect speech data, layered text, the initial time boundary of the corrected phonemes, the boundary division comparison results and the speaker's metadata are correlated to construct a multi-dimensional associated dataset. Specifically, the core index uses the globally unique identifier of each speech segment as the primary key; it iterates through dialect speech data, layered text, initial time boundaries of phonemes after directional correction, boundary division comparison results, and speaker metadata to extract speech segment identification information carried in each data; it matches and binds all data corresponding to the same unique identifier to establish a bidirectional association between the data; it generates a structured association index table to record the data storage location and field mapping corresponding to each unique identifier; it performs consistency verification on the association dataset to finally form a multi-dimensional association dataset.

[0078] It should be explained that the core specifications for constructing multi-dimensional related datasets are as follows: Index Primary Key Specification: A globally unique and immutable speech segment identifier is used as the unique index primary key to ensure that each speech segment has a unique identity in the dataset, avoiding data confusion and duplicate associations; Full-Dimensional Association Specification: Ensure that the original speech data, text annotation data, boundary data, quality assessment data, and speaker information of each speech segment are all fully associated, with no missing data in any key dimensions; Consistency Verification Specification: Verify whether the time span, phoneme count, and other core attributes of all data under the same index are consistent, and immediately mark and trace the source of any mismatched data; Index Optimization Specification: Establish secondary indexes based on speech phenomenon type, speaker information, and quality level to improve the efficiency of subsequent data queries, statistical analysis, and rule optimization.

[0079] S53. In accordance with the unified format specifications of the dialect digital archive, the multi-dimensional associated dataset is encapsulated into a standardized archiving unit containing speech files, hierarchical text files, phoneme boundary files, quality assessment files and metadata files, and a globally unique archive number is generated. As a preferred embodiment, the process of encapsulating multi-dimensional associated datasets into standardized archival units containing speech files, hierarchical text files, phoneme boundary files, quality assessment files, and metadata files, and generating globally unique archive numbers, according to the unified format specifications of the dialect digital archive, includes the following steps: S531. According to the unified format specifications of the dialect digital archive, the dialect speech data in the multi-dimensional associated dataset is converted into standard digital audio format to generate speech files, the hierarchical text is converted into structured markup language format to generate hierarchical text files, and the initial time boundaries of the directionally corrected phonemes are converted into key-value pair format to generate phoneme boundary files. Specifically, the pre-constructed multi-dimensional associated dataset is retrieved, and format conversion and file generation are carried out strictly in accordance with the unified format specifications of the dialect digital archive. First, dialect speech data is extracted from the dataset, and transcoded according to the encoding and encapsulation standards required by the archive, converting it to the specified standard digital audio format to generate formal speech files. Next, the layered text content within the dataset is extracted, and the text hierarchy and annotation information are standardized, converting it to the corresponding structured markup language format to generate layered text files. Finally, the initial time boundary data of the phonemes after targeted correction is extracted, and the boundary information is reorganized into standard key-value pairs according to preset rules to generate phoneme boundary files. After all files are generated, the file content is initially verified to be consistent with the original data source.

[0080] It should be explained that the core specifications for format conversion and file generation are as follows: Overall Format Standards: The entire process adheres to the format requirements of the Dialect Digital Archive, maintaining consistency in file encapsulation, encoding methods, field structures, and naming rules to ensure generated files are directly compatible with the archive for storage and retrieval. Speech Conversion Standards: Audio format conversion only adjusts the encoding and encapsulation, fully preserving the acoustic features, duration information, and vocal tract data of the original speech without altering its content. Hierarchical Text Conversion Standards: During conversion to structured markup language, the hierarchical structure, phoneme annotations, and various speech phenomenon labels of the original text are fully preserved, ensuring the accuracy of hierarchical logic and annotation information. Phoneme Boundary Conversion Standards: Data fields are organized according to standard key-value pair rules, mapping phoneme content, timestamps, confidence levels, and other information one-to-one, unifying field names and data formats to ensure normal data parsing. Content Validation Standards: After file generation, the output file is compared one by one with the original associated dataset to check for missing data, content misalignment, information tampering, and other issues, ensuring the authenticity and reliability of the file content.

[0081] S532. Organize the boundary division comparison results into a quality assessment file containing a list of qualified segments and a list of unqualified segments, and integrate the speaker metadata with the timestamp of this processing, processing device information, and alignment model version information to generate a metadata file; Specifically, the complete boundary segmentation comparison results are retrieved, categorized and sorted according to the segment judgment status, and standardized lists of qualified and unqualified segments are compiled. Various auxiliary annotation information within the lists is improved, and finally, a standardized quality assessment document is generated. The speaker metadata corresponding to this processing task is extracted, and the task execution timestamp, running device information, and currently used alignment model version information are collected simultaneously. Multiple types of information are integrated and arranged in a unified manner to generate a matching metadata file. After the file is generated, the completeness of the content and field format of the two files are checked in turn, and the data relationship between the files is verified based on the unique identifier of the speech segment.

[0082] It should be explained that the core specifications for document preparation and information integration are as follows: Quality assessment document organization standards: Strictly distinguish between two types of segment lists based on boundary delineation results. Each entry retains key information such as segment identifier, assessment conclusion, and anomaly type to ensure traceability and accessibility of assessment information. Metadata integration standards: Speaker metadata, task time information, equipment information, and model version information are categorized and summarized, with each type of information arranged in a partitioned manner to ensure clear organization and prevent missing or incorrect information. File format standards: Standardize the storage format, field structure, and naming rules for quality assessment documents and metadata files to adapt to system read / write and long-term archiving requirements. Association verification standards: Use the unique identifier of the speech segment as the association link to verify the matching of corresponding entries in the two files, ensuring the relevance and consistency of the entire data set.

[0083] S533. A globally unique file number is generated using a globally unique identifier generation algorithm, and the speech file, layered text file, phoneme boundary file, quality assessment file and metadata file are packaged into a standardized archive unit.

[0084] Specifically, the system invokes a preset globally unique identifier generation algorithm to generate a unique globally unique file number for the entire set of processed data; it then sequentially collects the voice files, layered text files, phoneme boundary files, quality assessment files, and metadata files produced in this process; according to the system's established archive directory structure and packaging standards, it integrates and packages all files to form an independent and complete standardized archive unit; and it associates and marks the generated globally unique file number with this archive unit to complete the information binding in the packaging process.

[0085] It should be explained that the core specifications for generating file numbers and packaging files are as follows: Numbering Standards: Archive numbers are generated using a globally unique identifier algorithm, ensuring each number is unique across the entire archived data, preventing duplication and errors, and serving as the core identifier for archive retrieval. File Collection Standards: All associated files related to this task are collected completely, with file names, content, and quantities verified one by one to ensure no files are missing or misplaced. Packaging Standards: Packaging follows a unified compression format and directory layout rules, with internal files categorized and stored, maintaining the original field structure and data integrity, suitable for long-term storage and retrieval. Number Binding Standards: The globally unique archive number is synchronously recorded in the archive unit identifier and metadata file, achieving a two-way association between the number and the archive unit, facilitating rapid retrieval, location, and access to archives. Archive Unit Standards: Packaged archive units have a unified format and standard structure, meeting the platform's archive management requirements and can be directly incorporated into the overall archive database for unified management.

[0086] S54. Verify the data integrity and format compliance of the standardized archiving unit. After verification, import it into the dialect digital archive and establish a search index based on the archive number, dialect area, speaker information and speech phenomenon tags.

[0087] Specifically, each standardized archival unit that has been packaged is verified. On the one hand, it is checked whether all files in the unit are complete and whether the content is complete, thus completing the data integrity verification. On the other hand, it is checked whether the format, directory structure and naming rules of each file meet the archiving requirements, thus completing the format compliance verification. After both verifications are passed, the archival unit is officially imported into the dialect digital archive for centralized storage. Then, the global archive number, dialect region, speaker information and speech phenomenon tags corresponding to the archival unit are extracted. A retrieval index is built layer by layer based on various key fields, and the correspondence between the index and the archival data is verified.

[0088] It should be explained that the core specifications for archive verification, database insertion, and index building are as follows: Data Integrity Verification Standards: Each archive unit is individually checked for missing files, incomplete content, and data corruption to ensure the completeness and usability of the entire archived data. Format Compliance Verification Standards: Following unified archival management standards, file formats, internal fields, directory layout, and encapsulation methods are checked. Archive units with formatting issues are not accepted into the database. Archive Entry Standards: Only standardized archive units that pass dual verification are included in the dialect digital archive database, stored systematically according to the database's storage rules, preserving the original relationships between files. Search Index Construction Standards: A multi-dimensional index system is built using the global archive number as the primary search term, and dialect region, speaker information, and speech phenomenon tags as secondary search terms to meet diverse search needs. Index Association Verification Standards: After the index is built, the relationship between entries is verified to ensure accurate matching between the index and the corresponding archive unit, guaranteeing the normal operation of subsequent search operations.

[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for aligning and archiving speech and text for the protection of Chinese dialects, characterized in that, Includes the following steps: S1. Obtain dialect speech data, speaker metadata, and layered text containing the international phoneme layer, and mark chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments in the international phoneme layer. S2. Construct a forced alignment model based on the connection time-series classification algorithm, and force-align dialect speech data with hierarchical text input to generate initial time boundaries for phonemes; S3. Preset the dialect phonological rule base and boundary division threshold, and perform targeted correction on the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak reading segments and literary and colloquial reading segments; S4. Compare the confidence level in the initial time boundary of the phonemes after directional correction with the boundary division threshold, and feed back the boundary division comparison result to update the dialect phonology rule base. S5. Associate the boundary division comparison results with the speaker's metadata and generate standardized archival units to import into the dialect digital archive. The preset dialect phonological rule base and boundary division threshold, and the targeted correction of the initial time boundary of phonemes based on the dialect phonological rule base, including chord segments, weak pronunciation segments, and literary and colloquial pronunciation segments, include the following steps: S31. A dialect phonological rule library is preset, which includes rules for merging consonants, rules for the duration of weak pronunciations, and rules for mapping literary and colloquial pronunciations. Boundary division thresholds are set, which include duration deviation thresholds, boundary offset thresholds, and confidence thresholds. S32. Based on the international phoneme layer in the layered text, locate the time intervals and phoneme sequences corresponding to all chord segments, weak pronunciation segments, and literary / colloquial pronunciation segments from the initial time boundary of the phonemes. S33. Based on the merging rules in the dialect phonology rule base, calculate the duration deviation value of adjacent phonemes in the merging segment, compare the deviation with the duration deviation threshold, and then merge the time boundaries of adjacent phonemes according to the deviation comparison result, while adjusting the overall duration to the standard duration range of the merging segment. S34. Based on the weak reading duration rules in the dialect phonology rule base, calculate the offset of the phoneme boundary in the weak reading segment, compare it with the boundary offset threshold, compress the phoneme duration to the weak reading standard duration range according to the offset comparison result, and correct the boundary offset at the same time. S35. Based on the literary and colloquial pronunciation mapping rules in the dialect phonological rule base, replace the corresponding phonemes in the literary and colloquial pronunciation segments and recalculate the alignment confidence. Then, compare the confidence with the confidence threshold and redivide the time boundary according to the confidence comparison result. S36. Integrate all corrected phoneme time boundaries with uncorrected phoneme time boundaries to generate directional corrected initial phoneme time boundaries.

2. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 1, characterized in that, The process of constructing a forced alignment model based on a connection-time classification algorithm and forcibly aligning dialect speech data with hierarchical text input to generate initial time boundaries for phonemes includes the following steps: S21. Perform acoustic preprocessing on the dialect speech data to obtain a standardized acoustic feature sequence, extract the international phoneme sequence from the layered text, and construct a phoneme label mapping table. S22. Construct a basic forced alignment network based on the connection time-series classification algorithm. Input the standardized acoustic feature sequence and the international phoneme sequence into the basic forced alignment network for supervised training. Iteratively adjust the model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model. S23. Convert the dialect speech data into the corresponding standardized acoustic feature sequence, input it into the dialect-specific forced alignment model for temporal reasoning, generate the start timestamp, end timestamp and alignment confidence score for each phoneme, and form the initial time boundary of the phoneme. The process of constructing a basic forced alignment network based on a connection-time classification algorithm, inputting standardized acoustic feature sequences and international phoneme sequences into the basic forced alignment network for supervised training, and iteratively adjusting the model weights and bias parameters through an independent validation set to obtain a dialect-specific forced alignment model includes the following steps: S221. Construct a basic forced alignment network structure based on the connection-based temporal classification algorithm, which includes a convolutional feature extraction layer, a bidirectional long short-term memory network layer, and a connection-based temporal classification output layer, and initialize the weights and bias parameters of each layer. S222. Divide the standardized acoustic feature sequences and international phoneme sequences into a training set and an independent validation set; S223. The connection-time classification loss function is used as the target loss function, and the adaptive moment estimation optimizer is used to supervise the training of the basic forced alignment network based on the training set, and the model weights and bias parameters are iteratively adjusted. S224. After each round of training, the phoneme boundary alignment accuracy of the model is calculated using an independent validation set. When the alignment accuracy no longer improves for several consecutive rounds, training is stopped, and the optimal model weights are saved to obtain a dialect-specific forced alignment model.

3. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 1, characterized in that, The step of comparing the confidence level in the initial time boundary of the directionally corrected phonemes with the boundary segmentation threshold, and then updating the boundary segmentation comparison result to the dialect phonology rule base includes the following steps: S41. Extract the alignment confidence, duration deviation and boundary offset of each phoneme in the initial time boundary of the directionally corrected phonemes. S42. The alignment confidence, duration deviation value and boundary offset are compared with the confidence threshold, duration deviation threshold and boundary offset threshold respectively, and the boundary division comparison results containing qualified segments and unqualified segments are generated. S43. For the unqualified segments in the boundary segmentation comparison results, identify the corresponding speech phenomenon type, and extract the average duration, average boundary offset and average confidence statistics of the unqualified segments of this type. S44. Based on average confidence statistics, the weighted coefficients of the rules for merging consonants, the rules for weak pronunciation duration, and the rules for mapping literary and colloquial pronunciations are adjusted to automatically iterate and update the dialect phonological rule base.

4. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 1, characterized in that, The process of associating the boundary segmentation comparison results with the speaker's metadata and generating standardized archival units for import into the dialect digital archive includes the following steps: S51. Obtain the dialect speech data, layered text, initial time boundaries of the phonemes after directional correction and boundary division comparison results during this processing, and retrieve the corresponding speaker metadata. S52. Using the unique identifier of the speech segment as an index, the dialect speech data, layered text, the initial time boundary of the corrected phonemes, the boundary division comparison results and the speaker's metadata are correlated to construct a multi-dimensional associated dataset. S53. In accordance with the unified format specifications of the dialect digital archive, the multi-dimensional associated dataset is encapsulated into a standardized archiving unit containing speech files, hierarchical text files, phoneme boundary files, quality assessment files and metadata files, and a globally unique archive number is generated. S54. Verify the data integrity and format compliance of the standardized archiving unit. After verification, import it into the dialect digital archive and establish a search index based on the archive number, dialect area, speaker information and speech phenomenon tags.

5. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 1, characterized in that, The method of merging phonemes based on the dialect phonology rule base, calculating the duration deviation value of adjacent phonemes in a chord segment, comparing the deviation with a duration deviation threshold, merging the time boundaries of adjacent phonemes according to the deviation comparison result, and adjusting the overall duration to the standard duration range of the chord includes the following steps: S331. Extract the adjacent phoneme sequences corresponding to the chord segment and their respective start timestamps and end timestamps, retrieve the standard duration range of the chord corresponding to the chord from the dialect phonology rule base, calculate the sum of the individual durations of the adjacent phonemes, obtain the actual total duration of the chord segment, and calculate the difference between the actual total duration and the median of the standard duration range of the chord to obtain the duration deviation value. S332. Compare the duration deviation value with the duration deviation threshold to generate a deviation comparison result. Based on the deviation comparison result, merge the time boundaries of adjacent phonemes, set the start timestamp of the chorus segment to the start timestamp of the first phoneme, set the end timestamp to the end timestamp of the last phoneme, and then adjust the overall duration to within the standard duration range of the chorus.

6. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 3, characterized in that, The automatic iterative update of the dialect phonological rule base, based on average confidence statistics, involves adjusting the weighted coefficients of the rules for merging consonants, weak pronunciation duration, and literary / colloquial pronunciation mapping, and includes the following steps: S441. Based on the type of speech phenomenon corresponding to the unqualified segments, the statistical data of average duration, average boundary offset and average confidence are divided into statistical data of chords, statistical data of weak readings and statistical data of literary and colloquial pronunciations. S442. Based on the average duration and average confidence of the statistical data of phonological combinations, weak pronunciations, and literary / colloquial pronunciations, the gradient descent method is used to calculate the weighting coefficient adjustment of the phonological combination rules, weak pronunciation duration rules, and literary / colloquial pronunciation mapping rules. The weighting coefficients of the corresponding rules are updated using each adjustment, and the dialect phonological rule base is automatically iteratively updated.

7. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 4, characterized in that, The process of encapsulating multi-dimensional correlated datasets into standardized archival units containing speech files, hierarchical text files, phoneme boundary files, quality assessment files, and metadata files, and generating globally unique archive numbers, according to the unified format specifications of the dialect digital archive, includes the following steps: S531. According to the unified format specifications of the dialect digital archive, the dialect speech data in the multi-dimensional associated dataset is converted into standard digital audio format to generate speech files, the hierarchical text is converted into structured markup language format to generate hierarchical text files, and the initial time boundaries of the directionally corrected phonemes are converted into key-value pair format to generate phoneme boundary files. S532. Organize the boundary division comparison results into a quality assessment file containing a list of qualified segments and a list of unqualified segments, and integrate the speaker metadata with the timestamp of this processing, processing device information, and alignment model version information to generate a metadata file; S533. A globally unique file number is generated using a globally unique identifier generation algorithm, and the speech file, layered text file, phoneme boundary file, quality assessment file and metadata file are packaged into a standardized archive unit.

8. The method for aligning and archiving speech and text for the protection of Chinese dialects according to claim 6, characterized in that, The formula for calculating the weighted coefficient adjustment amount of the rules for merging consonants, weak pronunciation duration, and literary / colloquial pronunciation mapping using the gradient descent method is as follows: , In the formula, This is the adjustment amount for the weighting coefficients in the current iteration round t; For the current iteration round t, the weighted coefficient values ​​for each group corresponding to the merging rule of consonant sounds, the weak reading duration rule, and the literary and colloquial reading mapping rule are: The learning rate; For the loss function L with the current weighting coefficients gradient at; The momentum coefficient; The weighting coefficients are those from the previous iteration round t-1; And loss function The formula is: , In the formula, N is the total number of unqualified segments in the boundary segmentation comparison results; For the i-th non-compliant segment, the current weighting coefficient Alignment confidence; is the duration error penalty coefficient; M is the total number of harmonic segments and weak reading segments involved in the directional correction; Let j be the actual duration of the j-th segment under the current rule; This refers to the standard duration of the harmonic or the standard duration of the weak pronunciation corresponding to this segment.

Citation Information

Patent Citations

  • Voice evaluation method and device, equipment and storage medium

    CN117542346A

  • Cross-platform calling method and system for corpus training library

    CN121011171A