An intelligent voice-based english spelling phonetic sequence matching method
By segmenting English words into letter clusters and processing speech fragment sequences, combined with anchor domain unit comparison, the problem of matching phonological letter clusters with silent letter clusters in English phonics was solved, enabling detailed analysis and accurate feedback on learners' reading process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-14
AI Technical Summary
Existing speech sequence matching methods have difficulty accurately distinguishing between phonical and silent letter clusters when processing English phonics, resulting in unclear speech segment classification, inaccurate identification of silent letter clusters, difficulty in merging short-tailed segments, and insufficient detail in expressing unmatched results.
By extracting target English words and learner speech, letter cluster segmentation and preprocessing are performed to generate letter cluster sequences and speech segment sequences. Then, by using field mapping and anchor domain unit comparison, a letter cluster anchor domain synchronization chain is established. Character label consistency comparison and merging judgment are performed to generate spelling-speech matching results.
It achieves a clear correspondence between speech segments and letter clusters during learners' spelling process, can distinguish between phonical letter clusters and silent letter clusters, clearly expresses short-tailed segments and pauses, and improves the accuracy of spelling diagnosis and pronunciation feedback.
Smart Images

Figure CN122392573A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent speech recognition technology, specifically to a method for matching English spelling speech sequences based on intelligent speech recognition. Background Technology
[0002] In the field of intelligent speech recognition, the collection, recognition, and analysis of learners' speech have been increasingly applied to scenarios such as language learning, oral assessment, and pronunciation correction. For English learning, English words are typically composed of multiple letters or letter combinations, with different letter clusters corresponding to different pronunciations, and some letters are not pronounced in a word. Therefore, simply recognizing the entire word's speech is insufficient to accurately reflect the learner's pronunciation at specific phonics units. To facilitate the analysis of the correspondence between learners' pronunciation and target English words, it is necessary to break down the target word into letter clusters and match the learner's pronunciation segments with these letter clusters in chronological order. This involves English phonics speech sequence matching. This type of processing essentially belongs to speech sequence matching, that is, establishing positional relationships, attribution relationships, and matching results between a preset text sequence and the actual speech sequence.
[0003] Existing speech sequence matching methods typically focus on the overall comparison of whole words, phonemes, or continuous speech segments, with limited structural processing for the coexistence of "pronounced letter clusters" and "silent letter clusters" in English phonics. When learners pronounce words aloud, short ending sounds, pauses, mispronunciations, or omissions can easily lead to misalignment between speech segments and letter clusters. For example, a short-tailed speech segment might belong to the final sound of a preceding pronounced letter cluster, but during sequential comparison, it might be misjudged as belonging to a subsequent letter cluster. Similarly, when the subsequent letter cluster is a silent letter cluster, traditional matching methods might still attempt to assign it a speech segment, resulting in an unreasonable correspondence between the silent letter cluster and the actual pronounced segment. Therefore, existing methods suffer from problems such as unclear speech segment attribution, inaccurate identification of silent letter clusters, difficulty in merging short-tailed segments, and insufficient detail in expressing unmatched results when handling letter cluster-level phonics matching. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an English phonics speech sequence matching method based on intelligent speech, which solves the problems mentioned in the background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for matching English phonics speech sequences based on intelligent speech, comprising: Extract the target English words and the intelligent speech input by the learner for the target English words, segment the target English words into letter clusters to obtain letter cluster sequences, and distinguish the pronunciation letter clusters and silence letter clusters of the intelligent speech based on pronunciation tags. Then, preprocess the intelligent speech to generate a sequence of speech segments. The letter cluster sequence is mapped to generate several anchor domain units, and a letter cluster pointer is established to perform character label consistency comparison. The speech segments are written into the current anchor slot or adjacent slot respectively, and the writing position of the speech segments is recorded synchronously to form a letter cluster anchor domain synchronization chain. The adjacent slot records are read according to the arrangement order of the anchor domain units in the letter cluster anchor domain synchronization chain, and the merging determination is performed based on the segment morphology label of the speech segment in the adjacent slot and the letter cluster type of the next anchor domain unit. After the merging determination is completed, a letter cluster matching record is generated for each anchor domain unit in turn. Based on the writing status of this anchor slot, the letter cluster type and the merging mark, the spelling speech matching result is determined as this anchor matching, tail merging matching, unmatched pronunciation letter cluster or matched silent letter cluster.
[0006] Preferably, the target English word and the intelligent speech input by the learner for the target English word are extracted, and the target English word is segmented into letter clusters according to the spelling order in the target English word. Each segmented letter cluster is assigned a letter cluster number to obtain a letter cluster sequence. The letter cluster sequence includes at least one letter cluster, and each letter cluster includes a letter cluster number, letter cluster content, standard pronunciation label and letter cluster type. When the standard pronunciation label of a letter cluster is not empty, the letter cluster type of the current letter cluster is marked as a pronounced letter cluster; when the standard pronunciation label of a letter cluster is empty, the letter cluster type of the current letter cluster is marked as a silent letter cluster.
[0007] Preferably, the intelligent speech is preprocessed according to the learner's reading time sequence, and the speech frames of the preprocessed intelligent speech are extracted. The preprocessing includes removing the silent speech segments at the beginning and end of the intelligent speech and retaining the spoken speech segments of the intelligent speech. The preprocessed intelligent speech is divided into frames according to the set frame length to obtain several speech frames. When there is a non-speech frame sequence between two consecutive speech frame sequences, the speech interval corresponding to the non-speech frame sequence is determined as the pause interval, and the boundary position between the pause interval and the previous speech frame sequence and the boundary position between the pause interval and the next speech frame sequence are determined as the pause boundary. If there is a pause boundary between adjacent speech frames, the current pause boundary is used as the segmentation position of the two adjacent speech segments. The speech segments are segmented according to the segmentation position, and a segment sequence number is assigned to each segmented speech segment to obtain a speech segment sequence. The speech segment sequence includes at least one speech segment. Each speech segment includes a speech segment number, a pronunciation label to be tested, a segment sequence number, and a segment morphology label. The segment morphology label includes the main speech segment and the short tail segment.
[0008] Preferably, the letter cluster sequence is written into the letter cluster record according to the arrangement order of the letter clusters in the letter cluster sequence, and field mapping is performed on each letter cluster record to generate several anchor domain units. Then, the multiple anchor domain units are arranged in sequence according to the arrangement order of the anchor domain unit numbers to form an initial anchor domain chain. Field mappings include: Write the letter cluster number from the letter cluster record into the anchor domain cell number field of the anchor domain cell; Write the letter cluster number in the letter cluster record into the corresponding letter cluster number field of the anchor domain unit; Write the letter cluster content from the letter cluster record into the corresponding letter cluster content field of the anchor domain unit; Write the standard pronunciation tag from the letter cluster record into the corresponding standard pronunciation tag field of the anchor domain unit; Write the letter cluster type from the letter cluster record into the corresponding letter cluster type field of the anchor field unit; Create the anchor slot field in the anchor domain unit and initialize the anchor slot field to an empty slot; Create an adjacent slot field in the anchor domain unit and initialize the adjacent slot field to an empty slot.
[0009] Preferably, the audio segment sequence is read and arranged in ascending order of segment number; An anchor domain unit index table is generated based on the initial anchor domain chain, and the first anchor domain unit to be written is determined according to the smallest anchor domain unit number in the anchor domain unit index table. The index record corresponding to the first anchor domain unit to be written is used as the initial pointer record to establish a letter cluster pointer. The letter cluster pointer includes the current anchor domain unit number, the current anchor domain unit record address, the current anchor slot address, the adjacent slot address, and the backward index. Read the phonetic tag to be tested for the current speech segment and the standard phonetic tag in the anchor unit currently pointed to by the letter cluster pointer. Perform a character tag consistency comparison between the phonetic tag to be tested and the corresponding standard phonetic tag: output a consistent result when the two are completely consistent, and output an inconsistent result when the two are not completely consistent. When a consistent output result is obtained, the speech segment number of the current speech segment is written into the current anchor slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the current anchor slot; if the anchor unit currently pointed to by the letter cluster pointer has a backward index, the letter cluster pointer is moved to the anchor unit pointed to by the backward index. When an inconsistent output result is obtained, the speech segment number of the current speech segment is written into the adjacent slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the adjacent slot; wherein, the letter cluster pointer is kept in the current anchor unit; if the corresponding letter cluster type of the anchor unit currently pointed to by the letter cluster pointer is a silent letter cluster, then the corresponding standard pronunciation label of the current anchor unit is an empty label. The initial anchor domain chain that has been written with the voice segment number is determined as the letter cluster anchor domain synchronization chain.
[0010] Preferably, anchor domain units are selected as the current anchor domain units according to their arrangement order in the letter cluster anchor domain synchronization chain, and the adjacent slot records of the current anchor domain units are read. When the adjacent slot record of the current anchor domain unit is not empty, read the speech segment number in the adjacent slot record, and read the segment morphology label of the corresponding speech segment from the speech segment sequence according to the speech segment number; When the adjacent slot record of the current anchor domain unit is empty, keep the current anchor slot record and adjacent slot record of the current anchor domain unit unchanged, and read the adjacent slot record of the next anchor domain unit in the sorted order.
[0011] Preferably, the letter cluster type corresponding to the letter cluster of the next anchor unit is read, and a merging determination is performed on the speech segments in the adjacent slot records, including: When the segment morphology label of a speech segment is a short-tailed segment, and the letter cluster type of the letter cluster corresponding to the next anchor unit is a silent letter cluster, the speech segment number is deleted from the adjacent slot record of the current anchor unit. At the same time, the speech segment number is written into the current anchor slot record of the current anchor unit, and the merging mark of the speech segment is written as the tail merged segment, while keeping the anchor unit corresponding to the speech segment as the current anchor unit. When the segment morphology label of a speech segment is not a short-tailed segment or the letter cluster type of the corresponding letter cluster of the next anchor unit is not a silent letter cluster, the writing position of the speech segment number in the adjacent slot record of the current anchor unit is retained, the merging mark of the speech segment is written as a boundary-retained segment, and the speech segment number is not written to the current anchor slot record of the current anchor unit. When there is no subsequent anchor unit in the current anchor unit, the merge mark of the speech segment in the adjacent slot record of the current anchor unit is written as a boundary-preserving segment.
[0012] Preferably, after the merging determination is completed, each anchor domain unit is read sequentially and a letter cluster matching record is generated for each anchor domain unit; wherein, each letter cluster matching record includes letter cluster number, letter cluster content, standard pronunciation label, letter cluster type, the number of the speech segment matched by this anchor, the number of the adjacent preserved speech segment, and the spelling speech matching result.
[0013] Preferably, the current anchor slot field and the adjacent slot field of the current anchor domain element are read; When the current anchor slot field contains a speech segment number, the current speech segment number is written into the speech segment number matched by this anchor in the spelling speech matching result, and the matching result label is determined based on whether the current speech segment number has a merging mark for the tail merged segment: If the current speech segment number has a merge marker for the tail-merged segment, then the speech matching result is recorded as a tail-merged match; If the current speech segment number does not have a merge marker for the tail-merged segment, then the spelling speech matching result is recorded as this anchor match; When the voice segment number is not recorded in this anchor slot field, the corresponding letter cluster type field of the current anchor domain unit is read: If the corresponding letter cluster type field is a pronunciation letter cluster, then the number of the speech segment matched by this anchor will be left blank, and the spelling speech matching result will be recorded as a pronunciation letter cluster not matched; If the corresponding letter cluster type field is a silent letter cluster, then the number of the matched speech segment in this anchor will be left blank, and the spelling speech matching result will be recorded as a silent letter cluster match.
[0014] Preferably, when the adjacent slot field contains a speech segment number and the current speech segment number has a boundary-preserving segment merging mark, the current speech segment number is written into the adjacent-preserving speech segment number of the spelling speech matching result; When the adjacent slot field does not record the speech segment number, the adjacent reserved speech segment number of the spelling speech matching result is written as empty.
[0015] This invention provides a method for matching English phonics speech sequences based on intelligent speech. It has the following beneficial effects: (1) This method extracts the target English word and the intelligent speech input by the learner for the target English word. The scheme segments the target English word into letter clusters, forming a letter cluster sequence including letter cluster number, letter cluster content, standard pronunciation label, and letter cluster type. Based on whether the standard pronunciation label is empty, the letter cluster type is divided into pronunciation letter clusters and silence letter clusters. At the same time, the scheme preprocesses, frames, and segments the intelligent speech according to the learner's reading time sequence, generating a speech segment sequence including speech segment number, pronunciation label to be tested, segment sequence number, and segment morphology label. Thus, the spelling structure of the target English word and the learner's actual reading speech are respectively organized into data sequences with clear fields, undertaking tasks such as word spelling unit division, speech segment extraction, pronunciation label recording, and segment morphology marking, so that subsequent matching can be carried out around the correspondence between letter clusters and speech segments.
[0016] (2) This method writes the letter cluster sequence into a letter cluster record and maps the letter cluster number, letter cluster content, standard pronunciation label, and letter cluster type in the letter cluster record to fields. This scheme generates an anchor domain unit including the anchor domain unit number, the corresponding letter cluster number, the corresponding letter cluster content, the corresponding standard pronunciation label, the corresponding letter cluster type, the current anchor slot field, and the adjacent slot field. It further establishes a letter cluster pointer including the current anchor domain unit number, the current anchor domain unit record address, the current anchor slot address, the adjacent slot address, and the backward index. Compared with the technical means of comparing each item according to the order of speech segments, this scheme can write the speech segment number into the current anchor slot field when the pronunciation label to be tested is completely consistent with the corresponding standard pronunciation label; when the two are not completely consistent, the speech segment number is written into the adjacent slot field, and the speech segment writing position is recorded synchronously. In this way, there is not only a matching relationship between the speech segment and the letter cluster, but also the boundary relationship related to the adjacent anchor domain unit can be preserved, so that short-tailed segments, misread segments, and segments after pauses have clear data recording positions.
[0017] (3) In the merging determination and matching result generation stages, this method reads the adjacent slot record of the current anchor unit according to the arrangement order of the anchor units in the letter cluster anchor domain synchronization chain, and combines the segment morphology label corresponding to the speech segment number in the adjacent slot record with the letter cluster type of the letter cluster corresponding to the next anchor unit to perform merging determination on the speech segments. When the segment morphology label of the speech segment is a short-tailed segment and the letter cluster type of the letter cluster corresponding to the next anchor unit is a silent letter cluster, the speech segment number is merged from the adjacent slot record into the current anchor slot record, and the merging mark of the speech segment is written as the tail merged segment; if this condition is not met, the merging mark of the speech segment is written as the boundary preserved segment. Finally, this method generates a letter cluster matching record for each anchor unit, including the letter cluster number, letter cluster content, standard pronunciation label, letter cluster type, the speech segment number matched by this anchor, the adjacent preserved speech segment number, and the spelling speech matching result, and outputs results such as the current anchor matching, the tail merged matching, the pronunciation letter cluster not matched, and the silent letter cluster matching. Compared to methods that easily misjudge silent letter clusters as missed readings or short-tailed segments as the pronunciation of the next letter cluster, this solution can more clearly express the correspondence between the learner's actual pronunciation and the spelling structure of the target English word, facilitating subsequent spelling diagnosis, pronunciation feedback, and learning record analysis. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the steps of an English phonics speech sequence matching method based on intelligent speech according to the present invention. Figure 2 This is a logic block diagram for a method of matching English spelling speech sequences based on intelligent speech according to the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1
[0021] Please see Figure 1 This invention provides a method for matching English phonics speech sequences based on intelligent speech. To achieve the above objectives, this invention is implemented through the following technical solution: including: Extract the target English words and the intelligent speech input by the learner for the target English words, segment the target English words into letter clusters to obtain letter cluster sequences, and distinguish the pronunciation letter clusters and silence letter clusters of the intelligent speech based on pronunciation tags. Then, preprocess the intelligent speech to generate a sequence of speech segments. The letter cluster sequence is mapped to generate several anchor domain units, and a letter cluster pointer is established to perform character label consistency comparison. The speech segments are written into the current anchor slot or adjacent slot respectively, and the writing position of the speech segments is recorded synchronously to form a letter cluster anchor domain synchronization chain. The adjacent slot records are read according to the arrangement order of the anchor domain units in the letter cluster anchor domain synchronization chain, and the merging determination is performed based on the segment morphology label of the speech segment in the adjacent slot and the letter cluster type of the next anchor domain unit. After the merging determination is completed, a letter cluster matching record is generated for each anchor domain unit in turn. Based on the writing status of this anchor slot, the letter cluster type and the merging mark, the spelling speech matching result is determined as this anchor matching, tail merging matching, unmatched pronunciation letter cluster or matched silent letter cluster.
[0022] In this embodiment, taking an English phonics practice scenario as an example, the learner is shown the shape of the target word to be spelled. First, the target word is broken down into multiple letter clusters according to phonics rules, forming a sequence of letter clusters: the first letter cluster is sh, corresponding to the standard phoneme subsequence / ʃ / ; the second letter cluster is a, corresponding to the standard phoneme subsequence / eɪ / ; the third letter cluster is p, corresponding to the standard phoneme subsequence / p / ; the fourth letter cluster is e, corresponding to an empty standard phoneme subsequence, belonging to the silent letter cluster. Each letter cluster serves as a letter cluster anchor point, forming a letter cluster anchor point chain in the order sh / a / p / e. After the learner reads the target word, the learner's speech is converted into a sequence of phonemes to be tested, for example, the sequence of phonemes to be tested is / ʃ / , / eɪ / , / p / , / ə / , where the last phoneme to be tested, / ə / , is a speech fragment produced by the learner mispronouncing the silent letter cluster 'e'.
[0023] During the matching process, instead of directly performing whole-word matching between the entire sequence of phonemes to be tested and the entire target word to be spelled, local speech sequence matching is performed on a letter cluster anchor point basis. Specifically, the letter cluster pointer first points to the letter cluster anchor point sh, and only reads the relevant phonemes within the range of the current letter cluster anchor point sh and its adjacent letter cluster anchor point a. The phoneme to be tested / ʃ / is compared with the standard phoneme subsequence / ʃ / corresponding to the letter cluster sh for character tag consistency. If they match, the speech segment number corresponding to the phoneme to be tested is written into the current anchor slot field of the letter cluster sh, and the letter cluster pointer is moved to the next letter cluster anchor point a. Subsequently, the phoneme to be tested / eɪ / is compared with the standard phoneme subsequence / eɪ / corresponding to the letter cluster a. If they match, the speech segment number is written into the current anchor slot field of the letter cluster a. Next, the phoneme to be tested / p / is compared with the standard phoneme subsequence / p / corresponding to the letter cluster p. If they match, the speech segment number is written into the current anchor slot field of the letter cluster p.
[0024] When the phoneme / ə / is read, the letter cluster pointer already points to the letter cluster anchor point e. The standard phoneme subsequence of letter cluster e is empty, and the letter cluster type is a silent letter cluster. Therefore, the phoneme / ə / and the letter cluster e do not constitute a local anchor match. At this point, the entire word is not immediately judged as a pronunciation error. Instead, based on local matching rules, the speech segment number corresponding to the phoneme is written into the adjacent slot field corresponding to the adjacent boundary, and a merging determination is performed in conjunction with the segment morphology label of the speech segment. If the segment morphology label of the speech segment is a short-tailed segment, and the letter cluster type corresponding to the next letter cluster anchor point is a silent letter cluster, then the speech segment can be treated as the tail-merging segment of the previous letter cluster p, and marked as a tail-merging match in the letter cluster matching record. If the speech segment is not a short-tailed segment, but rather / ə / clearly pronounced by the learner, then it is retained as a boundary-preserving segment to indicate that the learner produced an additional pronunciation of the silent letter cluster e. This process can provide results for specific letter clusters, such as sh anchor matching, a anchor matching, p anchor matching or tail merging matching, e silent letter cluster matching or boundary preservation anomaly, so that the spelling analysis can locate the specific letter cluster, rather than just giving a rough conclusion of whether the whole word is correct or incorrect.
[0025] Example 2
[0026] Please refer to Figure 2 Specifically: extract the target English words and the intelligent speech input by the learner for the target English words, and divide the target English words into letter clusters according to the spelling order in the target English words, and assign a letter cluster number to each segmented letter cluster to obtain a letter cluster sequence, wherein the letter cluster sequence includes at least one letter cluster, and each letter cluster includes a letter cluster number, letter cluster content, standard pronunciation label and letter cluster type; The letter cluster number is used to identify the current letter cluster's position in the letter cluster sequence; the letter cluster content is used to represent a spelling unit consisting of one or more letters; the standard pronunciation tag is used to represent the spelling pronunciation of the current letter cluster in the target English word; the letter cluster type includes phonical letter clusters and silent letter clusters; When the standard pronunciation label of a letter cluster is not empty, the letter cluster type of the current letter cluster is marked as a pronounced letter cluster; when the standard pronunciation label of a letter cluster is empty, the letter cluster type of the current letter cluster is marked as a silent letter cluster.
[0027] The intelligent speech is preprocessed according to the learner's reading time sequence, and the speech frames of the preprocessed intelligent speech are extracted. The preprocessing includes removing the silent speech segments at the beginning and end of the intelligent speech and retaining the spoken speech segments of the intelligent speech. The preprocessed intelligent speech is divided into frames according to the set frame length to obtain several speech frames. When there is a non-speech frame sequence between two consecutive speech frame sequences, the speech interval corresponding to the non-speech frame sequence is determined as the pause interval, and the boundary position between the pause interval and the previous speech frame sequence and the boundary position between the pause interval and the next speech frame sequence are determined as the pause boundary. If there is a pause boundary between adjacent speech frames, the current pause boundary is used as the segmentation position of the two adjacent speech segments. The speech segments are segmented according to the segmentation position, and a segment sequence number is assigned to each segmented speech segment to obtain a speech segment sequence. The speech segment sequence includes at least one speech segment. Each speech segment includes a speech segment number, a pronunciation label to be tested, a segment sequence number, and a segment morphology label. The segment morphology label includes the main speech segment and the short tail segment. The speech segment number is used to identify the current speech segment; the pronunciation label is used to indicate the pronunciation content obtained after speech recognition of the current speech segment; the segment sequence number is used to indicate the position of the current speech segment in the speech segment sequence.
[0028] In this embodiment, the target English word and the learner's input intelligent speech for the target English word are first extracted. The target English word is then segmented into letter clusters according to its spelling order. Each segmented letter cluster is assigned a letter cluster number, forming a letter cluster sequence that includes the letter cluster number, letter cluster content, standard pronunciation tag, and letter cluster type. The letter cluster number identifies the current letter cluster's position in the letter cluster sequence; the letter cluster content represents a spelling unit composed of one or more letters; the standard pronunciation tag indicates the spelling pronunciation of the current letter cluster in the target English word; and the letter cluster type includes phonic letter clusters and silent letter clusters. When the standard pronunciation tag is not empty, the current letter cluster type is marked as a phonic letter cluster; when the standard pronunciation tag is empty, the current letter cluster type is marked as a silent letter cluster. Subsequently, the intelligent speech is preprocessed according to the learner's reading time sequence, removing silent speech segments at the beginning and end of the intelligent speech and retaining the spoken speech segments. Finally, the preprocessed intelligent speech is divided into frames according to a set frame length, obtaining several speech frames. When there is a non-voiced frame sequence between two consecutive vocalized frame sequences, the speech interval corresponding to the non-voiced frame sequence is determined as the pause interval, and the boundary position between the pause interval and the preceding vocalized frame sequence, and the boundary position between the pause interval and the following vocalized frame sequence are determined as the pause boundary. If there is a pause boundary between adjacent speech frames, the current pause boundary is used as the segmentation position of the two adjacent vocalized speech segments, the vocalized speech segments are segmented, and a segment sequence number is assigned to each segmented vocalized speech segment to form a speech segment sequence including a speech segment number, a pronunciation label to be tested, a segment sequence number, and a segment morphology label. The segment morphology label includes the main vocalized segment and the short tail segment. Through the above implementation method, the target English words are organized into spelling structure data with letter cluster numbers, letter cluster content, standard pronunciation labels, and letter cluster types. The learner's intelligent speech is organized into speech temporal data with speech segment numbers, test pronunciation labels, segment sequence numbers, and segment morphology labels. This allows subsequent local speech sequence matching based on letter cluster anchors to use letter clusters as the basic analysis object, rather than making overall judgments based solely on the pronunciation of the whole word. Compared to existing whole word matching or continuous phoneme overall comparison methods, this implementation method can clearly distinguish between phonological letter clusters and silent letter clusters. It also preserves the pauses, short ending sounds, and segment positions during the learner's reading process based on pause boundaries, segment sequence numbers, and segment morphology labels. This makes the correspondence between speech segments and letter clusters more refined, facilitating subsequent judgments of anchor matching, tail merging matching, phonological letter cluster non-matching, and silent letter cluster matching. This achieves the goal of locating the learner's spelling pronunciation by letter clusters and makes spelling diagnosis, pronunciation feedback, and learning record analysis more closely aligned with the learner's actual reading process.
[0029] Example 3
[0030] Please refer to Figure 2 Specifically: based on the order of letter clusters in the letter cluster sequence, the letter cluster sequence is written into the letter cluster record, and field mapping is performed on each letter cluster record to generate several anchor domain units. Then, according to the order of the anchor domain unit numbers, multiple anchor domain units are arranged in sequence to form an initial anchor domain chain. Field mappings include: Write the letter cluster number from the letter cluster record into the anchor domain cell number field of the anchor domain cell; Write the letter cluster number in the letter cluster record into the corresponding letter cluster number field of the anchor domain unit; Write the letter cluster content from the letter cluster record into the corresponding letter cluster content field of the anchor domain unit; Write the standard pronunciation tag from the letter cluster record into the corresponding standard pronunciation tag field of the anchor domain unit; Write the letter cluster type from the letter cluster record into the corresponding letter cluster type field of the anchor field unit; Create the anchor slot field in the anchor domain unit and initialize the anchor slot field to an empty slot; Create an adjacent slot field in the anchor domain element and initialize the adjacent slot field to an empty slot; Among them, an empty slot refers to a slot record that has not yet been written with a voice segment number; this anchor slot is used to write the voice segment number that matches the letter cluster corresponding to the current anchor domain unit; the adjacent slot is used to write the voice segment number that has not entered this anchor slot and needs to retain the affiliation relationship between the current anchor domain unit and the next anchor domain unit.
[0031] Read the audio segment sequence and arrange them in ascending order of segment number; An anchor domain unit index table is generated based on the initial anchor domain chain, and the first anchor domain unit to be written is determined according to the smallest anchor domain unit number in the anchor domain unit index table. The index record corresponding to the first anchor domain unit to be written is used as the initial pointer record to establish a letter cluster pointer. The letter cluster pointer includes the current anchor domain unit number, the current anchor domain unit record address, the current anchor slot address, the adjacent slot address, and the backward index. Read the phonetic tag to be tested for the current speech segment and the standard phonetic tag in the anchor unit currently pointed to by the letter cluster pointer. Perform a character tag consistency comparison between the phonetic tag to be tested and the corresponding standard phonetic tag: output a consistent result when the two are completely consistent, and output an inconsistent result when the two are not completely consistent. When a consistent output result is obtained, the speech segment number of the current speech segment is written into the current anchor slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the current anchor slot; if the anchor unit currently pointed to by the letter cluster pointer has a backward index, the letter cluster pointer is moved to the anchor unit pointed to by the backward index. When an inconsistent output result is obtained, the speech segment number of the current speech segment is written into the adjacent slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the adjacent slot; wherein, the letter cluster pointer is kept in the current anchor unit; if the corresponding letter cluster type of the anchor unit currently pointed to by the letter cluster pointer is a silent letter cluster, then the corresponding standard pronunciation label of the current anchor unit is an empty label. The initial anchor domain chain that has been written with the voice segment number is determined as the letter cluster anchor domain synchronization chain.
[0032] In this embodiment, each letter cluster is first written into a letter cluster record according to the order of letter clusters in the letter cluster sequence. The letter cluster number, letter cluster content, standard pronunciation tag, and letter cluster type in the letter cluster record are mapped to the anchor unit number field, the corresponding letter cluster number field, the corresponding letter cluster content field, the corresponding standard pronunciation tag field, and the corresponding letter cluster type field, respectively. At the same time, an anchor slot field and an adjacent slot field are established in each anchor unit, and the anchor slot field and adjacent slot field are initialized to empty slots, thereby forming an initial anchor chain arranged according to the anchor unit number. Subsequently, the preprocessed speech segment sequence of the learner's intelligent speech is read and arranged in ascending order of segment sequence number. An anchor unit index table is generated based on the initial anchor chain. The first anchor unit to be written is determined according to the smallest anchor unit number in the anchor unit index table, and a letter cluster pointer is established with the index record corresponding to the first anchor unit to be written. The letter cluster pointer simultaneously records the current anchor unit number, the current anchor unit record address, the anchor slot address, the adjacent slot address, and the backward index. During the speech segment writing process, the test pronunciation tag of the current speech segment is read, and the corresponding standard pronunciation tag in the anchor unit currently pointed to by the letter cluster pointer is read. A character tag consistency comparison is performed between the test pronunciation tag and the corresponding standard pronunciation tag. When the two are completely consistent, the speech segment number of the current speech segment is written to the current anchor slot field of the current anchor unit, and the writing position of the current speech segment is recorded as the current anchor slot. Then, the letter cluster pointer is moved according to the backward index. When the two are not completely consistent, the speech segment number of the current speech segment is written to the adjacent slot field of the current anchor unit, and the writing position of the current speech segment is recorded as the adjacent slot. At the same time, the letter cluster pointer is kept pointing to the current anchor unit. If the corresponding letter cluster type of the current anchor unit is a silence letter cluster, the corresponding standard pronunciation tag is an empty tag, so that the silence letter cluster is not treated as a normal pronunciation letter cluster. Through the above implementation method, the anchor matching relationship, boundary preservation relationship, and pointer movement state between speech segments and letter clusters are all written into the same chain structure, forming a letter cluster anchor domain synchronization chain. This letter cluster anchor domain synchronization chain records both the speech segment numbers that directly match the current letter cluster and the speech segment numbers that have not yet been assigned to the current anchor slot but have a belonging relationship with the current anchor domain unit and the next anchor domain unit. Compared with the method of sequentially comparing only whole words or whole segments of speech, this implementation method can place normal pronunciation, suspected misreading, short-tailed segments, and silent letter cluster boundary segments in the learner's reading into corresponding fields, achieving the purpose of organizing speech segments by letter cluster anchor points, saving speech segments to be judged by local boundaries, and controlling the matching progress by pointer state. This facilitates subsequent merging and judgment based on segment morphology labels and the letter cluster type of the next anchor domain unit, and makes the final generated letter cluster matching record more suitable for expressing the learner's spelling state on specific letter clusters.
[0033] Example 4
[0034] Please refer to Figure 2 Specifically: based on the order of the anchor domain units in the letter cluster anchor domain synchronization chain, the anchor domain units are selected as the current anchor domain units in sequence, and the adjacent slot records of the current anchor domain units are read. When the adjacent slot record of the current anchor domain unit is not empty, read the speech segment number in the adjacent slot record, and read the segment morphology label of the corresponding speech segment from the speech segment sequence according to the speech segment number; When the adjacent slot record of the current anchor domain unit is empty, keep the current anchor slot record and adjacent slot record of the current anchor domain unit unchanged, and read the adjacent slot record of the next anchor domain unit in the sorted order.
[0035] Read the letter cluster type of the letter cluster corresponding to the next anchor unit, where the next anchor unit is the anchor unit that is after the current anchor unit in the letter cluster anchor synchronization chain and is adjacent to the current anchor unit. Perform a merge determination on the speech segments in the adjacent slot records, including: When the segment morphology label of a speech segment is a short-tailed segment, and the letter cluster type of the letter cluster corresponding to the next anchor unit is a silent letter cluster, the speech segment number is deleted from the adjacent slot record of the current anchor unit. At the same time, the speech segment number is written into the current anchor slot record of the current anchor unit, and the merging mark of the speech segment is written as the tail merged segment, while keeping the anchor unit corresponding to the speech segment as the current anchor unit. When the segment morphology label of a speech segment is not a short-tailed segment or the letter cluster type of the corresponding letter cluster of the next anchor unit is not a silent letter cluster, the writing position of the speech segment number in the adjacent slot record of the current anchor unit is retained, the merging mark of the speech segment is written as a boundary-retained segment, and the speech segment number is not written to the current anchor slot record of the current anchor unit. When there is no subsequent anchor unit in the current anchor unit, the merge mark of the speech segment in the adjacent slot record of the current anchor unit is written as a boundary-preserving segment.
[0036] In this embodiment, based on the order of anchor units in the letter cluster anchor domain synchronization chain, each anchor unit is selected as the current anchor unit, and the adjacent slot record of the current anchor unit is read. When the adjacent slot record of the current anchor unit is not empty, the speech segment number recorded in the adjacent slot record is read, and the segment morphology label of the corresponding speech segment is retrieved from the speech segment sequence according to the speech segment number. At the same time, the letter cluster type of the letter cluster corresponding to the next anchor unit is read. When the segment morphology label of the speech segment is a short-tailed segment, and the letter cluster type of the letter cluster corresponding to the next anchor unit is a silent letter cluster, it indicates that the speech segment is more likely to belong to the phonological letter cluster corresponding to the current anchor unit. The speech segment number is deleted from the adjacent slot record of the current anchor unit and written to the current anchor slot record, as it is not an independent pronunciation of the following silent letter cluster. Simultaneously, the merging mark of this speech segment is written as a tail-merged segment. When the segment morphology label of a speech segment is not a short-tailed segment, or the letter cluster type of the corresponding letter cluster of the following anchor unit is not a silent letter cluster, the writing position of the speech segment number in the adjacent slot record of the current anchor unit is retained, and the merging mark of this speech segment is written as a boundary-preserved segment. When the current anchor unit does not have a following anchor unit, the speech segment in the adjacent slot record of the current anchor unit is also marked as a boundary-preserved segment. Through the above processing, the current anchor slot record, adjacent slot record, segment morphology label, letter cluster type of the corresponding letter cluster of the following anchor unit, and merging mark can be jointly judged to distinguish short-tailed pronunciations, silent letter cluster boundaries, extra pronunciation segments, and boundary-preserved segments that occur during the learner's reading. The purpose of this implementation is to solve the problems in existing whole-word matching or sequential matching that easily misjudge short-tailed segments as pronunciations of the following letter clusters, misjudge silent letter clusters as missed pronunciations, and misjudge boundary residual speech as independent incorrect pronunciations, so as to clarify the attribution relationship between speech segments and letter clusters. Compared with the technical means of simply comparing the test pronunciation tag with the standard pronunciation tag item by item, this solution introduces adjacent slot recording and merging judgment in the letter cluster anchor domain synchronization chain, so that speech segments that do not directly enter the current anchor slot record can still retain their boundary information, and further classification judgment is made based on the segment morphology tag and the letter cluster type of the letter cluster corresponding to the next anchor domain unit.
[0037] Example 5
[0038] Please refer to Figure 2 Specifically: After completing the merging determination, each anchor domain unit is read sequentially and a letter cluster matching record is generated for each anchor domain unit; each letter cluster matching record includes the letter cluster number, letter cluster content, standard pronunciation label, letter cluster type, the number of the matched speech segment in this anchor, the number of the adjacent preserved speech segment, and the spelling speech matching result.
[0039] Read the current anchor slot field and the adjacent slot field of the current anchor domain element; When the current anchor slot field contains a speech segment number, the current speech segment number is written into the speech segment number matched by this anchor in the spelling speech matching result, and the matching result label is determined based on whether the current speech segment number has a merging mark for the tail merged segment: If the current speech segment number has a merge marker for the tail-merged segment, then the speech matching result is recorded as a tail-merged match; If the current speech segment number does not have a merge marker for the tail-merged segment, then the spelling speech matching result is recorded as this anchor match; When the voice segment number is not recorded in this anchor slot field, the corresponding letter cluster type field of the current anchor domain unit is read: If the corresponding letter cluster type field is a pronunciation letter cluster, then the number of the speech segment matched by this anchor will be left blank, and the spelling speech matching result will be recorded as a pronunciation letter cluster not matched; If the corresponding letter cluster type field is a silent letter cluster, then the number of the matched speech segment in this anchor will be left blank, and the spelling speech matching result will be recorded as a silent letter cluster match.
[0040] When the adjacent slot field contains a speech segment number and the current speech segment number has a boundary-preserving segment merging mark, the current speech segment number is written into the adjacent-preserving speech segment number of the spelling speech matching result. When the adjacent slot field does not record the speech segment number, the adjacent reserved speech segment number of the spelling speech matching result is written as empty.
[0041] In this embodiment, after the merging determination is completed, the anchor slot field and adjacent slot field of each anchor unit are read sequentially according to the arrangement order of the letter cluster numbers in the letter cluster anchor domain synchronization chain, and a letter cluster matching record is generated for each anchor unit. The letter cluster matching record includes the letter cluster number, letter cluster content, standard pronunciation tag, letter cluster type, the number of the matched speech segment in this anchor, the number of the adjacent retained speech segment, and the spelling speech matching result. If the current anchor slot field of the current anchor unit contains a speech segment number, then the speech segment number is written into the speech segment number matched by the current anchor, and it is further determined whether the speech segment number has a merging mark for the tail-merging segment; if it has a merging mark for the tail-merging segment, then the spelling speech matching result is recorded as a tail-merging match, indicating that although the speech segment was once at the boundary of the adjacent slot, after the segment morphology label and the letter cluster type of the next anchor unit are determined, it should belong to the current letter cluster; if it does not have a merging mark for the tail-merging segment, then the spelling speech matching result is recorded as the current anchor match, indicating that the learner's speech segment directly corresponds to the standard pronunciation label of the current letter cluster. If the current anchor slot field of the current anchor unit does not record a speech segment number, then the corresponding letter cluster type field of the current anchor unit is read. When the corresponding letter cluster type field is a phonological letter cluster, the speech segment number matched by this anchor is written as empty, and the spelling speech matching result is recorded as phonological letter cluster not matched, indicating that the learner did not read the speech content corresponding to the phonological letter cluster. When the corresponding letter cluster type field is a silent letter cluster, the speech segment number matched by this anchor is written as empty, and the spelling speech matching result is recorded as silent letter cluster matched, indicating that the letter cluster itself does not need a corresponding phoneme. For the adjacent slot field, if it records a speech segment number and the speech segment number has a boundary-preserving segment merging mark, then the speech segment number is written to the adjacent preserved speech segment number, used to preserve speech segments with boundary abnormalities, extra pronunciations, or not included in this anchor slot during the learner's reading. If the adjacent slot field does not record a speech segment number, then the adjacent preserved speech segment number is written as empty. Through the above processing, the write state, merge state, and letter cluster type in the letter cluster anchor domain synchronization chain can be transformed into structured phonics speech matching results. Compared with the recognition method that only outputs the whole word as correct or incorrect, this scheme can express states such as direct matching, tail merging, pronunciation loss, normal silence, and boundary preservation, so that the learner's phonics problem can be located to the specific letter cluster, which is convenient for subsequent pronunciation correction, phonics training feedback, and learning process recording and analysis.
[0042] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended technical solutions and their equivalents.
Claims
1. A method for matching English phonics speech sequences based on intelligent speech, characterized in that, include: Extract the target English words and the intelligent speech input by the learner for the target English words, segment the target English words into letter clusters to obtain letter cluster sequences, and distinguish the pronunciation letter clusters and silence letter clusters of the intelligent speech based on pronunciation tags. Then, preprocess the intelligent speech to generate a sequence of speech segments. The letter cluster sequence is mapped to generate several anchor domain units, and a letter cluster pointer is established to perform character label consistency comparison. The speech segments are written into the current anchor slot or adjacent slot respectively, and the writing position of the speech segments is recorded synchronously to form a letter cluster anchor domain synchronization chain. The adjacent slot records are read according to the arrangement order of the anchor domain units in the letter cluster anchor domain synchronization chain, and the merging determination is performed based on the segment morphology label of the speech segment in the adjacent slot and the letter cluster type of the next anchor domain unit. After the merging determination is completed, a letter cluster matching record is generated for each anchor domain unit in turn. Based on the writing status of this anchor slot, the letter cluster type and the merging mark, the spelling speech matching result is determined as this anchor matching, tail merging matching, unmatched pronunciation letter cluster or matched silent letter cluster.
2. The method for matching English phonics speech sequences based on intelligent speech according to claim 1, characterized in that, Extract the target English words and the intelligent speech input by the learner for the target English words, and segment the target English words into letter clusters according to the spelling order in the target English words. Assign a letter cluster number to each segmented letter cluster to obtain a letter cluster sequence. The letter cluster sequence includes at least one letter cluster, and each letter cluster includes a letter cluster number, letter cluster content, standard pronunciation label and letter cluster type. When the standard pronunciation label of a letter cluster is not empty, the letter cluster type of the current letter cluster is marked as a pronounced letter cluster; when the standard pronunciation label of a letter cluster is empty, the letter cluster type of the current letter cluster is marked as a silent letter cluster.
3. The method for matching English phonics speech sequences based on intelligent speech according to claim 2, characterized in that, The intelligent speech is preprocessed according to the learner's reading time sequence, and the speech frames of the preprocessed intelligent speech are extracted. The preprocessing includes removing the silent speech segments at the beginning and end of the intelligent speech and retaining the spoken speech segments of the intelligent speech. The preprocessed intelligent speech is divided into frames according to the set frame length to obtain several speech frames. When there is a non-speech frame sequence between two consecutive speech frame sequences, the speech interval corresponding to the non-speech frame sequence is determined as the pause interval, and the boundary position between the pause interval and the previous speech frame sequence and the boundary position between the pause interval and the next speech frame sequence are determined as the pause boundary. If there is a pause boundary between adjacent speech frames, the current pause boundary is used as the segmentation position of the two adjacent speech segments. The speech segments are segmented according to the segmentation position, and a segment sequence number is assigned to each segmented speech segment to obtain a speech segment sequence. The speech segment sequence includes at least one speech segment. Each speech segment includes a speech segment number, a pronunciation label to be tested, a segment sequence number, and a segment morphology label. The segment morphology label includes the main speech segment and the short tail segment.
4. The method for matching English phonics speech sequences based on intelligent speech according to claim 3, characterized in that, Based on the order of letter clusters in the letter cluster sequence, the letter cluster sequence is written into the letter cluster record, and field mapping is performed on each letter cluster record to generate several anchor domain units. Then, according to the order of the anchor domain unit numbers, the multiple anchor domain units are arranged in sequence to form the initial anchor domain chain. Field mappings include: Write the letter cluster number from the letter cluster record into the anchor domain cell number field of the anchor domain cell; Write the letter cluster number in the letter cluster record into the corresponding letter cluster number field of the anchor domain unit; Write the letter cluster content from the letter cluster record into the corresponding letter cluster content field of the anchor domain unit; Write the standard pronunciation tag from the letter cluster record into the corresponding standard pronunciation tag field of the anchor domain unit; Write the letter cluster type from the letter cluster record into the corresponding letter cluster type field of the anchor field unit; Create the anchor slot field in the anchor domain unit and initialize the anchor slot field to an empty slot; Create an adjacent slot field in the anchor domain unit and initialize the adjacent slot field to an empty slot.
5. The method for matching English phonics speech sequences based on intelligent speech according to claim 4, characterized in that, Read the audio segment sequence and arrange them in ascending order of segment number; An anchor domain unit index table is generated based on the initial anchor domain chain, and the first anchor domain unit to be written is determined according to the smallest anchor domain unit number in the anchor domain unit index table. The index record corresponding to the first anchor domain unit to be written is used as the initial pointer record to establish a letter cluster pointer. The letter cluster pointer includes the current anchor domain unit number, the current anchor domain unit record address, the current anchor slot address, the adjacent slot address, and the backward index. Read the phonetic tag to be tested for the current speech segment and the standard phonetic tag in the anchor unit currently pointed to by the letter cluster pointer. Perform a character tag consistency comparison between the phonetic tag to be tested and the corresponding standard phonetic tag: output a consistent result when the two are completely consistent, and output an inconsistent result when the two are not completely consistent. When a consistent output result is obtained, the speech segment number of the current speech segment is written into the current anchor slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the current anchor slot; if the anchor unit currently pointed to by the letter cluster pointer has a backward index, the letter cluster pointer is moved to the anchor unit pointed to by the backward index. When an inconsistent output result is obtained, the speech segment number of the current speech segment is written into the adjacent slot of the anchor unit currently pointed to by the letter cluster pointer, and the writing position of the current speech segment is recorded as the adjacent slot; wherein, the letter cluster pointer is kept in the current anchor unit; if the corresponding letter cluster type of the anchor unit currently pointed to by the letter cluster pointer is a silent letter cluster, then the corresponding standard pronunciation label of the current anchor unit is an empty label. The initial anchor domain chain that has been written with the voice segment number is determined as the letter cluster anchor domain synchronization chain.
6. The method for matching English phonics speech sequences based on intelligent speech according to claim 5, characterized in that, Based on the order of the anchor domain units in the letter cluster anchor domain synchronization chain, select the anchor domain units as the current anchor domain units in sequence, and read the adjacent slot records of the current anchor domain units. When the adjacent slot record of the current anchor domain unit is not empty, read the speech segment number in the adjacent slot record, and read the segment morphology label of the corresponding speech segment from the speech segment sequence according to the speech segment number; When the adjacent slot record of the current anchor domain unit is empty, keep the current anchor slot record and adjacent slot record of the current anchor domain unit unchanged, and read the adjacent slot record of the next anchor domain unit in the sorted order.
7. The method for matching English phonics speech sequences based on intelligent speech according to claim 6, characterized in that, Read the letter cluster type of the letter cluster corresponding to the next anchor domain unit, and perform a merge determination on the speech segments in the adjacent slot records, including: When the segment morphology label of a speech segment is a short-tailed segment, and the letter cluster type of the letter cluster corresponding to the next anchor unit is a silent letter cluster, the speech segment number is deleted from the adjacent slot record of the current anchor unit. At the same time, the speech segment number is written into the current anchor slot record of the current anchor unit, and the merging mark of the speech segment is written as the tail merged segment, while keeping the anchor unit corresponding to the speech segment as the current anchor unit. When the segment morphology label of a speech segment is not a short-tailed segment or the letter cluster type of the corresponding letter cluster of the next anchor unit is not a silent letter cluster, the writing position of the speech segment number in the adjacent slot record of the current anchor unit is retained, the merging mark of the speech segment is written as a boundary-retained segment, and the speech segment number is not written to the current anchor slot record of the current anchor unit. When there is no subsequent anchor unit in the current anchor unit, the merge mark of the speech segment in the adjacent slot record of the current anchor unit is written as a boundary-preserving segment.
8. The method for matching English phonics speech sequences based on intelligent speech according to claim 7, characterized in that, After the merging decision is completed, each anchor domain unit is read sequentially and a letter cluster matching record is generated for each anchor domain unit. Each letter cluster matching record includes the letter cluster number, letter cluster content, standard pronunciation label, letter cluster type, the number of the matched speech segment in this anchor, the number of the adjacent preserved speech segment, and the spelling speech matching result.
9. The method for matching English phonics speech sequences based on intelligent speech according to claim 8, characterized in that, Read the current anchor slot field and the adjacent slot field of the current anchor domain element; When the current anchor slot field contains a speech segment number, the current speech segment number is written into the speech segment number matched by this anchor in the spelling speech matching result, and the matching result label is determined based on whether the current speech segment number has a merging mark for the tail merged segment: If the current speech segment number has a merge marker for the tail-merged segment, then the speech matching result is recorded as a tail-merged match; If the current speech segment number does not have a merge marker for the tail-merged segment, then the speech matching result is recorded as this anchor match; When the voice segment number is not recorded in this anchor slot field, the corresponding letter cluster type field of the current anchor domain unit is read: If the corresponding letter cluster type field is a pronunciation letter cluster, then the number of the speech segment matched by this anchor will be left blank, and the spelling speech matching result will be recorded as a pronunciation letter cluster not matched; If the corresponding letter cluster type field is a silent letter cluster, then the number of the matched speech segment in this anchor will be left blank, and the spelling speech matching result will be recorded as a silent letter cluster match.
10. A method for matching English phonics speech sequences based on intelligent speech according to claim 9, characterized in that, When the adjacent slot field contains a speech segment number and the current speech segment number has a boundary-preserving segment merging mark, the current speech segment number is written into the adjacent-preserving speech segment number of the spelling speech matching result. When the adjacent slot field does not record the speech segment number, the adjacent reserved speech segment number of the spelling speech matching result is written as empty.