Speech recognition method based on speech speed perception air traffic control speech recognition model

CN122531407APending Publication Date: 2026-08-07FUJIAN SANQINGNIAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN SANQINGNIAO TECH CO LTD
Filing Date
2026-07-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

然而,在实际空中交通管制通信过程中,当说话人在短时间内语速显著加快时,相邻指令之间的停顿特征会明显减弱甚至消失,导致基于停顿或固定时间窗口的语音分段方法难以准确识别指令边界

Benefits of technology

本发明通过在语音时间序列中引入语速变化情况与停顿时长信息的联合刻画,并围绕语速突然加快区段构建连续变化记录,在时间维度上恢复被压缩的语音边界结构,使原本因语速变化导致难以区分的语音间隔得到重构,从而在后续处理过程中形成具有可区分间隔的边界候选区段;基于该边界候选区段对连续语音内容进行重新划分处理,使紧密衔接的语音片段能够按照调整后的时间间隔进行拆分,从而提升语音分段的准确性,并保证各语音片段之间具备清晰的时间间隔关系。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531407A_ABST
    Figure CN122531407A_ABST
Patent Text Reader

Abstract

The application discloses a speech recognition method based on a speech speed perception air traffic control voice recognition model, relates to the technical field of speech recognition, and comprises the following steps: collecting speech time sequences in a continuous call process, recording speech speed change conditions and pause duration information in combination with time advancement, and arranging continuous change records around sections in which the speech speed suddenly increases. The application constructs a continuous change record through joint analysis of speech speed change and pause duration, recovers compressed speech boundaries and forms boundary candidate sections, so that accurate division of continuous speech is realized. Meanwhile, through sequential arrangement and connection adjustment of speech segments and synchronous optimization of output rhythm, the semantic expression and time structure are kept consistent, and the content splicing and mixing problem caused by sudden speech speed change is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and more specifically to a speech recognition method based on an air traffic control speech recognition model with speech rate perception. Background Technology

[0002] Speech recognition based on speech rate perception refers to the speech recognition model used in air traffic control communication scenarios. Addressing the characteristics of speech rates, such as significant variations, uneven rhythms, and high noise interference in the speech of controllers and pilots, a dynamic perception mechanism for speech rate changes is introduced into the speech recognition process. By continuously analyzing the rate of change of the speech signal over time, a correlation between speech rate and speech features is established. Based on this, the speech segmentation method, feature extraction rhythm, and recognition matching process are adaptively adjusted. This allows the recognition process to be optimized synchronously with changes in speech rate, improving the accuracy and stability of recognition in rapid instructions, continuous conversations, and complex communication environments. Ultimately, this achieves highly reliable reconstruction and understanding of air traffic control speech content.

[0003] The existing technology has the following shortcomings: In existing technologies, air traffic control speech recognition systems typically delineate command boundaries based on pause features, energy changes, or rhythm information in the speech signal. However, in actual air traffic control communications, when a speaker significantly increases their speaking speed within a short period, the pause features between adjacent commands weaken or even disappear, making it difficult for speech segmentation methods based on pauses or fixed time windows to accurately identify command boundaries. In this situation, speech recognition systems are prone to misinterpreting multiple semantically independent commands as continuous speech content, leading to command splicing problems. Furthermore, this problem can cause confusion in the parameter correspondence between different operational commands, such as parameter mismatch between heading and altitude commands, resulting in command parsing errors and reducing the security and reliability of the air traffic control communication system.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a speech recognition method based on a speech rate-aware air traffic control speech recognition model to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a speech recognition method based on a speech rate-sensing air traffic control speech recognition model, comprising the following steps: The speech time series during continuous calls is collected, and the changes in speech rate and pause duration are recorded in combination with the time progression. The continuous change record is formed around the sections where the speech rate suddenly increases. Based on continuous change records, the pause duration in the section where the speech rate suddenly increases is adjusted, and the compressed pause segments are stretched in time according to the speech rate change to obtain candidate boundary segments with distinguishable intervals. Based on the candidate boundary segments, the continuous speech content is re-divided and processed. Closely connected speech segments are split according to the adjusted time interval to form a set of speech segments with clear interval relationships. Based on a set of speech segments, the sequential relationships between the segments are organized, and the connection points between segments are adjusted accordingly based on changes in speech rate to obtain a continuous and consistent semantic expression result. Based on the semantic expression results, the output rhythm in the speech recognition process is adjusted synchronously to keep the set of speech segments consistent with the semantic expression results during the time progression, thereby avoiding the problem of content splicing and mixing when the speech rate changes abruptly.

[0007] Preferably, the steps for characterizing the continuous correlation between speech rate changes and pause duration information in a speech time series and generating a continuous change record reflecting the speech boundary compression state are as follows: The system receives voice signals segment by segment during continuous calls to form a voice time sequence. It assigns start and end time markers to each voice segment and counts the number and duration of voice segments per unit time to obtain speech rate changes. It also identifies silent segments to obtain pause duration information. The speech rate changes are continuously tracked along the time progression direction, the segments where the speech rate suddenly increases are extracted, and the corresponding speech segments and pause duration information are extracted by combining the speech time series to form a corresponding set within the time range; The corresponding sets are integrated and processed by pairing and combining information on time location, speech rate changes and pause duration to form a continuous change trajectory, so as to express the correspondence between speech rate changes and pause duration. The continuous change trajectory is organized and arranged in chronological order to form a continuous change record, which reflects the compression state of the speech boundary during the time process.

[0008] Preferably, the extraction process of segments with sudden acceleration of speech rate is determined by combining the speech rate changes of adjacent time slices and synchronously associating pause duration information to limit the segment range; during the construction of continuous change trajectory, the speech rate changes and pause duration information are arranged in a one-to-one correspondence, thereby ensuring that the continuous change record can accurately reflect the boundary compression state in the speech time series.

[0009] Preferably, the temporal reconstruction of pause segments within a segment where the speech rate suddenly increases to form distinguishable boundary candidate segments includes the following steps: Extract the segments in the continuous change record that suddenly increase the speech rate, and read the speech rate change and pause duration information corresponding to each time node in the time progression order to construct a pause segment sequence containing time position and pause duration; The pause segments in the pause segment sequence are refined, and the pause segments are divided into continuous sub-segments, while maintaining the correspondence between the speech rate changes and the pause duration information in each sub-segment; Adjust the duration of continuous sub-segments, extend the time of each sub-segment according to the changes in speech rate, and re-splice them to form pause segments, thus constructing a time-stretched structure with non-uniform distribution. The pause segments are mapped to the corresponding time axis of continuous change records, and then arranged and segmented in combination with the speech time series to complete the speech time series division and generate boundary candidate segments.

[0010] Preferably, the duration of consecutive sub-segments is allocated and adjusted during the time extension process to ensure that the speech rate changes corresponding to consecutive sub-segments are consistent, and the pause segments formed by splicing maintain the continuity of intervals in the time axis. Furthermore, the boundary positions of adjacent speech segments are defined by dividing segments, thereby stabilizing the time distribution structure of the boundary candidate segments.

[0011] Preferably, the process of splitting continuous speech content into a set of speech segments with clear intervals by time intervals includes the following steps: The start and end times of corresponding boundary candidate segments in continuous speech content are marked, and the segments are divided in chronological order to form candidate speech segments with time ranges. The connection positions between candidate speech segments are refined, the time interval information corresponding to the boundary candidate segments is read and applied to the adjacent positions, and the closely connected speech segments are split to obtain multiple speech sub-segments. Organize the temporal distribution of speech segments, mark the start and end times of each speech segment and arrange them in chronological order, and match the time interval information corresponding to the boundary candidate segments. Standardize the temporal relationships between speech segments, combine time interval information to make unified adjustments and complete the arrangement, forming a set of speech segments with clear interval relationships.

[0012] Preferably, the start and end time positions of each speech segment are constrained according to the time interval relationship of the speech segment set, so as to maintain the consistency between the time interval between adjacent speech segments and the correspondence between the boundary candidate segments, and to limit the arrangement order of the speech segments on the time axis, so as to ensure that the speech segment set forms a continuous and clearly spaced distribution structure during the time progression.

[0013] Preferably, the semantic expression result is constructed by sequentially organizing the speech segment set and combining it with changes in speech rate, including the following steps: Extract the start and end times of each speech segment from the speech segment set, arrange them in chronological order to form a sequence of segments with a temporal relationship, and record the time interval information between adjacent speech segments. The speech rate changes at corresponding time positions are introduced from the continuous change records. The speech rate changes are matched with the segment sequence, and the connection positions between adjacent speech segments are adjusted based on the speech rate changes to form an adjusted segment sequence. The adjusted segment sequence is renumbered according to time sequence, and the interval relationship between adjacent speech segments is bound in combination with the speech rate change, forming a segment sequence structure with time and speech rate correlation. The speech segments in the splicing sequence are continuously combined in chronological order, while maintaining the time interval between segments and the change in speech rate to obtain a continuous and consistent semantic expression result.

[0014] Preferably, the time interval between adjacent speech segments in the segment sequence is constrained and adjusted in combination with the speech rate variation, and the start and end time positions of the speech segments are simultaneously limited, so that the distribution of each speech segment on the time axis in the segment sequence is consistent with the speech rate variation, thereby maintaining the continuity and consistency of the semantic expression results.

[0015] Preferably, adjusting the rhythm of speech recognition output to maintain consistency between the speech segment set and the semantic expression result includes the following steps: The semantic expression results are associated with a set of speech segments. The semantic expression results are divided into semantic units and the start and end time positions are marked. At the same time, the corresponding speech segments are matched to form a time-aligned structure. The speech recognition results are divided into output sequences, and output segments are formed by combining the time range of semantic units, while keeping the time interval between output segments consistent with the set of speech segments; By introducing speech rate changes from continuous change records, the time intervals between output segments are adjusted to form an output rhythm structure that matches the speech rate changes. The speech recognition output process is controlled by advancing the output according to the time alignment structure and keeping the set of speech segments consistent with the semantic expression result, thereby avoiding the problem of content splicing and mixing.

[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention introduces a joint characterization of speech rate changes and pause duration information into the speech time series, and constructs continuous change records around segments where the speech rate suddenly accelerates. This restores the compressed speech boundary structure in the time dimension, reconstructing speech intervals that were originally difficult to distinguish due to speech rate changes. This results in the formation of candidate boundary segments with distinguishable intervals in subsequent processing. Based on these candidate boundary segments, continuous speech content is re-divided, allowing closely connected speech segments to be split according to adjusted time intervals. This improves the accuracy of speech segmentation and ensures clear time interval relationships between each speech segment.

[0017] After forming a set of speech segments, this invention organizes the sequential relationships between the segments and adjusts the connection positions of the segments according to changes in speech rate, so that the speech segments form a continuous and consistent semantic expression result on the time axis. At the same time, it further adjusts the output rhythm in the speech recognition process synchronously according to the semantic expression result, so that the set of speech segments and the semantic expression result remain consistent in the time process, thereby avoiding the splicing or mixing of semantic content under sudden changes in speech rate, and ensuring the consistency of speech recognition results in both temporal and semantic structure. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0019] Figure 1 This is a flowchart of the speech recognition method based on the speech rate perception-based air traffic control speech recognition model of the present invention.

[0020] Figure 2 This is a flowchart illustrating the process of generating candidate boundary sections in this invention.

[0021] Figure 3 This is a diagram illustrating the time extension process of continuous sub-segments in this invention. Detailed Implementation

[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0023] This invention provides, for example Figure 1 The speech recognition method based on the speech rate-aware air traffic control speech recognition model shown includes the following steps: The speech time series during continuous calls is collected, and the changes in speech rate and pause duration are recorded in combination with the time progression. The continuous change record is formed around the sections where the speech rate suddenly increases. In the continuous communication environment of air traffic control, voice information unfolds continuously along the time axis, with variations in speech rate and pause distribution intertwined, easily leading to boundary compression in local segments. By meticulously characterizing the speech rate variations and pause durations within the speech time series and constructing a continuous record reflecting the boundary compression state, a reliable foundation can be provided for subsequent processing. The specific implementation steps are as follows: After the continuous call begins, the voice signal is received segment by segment along the time progression direction. The received voice segments are arranged in chronological order to form a voice time series. Each voice segment in the voice time series is assigned a start time mark and an end time mark to determine the position of the voice segment in the overall time axis. Around the time interval between adjacent voice segments, the frequency of speech in the voice signal within the continuous time range is statistically analyzed segment by segment. By recording the number of voice segments appearing per unit time and the changes in the duration of voice segments, a specific description of the speech rate changes is formed. At the same time, silent segments are identified one by one in the voice time series. The start time and end time of each silent segment are marked, and the duration of the segment is recorded to form pause duration information. On the same time axis, the speech rate changes and pause duration information are arranged synchronously so that any time position corresponds to a set of speech rate change data and pause duration data, thereby establishing the correspondence between the voice time series, speech rate changes, and pause duration information.

[0024] The speech rate changes are continuously tracked along the time progression direction. The speech rate change data are compared point by point between adjacent time slices. When the number of speech segments per unit time in a certain time range increases continuously and the interval between speech segments decreases, the time range is marked as the segment where the speech rate suddenly increases, and the start and end time positions of the segment are recorded.

[0025] Around the marked segments where the speech rate suddenly increases, all speech segments corresponding to the segment are extracted from the speech time series, and the pause duration information within the segment is extracted synchronously, so that the speech rate change and pause duration information form a complete corresponding set within the same time range. By expanding this set in chronological order, the continuous distribution of speech rate change and pause duration information within the segment where the speech rate suddenly increases is obtained, providing a complete data source for the construction of subsequent continuous change records.

[0026] For the speech rate changes and pause duration information in the section where the speech rate suddenly increases, the data arranged continuously on the time axis are integrated and processed. The speech rate change data and pause duration data corresponding to each time position are paired and combined to form a three-dimensional data sequence containing time position, speech rate change value and pause duration value. The three-dimensional data sequence is then concatenated along the time sequence so that the speech rate change and pause duration information form a continuous change trajectory in the time dimension.

[0027] In this trajectory, the correspondence between speech rate change values ​​and pause duration values ​​is uniformly expressed, so that the positions where the speech rate change value increases and the positions where the pause duration value decreases are in a one-to-one correspondence on the time axis, thereby constructing a process of gradual compression of the speech boundary in the overall time series; by organizing this process as a whole, the continuous change trajectory is transformed into a structured continuous change record, so that the record can fully reflect the compression state of the speech boundary in the process of time progression.

[0028] The generated continuous change records are standardized and organized by arranging the time nodes on the timeline sequentially. At each time node, the speech rate change value and pause duration value are recorded simultaneously, ensuring that the continuous change records remain continuous and consistent in the direction of time progression. For sections where the speech rate suddenly increases, the corresponding time range is clearly marked in the continuous change records, and the correspondence between the speech rate change value and the pause duration value is maintained within this range without any break, thus ensuring that the state of speech boundary compression is fully presented in the continuous change records. By uniformly arranging the data of each time node in the continuous change records, any time position can reflect the corresponding speech rate change and pause duration information, thereby providing a clear basis for subsequent adjustment processing of pause segments and ensuring that the continuous change records can continuously reflect the state of speech boundary compression during subsequent retrieval.

[0029] Based on continuous change records, the pause duration in the section where the speech rate suddenly increases is adjusted, and the compressed pause segments are stretched in time according to the speech rate change to obtain candidate boundary segments with distinguishable intervals. Given that the continuous change record has been constructed, a temporal correspondence has been established between the pause duration and the change in speech rate within the segment where the speech rate suddenly increases. By making in-depth use of this correspondence and performing time-stretching processing on the compressed pause segments, the distinguishability of speech boundaries on the time axis can be restored, thereby obtaining candidate boundary segments with distinguishable intervals. The specific implementation steps are as follows: Around the section of speech rate that suddenly increases in the continuous change record, all time nodes in the section are traversed one by one in the order of time progression. At each time node, the corresponding speech rate change and pause duration information are read, and a pause segment sequence is constructed with time nodes as the unit. In the pause segment sequence, the start time position, end time position and corresponding pause duration of each pause segment are uniformly organized.

[0030] Simultaneously, by combining the speech rate changes at corresponding time points in the continuous change records, the distribution density of pause segments in the areas where the speech rate suddenly increases is described. This ensures that the pause segment sequence not only contains time location information but also the distribution relationship corresponding to the speech rate changes, thus providing a complete data foundation for subsequent pause duration adjustment processing.

[0031] After obtaining the sequence of pause segments and the corresponding changes in speech rate, the original pause duration of each pause segment is further refined and divided into several continuous sub-segments according to the direction of time progression. Within each sub-segment, the correspondence between the changes in speech rate and the pause duration information is maintained. In this way, the pause segments that originally existed as a whole are transformed into a refined structure composed of multiple continuous sub-segments.

[0032] In this refined structure, the time proportion of each sub-segment is redistributed based on the changes in speech rate, so that the differences in speech rate at different time positions can be reflected in the internal structure of the pause segment. This creates a close relationship between the internal time distribution of the pause segment and the changes in speech rate, providing a segmentation basis for time stretching processing.

[0033] For the pause segments that have been refined and segmented, the duration of each segment is extended based on the speech rate changes at the corresponding time positions. During the extension process, the trend of speech rate changes in the continuous change record is used as the basis. Sub-segments corresponding to time positions with higher speech rate changes are allocated a larger time extension, while sub-segments corresponding to time positions with relatively stable speech rate changes maintain the original time proportion. This creates a non-uniform time stretching structure within the overall pause segment range.

[0034] In this structure, each sub-segment is expanded and then reassembled into a complete pause segment, so that the expanded pause segment occupies a more defined interval position on the time axis and forms a time segment with a clear separating function in the speech time series, thereby realizing the recovery processing of the compressed pause segment.

[0035] After completing the time stretching process for all pause segments, the expanded pause segments are remapped back to the time axis corresponding to the continuous change records, and rearranged according to the time order with the speech segments in the speech time series, so that each expanded pause segment is located between adjacent speech segments, thereby forming multiple segmented sections with clear intervals in the overall time series.

[0036] Around these segmented sections, the temporal relationships between adjacent speech segments are uniformly organized, so that the expanded pause segments form stable boundary markers on the time axis. Based on these boundary markers, the speech time series is divided into several speech segments with independent time intervals, thereby obtaining candidate boundary segments with distinguishable intervals, providing a clear basis for subsequent re-segmentation of speech content.

[0037] Based on the candidate boundary segments, the continuous speech content is re-divided and processed. Closely connected speech segments are split according to the adjusted time interval to form a set of speech segments with clear interval relationships. With clearly defined boundary candidate segments, continuous speech content has operable segmentation criteria on the time axis. By meticulously utilizing time intervals, the originally continuously distributed speech content can be re-divided, and closely connected speech segments can be clearly separated in the time dimension, thereby constructing a set of speech segments with stable interval relationships. The specific implementation method is as follows: The continuous speech content is traversed segment by segment along the time progression direction. During the traversal, the start and end time positions of the boundary candidate segments are used as the segmentation reference points to mark and divide the continuous speech content on the time axis. A separation mark is set at each boundary candidate segment so that the speech content is divided into the preceding speech part and the following speech part at the corresponding position.

[0038] After marking all candidate boundary segments, the continuous speech content is divided into multiple candidate speech segments in chronological order. Each candidate speech segment corresponds to a specific time range, and all speech data contained within that time range is recorded. At the same time, the continuity of the time order is maintained during the recording process, so that all candidate speech segments are arranged in the order of the original speech time sequence, thus forming a preliminary division result.

[0039] Around the already formed candidate speech segments, the connection positions between adjacent candidate speech segments are refined. During the processing, the corresponding time interval value in the boundary candidate segment is read and used as a splitting reference scale. This value is then applied to the boundary positions of adjacent candidate speech segments to further split the originally closely connected speech segments.

[0040] In the specific processing, near the end time position of each candidate speech segment, new boundary positions are defined forward and backward according to the time interval value, so that the speech segment is decomposed into multiple continuous but independent speech sub-segments within this range. At the same time, it is ensured that each speech sub-segment contains complete speech waveform information and corresponds to a clear time interval, so that the speech content that was originally continuously distributed on the time axis is split into multiple speech sub-segment structures with interval distribution.

[0041] For the segmented speech data, the distribution of each speech segment on the time axis is uniformly organized. During the organization process, the start and end time positions of each speech segment are re-marked and arranged in chronological order so that all speech segments form a continuous and orderly arrangement.

[0042] During the arrangement process, the time interval between adjacent speech segments is matched one by one with the corresponding time interval in the boundary candidate segments, so that the interval distribution between speech segments can directly reflect the separation information provided by the boundary candidate segments. At the same time, the speech content corresponding to each speech segment is completely preserved, so that the speech segments have clear time boundaries while maintaining the original speech features, thus forming a prototype of a speech segment set with clear interval relationships.

[0043] Based on the completed set of speech segments, the temporal relationships within the set are standardized. During the standardization process, the time intervals between each adjacent speech segment are made consistent, so that the intervals between all speech segments are consistent with the corresponding boundary candidate segments. The entire set of speech segments is calibrated on the time axis, so that the speech segments maintain the continuity of temporal order and form a stable interval distribution between adjacent segments during the arrangement process.

[0044] This standardized processing ensures that each speech segment in the speech segment set has a clear start and end time boundary and can be clearly distinguished from adjacent speech segments as time progresses. This completes the re-division of continuous speech content and ultimately yields a speech segment set with clear interval relationships, providing a complete and orderly data foundation for the subsequent construction of semantic expression results.

[0045] Based on a set of speech segments, the sequential relationships between the segments are organized, and the connection points between segments are adjusted accordingly based on changes in speech rate to obtain a continuous and consistent semantic expression result. Given a set of speech segments with clear time intervals, by organizing the order of these segments and adjusting the transitions based on speech rate variations, the scattered speech segments can form a continuous and unified expression in both time and semantic dimensions, resulting in a consistent semantic expression. The specific implementation method is as follows: Based on the arrangement order of the speech segments on the timeline, the start and end times of each speech segment are extracted one by one, and all speech segments are arranged sequentially according to the direction of time progression. During the arrangement process, the existing temporal order of the speech segments is maintained. At the same time, the time intervals between adjacent speech segments are recorded synchronously, so that each speech segment not only has its own time range, but also has a temporal relationship with the preceding and following speech segments.

[0046] After chronological arrangement, the set of speech segments is transformed into a sequence of segments with continuous time markers, so that the sequence of segments can fully reflect the unfolding process of speech content on the timeline, thus providing a unified basis for the subsequent fine-tuning of the relationship between segments.

[0047] For a sequence of audio segments arranged chronologically, a detailed analysis of the connection relationships between segments is conducted based on changes in speech rate. During the analysis, speech rate changes at corresponding time positions in continuous change records are incorporated into the segment sequence, ensuring that each audio segment's position on the timeline corresponds to a set of speech rate change data. Using this speech rate change data, the connection positions between adjacent audio segments are reassessed. For audio segments located in areas of sudden speech rate acceleration, the time boundaries at their connection positions are redefined based on the speech rate changes, clearly separating segments that might have overlapped or been closely connected on the timeline. Simultaneously, for audio segments located in areas of gradual speech rate change, their original time connection relationships are maintained without shifting, thus creating an overall connection structure in the segment sequence that matches the changes in speech rate.

[0048] After adjusting the segment connection positions, the relationships between segments in the sequence are uniformly organized. During this process, the direction of time progression is the main thread; each speech segment is renumbered according to the adjusted time boundaries, and the relationships between segments are determined based on the numbering order, creating a logically continuous speech flow structure. The interval information between speech segments is then linked to changes in speech rate, ensuring that the interval between each pair of adjacent segments reflects the speech rate changes within their corresponding time range. This provides a unified expression for the segment sequence in both the time and speech rate dimensions, offering a stable foundation for constructing the semantic representation.

[0049] Based on the organized sequence of audio segments, the segments are sequentially assembled according to their adjusted order. During this assembly, the segments are combined continuously in chronological order, maintaining consistent spacing and speech rate variations between segments. This ensures the assembled audio content presents a continuous and reasonably spaced structure on the timeline. This method transforms a collection of scattered audio segments into continuously expressed audio content, maintaining semantic coherence and temporal consistency throughout the expression process. This results in a consistent semantic expression, providing a stable foundation for subsequent speech recognition output.

[0050] Based on the semantic expression results, the output rhythm in the speech recognition process is adjusted synchronously to keep the set of speech segments consistent with the semantic expression results during the time progression, thereby avoiding the problem of content splicing and mixing when the speech rate changes abruptly; Given that the semantic expression result has already established a temporal correspondence with the speech segment set, by refining the control of the output rhythm during the speech recognition process, the consistency between the speech segment set and the semantic expression result can be maintained during the time progression, ensuring that the speech recognition result maintains a stable expression structure even when the speech rate changes. The specific implementation method is as follows: Based on the way the semantic expression results unfold along the time axis, each speech segment in the speech segment set is associated one by one. During the processing, the semantic expression results are divided into multiple consecutively arranged semantic units according to the time progression, and the start and end time positions are marked for each semantic unit. At the same time, the speech segments in the speech segment set that correspond to the time range are searched, and the semantic units are bound to the speech segments one by one. After completing the binding of all semantic units to speech segments, the semantic expression results and speech segment sets are arranged uniformly on the same time axis, so that any time position has both semantic content and speech segment information, thus forming a complete time alignment structure and providing a clear reference for subsequent output rhythm adjustment.

[0051] Based on the established time alignment structure, the output sequence in the speech recognition process is segmented. During the segmentation process, the time range of the semantic unit is used as the dividing criterion to split the continuously output speech recognition results into multiple output segments. Each output segment corresponds to a semantic unit, and the time interval between the output segments is kept consistent with the speech segment set.

[0052] In specific processing, the start and end times of each output segment are clearly defined, so that the output segment occupies the same time range as the corresponding speech segment on the time axis. At the same time, time gaps consistent with the intervals of speech segments are inserted between adjacent output segments, so that the output rhythm is arranged in a consistent time structure with the set of speech segments.

[0053] The output rhythm is further refined by introducing speech rate variation. During the processing, speech rate variation data corresponding to each time position in the continuous variation record is read, and the speech rate variation data is aligned point by point with the time range of the semantic unit so that each semantic unit corresponds to a complete speech rate variation interval.

[0054] Around the range of speech rate changes, the time intervals between output segments are redistributed. In segments where the speech rate suddenly increases, the intervals between adjacent output segments are expanded, so that the output segments that were originally closely arranged on the time axis are clearly separated, while maintaining the consistency between the time range within each output segment and the corresponding semantic unit without deviation. In segments where the speech rate changes smoothly, the time intervals between the original output segments are maintained unchanged, so that the overall output rhythm can be consistent with the speech rate changes and form a stable rhythm distribution as time progresses.

[0055] After adjusting the overall output rhythm, the speech recognition results are continuously controlled for output. During the output process, the output is strictly carried out according to the time alignment structure, so that the output time range of each semantic unit is consistent with the time interval of the corresponding speech segment. At the same time, the time interval between adjacent semantic units and the speech rate changes are kept consistent. Throughout the output process, the correspondence between the speech segment set and the semantic expression result is maintained, so that the semantic content unfolds continuously in a predetermined order and forms a clear separation structure on the time axis. This avoids splicing or mixing of semantic content when the speech rate changes, and ultimately achieves consistent expression of speech recognition results in the time and semantic dimensions.

[0056] This invention introduces a joint characterization of speech rate changes and pause duration information into the speech time series, and constructs continuous change records around segments where the speech rate suddenly accelerates. This restores the compressed speech boundary structure in the time dimension, reconstructing speech intervals that were originally difficult to distinguish due to speech rate changes. This results in the formation of candidate boundary segments with distinguishable intervals in subsequent processing. Based on these candidate boundary segments, continuous speech content is re-divided, allowing closely connected speech segments to be split according to adjusted time intervals. This improves the accuracy of speech segmentation and ensures clear time interval relationships between each speech segment.

[0057] After forming a set of speech segments, this invention organizes the sequential relationships between the segments and adjusts the connection positions of the segments according to changes in speech rate, so that the speech segments form a continuous and consistent semantic expression result on the time axis. At the same time, it further adjusts the output rhythm in the speech recognition process synchronously according to the semantic expression result, so that the set of speech segments and the semantic expression result remain consistent in the time process, thereby avoiding the splicing or mixing of semantic content under sudden changes in speech rate, and ensuring the consistency of speech recognition results in both temporal and semantic structure.

[0058] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A speech recognition method based on a speech rate perception-based air traffic control speech recognition model, characterized in that, Includes the following steps: The speech time series during continuous calls is collected, and the changes in speech rate and pause duration are recorded in combination with the time progression. The continuous change record is formed around the sections where the speech rate suddenly increases. Based on continuous change records, the pause duration in the section where the speech rate suddenly increases is adjusted, and the compressed pause segments are stretched in time according to the speech rate changes to obtain the boundary candidate segments. Based on the candidate boundary segments, the continuous speech content is re-divided and processed. Closely connected speech segments are split according to the adjusted time interval to form a set of speech segments. Based on a set of speech segments, the sequential relationships between the segments are arranged, and the connection points between segments are adjusted accordingly based on changes in speech rate to obtain semantic expression results; Based on the semantic expression results, the output rhythm in the speech recognition process is adjusted synchronously to keep the set of speech segments consistent with the semantic expression results as time progresses.

2. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 1, characterized in that, The steps to characterize the continuous correlation between speech rate changes and pause duration information in speech time series and generate continuous change records reflecting the speech boundary compression state are as follows: The system receives voice signals segment by segment during continuous calls to form a voice time sequence. It assigns start and end time markers to each voice segment and counts the number and duration of voice segments per unit time to obtain speech rate changes. It also identifies silent segments to obtain pause duration information. The speech rate changes are continuously tracked along the time progression direction, the segments where the speech rate suddenly increases are extracted, and the corresponding speech segments and pause duration information are extracted by combining the speech time series to form a corresponding set within the time range; The corresponding sets are integrated and processed by pairing and combining information on time location, speech rate changes and pause duration to form a continuous change trajectory. The continuous change trajectory is organized and arranged in chronological order to form a continuous change record, which reflects the compression state of the speech boundary during the time process.

3. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 2, characterized in that, The extraction process for segments with sudden increases in speech rate is determined by combining the speech rate changes in adjacent time slices and simultaneously associating pause duration information to limit the segment range; during the construction of continuous change trajectories, the speech rate changes and pause duration information are arranged in a one-to-one correspondence to ensure that the continuous change records can accurately reflect the boundary compression state in the speech time series.

4. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 2, characterized in that, Temporal reconstruction is performed on pauses within segments where the speech rate suddenly increases to form distinguishable boundary candidate segments, including the following steps: Extract the segments in the continuous change record that suddenly increase the speech rate, and read the speech rate changes and pause duration information corresponding to each time point in the time progression sequence to construct a pause segment sequence; The pause segments in the pause segment sequence are refined, and the pause segments are divided into continuous sub-segments, while maintaining the correspondence between the speech rate changes and the pause duration information in each sub-segment; Adjust the duration of consecutive sub-segments, extend the time of each sub-segment according to the changes in speech rate, and re-splice them to form pause segments, thus constructing a time-stretched structure. The pause segments are mapped to the corresponding time axis of continuous change records, and then arranged and segmented in combination with the speech time series to complete the speech time series division and generate boundary candidate segments.

5. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 4, characterized in that, During the time extension process, the duration of consecutive sub-segments is allocated and adjusted to ensure that the speech rate changes corresponding to consecutive sub-segments are consistent, and the pause segments formed by splicing maintain the continuity of intervals in the time axis. The boundary positions of adjacent speech segments are defined by dividing segments.

6. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 4, characterized in that, The process of splitting continuous speech content into a set of speech segments with clear intervals includes the following steps: The start and end times of corresponding boundary candidate segments in continuous speech content are marked, and the segments are divided in chronological order to form candidate speech segments. The connection positions between candidate speech segments are refined, the time interval information corresponding to the boundary candidate segments is read and applied to the adjacent positions, and the closely connected speech segments are split to obtain multiple speech sub-segments. Organize the temporal distribution of speech segments, mark the start and end times of each speech segment and arrange them in chronological order, and match the time interval information corresponding to the boundary candidate segments. Standardize the temporal relationships between speech segments, combine time interval information to make unified adjustments and complete the arrangement, forming a set of speech segments with clear interval relationships.

7. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 6, characterized in that, Based on the time interval relationship of the speech segment set, the start and end time positions of each speech sub-segment are constrained to maintain the consistency between the time interval between adjacent speech sub-segments and the correspondence between the boundary candidate segments. The arrangement order of the speech sub-segments on the time axis is also limited to ensure that the speech segment set forms a continuous and clearly spaced distribution structure as time progresses.

8. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 6, characterized in that, The semantic representation of a collection of speech segments is constructed by sequentially organizing the segments and considering variations in speech rate, including the following steps: Extract the start and end times of each speech segment from the speech segment set, arrange them in chronological order to form a sequence of segments with a temporal relationship, and record the time interval information between adjacent speech segments. The speech rate changes at corresponding time positions are introduced from the continuous change records. The speech rate changes are matched with the segment sequence, and the connection positions between adjacent speech segments are adjusted based on the speech rate changes to form an adjusted segment sequence. The adjusted segment sequence is reorganized, each speech segment is renumbered according to time sequence, and the interval relationship between adjacent speech segments is bound in combination with the speech rate change to form a segment sequence structure; The speech segments in the splicing sequence are continuously combined in chronological order, while maintaining the time interval between segments and the change in speech rate to obtain a continuous and consistent semantic expression result.

9. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 8, characterized in that, The time interval between adjacent speech segments in the segment sequence is constrained and adjusted based on the speech rate variation, and the start and end time positions of the speech segments are simultaneously limited, so that the distribution of each speech segment on the time axis in the segment sequence is consistent with the speech rate variation.

10. The speech recognition method based on the speech rate perception-based air traffic control speech recognition model according to claim 8, characterized in that, Adjusting the rhythm of speech recognition output to maintain consistency between the speech segment set and the semantic expression result includes the following steps: The semantic expression results are associated with a set of speech segments. The semantic expression results are divided into semantic units and the start and end time positions are marked. At the same time, the corresponding speech segments are matched to form a time-aligned structure. The speech recognition results are divided into output sequences, and output segments are formed by combining the time range of semantic units, while keeping the time interval between output segments consistent with the set of speech segments; By introducing speech rate changes from continuous change records, the time intervals between output segments are adjusted to form an output rhythm structure that matches the speech rate changes. Control the output process of speech recognition results, advance the output according to the time alignment structure, and keep the set of speech segments consistent with the semantic expression results.