AI-based multilingual adaptive recognition method
By generating a rhythmic base script and dynamically adjusting the time window, the semantic fragmentation problem of multilingual speech recognition systems in dynamic speech speed scenarios is solved, achieving speech recognition continuity and accuracy under different speech speed conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUJIAN SANQINGNIAO TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-04-17
AI Technical Summary
Existing multilingual speech recognition systems struggle to adapt to varying speech rates in dynamic speech scenarios, leading to premature overlap or segmentation of speech segments. This results in misaligned recognition content, semantic breaks, and disordered sentence output, affecting the accuracy and continuity of recognition.
By collecting rhythmic features and pause nodes in continuous speech streams, a rhythmic draft is generated. The step size and coverage width of the time window are dynamically adjusted, and semantically misaligned segments and broken phrases are corrected in real time to ensure that the speech recognition process maintains temporal continuity and semantic integrity under different speech rates.
It effectively avoids the problem of time superposition or segmentation of speech segments under speech rate fluctuations, maintains clear semantic boundaries of recognition results, and improves the continuity and overall accuracy of multilingual adaptive recognition.
Smart Images

Figure CN121708902B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language information processing technology, and more specifically to an AI-based multilingual adaptive recognition method. Background Technology
[0002] AI-based multilingual adaptive recognition refers to the process of intelligently recognizing and dynamically learning from multiple languages' speech or text inputs within a unified model framework using artificial intelligence technology. This method extracts acoustic and linguistic features from speech signals through deep neural networks and uses semantic modeling mechanisms to map and fuse grammatical and semantic differences between different languages, thereby achieving adaptive language conversion and accurate recognition. In practical applications, the system can automatically adjust recognition parameters and feature weights based on the user's language environment, accent characteristics, and speech rate changes, continuously optimizing recognition accuracy and response speed, ultimately achieving cross-language, cross-regional, and cross-contextual human-computer speech understanding and interaction.
[0003] The existing technology has the following shortcomings:
[0004] In existing technologies, multilingual speech recognition systems generally rely on fixed time windows to segment and analyze speech signals. When the speaking speed frequently changes between fast and slow, the model often fails to adjust the length and step size of the time window in time, easily causing speech segments to overlap or break in time. Once the speaking speed increases, acoustic features are over-compressed, and the recognition system is prone to semantic overlap of entire segments; while when the speaking speed suddenly slows down, the speech signal is segmented prematurely, and the model is prone to missed recognition. These problems are more prominent in dynamic speech speed scenarios, directly leading to misalignment of recognized content, semantic breaks, and disordered sentence output, affecting the accuracy and continuity of multilingual adaptive recognition.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide an AI-based multilingual adaptive recognition method to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an AI-based multilingual adaptive recognition method, comprising the following steps:
[0008] The rhythm features and pause nodes in the continuous speech stream are collected, the speech rate change trajectory is extracted, and a rhythmic basic draft is generated to characterize the rhythmic change trend of the speech signal in the time dimension and to provide rhythmic basis for subsequent speech rate segment analysis.
[0009] Based on the rhythmic foundation, the speech signal is decomposed into speech rate changes to determine the speech rate increase and decrease regions, generate a time adjustment table, and record the duration and rhythmic span of each speech rate change segment for subsequent calculation of time window parameters.
[0010] Based on the time adjustment table, the semantic density of the speech signal is analyzed to distinguish between semantically dense and semantically sparse regions, determine the candidate range of the time window, generate a draft time window, and assign corresponding weight parameters to each candidate window to guide the rhythm allocation of the time window.
[0011] Based on the initial draft of the time window, the semantic connection relationship of the speech signal is detected, semantic misalignment segments and broken word groups are identified, a rhythm abnormality record table is generated, and the time position where dynamic rhythm correction needs to be performed is marked, which serves as the benchmark for subsequent window dynamic adjustment.
[0012] Dynamic rhythm adjustment is performed based on the marked positions in the rhythm anomaly record table. Time extension is performed in the speech rate increase zone, and rhythm forward is performed in the speech rate decrease zone. The step size and coverage width of the time window are adjusted in real time to ensure that the speech recognition process maintains temporal continuity and semantic integrity under different speech rate changes.
[0013] The preferred steps for generating the rhythmic basic draft are as follows:
[0014] The continuous speech stream is unfolded over time to establish a time index sequence of the speech signal. The speech energy value and fundamental frequency change value are recorded in time order to construct the speech energy curve and speech fundamental frequency curve. The time boundaries of the vocal segment and the silent pause segment are calibrated according to the energy change.
[0015] Using pause nodes as time boundaries, the speech stream is divided into multiple rhythm units. The time span and phoneme density of each rhythm unit are recorded. The time interval changes of adjacent rhythm units are compared to form a speech rate change sequence.
[0016] Arrange the speech rate increase zone, speech rate decrease zone, and speech rate steady zone in chronological order, record the duration, start and end time, and rhythm span of each zone, and correspond them with the speech energy curve to form a rhythm map.
[0017] Based on the time index, integrate the trajectory of speech rate changes, record the direction of speech rate, rhythm intervals and energy change trends, and generate a basic rhythmic draft.
[0018] Preferably, when establishing the time index sequence of the speech signal, the continuous changes of the speech energy curve are analyzed, and pause nodes are recorded when the energy drops to the silence threshold range. The time range of the pause nodes is determined by combining the fluctuation trend of the fundamental frequency curve. The rhythm intervals are continuously compared in the speech rate change sequence to calibrate the time boundaries between the speech rate increase area and the speech rate decrease area, so as to ensure the temporal continuity of speech rate changes in the rhythm map.
[0019] Preferably, the steps for generating the time adjustment table are as follows:
[0020] Based on the time index data in the rhythm foundation, the time series of the speech signal is expanded, the start and end times of pause nodes and adjacent speech segments are recorded, the time interval between adjacent time indices is calculated, and the trend of speech rate change is determined.
[0021] The continuous speech intervals were grouped according to the direction of speech rate change, the start and end positions of the speech rate increase and decrease areas were marked, and the duration, pause position and rhythm span of the speech units in each speech rate change segment were counted.
[0022] Extract and integrate the rhythm span within the speech rate variation segment, record the start time, end time, duration, rhythm span, speech rate direction and rhythm interval change rate, and generate a time adjustment table;
[0023] Connect the various speech rate change segments in the time adjustment table in chronological order to form a continuous speech rate change chain, and record the time span, rhythm span, and speech rate direction of each segment.
[0024] The preferred steps for generating the initial draft within the time window are as follows:
[0025] Based on the time index sequence in the time adjustment table, the rhythmic features of the speech signal are expanded. The start time, end time, duration, rhythmic span and speech direction of the speech rate change segment are recorded in time order. The semantic density distribution is extracted through energy change and vocalization duration.
[0026] Based on the trend of rhythm span change, the boundary between semantically dense and semantically sparse regions is identified, and the start time, end time, duration, average rhythm span and energy change value are recorded to establish a semantic density interval index table.
[0027] Based on the semantic density interval index table, the candidate range of the time window is defined, and the time boundary, rhythm span, semantic interval type and energy average are recorded to form a candidate set of time windows;
[0028] Assign weight parameters to each time window in the candidate time window set, record the time range, rhythm span, speech rate change direction and semantic density type, and generate a first draft of the time window.
[0029] Preferably, the steps for generating the rhythm anomaly record table are as follows:
[0030] Based on the time index order in the initial draft of the time window, the start time, end time, coverage area and weight parameters of the time window are arranged in chronological order. Semantic coherence detection is performed on the boundary area of adjacent time windows, and semantic misalignment segments and semantic overlap segments are marked according to semantic density parameters and rhythm span data.
[0031] Based on the semantic density distribution in the initial draft of the time window, the speech signal is divided into speech units, the start time and end time of each speech unit are recorded, and broken phrases and semantically misaligned segments are identified by comparing the time interval and semantic content of adjacent speech units.
[0032] The data of semantic misaligned segments and broken phrases are summarized, and the abnormal start time, end time, duration, rhythm span deviation value, semantic density parameter and time window weight parameter are recorded to generate a rhythm abnormality record table.
[0033] Extract the abnormal records from the rhythm abnormality record table into a time base set, record the key time nodes of semantic misalignment, broken phrases and semantically overlapping segments, and establish a dynamic correction reference time table.
[0034] Preferably, the time reference set in the rhythm anomaly record table establishes a dynamic correction reference time table by extending the start time of the semantically misaligned segment by a rhythm span, shifting the end time of the broken phrase by a rhythm span, and recording the midpoint time of the semantically overlapping segment, so as to achieve precise positioning of the time window at the semantic anomaly position and rhythm synchronization adjustment.
[0035] Preferably, dynamic rhythm adjustment is performed based on the marked positions in the rhythm anomaly record table, extending the time in the speech rate increase zone and shifting the rhythm forward in the speech rate decrease zone. The steps for adjusting the step size and coverage width of the time window in real time are as follows:
[0036] Based on the time information and semantic anomaly types in the rhythm anomaly record table, the speech rate change intervals are located and classified. The start time, end time, duration, rhythm span deviation, semantic density parameters and rhythm weight values are extracted and matched with the time index of the initial draft of the time window to establish a mapping relationship between anomaly locations and time windows.
[0037] In the speech rate enhancement zone, a time extension operation is performed, extending the end time of the time window by one rhythm span based on the end time, keeping the start time unchanged, and adjusting the time window step size parameter to match the speech rate change rate.
[0038] In the area where the speech rate slows down, a rhythm forward operation is performed, shifting the start time of the time window forward by one rhythm span, while keeping the end time unchanged, and adjusting the boundaries of adjacent windows to maintain the time connection.
[0039] The adjusted time windows are rearranged according to the time index, the weight parameters and step parameters are updated, the sampling intervals between time windows are balanced, and the integrity of time connection is checked.
[0040] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0041] This invention introduces a dual-layer time control mechanism—comprising a rhythmic draft and a time adjustment table—into the speech recognition process, enabling dynamic mapping of rhythmic changes in the speech signal across time. By continuously extracting the trajectory of speech rate changes and generating a rhythmic draft, the system can perceive the increasing and decreasing trends of speech rate in real time, thereby automatically adjusting the segmentation method and parameter settings of the time window to maintain temporal continuity in the feature extraction process of the speech signal. This method effectively avoids the problems of time superposition or premature segmentation of speech segments under speech rate fluctuations, making the rhythmic structure of speech recognition more consistent with the rhythmic characteristics of natural language expression, and resulting in clearer semantic boundaries in the recognition results.
[0042] This invention employs a dynamic rhythm adjustment mechanism to adaptively adjust the time window step size and coverage width, enabling the speech recognition process to maintain semantic stability under varying speech rates. By extending the time frame in the speech rate acceleration zone and advancing the rhythm frame in the speech rate deceleration zone, the system can correct semantically misaligned segments and broken phrases in real time, maintaining the logical coherence and semantic integrity of the speech recognition. This approach allows for accurate alignment and stable output of speech recognition results even in multilingual environments with frequent speech rate changes, significantly improving the continuity and overall accuracy of multilingual adaptive recognition. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0044] Figure 1 This is a flowchart of the AI-based multilingual adaptive recognition method of the present invention. Detailed Implementation
[0045] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0046] This invention provides, for example Figure 1The AI-based multilingual adaptive recognition method shown includes the following steps:
[0047] The rhythm features and pause nodes in the continuous speech stream are collected, the speech rate change trajectory is extracted, and a rhythmic basic draft is generated to characterize the rhythmic change trend of the speech signal in the time dimension and to provide rhythmic basis for subsequent speech rate segment analysis.
[0048] By unfolding the temporal dimension and analyzing the rhythmic structure of continuous speech streams, we meticulously extract rhythmic features and pauses from the speech signal to construct a complete trajectory of speech rate changes, ultimately generating a rhythmic foundation for time analysis. The entire process revolves around the temporal continuity and rhythmic integrity of speech, employing multi-level processing based on speech energy variation patterns, phoneme distribution intervals, intonation trends, and the natural connection of speech pause sequences. This comprehensively reveals the rhythmic trends of speech in the temporal dimension, providing a stable rhythmic basis for subsequent speech rate segmentation. The specific implementation steps are as follows:
[0049] A time-expanded sequence of continuous speech signals is constructed by performing temporal unfolding on the continuous speech stream. The original speech signal is divided into time segments according to the sampling order, and the speech energy value and fundamental frequency change value are recorded point by point in chronological order, constructing speech energy curves and fundamental frequency curves. During this process, the energy rise, energy fall, and energy plateau state at each sampling point on the time axis are classified and labeled to identify the time boundaries between speech production segments and silence pauses. To ensure the integrity of rhythmic features, when a region in the energy curve is detected where the energy continuously decreases to the silence threshold, the start and end points of this region are recorded as pause nodes, and the accurate time range of the pause is confirmed by combining the fluctuation trend of the fundamental frequency curve. In this way, the distribution structure of speech production can be accurately depicted on the time axis, giving the start point, duration, and end point of each speech segment quantifiable temporal characteristics, thus forming a preliminary framework for the distribution of speech rhythm.
[0050] Based on the obtained rhythm distribution framework, the speech signal is divided into rhythm units. Using pause nodes as time boundaries, the entire speech stream is divided into multiple independent rhythm units, each corresponding to a continuous vocal segment. The time span within each rhythm unit is statistically analyzed, the vocal duration is calculated, and the phoneme density is recorded. To ensure the temporal continuity between rhythm units, a temporal mapping relationship is established between the end point of each rhythm unit and the beginning point of the next rhythm unit, thus forming a continuous transition structure of the speech stream in the time dimension. Subsequently, difference analysis is performed on the time spans between rhythm units to extract the rate of change of time intervals between adjacent rhythm units, determining the preliminary trend of speech rate changes. When the time span of adjacent rhythm units continuously shortens, the time interval is defined as the speech rate increase zone; when the time span continuously lengthens, the time interval is defined as the speech rate decrease zone; when the time span remains stable, it is defined as the speech rate plateau zone. In this way, a speech rate change sequence in the time dimension of the speech signal can be established, providing a reliable data source for the formation of the speech rate trajectory in the next step.
[0051] Based on the speech rate variation sequence, the rhythmic fluctuations in the time dimension are structurally organized. Speech rate increase zones, decrease zones, and stable zones are arranged chronologically to form a speech rate variation trajectory. To ensure the trajectory comprehensively reflects the overall rhythmic variation pattern, the duration, start and end times, and rhythmic span of each speech rate segment are recorded, and these data are continuously connected along a time axis. Subsequently, the speech rate variation trajectory is fused with the speech signal's energy distribution curve, so that the trend of speech energy variation corresponds to the rhythmic distribution of the speech rate trajectory. During the fusion process, when speech energy shows a concentrated upward or downward trend, a temporal correlation is established between the variation segment and the corresponding speech rate segment to clarify the relationship between speech rate variation and speech energy fluctuations. In this way, a complete rhythmic map containing speech rate variation trends, rhythmic intervals, pause positions, and corresponding energy fluctuations can be obtained, forming a continuous and visible temporal path for the rhythmic changes of the speech signal in the time dimension.
[0052] Based on the rhythm map, the speech rate change trajectory is sequentially integrated according to the time index to generate a basic rhythm draft. The basic rhythm draft uses the time index as the main axis and the speech rate change trajectory as the core, uniformly arranging all pause nodes, speech rate change segments, and rhythm span information in the speech signal, thus forming a structured hierarchical description of speech rhythm in the time dimension. Specifically, the start point, end point, duration, direction of speech rate change, rhythm interval, and energy change trend of each speech rate segment are integrated as continuous data, and an index mapping relationship is established on the time axis, thereby giving the speech rhythm change trajectory complete temporal continuity. In this process, by sequentially arranging the various segments in the basic rhythm draft, the rhythmic fluidity of the speech signal in the time dimension is fully expressed. The basic rhythm draft not only records the continuous change process of speech rate in the speech signal but also describes the positional distribution of pause nodes in the time series, enabling a fully quantified representation of the speech rhythm information on the time axis. This rhythmic foundation provides a unified temporal basis for subsequent speech rate segmentation, time window adjustment, and semantic rhythm alignment, thereby ensuring that the speech recognition process maintains temporal consistency and rhythmic coherence under different speech rate environments.
[0053] Based on the rhythmic foundation, the speech signal is decomposed into speech rate changes to determine the speech rate increase and decrease regions, generate a time adjustment table, and record the duration and rhythmic span of each speech rate change segment for subsequent calculation of time window parameters.
[0054] By refining the time index sequence, rhythm interval data, and speech rate change curve recorded in the rhythmic foundation draft, the speech rate changes of the speech signal are continuously decomposed. This allows for a clear expression of the temporal distribution and rhythmic characteristics of the speech rate boost and deceleration zones in the time dimension. Based on this, a time adjustment table is established for the precise calculation of subsequent time window parameters. The entire process is based on the temporal continuity of the speech signal, using the speech rate change trajectory in the rhythmic foundation draft as the core reference. It sequentially completes time segment division, speech rate status confirmation, rhythmic span extraction, and time parameter summarization, forming an ordered distribution structure of speech rate change segments. The specific implementation steps are as follows:
[0055] Based on the time index data in the rhythmic foundation, the time series of the speech signal is continuously unfolded. In this process, each pause node recorded in the rhythmic foundation is mapped one-to-one with the start and end times of adjacent speech segments, constructing a time series index chain. Subsequently, the time intervals between adjacent time indices are continuously statistically analyzed to obtain the actual duration of each speech segment. By continuously comparing the time intervals, the time contraction and time extension trends of the speech signal in different segments on the time axis are determined. When the time interval between adjacent speech segments continuously decreases, it indicates that the speech rate is increasing within this time period; when the time interval continuously increases, it indicates that the speech rate is decreasing within this time period. In this way, a preliminary continuous distribution map of speech rate changes can be formed in the time dimension, giving the speech rate state of the speech signal in different time periods a clear temporal assignment, thus providing a basic time framework for subsequent speech rate segmentation.
[0056] Based on the resulting time distribution map, speech rate changes are precisely segmented. In this process, continuous speech intervals on the timeline are grouped according to their speech rate change trends, with each group representing an independent speech rate change segment. For speech rate increase segments, starting from the point where the speech rate begins to increase, the time intervals of speech segments are continuously recorded, decreasing until the speech rate change stops. For speech rate decrease segments, starting from the point where the speech rate change begins to decrease, the time intervals of speech segments are continuously recorded, increasing until the speech rate returns to a stable state. By calibrating the boundary times of each speech rate change segment, the start and end positions of the speech rate increase and decrease segments can be clearly identified on the timeline. Furthermore, the number of speech units within each speech rate change segment is counted, and the duration, pause position, and rhythmic span of each speech unit are recorded. This gives the speech rate change segments not only a duration attribute but also a detailed rhythmic description, providing data for subsequent rhythmic span calculations.
[0057] Based on defined speech rate variation segments, the rhythmic span of the speech signal is continuously extracted and integrated to generate a complete time adjustment table. During this process, the rhythmic interval data in the rhythmic foundation draft is used as a reference to read and statistically analyze the rhythmic span within each speech rate variation segment. When the speech rate increases, the rhythmic span gradually tightens; at this time, the reduction ratio of the rhythmic interval between each speech segment and the rate of change of the time interval need to be recorded. When the speech rate decreases, the rhythmic span gradually lengthens; at this time, the extension ratio of the rhythmic interval and the change value of the interval between speech segments need to be recorded. In this way, complete data on each speech rate variation segment in both time and rhythm dimensions can be obtained. Subsequently, all speech rate variation segments are arranged in chronological order, segment index numbers are established on the time axis, and the start time, end time, duration, rhythmic span, direction of speech rate change, and rate of change of rhythmic interval for each segment are entered into the time adjustment table. This time adjustment table can completely describe the distribution pattern of speech rate changes in the time dimension, providing a quantitative basis for subsequent time window adjustments.
[0058] After the time adjustment table is established, the data of each speech rate change segment recorded within it are summarized and integrated to ensure the consistency and coherence of time and rhythm information. In this process, the duration and rhythm span of all speech rate change segments in the time adjustment table are connected chronologically to form a continuous chain of speech rate changes. For each speech rate change segment, the duration represents the time range in which the speech rate change continues, and the rhythm span represents the rhythm density of the speech signal within that time range. When the speech rate is increasing, the duration and rhythm span recorded in the time adjustment table are negatively correlated; that is, the shorter the duration, the tighter the rhythm span. When the speech rate is decreasing, the duration and rhythm span are positively correlated; that is, the longer the duration, the looser the rhythm span. In this way, the time adjustment table not only reflects the direction of speech rate changes but also accurately depicts the changes in speech rhythm over time. Ultimately, the time adjustment table forms a structured set of time parameters, covering the time span, rhythm span, speech rate direction, and duration of each speech rate change segment. This provides a precise reference for the calculation of subsequent time window parameters, enabling the speech recognition process to dynamically adjust the time step and sampling coverage according to speech rate changes.
[0059] Based on the time adjustment table, the semantic density of the speech signal is analyzed to distinguish between semantically dense and semantically sparse regions, determine the candidate range of the time window, generate a draft time window, and assign corresponding weight parameters to each candidate window to guide the rhythm allocation of the time window.
[0060] By systematically processing the speech rate variation segments, rhythm span, duration, and speech rate direction information recorded in the time adjustment table, an in-depth analysis of the semantic density of the speech signal in the time dimension is performed, enabling the distinction between semantically dense and semantically sparse regions. Subsequently, based on the correlation between semantic density and rhythm span, candidate ranges for time windows are determined, and a draft time window is generated. Each candidate time window is assigned corresponding weight parameters to guide the subsequent rhythm allocation process. The entire process adheres to the principles of temporal continuity and semantic coherence of the speech signal, forming a quantifiable and traceable time window construction process by accurately characterizing the mapping relationship between semantic density changes and rhythm fluctuations. The specific implementation steps are as follows:
[0061] Based on the time index sequence of the time adjustment table, the rhythmic features of the speech signal in the time dimension are expanded. The start time, end time, duration, rhythmic span, and direction of speech rate change of each segment recorded in the time adjustment table are arranged chronologically to form a complete speech rate timeline. On this timeline, the energy changes, duration of articulation, and rhythmic intervals of the speech signal are analyzed point-by-point. By continuously comparing time intervals and articulation density, the temporal distribution features of semantic content are extracted. When the energy of the speech signal remains in a high amplitude range and the proportion of articulation time is relatively stable within a continuous time period, it indicates that the speech signal has concentrated expression within that time period, and the semantic information is relatively dense; this time period is marked as a semantically dense region. When the energy amplitude of the speech signal is dispersed and the pause intervals are significantly lengthened, it indicates that the semantic content is relatively sparse; this time period is marked as a semantically sparse region. Through this process, a preliminary distribution curve of semantic density can be established on the time axis, providing basic data support for subsequent time window delineation.
[0062] After obtaining the semantic density distribution curve, the boundary between semantically dense and semantically sparse regions is further refined. Using the rhythm span variation trend recorded in the time adjustment table as a reference, the rhythm span difference between adjacent speech rate variation segments is calculated and compared. When the rhythm span of two adjacent segments shows a continuous contraction trend, it indicates that the concentration of semantic content is increasing, forming the expansion boundary of the semantically dense region; when the rhythm span of two adjacent segments shows a continuous extension trend, it indicates that the distribution of semantic content is gradually diluting, forming the expansion boundary of the semantically sparse region. Subsequently, the start time, end time, duration, average rhythm span, and speech energy variation value of all semantically dense and semantically sparse regions are systematically recorded, and a semantic density interval index table is established according to the time sequence. This index table clearly defines the distribution position of each semantic interval in the time dimension, giving semantically dense and semantically sparse regions independent time identifiers and rhythmic attributes, thus laying the foundation for defining the range of the time window.
[0063] Based on the semantic density interval index table, the candidate range of the time window is defined. The temporal information and rhythmic span data of semantically dense and sparse regions are used as core inputs to analyze the semantic distribution along the time axis. For semantically dense regions, the candidate range of the time window covers the start and end times of the semantic region to ensure complete capture of semantic content within that time range. To prevent the omission of semantic boundary information, the start time of the time window is extended by one rhythmic interval before the start of the semantically dense region, and the end time is extended by one rhythmic interval after the end of the semantically dense region, thus ensuring complete semantic coverage. For semantically sparse regions, to improve the efficiency of time sampling, the time period is divided into several equally spaced time units based on the rhythmic span value, with each time unit corresponding to a candidate time window. Subsequently, all candidate time windows are arranged in chronological order to form a candidate time window set. For each candidate time window, its start time, end time, coverage range, corresponding semantic interval type, rhythmic span value, and average speech energy are recorded, thus giving each candidate window clear temporal attributes and semantic features.
[0064] After determining the candidate time window set, weight parameters are assigned to each candidate time window to generate an initial draft of the time window. The weight parameters are set based on the comprehensive relationship between semantic density and rhythm span. In semantically dense areas, the semantic load is high, so the weight parameters are primarily based on semantic density, increasing sequentially according to the degree of semantic density, giving time windows in semantically dense areas a higher priority in subsequent recognition. In semantically sparse areas, the rhythm span varies more, so the weight parameters are allocated based on the compactness of the rhythm span. When the rhythm span changes frequently, the weight parameters are increased accordingly to ensure a timely response to rhythm fluctuations; when the rhythm span changes tend to be stable, the weight parameters are relatively decreased to maintain the balance of time sampling. Subsequently, the weight parameters of each candidate time window are bound to its time range, rhythm span, direction of speech rate change, and semantic density type to form the initial draft of the time window. The initial drafts of the time windows are arranged in chronological order, recording the time boundary, coverage area, weight value, and semantic attributes of each window, so that subsequent rhythm allocation can be scheduled according to weight priority. In this way, the time window not only has rhythm-adaptive features, but can also dynamically adapt to the semantic density distribution of the speech signal, so that the speech recognition process can maintain continuity and stability under different speech speeds and semantic environments.
[0065] Based on the initial draft of the time window, the semantic connection relationship of the speech signal is detected, semantic misalignment segments and broken word groups are identified, a rhythm abnormality record table is generated, and the time position where dynamic rhythm correction needs to be performed is marked, which serves as the benchmark for subsequent window dynamic adjustment.
[0066] By comprehensively analyzing the semantic coherence of the speech signal using the time boundaries, coverage, weight parameters, and semantic density distribution data in the initial draft of the time window, the specific locations of semantically misaligned segments and broken phrases are detected. This results in the generation of a rhythm anomaly record table, which marks the time points where dynamic rhythm correction must be performed, providing a precise reference basis for subsequent dynamic adjustments to the time window. The entire process focuses on the temporal continuity, semantic coherence, and rhythmic consistency of the speech signal. Through detailed temporal comparison and semantic coherence tracking, it ensures that each anomaly can be accurately identified in the temporal dimension. The specific implementation steps are as follows:
[0067] Based on the time index order recorded in the initial draft of the time windows, the temporal distribution of the speech signal is unfolded segment by segment. The start time, end time, coverage area, and weight parameters of each time window in the initial draft are arranged in chronological order, forming a continuous time partition structure of the speech signal on the time axis. On this basis, semantic coherence detection is performed on the boundary regions between adjacent time windows. During detection, the semantic density parameters and rhythmic span data of adjacent time windows are used as the basis. By comparing the differences in rhythmic change direction and semantic distribution within the overlapping time area, the continuity of semantic content is determined. When the semantic content of adjacent time windows has a logical interruption within the overlapping time area, or when the semantic expression is misaligned at rhythmic transitions, the area is marked as a semantically misaligned segment. When adjacent time windows overlap in terms of temporal coverage, and the semantic content within the overlapping area appears repeatedly in the preceding and following time windows, it is marked as a semantically overlapping segment. Through this process, the range of regions with abnormal semantic coherence can be initially identified on the time axis, laying the temporal localization foundation for subsequent semantic misalignment and break detection.
[0068] After initially locating semantic connection anomalies, the semantic details within each time window are precisely decomposed. Based on the semantic density distribution in the initial draft of the time window, the speech signal within each time window is divided according to the rhythmic span distribution, forming a group of speech units. Each speech unit corresponds to a continuous semantic expression and has start and end timestamps. Subsequently, according to the order of the time windows, the time interval between adjacent speech units is compared with the semantic content. When the time interval between adjacent speech units is significantly greater than the average rhythmic span within the time window, it indicates a semantic interruption, which is judged as a broken phrase; when the semantic content between two adjacent speech units has logical overlap or reversed order, it is judged as a semantic misalignment segment. To ensure positioning accuracy, the start and end times, semantic span, rhythmic span deviation, and duration of adjacent speech units are recorded for each semantic misalignment and broken phrase. Through this process, the type, location, and time parameters of semantic anomalies can be structurally labeled, providing detailed data for the construction of the rhythm anomaly recording table.
[0069] After obtaining complete data on semantically misaligned segments and broken phrases, the abnormal information is summarized and classified to generate a rhythm anomaly record table. The rhythm anomaly record table is organized chronologically, arranging all anomalies in chronological order. The table records the anomaly type (semantic misalignment, semantic overlap, or broken phrases), anomaly start time, anomaly end time, anomaly duration, rhythm span deviation value, semantic density parameters, speech energy changes, and the weight parameters of the corresponding time window. During recording, if multiple semantic anomaly types exist within the same time period, they are separately identified as composite anomaly areas in the rhythm anomaly record table, and the time span and rhythm deviation value of each anomaly are recorded separately. To ensure the traceability of the anomaly information, each record in the rhythm anomaly record table maintains a one-to-one correspondence with the time index of the initial draft time window, enabling the anomaly information to be accurately mapped to the corresponding time window in the subsequent dynamic correction stage. In this way, the rhythm anomaly record table not only records the specific time location of the anomaly occurrence but also fully reflects the relationship between semantic breaks and misalignments and the time window structure, thus providing reliable basic data for dynamic rhythm correction.
[0070] After the rhythm anomaly record table is generated, all anomaly records are filtered and labeled to form a time reference set for dynamic rhythm correction. During this process, the start and end times of each anomaly record are extracted to establish a set of key time nodes. For semantically misaligned segments, the starting point is included in the time reference set by extending the time by one rhythm span, ensuring that the time window expansion can be initiated before that moment during subsequent rhythm correction, thus guaranteeing semantic continuity. For broken phrases, the shifted time point is included in the time reference set by shifting the time by one rhythm span, allowing for early capture of the next phonation start point during subsequent rhythm correction, avoiding semantic loss. For semantically overlapping segments, the midpoint time of the overlapping area is recorded in the time reference set to ensure a balanced distribution of the time window coverage during dynamic rhythm adjustment, preventing duplicate sampling. Subsequently, all time reference points are organized chronologically to establish a dynamic correction reference timetable. This timetable clearly defines the correction start time, adjustment duration, and rhythm offset direction for each semantic anomaly, providing a complete execution benchmark for dynamic rhythm correction. In this way, the time reference set achieves precise mapping of semantic anomalies, enabling dynamic adjustments to be implemented for specific time periods, thereby restoring the rhythmic balance and semantic continuity of the speech signal in the time dimension.
[0071] Dynamic rhythm adjustment is performed based on the marked positions in the rhythm anomaly record table. In the speech rate increase zone, time extension is performed, and in the speech rate decrease zone, rhythm forward is performed. The step size and coverage width of the time window are adjusted in real time to ensure that the speech recognition process maintains temporal continuity and semantic integrity under different speech rate changes.
[0072] By precisely analyzing the time stamp information recorded in the rhythm anomaly log, the rhythmic structure of the speech signal in the time dimension is dynamically adjusted, allowing the step size and coverage width of the time window to be adjusted in real time according to changes in speech rate. The entire process is based on the speech rate increase / decrease areas and abnormal time nodes recorded in the rhythm anomaly log, combined with the time range and weight parameters in the initial draft of the time window, forming a dynamic rhythm adjustment process. By performing time extension operations in the speech rate increase area and rhythm forward operations in the speech rate decrease area, the temporal continuity and semantic integrity of the speech recognition process are achieved, ensuring a stable temporal rhythm distribution of the speech signal under varying speech rate conditions. The specific implementation steps are as follows:
[0073] Based on the time information and semantic anomaly types in the rhythm anomaly record table, the speech rate change intervals are located and categorized in time. All abnormal time markers in the rhythm anomaly record table are arranged chronologically, and the anomaly type corresponding to each marker is classified: semantic misalignment segments and rhythm compression correspond to speech rate increase areas, while semantic break segments and rhythm extension correspond to speech rate decrease areas. Subsequently, the time data in the rhythm anomaly record table is matched with the time index in the initial draft of the time window to confirm the time window range corresponding to each anomaly point. During the matching process, the start time, end time, duration, rhythm span deviation, semantic density parameter, and rhythm weight value are extracted from the rhythm anomaly record table and compared with the window boundaries and coverage areas recorded in the initial draft of the time window to establish a mapping relationship between the abnormal time location and the time window. After completing the mapping, all speech rate increase and decrease areas are stored in a time series set, and the number of windows and the direction of rhythm span change within each time interval are recorded. This process clarifies the specific distribution of speech rate change areas on the time axis, providing a time-based basis for subsequent dynamic rhythm adjustment.
[0074] After locating the speech rate variation range, a time extension operation is performed on the speech rate increase region. This time extension operation aims to address the time compression issue caused by speech rate increase, expanding the coverage of the time window in the high speech rate region to prevent semantic information loss. Using the start and end times of the speech rate increase region as boundaries, all time windows within this time range are sequentially extended. The specific process is as follows: using the end time of each time window as a baseline, its end time is extended backward by a full rhythmic span, while keeping the start time unchanged, thus increasing the coverage width of the time window over time. The extended time window requires recalculating the rhythmic span distribution within its coverage area to ensure that rhythmic fluctuations within the time window are consistent with the direction of speech rate change. For sections where adjacent time windows overlap, the boundary positions are adjusted to ensure that each window maintains an independent coverage range on the time axis, avoiding overlapping sampling. During the time extension process, the window step parameter also needs to be updated synchronously to ensure that the sampling interval between adjacent windows is proportional to the speech rate change rate, thereby maintaining the continuity of rhythmic flow. This operation can extend the sampling time range under conditions of increased speech rate, so that the semantic content of the speech signal can be fully captured, preventing recognition errors caused by semantic compression.
[0075] After performing time extension operations in the speech rate boosting zone, a rhythm shift operation is performed in the speech rate slowing zone. The purpose of rhythm shifting is to adjust the starting position of the time window, allowing it to enter the semantically dense time period earlier, thereby preventing premature segmentation of speech caused by speech rate slowing. During rhythm shifting, the starting and ending times of the speech rate slowing zone are used as the basis for adjusting the time position of each time window within the interval. The specific adjustment process is as follows: the starting time of each time window is shifted forward by the length of a complete rhythm span, while the ending time remains unchanged, so that the sampling starting point of the time window enters the semantic expression zone earlier. During the shifting process, the coverage area of the window is recalculated to ensure that the step size of the time window is consistent with the rhythm interval. Simultaneously, the boundaries of adjacent windows are aligned to maintain temporal continuity between the shifted windows, preventing sampling gaps or overlaps. When a shifted window intersects with the previous time window, the step size parameter is adjusted by balancing the temporal boundaries of both to maintain a balanced rhythm distribution. In this way, the time window of the speech rate slowdown zone can respond to the rhythm changes of the speech signal in advance, so that the rhythm remains continuous and stable during the slowdown phase, avoiding semantic breaks and thus maintaining the logical coherence of the speech recognition results in the time dimension.
[0076] After completing the time extension operations in the speech rate increase zone and the rhythm shift operations in the speech rate decrease zone, the time windows across the entire time series are globally and dynamically optimized to achieve a stable distribution of window step size and coverage width. Then, all adjusted time windows are rearranged according to their time index order, forming a new time window distribution sequence. For each window, the weight parameters are updated based on its corresponding speech rate change type. The window weight values in the speech rate increase zone are adjusted according to the increased coverage area after the extension, while the window weight values in the speech rate decrease zone are corrected based on the advance sampling amount after the shift. Subsequently, the step size parameters of each window are uniformly balanced to ensure a linear transition between the sampling intervals of the time windows. For segments with drastic speech rate changes, the sampling density is increased by reducing the step size parameter; for segments with stable speech rates, the sampling density is decreased by increasing the step size parameter, thus achieving a balanced rhythm distribution on the time axis. After completing the global adjustment, time continuity checks are performed on the boundaries of all time windows to ensure no time discontinuities between windows, and the coverage width is adjusted as necessary to fill potential gaps. In this way, the speech signal achieves dynamic adaptation of the time window under different speech rates, so that the speech sampling covers the time axis to form a continuous distribution. The rhythm changes of the speech recognition process are synchronized with the semantic distribution, thereby maintaining a stable time sequence structure and complete semantic expression effect during the recognition process.
[0077] This invention introduces a dual-layer time control mechanism—comprising a rhythmic draft and a time adjustment table—into the speech recognition process, enabling dynamic mapping of rhythmic changes in the speech signal across time. By continuously extracting the trajectory of speech rate changes and generating a rhythmic draft, the system can perceive the increasing and decreasing trends of speech rate in real time, thereby automatically adjusting the segmentation method and parameter settings of the time window to maintain temporal continuity in the feature extraction process of the speech signal. This method effectively avoids the problems of time superposition or premature segmentation of speech segments under speech rate fluctuations, making the rhythmic structure of speech recognition more consistent with the rhythmic characteristics of natural language expression, and resulting in clearer semantic boundaries in the recognition results.
[0078] This invention employs a dynamic rhythm adjustment mechanism to adaptively adjust the time window step size and coverage width, enabling the speech recognition process to maintain semantic stability under varying speech rates. By extending the time frame in the speech rate acceleration zone and advancing the rhythm frame in the speech rate deceleration zone, the system can correct semantically misaligned segments and broken phrases in real time, maintaining the logical coherence and semantic integrity of the speech recognition. This approach allows for accurate alignment and stable output of speech recognition results even in multilingual environments with frequent speech rate changes, significantly improving the continuity and overall accuracy of multilingual adaptive recognition.
[0079] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. An AI-based multilingual adaptive recognition method, characterized by, Includes the following steps: The rhythm features and pauses in a continuous speech stream are collected, the speech rate change trajectory is extracted, and a rhythmic basic text is generated to characterize the rhythmic change trend of the speech signal in the time dimension. Based on the rhythmic script, the speech signal is decomposed into speech rate changes, the speech rate increase zone and speech rate decrease zone are determined, a time adjustment table is generated, and the duration and rhythm span of each speech rate change zone are recorded. Based on the time adjustment table, the semantic density of the speech signal is analyzed to distinguish between semantically dense and semantically sparse regions, the candidate range of the time window is determined, a draft time window is generated, and corresponding weight parameters are assigned to each candidate window. The steps for generating the initial draft within the time window are as follows: Based on the time index sequence in the time adjustment table, the rhythmic features of the speech signal are expanded. The start time, end time, duration, rhythmic span and speech direction of the speech rate change segment are recorded in time order. The semantic density distribution is extracted through energy change and vocalization duration. Based on the trend of rhythm span change, the boundary between semantically dense and semantically sparse regions is identified, and the start time, end time, duration, average rhythm span and energy change value are recorded to establish a semantic density interval index table. Based on the semantic density interval index table, the candidate range of the time window is defined, and the time boundary, rhythm span, semantic interval type and energy average are recorded to form a candidate set of time windows; Assign weight parameters to each time window in the candidate time window set, record the time range, rhythm span, speech rate change direction and semantic density type, and generate a first draft of the time window; Based on the initial draft of the time window, the semantic connection relationship of the speech signal is detected, semantic misalignment segments and broken word groups are identified, a rhythm anomaly record table is generated, and the time position where dynamic rhythm correction needs to be performed is marked. Dynamic rhythm adjustment is performed based on the marked positions in the rhythm anomaly record table. Time extension is performed in the speech rate increase zone, and rhythm forward is performed in the speech rate decrease zone. The step size and coverage width of the time window are adjusted in real time to ensure that the speech recognition process maintains temporal continuity and semantic integrity under different speech rate changes. 2.The AI-based multi-language adaptive recognition method of claim 1, wherein, The steps to generate a basic rhythm script are as follows: The continuous speech stream is unfolded over time to establish a time index sequence of the speech signal. The speech energy value and fundamental frequency change value are recorded in time order to construct the speech energy curve and speech fundamental frequency curve. The time boundaries of the vocal segment and the silent pause segment are calibrated according to the energy change. Using pause nodes as time boundaries, the speech stream is divided into multiple rhythm units. The time span and phoneme density of each rhythm unit are recorded. The time interval changes of adjacent rhythm units are compared to form a speech rate change sequence. Arrange the speech rate increase zone, speech rate decrease zone, and speech rate steady zone in chronological order, record the duration, start and end time, and rhythm span of each zone, and correspond them with the speech energy curve to form a rhythm map. Based on the time index, integrate the trajectory of speech rate changes, record the direction of speech rate, rhythm intervals and energy change trends, and generate a basic rhythmic draft. 3.The AI-based multi-language adaptive recognition method of claim 2, wherein, When establishing the time index sequence of speech signals, the continuous changes of the speech energy curve are analyzed. When the energy drops to the silence threshold range, the pause node is recorded. The time range of the pause node is determined by combining the fluctuation trend of the fundamental frequency curve. The rhythm interval is continuously compared in the speech rate change sequence to calibrate the time boundary between the speech rate increase zone and the speech rate decrease zone. 4.The AI-based multi-language adaptive recognition method of claim 2, wherein, The steps to generate a time adjustment table are as follows: Based on the time index data in the rhythm foundation, the time series of the speech signal is expanded, the start and end times of pause nodes and adjacent speech segments are recorded, the time interval between adjacent time indices is calculated, and the trend of speech rate change is determined. The continuous speech intervals were grouped according to the direction of speech rate change, the start and end positions of the speech rate increase and decrease areas were marked, and the duration, pause position and rhythm span of the speech units in each speech rate change segment were counted. Extract and integrate the rhythm span within the speech rate variation segment, record the start time, end time, duration, rhythm span, speech rate direction and rhythm interval change rate, and generate a time adjustment table; Connect the various speech rate change segments in the time adjustment table in chronological order to form a continuous speech rate change chain, and record the time span, rhythm span, and speech rate direction of each segment.
5. The AI-based multilingual adaptive recognition method according to claim 1, characterized in that, The steps for generating the rhythm anomaly record table are as follows: Based on the time index order in the initial draft of the time window, the start time, end time, coverage area and weight parameters of the time window are arranged in chronological order. Semantic coherence detection is performed on the boundary area of adjacent time windows, and semantic misalignment segments and semantic overlap segments are marked according to semantic density parameters and rhythm span data. Based on the semantic density distribution in the initial draft of the time window, the speech signal is divided into speech units, the start time and end time of each speech unit are recorded, and broken phrases and semantically misaligned segments are identified by comparing the time interval and semantic content of adjacent speech units. The data of semantic misaligned segments and broken phrases are summarized, and the abnormal start time, end time, duration, rhythm span deviation value, semantic density parameter and time window weight parameter are recorded to generate a rhythm abnormality record table. Extract the abnormal records from the rhythm abnormality record table into a time base set, record the key time nodes of semantic misalignment, broken phrases and semantically overlapping segments, and establish a dynamic correction reference time table.
6. The AI-based multilingual adaptive recognition method according to claim 5, characterized in that, The time reference set in the rhythm anomaly record table establishes a dynamic correction reference timetable by extending the start time of semantically misaligned segments by a rhythm span, shifting the end time of broken phrases by a rhythm span, and recording the midpoint time of semantically overlapping segments.
7. The AI-based multilingual adaptive recognition method according to claim 5, characterized in that, Based on the marked positions in the rhythm anomaly record table, dynamic rhythm adjustment is performed, extending the time in the speech rate increase zone and shifting the rhythm forward in the speech rate decrease zone. The steps for adjusting the step size and coverage width of the time window in real time are as follows: Based on the time information and semantic anomaly types in the rhythm anomaly record table, the speech rate change intervals are located and classified. The start time, end time, duration, rhythm span deviation, semantic density parameters and weight parameters are extracted and matched with the time index of the initial draft of the time window to establish a mapping relationship between anomaly locations and time windows. In the speech rate enhancement zone, a time extension operation is performed, extending the end time of the time window by one rhythm span based on the end time, keeping the start time unchanged, and adjusting the time window step size parameter to match the speech rate change rate. In the area where the speech rate slows down, a rhythm forward operation is performed, shifting the start time of the time window forward by one rhythm span, while keeping the end time unchanged, and adjusting the boundaries of adjacent windows to maintain the time connection. The adjusted time windows are rearranged according to the time index, the weight parameters and step parameters are updated, the sampling intervals between time windows are balanced, and the integrity of time connection is checked.
Citation Information
Patent Citations
Speech translation system based on Bluetooth earphone
CN120164455A
Extracting content from speech prosody
WO2019102477A1