Audio text alignment method and device, equipment and storage medium
By combining audio rhythm characteristics and text semantics, using rhythm change rate index and semantic analysis technology, high-precision alignment between audio and text is achieved, solving the problem of insufficient alignment accuracy for complex natural speech in the prior art.
Patent Information
- Application Number
- CN202510526402.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When the prior art realizes the alignment of audio text timestamps, there are problems of sparse data distribution and inaccurate keyword timestamps, and it is impossible to effectively deal with natural speech that includes complex factors such as speech speed changes, pauses, ups and downs of speech tone, and emotional expression.
By combining the rhythm characteristics of audio and the semantics of text, the rhythm rate change index is used to dynamically allocate timestamps, and semantic analysis is used to identify important semantic nodes in the text as alignment anchors, thereby achieving the alignment of audio and text.
It improves the accuracy of audio and text alignment, can more accurately reflect the rhythm characteristics of real voice, and enhances the system's ability to adapt to different language styles and content types.
Smart Images

Figure CN120104759A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to an audio text alignment method, device, equipment and storage medium. Background Art
[0002] In the digital age, accurate synchronization of audio content and text has become a core requirement for many applications. From video subtitles, speech transcription to language learning platforms, accurate timestamp alignment is the key to providing a high-quality user experience. The core task of timestamp alignment technology is to determine the exact time point of each element (word, phrase or character) in the text and the corresponding segment in the audio stream, thereby establishing a mapping relationship between the two.
[0003] Currently, the end-to-end model of the CTC algorithm (Connectionist Temporal Classification) and the Transformer-based attention mechanism are usually used to achieve audio-text timestamp alignment. However, this method has the problems of sparse distribution of key data and inaccurate timestamps of detected keywords. It cannot guarantee the accuracy of aligning natural speech and related texts that contain complex factors such as changes in speaking speed, pauses, intonation, and emotional expressions. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an audio-text alignment method, device, equipment and storage medium, which can align audio and text by combining the rhythm characteristics of audio and the semantics of text, thereby ensuring the alignment accuracy of audio and text. The specific scheme is as follows:
[0005] In a first aspect, the present application provides an audio text alignment method, comprising:
[0006] Acquire initial audio data and a target transcribed text corresponding to the initial audio data, acquire the rhythm change rate index corresponding to the initial audio data at different time nodes, and perform semantic analysis on the target transcribed text to acquire the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data;
[0007] Determine, from each of the initial semantic units, a target semantic unit whose importance is greater than a preset importance threshold according to each of the importance levels, and preliminarily match each of the target semantic units with the initial audio data to determine an anchor point position of each of the target semantic units in the initial audio data;
[0008] Based on each of the rhythm change rate indices, a timestamp is dynamically assigned to the initial audio data to obtain corresponding target audio data, the target audio data is divided into different target audio segments based on the anchor point positions corresponding to each of the target semantic units, and each of the target audio segments is aligned with the target transcription text based on each of the timestamps.
[0009] Optionally, obtaining the rhythm change rate indexes corresponding to the initial audio data at different time nodes respectively includes:
[0010] Obtaining speech rate changes, tone information, and pauses corresponding to the initial audio data at different time nodes, and determining target weights corresponding to the speech rate changes, the tone information, and the pauses, respectively;
[0011] The speech rate change, the pitch information and the pause conditions corresponding to different time nodes are weightedly fused based on the target weights to obtain the rhythm change rate index corresponding to the initial audio data at different time nodes.
[0012] Optionally, obtaining the rhythm change rate indexes corresponding to the initial audio data at different time nodes respectively includes:
[0013] The speech scene and speech style corresponding to the initial audio data are obtained, the target parameters corresponding to the rhythm change rate index are automatically adjusted according to the speech scene and the speech style, and the rhythm change rate index corresponding to the initial audio data at different time nodes is obtained according to the corresponding adjusted parameters.
[0014] Optionally, the performing semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text includes:
[0015] Performing semantic analysis on the target transcribed text to obtain a grammatical centrality score, a semantic density score, and a topic relevance score corresponding to each of the initial semantic units in the target transcribed text; wherein the grammatical centrality score represents the position of each of the initial semantic units in the corresponding syntactic structure, the semantic density score represents the amount of information contained in each of the initial semantic units, and the topic relevance score represents the degree of association between each of the initial semantic units and the core topic corresponding to the target transcribed text;
[0016] The grammatical centrality score, the semantic density score, and the topic relevance score are weightedly fused to obtain the importance of each initial semantic unit in the target transcription text.
[0017] Optionally, the obtaining of the grammatical centrality score, the semantic density score, and the topic relevance score corresponding to each of the initial semantic units in the target transcription text includes:
[0018] Acquire the grammatical centrality score corresponding to each of the initial semantic units based on the part of speech of each of the initial semantic units;
[0019] Obtaining the semantic density scores respectively corresponding to the initial semantic units according to the inverse document frequencies corresponding to the initial semantic units and the occurrence frequencies of the initial semantic units in the target transcribed text;
[0020] Map each of the initial semantic units and the core topic to the target semantic space to obtain a first vector corresponding to each of the initial semantic units and a second vector corresponding to the core topic, obtain the cosine similarity between each of the first vectors and the second vectors, and obtain the topic relevance score corresponding to each of the initial semantic units based on each of the cosine similarities.
[0021] Optionally, dividing the target audio data into different target audio segments based on the anchor point positions corresponding to the target semantic units includes:
[0022] Determine the information complexity corresponding to different parts of the target audio data, and dynamically set the target paragraph lengths corresponding to different parts of the target audio data according to each information complexity and each anchor point position; wherein the information complexity is used to characterize the amount of information contained in different parts of the target audio data;
[0023] The target audio data is divided into different target audio segments based on the length of each target paragraph; wherein the audio data contained in the head of any target audio segment is the same as the audio data contained in the tail of the previous adjacent target audio segment, and the audio data contained in the tail of any target audio segment is the same as the audio data contained in the head of the next adjacent target audio segment.
[0024] Optionally, aligning each of the target audio segments with the target transcription text based on each of the timestamps includes:
[0025] Constructing a similarity matrix between the target audio data and the target transcribed text, and obtaining a matching pattern between the target audio data and the target transcribed text according to the similarity matrix;
[0026] Based on the matching pattern, a difference area between the target audio data and the target transcription text is obtained, a difference type corresponding to the difference area is obtained, and each of the target audio segments is aligned with the target transcription text according to an alignment strategy corresponding to the difference type.
[0027] In a second aspect, the present application provides an audio text alignment device, comprising:
[0028] A semantic analysis module is used to obtain initial audio data and a target transcribed text corresponding to the initial audio data, obtain the rhythm change rate index corresponding to the initial audio data at different time nodes, and perform semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data;
[0029] An anchor point position determination module, used to determine, from each of the initial semantic units, a target semantic unit whose importance is greater than a preset importance threshold according to each of the importance levels, and preliminarily match each of the target semantic units with the initial audio data to determine the anchor point position of each of the target semantic units in the initial audio data;
[0030] An audio-text alignment module is used to dynamically assign timestamps to the initial audio data based on each rhythm change rate index to obtain corresponding target audio data, divide the target audio data into different target audio segments based on each anchor point position corresponding to each target semantic unit, and align each target audio segment with the target transcription text based on each timestamp.
[0031] In a third aspect, the present application provides an electronic device, including:
[0032] Memory, used to store computer programs;
[0033] A processor is used to execute the computer program to implement the aforementioned audio-text alignment method.
[0034] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which implements the aforementioned audio-text alignment method when executed by a processor.
[0035] In the present application, initial audio data and a target transcribed text corresponding to the initial audio data are first obtained, rhythm change rate indexes corresponding to the initial audio data at different time nodes are obtained, and semantic analysis is performed on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data, and then, according to each importance, a target semantic unit whose importance is greater than a preset importance threshold is determined from each initial semantic unit, and each target semantic unit is preliminarily matched with the initial audio data to determine the anchor point position of each target semantic unit in the initial audio data, and finally, based on each rhythm change rate index, a timestamp is dynamically assigned to the initial audio data to obtain the corresponding target audio data, and based on each anchor point position corresponding to each target semantic unit, the target audio data is divided into different target audio segments, and each target audio segment is aligned with the target transcribed text based on each timestamp. It can be seen that this application can obtain complex factors such as speech rate changes, pauses, intonation fluctuations and emotional expressions of audio data through in-depth analysis of the rhythmic characteristics of audio data and the corresponding semantics of the transcribed text, thereby improving the accuracy of audio and text alignment; by allocating timestamps to audio data according to the rhythm change rate index corresponding to the audio data, it avoids the use of a simple uniform distribution strategy, and can ensure that the timestamps of highly complex segments in the audio are more dense, thereby ensuring the alignment accuracy of highly complex positions in the audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0037] Figure 1 A flowchart of an audio-text alignment method disclosed in this application;
[0038] Figure 2 A flowchart of a specific audio-text alignment method disclosed in this application;
[0039] Figure 3 A flowchart of text semantic analysis disclosed in this application;
[0040] Figure 4 A flowchart of a specific audio-text alignment method disclosed in this application;
[0041] Figure 5A flowchart of audio-text mismatch area detection disclosed in this application;
[0042] Figure 6 This is a schematic diagram of the structure of an audio-text alignment device disclosed in this application;
[0043] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] The current audio-text alignment method cannot guarantee the accuracy of audio and text alignment. To this end, the present application provides an audio-text alignment method, which aligns audio and text by combining the rhythm characteristics of the audio and the semantics of the text, thereby ensuring the alignment accuracy of the audio and text.
[0046] See also Figure 1 As shown, an embodiment of the present invention discloses an audio text alignment method, comprising:
[0047] Step S11, obtaining initial audio data and a target transcribed text corresponding to the initial audio data, obtaining the rhythm change rate index corresponding to the initial audio data at different time nodes, and performing semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data.
[0048] The audio-text alignment method provided in this embodiment is applied to the SRATA system (Speech Rhythm-Aware and Semantic-Guided Adaptive Timestamp Alignment System), hereinafter referred to as the system, which is derived from an in-depth study of the human speech comprehension process. Studies have shown that when humans process speech information, they not only pay attention to the pronunciation content, but also subconsciously perceive the speech rhythm, stress position and semantic structure. These factors jointly affect the perception and understanding of the speech time distribution. Traditional alignment algorithms often ignore these natural characteristics, resulting in poor performance when processing real speech. This system redefines alignment as a comprehensive process of speech rhythm understanding and semantic structure recognition, rather than a simple signal matching problem. The system design solves several key technical problems: how to quantify speech rhythm features (i.e., rhythm change rate index) and convert them into a basis for timestamp allocation; how to identify semantically important nodes in the text and use them as alignment anchors (i.e., anchor points); how to design an alignment algorithm that can both ensure accuracy and control computational complexity. Among them, the overall process of the audio-text alignment method in this implementation is as follows: Figure 2 As shown, first, the text needs to be semantically analyzed separately, and the rhythm change rate index of the audio needs to be extracted in different time windows, and then the audio and text are aligned based on the semantic analysis results of the text and the rhythm change rate index of the audio. The overall pseudo code of the audio text alignment method in this embodiment is as follows:
[0049] Algorithm: SRATA;
[0050] Input: audio file A, transcription text T;
[0051] Output: Aligned timestamps for each word in T;
[0052] / / Step 1: Preprocess audio and text;
[0053] preprocessed_audio=preprocess(A)
[0054] preprocessed_text=preprocess(T)
[0055] / / Step 2: Extract multi-scale rhythm features from audio;
[0056] rhythm_features=extract_rhythm_features(preprocessed_audio)
[0057] / / Step 3: Calculate the rhythm change rate index;
[0058] RVI=calculate_RVI(rhythm_features)
[0059] / / Step 4: Analyze the semantic structure of the text;
[0060] semantic_units=analyze_semantics(preprocessed_text)
[0061] / / Step 5: Calculate semantic importance score;
[0062] importance_scores=calculate_semantic_importance(semantic_units)
[0063] / / Step 6: Identify semantic anchor points;
[0064] anchor_points=identify_anchors(semantic_units,importance_scores)
[0065] / / Step 7: Use anchor points for rough alignment;
[0066] segments=coarse_align(preprocessed_audio, preprocessed_text, anchor_points)
[0067] / / Step 8: Perform fine alignment on each paragraph;
[0068] aligned_timestamps=[]
[0069] for segment in segments:
[0070] / / Application adaptive fault tolerance mechanism;
[0071] segment_with_tolerance=apply_error_tolerance(segment)
[0072] / / Perform rhythm-aware fine alignment;
[0073] segment_timestamps=fine_align(segment_with_tolerance,RVI)
[0074] aligned_timestamps.append(segment_timestamps)
[0075] / / Step 9: Apply global consistency constraints;
[0076] final_timestamps=apply_global_constraints(aligned_timestamps)
[0077] return final_timestamps
[0078] In this embodiment, it is necessary to obtain the rhythm change rate index of the initial audio data at different time points. It should be noted that speech rhythm is one of the core characteristics of human language expression, which is reflected in many aspects such as speech rate change, pause distribution, and stress pattern. Traditional alignment algorithms usually assume that speech has a constant rate, resulting in insufficient accuracy when processing real speech. SRATA has designed a speech rhythm perception mechanism, which can dynamically adjust the timestamp density distribution by analyzing the speech rate changes, pause patterns and intonation contours in the audio. This mechanism is based on an in-depth study of the expression of speech rhythm: changes in speech rate usually reflect the importance and complexity of the content; pauses often mark the boundaries of semantic units; and stress emphasizes key information points. The system constructs a mathematical description model of speech rhythm through multi-dimensional speech rhythm feature extraction.
[0079] The system innovatively introduces the Rhythm Variation Index (RVI), which is a quantitative indicator that comprehensively reflects the degree of change in speech rhythm. RVI is calculated by weighted fusion of factors such as speech rate change, pause distribution, and stress pattern. Based on RVI, the system constructs a timestamp density mapping function to assign appropriate timestamp density values to each time point in the audio. This non-uniform allocation strategy makes the timestamp distribution more consistent with the characteristics of real speech and significantly improves the alignment accuracy. Specifically, the system extracts multidimensional features from the audio signal, including short-term energy changes, pause point recognition, tone and pitch contour analysis, and local speech rate measurement. These features together constitute a comprehensive description of speech rhythm. Based on these features, the system calculates the rhythm variation rate index (RVI), which is a quantitative indicator that comprehensively reflects the degree of change in speech rhythm. The rhythm variation rate index is a quantitative indicator that reflects the degree of change in speech rhythm, and its mathematical expression is as follows:
[0080] ;
[0081] in: represents the rhythm change rate index at time point t, represents the normalized speech rate change at time point t, represents the normalized pause density at time point t, represents the normalized stress pattern at time point t, is the weight factor, satisfying ; The calculation of each component is as follows:
[0082] ;
[0083] in, is the local speaking rate at time point t (e.g., syllables per second), is the average speaking speed of the entire audio, is the maximum speech rate deviation observed in the audio.
[0084] ;
[0085] in, is the local pause density at time point t (e.g., the total pause duration in the window around t), is the maximum pause density observed in the audio.
[0086] ;
[0087] in, is the local energy change at time t (representing the stress pattern), is the maximum energy change observed in the audio.
[0088] To balance local and global rhythm features, the system adopts a multi-scale rhythm analysis method, considers the rhythm features of different time windows at the same time, and obtains the final rhythm description through weighted fusion. This method not only retains the sensitivity to local changes, but also maintains global consistency, and solves the difficulties of traditional methods in processing different scales of speech changes. Correspondingly, in the present embodiment, the process of obtaining the rhythm change rate index corresponding to the initial audio data at different time nodes can specifically include: obtaining the speech speed change situation, tone information and pause situation corresponding to the initial audio data at different time nodes, and determining the target weights corresponding to the speech speed change situation, tone information and pause situation respectively; Based on each target weight, the speech speed change situation, tone information and pause situation corresponding to different time nodes are weighted fused to obtain the rhythm change rate index corresponding to the initial audio data at different time nodes. Specifically, the system adopts a multi-scale analysis method and simultaneously considers the rhythmic features in different time windows. The short window (100-300ms) captures local speech rate changes, the medium window (0.5-2s) captures the structural features of semantic units, and the long window (3-10s) reflects the overall expression pattern. By weighted fusion of features at different scales, the system obtains a comprehensive and accurate description of the rhythm.
[0089] In addition, due to the significant differences in speech rhythm features between different speakers and different scenes (i.e., speech scenes), such as speeches, conversations, and readings, the embodiment of the present application also considers the adaptability of different speaking styles (i.e., speech styles); accordingly, the process of obtaining the rhythm change rate index corresponding to the initial audio data at different time nodes may specifically include: obtaining the speech scene and speech style corresponding to the initial audio data, automatically adjusting the target parameters corresponding to the rhythm change rate index according to the speech scene and the speech style, and obtaining the rhythm change rate index corresponding to the initial audio data at different time nodes according to the corresponding adjusted parameters; specifically, the system designs a style adaptation mechanism, which automatically adjusts the rhythm analysis parameters by analyzing the overall characteristics of the audio content to adapt to different speech styles. For example, for fast speeches, the system adjusts the sensitivity of speech speed changes; for content with rich emotional expression, the system increases the weight of pitch changes. By integrating semantic analysis into the alignment process, alignment is no longer a simple signal matching, but an intelligent process with language understanding capabilities. This method not only improves the alignment accuracy, but also enhances the system's adaptability to different language styles and content types.
[0090] Step S12: determine, from the initial semantic units, target semantic units whose importance is greater than a preset importance threshold according to the importance levels, and preliminarily match the target semantic units with the initial audio data to determine the anchor point position of each target semantic unit in the initial audio data.
[0091] Semantic anchor technology is one of the core innovations of this system. It aims to identify semantically important nodes in text and use them as reliable reference points for alignment. This technology is derived from the observation of the human speech comprehension process: when people listen to speech content, they often capture key words first, and then understand the complete content around these key points. Through in-depth research on linguistic principles and natural language processing technology, the system has developed a multi-dimensional semantic importance assessment method that comprehensively considers the three dimensions of grammatical centrality, semantic density, and topic relevance. In addition, matching semantic anchors with audio features is another key challenge. The system has designed a semantic-acoustic feature matching algorithm to find the most likely matching point by analyzing the possible acoustic performance of semantically important nodes in audio. In order to improve the matching accuracy, the system adopts a multi-feature fusion strategy that comprehensively considers energy features, frequency features, and time domain features.
[0092] In this embodiment, the semantic anchor technology first performs semantic unit analysis, wherein the process of scoring the importance of the semantic unit is as follows: Figure 3As shown in the figure, it includes syntactic analysis of the text to identify core grammatical components such as subject, predicate and object; semantic role labeling to determine the function of each word in the semantic structure; named entity recognition to find special entities. On this basis, the system evaluates the importance of each semantic unit and quantifies the importance of semantic units from three dimensions: grammatical centrality, semantic density and topic relevance. Among them, grammatical centrality reflects the status of vocabulary in the syntactic structure, and core components such as subject and predicate have high grammatical centrality; semantic density measures the amount of information carried by vocabulary, and professional terms, key concepts, etc. often have high semantic density; topic relevance evaluates the degree of association between vocabulary and core topics, and vocabulary directly related to the central topic has high topic relevance; specifically, the semantic importance scoring model in this system evaluates the importance of semantic units from three dimensions:
[0093] ;
[0094] in, is the semantic importance score of word or phrase w, is the grammatical centrality score, is the semantic density score, is the topic relevance score, is a weight factor that satisfies The components are calculated as follows:
[0095] ;
[0096] in, is the importance level of grammatical role i (e.g., subject = 1.0, predicate = 0.8, etc.), is a binary indicator (1 if w has role i, 0 otherwise).
[0097] ;
[0098] in, is the frequency of word w in the document, is the inverse document frequency of w in the reference corpus, It is a specificity score that measures the professionalism / technicality of the term.
[0099] ;
[0100] in, is the vector representation of word w in the semantic space, is the vector representation of the topic, is the cosine similarity function.
[0101] After identifying the semantically important nodes, the system matches them with the audio features to determine the anchor point location. The system designs a semantic and acoustic feature matching algorithm to analyze the possible acoustic performance of semantically important nodes in the audio, such as stress position, pitch change, energy peak, etc., and combines contextual information to find the most likely matching point.
[0102] To improve matching accuracy, the system adopts a multi-feature fusion strategy, taking into account energy features, frequency features, and time domain features. Energy features reflect changes in speech intensity, and important words often have higher energy values; frequency features capture pitch and tone contours, and important content is often accompanied by specific tone changes; time domain features analyze duration and rhythm patterns, and the expression of core content often has a unique time structure. By weighted fusion of these features, the system can accurately find the location of semantically important nodes in the audio.
[0103] Accordingly, the process of performing semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text may specifically include: performing semantic analysis on the target transcribed text to obtain the grammatical centrality score, semantic density score and topic relevance score corresponding to each initial semantic unit in the target transcribed text; wherein the grammatical centrality score represents the position of each initial semantic unit in the corresponding syntactic structure, the semantic density score represents the amount of information contained in each initial semantic unit, and the topic relevance score represents the degree of association between each initial semantic unit and the core topic corresponding to the target transcribed text; weighted fusion of the grammatical centrality score, the semantic density score and the topic relevance score is performed to obtain the importance of each initial semantic unit in the target transcribed text.
[0104] Among them, the process of obtaining the grammatical centrality score, semantic density score and topic relevance score corresponding to each initial semantic unit in the target transcription text may specifically include: obtaining the grammatical centrality score corresponding to each initial semantic unit based on the part of speech of each initial semantic unit; obtaining the semantic density score corresponding to each initial semantic unit according to the inverse document frequency corresponding to each initial semantic unit and the frequency of occurrence of each initial semantic unit in the target transcription text; mapping each initial semantic unit and the core topic to the target semantic space to obtain the first vector corresponding to each initial semantic unit and the second vector corresponding to the core topic, obtaining the cosine similarity between each first vector and the second vector, and obtaining the topic relevance score corresponding to each initial semantic unit according to each cosine similarity. By identifying the semantically important nodes in the text and using them as reliable reference points for alignment, the alignment process of semantics and audio text is combined, thereby improving the accuracy of audio and text alignment.
[0105] Step S13, dynamically assigning timestamps to the initial audio data based on each of the rhythm change rate indices to obtain corresponding target audio data, dividing the target audio data into different target audio segments based on each of the anchor point positions corresponding to each of the target semantic units, and aligning each of the target audio segments with the target transcription text based on each of the timestamps.
[0106] In this embodiment, based on the aforementioned RVI, this embodiment further defines a timestamp density allocation function: , and assign timestamps accordingly:
[0107] ;
[0108] in, is the timestamp density at time point t, is the base timestamp density, is the RVI impact factor (value range [0,1]), is the rhythm change rate index at time point t.
[0109] Timestamp density determines the number of timestamps allocated per unit time. In areas with large speech rate changes, many pauses, or obvious stress, the RVI value is higher, resulting in a higher timestamp density, thus more accurately reflecting the speech rhythm characteristics.
[0110] For each text unit (word, phrase, etc.), its duration is calculated as:
[0111] ;
[0112] in, is the time length of unit w, is the content of unit w (e.g., number of syllables), is the cumulative timestamp density of the corresponding audio segment.
[0113] One challenge faced by the audio-text alignment method in this embodiment is how to control computational complexity while ensuring accuracy. Especially when processing long audio content, the computational complexity of traditional global alignment algorithms often grows quadratically, resulting in low processing efficiency. The system in this embodiment designs a hierarchical progressive alignment strategy, which significantly improves efficiency and accuracy through a divide-and-conquer approach. This strategy decomposes the alignment task into two stages: coarse alignment and fine alignment. The coarse alignment stage uses the previously identified semantic anchor points to establish a preliminary alignment framework and divides the long audio content into multiple relatively independent paragraphs. This step significantly reduces the overall complexity of the algorithm, because subsequent fine alignment only needs to be performed in smaller paragraphs, avoiding the high complexity of global calculations. The pseudocode for the above process is as follows:
[0114] Algorithm:hierarchical_alignment;
[0115] Input: preprocessed audio A, preprocessed text T, anchor point;
[0116] Output: aligned timestamp;
[0117] function hierarchical_alignment(A, T, anchors):
[0118] / / Step 1: Use anchor points for coarse segmentation;
[0119] segments=segment_using_anchors(A, T, anchors)
[0120] / / Step 2: Initialize the final alignment timestamp;
[0121] final_timestamps = empty array of length T
[0122] / / Step 3: Process each paragraph independently;
[0123] for segment in segments:
[0124] audio_segment = segment.audio
[0125] text_segment = segment.text
[0126] / / Step 3a: Check for potential mismatches;
[0127] mismatches=detect_mismatches(audio_segment, text_segment)
[0128] / / Step 3b: Apply adaptive fault tolerance mechanism;
[0129] if mismatches:
[0130] audio_segment, text_segment=apply_error_tolerance(audio_segment, text_segment, mismatches)
[0131] Step 3c: extract the rhythm features of this paragraph;
[0132] rhythm_features=extract_rhythm_features(audio_segment)
[0133] / / Step 3d: Calculate RVI for this paragraph;
[0134] segment_RVI=calculate_RVI(rhythm_features)
[0135] / / Step 3e: Fine alignment using rhythm-aware timestamp assignment;
[0136] segment_timestamps=fine_align(audio_segment, text_segment, segment_RVI)
[0137] / / Step 3f: Add to final timestamp;
[0138] final_timestamps[segment.start_idx:segment.end_idx]=segment_timestamps
[0139] / / Step 4: Apply global consistency constraints;
[0140] final_timestamps=apply_global_constraints(final_timestamps)
[0141] return final_timestamps
[0142] The rough alignment stage first uses the semantic anchors identified earlier to establish a preliminary alignment framework. These anchors usually have high confidence and can provide a reliable reference for overall alignment. The system uses these anchors to calculate preliminary global mapping relationships and divide long audio content into multiple relatively independent segments. Although this step is not very accurate, it can provide a global structure and lay the foundation for subsequent fine alignment.
[0143] The key to coarse alignment is how to build a global mapping using limited anchor points. The system designs an interpolation-based mapping construction method. For the content between two anchor points, the system calculates the approximate mapping relationship of the middle point based on the time position of the anchor points and the content length. This method takes into account the changes in content density and avoids the deviation that may be caused by simple linear interpolation. For the content before or after the anchor point, the system uses extrapolation to estimate the mapping relationship of the boundary area based on the alignment characteristics of the adjacent area.
[0144] This coarse alignment method significantly reduces the overall complexity of the algorithm because it avoids detailed comparisons on a global scale and only needs to deal with the matching relationships of a small number of key points. Even for hours of long audio, the coarse alignment stage can be completed in a few seconds, providing a good initial value for the fine alignment stage.
[0145] The fine alignment stage applies a dedicated fine alignment algorithm to each paragraph. At this stage, the system combines speech rhythm features and local semantic structure to achieve precise alignment within the paragraph. Since the paragraph size is controlled, the overall efficiency is still acceptable even if a fine algorithm with higher computational complexity is used.
[0146] A key challenge of the hierarchical strategy is the granularity selection of paragraph division. The system designs an adaptive paragraph division method to dynamically adjust the paragraph size according to content complexity, speech clarity and anchor point distribution. In areas with simple content and clear speech, the system will divide larger paragraphs; while in areas with complex content or unclear speech, the system will divide smaller paragraphs to increase processing accuracy; accordingly, the aforementioned process of dividing the target audio data into different target audio segments based on the anchor point positions corresponding to each target semantic unit may specifically include: determining the information complexity corresponding to different parts of the target audio data, and dynamically setting the target paragraph lengths corresponding to different parts of the target audio data according to each information complexity and each anchor point position; wherein the information complexity is used to characterize the amount of information contained in different parts of the target audio data; dividing the target audio data into different target audio segments based on each target paragraph length; wherein the head of any target audio segment contains the same audio data as the tail of the previous adjacent target audio segment, and the tail of any target audio segment contains the same audio data as the head of the next adjacent target audio segment; that is, to ensure the coherence between paragraphs, the system applies special processing at the paragraph boundaries. One method is to set overlapping areas, that is, adjacent paragraphs share a small part of the content, and optimize the alignment consistency of these overlapping areas to eliminate the incoherence that may occur at the splicing of paragraphs. Another method is to apply global consistency constraints, requiring the alignment results of all paragraphs to meet the rationality of global order and time distribution, and adjust the relationship between paragraphs through iterative optimization.
[0147] In addition, for the case where the audio and the transcribed text are inconsistent, this embodiment designs a mismatch region identification method based on a local similarity matrix, and identifies possible mismatch regions by analyzing the local correspondence between text and audio features. This method can not only detect obvious inconsistencies (such as missing or additional content), but also identify more subtle inconsistencies (such as synonymous substitution, word order adjustment), and align according to the type of inconsistency between the audio and the transcribed text. By using semantic anchors to preliminarily align the audio and the transcribed text, and then accurately aligning them in paragraphs, the overall complexity of the algorithm is significantly reduced, and the alignment error of one paragraph will not propagate to other paragraphs, thereby improving the robustness of the system. This solves the problem of error propagation in traditional global alignment algorithms; by allocating timestamps according to the rhythm of speech, the timestamp allocation is more in line with the inherent laws of natural speech, thereby significantly improving the alignment accuracy, especially when processing complex real speech content. The effect is particularly significant. Experiments show that when facing complex real speech (including characteristics such as speech rate changes, pauses, repetitions, and corrections), the alignment accuracy of the audio-text alignment method in this embodiment is improved by 15-30% compared with the alignment accuracy of the traditional method, and the computational efficiency is improved by 40-60%. This system has broad application prospects in the field of audio-text alignment, especially in scenarios such as automatic video subtitle generation, speech-assisted learning platforms, oral history archive processing, and conference record automation. By implementing this system, the alignment accuracy can be effectively improved, especially in processing speech speed changes, spoken language characteristics, and long audio content.
[0148] It can be seen that the present application can more accurately reflect the real speech rhythm characteristics by using the non-uniform distribution strategy based on RVI; by integrating semantic analysis into the audio-text alignment process, it provides semantic guidance for the alignment process and improves the accuracy of alignment; by identifying the mismatching areas and mismatching types between audio and text, and selecting the corresponding adjustment method for audio-text alignment, the robustness of the alignment process is improved; through the coarse and fine two-stage alignment, the computational complexity is significantly reduced while maintaining high precision.
[0149] Based on the above embodiments, it can be seen that the present application describes the overall process of aligning audio with text. In order to make the technical solution in the present embodiment more complete, the present application will next describe the audio text alignment process when the audio and text do not match. Figure 4 As shown, the embodiment of the present invention discloses a specific audio text alignment method, including:
[0150] Step S21: construct a similarity matrix between the target audio data and the target transcribed text, and obtain a matching pattern between the target audio data and the target transcribed text according to the similarity matrix.
[0151] It is understandable that in practical applications, there are often inconsistencies between the transcribed text and the audio content, such as transcription errors, spoken language features (repetition, correction, modal particles), etc. These inconsistencies can cause traditional alignment algorithms to fail or produce erroneous results. This embodiment designs an adaptive fault-tolerant mechanism that can detect and intelligently handle various inconsistencies.
[0152] This embodiment designs a mismatch region identification method based on a local similarity matrix, and its specific process is as follows: Figure 5 As shown, first, it is necessary to detect the mismatching areas in the audio and text, and then make appropriate adjustments to the audio or text according to the corresponding mismatch type in order to align the audio with the text. By analyzing the local correspondence between text and audio features, possible mismatching areas can be identified. This method can not only detect obvious inconsistencies (such as missing or additional content), but also more subtle inconsistencies (such as synonymous substitutions and word order adjustments). Specifically, the system constructs a local similarity matrix of text and audio features and analyzes the matching patterns therein. Ideally, the values on the main diagonal should be high, indicating a sequential match; in areas where there are inconsistencies, the matrix will show specific patterns, such as diagonal breaks, discrete high-value points, etc. By analyzing these patterns, the system can accurately identify possible mismatching areas. Mismatching area identification is the key to the adaptive fault-tolerance mechanism. The system identifies mismatching areas through the local similarity matrix L:
[0153] ;
[0154] in, It is an audio segment and text units The local similarity of is a similarity function based on the matching degree between acoustic features and text representation.
[0155] Step S22: obtaining a difference region between the target audio data and the target transcription text based on the matching pattern, obtaining a difference type corresponding to the difference region, and aligning each of the target audio segments with the target transcription text according to an alignment strategy corresponding to the difference type.
[0156] In this embodiment, after detecting inconsistencies, the system classifies the inconsistencies into multiple types, such as transcription errors, oral corrections, repeated expressions, modal particles, etc. This classification is based on the structural features and contextual relationships of the mismatched regions, using linguistic knowledge and pattern matching technology. For different types of inconsistencies, the system uses special processing strategies to ensure that the alignment quality is not affected.
[0157] Specifically, mismatched areas appear as specific patterns in the matrix. The system analyzes these patterns to identify possible mismatch types:
[0158] ;
[0159] in, Is the mismatch type of judgment, is the probability that a mismatch belongs to a specific type given the local similarity matrix L and the context.
[0160] After identifying the mismatched area, the system performs mismatch type analysis. Different types of inconsistencies have different characteristic patterns and processing strategies. The system classifies inconsistencies into the following main types: Transcription error: the content in the text does not match the audio, which may be a recognition error or manual transcription error; Spoken correction: the speaker makes self-corrections during the expression process, such as "this week there will be a meeting on Wednesday, no, Thursday"; Repeated expression: the speaker repeats certain content to emphasize or clarify, such as "this time, this type of problem must be solved"; Modal particles and filler words: common modal particles, pause words, etc. in spoken language, such as "um", "that", "you know"; Content omission: some content in the audio is omitted in the transcription, or some content in the transcription is omitted in the audio.
[0161] This embodiment adopts a special processing strategy for type mismatches: for transcription errors, the system applies a fuzzy matching algorithm to find the actual text that the audio content may correspond to; for spoken corrections, the system identifies the content before and after the correction and adjusts the timestamp allocation appropriately; for repeated expressions, the system identifies repeated content units and looks for corresponding patterns that appear multiple times in the audio; for modal particles and filler words, the system identifies these special elements through a special model and gives them appropriate treatment during the alignment process.
[0162] In addition, this embodiment does not simply apply predetermined rules, but dynamically adjusts the processing strategy according to the specific features and context of the inconsistency. For example, when it is recognized that a specific speaker is accustomed to using certain filler words, the system will optimize the corresponding recognition parameters; when processing professional content, the system will adjust the matching sensitivity of professional terms; in transcriptions with uneven quality, the system will dynamically adjust the credibility weights of different regions. This adaptive fault-tolerant mechanism enables the system to handle various inconsistencies gracefully, significantly improving its robustness in real application scenarios. Even in the case of poor transcription quality or containing a large number of spoken features, the system can still maintain a high alignment accuracy, greatly expanding the scope of application scenarios.
[0163] It can be seen that this application can obtain complex factors such as speech rate changes, pauses, intonation fluctuations and emotional expressions of audio data through in-depth analysis of the rhythmic characteristics of audio data and the corresponding semantics of the transcribed text, thereby improving the accuracy of audio and text alignment; by allocating timestamps to audio data according to the rhythm change rate index corresponding to the audio data, it avoids the use of a simple uniform distribution strategy, and can ensure that the timestamps of highly complex segments in the audio are more dense, thereby ensuring the alignment accuracy of highly complex positions in the audio.
[0164] See also Figure 6 As shown, an embodiment of the present invention discloses an audio text alignment device, comprising:
[0165] The semantic analysis module 11 is used to obtain the initial audio data and the target transcribed text corresponding to the initial audio data, obtain the rhythm change rate index corresponding to the initial audio data at different time nodes, and perform semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of the speech rhythm in the initial audio data;
[0166] An anchor point position determination module 12, used to determine, from each of the initial semantic units, a target semantic unit whose importance is greater than a preset importance threshold according to each of the importance levels, and preliminarily match each of the target semantic units with the initial audio data to determine the anchor point position of each of the target semantic units in the initial audio data;
[0167] The audio-text alignment module 13 is used to dynamically assign timestamps to the initial audio data based on each rhythm change rate index to obtain corresponding target audio data, divide the target audio data into different target audio segments based on each anchor point position corresponding to each target semantic unit, and align each target audio segment with the target transcription text based on each timestamp.
[0168] It can be seen that this application can obtain complex factors such as speech rate changes, pauses, intonation fluctuations and emotional expressions of audio data through in-depth analysis of the rhythmic characteristics of audio data and the corresponding semantics of the transcribed text, thereby improving the accuracy of audio and text alignment; by allocating timestamps to audio data according to the rhythm change rate index corresponding to the audio data, it avoids the use of a simple uniform distribution strategy, and can ensure that the timestamps of highly complex segments in the audio are more dense, thereby ensuring the alignment accuracy of highly complex positions in the audio.
[0169] In some specific embodiments, the semantic analysis module 11 may specifically include:
[0170] A weight determination unit, used to obtain the speech rate change, tone information and pause conditions corresponding to the initial audio data at different time nodes, and determine the target weights corresponding to the speech rate change, the tone information and the pause conditions respectively;
[0171] The first data fusion unit is used to weightedly fuse the speech rate change, the pitch information and the pause situation corresponding to different time nodes based on each target weight, so as to obtain the rhythm change rate index corresponding to the initial audio data at different time nodes.
[0172] In some specific embodiments, the semantic analysis module 11 may specifically include:
[0173] A parameter adjustment unit is used to obtain the speech scene and speech style corresponding to the initial audio data, automatically adjust the target parameters corresponding to the rhythm change rate index according to the speech scene and the speech style, and obtain the rhythm change rate index corresponding to the initial audio data at different time nodes according to the corresponding adjusted parameters.
[0174] In some specific embodiments, the semantic analysis module 11 may specifically include:
[0175] A semantic analysis submodule, for performing semantic analysis on the target transcribed text to obtain a grammatical centrality score, a semantic density score and a topic relevance score corresponding to each of the initial semantic units in the target transcribed text; wherein the grammatical centrality score represents the position of each of the initial semantic units in the corresponding syntactic structure, the semantic density score represents the amount of information contained in each of the initial semantic units, and the topic relevance score represents the degree of association between each of the initial semantic units and the core topic corresponding to the target transcribed text;
[0176] The second data fusion unit is used to perform weighted fusion on the grammatical centrality score, the semantic density score and the topic relevance score to obtain the importance of each initial semantic unit in the target transcription text.
[0177] In some specific embodiments, the semantic analysis submodule may specifically include:
[0178] A first score acquisition unit, configured to acquire the grammatical centrality score corresponding to each of the initial semantic units based on the part of speech of each of the initial semantic units;
[0179] A second score acquisition unit acquires the semantic density score corresponding to each of the initial semantic units according to the inverse document frequency corresponding to each of the initial semantic units and the occurrence frequency of each of the initial semantic units in the target transcription text;
[0180] The third score acquisition unit is used to map each of the initial semantic units and the core topic to the target semantic space to obtain a first vector corresponding to each of the initial semantic units and a second vector corresponding to the core topic, obtain the cosine similarity between each of the first vectors and the second vectors, and obtain the topic relevance score corresponding to each of the initial semantic units according to each of the cosine similarities.
[0181] In some specific embodiments, the audio-text alignment module 13 may specifically include:
[0182] A paragraph length setting unit, used to determine the information complexity corresponding to different parts of the target audio data, and dynamically set the target paragraph lengths corresponding to different parts of the target audio data according to each information complexity and each anchor point position; wherein the information complexity is used to characterize the amount of information contained in different parts of the target audio data;
[0183] A paragraph division unit is used to divide the target audio data into different target audio segments based on the length of each target paragraph; wherein the audio data contained in the head of any target audio segment is the same as the audio data contained in the tail of the previous adjacent target audio segment, and the audio data contained in the tail of any target audio segment is the same as the audio data contained in the head of the next adjacent target audio segment.
[0184] In some specific embodiments, the audio-text alignment module 13 may specifically include:
[0185] A matching pattern acquisition unit, used to construct a similarity matrix between the target audio data and the target transcribed text, and acquire a matching pattern between the target audio data and the target transcribed text according to the similarity matrix;
[0186] The audio-text alignment unit is used to obtain the difference area between the target audio data and the target transcription text based on the matching pattern, obtain the difference type corresponding to the difference area, and align each of the target audio segments with the target transcription text according to the alignment strategy corresponding to the difference type.
[0187] Furthermore, the present application also discloses an electronic device. Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0188] Figure 7A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the audio-text alignment method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0189] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0190] In addition, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0191] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the audio text alignment method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0192] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned audio-text alignment method is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0193] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0194] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0195] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0196] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0197] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An audio text alignment method, characterized in that: include: Acquire initial audio data and a target transcribed text corresponding to the initial audio data, acquire the rhythm change rate index corresponding to the initial audio data at different time nodes, and perform semantic analysis on the target transcribed text to acquire the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data; Determine, from each of the initial semantic units, a target semantic unit whose importance is greater than a preset importance threshold according to each of the importance levels, and preliminarily match each of the target semantic units with the initial audio data to determine an anchor point position of each of the target semantic units in the initial audio data; Based on each of the rhythm change rate indices, a timestamp is dynamically assigned to the initial audio data to obtain corresponding target audio data, the target audio data is divided into different target audio segments based on the anchor point positions corresponding to each of the target semantic units, and each of the target audio segments is aligned with the target transcription text based on each of the timestamps.
2. The audio-text alignment method according to claim 1, characterized in that: The obtaining of the rhythm change rate indexes corresponding to the initial audio data at different time nodes includes: Obtaining speech rate changes, tone information, and pauses corresponding to the initial audio data at different time nodes, and determining target weights corresponding to the speech rate changes, the tone information, and the pauses, respectively; The speech rate change, the pitch information and the pause conditions corresponding to different time nodes are weightedly fused based on the target weights to obtain the rhythm change rate index corresponding to the initial audio data at different time nodes.
3. The audio-text alignment method according to claim 1, characterized in that: The obtaining of the rhythm change rate indexes corresponding to the initial audio data at different time nodes includes: The speech scene and speech style corresponding to the initial audio data are obtained, the target parameters corresponding to the rhythm change rate index are automatically adjusted according to the speech scene and the speech style, and the rhythm change rate index corresponding to the initial audio data at different time nodes is obtained according to the corresponding adjusted parameters.
4. The audio-text alignment method according to claim 1, characterized in that: The performing semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text includes: Performing semantic analysis on the target transcribed text to obtain a grammatical centrality score, a semantic density score, and a topic relevance score corresponding to each of the initial semantic units in the target transcribed text; wherein the grammatical centrality score represents the position of each of the initial semantic units in the corresponding syntactic structure, the semantic density score represents the amount of information contained in each of the initial semantic units, and the topic relevance score represents the degree of association between each of the initial semantic units and the core topic corresponding to the target transcribed text; The grammatical centrality score, the semantic density score, and the topic relevance score are weightedly fused to obtain the importance of each initial semantic unit in the target transcription text.
5. The audio-text alignment method according to claim 4, characterized in that: The obtaining of the grammatical centrality score, the semantic density score and the topic relevance score corresponding to each of the initial semantic units in the target transcription text includes: Acquire the grammatical centrality score corresponding to each of the initial semantic units based on the part of speech of each of the initial semantic units; Obtaining the semantic density scores respectively corresponding to the initial semantic units according to the inverse document frequencies corresponding to the initial semantic units and the occurrence frequencies of the initial semantic units in the target transcribed text; Map each of the initial semantic units and the core topic to the target semantic space to obtain a first vector corresponding to each of the initial semantic units and a second vector corresponding to the core topic, obtain the cosine similarity between each of the first vectors and the second vectors, and obtain the topic relevance score corresponding to each of the initial semantic units based on each of the cosine similarities.
6. The audio-text alignment method according to claim 1, characterized in that: The step of dividing the target audio data into different target audio segments based on the anchor point positions corresponding to the target semantic units includes: Determine the information complexity corresponding to different parts of the target audio data, and dynamically set the target paragraph lengths corresponding to different parts of the target audio data according to each information complexity and each anchor point position; wherein the information complexity is used to characterize the amount of information contained in different parts of the target audio data; The target audio data is divided into different target audio segments based on the length of each target paragraph; wherein the audio data contained in the head of any target audio segment is the same as the audio data contained in the tail of the previous adjacent target audio segment, and the audio data contained in the tail of any target audio segment is the same as the audio data contained in the head of the next adjacent target audio segment.
7. The audio-text alignment method according to any one of claims 1 to 6, characterized in that: The aligning each of the target audio segments with the target transcription text based on each of the timestamps includes: Constructing a similarity matrix between the target audio data and the target transcribed text, and obtaining a matching pattern between the target audio data and the target transcribed text according to the similarity matrix; Based on the matching pattern, a difference area between the target audio data and the target transcription text is obtained, a difference type corresponding to the difference area is obtained, and each of the target audio segments is aligned with the target transcription text according to an alignment strategy corresponding to the difference type.
8. An audio text alignment device, characterized in that: include: A semantic analysis module is used to obtain initial audio data and a target transcribed text corresponding to the initial audio data, obtain the rhythm change rate index corresponding to the initial audio data at different time nodes, and perform semantic analysis on the target transcribed text to obtain the importance of each initial semantic unit in the target transcribed text; wherein the rhythm change rate index is used to characterize the degree of change of speech rhythm in the initial audio data; An anchor point position determination module, used to determine, from each of the initial semantic units, a target semantic unit whose importance is greater than a preset importance threshold according to each of the importance levels, and preliminarily match each of the target semantic units with the initial audio data to determine the anchor point position of each of the target semantic units in the initial audio data; An audio-text alignment module is used to dynamically assign timestamps to the initial audio data based on each rhythm change rate index to obtain corresponding target audio data, divide the target audio data into different target audio segments based on each anchor point position corresponding to each target semantic unit, and align each target audio segment with the target transcription text based on each timestamp.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the audio-text alignment method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the audio-text alignment method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Speech synthesis method and device based on hierarchical emotion distribution, equipment and medium
CN119207372A
Audio-to-text conversion method and device, electronic equipment and storage medium
CN119380719A
Intelligent MV generation method, system and device based on AIGC and medium
CN119788886A
Automatic caption synchronization and positioning
EP3839953A1
User interface linking analyzed segments of transcripts with extracted key points
US20230223016A1