Speech rate adjustment method and system
By analyzing voiced and unvoiced segments in speech signals and adjusting the speech rate to meet time constraints, the problem of high audio description production costs has been solved, achieving more efficient audio description production.
Patent Information
- Application Number
- CN202111553476.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-06
- Filing Date
- 2021-12-17
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-12-27
AI Technical Summary
In existing technologies, the production of audio descriptions is costly and cumbersome, and traditional methods require repeated adjustments to the speech signal to adapt to time constraints.
By analyzing the voiced and unvoiced signal segments in the speech signal, the amount of speech frame adjustment is calculated, and the speech rate is adjusted according to the total adjustment duration to reduce or increase the number of speech frames in the voiced signal segments, thus forming the adjusted speech signal.
It effectively reduces the number of times recordings are repeated, thus lowering the production cost of audio descriptions.
Smart Images

Figure CN115705838B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to speech processing technology, and in particular, to a speech speed adjustment method and system. BACKGROUND
[0002] Audio descriptions of audiovisuals provide a spoken description of changes in events such as character actions and scene changes. Audio descriptions can improve the accessibility of visual images for blind, low vision, or other visually impaired persons.
[0003] The creation of audio descriptions is both expensive and cumbersome. Traditionally, producers of audiovisuals hire scriptwriters and voice talent to create audio descriptions. In this traditional approach, a scriptwriter identifies segments of the audiovisual that require audio description by watching the content of the audiovisual, determines the points in time at which the audio description is to be inserted and estimates the available time for the audio description, and creates a script for the descriptive audio based on the content of the segments. Then, a voice talent records the audio description to fit within the available time based on the script. Often, the scriptwriter and voice talent must repeat the foregoing process multiple times to optimize the resulting audio description under the time constraints of the available time. For example, the scriptwriter can modify the segments that require audio description to obtain new available times, or the scriptwriter can rewrite the script to accommodate shorter available times, or the voice talent can repeatedly adjust the speed of his or her speech to accommodate the available time. Because of these challenges, the price of traditional audio description services is quite high. SUMMARY
[0004] In view of the above problems of the prior art, the present application aims to provide a speech speed adjustment method and system, which is suitable for adjusting the speech speed of a speech signal to provide an audio track that meets time constraints, thereby reducing the number of repeated recordings and substantially reducing the production cost of the audio track.
[0005] One embodiment of the present application provides a speech speed adjustment method, which includes obtaining an original speech signal of a plurality of words and a total adjustment time length; analyzing the original speech signal to obtain voiced signal segments and unvoiced signal segments corresponding to each word; calculating a frame adjustment amount based on the total adjustment time length and a unit frame time length; and adjusting the number of frames of at least one voiced signal segment based on the frame adjustment amount to form an adjusted speech signal.
[0006] Another embodiment of the present application provides a speech speed adjustment system, which includes a storage unit, an analysis unit, and an adjustment unit. The analysis unit is coupled to the storage unit, and the adjustment unit is coupled to the storage unit and the analysis unit. The storage unit temporarily stores an original speech signal of a plurality of words. The analysis unit analyzes the original speech signal to obtain voiced signal segments and unvoiced signal segments corresponding to each word. The adjustment unit calculates a frame adjustment amount based on a total adjustment time length and a unit frame time length, and adjusts the number of frames of at least one voiced signal segment based on the frame adjustment amount to form an adjusted speech signal.
[0007] In summary, the speech speed adjustment method of any embodiment is applicable to adjust the speech speed of a speech signal to provide a sound track meeting a time limit, thereby reducing the number of repeated recording and substantially reducing the production cost of sound tracks.
[0008] The present application is described in detail below with reference to the accompanying drawings and specific embodiments, but is not limited thereto. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 Flowchart of the speech speed adjustment method of some embodiments;
[0010] Figure 2 Functional block diagram of the speech speed adjustment system of some embodiments;
[0011] Figure 3 Schematic diagram of the original speech signal of an embodiment;
[0012] Figure 4 Functional block diagram of the speech speed adjustment system of some embodiments;
[0013] Figure 5 Flowchart of the speech speed adjustment method of some embodiments;
[0014] Figure 6 Flowchart of the speech speed adjustment method of some embodiments;
[0015] Figure 7 Functional block diagram of the speech speed adjustment system of some embodiments;
[0016] Figure 8 Flowchart of the speech speed adjustment method of some embodiments.
[0017] Wherein, the reference signs
[0018] 10: Speech speed adjustment system
[0019] 110: Storage unit
[0020] 120: Analysis unit
[0021] 130: Adjustment unit
[0022] 140: Conversion unit
[0023] 150: Judgment unit
[0024] 160: Recognition unit
[0025] 170: Update unit
[0026] 180: Audio processing unit
[0027] 190: merging unit
[0028] Si: original speech signal
[0029] N1: total adjustment duration
[0030] So: adjusted speech signal
[0031] Wo: word
[0032] Sp1-Sp7: word sound signal
[0033] F1: speech waveform
[0034] F2: sound spectrum
[0035] Z1: unvoiced signal segment
[0036] Z2: voiced signal segment
[0037] Vi: original video signal
[0038] Vo: spoken video signal
[0039] Mi: silent continuous picture
[0040] Mo: voiced continuous picture
[0041] XL: description text
[0042] Tt: total duration of silent content
[0043] S21-S27: steps
[0044] S11-S15: steps
[0045] S14': step
[0046] S26': step DETAILED DESCRIPTION
[0047] The structural principles and working principles of the present application will be described in detail below in combination with the drawings:
[0048] Referring to Figure 1 With Figure 2 , an embodiment of the present application provides a speech speed adjustment system 10, comprising a storage unit 110, an analysis unit 120, and an adjustment unit 130. The analysis unit 120 is coupled to the storage unit 110, and the adjustment unit 130 is coupled to the storage unit 110 and the analysis unit 120. Here, the storage unit 110 temporarily stores an original speech signal of at least one word. The present application also provides a speech speed adjustment method, which can be implemented by the speech speed adjustment system. For clarity, the original speech signal of multiple words will be described below as an example.
[0049] In this regard, the speech speed adjustment method can be applied to adjust the speech speed of pronunciation of a single word or a sentence, or to adjust the speech speed of audio description of a video. Details of the speech speed adjustment method will be described later.
[0050] In one embodiment, the analysis unit 120 can obtain and analyze the original speech signal Si (step S21) to obtain voiced sound signal segments and unvoiced sound signal segments corresponding to each word (step S22). The adjustment unit 130 can obtain the total adjustment duration N1 (step S21), calculate the number of speech frames to be removed (i.e., the speech frame adjustment amount) according to the total adjustment duration N1 and the unit speech frame duration (step S23), and adjust the number of speech frames of at least one of the voiced sound signal segments of the words according to the speech frame adjustment amount to form the adjusted speech signal So (step S24). Here, a speech frame is the smallest signal segment in speech signal processing, and the unit speech frame duration is the time length of one smallest signal segment.
[0051] In some embodiments, the original speech signal Si includes word sound signals Sp1-Sp7 of the words Wo, as shown in Figure 3 Each word sound signal Sp1-Sp7 includes a plurality of speech frames (hereinafter referred to as original speech frames).
[0052] Here, in linguistics, the sound of vocal cord vibration during pronunciation is called voiced sound, the sound of vocal cord non-vibration is called unvoiced sound, and there is also a consonant that has both voiced sound and unvoiced sound. In this case, when a word with a consonant is encountered, the voiced sound and the unvoiced sound can be distinguished first, and then the speech speed adjustment is performed.
[0053] In some embodiments of step S22, the analysis unit 120 analyzes the original speech signal Si to find the word sound signals Sp1-Sp7 of each word Wo (i.e., the signal segments corresponding to each word), and then analyzes each word sound signal Sp1-Sp7 to find the voiced sound signal segments (i.e., the signal segments corresponding to voiced sound pronunciation) and the unvoiced sound signal segments (i.e., the signal segments corresponding to unvoiced sound pronunciation) in each word sound signal Sp1-Sp7. For example, with reference to Figure 3 , the analysis unit 120 can convert the original speech signal Si from a time versus amplitude speech waveform F1 to a time versus frequency sound spectrum F2, and identify the word sound signals Sp1-Sp7 of each word Wo according to the energy distribution state. In Figure 3In the middle, for the voice waveform F1, the horizontal axis is time (seconds), and the vertical axis is amplitude (decibels); for the sound spectrum F2, the horizontal axis is time (seconds), and the vertical axis is frequency (hertz (Hz)). Then, the analysis unit 120 further identifies the voiced signal segment Z2 and the unvoiced signal segment Z1 according to the energy distribution state in the word sound signal Sp of each word Wo.
[0054] In an embodiment of step S23, the adjustment unit 130 removes the original sound frames in the voiced signal segment Z2 in a manner of removing one original sound frame at a fixed interval to form the adjusted speech signal So with a faster speech speed relative to the original speech signal Si. Here, the adjustment unit 130 is a sound signal processing of sound frame deletion on the voiced signal segment Z2 with a total amount of sound frames greater than the sound frame adjustment amount.
[0055] For example, assuming that the sound frame adjustment amount is 20 original sound frames deleted per word. When the voiced signal segment Z2 of the word sound signal Sp1 of one word Wo has 100 original sound frames, the adjustment unit 130 performs sound signal processing of deleting one original sound frame every 5 original sound frames on this voiced signal segment Z2.
[0056] In some embodiments, the adjustment unit 130 can calculate the sound frame adjustment amount of the original sound frames to be removed per word Wo according to the total adjustment time length, the unit sound frame time length, and the processing quantity (i.e., the number of word sound signals Sp1-Sp7 with the voiced signal segment Z2), and then delete the sound frame adjustment amount of original sound frames from each word sound signal with the voiced signal segment Z2. In some embodiments, after the sound frame adjustment amount, the adjustment unit 130 first confirms whether the number of sound frames of the voiced signal segment Z2 of the word sound signal with the voiced signal segment Z2 is greater than the current sound frame adjustment amount. When it is greater, the adjustment unit 130 performs sound frame deletion. Otherwise, the adjustment unit 130 excludes the word sound signal with the voiced signal segment Z2 to obtain a new processing quantity and recalculate the sound frame adjustment amount.
[0057] In another embodiment of step S23, the adjustment unit 130 inserts the sound frame adjustment amount of a sound frame (hereinafter referred to as a supplementary sound frame) into the voiced signal segment Z2 in a manner of inserting a sound frame at a fixed interval to form the adjusted speech signal So with a slower speech speed relative to the original speech signal Si.
[0058] For example, assuming that the sound frame adjustment amount is 20 supplementary sound frames added per word. When the voiced signal segment Z2 of the word sound signal Sp1 of one word Wo has 100 original sound frames, the adjustment unit 130 performs sound signal processing of inserting one supplementary sound frame every 5 original sound frames on this voiced signal segment Z2.
[0059] In some embodiments, the inserted supplemental frame is related to at least one original frame adjacent to the insertion position in the voiced signal segment Z2. In one embodiment, the inserted supplemental frame can be an average of the original frames adjacent to the insertion position in the voiced signal segment Z2. For example, with reference to the previous example, the adjustment unit 130 inserts a supplemental frame obtained by averaging the 5th and 6th original frames between the 5th and 6th original frames in the voiced signal segment Z2. In another embodiment, the inserted supplemental frame can be the previous original frame of the insertion position in the voiced signal segment Z2. For example, with reference to the previous example, the adjustment unit 130 inserts a supplemental frame obtained by copying the 5th original frame between the 5th and 6th original frames in the voiced signal segment Z2. In yet another embodiment, the inserted supplemental frame can be the next original frame of the insertion position in the voiced signal segment Z2. For example, with reference to the previous example, the adjustment unit 130 inserts a supplemental frame obtained by copying the 6th original frame between the 5th and 6th original frames in the voiced signal segment Z2. In other words, the adjustment unit 130 can adjust the number of frames removed or inserted per interval as appropriate, and the present disclosure is not limited in this regard.
[0060] In some embodiments, the speech rate adjustment system 10 can further generate a spoken video from the adjusted speech signal So and a dynamic video. Here, the adjusted speech signal corresponds to the silent content in the dynamic video. In some embodiments, the dynamic video can be a silent continuous picture (e.g., a GIF, etc.) or an original video (e.g., a movie or an animation, etc.), and the generated spoken video can be a spoken continuous picture (e.g., a spoken GIF, etc.) or a spoken video (e.g., a spoken movie or a spoken animation, etc.).
[0061] In some embodiments, the original speech signal Si and the adjusted speech signal So can be used to provide event descriptions of the image content (i.e., silent content) in a segment of a silent video in the original video. Here, the silent video refers to a video in which no one is speaking and there is no sound effect (e.g., a door opening sound, or a vehicle approaching sound, etc.) that has a dramatic meaning. In other words, the original speech signal Si and the adjusted speech signal So provide speech descriptions of changes in events such as character actions and / or scene changes in the segment of the silent video at different speech rates.
[0062] In some embodiments, with reference to Figure 4 , the speech rate adjustment system 10 can further include a video processing unit 180. The video processing unit 180 is coupled to the storage unit 110 and the adjustment unit 130. Here, with reference to Figure 4 and Figure 5 , the video processing unit 180 can generate a spoken video Vo from the adjusted speech signal So and an original video Vi (step S26).
[0063] In some embodiments of step S21, the original speech signal Vi and the total adjustment duration N1 can be provided by an external device coupled to the speech rate adjustment system 10, and / or inputted by a user via a user interface to the speech rate adjustment system 10.
[0064] In some embodiments, referring to Figure 4 , the speech rate adjustment system 10 can further comprise a conversion unit 140 and a determination unit 150. The conversion unit 140 is coupled to the storage unit 110, and the determination unit 150 is coupled to the conversion unit 140, the analysis unit 120 and the adjustment unit 130.
[0065] In some embodiments of step S21, referring to Figure 4 and Figure 5 , the conversion unit 140 receives the description text XL corresponding to the silent content (step S11), and converts the description text XL into the original speech signal Si (step S12). The description text XL records an event description composed of words Wo, and the event description describes the silent content in the original audio-visual video Vi.
[0066] After the conversion, the conversion unit 140 temporarily stores the generated original speech signal Si in the storage unit 110, and the determination unit 150 compares the total duration of the original speech signal Si with the total duration Tt of the silent content (step S13).
[0067] In some embodiments, the determination unit 150 confirms whether the total duration of the original speech signal Si is greater than the total duration Tt of the silent content by the comparison step (step S13) (step S14).
[0068] When the total duration of the original speech signal Si is greater than the total duration Tt of the silent content, the determination unit 150 calculates the time difference between the total duration of the original speech signal Si and the total duration Tt of the silent content to obtain the total adjustment duration N1 (step S15), and provides the total adjustment duration N1 to the adjustment unit 130. In addition, the determination unit 150 also enables the analysis unit 120 to start analyzing the generated original speech signal Si, so that the adjustment unit 130 generates the adjusted speech signal So with a faster speech rate relative to the original speech signal Si according to the total adjustment duration N1 and the analysis result (i.e., steps S22-S24 are sequentially performed).
[0069] When the total duration of the original speech signal Si is not greater than the total duration Tt of the silent content, the determination unit 150 does not enable the analysis unit 120 (i.e., steps S22-S24 are not sequentially performed). At this time, if the speech rate adjustment system 10 has an audio-visual processing unit 180, the determination unit 150 causes the audio-visual processing unit 180 to generate a spoken audio-visual video Vo according to the original speech signal Si and the original audio-visual video Vi (step S26).
[0070] In some embodiments, with reference to Figure 4 and Figure 6 , the determination unit 150 confirms whether the total time length of the original speech signal Si is equal to the total time length Tt of the non-speech content (step S14') by comparing (step S13).
[0071] When the total time length of the original speech signal Si is not equal to the total time length Tt of the non-speech content, the determination unit 150 calculates the time difference between the total time length of the original speech signal Si and the total time length Tt of the non-speech content to obtain the total adjustment time length N1 (step S15), and provides the adjustment unit 130. Also, the determination unit 150 enables the analysis unit 120 to start analyzing the generated original speech signal Si, so that the adjustment unit 130 generates the adjusted speech signal So according to the total adjustment time length N1 and the analysis result (i.e., steps S22-S24 are sequentially executed). When the total time length of the original speech signal Si is greater than the total time length Tt of the non-speech content, the adjustment unit 130 generates the adjusted speech signal So with a faster speech rate relative to the original speech signal Si (step S24). When the total time length of the original speech signal Si is less than the total time length Tt of the non-speech content, the adjustment unit 130 generates the adjusted speech signal So with a slower speech rate relative to the original speech signal Si (step S24).
[0072] When the total time length of the original speech signal Si is equal to the total time length Tt of the non-speech content, the determination unit 150 does not enable the analysis unit 120 (i.e., steps S22-S24 are not sequentially executed). At this time, if the speech rate adjustment system 10 has the video processing unit 180, the determination unit 150 causes the video processing unit 180 to generate the spoken video Vo according to the original speech signal Si and the original video Vi (step S26).
[0073] In some embodiments, the video processing unit 180 can combine the speech signal (i.e., the original speech signal Si or the adjusted speech signal So) and the original video Vi in the form of mixing, replacing, or associating to form the spoken video Vo with the speech signal as the audio of the non-speech content.
[0074] In an embodiment, the video processing unit 180 receives the original video Vi and separates the original video Vi into an original audio track and a non-speech video. Then, the video processing unit 180 mixes the original audio track with the speech signal (i.e., the original speech signal Si or the adjusted speech signal So) to form an adjusted audio track, and combines the adjusted audio track and the non-speech video into the spoken video Vo by synchronizing the adjusted audio track and the non-speech video.
[0075] In another embodiment, the audio processing unit 180 receives the original audio-visual video Vi and separates the original audio-visual video Vi into the original audio track and the silent video. Then, the audio processing unit 180 replaces the audio track segment corresponding to the silent content in the original audio track with the speech signal to form the adjusted audio track, and then combines the adjusted audio track and the silent video into the spoken audio-visual video Vo by synchronizing the adjusted audio track and the silent video.
[0076] In yet another embodiment, the audio processing unit 180 receives the original audio-visual video Vi and finds out the audio track segment corresponding to the silent content in the original audio-visual video Vi. Then, the audio processing unit 180 creates a replacement signal of the speech signal for the audio track segment, and generates the spoken audio-visual video Vo containing the original audio-visual video Vi, the speech signal and the replacement signal. Assuming that the audio track segment corresponding to the silent content is between the first playback time and the second playback time in the original audio track. At this time, during the playback of the spoken audio-visual video Vo, the replacement signal is triggered at the first playback time to replace the original audio track with the speech signal, and then the original audio track is played again from the position of the second playback time.
[0077] In some embodiments, semantic recognition can be performed after the adjusted speech signal So is generated, and the adjusted speech signal So is outputted only when the semantic of the adjusted speech signal So can be recognized.
[0078] In some embodiments, referring to Figure 4 , the speech speed adjustment system 10 can further include a recognition unit 160 and an updating unit 170. The recognition unit 160 is coupled to the adjustment unit 130 and the updating unit 170.
[0079] Referring to Figure 4 and Figure 5 or Figure 4 and Figure 6 , after the adjustment unit 130 generates the adjusted speech signal So, the recognition unit 160 first detects the semantic of the adjusted speech signal So (step S25) to confirm whether the semantic can be recognized (i.e. to confirm whether the content of the speech played after the adjusted speech signal So can be recognized). Here, the semantic detection technology is well known to those skilled in the art, and therefore will not be described here.
[0080] When the semantic is not recognizable, the recognition unit 160 does not output the adjusted speech signal So, and causes the update unit 170 to update the description text to reduce the number of words constituting the event description (step S27). Then, the update unit 170 provides the updated description text to the conversion unit 140 for conversion (i.e. step S12 is performed subsequently). At this time, if the speech rate adjustment system 10 has the video processing unit 180, the recognition unit 160 does not output the generated adjusted speech signal So to the video processing unit 180.
[0081] When the semantic is recognizable, the recognition unit 160 outputs the adjusted speech signal So. At this time, if the speech rate adjustment system 10 has the video processing unit 180, the recognition unit 160 outputs the adjusted speech signal So to the video processing unit 180 for video processing (i.e. step S26 is performed).
[0082] In some embodiments, the storage unit 110, the analysis unit 120, the adjustment unit 130, the conversion unit 140, the judgment unit 150, the recognition unit 160, the update unit 170, and the video processing unit 180 can be implemented by single or multiple processing components.
[0083] In some embodiments, the original speech signal Si and the adjusted speech signal So can provide corresponding speech for the silent content in the silent continuous picture. In other words, the original speech signal Si and the adjusted speech signal So are speech providing the same content but at different speech rates. For example, the silent continuous picture can present a variation of mouth shape for the pronunciation of a word or a sentence, and the original speech signal Si and the adjusted speech signal So provide the pronunciation of the word or the sentence.
[0084] In some embodiments, referring to Figure 7 , the speech rate adjustment system 10 can further include a merging unit 190. The merging unit 190 is coupled to the adjustment unit 130. Referring to Figure 7 and Figure 8 , the merging unit 190 receives a silent continuous picture Mi, and combines the adjusted speech signal So with the silent continuous picture Mi into a voiced continuous picture Mo by synchronizing the adjusted speech signal So with the silent continuous picture Mi (step S26').
[0085] In some embodiments, the storage unit 110, the analysis unit 120, the adjustment unit 130, and the merging unit 190 can be implemented by single or multiple processing components.
[0086] In some embodiments, the storage unit 110 can be implemented by single or multiple memories. The aforementioned processing components can be microprocessors, microcontrollers, central processing units, programmable logic controllers, logic circuits, analog circuits, digital circuits, or any analog and / or digital devices operating signals based on operation instructions.
[0087] In some embodiments, the speech rate adjustment method of any embodiment can be implemented by a computer program product, such that when the computer program product is loaded into a computer and executed, the speech rate adjustment method of any embodiment can be completed. In some embodiments, the computer program product can be a non-transitory recording medium, and the above-mentioned program is stored in the non-transitory recording medium for loading into the computer. In some embodiments, the above-mentioned program itself can be a computer program product, and is transmitted to the computer through wired or wireless means.
[0088] In summary, the speech rate adjustment method of any embodiment is suitable for adjusting the speech rate of a speech signal to provide a sound file meeting a time limit, thereby reducing the number of repeated recordings and greatly reducing the production cost of the sound file.
[0089] Of course, the present application can have other various embodiments, and those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application. However, these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.
Claims
1. A speech rate adjustment method characterized by comprising: The method comprises: obtaining an original speech signal of a plurality of words and a total adjustment time length; analyzing the original speech signal to obtain a voiced signal segment and an unvoiced signal segment corresponding to each of the words; calculating an audio frame adjustment amount according to the total adjustment time length and a unit audio frame time length; and adjusting an audio frame number of at least one of the voiced signal segments of the plurality of words according to the audio frame adjustment amount to form an adjusted speech signal.
2. The speech rate adjustment method according to claim 1, characterized by, The step of adjusting the audio frame number of the at least one of the voiced signal segments of the plurality of words according to the audio frame adjustment amount to form the adjusted speech signal comprises: removing at least one original audio frame corresponding to the audio frame adjustment amount from the at least one voiced signal segment in a fixed interval manner, or inserting at least one supplementary audio frame corresponding to the audio frame adjustment amount, to form the adjusted speech signal with a different speech speed from the original speech signal, wherein the at least one supplementary audio frame is related to the at least one original audio frame adjacent to the insertion position.
3. The speech rate adjustment method according to claim 1, characterized by, The method further comprises: generating a spoken video according to the adjusted speech signal and a dynamic video, wherein the adjusted speech signal corresponds to a silent content in the dynamic video.
4. The speech rate adjustment method according to claim 3, characterized by, The steps of obtaining the original speech signal of the plurality of words and the total adjustment time length comprise: receiving a description text, wherein an event description of the silent content composed of the plurality of words is recorded in the description text; and converting the description text into the original speech signal, wherein the total adjustment time length is a time difference between a total time length of the original speech signal and a total time length of the silent content.
5. The speech rate adjustment method according to claim 4, characterized by, The method further comprises: judging whether a semantic of the adjusted speech signal is recognizable; when the semantic of the adjusted speech signal is not recognizable, reducing a word number of the plurality of words constituting the event description in the description text, converting an updated description text into the original speech signal, and then performing the steps of analyzing the original speech signal to obtain the voiced signal segment and the unvoiced signal segment corresponding to each of the words; and when the semantic is recognizable, outputting the adjusted speech signal.
6. A speech rate adjustment system characterized by comprising: The method comprises: a storage unit temporarily storing an original speech signal of a plurality of words; an analysis unit coupled to the storage unit, analyzing the original speech signal to obtain a voiced signal segment and an unvoiced signal segment corresponding to each of the words; and an adjustment unit coupled to the storage unit and the analysis unit, calculating an audio frame adjustment amount according to a total adjustment time length and a unit audio frame time length, and adjusting an audio frame number of at least one of the voiced signal segments of the plurality of words according to the audio frame adjustment amount to form an adjusted speech signal.
7. The speech rate adjustment system of claim 6, wherein, The method further comprises: a merging unit coupled to the adjustment unit, receiving a silent continuous picture, and combining the adjusted speech signal and the silent continuous picture into a voiced continuous picture.
8. The speech rate adjustment system of claim 6, wherein, The method further comprises: an audio processing unit coupled to the adjustment unit, generating a spoken audio video according to the adjusted speech signal and an original audio video, wherein the adjusted speech signal is used to provide an event description of a silent content in the original audio video.
9. The speech rate adjustment system of claim 8, wherein, The method further comprises: a conversion unit coupled to the storage unit, receiving a description text and converting the description text into the original voice signal, wherein the event description composed of the plurality of words is recorded in the description text, and the total adjustment time length is a time difference between a total time length of the original voice signal and a total time length of the non-speech content.
10. The speech rate adjustment system of claim 9, wherein, Further comprising: a recognition unit coupled to the adjustment unit, judging whether a semantic of the adjusted voice signal is recognizable; and an update unit coupled to the recognition unit and the conversion unit, updating the description text to reduce a number of words of the plurality of words constituting the event description when the semantic is not recognizable, and providing the updated description text to the conversion unit for conversion.
Citation Information
Patent Citations
Method for adjusting speech rate and system using the same
TWI790705B