Audio playing method and device, storage medium and electronic device
Patent Information
- Application Number
- CN202310800087.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-30
AI Technical Summary
[0003]本申请提供了一种音频播放方法、装置、存储介质以及电子设备,以解决采用固定语速的音频练习口语听力效果差的技术问题
[0020]本申请实施例提供的上述技术方案与现有技术相比具有如下优点:本申请实施例提供的该方法,通过播放不同语速等级的原始音频来确定目标对象的语速等级,然后根据目标对象的语速等级来确定要播放的目标音频,从而可以根据用户的等级来播放不同语速等级的音频,实现了准确播放音频的效果,进一步提升了用户练习口语听力的效果。
Smart Images

Figure CN116682449B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more particularly to an audio playback method, apparatus, storage medium, and electronic device. Background Technology
[0002] In existing technologies, users can practice their speaking and listening skills by listening to audio. However, the speaking speed in spoken and listening exercises is fixed, and since users have varying levels of speaking and listening ability, using audio with a fixed speaking speed to practice speaking and listening skills is ineffective. Summary of the Invention
[0003] This application provides an audio playback method, apparatus, storage medium, and electronic device to solve the technical problem of poor results in practicing spoken language listening comprehension using audio at a fixed speaking speed.
[0004] In a first aspect, this application provides an audio playback method, comprising: playing original audio at different speech rates to determine the speech rate level of a target audience listening to the original audio based on the speech rate level of the original audio; and determining the target audio to be played based on the speech rate level of the target audience.
[0005] As an optional example, determining the target audio to be played based on the speech rate level of the target object includes: determining a third recommended audio from the audio to be recommended that has a speech rate level greater than that of the target object; and determining the third recommended audio as the target audio.
[0006] As an optional example, before playing the original audio at different speech rates, the method further includes: acquiring audio to be adjusted at different speech rates; adjusting each of the audio to be adjusted into multiple adjusted audio at different speech rates, wherein the adjusted audio includes the audio to be adjusted; and determining the adjusted audio as the original audio.
[0007] As an optional example, adjusting each of the above-mentioned audio to be adjusted into multiple adjusted audios with different speech rates includes: performing the following operations on each of the above-mentioned audio to be adjusted: dividing the above-mentioned audio to be adjusted into multiple segments, wherein each of the above-mentioned segments includes a word list of words in the above-mentioned segment, the start and end time points of each word in the above-mentioned segment, and the start and end time points of the above-mentioned segment; lengthening or shortening each of the above-mentioned segments to obtain multiple adjusted audios of the above-mentioned audio to be adjusted, wherein the number of words played per unit time of each of the above-mentioned adjusted audios is the same as the number of words played per unit time corresponding to a speech rate level.
[0008] Secondly, this application provides an audio playback device, comprising: a playback module for playing original audio at different speech rates to determine the speech rate level of a target audience listening to the original audio based on the speech rate level of the original audio; and a determination module for determining the target audio to be played based on the speech rate level of the target audience.
[0009] As an optional example, the playback module includes: an acquisition unit, configured to acquire feedback results for each of the original audio files after playing the original audio files at different speech rate levels, wherein the feedback results represent the feedback operation performed by the target object after listening to the original audio files; and to determine the speech rate level of the target object based on the feedback results.
[0010] As an optional example, the feedback result is either a confirmation selection operation or a non-selection operation of the target object. The acquisition unit includes: a first determining subunit, configured to determine from the original audio that the feedback result is a first audio of the confirmation selection operation; determine a weighted average result of the speech rate level of the first audio; and determine the weighted average result as the speech rate level of the target object.
[0011] As an optional example, the feedback result is either a correct recognition operation or an incorrect recognition operation. The acquisition unit includes: a second determining subunit, used to determine from the original audio that the feedback result is a second audio that is a correct recognition operation; to count the proportion of audio at each speech rate level in the second audio; and to determine the speech rate level of the second audio with the highest speech rate level among the second audios with a proportion greater than a predetermined threshold as the speech rate level of the target object.
[0012] As an optional example, the above-mentioned determining module includes: a first determining unit, configured to determine a first recommended audio from the audio to be recommended that matches the speech rate level of the target object; and to determine the first recommended audio as the target audio.
[0013] As an optional example, the above-mentioned determining module includes: a second determining unit, configured to, when it is determined that a second recommended audio will be played, adjust the speech rate level of the second recommended audio to the speech rate level of the target object if the speech rate level of the second recommended audio does not match the speech rate level of the target object; and determine the second recommended audio after adjusting the speech rate level as the target audio.
[0014] As an optional example, the second determining unit includes: an adjustment subunit for determining the number of words played per unit time corresponding to the speech rate level of the target object; dividing the second recommended audio into multiple segments, wherein each segment includes a word list of words in the segment, the start and end times of each word in the segment, and the start and end times of the segment; lengthening or shortening each segment to make the number of words played per unit time of the second recommended audio the same as the number of words played per unit time corresponding to the speech rate level of the target object.
[0015] As an optional example, the above-mentioned determining module includes: a third determining unit, used to determine a third recommended audio from the audio to be recommended that has a speech rate level greater than that of the target object; and to determine the third recommended audio as the target audio.
[0016] As an optional example, the above-described apparatus further includes: an adjustment module for acquiring audio to be adjusted at different speech rates before playing original audio at different speech rates; adjusting each of the audio to be adjusted into multiple adjusted audio at different speech rates, wherein the adjusted audio includes the audio to be adjusted; and determining the adjusted audio as the original audio.
[0017] As an optional example, the above adjustment module includes: an adjustment unit for performing the following operations on each of the above-mentioned audio to be adjusted: dividing the above-mentioned audio to be adjusted into multiple segments, wherein each of the above-mentioned segments includes a word list of words in the above-mentioned segment, the start and end time points of each word in the above-mentioned segment, and the start and end time points of the above-mentioned segment; lengthening or shortening each of the above-mentioned segments to obtain multiple adjusted audios of the above-mentioned audio to be adjusted, wherein the number of words played per unit time of each of the above-mentioned adjusted audios is the same as the number of words played per unit time corresponding to a speech rate level.
[0018] Thirdly, this application provides an electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the memory stores a computer program, and the processor is configured to implement any of the above-described audio playback methods when executing the computer program.
[0019] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for performing any of the above-described audio playback methods of this application.
[0020] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application determines the speech rate level of the target object by playing original audio at different speech rate levels, and then determines the target audio to be played based on the speech rate level of the target object. Thus, different speech rate levels of audio can be played according to the user's level, achieving the effect of accurate audio playback and further improving the effect of the user's oral listening practice. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0024] Figure 1 A flowchart illustrating an audio playback method provided in this application embodiment;
[0025] Figure 2 A system flowchart of an audio playback method provided in an embodiment of this application;
[0026] Figure 3 A schematic diagram illustrating the determination of a user's speech rate level in an audio playback method provided in this application embodiment;
[0027] Figure 4 A schematic diagram illustrating an audio playback method according to the user's speech rate level, provided in an embodiment of this application;
[0028] Figure 5 A flowchart illustrating an audio playback method for using audio to train a user's spoken language, provided as an embodiment of this application;
[0029] Figure 6 A schematic diagram illustrating the adjustment of the speech rate level of the audio to be played in an audio playback method provided in an embodiment of this application;
[0030] Figure 7 A schematic diagram of a training audio recognition model for an audio playback method provided in an embodiment of this application;
[0031] Figure 8 A schematic diagram illustrating the changing speech rate level of an audio playback method provided in an embodiment of this application;
[0032] Figure 9 This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application;
[0033] Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0036] Figure 1 This is a flowchart illustrating an audio playback method provided in an embodiment of this application. Figure 1 As shown, the above audio playback method includes:
[0037] S102, Play the original audio at different speech rate levels to determine the speech rate level of the target audience listening to the original audio based on the speech rate level of the original audio.
[0038] S104, determine the target audio to be played based on the speech rate level of the target audience.
[0039] The above audio playback method can be implemented through a terminal. The terminal can be a mobile phone, tablet, laptop, desktop computer, smart hardware such as a smart bracelet, or smart learning machine, or it can be executed by software programs on the terminal or smart hardware. For example, users can learn spoken language through the terminal, regardless of the specific language. The terminal can play raw audio at different speech rates. The raw audio can be pre-acquired or stored in an audio library. Different speech rates refer to the speed of speech, specifically the reading of the same passage over different time intervals. Alternatively, the different number of words read within the same time interval can also be used as an indicator. In other words, audio at different speech rate levels has a different reading speed.
[0040] The terminal plays raw audio at different speech rate levels, which can be used to determine the target audience, i.e., the user listening to the raw audio. For example, the terminal plays raw audio at levels 1 to 10 without displaying the speech rate level; the user listens to the audio to determine their own speech rate level. The higher the audio level, the faster the speech rate.
[0041] Once the user's speech rate level is determined, the target audio to be played can be determined based on that level. The target audio can be obtained from an audio library or by adjusting existing audio within that library.
[0042] The method provided in this application determines the speech rate level of the target object by playing original audio at different speech rate levels, and then determines the target audio to be played based on the speech rate level of the target object. Thus, different speech rate levels of audio can be played according to the user's level, achieving the effect of accurate audio playback and further improving the user's practice of speaking and listening skills.
[0043] As an optional example, playing original audio at different speech rate levels to determine the speech rate level of the target object listening to the original audio based on the speech rate level of the original audio includes: after playing original audio at different speech rate levels, obtaining feedback results for each original audio, wherein the feedback results are used to represent the feedback operation of the target object after listening to the original audio; and determining the speech rate level of the target object based on the feedback results.
[0044] In this embodiment, when determining the user's speech rate level by playing the original audio, user feedback can be obtained. The feedback can be the user's action or content after listening to the original audio. Taking user feedback as an example, after each original audio clip is played, the user can choose whether they understand it or not; this is the feedback action. If the user understands, it means they can comprehend audio at that speech rate level. If they don't understand, it means they don't understand audio at that speech rate level. By checking which speech rate levels the user can understand, the user's speech rate level can be determined. For example, if a user can basically understand level 5 but cannot understand level 6, then the user's speech rate level is level 5.
[0045] As an optional example, the feedback result is either a confirmation or non-selection operation of the target object. Determining the speech rate level of the target object based on the feedback result includes: identifying a first audio recording from the original audio recording where the feedback result is a confirmation operation; determining a weighted average of the speech rate levels of the first audio recording; and determining the weighted average as the speech rate level of the target object.
[0046] In this embodiment, when determining the user's speech rate level by playing the original audio, user feedback can be obtained. The feedback result can be whether the user made a selection. The user's selection is an active action after listening to the original audio. It indicates that the user believes their level is similar to the speech rate level of the selected audio. For example, if audio with speech rate levels from 1 to 10 is played, and the user selects levels 4, 5, or 6 (the user doesn't know the speech rate level but determines whether it suits them based on their listening experience, whether it's too fast or too slow), it means the user feels their level is between 4 and 6. Therefore, in this embodiment, the weighted average of the speech rate levels of the selected audio is determined as the user's speech rate level. The weight can be determined based on the time taken for the user to select. The faster the user selects, the larger the weight; the slower the user selects, indicating a longer consideration time, the smaller the weight.
[0047] As an optional example, the feedback result is either a correct recognition operation or an incorrect recognition operation. Determining the speech rate level of the target object based on the feedback result includes: identifying a second audio file from the original audio file where the feedback result is a correct recognition operation; calculating the proportion of audio files with each speech rate level in the second audio file; and determining the speech rate level of the second audio file with the highest speech rate level among the second audio files with a proportion greater than a predetermined threshold as the speech rate level of the target object.
[0048] In this embodiment, after playing the original audio, corresponding exercises can be displayed for the user to answer. The feedback result is whether the answer is correct. After playing the original audio, the accuracy rate of the user's answers for each speech rate level is calculated. If the accuracy rate exceeds a predetermined threshold, it means that the user can understand the audio at that level. The highest speech rate level that the user can understand is determined as the user's speech rate level; any higher speech rate level will be incomprehensible to the user.
[0049] As an optional example, determining the target audio to be played based on the speech rate level of the target audience includes: determining a first recommended audio from the audio to be recommended that matches the speech rate level of the target audience; and determining the first recommended audio as the target audio.
[0050] Once the user's speaking speed level is determined, audio can be recommended to them. There are several recommendation scenarios. One scenario is recommending audio at the same speaking speed level to the user.
[0051] If a user's speech rate level is 5, and the audio library contains audio with speech rate levels from 1 to 10, the user can select the audio with a speech rate level of 5 to play. Each audio file in the audio library corresponds to a speech rate level, which is known.
[0052] As an optional example, determining the target audio to be played based on the speech rate level of the target audience includes: if it is determined that a second recommended audio will be played, and if the speech rate level of the second recommended audio does not match the speech rate level of the target audience, adjusting the speech rate level of the second recommended audio to match the speech rate level of the target audience; and determining the second recommended audio after adjusting the speech rate level as the target audio.
[0053] In this embodiment, the recommendation scenario can be to recommend audio that matches the user's speech rate level. The recommended audio in this embodiment is either the audio selected by the user or the audio recommended by the system. If the recommended audio matches the user's speech rate level, no processing is needed. If the recommended audio does not match the user's speech rate level, the speech rate level of the recommended audio needs to be adjusted to match the user's speech rate level. For example, if the user selects a level 4 audio to play, the level 4 audio is adjusted to a level 5 audio before playback. If the audio library contains a level 5 audio corresponding to the level 4 audio, no adjustment is needed, and the level 5 audio is directly retrieved from the audio library for playback. If the audio library does not contain a level 5 audio corresponding to the level 4 audio, the level 4 audio is adjusted to a level 5 audio before playback. In this case, the audio library retains both level 4 and level 5 audio for that purpose.
[0054] As an optional example, if it is determined that a second recommended audio will be played, and the speech rate level of the second recommended audio does not match the speech rate level of the target audience, adjusting the speech rate level of the second recommended audio to match the speech rate level of the target audience includes: determining the number of words played per unit time corresponding to the speech rate level of the target audience; dividing the second recommended audio into multiple segments, wherein each segment includes a word list of words in the segment, the start and end times of each word in the segment, and the start and end times of the segment; lengthening or shortening each segment to make the number of words played per unit time of the second recommended audio the same as the number of words played per unit time corresponding to the speech rate level of the target audience.
[0055] In this embodiment, adjusting the speech rate level of audio involves adjusting the speech rate of words within the audio. Specifically, the audio can first be segmented into multiple segments. The purpose of segmentation is to remove interference from content other than words. After segmentation, each segment is lengthened or shortened. When lengthening or shortening a segment, different segments of an audio are lengthened or shortened proportionally. For example, if a level 4 audio is divided into 3 segments, and the word playback speed of each segment is lengthened to 1.2 times the original speed, then a segment that originally took 3 seconds to play will take 3.6 seconds after lengthening. The lengthened segments constitute the lengthened audio; if the original audio is level 4, the lengthened audio might be level 5, level 6, etc. Each speech rate level has a requirement for the number of words played per unit time. By lengthening or shortening segments according to the number of words played per unit time, different speech rate levels of audio are obtained.
[0056] As an optional example, determining the target audio to be played based on the speech rate level of the target audience includes: identifying a third recommended audio from the audio to be recommended that has a speech rate level greater than that of the target audience; and determining the third recommended audio as the target audio.
[0057] In this embodiment, the recommended scenario can be to assist users in improving their speaking speed level. For example, if a user's speaking speed level is 5, audio clips with speaking speed levels of 5 and 6 can be recommended to the user. These audio clips with speaking speed levels of 5 and 6 can be played as a training library to improve the user's oral communication skills. After a long period of training, the user's speaking speed level may rise to level 6, thereby achieving the training goal.
[0058] As an optional example, before playing the original audio at different speech rates, the method further includes: acquiring audio to be adjusted at different speech rates; adjusting each audio to be adjusted into multiple adjusted audio at different speech rates, wherein the adjusted audio includes the audio to be adjusted; and determining the adjusted audio as the original audio.
[0059] In this embodiment, the original audio can be located in an audio library. The audio library contains audio at different speech rate levels. The original audio can be obtained by adjusting the audio to be adjusted. There can be one or more audio tracks to be adjusted. For example, an audio track at level 3 can be adjusted to obtain original audio at levels 1-10, and then the original audio is added to the audio library. Thus, a large number of original audio tracks at different speech rate levels can be obtained using a small number of audio tracks to be adjusted, enriching the audio library.
[0060] As an optional example, adjusting each audio to be adjusted into multiple adjusted audios at different speech rates includes: performing the following operations on each audio to be adjusted: dividing the audio to be adjusted into multiple segments, wherein each segment includes a word list of words in the segment, the start and end times of each word in the segment, and the start and end times of the segment; lengthening or shortening each segment to obtain multiple adjusted audios of the audio to be adjusted, wherein the number of words played per unit time in each adjusted audio is the same as the number of words played per unit time corresponding to a speech rate level.
[0061] When adjusting an audio file to multiple audio files with different speech rates, the audio file can be lengthened or shortened by different proportions. Different proportions correspond to different speech rates, thus obtaining the original audio files at different speech rates. Specific methods can be found in the process described above, and will not be repeated here.
[0062] The number of words in the original audio in this embodiment is not limited. The original audio may include word audio, sentence audio, music audio, article audio, and audio content such as reviews and recitations. Figure 2 This is a system flowchart for this embodiment.
[0063] Users can provide audio, which is then converted into multiple original audio files with different speaking speeds and added to an audio library. The system plays original audio files at random speaking speed levels for users to listen to. Users choose the voice they can understand or the one they deem appropriate. The user's speaking speed level is assessed based on a weighted average of the selected audio's speaking speed levels. This weighted averaging algorithm can also be implemented using other more efficient methods, such as training an AI model algorithm: inputting multiple audio sentences of varying lengths and their corresponding speaking speed levels to obtain a final listening comprehension speaking speed level. After a user's level is determined, different services can be provided, such as listening training, adjusting the speaking speed level of the audio in the learning materials to be played, or recommending audio files that match the user's level.
[0064] Figure 3 This diagram illustrates how to determine a user's speech rate level. First, the speech rate level is adjusted based on different audio sources to generate original audio. Then, a portion of the original audio is played, and user feedback is collected to determine the user's speech rate level.
[0065] Figure 4 This is an illustration of how learning materials are recommended based on the user's speaking speed level. The system selects suitable learning materials from a variety of speaking speed levels based on the user's speed level, such as recommending or playing materials that match the user's speed level.
[0066] Figure 5 This is a flowchart for using audio to train users' spoken language. Based on the user's speaking speed level, a plan is developed, and appropriately level audio is selected for playback to train the user's spoken language skills. Regular tests are conducted, and the training plan can be adjusted based on the test results.
[0067] Figure 6 This diagram illustrates how to adjust the speech rate level of the audio to be played. If the speech rate level of the audio to be played differs from the user's speech rate level, the process involves adjusting the speech rate level of the audio to match the user's speech rate level before playback.
[0068] The following methods can be used to adjust the speech rate level of audio.
[0069] First, the audio to be adjusted is segmented. Segmentation can be based on the concentration of words in the audio. The concentration of words, or word density, determines the playback duration of each word. Words with consecutive playback durations are grouped into a segment, ensuring a consistent playback speed for each segment. Each segment includes information such as the start and end times of each word and the duration of the segment.
[0070] Figure 7 This is a schematic diagram for training an Automatic Speech Recognition (ASR) model. The model takes a labeled audio segment as input and outputs the corresponding text after segmenting the segment.
[0071] Speech rate adjustment involves dividing the audio into multiple segments, which can then be adjusted using a formula. The speech rate formula refers to the number of words spoken per minute. It can be calculated using the following formula: Speech Rate (WPM) = Total Words ÷ Speaking Time (minutes). For example, if 200 words are spoken in two minutes, then the speech rate is 100 WPM.
[0072] Figure 8 This is a diagram illustrating the variation of speech rate levels. For example, increasing the speech rate of an audio file at a speech rate level of 5 to an audio file at a speech rate level of 8.
[0073] Figure 9 This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application. Figure 9 As shown, the aforementioned audio playback device includes:
[0074] The playback module 902 is used to play original audio at different speech rates to determine the speech rate of the target audience listening to the original audio based on the speech rate of the original audio.
[0075] The determination module 904 is used to determine the target audio to be played based on the speech rate level of the target object.
[0076] In this embodiment, the terminal plays raw audio at different speech rate levels, which can be used to determine the speech rate level of the target audience, i.e., the user listening to the raw audio. For example, the terminal plays raw audio at levels 1 to 10 without displaying the speech rate level; the user listens to the audio to determine their own speech rate level. The higher the audio level, the faster the speech rate.
[0077] Once the user's speech rate level is determined, the target audio to be played can be determined based on that level. The target audio can be obtained from an audio library or by adjusting existing audio within that library.
[0078] The method provided in this application determines the speech rate level of the target object by playing original audio at different speech rate levels, and then determines the target audio to be played based on the speech rate level of the target object. Thus, different speech rate levels of audio can be played according to the user's level, achieving the effect of accurate audio playback and further improving the user's practice of speaking and listening skills.
[0079] For other examples of this embodiment, please refer to the examples above, which will not be repeated here.
[0080] like Figure 10 As shown in the figure, this application provides an electronic device including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0081] Memory 113 is used to store computer programs;
[0082] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the audio playback method provided in any of the foregoing method embodiments, including:
[0083] Play the original audio at different speech rates to determine the speech rate of the target audience listening to the original audio based on the speech rate of the original audio.
[0084] The target audio to be played is determined based on the speech rate level of the target audience.
[0085] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the audio playback method provided in any of the foregoing method embodiments.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0088] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0089] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An audio playback method, characterized in that, include: Play original audio at different speech rates to determine the speech rate level of the target audience listening to the original audio based on the speech rate level of the original audio. The target audio to be played is determined based on the speech rate level of the target audience; The step of determining the target audio to be played based on the speech rate level of the target object includes: when it is determined that a second recommended audio will be played, if the speech rate level of the second recommended audio does not match the speech rate level of the target object, adjusting the speech rate level of the second recommended audio to match the speech rate level of the target object; and determining the second recommended audio after adjusting the speech rate level as the target audio. The step of adjusting the speech rate of the second recommended audio to match the speech rate of the target object when the speech rate level of the second recommended audio does not match the speech rate level of the target object, when it is determined that the second recommended audio will be played, includes: determining the number of words played per unit time corresponding to the speech rate level of the target object; dividing the second recommended audio into multiple segments, wherein each segment includes a word list of words in the segment, the start and end times of each word in the segment, and the start and end times of the segment; lengthening or shortening each segment to make the number of words played per unit time of the second recommended audio the same as the number of words played per unit time corresponding to the speech rate level of the target object; The step of playing original audio at different speech rates to determine the speech rate level of the target object listening to the original audio includes: after playing the original audio at different speech rates, obtaining a feedback result for each original audio, wherein the feedback result is used to represent the feedback operation of the target object after listening to the original audio; and determining the speech rate level of the target object based on the feedback result. Wherein, when the feedback result is a confirmation selection operation or a non-selection operation of the target object, determining the speech rate level of the target object based on the feedback result includes: determining a first audio from the original audio where the feedback result is the confirmation selection operation; determining a weighted average result of the speech rate level of the first audio; and determining the weighted average result as the speech rate level of the target object, with a larger weight for faster user selection.
2. The method according to claim 1, characterized in that, When the feedback result is a correct recognition operation or an incorrect recognition operation, determining the speech rate level of the target object based on the feedback result includes: The feedback result is determined from the original audio to be the second audio of the correct recognition operation; Calculate the proportion of audio at each speech rate level in the second audio; The speech rate level of the second audio with the highest speech rate level among the second audios whose proportion is greater than a predetermined threshold is determined as the speech rate level of the target object.
3. The method according to claim 1, characterized in that, The step of determining the target audio to be played based on the speech rate level of the target object includes: Determine a first recommended audio from the audio to be recommended that matches the speech rate level of the target object; The first recommended audio is determined as the target audio.
4. The method according to claim 1, characterized in that, The step of determining the target audio to be played based on the speech rate level of the target object includes: A third recommended audio is selected from the audio to be recommended, whose speech rate level is greater than that of the target object; The third recommended audio is determined as the target audio.
5. The method according to claim 1, characterized in that, Before playing the original audio at different speech rates, the method further includes: Obtain audio files to be adjusted at different speech rate levels; Each of the audio files to be adjusted is converted into multiple adjusted audio files at different speech rate levels, wherein the adjusted audio files include the audio files to be adjusted. The adjusted audio is identified as the original audio.
6. The method according to claim 5, characterized in that, The process of adjusting each of the audio files to be adjusted into multiple different speech rate levels includes: For each of the audio files to be adjusted, perform the following operations: The audio to be adjusted is divided into multiple segments, wherein each segment includes a word list of words in the segment, the start and end time points of each word in the segment, and the start and end time points of the segment. Each of the aforementioned segments is lengthened or shortened to obtain multiple adjusted audios of the audio to be adjusted, wherein the number of words played per unit time for each of the adjusted audios is the same as the number of words played per unit time corresponding to a speech rate level.
7. An audio playback device, characterized in that, include: The playback module is used to play original audio at different speech rates to determine the speech rate level of the target audience listening to the original audio based on the speech rate level of the original audio. The determination module is used to determine the target audio to be played based on the speech rate level of the target object; The step of determining the target audio to be played based on the speech rate level of the target object includes: when it is determined that a second recommended audio will be played, if the speech rate level of the second recommended audio does not match the speech rate level of the target object, adjusting the speech rate level of the second recommended audio to match the speech rate level of the target object; and determining the second recommended audio after adjusting the speech rate level as the target audio. The step of adjusting the speech rate of the second recommended audio to match the speech rate of the target object when the speech rate level of the second recommended audio does not match the speech rate level of the target object, when it is determined that the second recommended audio will be played, includes: determining the number of words played per unit time corresponding to the speech rate level of the target object; dividing the second recommended audio into multiple segments, wherein each segment includes a word list of words in the segment, the start and end times of each word in the segment, and the start and end times of the segment; lengthening or shortening each segment to make the number of words played per unit time of the second recommended audio the same as the number of words played per unit time corresponding to the speech rate level of the target object; The step of playing original audio at different speech rates to determine the speech rate level of the target object listening to the original audio includes: after playing the original audio at different speech rates, obtaining a feedback result for each original audio, wherein the feedback result is used to represent the feedback operation of the target object after listening to the original audio; and determining the speech rate level of the target object based on the feedback result. Wherein, when the feedback result is a confirmation selection operation or a non-selection operation of the target object, determining the speech rate level of the target object based on the feedback result includes: determining a first audio from the original audio where the feedback result is the confirmation selection operation; determining a weighted average result of the speech rate level of the first audio; and determining the weighted average result as the speech rate level of the target object, with a larger weight for faster user selection.
8. An electronic device, characterized in that, include: At least one communication interface; At least one bus connected to the at least one communication interface; At least one processor connected to the at least one bus; At least one memory connected to the at least one bus, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium storing computer-executable instructions for performing the method described in any one of claims 1 to 6 of this application.
Citation Information
Patent Citations
Computer-based training system and method for enhancing language listening comprehension
US20040243418A1