Method for Adjusting Playing Progress, Vehicle, and Audio Playing Device
By recognizing the user's voice prompts for adjusting playback progress in the audio playback device, the system automatically adjusts the playback progress of the audio file, solving the problem of inflexible playback progress adjustment in existing technologies and achieving more efficient playback progress control and improved user experience.
Patent Information
- Application Number
- CN202110641732.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-06-09
AI Technical Summary
The current audio playback devices offer limited flexibility in adjusting audio playback progress, making them inconvenient for users.
By collecting and recognizing the user's voice prompts for adjusting playback progress, the system identifies keywords in the target audio text and automatically adjusts the playback progress of the audio file, supporting voice-controlled playback progress adjustment.
It improves the flexibility and accuracy of playback progress adjustment, simplifies user operation, and enhances the user experience.
Smart Images

Figure CN115457948B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and particularly to a method for adjusting playing progress, a vehicle, and an audio playing device. Background Art
[0002] An audio playing device can play audio and display a playing progress bar of the audio on its display screen.
[0003] A user can touch the playing progress bar to adjust the playing progress of the audio. However, the flexibility of this way of adjusting the playing progress of the audio is relatively low. Summary of the Invention
[0004] This application provides a method for adjusting playing progress, a vehicle, and an audio playing device, which can solve the problem that the flexibility of the way of adjusting the playing progress of the audio in the related art is relatively low. The technical solutions are as follows:
[0005] On the one hand, a vehicle is provided, which is characterized in that the vehicle includes: a processor; and the processor is configured to:
[0006] During the process of playing a target audio file, collect a progress adjustment voice for the target audio file, where the target audio file is associated with a target audio text;
[0007] Recognize the progress adjustment voice to obtain a first voice recognition text;
[0008] If the first voice recognition text includes a content keyword belonging to the target audio text, start playing the target audio file from the time stamp corresponding to the content keyword.
[0009] On the other hand, an audio playing device is provided, and the audio playing device includes: a processor; and the processor is configured to:
[0010] During the process of playing a target audio file, collect a progress adjustment voice for the target audio file, where the target audio file is associated with a target audio text;
[0011] Recognize the progress adjustment voice to obtain a first voice recognition text;
[0012] If the first voice recognition text includes a content keyword belonging to the target audio text, start playing the target audio file from the time stamp corresponding to the content keyword.
[0013] Optionally, the audio playing device further includes a display screen; and the processor is configured to:
[0014] If the target audio text includes multiple content keywords, then display the target audio text on the display screen, and the display effects of the multiple content keywords in the target audio text are different from those of other words and sentences except the multiple content keywords;
[0015] In response to a selection operation on a target content keyword among the multiple content keywords, start playing the target audio file from the time stamp corresponding to the target content keyword.
[0016] Optionally, the processor is configured to:
[0017] Display multiple first text segments included in the target audio text on the display screen, and each first text segment includes one content keyword;
[0018] Wherein, the multiple first text segments are partial text segments of the target audio text.
[0019] Optionally, the processor is configured to:
[0020] If the first speech recognition text further includes a direction keyword for indicating a progress adjustment direction, and at least two of the multiple content keywords are located in a second text segment of the target audio text, then display the target audio text;
[0021] Wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, then the second text segment is the text segment of the target audio text after the reference keyword currently being played; if the progress adjustment direction indicated by the direction keyword is rewind, then the second text segment is the partial text segment of the target audio text before the reference keyword;
[0022] The display effects of the content keywords in the second text segment are different from those of other words and sentences.
[0023] Optionally, the processor is configured to:
[0024] If the target audio text includes multiple content keywords, the first speech recognition text further includes a direction keyword for indicating a progress adjustment direction, and one target content keyword among the multiple content keywords is located in a second text segment of the target audio text, then start playing the target audio text from the time stamp corresponding to the target content keyword;
[0025] Wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment in the target audio text that is after the reference keyword currently being played; if the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is a partial text segment in the target audio text that is before the reference keyword.
[0026] Optionally, the processor is configured to:
[0027] If the target audio text includes multiple content keywords, and the historical number of times of playing the target audio file starting from the timestamp corresponding to the target content keyword among the multiple content keywords is greater than the number threshold, then play the target audio file starting from the timestamp corresponding to the target content keyword.
[0028] Optionally, the processor is further configured to:
[0029] Recognize the collected playback voice to obtain a second speech recognition text;
[0030] If multiple audio files to be played are determined based on the second speech recognition text, display the file identifiers of the multiple audio files to be played;
[0031] In response to a selection operation on the file identifier of the target audio file among the multiple audio files, play the target audio file.
[0032] On the other hand, a method for adjusting the playback progress is provided, which is applied to a vehicle; the method includes:
[0033] During the playback of the target audio file, collect the progress adjustment voice for the target audio file, and the target audio file is associated with a target audio text;
[0034] Recognize the progress adjustment voice to obtain a first speech recognition text;
[0035] If the first speech recognition text includes a content keyword belonging to the target audio text, play the target audio file starting from the timestamp corresponding to the content keyword.
[0036] On yet another hand, a method for adjusting the playback progress is provided, which is applied to an audio playback device; the method includes:
[0037] During the playback of the target audio file, collect the progress adjustment voice for the target audio file, and the target audio file is associated with a target audio text;
[0038] Recognize the progress adjustment voice to obtain a first speech recognition text;
[0039] If the content keyword belonging to the target audio text is included in the first speech recognition text, play the target audio file starting from the time stamp corresponding to the content keyword.
[0040] Optionally, the playing of the target audio file starting from the time stamp corresponding to the content keyword includes:
[0041] If the target audio text includes multiple content keywords, display the target audio text, and the display effect of the multiple content keywords in the target audio text is different from the display effect of other sentences except the multiple content keywords;
[0042] In response to a selection operation on a target content keyword among the multiple content keywords, play the target audio file starting from the time stamp corresponding to the target content keyword.
[0043] Optionally, the displaying of the audio text includes:
[0044] Display multiple first text segments included in the target audio text, and each first text segment includes one content keyword;
[0045] Wherein, the multiple first text segments are partial text segments of the target audio text.
[0046] In another aspect, an audio playback device is provided, and the audio playback device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for adjusting the playback progress as described in the above aspect is implemented.
[0047] In another aspect, a computer-readable storage medium is provided, and a computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the method for adjusting the playback progress as described in the above aspect.
[0048] In another aspect, a computer program product including instructions is provided. When the computer program product runs on an audio playback device, the audio playback device is enabled to execute the method for adjusting the playback progress as described in the above aspect.
[0049] The beneficial effects brought by the technical solution provided by this application at least include:
[0050] The present application provides a method for adjusting the playback progress, a vehicle, and an audio playback device. During the playback of a target audio file, the audio playback device can recognize the collected progress adjustment speech to obtain a first speech recognition text, and when it is determined that the first speech recognition text includes a content keyword in the target audio text associated with the target audio file, start playing the target audio file from the content keyword. Thus, it can be seen that the audio playback device can respond to the progress adjustment speech and adjust the playback progress of the target audio file, thereby effectively improving the flexibility of adjusting the playback progress.
[0051] Moreover, since it is easier to remember the content keyword than to remember the timestamp of a certain content keyword in the target audio text, when the user adjusts the playback progress of the audio file through the progress adjustment speech containing the content keyword, the accuracy of adjusting the playback progress can be ensured, and the user experience is better. Also, since there is no need for the user to touch the playback progress bar when adjusting the playback progress, the user operation is effectively simplified, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a schematic diagram of the implementation environment involved in the method for adjusting the playback progress provided by the embodiment of the present application;
[0054] Figure 2 It is a flowchart of a method for adjusting the playback progress provided by the embodiment of the present application;
[0055] Figure 3 It is a flowchart of another method for adjusting the playback progress provided by the embodiment of the present application;
[0056] Figure 4 It is a schematic diagram of selecting a target audio file provided by the embodiment of the present application;
[0057] Figure 5 It is a flowchart of a method for playing a target audio file provided by the embodiment of the present application;
[0058] Figure 6 It is a schematic diagram of selecting a target content keyword provided by the embodiment of the present application;
[0059] Figure 7 It is a block diagram of the structure of an audio playback device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0061] An embodiment of the present application provides an audio playback device, in which an audio playback application is installed, and the audio playback device has a voice interaction system. Among them, the voice interaction system may include: a speaker and a microphone. The audio playback device can play an audio file through the speaker and can collect voices through the microphone.
[0062] Optionally, the audio playback device may be a vehicle, a smart phone, a smart speaker, a computer, a Moving Picture Experts Group Audio Layer III (MP3) player, or a Moving Picture Experts Group Audio Layer IV (MP4) player.
[0063] For example, as Figure 1 shown, the audio playback device may be a vehicle 110. The vehicle 110 may be a sedan, a bus, or a truck. In this scenario, the audio playback device also has a central control system, and the audio playback application may be installed in the central control system.
[0064] An embodiment of the present application provides a method for adjusting the playback progress, which can be applied to an audio playback device, such as Figure 1 the audio playback device shown. Refer to Figure 1 , the method includes:
[0065] Step 101, during the process of playing a target audio file, collect progress adjustment voices for the target audio file.
[0066] During the process of the audio playback device running the audio playback application to play the target audio file, if the user needs the audio playback device to adjust the playback progress of the target audio file, the user can express the progress adjustment voices for the target audio file. Correspondingly, the audio playback device can collect the progress adjustment voices through the microphone in its voice interaction system.
[0067] Among them, the target audio file is associated with a target audio text. The association between the audio file and the audio text may mean that the text obtained by recognizing the audio voice of the audio file is the same as the audio text.
[0068] For example, the target audio text can be the audio text encapsulated in the target audio file. Alternatively, the target audio text can be a text file with the same file identifier as that of the target audio text. The file identifier can include at least one of the file name, the author's name, and the singer's name.
[0069] Step 102: Recognize the progress adjustment speech to obtain a first speech recognition text.
[0070] The audio playback device can use a speech recognition algorithm to recognize the collected playback speech, thereby obtaining a first speech recognition text.
[0071] Step 103: If the first speech recognition text includes content keywords belonging to the target audio text, start playing the target audio file from the time stamp corresponding to the content keywords.
[0072] After obtaining the first speech recognition text, the audio playback device can search in the target audio text based on the first speech recognition text to detect whether there are the same words and sentences in the target audio text and the first speech recognition text. If the audio playback device determines that there are the same words and sentences in the target audio text and the first speech recognition text, for example, the target audio text includes the first speech recognition text, or the target audio text includes part of the text in the first speech recognition text, it can be determined that the first speech recognition text includes content keywords belonging to the target audio text, and then the target audio file can be started playing from the time stamp corresponding to the content keywords.
[0073] That is to say, the audio playback device can adjust the playback progress of the target audio file to the position of the time stamp corresponding to the content keywords of the first speech recognition text.
[0074] In summary, the embodiment of the present application provides a method for adjusting the playback progress. During the process of playing the target audio file, the audio playback device can recognize the collected progress adjustment speech to obtain a first speech recognition text, and when it is determined that the first speech recognition text includes content keywords belonging to the target audio text associated with the target audio file, start playing the target audio file from the content keywords. It can be seen that the audio playback device can adjust the playback progress of the target audio file in response to the progress adjustment speech, thus effectively improving the flexibility of adjusting the playback progress.
[0075] Moreover, since it is easier to remember a content keyword of the target audio text than to remember its timestamp, when the user adjusts the playback progress of the audio file by a progress adjustment voice containing the content keyword, the accuracy of the playback progress adjustment can be ensured, and the user experience is better. Also, since there is no need for the user to touch the playback progress bar when adjusting the playback progress, the user operation is effectively simplified, and the user experience is improved.
[0076] Figure 3 This is another method for adjusting the playback progress provided by the embodiments of the present application. This method can be applied to an audio playback device, such as Figure 1 the audio playback device shown. As Figure 3 shown, this method may include:
[0077] Step 201, collect the playback voice.
[0078] If the user of the audio playback device needs the audio playback device to play the target audio file, the user can express the playback voice for instructing the audio playback device to play the target audio file. Correspondingly, the audio playback device can collect the playback voice through the microphone in its voice interaction system.
[0079] Wherein, the speech recognition text corresponding to the playback voice may include: any content keyword in the target audio text associated with the target audio file, and / or, the file identifier of the target audio file. The file identifier of the target audio file may include at least one of: the file name, the author's name, and the singer's name.
[0080] Optionally, the target audio file may be a song file, and the target audio text associated with the song file is the lyrics text. Correspondingly, the content keyword is a certain sentence of lyrics in the song. The file identifier of the song file may include at least one of: the song name, the name of the lyricist, the name of the composer, and the singer's name.
[0081] Step 202, recognize the playback voice to obtain a second speech recognition text.
[0082] The audio playback device may use a speech recognition algorithm to recognize the collected playback voice, so as to obtain a second speech recognition text.
[0083] Step 203, detect the number of audio files to be played determined based on the second speech recognition text.
[0084] After obtaining the second speech recognition text, the audio playback device can determine the audio file to be played that matches the second speech recognition text. Among them, the audio file to be played that matches the second speech recognition text can refer to: the audio file whose file identifier is included in the second speech recognition text, and / or, the audio file whose content keywords of the associated audio text are included in the second speech recognition text.
[0085] For example, if the second speech recognition text includes content keywords, the associated audio text of the audio file to be played also includes the content keywords. If the second speech recognition text includes the file identifier of the audio file, the file identifier of the audio file to be played is the same as the file identifier in the second speech recognition text.
[0086] Since there may be multiple different audio files including the same content keywords and / or the same file identifier, in order to ensure accurate playback of the target audio file, after obtaining the audio file to be played that matches the second speech recognition text, the audio playback device can detect the number of audio files determined based on the second speech recognition text. If the audio playback device determines multiple audio files to be played, it can execute step 204. If the audio playback device determines only one audio file to be played, it can determine this audio file as the target audio file and execute step 206.
[0087] Optionally, the audio playback device can directly search for the audio file that matches the second speech recognition text from its local database. Or, the audio playback device can send an audio search instruction to the audio server, and the audio search instruction carries the second speech recognition text. Correspondingly, the audio server can respond to the audio search instruction, search for the audio file that matches the second speech recognition text, and send the searched audio file to the audio playback device after the search is completed. The audio playback device can then obtain the audio file to be played that matches the second speech recognition text.
[0088] Optionally, the second speech recognition text can include a verb for indicating playing the target audio file, and any content keyword in the target audio file, and / or, the file identifier of the target audio file. And usually the content keyword and / or the file identifier are used as search terms for the search. Therefore, after obtaining the second speech recognition text, the audio playback device can perform semantic analysis on the second speech recognition text to obtain search terms (i.e., the content keywords and / or the file identifier of the target audio file). Then, the audio playback device can search for the audio file that matches the second speech recognition text based on the search terms. Optionally, the audio playback device can use a preset semantic analysis rule to perform semantic analysis on the second speech recognition text.
[0089] It can be understood that the audio playback device can also directly send the second speech recognition text to the audio server, so that the audio server performs semantic recognition on the second speech recognition text and conducts a search based on the search terms obtained after semantic recognition. Thus, the processing resources of the audio playback device can be saved.
[0090] Step 204: Display the file identifiers of multiple audio files to be played.
[0091] If the audio playback device determines multiple audio files to be played based on the second speech recognition text, it can display the file identifiers of the multiple audio files to be played on its display screen.
[0092] For example, taking the audio playback device as a vehicle and the display screen of the vehicle as the central control display screen, if the vehicle searches for multiple audio files to be played based on the content keyword "A" in the second speech recognition text, it can display the Figure 4 interface shown.
[0093] From Figure 4 it can be seen that the file identifiers of the multiple audio files are in sequence: xinxxx, xxxing, xxnihexxling. And, this interface can also include: a prompt 01 for reminding the user to select a target audio file. This prompt 01 can be the text: Please select the target audio file.
[0094] Step 205: In response to a selection operation on the file identifier of the target audio file among the multiple audio files, determine the target audio file.
[0095] The user can select the file identifier of the target audio file from the multiple audio files. Correspondingly, the audio playback device can, in response to the user's selection operation on the identifier of the target audio file, determine the target audio file and execute Step 206.
[0096] Optionally, the user can select the file identifier of the target audio file by touching the display screen of the audio playback device. Or, the user can select the file identifier of the target audio file by voice.
[0097] For example, assuming the file identifier of the target audio file is "xinxxx", please continue to refer to Figure 4 , the user touches the file identifier "xinxxx". Correspondingly, the audio playback device can, in response to the user's touch operation on the file identifier "xinxxx", determine the audio file indicated by the file identifier "xinxxx" as the target audio file.
[0098] Step 206: Play the target audio file.
[0099] If the audio playback device determines an audio file to be played based on the second speech recognition text, the audio file can be determined as the target audio file, and the installed audio playback application can be automatically run to play the target audio file.
[0100] Alternatively, after the audio playback device determines the target audio file in response to a selection operation on the file identifier of the target audio file among multiple audio files, it can run the installed audio playback application to play the target audio file.
[0101] For example, please continue to refer to Figure 4 , when the audio playback device plays the target audio file, it can display a playback interface. From Figure 4 It can be seen that the playback interface may include: a progress bar 02 of the target audio file, and a pause control 03.
[0102] The audio playback device can adjust the playback progress of the target audio file in response to a touch operation by the user on the progress bar 02. In addition, the audio playback device can also stop playing the target audio file in response to a touch operation by the user on the pause control 03.
[0103] Step 207: During the process of playing the target audio file, collect progress adjustment speech for the target audio file.
[0104] During the process of the audio playback device playing the target audio file, if the user needs the audio playback device to adjust the playback progress of the target audio file, the user can express the progress adjustment speech for the target audio file. Correspondingly, the audio playback device can collect the progress adjustment speech through the microphone.
[0105] For example, the target audio text is: ACABCDECB. Assuming the user wants to listen to "C", the expressed progress adjustment speech can be: Jump to "C", Fast forward to "C", Rewind to "C", Play "C" or I want to listen to "C".
[0106] Step 208: Recognize the progress adjustment speech to obtain a first speech recognition text.
[0107] The audio playback device can use a speech recognition algorithm to recognize the progress adjustment speech, thereby obtaining a first speech recognition text.
[0108] Optionally, the progress adjustment speech may include: a verb for indicating the adjustment of the playback progress, and words and sentences other than the verb. Based on this, the audio playback device can also perform semantic analysis on the first speech recognition text to obtain the words and sentences other than the verb in the first speech recognition text.
[0109] For example, assume that the first speech recognition text is: Jump to A. Then, the words and phrases other than the verb used to indicate adjusting the playback progress obtained by the audio playback device through semantic analysis of the first speech recognition text are "A".
[0110] Step 209: If the first speech recognition text includes content keywords belonging to the target audio text, start playing the target audio file from the timestamp corresponding to the content keywords.
[0111] After obtaining the first speech recognition text, the audio playback device can detect whether the first speech recognition text includes content keywords belonging to the target audio text. If the audio playback device determines that the first speech recognition text includes content keywords belonging to the target audio text, it can start playing the target audio file from the timestamp corresponding to the content keywords. If the audio playback device determines that the first speech recognition text does not include content keywords belonging to the target audio text, it can issue a prompt message. The prompt message can be used to remind the user to re-express the progress adjustment speech.
[0112] In the embodiments of the present application, after obtaining the first speech recognition text, the audio playback device can search in the target audio text based on the first speech recognition text to detect whether there are the same words and phrases in the target audio text and the first speech recognition text. For example, the audio playback device can use the first speech recognition text as the search term to search in the target audio text. Alternatively, the audio playback device can use the words and phrases other than the verb used to indicate adjusting the playback progress in the first speech recognition text as the search term to search in the target audio text.
[0113] After completing the search, if the audio playback device determines that there are the same words and phrases in the target audio text and the first speech recognition text, for example, the target audio text includes the first speech recognition text, or the target audio text includes part of the text in the first speech recognition text, it can be determined that the first speech recognition text includes content keywords belonging to the target audio text, that is, both the first speech recognition text and the target audio text include the content keywords. If the audio playback device determines that there are no same words and phrases in the target audio text and the first speech recognition text, it can be determined that the first speech recognition text does not include content keywords belonging to the target audio text.
[0114] In the embodiments of the present application, before searching in the target audio text based on the first speech recognition text, the audio playback device can obtain the target audio text associated with the currently played target audio file, and the target audio text includes timestamps. Taking the target audio file as a song file and the target audio text as the lyrics text as an example, an exemplary description of the process for the audio playback device to obtain the target audio text is as follows:
[0115] The audio playback device can first read the ID3 tag of the target audio file to detect whether the target audio text is encapsulated in the ID3 tag. If the audio playback device reads the target audio text from the ID3 tag, it can obtain the target audio text associated with the target audio file. If the audio playback device does not read the target audio text from the ID3 tag, it can detect whether there is a text file in the memory of the audio playback device with the same identifier as the target audio file (such as a text file with the same name as the target audio file), and the file format of the text file can be.irc.
[0116] If the audio playback device determines that the text file exists in the memory, it can directly obtain the target audio text. If the audio playback device determines that the text file does not exist in the memory, it can send a request to obtain the target audio text to the audio server, and the request to obtain can carry the identifier of the target audio file. Correspondingly, the audio server responds to the request to obtain, and searches for a text file with the same identifier as the target audio file. If the audio server searches for a text file with the same identifier as the target audio file, it can directly send the text file to the audio playback device, and the audio playback device can obtain the target audio text. If the audio server does not search for a text file with the same identifier as the target audio file, it can send a prompt message indicating that no search result is found to the audio playback device. After that, the audio playback device can perform speech recognition on the target audio file to obtain the target audio text.
[0117] Among them, there are multiple tag frames in the ID3 tag of the audio file. The target audio text associated with the target audio file can be encapsulated in the synchronized lyric / text (SYLT) tag frame, the unsychronized lyric / text transcription (USLT) tag frame, the original lyricist(s) / text writer(s) (TOLY) tag frame, and the lyricist / text writer (TEXT) tag frame. Correspondingly, the audio playback device can read the target audio text from the SYLT tag frame, the USLT tag frame, the TOLY tag frame, or the TEXT tag frame.
[0118] In the embodiment of the present application, if the audio playback device determines that the first speech recognition text includes a content keyword belonging to the target audio text, and there is one such content keyword in the target audio text, the audio playback device can directly start playing the target audio file from the time stamp corresponding to the content keyword.
[0119] For example, assume the target audio text is: ACABCDECB. The first speech recognition text is "Jump to E". Since there is only one E in the target audio text, the audio playback device can directly start playing the target audio file from the timestamp corresponding to E.
[0120] If the audio playback device determines that the first speech recognition text includes content keywords belonging to the target audio text, and the target audio text includes multiple such content keywords, then refer to Figure 5 , the audio playback device can play the target audio file through the following process.
[0121] Step 2091: Detect whether the first speech recognition text includes direction keywords.
[0122] When the audio playback device determines that there are multiple content keywords in the target audio text, it can detect whether the first speech recognition text includes direction keywords for indicating the progress adjustment direction. The progress adjustment direction is fast forward or rewind. If the audio playback device determines that the first speech recognition text includes direction keywords, the audio playback device can execute Step 2092. If the audio playback device determines that the first speech recognition text does not include direction keywords, it can execute Step 2093.
[0123] In the embodiments of the present application, after the audio playback device performs semantic analysis on the first speech recognition text to obtain a verb for indicating the adjustment of the playback progress, it can detect whether the verb matches a certain phrase in the direction word set. If the audio playback device determines that the verb matches a certain phrase in the direction word set, it can determine that the first speech recognition text includes direction keywords. If the audio playback device determines that the verb does not match any phrase in the direction word set, it can determine that the first speech recognition text does not include direction keywords.
[0124] For example, assume the direction word set includes: fast forward and rewind. The verb obtained by the audio playback device after performing semantic analysis on the first speech recognition text is "fast forward". Since the direction word set includes a phrase identical to this verb, the audio playback device can determine that the first speech recognition text includes direction keywords.
[0125] Assume the verb obtained by the audio playback device after performing semantic analysis on the first speech recognition text is "play". Since this verb is not the same as any phrase in the direction word set, the audio playback device can determine that the first speech recognition text does not include direction keywords.
[0126] Step 2092: Detect the number of content keywords included in the second text segment of the target audio text.
[0127] After determining that the first speech recognition text includes a direction keyword for indicating a progress adjustment direction, the audio playback device may further detect the number of content keywords included in a second text segment of the target audio text. If the audio playback device determines that the second text segment includes at least two content keywords, it may perform step 2093. If the audio playback device determines that the second text segment includes one content keyword, it may use the content keyword as the target content keyword and perform step 2094.
[0128] Wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment of the target audio text after the reference keyword currently being played. If the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is the text segment of the target audio text before the reference keyword currently being played. The reference keyword currently being played may refer to: a keyword belonging to the target audio text obtained by recognizing the sound currently being played in the target audio file.
[0129] Exemplarily, the target audio text is: ACABCDECB, and the reference keyword currently being played is D. Assuming the first speech recognition text is "rewind to C", the second text segment is: ACABC. And this second text segment includes two content keywords C, so the audio playback device may perform step 2093.
[0130] Assuming the first speech recognition text is "fast forward to C", the second text segment is: ECB. And this second text segment only includes one content keyword, so the audio playback device may perform step 2094.
[0131] Step 2093: Detect whether the historical number of times of playing the target audio file starting from the timestamp corresponding to any one of the multiple content keywords is greater than the number threshold.
[0132] During each playback of the target audio file, the audio playback device may record the position and timestamp of the progress adjustment. If the audio playback device determines that the first speech recognition text does not include a direction keyword, or if it determines that the first speech recognition text includes a direction keyword and the second text segment includes multiple content keywords, for each content keyword, the audio playback device may detect whether the historical number of times of starting to play the target audio file from the timestamp corresponding to the content keyword is greater than the number threshold. Wherein, the number threshold may be pre-stored in the audio playback device. For example, the number threshold may be 10.
[0133] If the historical number of times the audio playback device determines to start playing the target audio file from the time stamp corresponding to the target content keyword among multiple content keywords is greater than the number threshold, step 2094 can be executed. If the audio playback device determines that the historical number of times of starting to play the target audio file from the time stamp corresponding to any content keyword is less than or equal to the number threshold, step 2095 can be executed.
[0134] It should be noted that for the scenario where the audio playback device determines that the first speech recognition text does not include a direction keyword, the multiple content keywords are the content keywords in the entire target audio text. For the scenario where the audio playback device determines that the first speech recognition text includes a direction keyword and the second text segment includes multiple content keywords, the multiple content keywords are the content keywords in the second text segment.
[0135] For example, if the target audio text includes 3 content keywords C, that is, C appears 3 times in the target audio text. The time stamp corresponding to the first occurrence of C is: 00:05, the time stamp corresponding to the second occurrence of C is 00:40, and the time stamp corresponding to the third occurrence of C is 01:10.
[0136] Assume that the number threshold is 10. If the audio playback device determines that the historical number of times of starting to play the target audio file from the time stamp 00:05 is 4, the historical number of times of starting to play the target audio file from the time stamp 00:40 is 12, and the historical number of times of starting to play the target audio file from the time stamp 01:10 is 5. Since the historical number of times of starting to play the target audio file from the time stamp 00:40, which is 12, is greater than the number threshold 10, the audio playback device can determine the C corresponding to the time stamp 00:40 as the target content keyword and execute step 2094.
[0137] If the audio playback device determines that the historical number of times of starting to play the target audio file from the time stamp 00:05 is 4, the historical number of times of starting to play the target audio file from the time stamp 00:40 is 6, and the historical number of times of starting to play the target audio file from the time stamp 01:10 is 5. Since the historical number of times of starting to play the target audio file from any time stamp is less than the number threshold 10, the audio playback device can execute step 2095.
[0138] Step 2094: Start playing the target audio file from the time stamp corresponding to the target content keyword.
[0139] If the audio playback device determines that the second text segment only includes one keyword, the content keyword can be used as the target content keyword, and the target audio file can be automatically started to be played from the time stamp corresponding to the target content keyword.
[0140] Alternatively, if the number of historical times when the audio playback device starts playing the target audio file from the timestamp corresponding to the target content keyword among multiple content keywords is greater than the threshold number of times, it can automatically start playing the target audio file from the timestamp corresponding to the target content keyword.
[0141] The number of historical times of starting to play the target audio file from the timestamp corresponding to the target content keyword being greater than the threshold number of times indicates that the user often adjusts the playback progress of the target audio file to the timestamp corresponding to the target content keyword. Therefore, when the audio playback device automatically starts playing the target audio file from this timestamp, it can meet the user's listening preferences, thereby effectively improving the user experience.
[0142] Step 2095: Display the audio text.
[0143] If the audio playback device determines that the number of historical times of starting to play the target audio file from the timestamp corresponding to any content keyword among multiple content keywords is not greater than the threshold number of times, it can display the audio text.
[0144] It should be noted that if the audio playback device determines that the first speech recognition text does not include a direction keyword, and the number of historical times of starting to play the target audio file from the timestamp corresponding to any content keyword among multiple content keywords is less than the threshold number of times, the display effects of multiple content keywords in the entire target audio text displayed by the audio playback device are different from those of other words and sentences except these multiple content keywords.
[0145] If the audio playback device determines that the first speech recognition text includes a direction keyword, and the second text segment includes multiple content keywords, and the number of historical times of starting to play the target audio file from the timestamp corresponding to any content keyword among the multiple content keywords included in the second text segment is less than the threshold number of times, the display effects of the multiple content keywords in the second text segment displayed by the audio playback device are different from those of other words and sentences except these multiple content keywords. '
[0146] Among them, the display effects of multiple content keywords being different from those of other words and sentences except multiple content keywords may mean that: the font, and / or, font size, and / or, stroke thickness, and / or, font color of the content keywords are different. For example, the font, font size, stroke thickness, and font color of the content keywords are all different from those of the other words and sentences.
[0147] Since the display effects of the multiple content keywords are different from those of the other words and sentences, it is convenient for the user to quickly locate the multiple content keywords, and then quickly locate the target content keyword. On the one hand, it improves the efficiency of the audio playback device in playing the target audio file. On the other hand, it is not necessary for the user to search for the multiple content keywords one by one from the target audio text, effectively improving the user experience.
[0148] Optionally, the audio playback device may display the complete target audio text. Alternatively, the audio playback device may display multiple first text segments included in the target audio text, each first text segment including a content keyword. Among them, the multiple first text segments are partial text segments of the target audio text. That is, the sum of the storage spaces occupied by the multiple first text segments may be smaller than the storage space occupied by the target audio text. In this way, the display resources of the audio playback device can be avoided from being wasted.
[0149] Taking the target audio file as a song file as an example, each first text segment may be a lyric segment including a content keyword. The lyric segment may refer to: the lyric including the content keyword and a segment of at least one adjacent lyric. That is to say, the audio playback device only displays specific several lyrics.
[0150] Exemplarily, assume that the target audio text is: ACABCDECB. The first speech recognition text is "Jump to C". It can be seen that the first speech recognition text does not include a direction keyword, and the first speech recognition text includes the content keyword C belonging to the target audio text.
[0151] Assume that the historical number of times of playing the target audio file starting from the timestamp corresponding to each C is less than the number threshold, then as Figure 6 shown, the vehicle can display three text segments on its central control display screen. The three text segments may be: "AC", "CD", and "EC" respectively. And, it can also be seen from Figure 6 that the font size of "C" is larger than that of other words and sentences.
[0152] Step 2096: In response to a selection operation on a target content keyword among multiple content keywords, determine the target content keyword.
[0153] The user can select a target content keyword from multiple content keywords displayed on the display screen of the audio playback device. Correspondingly, the audio playback device can, in response to the user's selection operation on the target content keyword, determine the target content keyword, and execute step 2094, that is, start playing the target audio file from the timestamp corresponding to the target content keyword.
[0154] Optionally, the user can select the target content keyword by touching the display screen of the audio playback device. Alternatively, the user can select the target content keyword by voice.
[0155] Exemplarily, please continue to refer to Figure 6, the user selects the segment "AC", and the timestamp corresponding to "A" in this segment is 00:40. Correspondingly, the vehicle can determine "A" in this segment as the target content keyword and start playing the target audio file from the timestamp 00:40 corresponding to "A".
[0156] It can be understood that in the above process, the audio playback device can omit steps 2091 to 2093. That is, if the audio playback device determines that the target audio file includes multiple content keywords, it can directly display the audio text, and then in response to the selection operation for the target content keyword among the multiple content keywords, start playing the target audio file from the timestamp corresponding to this target content keyword.
[0157] Alternatively, the audio playback device can omit step 2093, that is, if the audio playback device determines that the target audio file includes multiple content keywords and determines that the second text segment includes at least two keywords, it can display the audio text.
[0158] If the audio playback device determines that the target audio file includes multiple content keywords and determines that the second text segment includes one target keyword audio playback device, it can directly start playing the target audio file from the timestamp corresponding to this target content keyword.
[0159] Or, the audio playback device can omit steps 2091 to 2092, that is, if the audio playback device determines that the target audio file includes multiple content keywords and the historical number of times of starting to play the target audio file from the timestamp corresponding to the target content keyword among these multiple content keywords is greater than the number threshold, it can directly start playing the target audio file from the timestamp corresponding to this target content keyword.
[0160] It should be noted that the order of the steps of the method for adjusting the playback progress provided in the embodiments of the present application can be appropriately adjusted, and the steps can also be increased or decreased accordingly according to the situation. For example, steps 201 to 206 can be deleted according to the situation. Or, steps 203 to 205 can be deleted according to the situation. For example, after the audio playback device obtains the audio file to be played based on the second speech recognition text, it can directly play the first obtained audio file. Any person skilled in the art within the technical scope disclosed in the present application can easily think of the changed methods, which should all be covered within the protection scope of the present application, so details are not described herein again.
[0161] In summary, the embodiments of the present application provide a method for adjusting the playback progress. During the playback of a target audio file, an audio playback device can recognize the collected progress adjustment speech to obtain a first speech recognition text, and when it is determined that the first speech recognition text includes a content keyword in the target audio text associated with the target audio file, start playing the target audio file from the content keyword. Thus, it can be seen that the audio playback device can adjust the playback progress of the target audio file in response to the progress adjustment speech, thereby effectively improving the flexibility of adjusting the playback progress.
[0162] Moreover, since it is easier to remember the content keyword than to remember the timestamp of a certain content keyword in the target audio text, when the user adjusts the playback progress of the audio file through the progress adjustment speech containing the content keyword, the accuracy of adjusting the playback progress can be ensured, and the user experience is better. Also, since there is no need for the user to touch the playback progress bar when adjusting the playback progress, the user operation is effectively simplified, and the user experience is improved.
[0163] The embodiments of the present application also provide an audio playback device, which can be used to execute the method for adjusting the playback progress provided in the above method embodiments. Refer to Figure 7 , the audio playback device 110 includes: a processor 1101, and the processor 1101 is used for:
[0164] During the playback of a target audio file, collect the progress adjustment speech for the target audio file, and the target audio file is associated with the target audio text;
[0165] Recognize the progress adjustment speech to obtain a first speech recognition text;
[0166] If the first speech recognition text includes a content keyword belonging to the target audio text, start playing the target audio file from the timestamp corresponding to the content keyword.
[0167] Optionally, as Figure 7 shown, the audio playback device 110 further includes a microphone 1102 and a speaker 1103. The processor 1101 can be used for:
[0168] Control the microphone to collect the progress adjustment speech for the target audio file;
[0169] Control the speaker to start playing the target audio file from the timestamp corresponding to the content keyword.
[0170] Optionally, refer to Figure 7 , the audio playback device 110 further includes a display screen 1104. The processor 1101 can be used for:
[0171] If the target audio text includes multiple content keywords, the target audio text is displayed on the display screen, and the display effects of the multiple content keywords in the target audio text are different from those of other words and sentences except the multiple content keywords;
[0172] In response to a selection operation on a target content keyword among the multiple content keywords, the target audio file is played starting from the timestamp corresponding to the target content keyword.
[0173] Optionally, the processor 1101 can be used for:
[0174] Display multiple first text segments included in the target audio text on the display screen, and each first text segment includes one content keyword;
[0175] Among them, the multiple first text segments are partial text segments of the target audio text.
[0176] Optionally, the processor 1101 can be used for:
[0177] If the first speech recognition text further includes a direction keyword for indicating a progress adjustment direction, and at least two content keywords among the multiple content keywords are located in the second text segment of the target audio text, then the target audio text is displayed;
[0178] Among them, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment of the target audio text after the current playing reference keyword; if the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is the partial text segment of the target audio text before the reference keyword;
[0179] The display effects of the content keywords in the second text segment are different from those of other words and sentences.
[0180] Optionally, the processor 1101 can be used for:
[0181] If the target audio text includes multiple content keywords, the first speech recognition text further includes a direction keyword for indicating a progress adjustment direction, and one target content keyword among the multiple content keywords is located in the second text segment of the target audio text, then the target audio text is played starting from the timestamp corresponding to the target content keyword;
[0182] Among them, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment of the target audio text after the current playing reference keyword; if the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is the partial text segment of the target audio text before the reference keyword.
[0183] Optionally, the processor 1101 can be used for:
[0184] If the target audio text includes multiple content keywords, and the historical number of times of playing the target audio file starting from the timestamp corresponding to the target content keyword among the multiple content keywords is greater than the number threshold, then play the target audio file starting from the timestamp corresponding to the target content keyword.
[0185] Optionally, the processor 1101 can also be used for:
[0186] Recognize the collected playback voice to obtain a second voice recognition text;
[0187] If multiple audio files to be played are determined based on the second voice recognition text, then display the identifiers of the multiple audio files to be played;
[0188] In response to a selection operation on the identifier of the target audio file among the multiple audio files, play the target audio file.
[0189] In summary, the embodiment of the present application provides an audio playback device. During the process of playing a target audio file, the audio playback device can recognize the collected progress adjustment voice to obtain a first voice recognition text, and when it is determined that the first voice recognition text includes a content keyword in the target audio text associated with the target audio file, play the target audio file starting from the content keyword. Thus, it can be seen that the audio playback device can respond to the progress adjustment voice to adjust the playback progress of the target audio file, thereby effectively improving the flexibility of adjusting the playback progress.
[0190] Moreover, since it is easier to remember the content keyword than to remember the timestamp of a certain content keyword in the target audio text, the user can ensure the accuracy of adjusting the playback progress by using the progress adjustment voice containing the content keyword, and the user experience is better. Also, since there is no need for the user to touch the playback progress bar when adjusting the playback progress, the user operation is effectively simplified and the user experience is improved.
[0191] The embodiment of the present application provides an audio playback device, which may include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for adjusting the playback progress provided in the above embodiment, such as Figure 2 or Figure 3 the method shown.
[0192] The embodiment of the present application provides a vehicle, which may include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for adjusting the playback progress provided in the above embodiment, such as Figure 2or Figure 3 the method shown above.
[0193] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to perform the method for adjusting the playback progress provided in the above embodiment. For example Figure 2 or Figure 3 the method shown above.
[0194] An embodiment of the present application further provides a computer program product containing instructions. When the computer program product runs on a processor, the processor is caused to execute the method for adjusting the playback progress provided in the above method embodiment. For example Figure 2 or Figure 3 the method shown above.
[0195] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The described program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.
[0196] It should be understood that the “and / or” mentioned in this article indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character “ / ” generally represents an “or” relationship between the associated objects before and after. Also, the meaning of the term “at least one” in the present application is one or more, and the meaning of the term “multiple” in the present application is two or more.
[0197] The terms “first”, “second” and other such words in the present application are used to distinguish the same items or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependence between “first”, “second”, “nth”, and there is no limitation on the quantity and execution order. For example, without departing from the scope of various examples, the first text fragment can be called the second text fragment, and similarly, the second text fragment can be called the first text fragment.
[0198] The above are only exemplary embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A vehicle, characterized in that, The vehicle includes: a processor and a display screen; the processor is configured to: During the playback of a target audio file, collect progress adjustment speech for the target audio file, where the target audio file is associated with a target audio text, and the target audio text includes multiple content keywords; Recognize the progress adjustment speech to obtain a first speech recognition text; If the first speech recognition text includes a content keyword belonging to the target audio text and a direction keyword for indicating a progress adjustment direction, then detect the number of content keywords included in a second text segment of the target audio text; wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment of the target audio text after a reference keyword currently being played; if the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is the text segment of the target audio text before the reference keyword; if the second text segment includes at least two content keywords, then detect whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword in the at least two content keywords is greater than a number threshold. If the historical number of times of playing the target audio file starting from the timestamp corresponding to the first content keyword in the at least two content keywords is greater than the number threshold, then start playing the target audio file from the timestamp corresponding to the first content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword in the at least two content keywords is less than or equal to the number threshold, then display the target audio text on the display screen, and the display effect of the multiple content keywords in the target audio text is different from the display effect of other words and sentences except the multiple content keywords; if the second text segment includes one content keyword, then start playing the target audio file from the timestamp corresponding to the content keyword; If the first speech recognition text includes a content keyword belonging to the target audio text and does not include the direction keyword, then detect whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword in the multiple content keywords is greater than the number threshold; if the historical number of times of playing the target audio file starting from the timestamp corresponding to the second content keyword in the multiple content keywords is greater than the number threshold, then start playing the target audio file from the timestamp corresponding to the second content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword in the multiple content keywords is less than or equal to the number threshold, then display the target audio text; In the case of displaying the target audio text, in response to a selection operation on a target content keyword among the multiple content keywords, start playing the target audio file from the timestamp corresponding to the target content keyword; If the first speech recognition text does not include the content keywords belonging to the target audio text, then a prompt message is output, and the prompt message is used to prompt the user to re-express the progress adjustment speech.
2. An audio playback device, characterized in that, The audio playback device includes: a processor and a display screen; the processor is configured to: During the playback of the target audio file, collect the progress adjustment speech for the target audio file, where the target audio file is associated with a target audio text, and the target audio text includes multiple content keywords; Recognize the progress adjustment speech to obtain a first speech recognition text; If the first speech recognition text includes the content keywords belonging to the target audio text and the direction keywords for indicating the progress adjustment direction, then detect the number of content keywords included in the second text segment of the target audio text; wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, then the second text segment is the text segment of the target audio text after the reference keyword currently being played; if the progress adjustment direction indicated by the direction keyword is rewind, then the second text segment is the text segment of the target audio text before the reference keyword; if the second text segment includes at least two content keywords, then detect whether the historical number of times of playing the target audio file starting from the time stamp corresponding to each content keyword in the at least two content keywords is greater than the number threshold. If the historical number of times of playing the target audio file starting from the time stamp corresponding to the first content keyword in the at least two content keywords is greater than the number threshold, then start playing the target audio file from the time stamp corresponding to the first content keyword; if the historical number of times of playing the target audio file starting from the time stamp corresponding to any content keyword in the at least two content keywords is less than or equal to the number threshold, then display the target audio text on the display screen, and the display effect of the multiple content keywords in the target audio text is different from the display effect of other words and sentences except the multiple content keywords; if the second text segment includes one content keyword, then start playing the target audio file from the time stamp corresponding to the content keyword; If the first speech recognition text includes the content keywords belonging to the target audio text and does not include the direction keyword, then detect whether the historical number of times of playing the target audio file starting from the time stamp corresponding to each content keyword in the multiple content keywords is greater than the number threshold; if the historical number of times of playing the target audio file starting from the time stamp corresponding to the second content keyword in the multiple content keywords is greater than the number threshold, then start playing the target audio file from the time stamp corresponding to the second content keyword; if the historical number of times of playing the target audio file starting from the time stamp corresponding to any content keyword in the multiple content keywords is less than or equal to the number threshold, then display the target audio text; When the target audio text is being displayed, in response to a selection operation on a target content keyword among multiple content keywords, play the target audio file starting from the timestamp corresponding to the target content keyword; If the first speech recognition text does not include a content keyword belonging to the target audio text, output a prompt message for prompting the user to re-express the progress adjustment speech.
3. The audio playback device according to claim 2, wherein, The processor is specifically configured to: Display, in the display screen, a plurality of first text segments included in the target audio text, each of the first text segments including one of the plurality of content keywords; Wherein, the plurality of first text segments are partial text segments of the target audio text.
4. The audio playback device according to claim 2 or 3, characterized in that The processor is further configured to: Recognize the collected playback speech to obtain a second speech recognition text; If a plurality of audio files to be played are determined based on the second speech recognition text, display the file identifiers of the plurality of audio files to be played; In response to a selection operation on the file identifier of the target audio file among the plurality of audio files, play the target audio file.
5. A method for adjusting the playing progress, characterized in that, Applied to a vehicle and a display screen; the method includes: During the process of playing a target audio file, collect progress adjustment speech for the target audio file, the target audio file being associated with a target audio text, the target audio text including a plurality of content keywords; Recognize the progress adjustment speech to obtain a first speech recognition text; If the first speech recognition text includes content keywords belonging to the target audio text and direction keywords for indicating the progress adjustment direction, then detect the number of content keywords included in the second text segment of the target audio text; wherein, if the progress adjustment direction indicated by the direction keywords is fast forward, then the second text segment is the text segment of the target audio text after the reference keyword currently being played; if the progress adjustment direction indicated by the direction keywords is rewind, then the second text segment is the text segment of the target audio text before the reference keyword; if the second text segment includes at least two content keywords, then detect whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword among the at least two content keywords is greater than the number threshold. If the historical number of times of playing the target audio file starting from the timestamp corresponding to the first content keyword among the at least two content keywords is greater than the number threshold, then start playing the target audio file from the timestamp corresponding to the first content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword among the at least two content keywords is less than or equal to the number threshold, then display the target audio text on the display screen, and the display effect of the multiple content keywords in the target audio text is different from the display effect of other words and sentences except the multiple content keywords; if the second text segment includes one content keyword, then start playing the target audio file from the timestamp corresponding to the content keyword; If the first speech recognition text includes content keywords belonging to the target audio text and does not include the direction keywords, then detect whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword among the multiple content keywords is greater than the number threshold; if the historical number of times of playing the target audio file starting from the timestamp corresponding to the second content keyword among the multiple content keywords is greater than the number threshold, then start playing the target audio file from the timestamp corresponding to the second content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword among the multiple content keywords is less than or equal to the number threshold, then display the target audio text; In the case of displaying the target audio text, in response to a selection operation on a target content keyword among the multiple content keywords, start playing the target audio file from the timestamp corresponding to the target content keyword; If the first speech recognition text does not include content keywords belonging to the target audio text, then output a prompt message for prompting the user to re - express the progress adjustment speech.
6. A method for adjusting the playing progress, characterized in that Applied to an audio playback device and a display screen; the method includes: During the process of playing a target audio file, progress adjustment speech for the target audio file is collected. The target audio file is associated with a target audio text, and the target audio text includes multiple content keywords; The progress adjustment speech is recognized to obtain a first speech recognition text; If the first speech recognition text includes a content keyword belonging to the target audio text and a direction keyword for indicating a progress adjustment direction, the number of content keywords included in a second text segment of the target audio text is detected. Wherein, if the progress adjustment direction indicated by the direction keyword is fast forward, the second text segment is the text segment of the target audio text after a reference keyword that is currently being played; if the progress adjustment direction indicated by the direction keyword is rewind, the second text segment is the text segment of the target audio text before the reference keyword. If the second text segment includes at least two content keywords, it is detected whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword among the at least two content keywords is greater than a number threshold. If the historical number of times of playing the target audio file starting from the timestamp corresponding to the first content keyword among the at least two content keywords is greater than the number threshold, the target audio file is played starting from the timestamp corresponding to the first content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword among the at least two content keywords is less than or equal to the number threshold, the target audio text is displayed on the display screen, and the display effect of the multiple content keywords in the target audio text is different from the display effect of other words and sentences except the multiple content keywords; if the second text segment includes one content keyword, the target audio file is played starting from the timestamp corresponding to the content keyword; If the first speech recognition text includes a content keyword belonging to the target audio text and does not include the direction keyword, it is detected whether the historical number of times of playing the target audio file starting from the timestamp corresponding to each content keyword among the multiple content keywords is greater than the number threshold; if the historical number of times of playing the target audio file starting from the timestamp corresponding to the second content keyword among the multiple content keywords is greater than the number threshold, the target audio file is played starting from the timestamp corresponding to the second content keyword; if the historical number of times of playing the target audio file starting from the timestamp corresponding to any content keyword among the multiple content keywords is less than or equal to the number threshold, the target audio text is displayed; In the case of displaying the target audio text, in response to a selection operation on a target content keyword among the multiple content keywords, the target audio file is played starting from the timestamp corresponding to the target content keyword; If the first speech recognition text does not include content keywords belonging to the target audio text, a prompt message is output, and the prompt message is used to prompt the user to re-express the progress adjustment speech.
Citation Information
Patent Citations
Method and device for pushing songs on basis of voice recognition
CN103685520A
Video playing method and device, terminal device and storage medium
CN109246472A
Play progress control method, smart wearable equipment and multi-media display equipment
CN109767771A