A video playback method, apparatus, device, and storage medium

By acquiring and displaying a set of text information at the moment of video playback, the problem of users having difficulty accurately locating video segments is solved, enabling more efficient and accurate video progress adjustment.

CN115396738BActive Publication Date: 2025-12-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110571850.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-25
Publication Date
2025-12-02
Estimated Expiration
2041-05-25

AI Technical Summary

Technical Problem

In existing technologies, it is difficult for users to accurately locate video segments, resulting in low efficiency in adjusting the video viewing progress, especially on small-screen devices or with long videos, requiring multiple swipes of the progress bar or viewing of unnecessary segments.

Method used

By obtaining a set of text information about the current playback moment of the video, displaying it on the video playback interface, and responding to user selections, the system can accurately locate the target playback moment for video playback.

Benefits of technology

It improves the efficiency and accuracy of adjusting video viewing progress, allowing users to jump to video segments of interest more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115396738B_ABST
    Figure CN115396738B_ABST
Patent Text Reader

Abstract

This application provides a video playback method, apparatus, device, and storage medium, relating to the field of computer technology. The method includes: first, obtaining a first playback moment based on a user's target adjustment operation triggered by the playback progress of a target video; then, acquiring and displaying a set of text information corresponding to the first playback moment, allowing the user to intuitively understand the video content near the first playback moment from the text information set. Next, in response to a selection operation triggered by target text information in each set of text information, obtaining a second playback moment corresponding to the target text information in the target video, and starting playback of the target video from the second playback moment. This precisely locates the second playback moment in the target video needed by the user and starts playback from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a video playback method, apparatus, device and storage medium. Background Technology

[0002] With the development of video technology, people's demands for viewing experience are also increasing. Currently, users can adjust the viewing progress of videos as needed, allowing them to watch only the segments they are interested in. For this purpose, video platforms provide a progress bar. Users can slide the progress bar to any point in the video to adjust the viewing progress.

[0003] However, the method of sliding the video progress bar is difficult to accurately locate the video segment that the user needs. Users often need to slide the video progress bar multiple times, or watch unnecessary video segments to naturally transition to the desired video segment, resulting in low efficiency in adjusting the video viewing progress. Summary of the Invention

[0004] This application provides a video playback method, apparatus, device, and storage medium to improve the efficiency and accuracy of adjusting video viewing progress.

[0005] On one hand, embodiments of this application provide a video playback method, the method comprising:

[0006] In response to a target adjustment operation triggered by the playback progress of the target video, the first playback moment of the adjusted target video is determined;

[0007] Obtain a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information;

[0008] Display at least one text message in the video playback interface;

[0009] In response to a selection operation triggered for target text information in the at least one text information, a second playback time corresponding to the target text information in the target video is obtained, and the target video is played starting from the second playback time.

[0010] On one hand, embodiments of this application provide a video playback method, the method comprising:

[0011] In response to a target adjustment operation triggered by the playback progress of a target video, a set of text information is displayed in the video playback interface, wherein the set of text information includes at least one piece of text information, and the target adjustment operation is used to determine a first playback moment in the playback progress of the target video, and the set of text information corresponds to the first playback moment.

[0012] In response to a selection operation triggered for target text information in the at least one text information, the target video is played starting from a second playback time, wherein the second playback time is the playback time associated with the target text information.

[0013] Optionally, the target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface;

[0014] The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

[0015] Optionally, displaying the text information set in the video playback interface includes:

[0016] In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

[0017] Optionally, it also includes:

[0018] The video playback interface displays the information of the person associated with each of the at least one text information in the target video.

[0019] On one hand, embodiments of this application provide a video playback device, the device comprising:

[0020] An adjustment response module is used to respond to a target adjustment operation triggered by the playback progress of the target video and determine the first playback moment of the adjusted target video.

[0021] The acquisition module is used to acquire a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information;

[0022] A display module is used to display at least one text message in a video playback interface;

[0023] A video playback module is configured to respond to a selection operation triggered by target text information in the at least one text information, obtain a second playback moment corresponding to the target text information in the target video, and start playing the target video from the second playback moment.

[0024] Optionally, the display module is specifically used for:

[0025] In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

[0026] Optionally, the display module is specifically used for:

[0027] The video playback interface displays the information of the person associated with each of the at least one text information in the target video.

[0028] Optionally, the acquisition module is specifically used for:

[0029] Obtain the set of first preview image and text information corresponding to the first playback moment in the target video;

[0030] The display module is specifically used for:

[0031] Based on the temporal relationship between the at least one text information and the first preview image, the first preview image and the at least one text information are displayed in the video playback interface.

[0032] Optionally, the display module is specifically used for:

[0033] According to the time sequence relationship between the at least one text information and the first preview image, the at least one text information is displayed in the first area of ​​the video playback interface;

[0034] The first preview image is displayed in the second area of ​​the video playback interface.

[0035] Optionally, the display module is further used for:

[0036] In the first area, the personal information associated with each of the at least one text information in the target video is displayed.

[0037] In the second area, information about the person associated with the first preview image in the target video is displayed.

[0038] Optionally, the set of text information includes first text information played at the first playback time, and at least two second text information played before and after the first playback time.

[0039] Optionally, the acquisition module is further configured to:

[0040] In response to a target adjustment operation triggered by the playback progress of a target video, after determining the first playback moment of the adjusted target video, at least one second preview image is obtained that satisfies a preset condition for the playback time interval between the target video and the first preview image, and has a similarity to the first preview image that is less than a preset threshold.

[0041] The display module is also used for:

[0042] In the second region, the at least one second preview image is displayed according to its temporal association with the first preview image.

[0043] Optionally, the display module is further used for:

[0044] In the second area, the person information associated with each of the at least one second preview image in the target video is displayed.

[0045] Optionally, the video playback module is further configured to:

[0046] According to the temporal association relationship between at least one text information in the text information set and the first preview image, after displaying the at least one text information in the video playback interface, in response to the selection operation triggered for the target preview image in the preview image set, the third playback time corresponding to the target preview image in the target video is obtained, and the target video is played from the third playback time, wherein the preview image set includes the first preview image and the at least one second preview image.

[0047] Optionally, the video playback module is further configured to:

[0048] According to the temporal relationship between at least one text information in the text information set and the first preview image, after displaying the at least one text information in the video playback interface, if no operation is responded to within a preset time period, the target video will start playing from the first playback time.

[0049] Optionally, the target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface. The maximum number of text information contained in the text information set is the number of text information corresponding to the video segment to which the target video jumps when the video progress bar is dragged a unit distance. The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

[0050] On one hand, embodiments of this application provide a video playback device, the device comprising:

[0051] An adjustment response module is used to respond to a target adjustment operation triggered by the playback progress of a target video, and to display a set of text information in the video playback interface, wherein the set of text information includes at least one piece of text information, and the target adjustment operation is used to determine a first playback moment in the playback progress of the target video, and the set of text information corresponds to the first playback moment.

[0052] A video playback module is configured to play the target video starting from a second playback time in response to a selection operation triggered for target text information in the at least one text information, wherein the second playback time is a playback time associated with the target text information.

[0053] Optionally, the target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface;

[0054] The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

[0055] Optionally, the adjustment response module is specifically used for:

[0056] In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

[0057] Optionally, the adjustment response module is specifically used for:

[0058] The video playback interface displays the information of the person associated with each of the at least one text information in the target video.

[0059] On one hand, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described video playback method.

[0060] On one hand, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the video playback method described above.

[0061] In this embodiment, the first playback moment is obtained based on the user's target adjustment operation triggered by the playback progress of the target video. Then, the text information set corresponding to the first playback moment is acquired and displayed. Thus, the user can learn about the video content near the first playback moment from the text information set. Afterward, by selecting target text information from the text information set, the second playback moment in the target video that the user needs is precisely located, and the target video is played from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 A schematic diagram of a system architecture provided for an embodiment of this application;

[0064] Figure 2 A schematic diagram of a video playback interface provided in an embodiment of this application;

[0065] Figure 3 A schematic diagram of a video playback interface provided in an embodiment of this application;

[0066] Figure 4a A schematic diagram of a video playback interface provided in an embodiment of this application;

[0067] Figure 4b A schematic diagram of a video playback interface provided in an embodiment of this application;

[0068] Figure 5 A flowchart illustrating a video playback method provided in an embodiment of this application;

[0069] Figure 6 A schematic diagram illustrating the adjustment of video progress as provided in an embodiment of this application;

[0070] Figure 7 A schematic diagram illustrating the adjustment of video progress as provided in an embodiment of this application;

[0071] Figure 8 A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0072] Figure 9 A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0073] Figure 10a A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0074] Figure 10b A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0075] Figure 10c A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0076] Figure 10d A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0077] Figure 11a A schematic diagram illustrating the display of dialogue information provided in an embodiment of this application;

[0078] Figure 11b A schematic diagram illustrating the adjustment of video progress as provided in an embodiment of this application;

[0079] Figure 12 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0080] Figure 13 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0081] Figure 14 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0082] Figure 15 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0083] Figure 16a A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0084] Figure 16b A schematic diagram of a video playback interface provided in an embodiment of this application;

[0085] Figure 17 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0086] Figure 18 A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0087] Figure 19 A flowchart illustrating a video playback method provided in an embodiment of this application;

[0088] Figure 20a A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0089] Figure 20b A schematic diagram illustrating the display of dialogue information and a preview image, provided as an embodiment of this application;

[0090] Figure 21 This is a schematic diagram of the structure of a video playback device provided in an embodiment of this application;

[0091] Figure 22 This is a schematic diagram of the structure of a video playback device provided in an embodiment of this application;

[0092] Figure 23 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0093] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0094] For ease of understanding, the terms used in the embodiments of this invention are explained below.

[0095] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0096] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0097] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0098] Automatic speech recognition (ASR) is a technology that converts human speech into text. It involves multiple disciplines such as acoustics, phonetics, and linguistics, and is now widely integrated into daily life. In this embodiment, text information from a target video can be obtained using ASR.

[0099] Voiceprint: It is a type of biometric feature that is extracted when a speaker makes a sound and can be used as a representation and identifier of the speaker.

[0100] Dialogue: refers to the text data corresponding to the audio information in a video.

[0101] The design concept of the embodiments of this application will be introduced below.

[0102] Currently, users can adjust the viewing progress of videos as needed, allowing them to watch only the segments they are interested in. For this adjustment, relevant video platforms provide a video progress bar. Users can slide the progress bar to any point in the video to adjust the viewing progress.

[0103] However, using a swipe video progress bar makes it difficult to accurately locate the video segment a user needs. For example, when the screen of the mobile device watching the video is small, or when the video is long, a small drag of the progress bar might correspond to several minutes of video content. In such cases, users often need to swipe the progress bar multiple times, or watch unnecessary video segments to naturally transition to the desired segment, resulting in low efficiency in adjusting the video viewing progress.

[0104] Analysis revealed that preview images of a video at the current playback moment allow users to intuitively understand the video content at that moment. However, some videos may maintain a single shot for an extended period, such as presentation videos where the speaker remains in focus for a long time. This results in minimal differences between preview images at different playback moments, making it difficult for users to locate the desired video segment based solely on the preview images.

[0105] Further analysis revealed that while some videos may maintain a single shot for extended periods, the accompanying dialogue differs at different playback points. For instance, speech videos may feature a prolonged shot of the speaker, but the content of the speech varies at different playback points. By relying on the dialogue within the video, users can more accurately understand the content at various moments and thus locate the desired video segment.

[0106] In view of this, embodiments of this application provide a video playback method, the method comprising: in response to a target adjustment operation triggered by the playback progress of a target video, determining a first playback moment of the adjusted target video; then obtaining a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information; displaying at least one piece of text information in a video playback interface; and in response to a selection operation triggered by a target text information among the at least one piece of text information, obtaining a second playback moment corresponding to the target text information in the target video, and starting playback of the target video from the second playback moment.

[0107] In this embodiment, the first playback moment is obtained based on the user's target adjustment operation triggered by the playback progress of the target video. Then, the text information set corresponding to the first playback moment is acquired and displayed. Thus, the user can learn about the video content near the first playback moment from the text information set. Afterward, by selecting target text information from the text information set, the second playback moment in the target video that the user needs is precisely located, and the target video is played from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress.

[0108] refer to Figure 1 This is a system architecture diagram applicable to the embodiments of this application. The system architecture includes at least a terminal device 101 and a server 102.

[0109] The terminal device 101 has video applications pre-installed, including video playback applications, short video applications, live streaming applications, etc. The types of video applications include client applications, web-based applications, and mini-program applications. The terminal device 101 may include one or more processors 1011, memory 1012, I / O interfaces 1013 for interacting with the server 102, and a display panel 1014, etc. The terminal device 101 may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, in-vehicle device, etc., but is not limited to these.

[0110] Server 102 is the backend server for the video application. Server 102 may include one or more processors 1021, memory 1022, and I / O interfaces 1023 for interacting with terminal device 101. Furthermore, server 102 may be configured with a database 1024. Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal device 101 and server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0111] In response to a target adjustment operation triggered by the playback progress of a target video in a video application, terminal device 101 determines the first playback moment of the adjusted target video. Then, it retrieves a set of text information corresponding to the first playback moment in the target video from server 102, wherein the set of text information includes at least one piece of text information. Terminal device 101 displays at least one piece of text information in the video playback interface of the video application. In response to a selection operation triggered by a target piece of text information in the video application, terminal device 101 obtains the second playback moment corresponding to the target piece of text information in the target video, and starts playing the target video from the second playback moment.

[0112] For example, the video playback interface of a video application is as follows: Figure 2 As shown, the target video currently playing in the video application is a football match video, and the current video progress bar 201 is at the playback time of "31 minutes and 21 seconds".

[0113] The user slides the video progress bar 201 to the playback time "54 minutes and 20 seconds". Terminal device 101 responds to the target adjustment operation triggered in the video application for the playback progress of the football match video, determines the adjusted playback time "54 minutes and 20 seconds", and retrieves the first preview image and corresponding text information set corresponding to the playback time "54 minutes and 20 seconds" from server 102. The playback time corresponding to the first preview image is 54 minutes and 20 seconds. The text information set includes {Team A has scored first, the score is now 1:0, and Team B may be under pressure after conceding a goal}. Specifically, the playback time corresponding to text information 1 "Team A has scored first" is "54 minutes and 18 seconds", text information 2 "The score is now 1:0" corresponds to "54 minutes and 20 seconds", and text information 3 "Team B may be under pressure after conceding a goal" corresponds to "54 minutes and 22 seconds".

[0114] Terminal device 101 displays at least one text message in the video playback interface of a video application, arranged according to the playback time associated with at least one text message, while simultaneously displaying a first preview image. At this time, the video playback interface of the video application appears as follows: Figure 3 As shown, the video progress bar 201 is at the playback time "54 minutes and 20 seconds". Text information 1, text information 2, and text information 3 are displayed in the first area 301 of the video playback interface. The font size of text information 2 is larger than that of text information 1 and text information 3, and text information 2 is located at the target selection position 303. The second area 302 of the video playback interface displays the first preview image and its playback time "54 minutes and 20 seconds".

[0115] The user swipes down in the first area 301, moving the text information 1 to the target selection position 303, such as... Figure 4a As shown. In response to a selection operation triggered by text information 1 in a video application, terminal device 101 obtains that the playback time corresponding to text information 1 in the football match video is "54 minutes and 18 seconds", and starts playing the football match video from 54 minutes and 18 seconds, as... Figure 4b As shown.

[0116] based on Figure 1 The system architecture diagram shown in this application illustrates a flowchart of a video playback method. Figure 5 As shown, the process of this method is executed by a computer device, which can be... Figure 1 The terminal device 101 shown includes the following steps:

[0117] Step S501: In response to the target adjustment operation triggered by the playback progress of the target video, determine the first playback moment of the adjusted target video.

[0118] Specifically, the target adjustment operation can be a dragging operation along the video progress bar in the video playback interface. The direction of dragging the video progress bar corresponds to the direction of video playback adjustment, and the drag distance corresponds to the length of the video segment to be jumped to. In practice, the duration of the video segment to be jumped to is determined based on the total video duration and the total drag range of the video progress bar.

[0119] For example, on a terminal device with 390 pixels × 844 pixels, the total length of the drag range of the video progress bar is 844 pixels, with each pixel corresponding to a video segment. If the total duration of the target video is 2 hours, then dragging the video progress bar by one pixel will jump to a video segment with a duration of approximately 8.5 seconds.

[0120] like Figure 6 As shown, the current playback time is set to 35 minutes and 18 seconds. Users can drag the video progress bar two pixels to the right to jump to the first playback time "35 minutes and 35 seconds" after 35 minutes and 18 seconds. Users can also drag the video progress bar two pixels to the left to jump to the first playback time "35 minutes and 01 seconds" before 35 minutes and 18 seconds.

[0121] The target adjustment operation can also be performed by sliding directly on the video playback interface. The sliding direction corresponds to the direction of video playback adjustment, and the sliding distance corresponds to the length of the video to be jumped to.

[0122] For example, such as Figure 7As shown, the current playback time is set to 35 minutes and 18 seconds. Users can swipe right on the video playback interface to jump to the first playback time after 35 minutes and 18 seconds, or swipe left on the video playback interface to jump to the first playback time before 35 minutes and 18 seconds.

[0123] The target adjustment operation can also be an operation of the fast forward or rewind buttons used to control the video playback progress. These fast forward or rewind buttons can be virtual buttons on the video playback interface or hardware buttons on the terminal device. It should be noted that the target adjustment operation can also be other settings operations, which will not be elaborated here.

[0124] Step S502: Obtain the set of text information corresponding to the first playback moment in the target video. The text information in the set is the video text in the target video. The video text can be dialogue, subtitles, bullet comments, narration, etc. A piece of text information can refer to a sentence, a word, or a paragraph, etc. A piece of text information can be associated with a video segment in the target video. For example, the text information can be the dialogue of a character playing in the video segment, or the narration text of the content of the video segment, etc. The video segment can include at least one video frame. All video text in the target video can be obtained through speech recognition, image recognition, or manual methods. The speech recognition method in this application embodiment includes, but is not limited to, using Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM), or acoustic models such as recurrent neural networks, encoder-decoder frameworks, and attention mechanisms based on deep learning. For example, videos such as TV series, variety shows, and movies have less audio noise. Based on deep learning speech recognition technology, combined with manual correction, a list of dialogue data for the video can be generated. The dialogue data format can be represented as {"time point": "dialogue information"}, and the dialogue information in the dialogue data list is sorted in order of time points.

[0125] In some embodiments, text information may be associated with a point in time or a time period in the target video, at which the video segment associated with the text information is played. The point in time associated with the text information may be the playback time of the text information, or the start or end of the time period associated with the text information, or any of these point in time may be the playback time of the text information.

[0126] For example, in a TV drama scene, if a character is saying a line in a specific video segment (35 minutes 01 seconds to 35 minutes 25 seconds), that line can be associated with that time segment (35 minutes 01 seconds to 35 minutes 25 seconds). When the user adjusts the playback to a point between 35 minutes 01 seconds and 35 minutes 25 seconds (i.e., the first playback moment), and this first playback moment falls within the time segment associated with the line, then the line associated with that time segment is identified as the text information corresponding to the first playback moment.

[0127] The text information set may include at least one text message, each of which may be related to a first playback moment. In some embodiments, each text message in the text information set may be a predetermined number of text messages that are sequentially adjacent to the text message corresponding to the first playback moment. For example, the text information set may include at least one of the text message corresponding to the first playback moment, M text messages preceding the text message corresponding to the first playback moment in the playback sequence, and M text messages following the text message corresponding to the first playback moment in the playback sequence, where M is greater than or equal to 1 and can be set according to the height of the display terminal. In this embodiment, the number of text messages in the text information set (i.e., the size of M) can be dynamically adjusted according to the video length. It is understood that the longer the video length, the lower the accuracy of adjusting the progress through the horizontal video progress bar. Therefore, the number of text messages in the text information set can be dynamically adjusted according to the video length. The longer the video length, the more text messages in the text information set, and the wider the range of precise adjustment.

[0128] In other embodiments, the difference between the time point or time period associated with each piece of text information included in the text information set and the first playback time is within a predetermined threshold. For example, the text information set may include at least one of the following: text information corresponding to the first playback time, text information associated with a time interval preceding the first playback time in the playback sequence, and text information associated with a time interval following the first playback time in the playback sequence. The time interval may include at least one time period, and each time period in the at least one time period is associated with text information. In this embodiment, the time interval can be dynamically adjusted according to the video length. It is understood that the longer the video length, the lower the precision of adjusting the progress using a horizontal video progress bar. Therefore, the number of text information pieces in the text information set can be dynamically adjusted according to the video length. The longer the video length, the longer the time interval, and the more text information pieces in the text information set, thereby increasing the precisely adjustable time interval.

[0129] Optionally, when the target adjustment operation is a drag operation of the video progress bar along the direction of the video progress bar in the video playback interface, the maximum number of text information contained in the text information set is the number of text information corresponding to the video segment to which the target video jumps when the video progress bar is dragged a unit distance.

[0130] Specifically, first, the target duration of the video segment to which the target video jumps when the video progress bar is dragged a unit distance is determined. Then, based on the target duration, the target video is divided into multiple video segments, and the number of text information in each video segment is determined. The average number of text information in each video segment is taken as the maximum number of text information contained in the text information set. Of course, the maximum number of text information contained in the text information set can also be the number of text information corresponding to the video segment to which the target video jumps when the video progress bar is dragged a preset distance; this application does not specifically limit this. Optionally, in this embodiment, when the average number of text information in each video segment exceeds a threshold, one piece of text information can be extracted every L text information according to the playback sequence of the target video as the text information in the text information set, wherein the threshold can be determined according to the height of the display interface.

[0131] For example, for a target video with a total duration of 2 hours, on a terminal device with a resolution of 390 pixels × 844 pixels, the total drag range of the video progress bar is 844 pixels, and the duration of the video segment corresponding to each pixel is approximately 8.5 seconds. The average number of lines in the 8.5-second video segment of the target video is calculated, and then this average number of lines is used as the maximum number of lines contained in the dialogue information set.

[0132] Step S503: Display at least one text message in the video playback interface.

[0133] Specifically, text information preceding and following the first playback moment can be displayed separately. Alternatively, at least one text message can be arranged and displayed according to its associated playback moment within the video playback interface. Furthermore, only a portion of the text information can be displayed, with the rest hidden. Users can drag the dialogue progress bar perpendicular to the video progress bar or simply slide it in that direction to reveal the hidden text. The display style of each text message can be customized based on various factors, including font, color, font size, display angle, position, and display method.

[0134] For example, the text information set is set as the dialogue information set, including five dialogue information, namely: {First time point "35 minutes 18 seconds": First dialogue information "Team A has scored first"}, {Second time point "35 minutes 20 seconds": Second dialogue information "The score on the field is now 1:0"}, {Third time point "35 minutes 22 seconds": Third dialogue information "Team B may be under pressure after conceding a goal"}, {Fourth time point "35 minutes 24 seconds": Fourth dialogue information "Team B needs to cheer up"}, {Fifth time point "35 minutes 26 seconds": Fifth dialogue information "Team B starts to attack"}.

[0135] One possible implementation, such as Figure 8 As shown, the video playback interface displays the first, second, third, fourth, and fifth lines of dialogue from left to right.

[0136] One possible implementation, such as Figure 9 As shown, on the left side of the video playback interface, the first line information, the second line information, the third line information, the fourth line information, and the fifth line information are displayed from top to bottom.

[0137] One possible implementation, such as Figure 10a As shown, on the left side of the video playback interface, the second, third, and fourth lines of dialogue are displayed from top to bottom, while the first and fifth lines of dialogue are hidden.

[0138] In response to the user's downward swipe, the first, second, and third lines of dialogue are displayed on the video playback screen, while the fourth and fifth lines are hidden. (See details below.) Figure 10b As shown.

[0139] In response to the user dragging the vertical dialogue progress bar 1001 upwards, the third, fourth, and fifth dialogue information are displayed on the video playback interface, while the first and second dialogue information are hidden. Specifically, as follows... Figure 10c As shown.

[0140] Optionally, the video playback interface may display information about at least one person associated with each text message in the target video.

[0141] Specifically, a trained voiceprint model can be used to extract voiceprint features from a video. Then, based on these features and a pre-defined list of personal information, the personal information corresponding to each piece of text in the target video can be determined. Alternatively, the personal information corresponding to each piece of text in the target video can be determined manually. Personal information can include character names, actor names, genders, ages, etc. After obtaining the personal information, it can be associated with the text information and included as part of the text information set. When displaying each piece of text information, the associated personal information is also displayed simultaneously.

[0142] For example, such as Figure 10d As shown, on the left side of the video playback interface, the second, third, and fourth lines of dialogue information, along with the corresponding character information, are displayed from top to bottom. The character information for the second, third, and fourth lines of dialogue information is "Narrator A".

[0143] It should be noted that, in the embodiments of this application, the forms of various text information in the video playback interface are not limited to the few examples mentioned above, and may also be other forms, which are not specifically limited in this application.

[0144] Step S504: In response to the selection operation triggered for the target text information in at least the text information, obtain the second playback time corresponding to the target text information in the target video, and start playing the target video from the second playback time.

[0145] Specifically, the selection operation can be a swipe perpendicular to the video progress bar, or dragging the dialogue progress bar perpendicular to the video progress bar, or it can be a click, double-click, long-press, etc. In practice, users can use one finger to adjust the target, and after the adjustment is complete, keep that finger still and use another finger to make the selection. Alternatively, users can use one finger to adjust the target, and after the adjustment is complete, the finger is released, the video application locks the first playback moment after the adjustment, and the user can then use that finger to make the selection. The second playback moment can be the first playback moment, or a playback moment before or after the first playback moment. Optionally, if no operation is responded to within a preset time, the target video will start playing from the first playback moment.

[0146] As mentioned above, in some embodiments, the target text information can be associated with a time point, so the second playback time can be the time point associated with the target text information. In other embodiments, the target text information can be associated with a time period, so the second playback time can be the start, end, or any time point within the time period associated with the target text information.

[0147] For example, the selection operation is set as a sliding operation perpendicular to the video progress bar, and the dialogue information set is as follows: Figure 10a As shown, it will not be elaborated further here.

[0148] After the user drags the video progress bar to the right, as Figure 11a As shown, on the left side of the video playback interface, the second, third, and fourth lines of dialogue are displayed from top to bottom. The font size of the third line, "After conceding a goal, Team B may be under pressure," is larger than the font size of the other lines, and the third line is located at the target selection position 1101.

[0149] If the user swipes down on the video playback screen and slides the second line of dialogue, "The current score is 1:0," to the target selection position 1101, then the video playback screen will display the first, second, and third lines of dialogue. The font size of the second line of dialogue will be larger than the other lines. Figure 11b As shown. The playback time corresponding to the second line of dialogue is "35 minutes and 20 seconds", so the target video will start playing from 35 minutes and 20 seconds.

[0150] If the user does not perform any operation within the preset time period, the target video will start playing from the first playback time of 35 minutes and 22 seconds.

[0151] Optionally, when a user performs a swipe or drag operation on the dialogue progress bar in the video playback interface, the swipe distance corresponding to each dialogue entry is determined based on the maximum number of dialogue entries in the dialogue entry set and the total length of the swipe range. For example, if the maximum number of dialogue entries in the dialogue entry set is t, and the total length of the drag range of the dialogue progress bar is h pixels, then the swipe distance s corresponding to each dialogue entry is s = h / t.

[0152] In this embodiment, the first playback moment is obtained based on the user's target adjustment operation triggered by the playback progress of the target video. Then, the text information set corresponding to the first playback moment is acquired and displayed. Thus, the user can learn about the video content near the first playback moment from the text information set. Afterward, by selecting target text information from the text information set, the second playback moment in the target video that the user needs is precisely located, and the target video is played from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress.

[0153] Optionally, since the preview image of the video at the current playback moment allows users to intuitively understand the video content corresponding to the current playback moment, simultaneously displaying the video preview image and the dialogue in the video can help users more accurately understand the content of each moment in the video, thereby precisely locating the required video segment. Therefore, in this embodiment, a first preview image and a set of text information corresponding to the first playback moment in the target video are obtained, and the first preview image and the at least one text information are displayed in the video playback interface according to their temporal association with the first preview image.

[0154] Specifically, the first preview image can be generated based on video frames played at the first playback time, or it can be generated based on video frames played at the first playback time and video frames before and / or after the first playback time. The time interval between the video frames before and after the first playback time and the video frames played at the first playback time is less than a first threshold. The first preview image is the preview image corresponding to the first playback time, and the temporal association relationship refers to the relationship between the playback time of each text message and the first playback time. The first preview image can be associated with a point in time or a time period. If the first preview image is associated with a point in time, then the first playback time is the point in time associated with the first preview image. If the first preview image is associated with a time period, then the first playback time can be the beginning, end, or any point in time within the time period associated with the first preview image.

[0155] Optionally, the first preview image and various text information can be displayed in separate sections, specifically including the following implementation methods:

[0156] Implementation Method 1: Based on the temporal association between at least one text message and the first preview image, at least one text message is displayed in the first area of ​​the video playback interface, and the first preview image is displayed in the second area of ​​the video playback interface.

[0157] Specifically, the first region and the second region can be two non-intersecting regions or two partially intersecting regions. The shapes of the first and second regions and their positions on the video playback interface can be set according to actual needs.

[0158] For example, the target video is set as a football match video. When the user slides the video progress bar to the first playback time "35 minutes and 22 seconds", the first playback time corresponding to the first preview image is "35 minutes and 22 seconds". The text information set is a set of dialogue information, including five dialogue information, namely: {first time point "35 minutes and 18 seconds": first dialogue information "Team A has scored first"}, {second time point "35 minutes and 20 seconds": second dialogue information "the score on the field is now 1:0"}, {third time point "35 minutes and 22 seconds": third dialogue information "with a goal conceded, Team B may be under pressure"}, {fourth time point "35 minutes and 24 seconds": fourth dialogue information "Team B needs to cheer"}, {fifth time point "35 minutes and 26 seconds": fifth dialogue information "Team B starts to attack"}.

[0159] like Figure 12 As shown, in the first area 1201, the second, third, and fourth lines of dialogue are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. The font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1203. The first preview image is displayed in the second area 1202.

[0160] In this application, the first preview image and text information set are displayed in different areas, allowing users to intuitively understand the video content of the target video near the first playback time from the first preview image and text information set. At the same time, it is convenient for users to select each text information in the text information set, so as to achieve precise positioning of the video playback progress and improve the efficiency and accuracy of adjusting the video viewing progress.

[0161] Implementation Method 2: Based on the temporal association relationship between at least one text information and the first preview image, at least one text information and the person information associated with each of the at least one text information in the target video are displayed in the first area of ​​the video playback interface, and the first preview image and the person information associated with the first preview image in the target video are displayed in the second area of ​​the video playback interface.

[0162] Specifically, voiceprint features can be extracted from videos using a trained voiceprint model. Then, based on these voiceprint features and a pre-defined person information table, the person information corresponding to each text message in the target video can be determined. Alternatively, the person information corresponding to each text message in the target video can be determined manually. Person information can include character names, actor names, genders, ages, etc. After obtaining the person information, it can be associated with the text information and included as part of the text information set. When displaying each text message in the first area, the associated person information is also displayed. Similarly, person information can be associated with the first preview image; when displaying the first preview image, the corresponding person information is also displayed.

[0163] For example, if the target video is set to a TV series, and the user slides the video progress bar to the first playback time "35 minutes and 22 seconds", the character information associated with the first preview image is "Teacher Li" in the target video.

[0164] The text information set is a set of dialogue information, including five dialogue information sets, namely: {First time point "35 minutes 18 seconds": First dialogue information "Students, today is Arbor Day": Character information "Teacher Li"}, {Second time point "35 minutes 20 seconds": Second dialogue information "The school organizes tree planting activities": Character information "Teacher Li"}, {Third time point "35 minutes 22 seconds": Third dialogue information "Work in pairs": Character information "Teacher Li"}, {Fourth time point "35 minutes 24 seconds": Fourth dialogue information "Everyone, please pay attention to safety": Character information "Teacher Li"}, {Fifth time point "35 minutes 26 seconds": Fifth dialogue information "Okay": Character information "Student"}.

[0165] like Figure 13 As shown, in the first area 1301, the second, third, and fourth lines of dialogue, along with their corresponding character information, are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. In the second area 1302, the first preview image and its associated character information are displayed. During display, the font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1303.

[0166] It should be noted that in this embodiment, the first area of ​​the video playback interface may display each piece of text information and the person information associated with each piece of text information in the target video. In the second area of ​​the video playback interface, only the first preview image is displayed, without displaying the person information associated with the first preview image in the target video. Alternatively, the first area of ​​the video playback interface may only display each piece of text information, without displaying the person information associated with each piece of text information in the target video, while the second area of ​​the video playback interface displays the first preview image and the person information associated with the first preview image in the target video. This application does not impose specific limitations on this approach.

[0167] In this embodiment of the application, when displaying the first preview image and text information set, the first preview image and / or text information set, along with the associated character information in the target video, are displayed simultaneously. This makes it easier for the user to locate the character corresponding to the preview image or text information, thereby facilitating the user to select each text information in the text information set. This enables precise positioning of the video playback progress and improves the efficiency and accuracy of adjusting the video viewing progress.

[0168] Implementation Method 3: According to the temporal association relationship between at least one text information and the first preview image, at least one text information and the person information associated with each of the at least one text information in the target video are displayed in the first area of ​​the video playback interface, and the first playback time and the first preview image are displayed in the second area of ​​the video playback interface.

[0169] For example, a set of dialogue information and a first preview image are set, and... Figure 13 The set of dialogue information shown is the same as that in the first preview image, so it will not be repeated here.

[0170] like Figure 14 As shown, in the first area 1401, the second, third, and fourth lines of dialogue, along with their corresponding character information, are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. In the second area 1402, the first preview image and its playback time "35 minutes and 22 seconds" are displayed. During the display, the font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1403.

[0171] It should be noted that, in the embodiments of this application, the playback time of each text message or the playback time of text messages located at the target selection position may also be displayed. In this regard, this application does not make specific limitations.

[0172] In this embodiment, when displaying the first preview image and the set of text information, the playback time of the first preview image is shown, allowing the user to intuitively understand the playback time of the target video after initial adjustments. Simultaneously, the information about the characters in the target video associated with each piece of text in the text information set is displayed, facilitating the user's location of the characters speaking the various text messages in the target video. This allows the user to select individual text messages from the text information set, enabling precise positioning of the video playback progress and improving the efficiency and accuracy of adjusting the video viewing progress.

[0173] Implementation Method 4: Based on the temporal association relationship between at least one text information and the first preview image, display at least one text information, the person information and content tags associated with each of the at least one text information in the target video in the first area of ​​the video playback interface, and display the first preview image, the person information and content tags associated with the first preview image in the target video in the second area of ​​the video playback interface.

[0174] Specifically, content tags are used to characterize the main or highlight content of a video segment within a target video. The target video can be pre-divided into multiple video segments, and then the highlight content of each video segment can be determined using a neural network model or manual methods. Content tags are then assigned to each video segment based on this highlight content. For any given text information, the corresponding video segment is first determined based on the playback time of the text information, and then the content tag corresponding to the video segment is used as the content tag for the text information. Similarly, the corresponding video segment is first determined based on the playback time of the first preview image, and then the content tag corresponding to the video segment is used as the content tag for the first preview image.

[0175] For example, if the target video is set to a TV series, and the user slides the video progress bar to the first playback time "35 minutes and 22 seconds", the character information associated with the first preview image is "Teacher Li" in the target video, and the content tag associated with the first preview image is "Arbor Day".

[0176] The text information set is a set of dialogue information, including five dialogue information sets, namely: {First time point "35 minutes 18 seconds": First dialogue information "Students, today is Arbor Day": Character information "Teacher Li": Content tag "Arbor Day"}, {Second time point "35 minutes 20 seconds": Second dialogue information "The school organizes tree planting activities": Character information "Teacher Li": Content tag "Arbor Day"}, {Third time point "35 minutes 22 seconds": Third dialogue information "Do it in pairs": Character information "Teacher Li": Content tag "Arbor Day"}, {Fourth time point "35 minutes 24 seconds": Fourth dialogue information "Everyone, please pay attention to safety": Character information "Teacher Li": Content tag "Arbor Day"}, {Fifth time point "35 minutes 26 seconds": Fifth dialogue information "Okay": Character information "Student": Content tag "Arbor Day"}.

[0177] like Figure 15 As shown, in the first area 1501, the second, third, and fourth lines of dialogue are displayed sequentially from top to bottom, along with corresponding character information and content tags. The first and fifth lines of dialogue are hidden. In the second area 1502, the first preview image and its associated character information and content tags are displayed. When displayed, the font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1503.

[0178] In this embodiment, when displaying the first preview image and text information set, the first preview image and / or text information set, along with the associated character information and content tags in the target video, are simultaneously displayed. This facilitates the user's location of the character corresponding to the preview image or text information and the corresponding video content. Consequently, it allows the user to select each piece of text information in the text information set, enabling precise positioning of the video playback progress and improving the efficiency and accuracy of adjusting the video viewing progress.

[0179] It should be noted that the implementation methods for partitioning the first preview image and various text information in this application are not limited to the above-mentioned methods, and other display styles are also possible. Furthermore, the displayed content is not limited to character information, playback time, and content tags corresponding to the first preview image and text information set; it can also include user view counts, user discussion counts, etc., which are not specifically limited in this application. Additionally, after partitioning the first preview image and various text information using any of the above implementation methods, users can drag the dialogue progress bar in the first area or directly slide in the first area to display the hidden text information on the video playback interface. Simultaneously, users can also select the second playback time through clicks, double-clicks, long presses, swipes, etc., and start playing the target video from the second playback time.

[0180] For example, after a user slides the horizontal video progress bar to the first playback point "35 minutes and 22 seconds", the video playback interface will look like this: Figure 14 As shown. In the first area 1401, slide downwards to move the second dialogue message "The school organizes a tree planting activity" to the target selection position 1403, as shown. Figure 16a As shown, the video playback interface displays the first, second, and third lines of dialogue, with the font size of the second line of dialogue being larger than that of the other lines.

[0181] The terminal device obtains the playback time "35 minutes and 20 seconds" of the second dialogue information located at the target selection position 1403, and starts playing the TV series from 35 minutes and 20 seconds. For example... Figure 16b As shown, the video playback interface displays a video frame at 35 minutes and 20 seconds.

[0182] Optionally, in step S501 above, in response to a target adjustment operation triggered by the playback progress of the target video, after determining the first playback moment of the adjusted target video, at least one second preview image is acquired that satisfies a preset condition regarding the playback time interval between the second preview image and the first preview image, and that the similarity between the second preview image and the first preview image is less than a preset threshold. Then, in the second region, at least one second preview image is displayed according to the temporal association relationship between each of the at least one second preview image and the first preview image.

[0183] Specifically, the playback time interval between the second preview image and the first preview image meets a preset condition, which can be that the playback time interval between the second preview image and the first preview image is less than a preset threshold. At least one of the second preview images can be a preview image played before or after the first playback time, or it can be a preview image played before and after the first playback time.

[0184] For example, if the target video is a TV series, after the user slides the video progress bar to the first playback time "35 minutes and 22 seconds", the terminal device obtains a first preview image and a second preview image. The first playback time corresponding to the first preview image is "35 minutes and 22 seconds", and the first playback time corresponding to the second preview image is "35 minutes and 26 seconds".

[0185] The text information set is a set of dialogue information, including five dialogue information sets, namely: {First time point "35 minutes 18 seconds": First dialogue information "Students, today is Arbor Day": Character information "Teacher Li"}, {Second time point "35 minutes 20 seconds": Second dialogue information "The school organizes tree planting activities": Character information "Teacher Li"}, {Third time point "35 minutes 22 seconds": Third dialogue information "Work in pairs": Character information "Teacher Li"}, {Fourth time point "35 minutes 24 seconds": Fourth dialogue information "Everyone, please pay attention to safety": Character information "Teacher Li"}, {Fifth time point "35 minutes 26 seconds": Fifth dialogue information "Okay": Character information "Student"}.

[0186] like Figure 17 As shown, in the first area 1701, the second, third, and fourth lines of dialogue, along with the corresponding character information, are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. The font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1703.

[0187] In the second area 1702, the first preview image and the second preview image are displayed from left to right, along with the playback times corresponding to the first preview image and the second preview image, respectively.

[0188] In this embodiment of the application, when displaying the first preview image and the set of text information, at least one second preview image that is near the first preview image and has a similarity of less than a preset threshold to the first preview image is also displayed. Therefore, based on the first preview image, the second preview image and the set of text information, the user can more intuitively understand the video content near the first playback time, which makes it easier for the user to more precisely locate the video playback progress and improve the efficiency and accuracy of adjusting the video viewing progress.

[0189] Optionally, in the second area, information about the person associated with each of the two second preview images in the target video is displayed.

[0190] For example, setting the target video as a TV series, after the user slides the video progress bar to the first playback time "35 minutes and 22 seconds", the terminal device obtains a first preview image and a second preview image. The first preview image corresponds to the first playback time "35 minutes and 22 seconds", and the second preview image corresponds to the first playback time "35 minutes and 26 seconds". The character information associated with the first preview image is "Teacher Li" in the target video, and the character information associated with the second preview image is "student" in the target video. The dialogue information set and... Figure 17 The set of dialogue texts shown is the same, so it will not be repeated here.

[0191] like Figure 18As shown, in the first area 1701, the second, third, and fourth lines of dialogue, along with the corresponding character information, are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. The font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 1703.

[0192] In the second area 1702, from left to right, the first preview image and the second preview image are displayed, along with the person information associated with the first preview image and the second preview image, respectively.

[0193] It should be noted that when displaying the second preview image, it is not only necessary to display the information of the person associated with the second preview image, but also the playback time, content tags, number of user views, number of user discussions, etc. The applicant does not make any specific restrictions on this.

[0194] In this embodiment, while displaying the first preview image and the text information set, at least one second preview image that is near the first preview image and has a similarity to the first preview image of less than a preset threshold is also displayed. In addition, the first preview image, the second preview image, and the text information set, along with the person information associated with each of them, are also displayed. Therefore, based on the person information, the user can quickly locate the person corresponding to the preview image or text information, which facilitates the user to select each piece of text information in the text information set, thereby achieving precise positioning of the video playback progress and improving the efficiency and accuracy of adjusting the video viewing progress.

[0195] Optionally, since each second preview image corresponds to a playback moment, the user can also adjust the video playback progress by selecting a second preview image. Specifically, in response to a selection operation triggered for a target preview image in the preview image set, the third playback moment corresponding to the target preview image in the target video is obtained, and the target video is played starting from the third playback moment, wherein the preview image set includes a first preview image and at least one second preview image.

[0196] Specifically, the selection operation can be a click, double-click, long press, swipe, etc. The target preview image can be a first preview image or a second preview image. The third playback time can be the first playback time, or a playback time before or after the first playback time. As mentioned above, in some embodiments, the target preview image can be associated with a time point, then the third playback time can be the time point associated with the target preview image. In other embodiments, the target preview image can be associated with a time period, then the third playback time can be the start, end, or any time point within the time period associated with the target preview image. Optionally, if no operation is responded to within a preset duration, the target video is played starting from the first playback time.

[0197] For example, in Figure 17 Based on the video playback interface shown, the first preview image corresponds to a first playback time of "35 minutes and 22 seconds," and the second preview image corresponds to a first playback time of "35 minutes and 26 seconds." If the user clicks on the second preview image in the video playback interface, the video application will start playing the target video at 35 minutes and 26 seconds.

[0198] In this application, since each preview image corresponds to a playback moment, users can also adjust the video playback progress by selecting a preview image, giving users more ways to adjust the video playback progress and thus improving the convenience of adjusting the video progress.

[0199] based on Figure 1 The system architecture diagram shown illustrates a video playback method provided in this embodiment. This method is executed by a computer device, which may be... Figure 1 The terminal device 101 shown includes the following steps:

[0200] In response to a target adjustment operation triggered by the playback progress of the target video, a set of text information is displayed in the video playback interface. The set of text information includes at least one piece of text information, and the target adjustment operation is used to determine a first playback moment in the playback progress of the target video. The set of text information corresponds to the first playback moment. In response to a selection operation triggered by a target text information in the at least one piece of text information, the target video is played starting from a second playback moment, where the second playback moment is the playback moment associated with the target text information.

[0201] One possible implementation includes a target adjustment operation that involves dragging the video progress bar along its direction in the video playback interface, and a selection operation triggered for a target text information in at least one text message that includes a sliding operation perpendicular to the video progress bar direction.

[0202] Specifically, the direction in which the video progress bar is dragged corresponds to the direction of video playback adjustment, and the drag distance of the video progress bar corresponds to the length of the video segment to be jumped to. In practice, the duration of the video segment to be jumped to is determined based on the total video duration and the total drag range of the video progress bar. Target adjustment operations can also be performed by sliding directly on the video playback interface, or by using fast forward or rewind buttons to control the video playback progress. The selection operation triggered by the target text information in at least one text message can also be by dragging a dialogue progress bar perpendicular to the direction of the video progress bar.

[0203] One possible implementation is to arrange and display at least one text message in the video playback interface according to the playback time associated with at least one text message.

[0204] One possible implementation involves displaying, in the video playback interface, at least one text message associated with a person in the target video.

[0205] Specifically, a trained voiceprint model can be used to extract voiceprint features from a video. Then, based on these features and a pre-defined list of personal information, the personal information corresponding to each piece of text in the target video can be determined. Alternatively, the personal information corresponding to each piece of text in the target video can be determined manually. Personal information can include character names, actor names, genders, ages, etc. After obtaining the personal information, it can be associated with the text information and included as part of the text information set. When displaying each piece of text information, the associated personal information is also displayed simultaneously.

[0206] In this embodiment, the first playback moment is obtained based on the user's target adjustment operation triggered by the playback progress of the target video. Then, the text information set corresponding to the first playback moment is acquired and displayed. Thus, the user can learn about the video content near the first playback moment from the text information set. Afterward, by selecting target text information from the text information set, the second playback moment in the target video that the user needs is precisely located, and the target video is played from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress.

[0207] To better explain the embodiments of this application, a video playback method provided by the embodiments of this application is described below in conjunction with a specific implementation scenario. This method is executed interactively by a user, a terminal device, and a server, and includes the following steps: Figure 19 As shown:

[0208] In step S1901, the user clicks the icon of the target video in the video application on the terminal device.

[0209] Step S1902: The terminal device plays the target video.

[0210] In step S1903, the terminal device sends a request for dialogue data of the target video to the server.

[0211] In step S1904, the server sends the dialogue data of the target video to the terminal device.

[0212] Specifically, the server extracts the audio information of the target video offline, and then generates the dialogue information of the target video through speech recognition and manual correction. Further, using a trained voiceprint model, voiceprint features are extracted from the target video, and combined with a pre-set character information table, the character information corresponding to each line of dialogue in the target video is identified. The dialogue data format can be represented as {"time point": "dialogue information": character information}, sorted sequentially by time point.

[0213] Step S1905: The terminal device loads and saves the dialogue data of the target video locally.

[0214] Specifically, the terminal device stores the dialogue data of the target video locally on the hard drive or in memory. Since the dialogue data is text data and occupies little space, obtaining the dialogue data and saving it locally before scrolling the video progress bar does not require a lot of storage space on the terminal device. At the same time, the corresponding dialogue data can be quickly obtained when the user scrolls the video progress bar, and it also avoids consuming bandwidth resources when requesting preview images in real time.

[0215] Step S1906: The user slides the video progress bar horizontally in the video playback interface.

[0216] In step S1907, the terminal device responds to the operation of sliding the video progress bar and obtains the first playback moment of the adjusted target video.

[0217] In step S1908, the terminal device sends a preview image acquisition request to the server.

[0218] The preview image retrieval request includes the first playback moment.

[0219] In step S1909, the server sends the first preview image corresponding to the first playback moment to the terminal device.

[0220] Step S1910: The terminal device obtains the set of dialogue information corresponding to the first playback moment.

[0221] The maximum number of dialogue information items contained in the dialogue information set is set to the number of dialogue information items corresponding to the video segment to which the target video jumps when the video progress bar slides by one pixel.

[0222] In step S1911, the terminal device displays the dialogue information and related character information in the dialogue information set in the first area, and displays the first preview image and the first playback moment in the second area.

[0223] In step S1912, the user selects a target line from the various line information.

[0224] In step S1913, the terminal device obtains the second playback time corresponding to the target dialogue information and starts playing the target video from the second playback time.

[0225] For example, if the target video is a 2-hour TV series, on a 390-pixel x 844-pixel terminal device, the video progress bar needs to be mapped to 844 video segments across 844 pixels. Each horizontal pixel movement corresponds to approximately 8.5 seconds of the target video. If we set the average number of lines per 8.5-second video segment in the target video to 't', then within 390 pixels, a maximum of 't' lines of dialogue will be mapped. Each time a line of dialogue is scrolled, the vertical progress bar needs to be dragged a distance of 390 / 't' pixels.

[0226] like Figure 20a As shown, the user drags the horizontal video progress bar 2001 to the first playback time "35 minutes and 22 seconds" to obtain the first preview image corresponding to the first playback time. The character information associated with the first preview image is "Teacher Li" in the target video. The dialogue information set includes five dialogue information, namely: {First time point "35 minutes and 18 seconds": First dialogue information "Students, today is Arbor Day": Character information "Teacher Li"}, {Second time point "35 minutes and 20 seconds": Second dialogue information "The school organizes a tree planting activity": Character information "Teacher Li"}, {Third time point "35 minutes and 22 seconds": Third dialogue information "Work in pairs": Character information "Teacher Li"}, {Fourth time point "35 minutes and 24 seconds": Fourth dialogue information "Everyone, please pay attention to safety": Character information "Teacher Li"}, {Fifth time point "35 minutes and 26 seconds": Fifth dialogue information "Okay": Character information "Student"}.

[0227] In the first area 2002, the second, third, and fourth lines of dialogue, along with their corresponding character information, are displayed sequentially from top to bottom. The first and fifth lines of dialogue are hidden. In the second area 2003, the first preview image and the first playback moment are displayed. During the display, the font size of the third line of dialogue is larger than that of the other lines of dialogue, and the third line of dialogue is located in the target selection position 2004.

[0228] The user drags the vertical dialogue progress bar 2005 downwards in the first area 2002, and slides the second dialogue information "The school organizes a tree planting activity" to the target selection position 2004. For example... Figure 20b As shown, the video playback interface displays the first, second, and third lines of dialogue, with the font size of the second line of dialogue being larger than that of the other lines.

[0229] The terminal device obtains the playback time "35 minutes and 20 seconds" of the second dialogue information located at the target selection position 2004, and starts playing the TV series from 35 minutes and 20 seconds. The video playback interface displays the video screen at 35 minutes and 20 seconds.

[0230] In this embodiment, when the user drags the video progress bar, the dialogue information before and after the current moment on the progress bar is dynamically displayed. This dialogue preview allows for a quick understanding of the video content corresponding to the current moment. Compared to a video preview image, pre-acquiring and storing the dialogue information locally allows for real-time display without any delay, avoiding the common problem of video preview images failing to clearly show the video content. Furthermore, the vertical progress bar dragging method ensures that while the horizontal video progress bar can be adjusted significantly, it also allows for minute-level adjustments, improving the efficiency and accuracy of adjusting the video viewing progress. This provides users with a refined progress dragging experience, effectively enhancing the viewing experience in scenarios such as rewinding or fast-forwarding.

[0231] Based on the same technical concept, embodiments of this application provide a video playback device, such as... Figure 21 As shown, the device 2100 includes:

[0232] The adjustment response module 2101 is used to respond to a target adjustment operation triggered by the playback progress of the target video and determine the first playback moment of the adjusted target video.

[0233] The acquisition module 2102 is used to acquire a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information;

[0234] Display module 2103 is used to display the at least one text information in the video playback interface;

[0235] The video playback module 2104 is configured to respond to a selection operation triggered by target text information in the at least one text information, obtain a second playback time corresponding to the target text information in the target video, and start playing the target video from the second playback time.

[0236] Optionally, the display module 2103 is specifically used for:

[0237] In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

[0238] Optionally, the display module 2103 is specifically used for:

[0239] The video playback interface displays the information of the person associated with each of the at least one text information in the target video.

[0240] Optionally, the acquisition module 2102 is specifically used for:

[0241] Obtain the set of first preview image and text information corresponding to the first playback moment in the target video;

[0242] The display module 2103 is specifically used for:

[0243] Based on the temporal relationship between the at least one text information and the first preview image, the first preview image and the at least one text information are displayed in the video playback interface.

[0244] Optionally, the display module 2103 is specifically used for:

[0245] According to the time sequence relationship between the at least one text information and the first preview image, the at least one text information is displayed in the first area of ​​the video playback interface;

[0246] The first preview image is displayed in the second area of ​​the video playback interface.

[0247] Optionally, the display module 2103 is further configured to:

[0248] In the first area, the personal information associated with each of the at least one text information in the target video is displayed.

[0249] In the second area, information about the person associated with the first preview image in the target video is displayed.

[0250] Optionally, the set of text information includes first text information played at the first playback time, and at least two second text information played before and after the first playback time.

[0251] Optionally, the acquisition module 2102 is further configured to:

[0252] In response to a target adjustment operation triggered by the playback progress of a target video, after determining the first playback moment of the adjusted target video, at least one second preview image is obtained that satisfies a preset condition for the playback time interval between the target video and the first preview image, and has a similarity to the first preview image that is less than a preset threshold.

[0253] The display module 2103 is also used for:

[0254] In the second region, the at least one second preview image is displayed according to its temporal association with the first preview image.

[0255] Optionally, the display module 2103 is further configured to:

[0256] In the second area, the person information associated with each of the at least one second preview image in the target video is displayed.

[0257] Optionally, the video playback module 2104 is further configured to:

[0258] According to the temporal association relationship between at least one text information in the text information set and the first preview image, after displaying the at least one text information in the video playback interface, in response to the selection operation triggered for the target preview image in the preview image set, the third playback time corresponding to the target preview image in the target video is obtained, and the target video is played from the third playback time, wherein the preview image set includes the first preview image and the at least one second preview image.

[0259] Optionally, the video playback module 2104 is further configured to:

[0260] According to the temporal relationship between at least one text information in the text information set and the first preview image, after displaying the at least one text information in the video playback interface, if no operation is responded to within a preset time period, the target video will start playing from the first playback time.

[0261] Optionally, the target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface. The maximum number of text information contained in the text information set is the number of text information corresponding to the video segment to which the target video jumps when the video progress bar is dragged a unit distance. The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

[0262] In this embodiment, a first playback moment is obtained based on the user's target adjustment operation triggered by the playback progress of the target video. Then, a first preview image and a corresponding set of text information corresponding to the first playback moment are acquired and displayed. Therefore, the user can intuitively understand the video content near the first playback moment from the first preview image and the set of text information. Then, by selecting target text information from the set of text information, the user can precisely locate the second playback moment in the target video and start playing the target video from the second playback moment, thereby improving the efficiency and accuracy of adjusting the video viewing progress.

[0263] Based on the same technical concept, embodiments of this application provide a video playback device, such as... Figure 22 As shown, the device 2200 includes:

[0264] The adjustment response module 2201 is used to respond to a target adjustment operation triggered by the playback progress of the target video and display a set of text information in the video playback interface. The set of text information includes at least one piece of text information. The target adjustment operation is used to determine a first playback moment in the playback progress of the target video. The set of text information corresponds to the first playback moment.

[0265] The video playback module 2202 is configured to play the target video starting from a second playback time in response to a selection operation triggered for target text information in the at least one text information, wherein the second playback time is the playback time associated with the target text information.

[0266] Optionally, the target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface;

[0267] The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

[0268] Optionally, the adjustment response module 2201 is specifically used for:

[0269] In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

[0270] Optionally, the adjustment response module 2201 is specifically used for:

[0271] The video playback interface displays the information of the person associated with each of the at least one text information in the target video.

[0272] Based on the same technical concept, embodiments of this application provide a computer device, which may be a terminal or a server, such as... Figure 23 As shown, it includes at least one processor 2301 and a memory 2302 connected to at least one processor. In this embodiment, the specific connection medium between the processor 2301 and the memory 2302 is not limited. Figure 23 Taking the connection between the processor 2301 and the memory 2302 via a bus as an example, the bus can be divided into address bus, data bus, control bus, etc.

[0273] In this embodiment of the application, the memory 2302 stores instructions that can be executed by at least one processor 2301. By executing the instructions stored in the memory 2302, at least one processor 2301 can perform the steps included in the above-described video playback method.

[0274] The processor 2301 is the control center of the computer device. It can connect to various parts of the computer device using various interfaces and lines, and controls video playback by running or executing instructions stored in the memory 2302 and calling data stored in the memory 2302. Optionally, the processor 2301 may include one or more processing units. The processor 2301 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 2301. In some embodiments, the processor 2301 and the memory 2302 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0275] Processor 2301 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0276] Memory 2302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 2302 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 2302 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 2302 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0277] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the video playback method described above.

[0278] Those skilled in the art will understand that embodiments of the present invention can be provided as methods or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0279] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0280] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0281] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0282] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0283] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A video playback method, characterized in that, include: In response to a target adjustment operation triggered by the playback progress of the target video, the first playback moment of the adjusted target video is determined; Obtain a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information; the difference between the time point associated with each piece of text information in the set of text information and the first playback moment is within a predetermined threshold; the text information in the set of text information is dialogue information. Display at least one text message in the video playback interface; In response to a selection operation triggered for target text information in the at least one text information, a second playback time corresponding to the target text information in the target video is obtained, and the target video is played starting from the second playback time.

2. The method as described in claim 1, characterized in that, Displaying at least one text message in the video playback interface includes: In the video playback interface, the at least one text information is arranged and displayed according to the playback time associated with the at least one text information.

3. The method as described in claim 2, characterized in that, Also includes: The video playback interface displays the person information associated with each of the at least one text information in the target video.

4. The method as described in claim 1, characterized in that, The step of obtaining the set of text information corresponding to the first playback moment in the target video and displaying the at least one piece of text information in the video playback interface includes: Obtain the set of first preview images and text information corresponding to the first playback moment in the target video; Based on the temporal relationship between the at least one text information and the first preview image, the first preview image and the at least one text information are displayed in the video playback interface.

5. The method as described in claim 4, characterized in that, The step of displaying the first preview image and the at least one text information in the video playback interface according to their respective temporal association with the first preview image includes: According to the time sequence relationship between the at least one text information and the first preview image, the at least one text information is displayed in the first area of ​​the video playback interface; The first preview image is displayed in the second area of ​​the video playback interface.

6. The method as described in claim 5, characterized in that, Displaying the at least one text message in the first area of ​​the video playback interface further includes: In the first area, the personal information associated with each of the at least one text information in the target video is displayed; The second area in the video playback interface, which displays the first preview image, further includes: In the second area, information about the person associated with the first preview image in the target video is displayed.

7. The method as described in claim 5, characterized in that, After determining the first playback moment of the adjusted target video in response to a target adjustment operation triggered by the playback progress of the target video, the method further includes: Obtain at least one second preview image whose playback time interval with the first preview image meets a preset condition and whose similarity to the first preview image is less than a preset threshold; In the second region, the at least one second preview image is displayed according to its temporal association with the first preview image.

8. The method as described in claim 7, characterized in that, The display of the at least one second preview image also includes: In the second area, the person information associated with each of the at least one second preview image in the target video is displayed.

9. The method as described in claim 7, characterized in that, After displaying the first preview image and the at least one text information in the video playback interface according to their respective temporal association with the first preview image, the method further includes: In response to a selection operation triggered for a target preview image in the preview image set, a third playback moment corresponding to the target preview image in the target video is obtained, and the target video is played starting from the third playback moment, wherein the preview image set includes the first preview image and the at least one second preview image.

10. The method as described in claim 7, characterized in that, After displaying the first preview image and the at least one text information in the video playback interface according to their respective temporal association with the first preview image, the method further includes: If no operation is responded to within the preset time period, the target video will be played starting from the first playback moment.

11. The method according to any one of claims 1-10, characterized in that, The target adjustment operation includes a drag operation of dragging the video progress bar along the direction of the video progress bar in the video playback interface. The maximum number of text information contained in the text information set is the number of text information corresponding to the video segment to which the target video jumps when the video progress bar is dragged a unit distance. The selection operation triggered for the target text information in the at least one text information includes a sliding operation perpendicular to the direction of the video progress bar.

12. A video playback method, characterized in that, include: In response to a target adjustment operation triggered by the playback progress of a target video, a set of text information is displayed in the video playback interface. The set of text information includes at least one piece of text information. The target adjustment operation is used to determine a first playback moment in the playback progress of the target video. The set of text information corresponds to the first playback moment. The difference between the time point associated with each piece of text information in the set and the first playback moment is within a predetermined threshold. The text information in the set is dialogue information. In response to a selection operation triggered for target text information in the at least one text information, the target video is played starting from a second playback time, wherein the second playback time is the playback time associated with the target text information.

13. A video playback device, characterized in that, include: An adjustment response module is used to respond to a target adjustment operation triggered by the playback progress of a target video and determine the first playback moment of the adjusted target video. The acquisition module is used to acquire a set of text information corresponding to the first playback moment in the target video, wherein the set of text information includes at least one piece of text information; the difference between the time point associated with each piece of text information in the set of text information and the first playback moment is within a predetermined threshold; and the text information in the set of text information is dialogue information. A display module is used to display at least one text message in the video playback interface; A video playback module is configured to respond to a selection operation triggered by target text information in the at least one text information, obtain a second playback moment corresponding to the target text information in the target video, and start playing the target video from the second playback moment.

14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Jump navigation method for video playing

    CN111212317A

  • Ultrasonic video image processing method, system and device and storage medium

    CN112580613A