Video clip playing method and device, medium and product

Generating and storing video clips through the object detection model solves the problem that users find it difficult to play specific objects efficiently, and realizes an efficient method of quickly positioning and playing specific video clips, which improves the consistency of user experience and video playback.

CN120358395APending Publication Date: 2025-07-22SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510628265.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing video playback mode, it is difficult for users to efficiently locate and play video clips of specific objects, and they need to manually drag the progress bar to search, resulting in cumbersome and inefficient operations.

Method used

The object detection model recognizes the occurrence time information of objects in the video, generates and stores video clips, and directly extracts and plays video clips from specific objects from the database in response to user instructions, including regenerating and optimizing clips when no clips are stored in the database.

Benefits of technology

It enables users to quickly locate and play video clips of specific objects without manual search, improve playback efficiency and user experience, and ensure the consistency and quality of video clips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358395A_ABST
    Figure CN120358395A_ABST
Patent Text Reader

Abstract

The invention discloses a video clip playing method and device, a medium and a product, and relates to the technical field of video playing, and the method comprises the steps: in response to a playing instruction for a first object in an original video, extracting a first video clip corresponding to the first object in the original video from a preset database, the first video clip is generated according to appearance time information, obtained through the object detection model, of a first object in the original video; and playing the first video clip. According to the invention, the high-efficiency playing requirement of a user on the video clip of the specific object is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video playback, and particularly to a method, device, medium, and product for playing video segments. Background Art

[0002] In the conventional video playback mode, viewers can only watch according to the pre-set content and order of the video. However, with the continuous enrichment and diversification of video content, people's viewing needs for video content are also becoming increasingly diverse. Especially in videos such as TV dramas, reality shows, and sports events with numerous characters and scenes, the need for viewers to accurately watch segments related to specific objects is becoming more prominent. However, in the conventional video playback mode, viewers need to fast forward or pull the progress bar multiple times to find the segments related to specific tasks. Therefore, how to efficiently provide the playback of specific segments according to user needs during video playback has become an urgent problem to be solved.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, medium, and product for playing video segments, aiming to solve the technical problem of how to efficiently provide the playback of specific segments according to user needs during video playback.

[0005] To achieve the above purpose, this application proposes a method for playing video segments, and the method for playing video segments includes:

[0006] In response to a playback instruction for a first object in the original video, extract a first video segment corresponding to the first object in the original video from a preset database, where the first video segment is generated according to the appearance time information of the first object in the original video obtained through an object detection model;

[0007] Play the first video segment.

[0008] In one embodiment, after the step of responding to the playback instruction for the first object in the original video, it further includes:

[0009] In the case where the first video segment is not in the database, input the feature information of the first object in the original video into the object detection model to obtain the appearance time information of the first object;

[0010] Determine a target sub-segment corresponding to the appearance time information from the original video;

[0011] Generate the first video segment according to the target sub-segment.

[0012] In one embodiment, the step of generating the first video segment according to the target sub-segment includes:

[0013] Obtain the voiceprint information of the first object;

[0014] Adjust the target sub-segment according to the voiceprint information;

[0015] Generate the first video segment according to the adjusted target sub-segment.

[0016] In one embodiment, the step of generating the first video segment according to the target sub-segment includes:

[0017] Determine adjacent first and second sub-segments according to the time interval between adjacent sub-segments in the target sub-segment, wherein the time interval between the first sub-segment and the second sub-segment is greater than a preset interval threshold;

[0018] Generate a transition video segment based on the video segment between the first sub-segment and the second sub-segment in the original video;

[0019] Generate the first video segment according to the transition video segment and the target sub-segment.

[0020] In one embodiment, the video segment playing method further includes:

[0021] When the original video being currently played is paused, obtain the first current frame displayed;

[0022] Identify the current object included in the first current frame and generate an information list of the current object;

[0023] Display the information list of the current object.

[0024] In one embodiment, after the step of playing the first video segment, it further includes:

[0025] In response to a play instruction for a second object in the original video, obtain the current time point corresponding to the second current frame displayed in the original video;

[0026] Extract the second video segment corresponding to the second object from the database;

[0027] Determine a third video segment from the second video segment according to the current time point, wherein the start moment of the third video segment in the original video is later than or equal to the current time point.

[0028] In one embodiment, after the step of playing the first video segment, it further includes:

[0029] Obtain the playback parameters during the playback of the first video segment;

[0030] Optimize the playback process based on the playback parameters.

[0031] In addition, to achieve the above object, the present application also proposes a video segment playback device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the video segment playback method as described above.

[0032] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the video segment playback method as described above.

[0033] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the video segment playback method as described above.

[0034] In response to a playback instruction for a first object in an original video, the present application obtains the actual playback requirements of the user for the first object, and then extracts the first video segment corresponding to the first object in the original video from a preset database. In the embodiments of the present application, the object detection model is used in advance to identify the first object in the original video to obtain the appearance time information of the first object in the original video, and according to the appearance time information of the first object, the first video segment in which the first object appears in the original video is generated, and then stored in a preset database. When the user has a playback requirement of "only watching" the first object, the first video segment can be directly extracted and played from the preset database, so that the user does not need to manually drag the progress bar to search, saving time and operation steps, realizing efficient playback of specific video segments according to user requirements, and effectively improving the user experience. Description of the Drawings

[0035] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0037] Figure 1Schematic flowchart of the video clip playing method according to the embodiments of the present application;

[0038] Figure 2 Schematic diagram of the scenario for playing the first video clip according to an embodiment of the present application;

[0039] Figure 3 Schematic diagram of the scenario for playing the second video clip according to an embodiment of the present application;

[0040] Figure 4 Schematic diagram of the module structure of the video clip playing device according to the embodiments of the present application;

[0041] Figure 5 Schematic diagram of the structure of the video clip playing device involved in the video clip playing method according to the embodiments of the present application.

[0042] Explanation of the reference numerals in the drawings:

[0043] 100, terminal; 101, video playback interface; 111, video name column;

[0044] 112, "Watch Only Him" button; 113, video operation button bar;

[0045] 114, video progress bar; 115, object list window;

[0046] 116, current playback progress; 117, first object; 118, second object.

[0047] The realization of the purpose, functional features and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0048] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0049] For a better understanding of the technical solutions of the present application, the following will be described in detail in conjunction with the drawings in the specification and the specific implementation manners.

[0050] In the related art, most video playback methods play videos in the original order, and only a few popular videos or TV dramas can implement the function of only watching video clips of the main object. Therefore, there is an urgent need to provide an efficient video playback method that can meet the user's viewing needs for specific objects.

[0051] In this embodiment, the appearance time information of each object in the original video can be pre-identified through an object detection model, and video segments corresponding to each object can be generated according to the appearance time information of each object and stored in a preset database. Furthermore, when the user has a "watch only" requirement for a specific object, that is, the first object, in the original video, the user can trigger a play instruction for the first object. Then, the terminal responds to the play instruction for the first object, filters out the video segment corresponding to the first object from the preset database and plays it, thus saving the user's operation steps and process, enabling the user to quickly find the video segment corresponding to the first object without repeatedly dragging the progress bar, improving the efficiency of video segment playback, and meeting the user's high-efficiency playback requirement for the video segment of a specific object.

[0052] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, a television, etc., or an electronic device, a terminal, etc. that can implement the above functions. Hereinafter, the terminal will be taken as an example to illustrate this embodiment and the following embodiments.

[0053] Based on this, an embodiment of the present application provides a method for playing video segments, referring to Figure 1 . In this embodiment, the method for playing video segments includes steps S10 to S20:

[0054] Step S10, in response to a play instruction for the first object in the original video, extract the first video segment corresponding to the first object in the original video from a preset database, where the first video segment is generated according to the appearance time information of the first object in the original video obtained through an object detection model;

[0055] In a feasible embodiment, a database is preset, and the video segments of each object included in the original video are stored in the database; the appearance time information of the first object in the original video can be obtained through an object detection model, and then video segments corresponding to each object can be generated according to the appearance time information of each object. When the user determines the first object and has a play requirement for the video segment of the first object, the terminal responds to the play instruction and extracts the first video segment corresponding to the first object from the preset database.

[0056] Optionally, the original video refers to a video file that has not been processed or has only been basically encoded and compressed, and includes complete video pictures and audio information. Optionally, the original video can include movies, TV series, short videos, live videos, teaching videos, sports event videos, etc.

[0057] Optionally, the first object refers to a specific object that the user is concerned about or designates in a specific scenario. This object is related to the video content and may be a person, thing, event, etc. In the context of video processing and playback, the first object is the key element that triggers the playback operation, and the terminal will filter and extract relevant video segments based on the first object.

[0058] Optionally, the play instruction refers to a command or request output by the user to start video playback. The play instruction can be triggered in various ways, such as clicking the play button, voice command, etc.

[0059] Optionally, the database refers to organizing, storing, and managing data according to a data structure, which is used to efficiently store, query, and manage a large amount of data. The database can store each video segment corresponding to the appearance time information of the first object. In response to the play instruction of the first object, each video segment is extracted from the database, integrated into the first video segment and then played; or the database stores the first video segment obtained by integrating each video segment corresponding to the appearance time information of the first object. In response to the play instruction of the first object, the first video segment is directly extracted from the database and then played.

[0060] Optionally, an intelligent distribution server background can be set up to plan the videos that each terminal needs to identify and distribute tasks. Each terminal performs identification tasks through an object detection model, and stores the video segments corresponding to each identified object in a preset database for subsequent extraction processes.

[0061] Optionally, the first video segment refers to a specific video part extracted from the original video and related to the first object. It is a video segment filtered and intercepted from the original video according to the characteristics of the first object and the user's needs, and only contains content related to the first object. For example, in a movie, if the first object is an exciting fight scene of the protagonist, then the intercepted fight video from the movie is the first video segment; in a football game, if the first object is a wonderful shot of a certain player, the corresponding shot video segment is the first video segment.

[0062] Optionally, the object detection model refers to a computer vision algorithm based on deep learning. Its core function is to identify the position and appearance time of specific objects (such as people, objects, actions) in the video, and generate structured data (such as bounding boxes, class labels, timestamps) to support the accurate extraction of video segments.

[0063] Optionally, the object detection model can be a Faster R-CNN model (Faster Region-based Convolutional Neural Network), a YOLO model (You Only Look Once), and an SSD model (Single Shot MultiBox Detector).

[0064] Optionally, the appearance time information refers to the time range data of the appearance of a specific object (such as a person, an object, etc.) identified after analyzing the original video through the object detection model. It is used to accurately locate the existence period of the object in the video. The object detection model analyzes each frame of the original video and identifies each object contained therein. During the identification process, the time point when each object first appears in the video (which can be represented by the time stamp of the video playback, for example, the 5th second after the start of the video playback) and the time point when the object last appears in the video (such as the 20th second after the start of the video playback) are recorded. These two time points determine a continuous appearance time period of the object in the video. If the object appears discontinuously in the video, disappearing and reappearing midway, multiple such time periods will be recorded, and this time period information is the appearance time information.

[0065] Optionally, extraction refers to the operation process of finding and obtaining the first video segment corresponding to the first object from a preset database according to specific conditions (such as the playback instruction for the first object in the original video). When the terminal receives the playback instruction for the first object in the original video, it will search in the preset database according to the requirements of the instruction.

[0066] Step S20, play the first video segment.

[0067] In a feasible embodiment, a "Watch Only Him" button can be added below or in the sidebar of the player. Then, when the terminal is in the video playback interface and the user clicks the "Watch Only Him" button below or in the sidebar of the video playback interface, a watch-only window can be displayed. The watch-only window provides a list of object information for the user to select the object they want to watch. And the terminal finds the corresponding first video segment from the pre-set database according to the user's selection and then plays it.

[0068] Optionally, playing refers to the process of continuously presenting the video segment corresponding to the first object in chronological order of its video and audio content. It is to display only the video paragraphs related to the first object according to the user's instruction, rather than the complete original video.

[0069] Exemplarily, see Figure 2, there is a video playback interface 101 on the terminal 100. There are a video name bar 111, a "Watch Only Him" button 112, a video operation button bar 113, and a video progress bar 114 on the video playback interface 101. When the "Watch Only Him" button 112 is clicked, an object list window 115 pops up on the video playback interface 101. Each object in the video is listed and displayed on the object list window 115. When the user clicks on the list, the terminal can play the video clip of that object.

[0070] In this embodiment, the original video is analyzed by an object detection model to obtain the appearance time information of each object, and corresponding video clips are generated accordingly and stored in the database. This enables accurate positioning of the video content containing a specific object. When the user issues a playback instruction for the first object, the first video clip can be quickly extracted from the database, saving the user's time in searching for specific content and improving the viewing efficiency. Secondly, only the video segments related to the first object are displayed during playback, avoiding the user from watching a large amount of irrelevant content and allowing the user to focus on the parts they are interested in, enhancing the viewing experience.

[0071] Further, based on any of the above embodiments of the present application, a second embodiment of the video clip playback method is proposed. After the step S10 of responding to the playback instruction for the first object in the original video, the following steps are further included:

[0072] Step S101, in the case where the first video clip is not in the database, input the feature information of the first object in the original video into the object detection model to obtain the appearance time information of the first object;

[0073] In a feasible embodiment, when a playback instruction for the first object in the original video is received, the first video clip corresponding to the first object can be preferentially searched in the database. When the first video clip corresponding to the first object is not stored in the database, the feature information of the first object in the original video is input into the object detection model. The feature information of the first object is a series of data that can describe the unique attributes of the object. Different types of objects have different feature information. The object detection model is a computer vision algorithm based on deep learning. By comparing and matching the feature information of the first object with the content in the original video, the video clip containing the first object in the original video is identified. During the identification process, the time point when the first object first appears in the original video and the time point when the object last appears in the original video are recorded. If the object appears discontinuously in the video, then multiple such time periods will be recorded. These time period information is the appearance time information of the first object.

[0074] Optionally, the feature information refers to visual, semantic, or structured data used to uniquely or significantly identify the first object, which is extracted by the object detection model and used to match and identify the appearance time information of the first object in the original video. Different types of first objects have different features, and the object detection model judges whether the first object exists in the original video by learning and analyzing these features. For human objects, the feature information may include appearance features (such as facial contours, hairstyles, skin colors, etc.), dressing features (such as clothing colors, styles, accessories, etc.); for object objects, the feature information may include shapes, colors, textures, etc.; for event objects, the feature information may include specific action patterns, combinations of scene elements, etc.

[0075] Optionally, the video segments corresponding to the object are not stored in the database, possibly because the original video was newly uploaded and there has not been enough time to analyze each object in it and generate corresponding video segments; or there are problems such as data loss during the database update process.

[0076] Step S102, determine the target sub-segment corresponding to the appearance time information from the original video;

[0077] In a feasible embodiment, the appearance time information of each object is obtained through the object detection model, and the corresponding target sub-segment is intercepted from the original video based on the appearance time information. The target sub-segment contains the complete actions and scenes of the first object.

[0078] Optionally, the target sub-segment refers to multiple video segments extracted from the original video according to the appearance time information (such as start and end time points) of the first object, containing the complete actions or scenes of the first object. In the original video, the first object may appear multiple times, and the time periods of each appearance are different. After obtaining the appearance time information of the first object through the object detection model, video segments containing the first object can be divided from the original video according to these time information, and these video segments are the target sub-segments. Each target sub-segment has its own independent start time and end time, and they jointly cover all the appearance periods of the first object in the original video.

[0079] Step S103, generate the first video segment according to the target sub-segment.

[0080] In a feasible embodiment, it is a process of combining multiple target sub-segments together according to certain rules to form a complete first video segment. After obtaining multiple target sub-segments, in order to facilitate users to watch and obtain complete information about the first object, the target sub-segments need to be combined. The integration rules can be set according to actual needs, and can be integrated in the time order of the target sub-segments in the original video.

[0081] In a feasible embodiment, step S103 of generating the first video segment according to the target sub-segment further includes:

[0082] Step A10, obtaining the voiceprint information of the first object;

[0083] In a feasible embodiment, through audio separation technology, the audio in the target sub-segment is extracted separately, and the extracted audio is analyzed by a voiceprint recognition algorithm to extract the biometric features contained therein, thereby obtaining the voiceprint information of the first object.

[0084] Optionally, the voiceprint information refers to the voice feature data that can uniquely or significantly identify the first object, and usually includes biometric features such as timbre, intonation, formant, and speech rate.

[0085] Step A20, adjusting the target sub-segment according to the voiceprint information;

[0086] In a feasible embodiment, the target sub-segment is optimized or modified in combination with the voiceprint characteristics of the first object. Specifically, the adjustment can be to ensure that the audio of the target sub-segment contains all the speech of the first object (for example, when the first object is not in the current video playback screen, retain the video segment corresponding to the voiceprint information of the first object); improve the clarity of the first object's speech (such as noise reduction, volume balance); correct the problem of audio-visual out-of-sync (such as the delay caused by editing). Further, the audio of the target sub-segment can be optimized according to the characteristics such as timbre and intonation in the voiceprint information. For example, if the voiceprint information shows that the voice of the first object is relatively clear, then the clarity of the high-frequency part of the audio can be enhanced to make the voice brighter; if the speech rate is fast, the playback speed of the audio can be appropriately adjusted to make it more understandable. The voiceprint information can also provide a reference for the adjustment of the video picture. For example, according to the emotional style reflected by the voiceprint information, adjust parameters such as the color, brightness, and contrast of the video. If the voiceprint information conveys a happy emotion, the saturation and brightness of the video picture can be increased to create a more pleasant visual atmosphere.

[0087] Step A30, generating the first video segment according to the adjusted target sub-segment.

[0088] In a feasible embodiment, the audio and video parts of the adjusted target sub-segment are synchronized and synthesized. Ensure that the time axes of the audio and video are consistent, and the picture and sound can match. Select appropriate video coding formats and parameters according to needs, and perform coding processing on the synthesized video to generate the first video segment.

[0089] This embodiment provides a basis for subsequent adjustments by extracting the voiceprint information of the first object from the target sub-segment, and the uniqueness of the voiceprint information ensures the targeted processing. In the stage of adjusting the target sub-segment, operations such as filtering irrelevant background sounds, improving voice clarity, and correcting audio and video asynchrony can significantly improve the playback quality of the first video segment. By processing the target sub-segment based on the voiceprint information, the integrity and accuracy of the target sub-segment extraction can be enhanced. When the first video segment is finally generated, the audio and video are synchronously synthesized and encoded to ensure the matching of the picture and the sound and maintain the playback quality of the first video segment, thereby improving the usability and viewing experience of the video and the user experience.

[0090] In another feasible embodiment, step S103 of generating the first video segment according to the target sub-segment further includes:

[0091] Step B10, determining adjacent first sub-segments and second sub-segments according to the time interval between adjacent sub-segments in the target sub-segment, wherein the time interval between the first sub-segment and the second sub-segment is greater than a preset interval threshold;

[0092] In a feasible embodiment, by setting a preset interval threshold (such as 10 seconds, 30 seconds), the terminal traverses all extracted target sub-segments and detects the time interval between two adjacent sub-segments. If the time interval exceeds the preset interval threshold, it is considered that there is an obvious time jump or content fault between the two, and the two adjacent sub-segments are the first sub-segment and the second sub-segment.

[0093] Optionally, the time interval refers to the difference in duration between two adjacent sub-segments in the video segment on the original video time axis, specifically, the time span between the end time of the first sub-segment and the start time of the second sub-segment.

[0094] Optionally, the first sub-segment and the second sub-segment refer to two adjacent sub-segments in the first video segment extracted from the original video, the first sub-segment is in front, the second sub-segment is in the back, and the time interval between them is greater than a preset interval threshold. In the first video segment, due to the discontinuity of the first object, multiple sub-segments are formed. When the time interval between adjacent sub-segments is too large, they are defined as the first sub-segment and the second sub-segment respectively. For example, in a football game video, the first object is a player, the preset interval threshold is 8 minutes, and the player has wonderful performances in the 10th to 15th minutes and the 25th to 30th minutes. The sub-segments corresponding to these two time periods are the first sub-segment and the second sub-segment respectively.

[0095] Optionally, the preset interval threshold refers to a preset time interval standard used to determine whether the time interval between adjacent sub-segments is too large. When the time interval between adjacent sub-segments exceeds the preset interval threshold, the sub-segments need to be transitioned.

[0096] Step B20, generating a transition video segment based on a video segment between the first sub-segment and the second sub-segment in the original video;

[0097] In a feasible embodiment, a video segment between the first sub-segment and the second sub-segment may be extracted by specific technical means to make it play a transitional role in terms of content and visual effects, making the connection between the first sub-segment and the second sub-segment more natural and smooth.

[0098] Optionally, the transition video segment refers to a video segment between the first sub-segment and the second sub-segment in the original video, which is generated by a specific technology or method (such as accelerated playback, key frame extraction, motion blur or fade-in and fade-out, etc.) to smoothly connect the two sub-segments. This can make the switching between two adjacent sub-segments more natural and coherent, and avoid abrupt direct switching.

[0099] Step B30: Generate a first video segment according to the transition video segment and the target sub-segment.

[0100] In a feasible embodiment, by dynamically fusing the transition video segment and the target sub-segment, the first sub-segment, the transition video segment and the second sub-segment are spliced in the time sequence in the original video to generate the first video segment.

[0101] This embodiment determines the first sub-segment and the second sub-segment according to the time interval between adjacent sub-segments, and accurately identifies the part of the video with a large time span and unnatural transition. When the time interval between adjacent sub-segments is greater than the preset interval threshold, the audience may feel abrupt and incoherent during the viewing process. However, by generating a transition video segment based on the video segment between the first sub-segment and the second sub-segment in the original video, and inserting the transition video segment between the first sub-segment and the second sub-segment, the adjacent sub-segments can be smoothly connected, making the video playback more smooth and natural, enhancing the viewing and coherence of the video, and improving the user's viewing experience.

[0102] In a feasible embodiment, during the process of generating the first video clip, the target sub-clip is refined based on voiceprint recognition technology. By obtaining the voiceprint information of the first object, intelligent noise reduction, volume equalization, and speech rate adjustment are performed on the audio in the target sub-clip to ensure clarity and coherence. At the same time, when the voiceprint information of the first object exists in the video clip that does not contain the first object, a corresponding video clip containing the voiceprint information of the first object is added to the first video clip to make the first video clip more complete. For adjacent sub-clips with time jumps (such as when the time interval between the first sub-clip and the second sub-clip exceeds the preset interval threshold), the intermediate video clip is extracted from the original video, and through specific technical means and combined with the audio processed based on the voiceprint information, a spatio-temporally coherent transition video clip is formed. Finally, the target sub-clip with optimized voiceprint information is fused with the intelligently generated transition video clip to generate the first video clip. This not only solves the problem of audio-visual disconnection in traditional editing but also eliminates the visual discontinuity caused by time jumps, making the generated first video clip have both playback quality and natural fluency.

[0103] In this embodiment, when there is no first video clip in the database, the appearance time information of the first object is obtained through the object detection model and the target sub-clip is determined, accurately positioning the video clip containing the first object in the original video, providing a basis for subsequent generation of the first video clip, avoiding manual search by users, and saving time and effort. During the process of generating the first video clip, extracting the voiceprint information of the first object and adjusting the target sub-clip accordingly can improve the video playback quality and optimize the user's auditory experience. Further, generating a transition video clip according to the time interval between adjacent sub-clips solves the problem of unnatural transition of video content between sub-clips when the time interval between adjacent sub-clips is too large, making the video playback more smooth and coherent, avoiding the audience being affected by large jumps when watching the video, and improving the user's viewing experience.

[0104] Further, based on any of the above embodiments of the present application, a third embodiment of the video clip playback method is proposed. The video clip playback method further includes:

[0105] Step S01, when the currently played original video is paused, obtain the displayed first current frame;

[0106] In a feasible embodiment, in the case where the current video playback is paused, at the moment when the picture freezes, the first frame of the picture displayed on the screen at this time is extracted, which is the obtained displayed first current frame.

[0107] Optionally, the first current frame refers to the frame image where the video picture freezes when the original video is paused. For example, when playing an animation and pausing, the picture displayed on the screen at this time is the first current frame.

[0108] Step S02: Identify the current objects included in the first current frame and generate a list of information about the current objects.

[0109] In a feasible embodiment, after obtaining the first current frame displayed when the video is paused, analyze the frame to identify each object present in it, and then organize the relevant information of each object into a list.

[0110] Optionally, the current objects refer to various entity objects included in the first current frame. They can be people, objects, scenes, etc. For example, in a certain frame of a movie, the cars, pedestrians, etc. in the picture can all belong to the current objects.

[0111] Optionally, the information list refers to a list formed by organizing the relevant information of the current objects. The information can include the name, attributes, relevant introductions, etc. of the objects.

[0112] Step S03: Display the list of information about the current objects.

[0113] In a feasible embodiment, display the generated information list on the user's terminal screen. Enable the user to intuitively see the relevant information of the current objects. For example, display the information list in the sidebar of the video playback interface.

[0114] In this embodiment, by obtaining the first current frame when the original video is paused, the paused frame in the first video segment is accurately captured. By identifying the current objects in the first current frame and generating an information list, an information list for the user to select from is provided. The user can, according to their own interests, select the objects they want to watch and the corresponding video segments from this information list to further understand the relevant content.

[0115] Furthermore, based on any of the above embodiments of the present application, a fourth embodiment of the video segment playback method is proposed. After the step of playing the first video segment in step S20, it further includes:

[0116] Step S201: In response to a playback instruction for a second object in the original video, obtain the current time point corresponding to the displayed second current frame in the original video.

[0117] In a feasible embodiment, the terminal first responds to the user's play instruction for the second object (such as clicking a button or a voice command), then obtains the second current frame displayed on the current screen, and determines the current time point corresponding to this frame on the original video timeline, ensuring that the second object video segment retrieved from the database can be consistent with the user's current viewing progress. For example, in a football game video, if the user is watching a segment of player A and requests to switch to player B, the terminal will record the time position of the current screen in the complete game, and subsequently only extract the action segments of player B at and after this time point, avoiding time confusion or content repetition, thereby achieving a smooth object focus switch and a coherent viewing experience.

[0118] Optionally, the second object refers to another object in the original video that is different from the first object and of interest to the user. For example, during the playback of the first video segment, if the user has a viewing requirement for another object that appears, then the other object is the second object.

[0119] Optionally, the second current frame refers to the current screen displayed by the terminal when the user selects the second object during the playback of the first video segment.

[0120] Optionally, the current time point refers to the time position corresponding to the second current frame in the original video.

[0121] Step S202, extract the second video segment corresponding to the second object from the database;

[0122] Optionally, the second video segment refers to the set of all video segments related to the second object extracted from the database.

[0123] Step S203, determine the third video segment from the second video segment according to the current time point, where the start moment of the third video segment in the original video is later than or equal to the current time point.

[0124] In a feasible embodiment, based on the time point when the user triggers the play instruction for the second object, the video segments corresponding to the second object that meet the time continuity requirement are screened out from the preset database. The terminal first obtains the current time point of the playing video when the user operates (such as the 35th minute of the game), and then screens out all the video segments from the second video segment whose start time is later than or equal to the current time point (such as the shooting segments of player B at the 36th minute and the 40th minute), and combines the eligible video segments into the third video segment. This makes the switched video content consistent with the user's current viewing progress in terms of the timeline, avoiding reverting to the video content that has been played.

[0125] Optionally, the third video segment refers to the video segment screened out from the second video segment whose start moment in the original video is later than or equal to the current time point.

[0126] Exemplarily, refer to Figure 3 , the upper sub - figure shows that there is a video playing interface 101 on the terminal 100. On the video playing interface 101, there are a video name bar 111, a "Watch Only Him" button 112, a video operation button bar 113, and a video progress bar 114. A first object 117 is displayed on the video playing interface 101, and there is a current playing progress 116 on the video progress bar 114. Click the "Watch Only Him" button 112 to enter the middle sub - figure. An object list window 115 pops up on the video playing interface 101. The object list window 115 lists and displays each object in the current interface. When the user clicks on the list, the terminal can play the video segment corresponding to the object, and then jumps to the lower sub - figure. The video segment of the second object 118 is displayed on the video playing interface 101, and the video progress bar 114 shows the playing progress of the video segment of the second object 118.

[0127] In this embodiment, by allowing the user to switch to the second object when playing the first video segment corresponding to the first object, the diverse viewing needs of the user are met. By obtaining the current time point, it is ensured that the video segment of the second object extracted from the database is consistent with the current viewing progress, avoiding time confusion and content repetition. Screening out the third video segment ensures the consistency of the video content timeline after switching, and will not roll back to the played content, realizing a smooth object focus switch and a coherent viewing experience, improving the fluency and satisfaction of the user when watching the video.

[0128] Further, based on any of the above - mentioned embodiments of the present application, a fifth embodiment of the video segment playing method is proposed. The video segment playing method further includes:

[0129] Step C10, obtain the playing parameters during the playing process of the first video segment;

[0130] In a feasible embodiment, relevant data generated during the playing process is collected when playing the first video segment corresponding to the first object.

[0131] Optionally, the playing parameters refer to the key technical indicators that affect the picture quality, fluency, and resource consumption during the video playing process, which can be resolution, frame rate, bit rate, buffering time, network bandwidth, etc.

[0132] Step C20, optimize the playing process based on the playing parameters.

[0133] In a feasible embodiment, playback parameters during the playback of the first video segment are obtained. These playback parameters are key indicators affecting video picture quality, smoothness, and resource consumption, including resolution, frame rate, bit rate, buffering time, network bandwidth, etc. Then, based on the obtained playback parameters, the playback process is optimized, and the playback strategy or resource allocation is dynamically adjusted according to the playback parameters to improve the overall effect of video playback, ensure the picture quality and smoothness of playback, and reasonably control resource consumption.

[0134] Optionally, optimization refers to dynamically adjusting the playback strategy or resource allocation based on the playback parameters.

[0135] In a feasible embodiment, according to user feedback and viewing data analysis, the "only watch him" function of the video is continuously optimized. More video files supporting the "only watch him" function are added, including video content such as education and training, and sports events in addition to movies and variety shows. The accuracy of person recognition is improved, the number of selectable persons is increased, and the playback smoothness is enhanced to meet the changing needs of users.

[0136] In this embodiment, by obtaining the playback parameters and optimizing the playback process based on them, the playback strategy and resource allocation can be dynamically adjusted according to key indicators such as resolution and frame rate, which can improve the overall effect of video playback, ensure the picture quality and smoothness, and reasonably control resource consumption to avoid unnecessary resource waste. Further, continuously optimizing the "only watch him" function of the video, increasing the types of video files supporting this function, improving the accuracy of person recognition, increasing the number of selectable persons, and enhancing the playback smoothness can better meet the changing needs of users and enhance the user viewing experience.

[0137] In a feasible embodiment, the original video can be a teaching video, the first object can be a knowledge point, and the video segment playback method can identify the knowledge point boundaries by parsing the original video subtitles or converting speech to text content; at the same time, detect PPT page turns, key frames of the blackboard writing, or teacher gestures to assist in segmentation, and store the video segments corresponding to each knowledge point in a preset database. When the user inputs a target knowledge point or makes a selection in the information list, the terminal quickly locates the target sub-segment corresponding to the target knowledge point. After extracting the corresponding target sub-segment, it is integrated into a complete video segment for playback. If the user needs to expand learning, the terminal can also recommend knowledge points similar to the target knowledge point.

[0138] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the video segment playback method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0139] This application also provides a video segment playback device. Please refer to Figure 4 , the video segment playback device includes:

[0140] A response module 10, configured to extract, in response to a playback instruction for a first object in an original video, a first video segment corresponding to the first object in the original video from a preset database, where the first video segment is generated according to appearance time information of the first object in the original video obtained by an object detection model;

[0141] A playback module 20, configured to play the first video segment.

[0142] The video segment playback device provided in this application adopts the video segment playback method in the above embodiment, and can solve the technical problem of video segment playback. Compared with the prior art, the beneficial effects of the video segment playback device provided in this application are the same as those of the video segment playback method provided in the above embodiment, and other technical features in the video segment playback device are the same as the features disclosed in the above embodiment method, and will not be elaborated here.

[0143] This application provides a video segment playback device. The video segment playback device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video segment playback method in the first embodiment above.

[0144] Next, refer to Figure 5 , which shows a schematic structural diagram of a video segment playback device suitable for implementing the embodiments of this application. The video segment playback device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The shown video segment playback device is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of this application.

[0145] As Figure 5As shown, the video clip playing device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the xxx device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the video clip playing device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a video clip playing device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems can be implemented or had.

[0146] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0147] The video clip playing device provided by the present application adopts the video clip playing method in the above embodiment and can solve the technical problems of video clip playing. Compared with the prior art, the beneficial effects of the video clip playing device provided by the present application are the same as those of the video clip playing method provided by the above embodiment, and other technical features in the video clip playing device are the same as the features disclosed in the previous embodiment method, which will not be elaborated here.

[0148] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0149] As described above, it is only a specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.

[0150] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the video segment playing method in the above embodiments.

[0151] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0152] The above computer-readable storage medium may be included in the video segment playing device; or it may exist alone without being assembled into the video segment playing device.

[0153] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the video segment playing device, the video segment playing device can, in response to a playing instruction for a first object in the original video, extract a first video segment corresponding to the first object in the original video from a preset database, where the first video segment is generated according to the appearance time information of the first object in the original video obtained by an object detection model; and play the first video segment.

[0154] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0157] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned video segment playback method, and can solve the technical problems of video segment playback. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the video segment playback method provided in the above embodiments, and will not be elaborated here.

[0158] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the video clip playing method as described above.

[0159] The computer program product provided by the present application can solve the technical problem of video clip playing. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the video clip playing method provided by the foregoing embodiments, and will not be elaborated herein.

[0160] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields shall be included within the patent protection scope of the present application.

Claims

1. A method for playing a video clip, characterized in that The video clip playing method includes: In response to a playing instruction for a first object in an original video, extracting a first video clip corresponding to the first object in the original video from a preset database, where the first video clip is generated according to appearance time information of the first object in the original video obtained by an object detection model; Playing the first video clip.

2. The video clip playing method according to claim 1, wherein After the step of responding to the playing instruction for the first object in the original video, it further includes: In a case where the first video clip is not in the database, inputting feature information of the first object in the original video into the object detection model to obtain the appearance time information of the first object; Determining a target sub-clip corresponding to the appearance time information from the original video; Generating the first video clip according to the target sub-clip.

3. The video clip playing method according to claim 2, wherein The step of generating the first video clip according to the target sub-clip includes: Obtaining voiceprint information of the first object; Adjusting the target sub-clip according to the voiceprint information; Generating the first video clip according to the adjusted target sub-clip.

4. The video clip playing method according to claim 2, wherein The step of generating the first video clip according to the target sub-clip includes: Determining an adjacent first sub-clip and a second sub-clip according to a time interval between adjacent sub-clips in the target sub-clip, where a time interval between the first sub-clip and the second sub-clip is greater than a preset interval threshold; Generating a transition video clip based on a video clip between the first sub-clip and the second sub-clip in the original video; Generating the first video clip according to the transition video clip and the target sub-clip.

5. The video clip playing method according to claim 1, characterized in that, The video clip playing method further includes: When the currently played original video is paused, obtaining a first current frame displayed; Identifying a current object included in the first current frame and generating an information list of the current object; Displaying the information list of the current object.

6. The video clip playing method according to claim 1, wherein After the step of playing the first video clip, it further includes: In response to a playing instruction for a second object in the original video, obtaining a current time point corresponding to a second current frame displayed in the original video; Extracting a second video clip corresponding to the second object from the database; Determining a third video clip from the second video clip according to the current time point, where a start moment of the third video clip in the original video is later than or equal to the current time point.

7. The video clip playing method according to claim 1, characterized in that After the step of playing the first video clip, it further includes: Obtaining playing parameters during the playing process of the first video clip; Optimizing the playing process based on the playing parameters.

8. A video clip playing device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the video clip playing method according to any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the video segment playing method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that, The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the video segment playing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Video stream processing method and device and electronic equipment

    CN122093619A