Method, apparatus, storage medium, and electronic device for determining key frames
By converting audio data to text data, video keyframes are determined using natural language processing technology, which solves the problem of manual annotation dependence in the existing technology, and realizes fast and accurate automatic extraction of video keyframes.
Patent Information
- Application Number
- CN202210096192.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The method of extracting video keyframes in the prior art lacks the utilization of information of different modalities, requires a large amount of manual annotation, which is costly and inefficient.
By extracting video audio data, converting it into text data, using natural language processing technology to determine keywords, and determining video clips and frames based on keyword time information, and automatically determining video keyframes.
Without manual annotation, the video keyframes are quickly and accurately determined, improving efficiency and finely distinguishing different scenarios, avoiding the omission of important information and redundant processing.
Smart Images

Figure CN114429606B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, storage medium, and electronic device for determining a key frame. Background Art
[0002] With the continuous development of information technology, video has become an indispensable medium in modern life. For long videos, summaries can be created to facilitate user browsing and searching. Extracting the most representative key frames from a video as a summary is an intuitive and effective method. Therefore, how to extract key frames from a video has become a pressing issue. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a method, device, storage medium and electronic device for determining a key frame.
[0004] According to a first aspect of an embodiment of the present disclosure, a method for determining a key frame is provided, the method comprising:
[0005] Extracting audio data from the video to be determined;
[0006] Determining text data corresponding to the audio data;
[0007] determining a plurality of keywords from the text data;
[0008] Determining multiple keywords corresponding to the video to be determined;
[0009] Determining a video segment corresponding to each of the keywords from the video to be determined;
[0010] A target frame in each of the video segments is determined, and a plurality of the target frames are used as key frames corresponding to the video to be determined.
[0011] In some embodiments, determining a plurality of keywords from the text data includes:
[0012] Performing sentence processing on the text data to obtain a plurality of sub-text data corresponding to the text data;
[0013] At least one keyword corresponding to each subtext data is determined.
[0014] In some embodiments, before determining the video segment corresponding to each keyword from the video to be determined, the method further includes:
[0015] For each of the subtext data, determining a target keyword from at least one of the keywords corresponding to the subtext data, to obtain a plurality of target keywords corresponding to the to-be-determined video;
[0016] Determining a video segment corresponding to each keyword from the to-be-determined video includes:
[0017] From the to-be-determined videos, a video segment corresponding to each of the target keywords is determined.
[0018] In some embodiments, determining a target keyword from at least one keyword corresponding to the subtext data includes:
[0019] For each of the keywords, if the keyword is the same as a pending keyword, the keyword is deleted from at least one of the keywords to obtain a target keyword corresponding to the subtext data, and the pending keyword is a keyword adjacent to the keyword before the keyword.
[0020] In some embodiments, determining a video segment corresponding to each keyword from the to-be-determined video includes:
[0021] Determining time information corresponding to each of the keywords;
[0022] For each keyword, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0023] In some embodiments, the time information includes the start time and the end time of the keyword in the video to be determined; before determining the video segment corresponding to the keyword according to the time information corresponding to the keyword, the method further includes:
[0024] Determining a time period corresponding to the keyword according to the start time and the end time;
[0025] The determining, according to the time information corresponding to the keyword, the video segment corresponding to the keyword includes:
[0026] When it is determined that the time period is greater than or equal to the preset time period threshold, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0027] In some embodiments, the method further comprises:
[0028] If it is determined that the time period is less than the preset time period threshold, determining a first preset time period and a second preset time period according to a difference between the time period and the preset time period threshold;
[0029] Determining a target starting time according to the starting time and the first preset time period;
[0030] Determining a target end time according to the end time and the second preset time period;
[0031] The determining, according to the time information corresponding to the keyword, the video segment corresponding to the keyword includes:
[0032] According to the target start time and the target end time, a video segment corresponding to the keyword is determined.
[0033] According to a second aspect of an embodiment of the present disclosure, a device for determining a key frame is provided, the device comprising:
[0034] An audio data extraction module is configured to extract audio data from the video to be determined;
[0035] a text data determination module, configured to determine text data corresponding to the audio data;
[0036] a keyword determination module, configured to determine a plurality of keywords from the text data;
[0037] a video segment determination module, configured to determine a video segment corresponding to each of the keywords from the video to be determined;
[0038] The key frame determination module is configured to determine a target frame in each of the video segments and use the multiple target frames as key frames corresponding to the video to be determined.
[0039] In some embodiments, the keyword determination module is further configured to:
[0040] Performing sentence processing on the text data to obtain a plurality of sub-text data corresponding to the text data;
[0041] At least one keyword corresponding to each subtext data is determined.
[0042] In some embodiments, the apparatus further comprises:
[0043] a target keyword determination module configured to determine, for each subtext data, a target keyword from at least one keyword corresponding to the subtext data, so as to obtain a plurality of target keywords corresponding to the to-be-determined video;
[0044] The video segment determination module is further configured to:
[0045] From the to-be-determined videos, a video segment corresponding to each of the target keywords is determined.
[0046] In some embodiments, the target keyword determination module is further configured to:
[0047] For each of the keywords, if the keyword is the same as a pending keyword, the keyword is deleted from at least one of the keywords to obtain a target keyword corresponding to the subtext data, and the pending keyword is a keyword adjacent to the keyword before the keyword.
[0048] In some embodiments, the video segment determination module is further configured to:
[0049] Determining time information corresponding to each of the keywords;
[0050] For each keyword, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0051] In some embodiments, the time information includes a start time and an end time of the keyword in the video to be determined; the apparatus further includes:
[0052] a time period determination module, configured to determine a time period corresponding to the keyword according to the start time and the end time;
[0053] The video segment determination module is further configured to:
[0054] When it is determined that the time period is greater than or equal to the preset time period threshold, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0055] In some embodiments, the apparatus further comprises:
[0056] a preset time period determination module, configured to, when it is determined that the time period is less than the preset time period threshold, determine a first preset time period and a second preset time period according to a difference between the time period and the preset time period threshold;
[0057] a target start time determination module, configured to determine the target start time according to the start time and the first preset time period;
[0058] a target termination time determination module, configured to determine the target termination time according to the termination time and the second preset time period;
[0059] The video segment determination module is further configured to:
[0060] According to the target start time and the target end time, a video segment corresponding to the keyword is determined.
[0061] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the method for determining key frames provided in the first aspect of the present disclosure are implemented.
[0062] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0063] a memory having a computer program stored thereon;
[0064] The processor is configured to execute the computer program in the memory to implement the steps of the method for determining key frames provided in the first aspect of the present disclosure.
[0065] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: extracting audio data from a video to be determined; determining text data corresponding to the audio data; determining multiple keywords from the text data; determining a video segment corresponding to each keyword from the video to be determined; determining a target frame in each video segment, and using the multiple target frames as key frames corresponding to the video to be determined. In other words, the present disclosure can determine the key frames corresponding to the video to be determined based on the multiple keywords corresponding to the video to be determined, thus eliminating the need for manual labeling and enabling more rapid and accurate identification of key frames in the video.
[0066] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0068] Figure 1 is a flowchart showing a method for determining a key frame according to an exemplary embodiment of the present disclosure;
[0069] Figure 2 is a flowchart illustrating another method for determining a key frame according to an exemplary embodiment of the present disclosure;
[0070] Figure 3 is a block diagram showing a device for determining a key frame according to an exemplary embodiment of the present disclosure;
[0071] Figure 4 is a block diagram of a second apparatus for determining a key frame according to an exemplary embodiment of the present disclosure;
[0072] Figure 5is a block diagram of a third apparatus for determining a key frame according to an exemplary embodiment of the present disclosure;
[0073] Figure 6 is a block diagram of a fourth apparatus for determining a key frame according to an exemplary embodiment of the present disclosure;
[0074] Figure 7 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0075] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0076] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0077] First, the application scenarios of the present disclosure are described. Currently, a purely visual method can be used to extract key frames in a video. For example, the image information in the video can be modeled, and a deep learning-based method can be used to extract image content features in the video by designing a deep neural network. Finally, training is performed based on the labeled data to predict which frame is a key frame. However, the inventors of the present disclosure have found that this method lacks the utilization of information from different modalities and often requires a large amount of manual annotation for assistance, which is costly.
[0078] In order to overcome the technical problems existing in the above-mentioned related technologies, the present disclosure provides a method, device, storage medium and electronic device for determining key frames, which can determine the key frames corresponding to the video to be determined based on multiple keywords corresponding to the video to be determined. In this way, the key frames in the video can be determined more quickly and accurately without manual labeling.
[0079] The present disclosure is described below with reference to specific embodiments.
[0080] Figure 1 This is a flowchart of a method for determining a key frame according to an exemplary embodiment of the present disclosure, which can be applied to electronic devices, which may include mobile devices such as smart phones, smart wearable devices, smart speakers, smart tablets, personal computers, etc., and may also include servers such as local servers or cloud servers. Figure 1As shown, the method may include:
[0081] S101: extracting audio data from a video to be determined.
[0082] The video to be determined may be a film or television drama, a sports event video, or a self-made video (long video, short video, etc.) launched on a video platform. The present disclosure does not limit the type of the video to be determined. The audio data may be the audio track information corresponding to the video to be determined.
[0083] In this step, the video to be determined can be obtained first, and then the audio track information in the video to be determined can be extracted by the existing method.
[0084] S102: Determine text data corresponding to the audio data.
[0085] In this step, after extracting the audio track information from the video to be determined, the audio track information can be converted into text data through the WaveNet algorithm in speech recognition (ASR).
[0086] S103: Determine multiple keywords from the text data.
[0087] In this step, after determining the text data corresponding to the audio data, the text data can be segmented to obtain multiple subtext data corresponding to the text data, and at least one keyword corresponding to each subtext data can be determined. For example, the text data can be segmented using existing methods, with each sentence in the text data treated as a scene to obtain multiple subtext data corresponding to the text data. Subsequently, the TextRank algorithm in NLP (Natural Language Processing) technology can be used to determine at least one keyword corresponding to each subtext data.
[0088] S104: Determine a video segment corresponding to each keyword from the video to be determined.
[0089] In this step, after determining multiple keywords corresponding to the video to be determined, the time information corresponding to each keyword can be determined. For each keyword, the video clip corresponding to the keyword can be determined according to the time information corresponding to the keyword.
[0090] S105: Determine target frames in each video segment, and use multiple target frames as key frames corresponding to the video to be determined.
[0091] In this step, after determining the video segment corresponding to each keyword, for each video segment, any frame in the video segment can be extracted as the target frame to obtain multiple target frames. Then, the multiple target frames are used as key frames corresponding to the video to be determined.
[0092] By adopting the above method, the key frames corresponding to the video to be determined can be determined based on multiple keywords corresponding to the video to be determined. In this way, the key frames in the video can be determined more quickly and accurately without manual labeling.
[0093] Figure 2 is a flowchart showing another method for determining a key frame according to an exemplary embodiment of the present disclosure. Figure 2 As shown, the method may include:
[0094] S201: Extract audio data from the video to be determined.
[0095] The video to be determined may be a film or television drama, a sports event video, or a self-made video (long video, short video, etc.) launched on a video platform. The present disclosure does not limit the type of the video to be determined. The audio data may be the audio track information corresponding to the video to be determined.
[0096] S202: Determine text data corresponding to the audio data.
[0097] S203: Determine multiple keywords from the text data.
[0098] In a possible implementation, after multiple keywords are determined from the text data, a keyword sequence corresponding to the video to be determined can be obtained according to the order of the multiple keywords in the text data.
[0099] S204: Determine the time information corresponding to each keyword.
[0100] The time information may include the start time and the end time of the keyword in the video to be determined.
[0101] In this step, after multiple keywords are determined from the text data, for each keyword, the start and end times of the keyword in the video to be determined can be determined based on the playback time of the audio data to which the keyword belongs. For example, for each keyword, the target subtext data to which the keyword belongs can be first determined based on the keyword sequence, and then the pending time information corresponding to the target subtext data can be determined. Then, the start and end times of the keyword in the video to be determined can be determined based on the pending time information.
[0102] S205 : For each keyword, determine the video clip corresponding to the keyword according to the time information corresponding to the keyword.
[0103] It should be noted that there may be consecutively repeated keywords among the multiple keywords corresponding to the video to be determined, and the similarity between the two consecutive target frames obtained based on these two consecutively repeated keywords is also relatively high, resulting in redundancy in the key frames corresponding to the video to be determined. In addition, in the keyword sequence corresponding to the video to be determined, there may also be keywords that are not strongly relevant to the theme of the video to be determined, such as keywords such as "how" and "or". If the key frames corresponding to the video to be determined are determined based on the keywords with weak relevance, the accuracy of the key frames corresponding to the video to be determined may be relatively low. Based on the above possible situations, for each sub-text data, the target keyword can be determined from at least one keyword corresponding to the sub-text data to obtain multiple target keywords corresponding to the video to be determined. Afterwards, the video clip corresponding to each target keyword is determined from the video to be determined.
[0104] In the case where there may be repeated keywords among the multiple keywords corresponding to the video to be determined, in a first possible implementation, for each keyword, if it is identical to the keyword to be determined, it is deleted from at least one keyword to obtain the target keyword corresponding to the subtext data. The keyword to be determined is the keyword that is adjacent to the keyword before it. For example, if the second and third keywords in the keyword sequence corresponding to the video to be determined are both "kitten", the third keyword can be deleted from the keyword sequence.
[0105] It should be noted that when the sub-text data includes multiple keywords, the multiple keywords can also be input into a pre-trained target keyword determination model to obtain the target keyword corresponding to the sub-text data output by the target keyword determination model, wherein the target keyword determination model can be trained by the model training method of the existing technology, which will not be repeated here.
[0106] In view of the situation that there may be keywords in the keyword sequence corresponding to the video to be determined that are not strongly relevant to the theme of the video to be determined, in one possible implementation method, the keyword sequence corresponding to the video to be determined and the video to be determined can be input into a pre-trained keyword weight acquisition model to obtain the keyword weight corresponding to each keyword in the keyword sequence output by the keyword weight acquisition model. After that, the associated keywords in the keyword sequence whose keyword weight is greater than or equal to the preset weight threshold can be obtained, and the target keyword sequence corresponding to the video to be determined can be obtained according to the order of the associated keywords in the text data. Among them, the keyword weight acquisition model can be trained by the model training method of the existing technology, which will not be repeated here.
[0107] In another possible implementation, a preset keyword vocabulary can be used to filter out keywords that are not strongly relevant to the theme of the video to be determined from the keyword sequence corresponding to the video to be determined, so as to obtain a target keyword sequence corresponding to the video to be determined, wherein the preset keyword vocabulary can be pre-set according to the type of the video to be determined.
[0108] In this step, after determining the start and end times of each keyword in the video to be determined, the time period corresponding to the keyword can be determined based on the start and end times. If the time period is determined to be greater than or equal to a preset time period threshold, the video segment corresponding to the keyword is determined based on the time information corresponding to the keyword. The preset time period threshold can be determined based on pre-testing, and for example, can be 1 second. If the time period is greater than or equal to the preset time period threshold, it indicates that the video segment corresponding to the keyword is long enough, and the video segment corresponding to the keyword can be directly determined based on the time information corresponding to the keyword.
[0109] If the time period is determined to be less than the preset time period threshold, it indicates that the video clip corresponding to the keyword is relatively short and cannot accurately reflect the scene corresponding to the keyword, resulting in a relatively low accuracy rate of key frames extracted from the video clip. In this case, the length of the video clip can be increased, and the first preset time period and the second preset time period can be determined based on the difference between the time period and the preset time period threshold.
[0110] For example, if it is determined that the time period is less than the preset time period threshold, the difference between the time period and the preset time period threshold can be obtained, and the first preset time period and the second preset time period are determined based on the difference. The first preset time period and the second preset time period can be the same, or the first preset time period and the second preset time period can be different. For example, if the difference is 200ms, the first preset time period and the second preset time period can both be set to 100ms, or the first preset time period can be set to 50ms, and the second preset time period can be set to 150ms. The present disclosure does not limit the setting method of the first preset time period and the second preset time period.
[0111] Furthermore, after determining the first preset time period and the second preset time period, the target start time can be determined based on the start time and the first preset time period, and the target end time can be determined based on the end time and the second preset time period. Thereafter, the video clip corresponding to the keyword can be determined according to the target start time and the target end time.
[0112] S206: Determine target frames in each video segment, and use multiple target frames as key frames corresponding to the video to be determined.
[0113] Using the above method, the key frames corresponding to the video to be determined can be determined based on the multiple keywords corresponding to the video to be determined. This eliminates the need for manual labeling and allows for more rapid and accurate determination of key frames in the video. Furthermore, the keywords can distinguish key frames from different scenes in the video to be determined, making the key frames of the video to be determined more refined, avoiding the omission of important information and the repeated processing of redundant information, and further improving the efficiency of key frame determination. Furthermore, after determining the multiple keywords corresponding to the video to be determined, a target keyword can be determined from the multiple keywords, and the key frames corresponding to the video to be determined can be determined based on the target keyword, making the determined key frames more accurate.
[0114] Figure 3 is a block diagram of a device for determining a key frame according to an exemplary embodiment of the present disclosure. Figure 3 As shown, the device may include:
[0115] The audio data extraction module 301 is configured to extract audio data from the video to be determined;
[0116] A text data determination module 302 is configured to determine text data corresponding to the audio data;
[0117] A keyword determination module 303 is configured to determine a plurality of keywords from the text data;
[0118] The video segment determination module 304 is configured to determine a video segment corresponding to each keyword from the video to be determined;
[0119] The key frame determination module 305 is configured to determine a target frame in each of the video segments, and use the target frames as key frames corresponding to the video to be determined.
[0120] In some embodiments, the keyword determination module 302 is further configured to:
[0121] Perform sentence processing on the text data to obtain multiple sub-text data corresponding to the text data;
[0122] At least one keyword corresponding to each subtext data is determined.
[0123] In some embodiments, Figure 4 is a block diagram of a second apparatus for determining a key frame according to an exemplary embodiment of the present disclosure. Figure 4 As shown, the device also includes:
[0124] The target keyword determination module 306 is configured to determine, for each sub-text data, a target keyword from at least one keyword corresponding to the sub-text data, to obtain a plurality of target keywords corresponding to the video to be determined;
[0125] The video segment determination module 303 is further configured to:
[0126] From the video to be determined, a video segment corresponding to each target keyword is determined.
[0127] In some embodiments, the target keyword determination module 305 is further configured to:
[0128] For each keyword, if the keyword is the same as the pending keyword, the keyword is deleted from at least one of the keywords to obtain the target keyword corresponding to the subtext data, and the pending keyword is the keyword adjacent to the keyword before the keyword.
[0129] In some embodiments, the video segment determination module 303 is further configured to:
[0130] Determine the time information corresponding to each keyword;
[0131] For each keyword, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0132] In some embodiments, the time information includes the start time and the end time of the keyword in the video to be determined; Figure 5 is a block diagram of a third apparatus for determining a key frame according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the device also includes:
[0133] A time period determination module 307 is configured to determine a time period corresponding to the keyword based on the start time and the end time;
[0134] The video segment determination module 303 is further configured to:
[0135] When it is determined that the time period is greater than or equal to the preset time period threshold, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
[0136] In some embodiments, Figure 6 is a block diagram of a fourth apparatus for determining a key frame according to an exemplary embodiment of the present disclosure. Figure 6 As shown, the device also includes:
[0137] The preset time period determination module 308 is configured to determine a first preset time period and a second preset time period according to a difference between the time period and the preset time period threshold when it is determined that the time period is less than the preset time period threshold;
[0138] The target start time determination module 309 is configured to determine the target start time according to the start time and the first preset time period;
[0139] a target termination time determination module 310, configured to determine a target termination time according to the termination time and the second preset time period;
[0140] The video segment determination module 303 is further configured to:
[0141] According to the target start time and the target end time, a video segment corresponding to the keyword is determined.
[0142] Through the above device, the key frames corresponding to the video to be determined can be determined based on multiple keywords corresponding to the video to be determined. In this way, the key frames in the video can be determined more quickly and accurately without manual labeling.
[0143] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0144] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the method for determining a key frame provided by the present disclosure are implemented.
[0145] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for executing the above-mentioned method for determining a key frame when the computer program is executed by the programmable device.
[0146] Figure 7 7 is a block diagram of an electronic device 700 according to an exemplary embodiment of the present disclosure. For example, the electronic device 700 can be provided as a server. Figure 7 The electronic device 700 includes a processing component 722, which further includes one or more processors, and a memory resource represented by a memory 732 for storing instructions executable by the processing component 722, such as an application. The application stored in the memory 732 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 722 is configured to execute the instructions to perform the above-mentioned method for determining key frames.
[0147] The electronic device 700 may further include a power supply component 726 configured to perform power management of the electronic device 700, a wired or wireless network interface 750 configured to connect the electronic device 700 to a network, and an input / output (I / O) interface 758. The electronic device 700 may operate based on an operating system stored in the memory 732, such as Windows Server 2000. TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or similar.
[0148] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0149] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for determining a key frame, characterized in that: The method comprises: Extracting audio data from the video to be determined; Determining text data corresponding to the audio data; determining a plurality of keywords from the text data; Determining a video segment corresponding to each of the keywords from the to-be-determined videos; Determining a target frame in each of the video segments, and using a plurality of the target frames as key frames corresponding to the video to be determined; Determining a plurality of keywords from the text data includes: Perform sentence processing on the text data, taking each sentence of the text data as a scene, and obtaining a plurality of sub-text data corresponding to the text data; determining at least one keyword corresponding to each subtext data; Determining a video segment corresponding to each keyword from the to-be-determined video includes: Determining time information corresponding to each of the keywords; For each keyword, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
2. The method according to claim 1, characterized in that Before determining the video segment corresponding to each keyword from the video to be determined, the method further includes: For each of the subtext data, determining a target keyword from at least one of the keywords corresponding to the subtext data, to obtain a plurality of target keywords corresponding to the to-be-determined video; Determining a video segment corresponding to each keyword from the to-be-determined video includes: From the to-be-determined videos, a video segment corresponding to each of the target keywords is determined.
3. The method according to claim 2, characterized in that The determining of a target keyword from at least one keyword corresponding to the subtext data includes: For each of the keywords, if the keyword is the same as a pending keyword, the keyword is deleted from at least one of the keywords to obtain a target keyword corresponding to the subtext data, and the pending keyword is a keyword adjacent to the keyword before the keyword.
4. The method according to claim 1, wherein The time information includes the start time and the end time of the keyword in the video to be determined; before determining the video segment corresponding to the keyword according to the time information corresponding to the keyword, the method further includes: Determining a time period corresponding to the keyword according to the start time and the end time; The determining, according to the time information corresponding to the keyword, the video segment corresponding to the keyword includes: When it is determined that the time period is greater than or equal to the preset time period threshold, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
5. The method according to claim 4, characterized in that The method further comprises: If it is determined that the time period is less than the preset time period threshold, determining a first preset time period and a second preset time period according to a difference between the time period and the preset time period threshold; Determining a target starting time according to the starting time and the first preset time period; Determining a target end time according to the end time and the second preset time period; The determining, according to the time information corresponding to the keyword, the video segment corresponding to the keyword includes: According to the target start time and the target end time, a video segment corresponding to the keyword is determined.
6. A device for determining a key frame, characterized in that: The device comprises: An audio data extraction module is configured to extract audio data from the video to be determined; a text data determination module, configured to determine text data corresponding to the audio data; a keyword determination module, configured to determine a plurality of keywords from the text data; a video segment determination module, configured to determine a video segment corresponding to each of the keywords from the video to be determined; a key frame determination module, configured to determine a target frame in each of the video segments, and use the target frames as key frames corresponding to the video to be determined; The keyword determination module is further configured to: Perform sentence processing on the text data, taking each sentence of the text data as a scene, and obtaining a plurality of sub-text data corresponding to the text data; determining at least one keyword corresponding to each subtext data; The video segment determination module is further configured to: Determining time information corresponding to each of the keywords; For each keyword, a video segment corresponding to the keyword is determined according to the time information corresponding to the keyword.
7. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video processing method and device, storage medium and electronic equipment
CN111767765A
Video content retrieval method and device, computer equipment and storage medium
CN112395420A
Key frame determination method and device, electronic equipment and readable storage medium
CN112784110A