Video brief introduction automatic generation method, system and device and storage medium

By constructing a popular keyword data set and image text generation model, comprehensively considering video subtitles and image information, a more accurate video introduction is generated, which solves the problem of low accuracy of video introduction in the existing technology and reduces manual editing costs.

CN120390127APending Publication Date: 2025-07-29UNICOM WOYUEDU TECH CULTURE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510418979.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The accuracy of the prior art is not high when generating video introductions, especially in long videos or daily update videos. Manual writing costs are high and algorithm generation is susceptible to video quality, language complexity or content diversity, and lacks contextual understanding of short video content.

Method used

By obtaining the subtitle text and image description text of the training video, building a hot keyword data set, filtering keywords and inputting a trained image text generation model, combining the image and text to generate a video introduction, comprehensively considering the key information implicit in the hot words and images.

Benefits of technology

It improves the accuracy of video introduction, reduces manual editing time, and the generated introduction is more in line with user needs, has a certain popularity and reflects the context of the video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390127A_ABST
    Figure CN120390127A_ABST
Patent Text Reader

Abstract

The invention discloses a video brief introduction automatic generation method, system and device and a storage medium, and the method comprises the steps: extracting a plurality of first keywords in a first subtitle text, and extracting a plurality of second keywords in a second subtitle text; constructing a popularity keyword data set based on the plurality of first keywords, and screening the plurality of second keywords based on the popularity keyword data set to obtain a plurality of first target keywords; screening the plurality of second images to obtain a screened second image; inputting the screened second images into a trained image text generation model to obtain a plurality of second image description texts; extracting a plurality of second target keywords in each second image description text; and generating a video brief introduction according to the plurality of first target keywords and the plurality of second target keywords. The accuracy of the generated video brief introduction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a method, system, device, and storage medium for automatically generating video introductions. Background Art

[0002] With the development of network technology, the video industry has been developing faster and faster. To attract users to watch videos, it is necessary to introduce the content of each video, that is, to construct a video introduction. The generation of video introductions can be done through manual editing or by automatically extracting key content from the video through algorithms to generate a concise text summary or representative segments to construct the video introduction.

[0003] If writing or editing video introductions manually requires additional time from the creators, especially for long videos or users with daily updates, the cost is relatively high, and the introductions may be affected by the subjective intentions of the creators (such as "clickbait" or over-simplification), resulting in the video introductions not matching the real content. To overcome the disadvantages of manually generating video introductions, existing researchers have automatically extracted key content from the video through algorithms, and automatically extracted content from key frames or video subtitles by relying on AI (such as NLP or computer vision, etc.) to generate video introductions. However, the algorithms may fail due to video quality, language complexity, or content diversity (such as dialects, technical terms), and existing related methods only extract information from the video content and lack an understanding of the context of short video content. Therefore, the accuracy of the generated video introductions is not high.

[0004] Therefore, using existing related methods to generate video introductions results in low accuracy of the generated video introductions. Summary of the Invention

[0005] The present application aims to propose a method, system, device, and storage medium for automatically generating video introductions, which can improve the accuracy of the generated video introductions.

[0006] In a first aspect, an embodiment of the present application provides a method for automatically generating a video introduction, the method comprising:

[0007] Obtaining a first subtitle text corresponding to a training video, multiple first images, and a first image description text corresponding to each of the first images, and obtaining a second subtitle text and multiple second images corresponding to a video to be generated;

[0008] Extracting multiple first keywords from the first subtitle text, and extracting multiple second keywords from the second subtitle text;

[0009] Constructing a popularity keyword dataset based on the multiple first keywords, and screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords;

[0010] screening the plurality of second images to obtain screened second images;

[0011] Inputting the filtered second image into a trained image-text generation model to obtain a plurality of second image description texts, wherein the trained image-text generation model is trained using the plurality of first images and the first image description texts;

[0012] extracting a plurality of second target keywords from each of the second image description texts;

[0013] A video introduction is generated according to the plurality of first target keywords and the plurality of second target keywords.

[0014] Compared with the prior art, the first aspect of the present application has the following beneficial effects:

[0015] The method obtains a first subtitle text corresponding to a training video, multiple first images, and a first image description text corresponding to each first image, and obtains a second subtitle text and multiple second images corresponding to a to-be-generated video; extracts multiple first keywords from the first subtitle text, and extracts multiple second keywords from the second subtitle text; constructs a hot keyword dataset based on the multiple first keywords, and based on the hot keyword dataset, screens multiple second keywords to obtain multiple first target keywords; screens multiple second images to obtain screened second images; inputs the screened second images into a trained image-text generation model to obtain multiple second image description texts, wherein the trained image-text generation model is trained by using multiple first images and first image description texts; extracts multiple second target keywords from each second image description text; and generates a video introduction based on the multiple first target keywords and the multiple second target keywords. In this way, based on the hot keyword data set, multiple second keywords are screened to obtain multiple first target keywords, and some keywords that are more popular with users can be obtained; by inputting the screened second image into the trained image-text generation model, multiple second image description texts are obtained, and then multiple second target keywords in each second image description text are extracted, that is, by understanding the context of the video content, the key information implicit in the video image can be further obtained, and finally, based on the multiple first target keywords and the multiple second target keywords, a video introduction is generated, which comprehensively considers the hot words and the key information implicit in the image, and can improve the accuracy of the generated video introduction.

[0016] In some implementations, constructing a hot keyword dataset based on the multiple first keywords includes:

[0017] Obtain the occurrence frequency of each of the first keywords and user interaction behavior information, where the user interaction behavior information is information on the user's behavioral interaction with videos containing the first keywords;

[0018] Calculate the keyword heat value of each of the first keywords according to the occurrence frequency of each of the first keywords and the user interaction behavior information;

[0019] Use the first keywords whose keyword heat values are greater than or equal to a first preset threshold as heat keywords, and construct a heat keyword dataset according to the heat keywords.

[0020] In some embodiments, based on the heat keyword dataset, screen a plurality of second keywords to obtain a plurality of first target keywords, including:

[0021] Vectorize each of the second keywords to obtain keyword vectors, and vectorize the second subtitle text to obtain text vectors;

[0022] Determine whether there are identical keywords for the second keywords in the heat keyword dataset;

[0023] If there are identical keywords for the second keywords in the heat keyword dataset, use the keyword heat value of the obtained identical keywords as the keyword heat value of the second keywords;

[0024] Calculate the correlation score between the keyword vector and the text vector according to the keyword heat value of the second keywords;

[0025] Screen a plurality of second keywords according to the correlation score to obtain a plurality of first target keywords.

[0026] In some embodiments, the calculating the correlation score between the keyword vector and the text vector according to the keyword heat value of the second keywords includes:

[0027]

[0028] where SC(d q ) represents the correlation score, H q∈A (q,t) represents the keyword heat value of the q-th second keyword at the current time t, w1 represents the first weight coefficient, D represents the text vector, d q represents the extracted keyword vector corresponding to the q-th second keyword, d p represents the keyword vector corresponding to the p-th second keyword that has been screened, R represents the set of second keywords that have been screened, sim(·) represents vector similarity calculation, and max(·) represents taking the maximum value.

[0029] In some embodiments, the image-text generation model includes an encoder, an attention mechanism, a decoder, and a discriminator, and the trained image-text generation model is obtained by training the image-text generation model using the multiple first images and the first image description text, including:

[0030] Constructing a first loss function, and constructing a second loss function;

[0031] extracting a plurality of third keywords from the first image description text;

[0032] Inputting the plurality of first images and the first image description text into an image-text generation model, and encoding each of the first images by the encoder to obtain an encoding result;

[0033] Processing the encoding result through the attention mechanism to obtain a weighted encoding result;

[0034] Decoding the weighted coding result by the decoder to obtain predicted image description text;

[0035] Minimizing the first loss function makes the predicted image description text approach the first image description text to obtain a target image description text;

[0036] The plurality of third keywords, the target image description text, and the first image description text are input into the discriminator, and the image-text generation model is trained by maximizing the second loss function.

[0037] In some embodiments, constructing the second loss function includes:

[0038]

[0039] Among them, LOSS D represents the second loss function, E represents the mean, B represents the first image, x represents the first image description text corresponding to the first image, M r represents the image description pairs matched in the image-text training set, D(·) represents the discriminator, w2 represents the second weight coefficient, Indicates the target image description text, M f represents a pair of description texts of the first image and the target image, s represents the number of third keywords contained in the description text of the target image, and z represents the total number of multiple third keywords.

[0040] In some implementations, generating a video introduction based on the plurality of first target keywords and the plurality of second target keywords includes:

[0041] Perform keyword screening on the multiple first target keywords and the multiple second target keywords to obtain multiple third target keywords;

[0042] Input the multiple third target keywords into a pre-trained large language model to generate a video summary.

[0043] In a second aspect, an embodiment of the present application further provides a video summary automatic generation system, and the system includes:

[0044] A data acquisition unit, configured to acquire a first subtitle text corresponding to a training video, multiple first images, and a first image description text corresponding to each of the first images, and to acquire a second subtitle text and multiple second images corresponding to the video to be generated;

[0045] A first keyword extraction unit, configured to extract multiple first keywords from the first subtitle text, and to extract multiple second keywords from the second subtitle text;

[0046] A keyword screening unit, configured to construct a popularity keyword dataset based on the multiple first keywords, and to screen the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords;

[0047] An image screening unit, configured to screen the multiple second images to obtain the screened second images;

[0048] A description text generation unit, configured to input the screened second images into a trained image text generation model to obtain multiple second image description texts, and the trained image text generation model is obtained by training the image text generation model with the multiple first images and the first image description texts;

[0049] A second keyword extraction unit, configured to extract multiple second target keywords from each of the second image description texts;

[0050] A video summary generation unit, configured to generate a video summary according to the multiple first target keywords and the multiple second target keywords.

[0051] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a video summary automatic generation method as described above.

[0052] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method for automatically generating a video introduction as described above.

[0053] It can be understood that the beneficial effects of the above-mentioned second to fourth aspects compared with the relevant technologies are the same as the beneficial effects of the above-mentioned first aspect compared with the relevant technologies. Please refer to the relevant description in the above-mentioned first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0055] Figure 1 This is a flowchart of an embodiment of the method for automatically generating a video introduction provided by the present application;

[0056] Figure 2 This is a structural diagram of an image-text generation model in the best embodiment of the method for automatically generating a video introduction provided by this application;

[0057] Figure 3 This is a schematic diagram of the structure of an embodiment of the video introduction automatic generation system provided by the present application;

[0058] Figure 4 It is a structural diagram of an embodiment of the electronic device provided by this application. DETAILED DESCRIPTION

[0059] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.

[0060] In the description of this application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0061] In the description of this application, it should be understood that descriptions involving orientation, such as the orientation or positional relationship indicated by up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.

[0062] In the description of this application, it should be noted that unless otherwise clearly defined, terms such as setting, installation, and connection should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.

[0063] With the development of network technology, the video industry is developing faster and faster. To attract users to watch videos, it is necessary to introduce the content of each video, that is, to construct a video summary. Video summary generation can be done through manual editing or by automatically extracting key content from the video through algorithms to generate a concise text summary or representative clips to construct the video summary.

[0064] If creating or editing a video summary manually requires additional time from the creator, especially for long videos or users with daily updates, the cost is relatively high, and the summary may be affected by the creator's subjective intentions (such as "clickbait" or over-simplification), resulting in the video summary not matching the real content. To overcome the shortcomings of manually generating video summaries, existing researchers automatically extract key content from videos through algorithms and rely on AI (such as NLP or computer vision, etc.) to automatically extract content from key frames or video subtitles to generate video summaries. However, the algorithms may fail due to video quality, language complexity, or content diversity (such as dialects, technical terms), and existing related methods only extract information from video content and lack an understanding of the context of short video content. Therefore, the accuracy of the generated video summaries is not high.

[0065] Therefore, using existing related methods to generate video summaries results in low accuracy of the generated video summaries.

[0066] To solve the problem of low accuracy of video summaries generated by existing related methods as described above, this application proposes a method, system, device, and storage medium for automatically generating video summaries.

[0067] Refer to Figure 1 , an embodiment of this application provides a method for automatically generating video summaries, and the method includes the following steps:

[0068] Step S100: Obtain the first subtitle text corresponding to the training video, multiple first images, and the first image description text corresponding to each first image, and obtain the second subtitle text and multiple second images corresponding to the video to be generated;

[0069] Step S200: Extract multiple first keywords from the first subtitle text, and extract multiple second keywords from the second subtitle text;

[0070] Step S300: Construct a popularity keyword dataset based on the multiple first keywords, and screen the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords;

[0071] Step S400: Screen multiple second images to obtain the screened second images;

[0072] Step S500: Input the screened second images into the trained image-text generation model to obtain multiple second image description texts, where the trained image-text generation model is trained by multiple first images and first image description texts;

[0073] Step S600: Extract multiple second target keywords from each second image description text;

[0074] Step S700: Generate a video summary based on the multiple first target keywords and multiple second target keywords.

[0075] In this embodiment, by obtaining the first subtitle text, multiple first images, and the first image description text corresponding to each first image of the training video, and by obtaining the second subtitle text and multiple second images corresponding to the video to be generated; extracting multiple first keywords from the first subtitle text, and extracting multiple second keywords from the second subtitle text; constructing a popularity keyword dataset based on the multiple first keywords, and screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords; screening the multiple second images to obtain the screened second images; inputting the screened second images into the trained image-text generation model to obtain multiple second image description texts, where the trained image-text generation model is trained by multiple first images and first image description texts; extracting multiple second target keywords from each second image description text; generating a video summary based on the multiple first target keywords and multiple second target keywords. In this way, screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords can obtain some keywords that are more concerned by users; by inputting the screened second images into the trained image-text generation model to obtain multiple second image description texts, and then extracting multiple second target keywords from each second image description text, that is, by understanding the context of the video content, the key information hidden in the video images can be further obtained. Finally, generating a video summary based on the multiple first target keywords and multiple second target keywords comprehensively considers the popularity words and the key information hidden in the images, and can improve the accuracy of the generated video summary.

[0076] The above-mentioned extraction of multiple first keywords in the first subtitle text and extraction of multiple second keywords in the second subtitle text can be performed by using the TF-IDF algorithm or the TextRank algorithm to extract keywords from the text, that is, extracting multiple first keywords from the first subtitle text and extracting multiple second keywords from the second subtitle text.

[0077] The above-mentioned screening of multiple second images to obtain the screened second image can be performed based on image features, or multiple second images can be screened using techniques known to those skilled in the art to obtain the screened second image, which is not specifically described or limited in this embodiment.

[0078] The above-mentioned generation of the video introduction based on the multiple first target keywords and the multiple second target keywords may be generating the video introduction based on the multiple first target keywords and the multiple second target keywords using a pre-trained large language model.

[0079] In some implementations, constructing a hot keyword dataset based on multiple first keywords includes:

[0080] Obtaining the frequency of occurrence of each first keyword and user interaction behavior information, where the user interaction behavior information is information about user interaction with the video containing the first keyword;

[0081] Calculate the keyword popularity value of each first keyword based on the occurrence frequency of each first keyword and user interaction behavior information;

[0082] A first keyword whose keyword heat value is greater than or equal to a first preset threshold is taken as a hot word, and a hot keyword data set is constructed based on the hot words.

[0083] In this embodiment, the frequency of occurrence and user interaction behavior information of each first keyword are obtained, where the user interaction behavior information is information about the user's interaction with the video containing the first keyword. Based on the frequency of occurrence and user interaction behavior information of each first keyword, the keyword popularity value of each first keyword is calculated. First keywords with keyword popularity values greater than or equal to a first preset threshold are regarded as hot words, and a hot keyword dataset is constructed based on the hot words. In this way, by constructing a hot keyword dataset, keywords with high keyword popularity values are collected, laying a good data foundation for subsequent keyword screening, thereby improving the accuracy of the generated video introduction.

[0084] In some implementations, based on the hot keyword dataset, multiple second keywords are screened to obtain multiple first target keywords, including:

[0085] Vectorizing each second keyword to obtain a keyword vector, and vectorizing the second subtitle text to obtain a text vector;

[0086] Determine whether the second keyword has the same keyword in the hot keyword dataset;

[0087] If the second keyword has the same keyword in the hot keyword dataset, the keyword hot value of the same keyword is used as the keyword hot value of the second keyword;

[0088] Calculate the relevance score between the keyword vector and the text vector based on the keyword heat value of the second keyword;

[0089] The plurality of second keywords are screened according to the relevance scores to obtain a plurality of first target keywords.

[0090] In this embodiment, each second keyword is vectorized to obtain a keyword vector, and the second subtitle text is vectorized to obtain a text vector; it is determined whether the second keyword has the same keyword in the hot keyword data set; if the second keyword has the same keyword in the hot keyword data set, the keyword heat value of the same keyword is obtained as the keyword heat value of the second keyword; based on the keyword heat value of the second keyword, the correlation score between the keyword vector and the text vector is calculated; based on the correlation score, multiple second keywords are screened to obtain multiple first target keywords. In this way, based on the hot keyword data set, multiple second keywords are screened to obtain multiple first target keywords that are both highly popular and relevant to the text, laying a good data foundation for the subsequent generation of accurate video introductions.

[0091] In some implementations, calculating a relevance score between the keyword vector and the text vector based on the keyword heat value of the second keyword includes:

[0092]

[0093] Among them, SC(d q ) represents the correlation score, H q∈A (q, t) represents the keyword heat value of the qth second keyword at the current time t, w1 represents the first weight coefficient, D represents the text vector, d q represents the extracted keyword vector corresponding to the qth second keyword, d q represents the keyword vector corresponding to the p-th second keyword that has been filtered, R represents the set of filtered second keywords, sim(·) represents vector similarity calculation, and max(·) represents taking the maximum value.

[0094] In some embodiments, the image-text generation model includes an encoder, an attention mechanism, a decoder, and a discriminator. The trained image-text generation model is obtained by training the image-text generation model with multiple first images and first image description texts, including:

[0095] Construct a first loss function and a second loss function;

[0096] Extract multiple third keywords from the first image description text;

[0097] Input the multiple first images and the first image description texts into the image-text generation model, and encode each first image through the encoder to obtain an encoding result;

[0098] Process the encoding result through the attention mechanism to obtain a weighted encoding result;

[0099] Decode the weighted encoding result through the decoder to obtain a predicted image description text;

[0100] Make the predicted image description text tend to the first image description text by minimizing the first loss function to obtain a target image description text;

[0101] Input the multiple third keywords, the target image description text, and the first image description text into the discriminator, and train the image-text generation model by maximizing the second loss function.

[0102] In this embodiment, by constructing a first loss function and a second loss function; extracting multiple third keywords from the first image description text; inputting the multiple first images and the first image description texts into the image-text generation model, encoding each first image through the encoder to obtain an encoding result; processing the encoding result through the attention mechanism to obtain a weighted encoding result; decoding the weighted encoding result through the decoder to obtain a predicted image description text; making the predicted image description text tend to the first image description text by minimizing the first loss function to obtain a target image description text; inputting the multiple third keywords, the target image description text, and the first image description text into the discriminator, and training the image-text generation model by maximizing the second loss function. In this way, through the constructed first loss function and second loss function, the generated target image description text is made as close as possible to the real description text, and the discriminator will continuously feedback the discrimination signal to the encoder, the attention mechanism, and the decoder, so that the text description generated by the image-text generation model becomes more and more accurate.

[0103] In some embodiments, constructing the second loss function includes:

[0104]

[0105] Among them, LOSS D represents the second loss function, E represents taking the mean, B represents the first image, x represents the first image description text corresponding to the first image, M r represents the image description pair matched in the image text training set, D(·) represents the discriminator, and w2 represents the second weight coefficient. represents the target image description text, M f represents the pair of the first image and the target image description text, s represents the number of the third keywords included in the target image description text, and z represents the total number of multiple third keywords.

[0106] In some embodiments, according to multiple first target keywords and multiple second target keywords, a video summary is generated, including:

[0107] Performing keyword screening on the multiple first target keywords and the multiple second target keywords to obtain multiple third target keywords;

[0108] Inputting the multiple third target keywords into a pre-trained large language model to generate a video summary.

[0109] In this embodiment, by performing keyword screening on the multiple first target keywords and the multiple second target keywords, multiple third target keywords are obtained; the multiple third target keywords are input into a pre-trained large language model to generate a video summary. In this way, by further screening the multiple first target keywords and the multiple second target keywords, redundant keywords and less relevant keywords can be removed, reducing the influence of redundant keywords and irrelevant keywords, so that the generated video summary is more accurate.

[0110] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments:

[0111] To solve the problem that the accuracy of the video summary generated by the existing related methods is not high, this embodiment proposes an automatic video summary generation method that comprehensively considers hot words and key information hidden in the image. This method makes the video summary more accurate and has a certain degree of popularity through multimodal analysis. The method of this embodiment specifically includes:

[0112] Step 1: First, obtain multiple training videos, and these multiple training videos can cover all types of videos. Then, use a subtitle extraction tool to extract the subtitles in each training video to obtain the first subtitle text corresponding to each training video. The subtitle extraction tool can be an OCR tool or a speech recognition tool. When there are subtitles in the training video, use the OCR tool, and when there are no subtitles in the training video, use the speech recognition tool. Obtain the video to be generated, and obtain the second subtitle text corresponding to the video to be generated in the same way.

[0113] Step 2: Obtain multiple first target keywords in the second subtitle text.

[0114] (1) Extract keywords from each first subtitle text to obtain multiple first keywords. Keyword extraction can adopt techniques well-known to those skilled in the art, and this embodiment will not be specifically described or limited.

[0115] (2) Construct a popularity keyword dataset based on the multiple first keywords. This popularity keyword dataset can be composed of popularity words related to each type of video.

[0116] Specifically, for emotional videos, there can be popularity keywords such as emotion, love, family affection, marriage, family, children, etc. These popularity keywords can be calculated by combining factors such as the frequency of the first keywords appearing in this type of video, time decay, and user interaction behaviors (including reading and clicking, forwarding and sharing, commenting, liking, and favoriting video behavior information), etc. The process of calculating the keyword popularity value is as follows:

[0117] H(i,t) = (α·f i + β·h i + γ·l i ) × e -λt

[0118] Among them, H(i,t) represents the keyword popularity value, α, β, and γ represent weight coefficients, α + β + γ = 1, f i represents the appearance frequency of the first keyword i, h i represents the number of times the first keyword i is read and clicked, l i represents the total number of times the first keyword i is forwarded and shared, commented, liked, and favorited. λ represents the time decay factor, and t represents time. The number of times the first keyword i is read and clicked, forwarded and shared, commented, liked, and favorited can be equal to the number of times the video corresponding to the first keyword i is read and clicked, forwarded and shared, commented, liked, and favorited.

[0119] Then, according to the keyword popularity value and the first preset threshold, select the keywords with high keyword popularity values as popularity words. That is, when the keyword popularity value is greater than or equal to the first preset threshold, the keyword corresponding to the keyword popularity value is used as a popularity word.

[0120] Finally, combine all the selected popularity words into a popularity keyword dataset.

[0121] (3) Perform keyword extraction on the second subtitle text to obtain multiple second keywords. Filter the multiple second keywords based on the hot keyword dataset to obtain multiple filtered second keywords. The multiple filtered second keywords are the multiple first target keywords in the obtained second subtitle text.

[0122] Specifically, any word vectorization model among existing models such as word2vec model, Skip-gram model and BERT model is selected to vectorize multiple second keywords to obtain keyword vectors; the same method is used to vectorize the second subtitle text to obtain a text vector.

[0123] Based on the hot keyword dataset, determine whether the extracted keyword has the same first keyword in the hot keyword dataset. If the same first keyword exists, the keyword heat value of the first keyword is used as the keyword heat value of the second keyword. Calculate the relevance score between the second keyword vector and the text vector. Based on the relevance score, filter multiple second keywords to obtain multiple filtered second keywords. The relevance score is calculated using the following formula:

[0124]

[0125] Among them, SC(d q ) represents the correlation score, H q∈A (q, t) represents the keyword heat value of the qth second keyword at the current time t. If q∈A, the keyword heat value of the qth second keyword is equal to the keyword heat value of the corresponding first keyword in the hot keyword dataset A. w1 represents the first weight coefficient, D represents the text vector, and d q represents the keyword vector corresponding to the qth second keyword, d p represents the keyword vector corresponding to the p-th filtered second keyword, and R represents the filtered second keyword set.

[0126] If the relevance score is greater than or equal to the second preset threshold, the second keyword corresponding to the relevance score is used as the filtered second keyword; if the relevance score is less than the second preset threshold, the second keyword corresponding to the relevance score is deleted.

[0127] Step 3. First, obtain multiple first images corresponding to each training video, extract frames from each training video, obtain multiple first images, and annotate the first object in each first image (the first object refers to a person or object in the image, etc.). Perform a text description based on the first object corresponding to each first image (which can describe the location information of each object, the positional relationship between objects, and background information, etc.), obtain the first image description text corresponding to each first image, and extract multiple third keywords from the first image description text, which include all first objects. The text description of this embodiment can be manually described, so that a more accurate image description text can be obtained. An image text training set is constructed using each first image and the corresponding first image description text.

[0128] Then, the image-text generation model is trained based on the image-text training set to obtain the trained image-text generation model. Figure 2 The image-text generation model includes an encoder, an attention mechanism, a decoder, and a discriminator. The encoder of the image-text generation model uses the ViT model (Vision Transformer model), and the decoder of the image-text generation model uses the Transformer model. The first loss function and the second loss function of the image-text generation model are constructed. The image-text generation model is trained based on the image text, the first loss function and the second loss function, and the training set to obtain a trained image-text generation model.

[0129] The first loss function is:

[0130]

[0131] Where L represents the first loss function, y represents the first image description text, Represents the predicted image description text, and N represents the number of predicted image description texts.

[0132] The second loss function is:

[0133]

[0134] Among them, LOSS D represents the second loss function, E represents the mean, B represents the first image, x represents the first image description text corresponding to the first image, M r represents the image description pairs matched in the image-text training set, D(·) represents the discriminator, w2 represents the second weight coefficient, Indicates the target image description text, M fDenote the first image and the target image description text pair, s represents the number of third keywords contained in the target image description text, and z represents the number of multiple third keywords. In this embodiment, the second loss function is improved by using the second weight coefficient, so that the target image description text contains as many third keywords as possible, making the target image description text closer to the first image description text, thereby enabling the image text generation model to generate more accurate image description text.

[0135] Specifically, the first image and the first image description text in the image text training set are input into the image text generation model, and the first image description text is used as the annotation of the first image. After the encoder encodes the first image, an encoded result is obtained; then the attention mechanism processes the encoded result to obtain a weighted encoded result; then the decoder decodes the weighted encoded result to obtain a predicted image description text. The first loss function (minimizing the first loss function) is used to make the predicted image description text continuously approach the true description text (i.e., the corresponding first image description text) to obtain the target image description text; multiple third keywords, the target image description text, and the first image description text are input into the discriminator, and the discriminator will feedback a discrimination signal to the encoder, the attention mechanism, and the decoder, that is, the discriminator and the encoder, the attention mechanism, and the decoder are jointly optimized. The discriminator is iteratively optimized based on the second loss function to train the image text generation model. By maximizing the second loss function, the target image description text is made as close as possible to the true description text, making the text description generated by the image text generation model more accurate.

[0136] Step 4: Obtain multiple second images corresponding to the video to be generated, and identify the second objects in each second image. Screen the multiple second images according to the identified second objects, and convert the screened second images into text information to obtain multiple second image description texts. Extract multiple second target keywords from each second image description text.

[0137] Specifically, a frame extraction tool is used to extract video frames from the video to be generated to obtain multiple second images corresponding to the video to be generated. An image object detection algorithm is used to identify the second objects in each second image, and the multiple second images are screened according to the identified second objects. Specifically:

[0138] Count the second objects in each second image, sort all the second objects according to the number of each second object, select multiple second objects ranked in the front, and obtain the images corresponding to the multiple second objects ranked in the front to obtain multiple initially screened second images;

[0139] The plurality of preliminary screening second images are sorted according to the number of second objects contained in each image, a preset number of images are selected from the sorted preliminary screening second images, and redundant images are deleted to obtain a screened second image; the method for deleting redundant images can use a CNN network model to extract features from the image to obtain feature vectors, and then calculate the similarity between the feature vectors, and remove images whose similarity is greater than a threshold.

[0140] The filtered second image is input into the trained image-text generation model and converted into text information corresponding to the filtered second image, thereby obtaining multiple second image description texts. The filtered second image is processed by the encoder, attention mechanism, and decoder in the trained image-text generation model to obtain multiple second image description texts.

[0141] A plurality of second target keywords in each second image description text may be extracted using natural language processing technology. For example, a TextCNN model or a Word2Vec model may be used to extract the second target keywords in each second image description text, which is not specifically limited or described in this embodiment.

[0142] Step 5: Generate a video introduction based on the multiple first target keywords in the second subtitle text obtained in step 2 and the multiple second target keywords obtained in step 4.

[0143] The multiple first target keywords in the second subtitle text obtained in step 2 and the multiple second target keywords obtained in step 4 are screened to obtain multiple third target keywords.

[0144] Specifically, keywords that appear in both the multiple first target keywords and the multiple second target keywords are retained as first retained keywords; the first retained keywords in the multiple first target keywords are eliminated to obtain first remaining keywords, and the first retained keywords in the multiple second target keywords are eliminated to obtain second remaining keywords; the first remaining keywords are judged for relevance with the second subtitle text, and keywords with high relevance are retained to obtain second retained keywords; the second remaining keywords are judged for relevance with the corresponding second image description text, and keywords with high relevance are retained to obtain third retained keywords; the first retained keywords, the second retained keywords, and the third retained keywords are merged to obtain multiple third target keywords.

[0145] Relevance judgment: Both keywords and texts are represented by vectors (such as the Word2Vec model, GloVe model, or BERT model), and then their cosine similarity is calculated. Keywords with high similarity between keywords and texts are retained.

[0146] Finally, input multiple third target keywords into the pre-trained large language model to automatically generate a video summary. It should be noted that the large language model can be a technology well-known to those skilled in the art, and this embodiment does not make specific descriptions and limitations.

[0147] In the later stage, users can modify the automatically generated video summary by themselves to make the video summary more in line with user needs. In the later stage, users modify based on the automatically generated video summary, which reduces the time for users to edit the video summary alone. Moreover, the automatically generated video summary comprehensively considers hot words and key information hidden in the images. Through multi-modal analysis, the video summary is more accurate and has a certain degree of popularity. In the later stage, through further polishing of the video summary by users, the efficiency and quality of video summary generation can be well balanced.

[0148] Refer to Figure 3 According to this, an embodiment of the present application also provides a video summary automatic generation system, which includes a data acquisition unit 100, a first keyword extraction unit 200, a keyword screening unit 300, an image screening unit 400, a description text generation unit 500, a second keyword extraction unit 600, and a video summary generation unit 700, where:

[0149] The data acquisition unit 100 is used to acquire the first subtitle text corresponding to the training video, multiple first images, and the first image description text corresponding to each first image, and to acquire the second subtitle text and multiple second images corresponding to the video to be generated;

[0150] The first keyword extraction unit 200 is used to extract multiple first keywords from the first subtitle text, and to extract multiple second keywords from the second subtitle text;

[0151] The keyword screening unit 300 is used to construct a hot keyword data set based on multiple first keywords, and to screen multiple second keywords based on the hot keyword data set to obtain multiple first target keywords;

[0152] The image screening unit 400 is used to screen multiple second images to obtain the screened second images;

[0153] The description text generation unit 500 is used to input the screened second images into the trained image text generation model to obtain multiple second image description texts, and the trained image text generation model is obtained by training the image text generation model with multiple first images and first image description texts;

[0154] The second keyword extraction unit 600 is used to extract multiple second target keywords from each second image description text;

[0155] The video introduction generating unit 700 is configured to generate a video introduction according to a plurality of first target keywords and a plurality of second target keywords.

[0156] In some implementations, the keyword screening unit 300 may be specifically configured to:

[0157] Obtaining the frequency of occurrence of each first keyword and user interaction behavior information, where the user interaction behavior information is information about user interaction with the video containing the first keyword;

[0158] Calculate the keyword popularity value of each first keyword based on the occurrence frequency of each first keyword and user interaction behavior information;

[0159] A first keyword whose keyword heat value is greater than or equal to a first preset threshold is taken as a hot word, and a hot keyword data set is constructed based on the hot words.

[0160] In some implementations, the keyword screening unit 300 may be specifically configured to:

[0161] Vectorizing each second keyword to obtain a keyword vector, and vectorizing the second subtitle text to obtain a text vector;

[0162] Determine whether the second keyword has the same keyword in the hot keyword dataset;

[0163] If the second keyword has the same keyword in the hot keyword dataset, the keyword hot value of the same keyword is used as the keyword hot value of the second keyword;

[0164] Calculate the relevance score between the keyword vector and the text vector based on the keyword heat value of the second keyword;

[0165] The plurality of second keywords are screened according to the relevance scores to obtain a plurality of first target keywords.

[0166] In some implementations, the keyword screening unit 300 may be specifically configured to:

[0167]

[0168] Among them, SC(d q ) represents the correlation score, H q∈A (q, t) represents the keyword heat value of the qth second keyword at the current time t, w1 represents the first weight coefficient, D represents the text vector, d q represents the extracted keyword vector corresponding to the qth second keyword, d pDenote the keyword vector corresponding to the p-th second keyword that has been screened, R denote the set of second keywords that have been screened, sim(·) denote vector similarity calculation, and max(·) denote taking the maximum value.

[0169] In some embodiments, the description text generation unit 500 may be specifically configured to:

[0170] Construct a first loss function and a second loss function;

[0171] Extract multiple third keywords from the first image description text;

[0172] Input multiple first images and the first image description text into the image text generation model, encode each first image through an encoder to obtain an encoding result;

[0173] Process the encoding result through an attention mechanism to obtain a weighted encoding result;

[0174] Decode the weighted encoding result through a decoder to obtain a predicted image description text;

[0175] Make the predicted image description text tend to the first image description text by minimizing the first loss function to obtain a target image description text;

[0176] Input the multiple third keywords, the target image description text, and the first image description text into a discriminator, and train the image text generation model by maximizing the second loss function.

[0177] In some embodiments, the description text generation unit 500 may be specifically configured to:

[0178]

[0179] Among them, LOSS D Denote the second loss function, E denote taking the mean, B denote the first image, x denote the first image description text corresponding to the first image, M r Denote the image description pair matched in the image text training set, D(·) denote the discriminator, w2 denote the second weight coefficient, Denote the target image description text, M f Denote the pair of the first image and the target image description text, s denote the number of third keywords included in the target image description text, and z denote the total number of multiple third keywords.

[0180] In some embodiments, the video summary generation unit 700 may be specifically configured to:

[0181] Perform keyword screening on multiple first target keywords and multiple second target keywords to obtain multiple third target keywords;

[0182] Input multiple third target keywords into the pre-trained large language model to generate a video introduction.

[0183] It should be noted that, since the automatic video introduction generation system in this embodiment and the automatic video introduction generation method described above are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment and will not be described in detail here.

[0184] Reference Figure 4 , an embodiment of the present application further provides an electronic device, the electronic device comprising:

[0185] at least one memory;

[0186] at least one processor;

[0187] at least one program;

[0188] The programs are stored in the memory, and the processor executes at least one program to implement the above-mentioned method for automatically generating a video introduction in the present disclosure.

[0189] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0190] The electronic device according to the embodiment of the present application is described in detail below.

[0191] The processor 1600 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure.

[0192] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called by the processor 1600 to execute the method for automatically generating a video introduction in the embodiments of this disclosure.

[0193] Input / output interface 1800, used for information input and output;

[0194] Communication interface 1900, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0195] bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 );

[0196] The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0197] The embodiment of the present disclosure further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned method for automatically generating a video introduction.

[0198] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0199] The embodiments described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems.

[0200] Those skilled in the art will understand that the technical solutions shown in the drawings do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than shown in the drawings, or a combination of certain steps, or different steps.

[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0202] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0203] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0204] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B may be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (individual) of the following" or similar expressions refer to any combination of these items, including any combination of single items (individuals) or plural items (individuals). For example, at least one (individual) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.

[0205] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in electrical, mechanical, or other forms.

[0206] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0207] In addition, each functional unit in various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs. The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.

[0209] The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.

Claims

1. A method for automatically generating a video summary, characterized in that, The method includes: Obtaining a first subtitle text corresponding to a training video, multiple first images, and a first image description text corresponding to each of the first images, and obtaining a second subtitle text and multiple second images corresponding to a video to be generated; Extracting multiple first keywords from the first subtitle text, and extracting multiple second keywords from the second subtitle text; Constructing a popularity keyword dataset based on the multiple first keywords, and screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords; Screening the multiple second images to obtain the screened second images; Inputting the screened second images into a trained image text generation model to obtain multiple second image description texts, where the trained image text generation model is obtained by training the image text generation model with the multiple first images and the first image description texts; Extracting multiple second target keywords from each of the second image description texts; Generating a video summary according to the multiple first target keywords and the multiple second target keywords.

2. The method for automatically generating a video summary according to claim 1, wherein, The constructing the popularity keyword dataset based on the multiple first keywords includes: Obtaining the occurrence frequency of each of the first keywords and user interaction behavior information, where the user interaction behavior information is information about user behavior interactions with videos containing the first keywords; Calculating the keyword popularity value of each of the first keywords according to the occurrence frequency of each of the first keywords and the user interaction behavior information; Taking the first keywords whose keyword popularity values are greater than or equal to a first preset threshold as popularity words, and constructing a popularity keyword dataset according to the popularity words.

3. The method for automatically generating a video summary according to claim 1, wherein The screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords includes: Vectorizing each of the second keywords to obtain keyword vectors, and vectorizing the second subtitle text to obtain a text vector; Determining whether there are identical keywords for the second keyword in the popularity keyword dataset; If there are identical keywords for the second keyword in the popularity keyword dataset, taking the keyword popularity value of the obtained identical keyword as the keyword popularity value of the second keyword; Calculating the correlation score between the keyword vector and the text vector according to the keyword popularity value of the second keyword; Screening the multiple second keywords according to the correlation score to obtain multiple first target keywords.

4. The method for automatically generating a video synopsis according to claim 3, wherein The calculating the correlation score between the keyword vector and the text vector according to the keyword popularity value of the second keyword includes: Among them, SC(d q ) represents the relevance score, H q∈A (q, t) represents the keyword popularity value of the q-th second keyword at the current time t, w1 represents the first weight coefficient, D represents the text vector, d q represents the extracted keyword vector corresponding to the q-th second keyword, d p represents the keyword vector corresponding to the p-th second keyword that has been screened, R represents the set of second keywords that have been screened, sim(·) represents the vector similarity calculation, and max(·) represents taking the maximum value.

5. The method for automatically generating a video profile according to claim 1, characterized in that The image text generation model includes an encoder, an attention mechanism, a decoder, and a discriminator. The trained image text generation model is obtained by training the image text generation model with the multiple first images and the first image description texts, and includes: Constructing a first loss function and constructing a second loss function; Extracting multiple third keywords from the first image description text; Input the multiple first images and the first image description text into an image-text generation model. Encode each of the first images through the encoder to obtain an encoding result; Process the encoding result through the attention mechanism to obtain a weighted encoding result; Decode the weighted encoding result through the decoder to obtain a predicted image description text; Minimize the first loss function to make the predicted image description text tend to the first image description text to obtain a target image description text; Input the multiple third keywords, the target image description text, and the first image description text into the discriminator, and maximize the second loss function to train the image-text generation model well.

6. The method for automatically generating a video summary according to claim 5, wherein, The construction of the second loss function includes: Among them, LOSS D represents the second loss function, E represents taking the mean, B represents the first image, x represents the first image description text corresponding to the first image, M r represents the image description pair matched in the image text training set, D(·) represents the discriminator, and w2 represents the second weight coefficient, represents the target image description text, M f represents the pair of the first image and the target image description text, s represents the number of third keywords included in the target image description text, and z represents the total number of multiple third keywords.

7. The method for automatically generating a video summary according to claim 1, wherein Generating a video summary according to the multiple first target keywords and the multiple second target keywords includes: Perform keyword screening on the multiple first target keywords and the multiple second target keywords to obtain multiple third target keywords; Input the multiple third target keywords into a pre-trained large language model to generate a video summary.

8. A video summary automatic generation system, characterized in that The system includes: A data acquisition unit for acquiring the first subtitle text corresponding to the training video, multiple first images, and the first image description text corresponding to each of the first images, and for acquiring the second subtitle text and multiple second images corresponding to the video to be generated; A first keyword extraction unit for extracting multiple first keywords from the first subtitle text and for extracting multiple second keywords from the second subtitle text; A keyword screening unit for constructing a popularity keyword dataset based on the multiple first keywords and for screening the multiple second keywords based on the popularity keyword dataset to obtain multiple first target keywords; An image screening unit for screening the multiple second images to obtain the screened second images; A description text generation unit for inputting the screened second images into the trained image-text generation model to obtain multiple second image description texts, where the trained image-text generation model is trained through the multiple first images and the first image description text; A second keyword extraction unit for extracting multiple second target keywords from each of the second image description texts; A video summary generation unit for generating a video summary according to the multiple first target keywords and the multiple second target keywords.

9. An electronic device, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the video summary automatic generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the video summary automatic generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Hotspot information monitoring and processing system and method

    CN121146846A