Video title generation method, device and storage medium
By matching video content features and hot keywords, calculating the overlap of video frame sets, and generating video titles that are closely related to current events, the problem of lack of innovation in titles in existing technologies is solved, and the efficiency of video generation and dissemination effects are improved.
Patent Information
- Application Number
- CN202111610447.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing automatic short video title generation methods lack distinctive features and cannot effectively combine video content, resulting in the generated titles being uninnovative and unattractive, making it difficult to meet the needs of rapid mass production.
By determining the content features of the target video, obtaining a set of hot content keywords in the previous cycle, matching the keywords with the video content features, calculating the overlap of the video frame set, and generating a video title that is closely related to current hot topics.
It improves the readability and fun of video titles, reduces the workload of operators, speeds up video production and online exposure, and enhances the spread of videos.
Smart Images

Figure CN114298018B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a method, device, and storage medium for generating a video title. Background Art
[0002] In the age of self-media, the importance of short video titles is self-evident. A good title can drive higher traffic, and even with the same content, different titles can have vastly different effects. Currently, short video titles are generally edited manually and are usually relevant to the content. Some automatically generated titles list a series of common words, which are then randomly added to the short video title by the backend system.
[0003] Manually editing titles requires high copywriting skills and sensitivity to hot topics from the video creator, video publisher or operator. If large quantities of short videos are produced, it will be even more difficult to meet the requirements of rapid video production and exposure.
[0004] The existing automatic title generation method lists a series of common words, and the background system randomly adds them to the short video titles. The generated titles are not characteristic, innovative, and cannot be well integrated with the video content. Although they meet the requirements of mass production of short videos, the stereotyped titles lack a positive role in promoting the exposure of videos. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, and storage medium for generating a video title, which can generate a video title that is closely integrated with current hot topics for a video, increase the readability and interest of the video title, and reduce the workload of operators.
[0006] A first aspect of the present application provides a method for generating a video title, which may include:
[0007] Determine the content features corresponding to the target video;
[0008] Get the keyword set of hot content in the previous period;
[0009] Matching each keyword in the keyword set with the content feature corresponding to the target video to obtain a matching result set;
[0010] The video title of the target video is determined according to the matching result set.
[0011] In one possible design, determining the video title of the target video according to the matching result set includes:
[0012] Determining a set of video frames in the target video that corresponds to the set of matching results;
[0013] Calculating the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set;
[0014] Determining a first keyword set corresponding to a first target coincidence degree, where the first target coincidence degree is a coincidence degree with a minimum coincidence value in the first coincidence degree set;
[0015] A video title of the target video is generated according to the first keyword set.
[0016] In one possible design, determining the set of video frames in the target video corresponding to the set of matching results includes:
[0017] Matching each video frame in the target video with each matching result in the matching result set to obtain a video frame set corresponding to each matching result in the matching result set;
[0018] The video frame sets corresponding to the matching results whose number of video frames is less than a preset threshold are eliminated from the matching result set to obtain the video frame set corresponding to the matching result set.
[0019] In one possible design, the method further includes:
[0020] If the number of video frames in the video frame set corresponding to each matching result in the matching result set is less than the preset threshold, determining the frame ratio of the video frames corresponding to each matching result in the matching result set;
[0021] Calculating the standard deviation value corresponding to each matching result in the matching result set;
[0022] Determining a weighted average value of each matching result in the matching result set according to the frame number ratio and the standard deviation value;
[0023] Calculating the degree of overlap between the set of video frames corresponding to the target matching result with the largest weighted average and the set of video frames corresponding to each matching result in the other matching result subsets to obtain a second set of overlaps, where the other matching result subsets are the sets of matching results in the matching result set excluding the target matching result;
[0024] Determining a second keyword set corresponding to a second target coincidence degree, where the second target coincidence degree is a coincidence degree with a minimum coincidence value in the second coincidence degree set;
[0025] A video title of the target video is generated according to the second keyword set.
[0026] In one possible design, the content features corresponding to the target video include a first set of character names and accuracy rates, a second set of character actions and accuracy rates, a third set of objects and accuracy rates, and a fourth set of scenes and accuracy rates corresponding to the target video. The method further includes:
[0027] If each keyword in the keyword set fails to match the first set, the second set, the third set, and the fourth set, determining, in the target video, a set of video frames corresponding to the first set, a set of video frames corresponding to the second set, a set of video frames corresponding to the third set, and a set of video frames corresponding to the fourth set;
[0028] calculating a degree of overlap between video frame sets corresponding to any two of the video frame sets corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set, to obtain a third degree of overlap set;
[0029] Determining a third keyword set corresponding to a third target coincidence degree, wherein the third target coincidence degree is a coincidence degree having a minimum coincidence value in the third coincidence degree set;
[0030] The video title of the target video is determined according to the third keyword set.
[0031] In one possible design, after determining the first keyword set corresponding to the first target overlap degree, the method further includes:
[0032] Determining a target video frame set corresponding to the first keyword set in the target video;
[0033] Calculating the degree of overlap between the target video frame set and other video frame sets to obtain a fourth degree of overlap set, where the other video frame sets are video frame sets in the video frame sets corresponding to the matching result set except the video frame set corresponding to the first target degree of overlap;
[0034] Determining a fourth keyword set corresponding to a fourth target coincidence degree, wherein the fourth target coincidence degree is a coincidence degree having a minimum coincidence value in the fourth coincidence degree set;
[0035] A video title of the target video is generated according to the first keyword set and the fourth keyword set.
[0036] In one possible design, calculating the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set includes:
[0037] The degree of overlap between any two video frame sets of matching results in the matching result set is calculated using the following formula:
[0038]
[0039] Wherein, PD(F1, F2) is the overlap between any two video frame sets of matching results in the matching result set, F1 is the video frame set corresponding to any matching result in the matching result set, F2 is any video frame set except F1 in the video frame set corresponding to the matching result set, and F x is any video frame included in both F1 and F2. x [F1] is F x Position in F1, F x [F2] is F x Position in F2.
[0040] A second aspect of the present application provides a video title generation device, comprising:
[0041] A first determining unit, configured to determine content features corresponding to a target video;
[0042] An acquisition unit, used to acquire a keyword set of hot content in the previous period;
[0043] a matching unit, configured to match each keyword in the keyword set with a content feature corresponding to the target video to obtain a matching result set;
[0044] A second determining unit is configured to determine a video title of the target video according to the matching result set.
[0045] In one possible design, the second determining unit is specifically configured to:
[0046] Determining a set of video frames in the target video that corresponds to the set of matching results;
[0047] Calculating the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set;
[0048] Determining a first keyword set corresponding to a first target coincidence degree, where the first target coincidence degree is a coincidence degree with a minimum coincidence value in the first coincidence degree set;
[0049] A video title of the target video is generated according to the first keyword set.
[0050] In one possible design, the second determining unit determines the set of video frames in the target video corresponding to the set of matching results, including:
[0051] Matching each video frame in the target video with each matching result in the matching result set to obtain a video frame set corresponding to each matching result in the matching result set;
[0052] The video frame sets corresponding to the matching results whose number of video frames is less than a preset threshold are eliminated from the matching result set to obtain the video frame set corresponding to the matching result set.
[0053] In one possible design, the second determining unit is further configured to:
[0054] If the number of video frames in the video frame set corresponding to each matching result in the matching result set is less than the preset threshold, determining the frame ratio of the video frames corresponding to each matching result in the matching result set;
[0055] Calculating the standard deviation value corresponding to each matching result in the matching result set;
[0056] Determining a weighted average value of each matching result in the matching result set according to the frame number ratio and the standard deviation value;
[0057] Calculating the degree of overlap between the set of video frames corresponding to the target matching result with the largest weighted average and the set of video frames corresponding to each matching result in the other matching result subsets to obtain a second set of overlaps, where the other matching result subsets are the sets of matching results in the matching result set excluding the target matching result;
[0058] Determining a second keyword set corresponding to a second target coincidence degree, where the second target coincidence degree is a coincidence degree with a minimum coincidence value in the second coincidence degree set;
[0059] A video title of the target video is generated according to the second keyword set.
[0060] In one possible design, the content features corresponding to the target video include a first set of character names and accuracy rates, a second set of character actions and accuracy rates, a third set of objects and accuracy rates, and a fourth set of scenes and accuracy rates corresponding to the target video. The second determining unit 204 is further configured to:
[0061] If each keyword in the keyword set fails to match the first set, the second set, the third set, and the fourth set, determining, in the target video, a set of video frames corresponding to the first set, a set of video frames corresponding to the second set, a set of video frames corresponding to the third set, and a set of video frames corresponding to the fourth set;
[0062] calculating a degree of overlap between video frame sets corresponding to any two of the video frame sets corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set, to obtain a third degree of overlap set;
[0063] Determining a third keyword set corresponding to a third target coincidence degree, wherein the third target coincidence degree is a coincidence degree having a minimum coincidence value in the third coincidence degree set;
[0064] The video title of the target video is determined according to the third keyword set.
[0065] In one possible design, the second determining unit is further configured to:
[0066] Determining a target video frame set corresponding to the first keyword set in the target video;
[0067] Calculating the degree of overlap between the target video frame set and other video frame sets to obtain a fourth degree of overlap set, where the other video frame sets are video frame sets in the video frame sets corresponding to the matching result set except the video frame set corresponding to the first target degree of overlap;
[0068] Determining a fourth keyword set corresponding to a fourth target coincidence degree, wherein the fourth target coincidence degree is a coincidence degree having a minimum coincidence value in the fourth coincidence degree set;
[0069] A video title of the target video is generated according to the first keyword set and the fourth keyword set.
[0070] In one possible design, the second determining unit calculates the overlap between any two video frame sets of matching results in the matching result set, and obtains a first overlap set including:
[0071] The degree of overlap between any two video frame sets of matching results in the matching result set is calculated using the following formula:
[0072]
[0073] Wherein, PD(F1, F2) is the overlap between any two video frame sets of matching results in the matching result set, F1 is the video frame set corresponding to any matching result in the matching result set, F2 is any video frame set except F1 in the video frame set corresponding to the matching result set, and F x is any video frame included in both F1 and F2. x [F1] is F x Position in F1, Fx [F2] is F x Position in F2.
[0074] A third aspect of the present application provides a computing device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0075] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the method for generating a video title.
[0076] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one executable instruction. When the executable instruction is executed on a computing device, the computing device executes the method for generating a video title as described in the first aspect of the present application.
[0077] A fifth aspect of the present application discloses a computer program product. When the computer program product is run on a computer, the computer is caused to execute the method for generating a video title as described in the first aspect of the present application.
[0078] In a sixth aspect, the present application discloses an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer executes the method for generating a video title described in the first aspect of the present application.
[0079] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0080] In the embodiment provided by this application, by crawling the keyword set of the hot content in the previous cycle, and determining the matching record between the keywords in the keyword set and each object in the video to be generated, and determining the video frame set corresponding to the matching result, the overlap between each video set can be calculated, and the keywords associated with it in the keyword set can be determined based on the overlap, and then the video title of the video can be generated based on the keywords. In this way, video titles that are closely integrated with current hot topics can be generated for batch videos, increasing the readability and fun of the video titles, reducing the workload of operators, and increasing the speed of video production and online exposure, thereby improving the spread of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:
[0082] Figure 1 Schematic diagram of the flow of the method for generating a video title in an embodiment of the present application;
[0083] Figure 2 This is a virtual structural diagram of a video title generating device in an embodiment of the present application;
[0084] Figure 3 This is a schematic diagram of the structure of the server in the embodiment of the present application. DETAILED DESCRIPTION
[0085] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All embodiments in the present invention should fall within the scope of protection of the present invention.
[0086] The following specifically describes the method for generating a video title provided in this application from the perspective of a video title generating device. The video title generating device may be a server or a service unit in the server.
[0087] See also Figure 1 , Figure 1 A schematic diagram of an embodiment of a method for generating a video title provided in an embodiment of the present application includes:
[0088] 101. Determine content features corresponding to the target video.
[0089] In this embodiment, when the creator publishes a new target video, the video title generation device can analyze the target video after receiving the target video to determine the content features corresponding to the target video, wherein the content features include a first set of target name (the target name is the name of the character appearing in the target video) and accuracy, a second set of target action (the target action is the action of the character appearing in the target video) and accuracy, a third set of object (the object is the object appearing in the target video, such as a basketball stand, a table tennis table lamp) and accuracy, and a fourth set of scene (the scene is the scene appearing in the target video, such as a football field, a basketball court, etc.) and accuracy. The accuracy is the accuracy of the character name, the accuracy of the target action, the accuracy of the object, and the accuracy of the scene obtained by recognizing each frame of the target video. The target video is the video for which the title is to be generated. Specifically, the video title generation device can input the target video into different neural networks respectively, and obtain the data sets of the characters, actions, scenes, objects and corresponding accuracy in the target video respectively, which are explained in detail below:
[0090] First, a face recognition algorithm based on a neural network is used to detect and identify the target video frame by frame, and a set of character names and accuracy rates with an accuracy rate of more than 95% is output, and the first set is marked as N g ={N1:P1,N2:P2,...,N x :P x}, where N is the character name, and P is the accuracy corresponding to the character name. Finally, the first set is stored in the background data cache.
[0091] Then, the regional 3D convolutional network based on temporal action detection detects and recognizes the target video frame by frame to find the potential action time interval in the target video and judge the action category. The set of character actions and accuracy with an accuracy rate of more than 95% in the video is output, which is the second set and marked as A. g ={A1:P1,A2:P2,...,A x :P x}, where A is the character action, P is the accuracy corresponding to the character action, and finally the second set is saved to the background data cache.
[0092] The target video is detected and identified frame by frame by the target detection model to obtain a data set of objects and accuracy rates with an accuracy rate of more than 95% in the target video, which is the third set, and the third set is marked as O g ={O1:P1,O2:P2,...,O x :P x}, where O is the object and P is the accuracy corresponding to the object, and the third set is saved in the background data cache.
[0093] Finally, the target video is detected and identified frame by frame through the scene classification algorithm to obtain the scenes appearing in the target video, and a data set of scenes and accuracy with an accuracy rate of more than 95% is obtained, which is the fourth set, and the fourth set is marked as S g ={S1:P1,S2:P2,...,S x :P x}, where S is the scene, P is the accuracy corresponding to the scene, and the fourth set is saved in the background data cache.
[0094] It should be noted that each of the above models can be obtained by training in advance through training samples. At the same time, the accuracy mentioned above can also be other values, such as 90%, which is not specifically limited.
[0095] 102. Obtain a keyword set of hot content in the previous period.
[0096] In this embodiment, the video title generation device can obtain the keyword set of the hot content in the previous period, that is, it can set a timer task to start every morning in the background to crawl the hot search list or the hot list of the previous day, and extract the keyword set corresponding to the hot content through natural language processing, and mark the keyword set as H g ={H1,H2,...,H x It is understandable that the period can be set to 1 day, or 12 hours, or other durations. In addition, hot content can be crawled from the hot search list or the popular list, or from other rankings. It can also be set by yourself, and there is no specific limitation.
[0097] It should be noted that, through step 101, the content features corresponding to the target video can be determined, and through step 102, a keyword set can be obtained. However, there is no restriction on the order of execution between these two steps. Step 101 can be executed first, or step 102 can be executed first, or they can be executed at the same time. There is no specific limitation.
[0098] 103. Match each keyword in the keyword set with the content feature corresponding to the target video to obtain a matching result set.
[0099] In this embodiment, after obtaining the keyword set of the previous cycle and the content features corresponding to the target video, the video title generation device can match each keyword in the keyword set with the first set, the second set, the third set, and the fourth set to obtain a matching result set. Specifically, the keyword set H can be traversed. g Each keyword in N g ,A g ,O g ,S g Each object in is processed to obtain a matching result set, such as H g The first keyword H1 in the g N1 in the example is "Zhan Moumou", and the matching result can be added to the matching result set R g In this case, the matching result set is R g ={N1:P1}, and so on, then H2 and N g ,A g ,O g ,S g Match each data in, assume that the matching result A2:P2 is obtained, and add the matching result to the matching result set R g In this case, the matching result set is R g={N1:P1,A2:P2}, match them in sequence to get the final
[0100] Matching result set R g The x in the equation can be any number, for example, x1=1, x2=2, x3=2, x4=2, then R g ={N1:P1,A1:P1,A2:P2,O1:P1,O2:P2,S1:P2,S2:P2}, which means that the keyword set of the hot content in the previous period includes one person, two actions, two objects, and two scenes.
[0101] 104. Determine the video title of the target video according to the matching result set.
[0102] In this embodiment, after determining the matching result set, the video title generating device can determine the video title of the target video based on the matching result set. The following is a detailed description of determining the video title of the target video based on the matching result set:
[0103] Step 1: Determine the video frame set in the target video that corresponds to the matching result set.
[0104] In this step, the video title generation device can determine the video frame set in the target video corresponding to the matching result set, that is, match each video frame in the target video with each matching result in the matching result set to obtain the video frame set corresponding to each matching result in the matching result set; and eliminate the video frame set corresponding to the matching result in which the number of video frames in the matching result set is less than a preset threshold to obtain the video frame set corresponding to the matching result set.
[0105] Below is R g ={N1:P1,A1:P1,A2:P2,O1:P1,O2:P2,S1:P2,S2:P2} is used as an example to illustrate how to determine the set of video frames corresponding to the matching result in the target video, as follows:
[0106] The video title generation device can traverse all video frames in the target video, detect tasks, actions, objects and scenes in each frame, mark all video frames where N1, A1, A2, O1, O2, S1, S2 are detected, and obtain the relationship set R between the video frames in the target video and N1, A1, A2, O1, O2, S1, S2 F , assuming that the number of all frames of the target video is F A , requiring the relation set R F The number of frames corresponding to tasks, actions, objects and scenes in the game should account for F AMore than 50% (of course, it can also be set according to actual conditions, such as 60%, which is not limited to this). For example, among N1, A1, A2, O1, O2, S1, and S2, the number of video frames corresponding to S2 accounts for less than 50%. Then, the video frames corresponding to S2 are filtered, and the filtered relationship set is finally obtained as follows:
[0107]
[0108] Among them, F is the video frame in the target video, corresponding to the position where the N1, A1, A2, O1, O2, S1 objects appear in the target video.
[0109] Step 2: Calculate the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set.
[0110] In this step, after obtaining the video frame set corresponding to the matching result set, the video title generation device can calculate the overlap between any two matching result video frame sets in the matching result set to obtain a first overlap set. Specifically, the overlap between any two matching result video frame sets in the matching result set can be calculated using the following formula:
[0111]
[0112] Among them, PD(F1, F2) is the overlap between any two matching result video frame sets in the matching result set, F1 is the video frame set corresponding to any matching result in the matching result set, and F2 is any video frame set except F1 in the video frame set corresponding to the matching result set. x is any video frame included in both F1 and F2. x [F1] is F x Position in F1, F x [F2] is F x Position in F2.
[0113] The following is an example of the matching result set N1, A1, A2, O1, O2, S1, S2 to calculate the video frame set F corresponding to N1. N1 ={F1,F2,...,F n The video frame set F corresponding to A1 A1 ={F1,F2,...,F n ,F n+2}, set F x In F N1 The position in is F x [F N1 ](0≤x≤n-1), Fx In F A1 The position in is F x [F A1 ](0≤x≤n+1), then F x [F N1 ]-F x [F A1 ] is the video frame F x The video frame set corresponding to N1 and F x The position difference in the video frame set corresponding to A1 finally gives the formula: That is, the video frame set F corresponding to N1 N1 The video frame set F corresponding to A1 A1 The average of the absolute values of the position differences of the same video frames between them is the video frame set F corresponding to N1. N1 The video frame set F corresponding to A1 A1 The smaller the overlap degree PD between the obtained video frame sets, the more frames N1 and A1 appear simultaneously in the target video.
[0114] Then calculate F using the following formula N1 and the video frame set F of A2 A2 The overlap between:
[0115]
[0116] Among them, PD(F N1 ,F A2 ) is F N1 and F A2 The overlap between them, F x [F N1 ]-F x [F A2 ] is the video frame F x The video frame set corresponding to N1 and F x The position difference in the set of video frames corresponding to A2.
[0117] Finally, according to the arrangement and combination method, PD(F N1 ,F O1 ),…,PD(F O2 ,F S1 ), which is the first coincidence degree set.
[0118] Step 3: Determine a first keyword set corresponding to the first target overlap degree.
[0119] In this step, after determining the first set of coincidence degrees, the video title generation device may first determine the first target coincidence degree, which is the coincidence degree with the smallest coincidence degree value in the first set of coincidence degrees. Then, it determines the first keyword set corresponding to the first target coincidence degree. For example, finally, the value of PD(F N1 ,F O1 ) is the smallest, N1 is "Mr. Zhan", and O1 is "basketball", which means that the probability of "Mr. Zhan" and "basketball" appearing simultaneously in the target video is the highest, and the video frames corresponding to "Mr. Zhan" and "basketball" account for more than 50% of the total duration of the target video. "Mr. Zhan" and "basketball" are the first keyword set.
[0120] Step 4: Generate the video title of the target video according to the first keyword set.
[0121] In this step, after obtaining the first keyword set, the video title generation device may generate a video title according to the first keyword set by using natural language understanding. The specific way of generating the video title is not limited here, as long as it can be generated through natural language understanding. In addition, after generating the video title of the target video, the target video can be launched, and when the search keyword input by the user matches the video title of the target video, the target video is pushed to the user to achieve accurate pushing.
[0122] It should be noted that with a title that fits the video content, when a user searches for videos related to a certain content, through the matching of the content input by the user and the title of our system, the content that the user wants to see can be better displayed, improving the relevance of the search result display. Or when a certain hot content appears recently, using the video materials in the existing resource library and combining with the video keywords to generate a title related to the hot content, so as to re-recommend the already out-of-date video content under the hot topic materials. For example, a user uploaded a video of a cute pet cat before, and the video content keywords (cat, demolishing the house) were identified. According to the Internet hot search topic "What bad intentions can a cat have", through the above steps, this video was matched in the background, and a new title was given to this historical video. So when a user searches for the topic "What bad intentions can a cat have", this video can be recommended to the user again, making full use of the video materials, fitting the user's search intention, and increasing the video exposure speed.
[0123] In one embodiment, the video title generation device further performs the following operations:
[0124] If the number of video frames in the video frame set corresponding to each matching result in the matching result set is less than the preset threshold, then determine the frame ratio of the video frames corresponding to each matching result in the matching result set;
[0125] Calculate the standard deviation value corresponding to each matching result in the matching result set;
[0126] Determine the weighted average of each matching result in the matching result set based on the frame ratio and the standard deviation value;
[0127] Calculate the degree of overlap between the set of video frames corresponding to the target matching result with the largest weighted average and the set of video frames corresponding to each matching result in the other matching result subsets, where the other matching result subsets are the sets of matching results in the matching result set excluding the target matching result;
[0128] Determining a second keyword set corresponding to a second target coincidence degree, where the second target coincidence degree is the coincidence degree with the smallest coincidence value in the second coincidence degree set;
[0129] Generate a video title for the target video based on the second keyword set
[0130] In this embodiment, if the number of video frames in the video frame set corresponding to each matching set in the matching result set is less than the preset threshold, it means that the number of video frames that can be matched by each matching result in the matching result set in the target video is less than the preset threshold. At this time, R e If it is empty, the standard deviation value corresponding to each matching result in the matching result set is calculated. For example, the matching result set R F The corresponding video frames are:
[0131]
[0132] Among them, the frame rate ratio of each object in N1, A1, A2, O1, O2, S1, and S2 is calculated in turn. Taking N1 as an example, the frame rate ratio of N1 is n is the number of frames in the video frame corresponding to N1, F A is the total number of frames of the target video. From this, we can calculate the frame ratio of each object N1, A1, A2, O1, O2, S1, and S2, and arrange the frame ratios in ascending order. For example, the obtained arrangement is: {1, 3, 2, 7, 5, 6, 4}.
[0133] Then, according to the standard deviation formula, calculate the standard deviation values corresponding to N1, A1, A2, O1, O2, S1, and S2. For example, the standard deviation value corresponding to N1 is And sort the standard deviations of all objects in N1, A1, A2, O1, O2, S1, S2 from small to large. The standard deviation rankings of N1, A1, A2, O1, O2, S1, S2 are {2, 5, 4, 1, 3, 7, 6}.
[0134] Then, the weighted average of each matching result in the matching result set is determined based on the frame rate and the standard deviation. Specifically, the weights of the frame rate and the standard deviation can be set to 30% and 70% respectively (of course, they can also be adjusted according to actual conditions, and there is no specific limit). Therefore, the weighted average of N1 can be calculated using the following formula:
[0135]
[0136] in, is the frame rate ratio of N1, is the standard deviation value of N1, is the weighted average of N1, and thus the weighted average of the remaining objects in the matching result set can be calculated based on this formula.
[0137] Then calculate the overlap between the video frame set corresponding to the target matching result with the largest weighted average and the video frame set corresponding to each matching result in the other matching result subsets to obtain a second overlap set. For example, the final result is The value of is the largest, then it is considered that the probability that S1 is one of the main contents of the target video is the highest. The overlap between the video frame set corresponding to S1 and the video frame sets of other objects can be calculated to obtain a second overlap set. The method of calculating the overlap has been described in detail above and will not be repeated here.
[0138] Finally, after obtaining the second overlap set, the overlap with the smallest overlap value among the second target overlaps can be selected as the second target overlap, and the second keyword set corresponding to the second target overlap can be determined, and then the video title of the target video can be generated according to the second keyword set.
[0139] In one embodiment, if H g With N g ,A g ,O g ,S g If there is no result in the set matching, that is, the result object analyzed from the target video has no intersection with the hot words in the previous cycle, the video title generation device can generate the video title of the target video in the following way:
[0140] If each keyword in the keyword set fails to match the first set, the second set, the third set, and the fourth set, determining the video frame set corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set in the target video;
[0141] Calculating the overlap between video frame sets corresponding to any two of the video frame sets corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set to obtain a third overlap set;
[0142] Determining a third keyword set corresponding to a third target coincidence degree, where the third target coincidence degree is the coincidence degree with the smallest coincidence value in the third coincidence degree set;
[0143] The video title of the target video is determined according to the third keyword set.
[0144] In this embodiment, if each keyword in the keyword set fails to match the first set, the second set, the third set and the fourth set, then the video frame set corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set and the video frame set corresponding to the fourth set in the target video are determined. Then, the video frames in the target video can be traversed to find the video frame set in the target video corresponding to the first set, the video frame set in the target video corresponding to the second set, the video frame set in the target video corresponding to the third set, and the video frame set in the target video corresponding to the fourth set. Afterwards, the degree of overlap between any two video frame sets in each video frame set is calculated to obtain a third degree of overlap set. The above has already described in detail the method for calculating the degree of overlap, and will not be repeated here.
[0145] Then, a third keyword set corresponding to a third target overlap with the smallest overlap value in the third overlap set is determined, and a video title of the target video is generated according to the third keyword set.
[0146] In one embodiment, after determining the first keyword set corresponding to the first target overlap degree, the video title generation device further performs the following operations:
[0147] Determining a target video frame set corresponding to the first keyword set in the target video;
[0148] Calculating the overlap between the target video frame set and the other video frame sets to obtain a fourth overlap set, where the other video frame sets are video frame sets in the video frame set corresponding to the matching result set except the video frame set corresponding to the first target overlap;
[0149] Determining a fourth keyword set corresponding to a fourth target coincidence degree, where the fourth target coincidence degree is the coincidence degree with the smallest coincidence value in the fourth coincidence degree set;
[0150] A video title of the target video is generated according to the first keyword set and the fourth keyword set.
[0151] In this embodiment, after determining the first keyword set corresponding to the first target overlap, the video title generating device can use the video frame set corresponding to the keywords in the first keyword set as a reference, calculate the overlap between it and the other videos in the video frame set corresponding to the matching result set except the video frames corresponding to the first keyword set, and obtain a fourth overlap set (the overlap calculation has been described in detail above and will not be repeated here), and then determine the fourth keyword set corresponding to the fourth target overlap with the smallest overlap value in the fourth overlap set, and generate a video title for the target video based on the fourth keyword set and the first keyword set, thereby obtaining more associated words for the target video, thereby generating a title that is more suitable for the video content.
[0152] In summary, in the embodiment provided by the present application, by crawling the keyword set of the hot content in the previous cycle, and determining the matching records between the keywords in the keyword set and each object in the video for which the title is to be generated, and determining the video frame set corresponding to the matching results, the overlap between each video set can be calculated, and the keywords associated with them in the keyword set can be determined based on the overlap, and then the video title of the video can be generated based on the keywords. In this way, video titles that are closely integrated with current hot topics can be generated for batch videos, increasing the readability and fun of the video titles, reducing the workload of operators, and increasing the speed of video production and online exposure, thereby improving the spread of the video.
[0153] The above describes the embodiment of the present application from the perspective of the method for generating a video title. The following describes the embodiment of the present application from the perspective of the video title generating device:
[0154] See also Figure 2 , Figure 2 This is a schematic diagram of an embodiment of a video title generation device provided in an embodiment of the present application. The video title generation device 200 includes:
[0155] A first determining unit 201 is configured to determine content features corresponding to a target video;
[0156] An acquisition unit 202 is configured to acquire a keyword set of hot content in the previous period;
[0157] A matching unit 203 is configured to match each keyword in the keyword set with a content feature corresponding to the target video to obtain a matching result set;
[0158] The second determining unit 204 is configured to determine the video title of the target video according to the matching result set.
[0159] In one possible design, the second determining unit 204 is specifically configured to:
[0160] Determining a set of video frames in the target video that corresponds to the set of matching results;
[0161] Calculating the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set;
[0162] Determining a first keyword set corresponding to a first target coincidence degree, where the first target coincidence degree is a coincidence degree with a minimum coincidence value in the first coincidence degree set;
[0163] A video title of the target video is generated according to the first keyword set.
[0164] In one possible design, the second determining unit 204 determines that the set of video frames in the target video corresponding to the set of matching results includes:
[0165] Matching each video frame in the target video with each matching result in the matching result set to obtain a video frame set corresponding to each matching result in the matching result set;
[0166] The video frame sets corresponding to the matching results whose number of video frames is less than a preset threshold are eliminated from the matching result set to obtain the video frame set corresponding to the matching result set.
[0167] In one possible design, the second determining unit 204 is further configured to:
[0168] If the number of video frames in the video frame set corresponding to each matching result in the matching result set is less than the preset threshold, determining the frame ratio of the video frames corresponding to each matching result in the matching result set;
[0169] Calculating the standard deviation value corresponding to each matching result in the matching result set;
[0170] Determining a weighted average value of each matching result in the matching result set according to the frame number ratio and the standard deviation value;
[0171] Calculating the degree of overlap between the set of video frames corresponding to the target matching result with the largest weighted average and the set of video frames corresponding to each matching result in the other matching result subsets to obtain a second set of overlaps, where the other matching result subsets are the sets of matching results in the matching result set excluding the target matching result;
[0172] Determining a second keyword set corresponding to a second target coincidence degree, where the second target coincidence degree is a coincidence degree with a minimum coincidence value in the second coincidence degree set;
[0173] A video title of the target video is generated according to the second keyword set.
[0174] In one possible design, the content features corresponding to the target video include a first set of character names and accuracy rates, a second set of character actions and accuracy rates, a third set of objects and accuracy rates, and a fourth set of scenes and accuracy rates corresponding to the target video. The second determining unit 204 is further configured to:
[0175] If each keyword in the keyword set fails to match the first set, the second set, the third set, and the fourth set, determining, in the target video, a set of video frames corresponding to the first set, a set of video frames corresponding to the second set, a set of video frames corresponding to the third set, and a set of video frames corresponding to the fourth set;
[0176] calculating a degree of overlap between video frame sets corresponding to any two of the video frame sets corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set, to obtain a third degree of overlap set;
[0177] Determining a third keyword set corresponding to a third target coincidence degree, wherein the third target coincidence degree is a coincidence degree having a minimum coincidence value in the third coincidence degree set;
[0178] The video title of the target video is determined according to the third keyword set.
[0179] In one possible design, the second determining unit 204 is further configured to:
[0180] Determining a target video frame set corresponding to the first keyword set in the target video;
[0181] Calculating the degree of overlap between the target video frame set and other video frame sets to obtain a fourth degree of overlap set, where the other video frame sets are video frame sets in the video frame sets corresponding to the matching result set except the video frame set corresponding to the first target degree of overlap;
[0182] Determining a fourth keyword set corresponding to a fourth target coincidence degree, wherein the fourth target coincidence degree is a coincidence degree having a minimum coincidence value in the fourth coincidence degree set;
[0183] A video title of the target video is generated according to the first keyword set and the fourth keyword set.
[0184] In one possible design, the second determining unit 204 calculates the overlap between any two video frame sets of matching results in the matching result set, and obtains a first overlap set including:
[0185] The degree of overlap between any two video frame sets of matching results in the matching result set is calculated using the following formula:
[0186]
[0187] Wherein, PD(F1, F2) is the overlap between any two video frame sets of matching results in the matching result set, F1 is the video frame set corresponding to any matching result in the matching result set, F2 is any video frame set except F1 in the video frame set corresponding to the matching result set, and F x is any video frame included in both F1 and F2. x [F1] is F x Position in F1, F x [F2] is F x Position in F2.
[0188] The present application also provides a computing device, which may be a server. Figure 3 , Figure 3 3 is a structural diagram of a server provided by an embodiment of the present invention. The server 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 322 (for example, one or more processors) and memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 can be temporary storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 322 can be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the server 300.
[0189] The server 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0190] The steps performed by the video title generating device in the above embodiment can be based on the Figure 3 The server structure shown.
[0191] The present application also provides a computer-readable storage medium, wherein the storage medium stores at least one executable instruction. When the executable instruction is executed on a computing device, the computing device executes the method for generating a video title described in any of the above embodiments.
[0192] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0193] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.
[0194] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0195] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0196] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0197] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0198] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0199] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating a video title, characterized in that: include: Determine the content features corresponding to the target video; Get the keyword set of hot content in the previous period; Matching each keyword in the keyword set with the content feature corresponding to the target video to obtain a matching result set; Matching each video frame in the target video with each matching result in the matching result set to obtain a video frame set corresponding to each matching result in the matching result set; removing the video frame sets corresponding to matching results whose number of video frames in the matching result set is less than a preset threshold to obtain a video frame set corresponding to the matching result set; calculating the overlap between any two video frame sets of the matching results in the matching result set to obtain a first overlap degree set; Determine a first keyword set corresponding to a first target coincidence, where the first target coincidence is a coincidence with a minimum coincidence value in the first coincidence set; and generate a video title for the target video according to the first keyword set.
2. The method according to claim 1, characterized in that The method further comprises: If the number of video frames in the video frame set corresponding to each matching result in the matching result set is less than the preset threshold, determining the frame ratio of the video frames corresponding to each matching result in the matching result set; Calculating the standard deviation value corresponding to each matching result in the matching result set; Determining a weighted average value of each matching result in the matching result set according to the frame number ratio and the standard deviation value; Calculating the degree of overlap between the set of video frames corresponding to the target matching result with the largest weighted average and the set of video frames corresponding to each matching result in the other matching result subsets to obtain a second set of overlaps, where the other matching result subsets are the sets of matching results in the matching result set excluding the target matching result; Determining a second keyword set corresponding to a second target coincidence degree, where the second target coincidence degree is a coincidence degree with a minimum coincidence value in the second coincidence degree set; A video title of the target video is generated according to the second keyword set.
3. The method according to claim 1, characterized in that The content features corresponding to the target video include a first set of character names and accuracy rates, a second set of character actions and accuracy rates, a third set of objects and accuracy rates, and a fourth set of scenes and accuracy rates corresponding to the target video. The method further includes: If each keyword in the keyword set fails to match the first set, the second set, the third set, and the fourth set, determining, in the target video, a set of video frames corresponding to the first set, a set of video frames corresponding to the second set, a set of video frames corresponding to the third set, and a set of video frames corresponding to the fourth set; calculating a degree of overlap between video frame sets corresponding to any two of the video frame sets corresponding to the first set, the video frame set corresponding to the second set, the video frame set corresponding to the third set, and the video frame set corresponding to the fourth set, to obtain a third degree of overlap set; Determining a third keyword set corresponding to a third target coincidence degree, wherein the third target coincidence degree is a coincidence degree having a minimum coincidence value in the third coincidence degree set; The video title of the target video is determined according to the third keyword set.
4. The method according to any one of claims 1 to 3, characterized in that After determining the first keyword set corresponding to the first target overlap degree, the method further includes: Determining a target video frame set corresponding to the first keyword set in the target video; Calculating the degree of overlap between the target video frame set and other video frame sets to obtain a fourth degree of overlap set, where the other video frame sets are video frame sets in the video frame sets corresponding to the matching result set except the video frame set corresponding to the first target degree of overlap; Determining a fourth keyword set corresponding to a fourth target coincidence degree, wherein the fourth target coincidence degree is a coincidence degree having a minimum coincidence value in the fourth coincidence degree set; A video title of the target video is generated according to the first keyword set and the fourth keyword set.
5. The method according to any one of claims 1 to 3, characterized in that Calculating the overlap between any two video frame sets of matching results in the matching result set to obtain a first overlap set includes: The degree of overlap between any two video frame sets of matching results in the matching result set is calculated using the following formula: Wherein, PD(F1, F2) is the overlap between any two video frame sets of matching results in the matching result set, F1 is the video frame set corresponding to any matching result in the matching result set, F2 is any video frame set except F1 in the video frame set corresponding to the matching result set, and F x is any video frame included in both F1 and F2. x [F1] is F x Position in F1, F x [F2] is F x Position in F2.
6. A video title generation device, characterized in that: include: A first determining unit, configured to determine content features corresponding to a target video; An acquisition unit, used to acquire a keyword set of hot content in the previous period; a matching unit, configured to match each keyword in the keyword set with a content feature corresponding to the target video to obtain a matching result set; The second determining unit is configured to match each video frame in the target video with each matching result in the matching result set, respectively, to obtain a video frame set corresponding to each matching result in the matching result set; eliminate video frame sets corresponding to matching results having a number of video frames less than a preset threshold in the matching result set, to obtain a video frame set corresponding to the matching result set; and calculate a degree of overlap between the video frame sets of any two matching results in the matching result set, to obtain a first degree of overlap set; Determine a first keyword set corresponding to a first target coincidence, where the first target coincidence is a coincidence with a minimum coincidence value in the first coincidence set; and generate a video title for the target video according to the first keyword set.
7. A computing device, characterized in that include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the method for generating a video title according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one executable instruction. When the executable instruction is executed on a computing device, the computing device executes the method for generating a video title according to any one of claims 1 to 5.
Citation Information
Patent Citations
Bullet screen information processing method, device and equipment
CN110166811A
Music recommendation method and device and readable storage medium
CN113569088A