Long video tag generation method and device, electronic equipment and storage medium
By using a large language model to generate plot semantic vectors and groupings of long videos, combined with the target prompt text, the problems of low efficiency and high granularity of manual editing and generating long video tags are solved, and more efficient and accurate video tag generation is achieved.
Patent Information
- Application Number
- CN202510119613.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
Video tags for long videos in the prior art are usually generated by manual editing, which has problems such as large granularity and low generation efficiency of video tags, which affects the user experience.
By obtaining preset long video sets and their plot description information, a large language model is used to generate the first plot semantic vector, video grouping is performed, and target prompt text is constructed to generate target video tags.
It improves the generation efficiency of long video tags, reduces the granularity of tags, makes the generated tags more in line with the plot theme, and improves the user experience.
Smart Images

Figure CN119988674A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a method, device, electronic device and storage medium for generating a long video tag. Background Art
[0002] On long video websites, video tags are the most important basic information for describing the long video itself. Through video tags, the content of long videos can be well classified, and users can also more conveniently filter and obtain the long videos they want through video tags. At present, video tags for long videos are usually generated through manual editing. However, the method of manually editing and generating video tags for long videos will have the problems of large granularity of video tags and low efficiency of video tag generation, which will affect the user experience. Summary of the invention
[0003] In view of this, in order to solve the above-mentioned technical problems or part of the technical problems, the embodiments of the present application provide a method, device, electronic device and storage medium for generating long video tags.
[0004] In a first aspect, the present application provides a method for generating a long video tag, comprising:
[0005] Obtain a preset long video set and plot description information corresponding to each preset long video in the preset long video set;
[0006] For each of the preset long videos in the preset long video set, inputting the plot description information corresponding to the preset long video into a first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video;
[0007] According to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set, grouping all the preset long videos in the preset long video set to obtain respective long video groups;
[0008] Construct target prompt text;
[0009] For each of the long video groups, the target prompt text and the plot description information corresponding to each of the preset long videos in the long video group are input into the first large language model, so that the first large language model outputs the target video label corresponding to the long video group.
[0010] In an optional implementation, grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group includes:
[0011] Performing dimensionality reduction processing on the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain the processed first plot semantic vector corresponding to each of the preset long videos;
[0012] Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a first clustering algorithm to obtain a first clustering result, wherein the first clustering result includes a plurality of first clustering clusters;
[0013] For each of the first clustering clusters in the first clustering result, determining the preset long video corresponding to each processed first plot semantic vector in the first clustering cluster;
[0014] All the preset long videos corresponding to the first cluster are determined as a long video group.
[0015] In an optional implementation, the construction target prompt text includes:
[0016] Get the preset thought chain prompt template;
[0017] The target prompt text is constructed according to the preset thought chain prompt template.
[0018] In an optional implementation, before executing the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group, the method further includes:
[0019] For each of the preset long videos in the preset long video set, inputting the plot description information corresponding to the preset long video into a second large language model, so that the second large language model outputs a second plot semantic vector corresponding to the preset long video, and the second large language model is different from the first large language model;
[0020] For each of the preset long videos in the preset long video set, determining a vector similarity between the first plot semantic vector corresponding to the preset long video and the second plot semantic vector corresponding to the preset long video;
[0021] According to the vector similarities corresponding to all the preset long videos in the preset long video set, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is performed.
[0022] In an optional implementation, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group comprises:
[0023] Determine from the preset long video set all the preset long videos that meet a first preset condition, where the first preset condition includes that the vector similarity is greater than a preset similarity threshold;
[0024] Determine a first number of all the preset long videos that meet the first preset condition and a second number of all the preset long videos in the preset long video set;
[0025] determining a target ratio between the first number and the second number;
[0026] When the target ratio is greater than a preset ratio threshold, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is executed.
[0027] In an optional implementation, after executing the step of outputting the target video label corresponding to the long video group by the first large language model, the method further includes:
[0028] Selecting a plurality of target long videos from each of the obtained long video groups;
[0029] Pushing the target video tags corresponding to the multiple target long videos to the target terminal corresponding to the first large language model, so that the target terminal displays the target video tags corresponding to the multiple target long videos;
[0030] Receiving first feedback information returned by the target terminal;
[0031] When the first feedback information includes target modification information corresponding to the target parameter in the first large language model, modify the target parameter in the first large language model using the target modification information to obtain an updated first large language model;
[0032] Using the updated first large language model, the step of inputting the plot description information corresponding to the preset long video into the first large language model is performed, so that the first feedback information does not include the target modification information corresponding to the target parameter in the first large language model.
[0033] In an optional implementation, before performing the step of determining the preset long video corresponding to each processed first plot semantic vector in the first cluster, the method further includes:
[0034] Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a second clustering algorithm to obtain a second clustering result, wherein the second clustering result includes a plurality of second clustering clusters, and the second clustering algorithm is different from the first clustering algorithm;
[0035] Determining whether a plurality of the first clusters correspond to a plurality of the second clusters one-to-one;
[0036] When a plurality of the first clusters correspond one-to-one to a plurality of the second clusters, the step of determining the preset long video corresponding to each processed first plot semantic vector in the first cluster is performed.
[0037] In a second aspect, the present application provides a device for generating a long video tag, comprising:
[0038] An acquisition module, used to acquire a preset long video set and plot description information corresponding to each preset long video in the preset long video set;
[0039] A generating module, for inputting the plot description information corresponding to each of the preset long videos in the preset long video set into a first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video;
[0040] A grouping module, configured to group all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set, so as to obtain respective long video groups;
[0041] A construction module, used to construct a target prompt text;
[0042] The generation module is also used to input the target prompt text and the plot description information corresponding to each of the preset long videos in the long video group into the first large language model for each of the long video groups, so that the first large language model outputs the target video label corresponding to the long video group.
[0043] In a third aspect, the present application provides an electronic device, comprising: a processor and a memory, wherein the processor is used to execute a long video tag generation program stored in the memory to implement the long video tag generation method as described above.
[0044] In a fourth aspect, the present application further provides a storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the method for generating long video tags as described above.
[0045] The above-mentioned technical solution provided by the embodiment of the present application has the following advantages compared with the prior art. The method provided by the embodiment of the present application includes: obtaining a preset long video set and plot description information corresponding to each preset long video in the preset long video set; for each preset long video in the preset long video set, inputting the plot description information corresponding to the preset long video into the first largest language model, so that the first largest language model outputs a first plot semantic vector corresponding to the preset long video; according to the first plot semantic vector corresponding to each preset long video in the preset long video set, all preset long videos in the preset long video set are grouped to obtain each long video group; constructing a target prompt text; for each long video group, inputting the target prompt text and the plot description information corresponding to each preset long video in the long video group into the first largest language model, so that the first largest language model outputs a target video label corresponding to the long video group. Through the above method, the embodiment of the present application uses a large language model to generate a corresponding first plot semantic vector based on the plot description information corresponding to each preset long video in the preset long video set, so as to use all the first plot semantic vectors produced to realize the grouping of all the preset long videos in the preset long video set, so as to generate a video label that conforms to the plot theme for each preset long video in the long video group based on the constructed target prompt text and the plot description information corresponding to each preset long video in the long video group, which not only improves the generation efficiency of the video label of the long video but also makes the granularity of the video label much smaller than the granularity of the video label generated by manual editing for each long video group. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0048] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0049] Figure 1 A schematic diagram of a flow chart of a method for generating a long video tag provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of a flow chart of another method for generating a long video tag provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of the structure of a device for generating a long video tag provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0053] In the above attached figure:
[0054] 10. Acquisition module; 20. Generation module; 30. Grouping module; 40. Construction module;
[0055] 400, electronic device; 401, processor; 42, memory; 4021, operating system; 4022, application; 403, user interface; 404, network interface; 405, bus system. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0057] The disclosure below provides many different embodiments or examples to implement different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0058] refer to Figure 1 , Figure 1 A flow chart of a method for generating a long video tag provided in an embodiment of the present application. A method for generating a long video tag provided in this embodiment includes the following steps:
[0059] S101: Obtain a preset long video set and plot description information corresponding to each preset long video in the preset long video set.
[0060] In this embodiment, the preset long video set includes multiple preset long videos, and the preset long videos can be selected according to actual needs. In this embodiment, there is no specific limitation on the preset long videos. After determining that the preset long video set is obtained, the long video information corresponding to each preset long video in the preset long video set can be determined, and the long video information includes the cataloged long video title, description information, content tags and age background. The video content description of each preset long video in the preset long video set can also be obtained from off-site information. After obtaining the above information, the content understanding platform can be used to use a large language model to summarize and generate the plot description information corresponding to each preset long video in the preset long video set, thereby obtaining the plot description information corresponding to each long video in the preset long video set. The plot description information is used to characterize the relevant information describing the plot of the preset long video.
[0061] S102: For each preset long video in the preset long video set, input the plot description information corresponding to the preset long video into the first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video.
[0062] In this embodiment, the first language model can be selected according to actual needs. Since the xiaobu-embedding-v2 language model has good Chinese classification and clustering performance, the first language model in this embodiment adopts the xiaobu-embedding-v2 language model. In this embodiment, the xiaobu-embedding-v2 language model can map high-dimensional and discrete data (such as words, user IDs, etc.) to a low-dimensional continuous vector space, so that the similarity can be calculated between discrete data that cannot be directly calculated. Therefore, in this embodiment, after obtaining the plot description information corresponding to each preset long video in the preset long video set, the first language model is used to map the plot description information corresponding to each preset long video in the preset long video set, so that the first language model outputs the first plot semantic vector (embedding) corresponding to each preset long video, and then the similarity can be calculated according to the first plot semantic vector corresponding to each preset long video, so as to realize the grouping of all preset long videos in the preset long video set.
[0063] S103: Grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each preset long video in the preset long video set to obtain each long video group.
[0064] In this embodiment, after obtaining the first plot semantic vector corresponding to each preset long video in the preset long video set, all the obtained first plot semantic vectors can be used to determine the similarity between each preset long video in the preset long video set, and then all the preset long videos in the preset long video set are grouped using the determined similarities to divide the preset long video group into a plurality of long video groups, wherein each long video group includes at least one preset long video.
[0065] Specifically, when grouping all the preset long videos in the preset long video set, a clustering algorithm may be used.
[0066] S104: Construct target prompt text.
[0067] In this embodiment, the target prompt text is used to guide the first language model to generate video tags corresponding to each preset long video in the preset long video set. Among them, the target prompt text can be implemented through prompt. Specifically, after the target prompt text is constructed, the plot description information corresponding to all long videos under each long video group can be handed over to the first language model. After analyzing these plot description information, the first language model will extract the common points and core features therein to accurately generate the video tags corresponding to each long video under the long video group.
[0068] S105: For each long video group, the target prompt text and the plot description information corresponding to each preset long video in the long video group are input into the first large language model, so that the first large language model outputs the target video label corresponding to the long video group.
[0069] In this embodiment, after obtaining the constructed target prompt text, for each long video group, the target prompt text and the plot description information corresponding to all preset long videos under the long video group can be input into the first large language model, so that the first large language model, under the guidance of the target prompt text, analyzes the plot description information corresponding to all preset long videos, thereby obtaining the target video label corresponding to the long video group, and using the target video label as the video label corresponding to each preset long video under the long video group. For example, for the preset long video A, the video label generated by the manual editor is war, while the video label obtained by the method provided in this embodiment can be special operations. It can be seen that the granularity of the video label generated by the present embodiment is much smaller than the granularity of the video label generated by the manual editor.
[0070] The present embodiment provides a method for generating long video labels, which uses a large language model to generate corresponding first plot semantic vectors based on plot description information corresponding to each preset long video in a preset long video set, so as to group all preset long videos in the preset long video set using all the generated first plot semantic vectors, thereby for each long video group, using a large language model to generate video labels that conform to the plot theme for each preset long video in the long video group based on the constructed target prompt text and the plot description information corresponding to each preset long video in the long video group, which not only improves the generation efficiency of video labels for long videos but also makes the granularity of video labels much smaller than the granularity of video labels generated by manual editing.
[0071] refer to Figure 2 , Figure 2 A flowchart of another method for generating a long video tag provided in an embodiment of the present application. A method for generating a long video tag provided in this embodiment includes the following steps:
[0072] S201: Obtain a preset long video set and plot description information corresponding to each preset long video in the preset long video set.
[0073] S202: For each preset long video in the preset long video set, input the plot description information corresponding to the preset long video into the first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video.
[0074] Regarding the above-mentioned step S201 and step S202, step S201 is consistent with the above-mentioned step S101, and step S202 is consistent with the above-mentioned step S102. For details, please refer to the above-mentioned step S101 and step S102, which will not be described in detail in this embodiment.
[0075] In this embodiment, before executing step S203, in order to ensure the accuracy of each long video group obtained subsequently and the accuracy of the obtained video tags, a method for generating a long video tag provided by this embodiment further includes the following steps:
[0076] For each preset long video in the preset long video set, inputting plot description information corresponding to the preset long video into the second largest language model, so that the second largest language model outputs a second plot semantic vector corresponding to the preset long video;
[0077] For each preset long video in the preset long video set, determining a vector similarity between a first plot semantic vector corresponding to the preset long video and a second plot semantic vector corresponding to the preset long video;
[0078] According to the vector similarities corresponding to all the preset long videos in the preset long video set, the following step S203 is performed.
[0079] In the above, the second largest language model is different from the first largest language model, and the second largest language model can be selected according to actual needs. In this embodiment, the second largest language model is not specifically limited. In order to avoid the problem of low accuracy when only the first largest language model is used to determine the target video label corresponding to the preset long video, after the first largest language model is used to determine the first plot semantic vector corresponding to each preset long video in the preset long video set, the plot description information corresponding to each preset long video in the preset long video set is input into the second largest language model, so that the second largest language model outputs the second plot semantic vector corresponding to each preset long video. For each preset long video, the vector similarity between the first plot semantic vector corresponding to the preset long video and the second plot semantic vector corresponding to the preset long video can be calculated using the existing distance calculation algorithm, so as to determine whether the first largest language model is accurate by using the vector similarity corresponding to all preset long videos in the preset long video set. If the first largest language model is determined to be accurate, the following S203 step is performed.
[0080] The distance calculation algorithm may be selected according to actual needs, and the distance calculation algorithm is not specifically limited in this embodiment.
[0081] In this embodiment, according to the vector similarities corresponding to all the preset long videos in the preset long video set, the following step S203 is performed, which specifically includes:
[0082] Determine all preset long videos that meet a first preset condition from a preset long video set;
[0083] Determine a first number of all preset long videos that meet a first preset condition and a second number of all long videos in a preset long video set;
[0084] determining a target ratio between the first number and the second number;
[0085] When the target ratio is greater than the preset ratio threshold, the following step S203 is executed.
[0086] In the above, the first preset condition includes that the vector similarity is greater than the preset similarity threshold. The vector similarity threshold represents the lower limit of the vector similarity accurately corresponding to the first largest language model. The vector similarity threshold can be set according to actual needs, and the specific value of the vector similarity is not limited in this embodiment. The preset ratio threshold represents the minimum value of the target ratio accurately corresponding to the first largest language model. The preset ratio threshold can be set according to actual needs, and the specific value of the preset ratio threshold is not limited in this embodiment.
[0087] Among them, in order to determine whether the selected first largest language model is accurate, after obtaining the vector similarity corresponding to each preset long video in the preset long video set, all preset long videos in the preset long video set whose vector similarity is greater than the preset similarity threshold are determined. If the target ratio between the first number of all the determined preset long videos and the second number of all the preset long videos in the preset long video set is greater than the preset ratio threshold, it indicates that the first largest language model is accurate, and at this time, the following S203 step can be directly executed.
[0088] Specifically, if the target ratio between the first number of all preset long videos determined and the second number of all preset long videos in the preset long video set is less than or equal to the preset ratio threshold, it indicates that the first language model may be inaccurate. Therefore, in order to ensure the accuracy of the first language model, the target ratio can be pushed to the target terminal corresponding to the first language model, so that the target terminal displays the target ratio. After viewing the target ratio, the user at the target terminal can choose to use the first language model or the second language model to perform subsequent steps, and feedback the second feedback information through the target terminal. After receiving the second feedback information, if the second feedback information is the target feedback information, the first language model is used to perform the subsequent S203 step; if the second feedback information is not the target feedback information, the second language model is updated to the first language model to perform the subsequent S203 step. Specifically, the target feedback information is to use the first language model to determine the video tag corresponding to the long video. It should be noted that the specific form of the target feedback information can be selected according to actual needs, and the specific form of the target feedback information is not limited in this embodiment. The target terminal may be a mobile phone, a computer, etc. The specific form of the target terminal may be selected according to actual needs. The specific form of the target terminal is not limited in this embodiment.
[0089] More specifically, if the target ratio between the first number of all determined preset long videos and the second number of all preset long videos in the preset long video set is less than or equal to the preset ratio threshold, the first plot semantic vector and the second plot semantic vector corresponding to all preset long videos in the preset long video set that meet the first preset condition are used to fine-tune the first language model to obtain an updated first language model, and the updated first language model is used to perform step S202 to ensure the accuracy of the first plot semantic vector corresponding to the preset long video output by the first language model, thereby ensuring the accuracy of the video label corresponding to the obtained long video.
[0090] S203: Performing dimensionality reduction processing on the first plot semantic vector corresponding to each preset long video in the preset long video set to obtain a processed first plot semantic vector corresponding to each preset long video.
[0091] S204: Clustering the processed first plot semantic vectors corresponding to each preset long video in the preset long video set using a first clustering algorithm to obtain a first clustering result.
[0092] S205: For each first clustering cluster in the first clustering result, determine a preset long video corresponding to each processed first plot semantic vector in the first clustering cluster.
[0093] S206: Determine all preset long videos corresponding to the first cluster as a long video group.
[0094] For the above-mentioned steps S203 to S206, the first clustering result includes multiple first clustering clusters. The first clustering algorithm is the KMeans clustering algorithm. The KMeans clustering algorithm is a distance-based partitioning clustering algorithm that determines how to cluster similar content based on a predetermined number of clusters K. In the distribution scenario of long videos, it is necessary to control the number of video groups and the number of video contents in each group (i.e., granularity). Since the KMeans clustering algorithm has more advantages in controlling the number of each group, the first clustering algorithm in this embodiment adopts the KMeans clustering algorithm.
[0095] Among them, after generating the first plot semantic vector corresponding to each preset long video in the preset long video set, each first plot semantic vector is a high-dimensional dense vector, but the entire semantic space will be very sparse, and the plot description information logically related to the semantics will be gathered in the space because of the close distance. The KMeans clustering algorithm is sensitive to the vector dimension. The higher the dimension, the slower the speed, so it is necessary to compress the original high-dimensional vector and reduce the dimension. The dimension reduction method can be selected according to actual needs. In this embodiment, the dimension reduction method is not specifically limited. For example, the dimension reduction method can select the UMAP (Uniform Manifold Approximation and Projection) dimension reduction method. By using the first plot semantic vector generated by the KMeans clustering algorithm in combination with the first language model, similar plot contents can be effectively clustered together to obtain each long video group, and each long video group generates a corresponding target video tag, which not only helps to improve the accuracy of generating plot tags, but also makes subsequent video recommendations and content searches more efficient and accurate.
[0096] In this embodiment, in order to ensure the accuracy of clustering using the first clustering algorithm, thereby improving the accuracy of the subsequently determined video tags, this embodiment further includes the following steps before executing step S205:
[0097] Clustering the processed first plot semantic vectors corresponding to each preset long video in the preset long video set using a second clustering algorithm to obtain a second clustering result, the second clustering result including a plurality of second clustering clusters, the second clustering algorithm being different from the first clustering algorithm;
[0098] Determining whether the plurality of first clusters correspond to the plurality of second clusters;
[0099] When the plurality of first clusters correspond one-to-one to the plurality of second clusters, step S205 is executed.
[0100] Among them, the second clustering algorithm can be selected according to actual needs. In this embodiment, the second clustering algorithm is not specifically limited. For example, the second clustering algorithm can be a DBSCAN clustering algorithm. In order to ensure the accuracy of clustering using the first clustering algorithm, after obtaining the first clustering result using the first clustering algorithm, the second clustering algorithm can be used to cluster the processed first plot semantic vectors corresponding to each preset long video in the preset long video set to obtain a second clustering result. If the first clustering cluster in the first clustering result corresponds to the second clustering result in the second clustering result, it indicates that the first clustering result obtained by the first clustering algorithm is accurate, and step S205 can be executed at this time. If there is a first clustering cluster in the first clustering result that is inconsistent with the second clustering cluster in the second clustering result, it indicates that the first clustering result obtained by the first clustering algorithm may be inaccurate. At this time, the first clustering result and the second clustering result are pushed to the target terminal corresponding to the first large language model, so that the target terminal displays the first clustering result and the second clustering result. The user of the target terminal can modify the first clustering result according to the displayed first clustering result and the second clustering result, and feedback the modified first clustering result. The modified first clustering result fed back by the target terminal is received, and the subsequent S205 step is performed using the modified first clustering result, so as to improve the accuracy of the subsequently determined video tags.
[0101] S207: Obtain a preset thought chain prompt template.
[0102] S208: Construct a target prompt text according to a preset thinking chain prompt template.
[0103] For the above-mentioned steps S207 and S208, in order to improve the accuracy of the determined video tags, a preset thought chain prompt template can be used when the target prompt text is provided. The preset thought chain prompt template can be set according to actual needs, and the preset thought chain prompt template is not limited in this embodiment.
[0104] S209: For each long video group, the target prompt text and the plot description information corresponding to each preset long video in the long video group are input into the first largest language model, so that the first largest language model outputs the target video label corresponding to the long video group.
[0105] In this embodiment, step S209 is consistent with the above-mentioned step S105. For details, please refer to the above-mentioned step S105, and this embodiment will not be described in detail here.
[0106] In this embodiment, after executing step S209, the method for generating a long video tag provided by this embodiment further includes the following steps:
[0107] Selecting multiple target long videos from each of the obtained long video groups;
[0108] Pushing target video tags corresponding to the multiple target long videos to the target terminal corresponding to the first largest language model, so that the target terminal displays the target video tags corresponding to the multiple target long videos;
[0109] Receiving first feedback information returned by the target terminal;
[0110] When the first feedback information includes target modification information corresponding to the target parameter in the first large language model, modify the target parameter in the first large language model using the target modification information to obtain an updated first large language model;
[0111] Using the updated first language model, the step of inputting the plot description information corresponding to the preset long video into the first language model is performed, so that the first feedback information does not include the target modification information corresponding to the target parameter in the first language model.
[0112] Among them, in order to verify whether the target video tag corresponding to the determined preset long video is accurate, after obtaining the target video tag corresponding to each preset long video in each long video group, multiple target long videos can be randomly selected from each long video group, and the target video tags corresponding to each selected target long video can be pushed to the target terminal corresponding to the first large language model, so that the target terminal displays the target video tags corresponding to the multiple target long videos. After the target terminal displays the target video tags corresponding to the multiple target long videos, the user at the target terminal can modify the target parameters in the first large language model based on the target video tags corresponding to the multiple target long videos displayed by the target terminal, and the target terminal returns the first feedback information. After receiving the first feedback information returned by the target terminal, if the first feedback information includes the target modification information corresponding to the target parameter in the first large language model, it indicates that the user needs to modify the first large language model. At this time, the target modification information is used to modify the target parameter in the first large language model to achieve the update of the first large language model, so as to use the updated first large language model to perform the step of inputting the plot description information corresponding to the preset long video into the first large language model, until the first feedback information does not include the target modification information corresponding to the target parameter in the first large language model. If the first feedback information does not include the target modification information corresponding to the target parameter in the first large language model, it indicates that the target video label corresponding to the determined preset long video is accurate.
[0113] The present embodiment provides a method for generating long video labels, which uses a large language model to generate corresponding first plot semantic vectors based on plot description information corresponding to each preset long video in a preset long video set, so as to group all preset long videos in the preset long video set using all the generated first plot semantic vectors, thereby for each long video group, using a large language model to generate video labels that conform to the plot theme for each preset long video in the long video group based on the constructed target prompt text and the plot description information corresponding to each preset long video in the long video group, which not only improves the generation efficiency of video labels for long videos but also makes the granularity of video labels much smaller than the granularity of video labels generated by manual editing.
[0114] refer to Figure 3 , Figure 3A schematic diagram of a structure of a device for generating a long video tag provided in an embodiment of the present application. A device for generating a long video tag provided in an embodiment of the present application includes: an acquisition module 10, a generation module 20, a grouping module 30 and a construction module 40. An acquisition module 10 is used to acquire a preset long video set and plot description information corresponding to each preset long video in the preset long video set; a generation module 20 is used to input the plot description information corresponding to the preset long video into a first large language model for each preset long video in the preset long video set, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video; a grouping module 30 is used to group all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each preset long video in the preset long video set to obtain each long video group; a construction module 40 is used to construct a target prompt text; the generation module 20 is also used to input the target prompt text and the plot description information corresponding to each preset long video in the long video group into the first large language model for each long video group, so that the first large language model outputs a target video label corresponding to the long video group.
[0115] In this embodiment, the grouping module 30 is further used for:
[0116] Performing dimensionality reduction processing on the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain the processed first plot semantic vector corresponding to each of the preset long videos;
[0117] Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a first clustering algorithm to obtain a first clustering result, wherein the first clustering result includes a plurality of first clustering clusters;
[0118] For each of the first clustering clusters in the first clustering result, determining the preset long video corresponding to each processed first plot semantic vector in the first clustering cluster;
[0119] All the preset long videos corresponding to the first cluster are determined as a long video group.
[0120] In this embodiment, the building block 40 is further used for:
[0121] Get the preset thought chain prompt template;
[0122] The target prompt text is constructed according to the preset thought chain prompt template.
[0123] In this embodiment, the grouping module 30 is further used for:
[0124] For each of the preset long videos in the preset long video set, inputting the plot description information corresponding to the preset long video into a second large language model, so that the second large language model outputs a second plot semantic vector corresponding to the preset long video, and the second large language model is different from the first large language model;
[0125] For each of the preset long videos in the preset long video set, determining a vector similarity between the first plot semantic vector corresponding to the preset long video and the second plot semantic vector corresponding to the preset long video;
[0126] According to the vector similarities corresponding to all the preset long videos in the preset long video set, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is performed.
[0127] In this embodiment, the grouping module 30 is further used for:
[0128] Determine from the preset long video set all the preset long videos that meet a first preset condition, where the first preset condition includes that the vector similarity is greater than a preset similarity threshold;
[0129] Determine a first number of all the preset long videos that meet the first preset condition and a second number of all the preset long videos in the preset long video set;
[0130] determining a target ratio between the first number and the second number;
[0131] When the target ratio is greater than a preset ratio threshold, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is executed.
[0132] In this embodiment, the generating module 20 is further used for:
[0133] Selecting a plurality of target long videos from each of the obtained long video groups;
[0134] Pushing the target video tags corresponding to the multiple target long videos to the target terminal corresponding to the first large language model, so that the target terminal displays the target video tags corresponding to the multiple target long videos;
[0135] Receiving first feedback information returned by the target terminal;
[0136] When the first feedback information includes target modification information corresponding to the target parameter in the first large language model, modify the target parameter in the first large language model using the target modification information to obtain an updated first large language model;
[0137] Using the updated first large language model, the step of inputting the plot description information corresponding to the preset long video into the first large language model is performed, so that the first feedback information does not include the target modification information corresponding to the target parameter in the first large language model.
[0138] In this embodiment, the grouping module 30 is further used for:
[0139] Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a second clustering algorithm to obtain a second clustering result, wherein the second clustering result includes a plurality of second clustering clusters, and the second clustering algorithm is different from the first clustering algorithm;
[0140] Determining whether a plurality of the first clusters correspond to a plurality of the second clusters one-to-one;
[0141] When a plurality of the first clusters correspond one-to-one to a plurality of the second clusters, the step of determining the preset long video corresponding to each processed first plot semantic vector in the first cluster is performed.
[0142] The present embodiment provides a device for generating long video labels, which uses a large language model to generate corresponding first plot semantic vectors based on plot description information corresponding to each preset long video in a preset long video set, so as to group all preset long videos in the preset long video set using all the generated first plot semantic vectors, thereby generating, for each long video group, a video label that conforms to the plot theme for each preset long video in the long video group based on the constructed target prompt text and the plot description information corresponding to each preset long video in the long video group using the large language model, which not only improves the generation efficiency of video labels for long videos but also makes the granularity of video labels much smaller than the granularity of video labels generated by manual editing.
[0143] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4The electronic device 400 shown includes: at least one processor 401, a memory 402, at least one network interface 404 and other user interfaces 403. The various components in the electronic device 400 are coupled together via a bus system 405. It is understood that the bus system 405 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 405 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 405 is not shown in FIG. Figure 4 Various buses are labeled as bus system 405 .
[0144] The user interface 403 may include a display, a keyboard, or a pointing device (eg, a mouse, a trackball, a touch pad, or a touch screen).
[0145] It can be understood that the memory 402 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 602 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0146] In some implementations, the memory 402 stores the following elements, executable units or data structures, or a subset thereof, or an extended set thereof: an operating system 4021 and application programs 4022 .
[0147] Among them, the operating system 4021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application 4022 includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present application can be included in the application 4022.
[0148] In the embodiment of the present application, by calling the program or instructions stored in the memory 402, specifically, the program or instructions stored in the application 4022, the processor 401 is used to execute the method steps provided by each method embodiment.
[0149] The method disclosed in the above embodiment of the present application can be applied to the processor 401, or implemented by the processor 401. The processor 401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 401 or the instruction in the form of software. The above processor 401 can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software units in the decoding processor can be executed. The software unit can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 402, and the processor 401 reads the information in the memory 402 and completes the steps of the above method in combination with its hardware.
[0150] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPDevice, DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application, or a combination thereof.
[0151] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0152] The electronic device provided in this embodiment may be Figure 4 The electronic device shown in FIG. 1 may perform the following steps: Figure 1 and Figure 2 All steps of the method for generating medium and long video labels, and then realize Figure 1 and Figure 2 For details, please refer to the technical effect of the method for generating long video tags shown in Figure 1 and Figure 2 For the sake of brevity, the relevant description is not repeated here.
[0153] The embodiment of the present application also provides a storage medium (computer-readable storage medium). The storage medium here stores one or more programs. The storage medium may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk or a solid-state drive; the memory may also include a combination of the above-mentioned types of memory.
[0154] When one or more programs in the storage medium can be executed by one or more processors, the long video tag generation method executed on the long video tag generation processing device side is implemented.
[0155] The processor is used to execute the long video tag generation program stored in the memory to implement the following steps of the long video tag generation method executed on the long video tag generation device side.
[0156] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0157] It should be noted that the phrases "one implementation", "an embodiment", "an exemplary embodiment", "some embodiments", etc. mentioned in the specification indicate that the described embodiments may include certain features, structures or characteristics, but not every embodiment may include the certain features, structures or characteristics. In addition, such phrases do not necessarily refer to the same embodiment. In addition, when describing certain features, structures or characteristics in conjunction with an embodiment, it is within the knowledge of those skilled in the art to implement such features, structures or characteristics in conjunction with other embodiments, whether explicitly or not explicitly described.
[0158] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating a long video tag, characterized in that: include: Obtain a preset long video set and plot description information corresponding to each preset long video in the preset long video set; For each of the preset long videos in the preset long video set, inputting the plot description information corresponding to the preset long video into a first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video; According to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set, grouping all the preset long videos in the preset long video set to obtain respective long video groups; Construct target prompt text; For each of the long video groups, the target prompt text and the plot description information corresponding to each of the preset long videos in the long video group are input into the first large language model, so that the first large language model outputs the target video label corresponding to the long video group.
2. The method according to claim 1, characterized in that The step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group includes: Performing dimensionality reduction processing on the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain the processed first plot semantic vector corresponding to each of the preset long videos; Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a first clustering algorithm to obtain a first clustering result, wherein the first clustering result includes a plurality of first clustering clusters; For each of the first clustering clusters in the first clustering result, determining the preset long video corresponding to each processed first plot semantic vector in the first clustering cluster; All the preset long videos corresponding to the first cluster are determined as a long video group.
3. The method according to claim 1, characterized in that: The build target prompt text includes: Get the preset thought chain prompt template; The target prompt text is constructed according to the preset thought chain prompt template.
4. The method according to claim 1, characterized in that Before executing the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group, the method further includes: For each of the preset long videos in the preset long video set, inputting the plot description information corresponding to the preset long video into a second large language model, so that the second large language model outputs a second plot semantic vector corresponding to the preset long video, and the second large language model is different from the first large language model; For each of the preset long videos in the preset long video set, determining a vector similarity between the first plot semantic vector corresponding to the preset long video and the second plot semantic vector corresponding to the preset long video; According to the vector similarities corresponding to all the preset long videos in the preset long video set, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is performed.
5. The method according to claim 4, characterized in that The step of grouping all the preset long videos in the preset long video set according to the vector similarities corresponding to all the preset long videos in the preset long video set to obtain each long video group comprises: Determine from the preset long video set all the preset long videos that meet a first preset condition, where the first preset condition includes that the vector similarity is greater than a preset similarity threshold; Determine a first number of all the preset long videos that meet the first preset condition and a second number of all the preset long videos in the preset long video set; determining a target ratio between the first number and the second number; When the target ratio is greater than a preset ratio threshold, the step of grouping all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set to obtain each long video group is executed.
6. The method according to claim 1, characterized in that After executing the step of outputting the target video label corresponding to the long video group by the first large language model, the method further includes: Selecting a plurality of target long videos from each of the obtained long video groups; Pushing the target video tags corresponding to the multiple target long videos to the target terminal corresponding to the first large language model, so that the target terminal displays the target video tags corresponding to the multiple target long videos; Receiving first feedback information returned by the target terminal; When the first feedback information includes target modification information corresponding to the target parameter in the first large language model, modify the target parameter in the first large language model using the target modification information to obtain an updated first large language model; Using the updated first large language model, the step of inputting the plot description information corresponding to the preset long video into the first large language model is performed, so that the first feedback information does not include the target modification information corresponding to the target parameter in the first large language model.
7. The method according to claim 2, characterized in that: Before executing the step of determining the preset long video corresponding to each processed first plot semantic vector in the first cluster, the method further includes: Clustering the processed first plot semantic vectors corresponding to each of the preset long videos in the preset long video set using a second clustering algorithm to obtain a second clustering result, wherein the second clustering result includes a plurality of second clustering clusters, and the second clustering algorithm is different from the first clustering algorithm; Determining whether a plurality of the first clusters correspond to a plurality of the second clusters one-to-one; When a plurality of the first clusters correspond one-to-one to a plurality of the second clusters, the step of determining the preset long video corresponding to each processed first plot semantic vector in the first cluster is performed.
8. A device for generating a long video tag, characterized in that: include: An acquisition module, used to acquire a preset long video set and plot description information corresponding to each preset long video in the preset long video set; A generating module, for inputting the plot description information corresponding to each of the preset long videos in the preset long video set into a first large language model, so that the first large language model outputs a first plot semantic vector corresponding to the preset long video; A grouping module, configured to group all the preset long videos in the preset long video set according to the first plot semantic vector corresponding to each of the preset long videos in the preset long video set, so as to obtain respective long video groups; A construction module, used to construct a target prompt text; The generation module is also used to input the target prompt text and the plot description information corresponding to each of the preset long videos in the long video group into the first large language model for each of the long video groups, so that the first large language model outputs the target video label corresponding to the long video group.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute a long video tag generation program stored in the memory to implement the long video tag generation method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method for generating a long video tag according to any one of claims 1 to 7.
Citation Information
Cited By
Long video structured label generation method and system
CN120340016A
A method and system for generating structured tags for long videos
CN120340016B