A short video optimization method and device based on a knowledge graph, a computer device, and a readable storage medium
By obtaining the basic information of short videos and using the Internet celebrity language model and knowledge graph to perform in-depth information extraction and generate optimization suggestions, the problem of lack of systematicness in short video optimization methods in existing technologies is solved, and the dissemination effect of short videos is improved.
Patent Information
- Application Number
- CN202411639917.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing short video optimization methods lack systematicness and comprehensiveness, making it difficult to fully tap the potential of short videos, resulting in poor communication effects.
By obtaining the link to the target short video, we call the influencer big data service to obtain basic information, and use the influencer big language model and influencer knowledge graph to extract information and generate optimization suggestions, including new titles, new tags and new lines.
It effectively improved the dissemination of short videos and increased key metrics such as views, likes, and favorites.
Smart Images

Figure CN119166902B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a short video optimization method and device based on a knowledge graph, a computer device and a readable storage medium. BACKGROUND
[0002] With the rapid development of the short video industry, the number of short videos is growing explosively. How to optimize short video content to improve its dissemination effect has become an important problem. Existing optimization methods often lack systematicness and comprehensiveness, making it difficult to fully tap the potential of short videos. SUMMARY
[0003] The purpose of the present application is to provide a short video optimization method and device based on a knowledge graph, a computer device and a readable storage medium.
[0004] In a first aspect, the present application provides a short video optimization method based on a knowledge graph, comprising:
[0005] Obtaining a target short video link, the target short video link pointing to a target short video;
[0006] Calling a preset influencer big data service to obtain basic information of the target short video;
[0007] Based on the basic information and a preset influencer knowledge graph, calling a preset influencer big language model to extract information from the target short video to obtain deep information;
[0008] Inputting the basic information and the deep information into the influencer big language model to obtain optimization suggestions for the target short video.
[0009] In a possible implementation, the calling of the preset influencer big data service to obtain the basic information of the target short video comprises:
[0010] Calling a preset influencer big data service to obtain effect data, basic content and background information of the target short video; the effect data includes the listing time, the number of plays, the number of likes, the number of comments and the number of collections of the target short video; the basic content includes the title, the label and the audio of the target short video; the background information includes the creator of the target short video, the head competitor video and the creator of the head competitor video;
[0011] The effect data, the basic content and the background information are taken as the basic information of the target short video.
[0012] In a possible implementation, the calling of the preset influencer big language model to extract information from the target short video based on the basic information comprises:
[0013] transcribe the audio into a script by invoking a preset tool set through the web celebrity large language model;
[0014] analyze the title, the label, and the script through the web celebrity large language model to obtain a theme of the target short video;
[0015] obtain the deep information from the preset web celebrity knowledge graph based on the effect data, the basic content, the background information, the script, and the theme through the web celebrity large language model.
[0016] In a possible implementation, the preset web celebrity knowledge graph is constructed by the following method, comprising:
[0017] obtain video itself information and account information of a plurality of stored videos stored by the preset web celebrity big data service through the web celebrity large language model;
[0018] structure the video itself information and the account information of each stored video to obtain a text block corresponding to each stored video, wherein the text block comprises the video itself information and the account information;
[0019] input a plurality of text blocks into the web celebrity large language model for multi-round extraction to obtain entities, relationships, and covariants corresponding to each plurality of text blocks;
[0020] perform multi-layer cluster extraction according to the entities, the relationships, and the covariants to obtain the preset web celebrity knowledge graph, wherein the preset web celebrity knowledge graph comprises a four-layer graph structure constructed from top to bottom of a large category cluster, a subcategory cluster, an attitude cluster, and a theme cluster.
[0021] In a possible implementation, the obtaining the deep information from the preset web celebrity knowledge graph based on the effect data, the basic content, the background information, the script, and the theme through the web celebrity large language model comprises:
[0022] perform large category cluster matching, subcategory cluster matching, attitude cluster matching, and theme cluster matching in the preset web celebrity knowledge graph based on the effect data, the basic content, the background information, the script, and the theme through the web celebrity large language model, and obtain target entities and target text blocks according to a preset relevance threshold;
[0023] take the target entities and the target text blocks as the deep information.
[0024] In a possible implementation, the preset influencer large language model is obtained by the following method, comprising:
[0025] obtaining a basic large language model and an advanced large language model, the model parameter quantity of the advanced large language model being greater than that of the basic large language model;
[0026] performing analysis, self-verification and self-picking extraction strategies on sample short videos based on a self-discovery mechanism through the advanced large language model, and generating knowledge extraction data through prompting engineering;
[0027] fine-tuning the basic large language model based on the knowledge extraction data in a LoRA manner to obtain the preset influencer large language model.
[0028] In a possible implementation, the method further comprises:
[0029] inputting the basic information and the depth information into the influencer large language model to obtain optimization suggestions for the target short video, the optimization suggestions comprising a new title, a new label and a new catchphrase.
[0030] In a second aspect, an embodiment of the present application provides a short video optimization device based on a knowledge graph, comprising:
[0031] an acquisition module configured to acquire a target short video link, the target short video link pointing to a target short video; call a preset influencer big data service to acquire basic information of the target short video; and call a preset influencer large language model to perform information extraction on the target short video based on the basic information and a preset influencer knowledge graph to obtain depth information;
[0032] an optimization module configured to input the basic information and the depth information into the influencer large language model to obtain optimization suggestions for the target short video.
[0033] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein the computer instructions are executed by the processor to cause the computer device to perform the method of the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides a readable storage medium, comprising a computer program, wherein the computer program is executed to control a computer device on which the readable storage medium is located to perform the method of the first aspect.
[0035] Compared with the prior art, the application provides the beneficial effects including: by adopting the short video optimization method, device, computer equipment and readable storage medium based on a knowledge graph, the target short video link and basic information are acquired, and a net red big data service is called to realize. Then, based on the basic information and the net red knowledge graph, deep information is extracted through a net red big language model. Finally, the basic and deep information are input into the model to obtain optimization suggestions. In this way, various technical resources are used to effectively mine the value of short videos to improve the transmission effect. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0037] Figure 1 The step flow diagram of the short video optimization method based on a knowledge graph provided by the embodiments of the present application;
[0038] Figure 2 The short video optimization system framework provided by the embodiments of the present application is shown in the schematic diagram;
[0039] Figure 3 The Prompt framework flow provided by the embodiments of the present application is shown in the schematic diagram;
[0040] Figure 4 The entity multi-round extraction provided by the embodiments of the present application is shown in the schematic diagram;
[0041] Figure 5 The multi-layer cluster extraction provided by the embodiments of the present application is shown in the schematic diagram;
[0042] Figure 6 The structure schematic diagram of the short video optimization device based on a knowledge graph provided by the embodiments of the present application is shown in the schematic diagram;
[0043] Figure 7 The structure schematic diagram of the computer equipment provided by the embodiments of the present application is shown in the schematic diagram. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0045] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0046] To solve the technical problems in the foregoing background art, Figure 1 The flowchart of the short video optimization method based on a knowledge graph provided by the embodiments of the present application is shown below, and the short video optimization method based on a knowledge graph is described in detail.
[0047] Step S201, a target short video link is obtained, the target short video link pointing to a target short video;
[0048] Step S202, a preset KOL big data service is called to obtain basic information of the target short video;
[0049] Step S203, based on the basic information and a preset KOL knowledge graph, a preset KOL big language model is called to perform information extraction on the target short video to obtain deep information;
[0050] Step S204, the basic information and the deep information are input into the KOL big language model to obtain optimization suggestions for the target short video.
[0051] In the embodiments of the present application, it is assumed that the server receives a request from a client (such as a content optimization platform used by a certain short video operator), and the request contains a short video link on the TikTok platform, for example, “https: / / www.tiktok.com / video / 123456789”. This link is the target short video link, which explicitly points to a specific short video. The server identifies and stores this link as a basis for subsequent operations.
[0052] After the server receives the target short video link, it starts to call the preset KOL big data service (KOL_BigData). For the acquisition of effect data, the server sends a query request to KOL_BigData, and KOL_BigData searches the vast video database (storing hundreds of millions of video data) for the record corresponding to the target short video link. For example, it is found that this short video was uploaded on May 1, 2023, and currently has a play count of 100,000, a like count of 5,000, a comment count of 1,000, and a collection count of 2,000.
[0053] For basic content, KOL_BigData returns that the title of the target short video is “Perfect Place for Summer Travel”, the tags are “Travel” “Summer” “Tourist Attractions”, and the audio is a cheerful background music plus some natural environmental sound effects.
[0054] For background information, the query found that the creator is a user named "Tourism Expert", and the head competitor videos are several popular short videos with the same travel theme. The creators of these head competitor videos are "Tourism Enthusiast" and "World Traveler".
[0055] The server integrates the effect data it has obtained (listing time (May 1, 2023), play count (100,000 times), like count (5,000 times), comment count (1,000 times), and collection count (2,000 times)), the basic content (title ("Perfect Summer Travel Destination"), tags ("Travel", "Summer", "Tourist Attraction"), and audio (upbeat background music + natural sound effects)), and the background information (creator ("Tourism Expert"), head competitor videos (competitor video links and related information), and head competitor video creators ("Tourism Enthusiast" and "World Traveler")) to form a basic information package for the target short video. This basic information package will be used for further analysis and processing in subsequent steps.
[0056] The server calls the preset KOL_LLM (KOL_LLM), which is based on the Llama3-70B 1M context window version fine-tuning. After receiving the basic information, KOL_LLM calls the transcription function in its toolset to transcribe the audio of the target short video. For example, the natural environmental sound effects in the audio are ignored, the background music is identified as atmosphere, and the commentary of the characters is transcribed as "Welcome to this charming summer travel destination, where there are blue seas, blue skies, golden beaches, and rich food" and other lines.
[0057] KOL_LLM begins to analyze the title "Perfect Summer Travel Destination", the tags "Travel", "Summer", and "Tourist Attraction", and the transcribed lines "Welcome to this charming summer travel destination, where there are blue seas, blue skies, golden beaches, and rich food". Through semantic analysis of these contents, KOL_LLM determines that the theme of this short video is "Summer Beach Travel Recommendation".
[0058] First, KOL_LLM operates in the preset KOL_Graph (KOL_Graph) based on all the information mentioned above (effect data, basic content, background information, lines, and theme). KOL_Graph is based on tens of millions of short videos.
[0059] For the big category cluster matching, since the topic is "summer beach travel recommendation", KOL_LLM matches it with the big category clusters in KOL_Graph, and it can be matched to the "travel" big category cluster. Then the sub-category cluster matching is performed under the "travel" big category cluster, and it can be matched to the "beach travel" sub-category cluster. Then the attitude cluster matching is performed, such as matching to the "full of passion for beach travel recommendation" attitude cluster. Finally, the topic cluster matching is performed, and the most relevant topic cluster to "summer beach travel recommendation" is found from the attitude cluster, such as the topic cluster containing specific beach locations, local food, and other related entities.
[0060] Through these matches, according to the preset relevance threshold, KOL_LLM obtains target entities such as the specific beach name "Bali Beach", the local specialty food "seafood barbecue", and related text blocks such as text blocks containing other user reviews of this beach from the topic cluster. These target entities and target text blocks constitute the depth information.
[0061] The server inputs the previously integrated basic information (including various information such as listing time, play count, etc.) and depth information (such as entities such as "Bali Beach" and "seafood barbecue" and related text blocks) into KOL_LLM. KOL_LLM analyzes and processes these information.
[0062] For example, the optimization suggestions that KOL_LLM can give include: the new title is "Bali Beach: the dream of summer travel, enjoy seafood barbecue", the new label increases more specific labels such as "Bali" and "seafood", and some descriptions about the unique features of Bali Beach can be added to the new catchphrase, such as "the sunset on Bali Beach is one of the most beautiful sunsets in the world, with fresh seafood barbecue, it is the ultimate experience of summer travel". These optimization suggestions will help improve the attractiveness of the short video, which may improve the play count, like count, and other indicators.
[0063] In the embodiment of the application, the calling of the preset KOL big data service to obtain the basic information of the target short video can be implemented by the following example.
[0064] The preset KOL big data service is called to obtain the effect data, basic content, and background information of the target short video. The effect data includes the listing time, play count, like count, comment count, and collection count of the target short video. The basic content includes the title, label, and audio of the target short video. The background information includes the creator of the target short video, the head competitor video, and the creator of the head competitor video.
[0065] The effect data, the basic content, and the background information are taken as the basic information of the target short video.
[0066] In an embodiment of the present application, after the server receives the link of the target short video, it starts to call the preset KOL big data service (KOL_BigData). The server sends a query request containing the target short video identifier (such as the unique identifier or link of the video) to KOL_BigData, and specifically queries the effect data part.
[0067] The video library of KOL_BigData stores a large amount of video information. After receiving the query request from the server, it quickly locates the record of the target short video in its data storage structure. For example, for a short video of food making, KOL_BigData queries that the video was uploaded on March 15, 2023. This date information is very important for analyzing the timeliness of the video and its propagation in different time periods.
[0068] Next, the play count data is queried and it is found that the video has currently had 80,000 plays. Play count is a key indicator of the popularity of a video, reflecting the spread of the video among the audience. The like count data is 3000, and likes are a direct expression of the audience's approval and love for the video content. The number of comments is 500, which reflects the degree of interaction between the audience and the video content, which may include the audience's questions about the food making steps, evaluations of the dish's taste, or discussions of the creator's creativity. The number of collections is 1500, indicating that a considerable number of viewers believe that this video has value for repeated viewing or reference, possibly because they want to try making the food in the video.
[0069] The server continues to obtain the basic content of the target short video through KOL_BigData.
[0070] For the title part, the title returned by KOL_BigData is "Make delicious spaghetti in ten minutes". This title clearly conveys the core content of the video, which is the time efficiency of making spaghetti and the name of the dish.
[0071] Looking at the tag part, the tags queried are "food making", "spaghetti", "ten-minute food", and "simple recipe". These tags classify and describe the video content from different dimensions, which helps to accurately position the video in search and recommendation systems. For example, when a user searches for "spaghetti" or "food that can be made in ten minutes" on the platform, videos with these tags are more likely to be recommended.
[0072] As for the audio part, KOL_BigData provides relevant information about the audio. The audio of this cooking video contains soft background music that matches the light and casual atmosphere of cooking, with sounds of cutting vegetables, boiling noodles, and stirring sauce. These sound elements together constitute the audio content of the video, providing a more rich viewing experience for the audience.
[0073] In terms of background information acquisition, the server obtains from KOL_BigData that the creator of the target short video is a user named "Food Little Chef". This creator identity information is important for understanding the style, quality and audience of the video. For example, "Food Little Chef" may be known on the platform for making simple and delicious home cooking, and her subscribers may have high expectations and preferences for such cooking videos.
[0074] At the same time, KOL_BigData also returns the relevant information of the head competitor video. The head competitor video refers to the hot video that competes with the target short video for attention and traffic under the same theme (cooking - Italian noodles). For example, one of the head competitor videos is titled "Super Authentic Italian Noodle Making Tutorial", and its creator is "Italian Food Master". This competitor video may have unique aspects in terms of production process, ingredient selection or video shooting techniques, attracting a large number of viewers. By analyzing the competitor video, we can find out the differences and advantages between the target short video and the competitor, and thus provide a reference for optimizing the target short video.
[0075] The server will integrate all the information obtained from KOL_BigData to form the basic information of the target short video.
[0076] First, the effect data part, the server will store the onboarding time (March 15, 2023), the number of plays (80,000), the number of likes (3,000), the number of comments (500), and the number of collections (1,500) according to a specific data structure. For example, represented in a JSON format data object:
[0077] {
[0078] "onboarding_time": "2023-03-15",
[0079] "plays": 80000,
[0080] "likes": 3000,
[0081] "comments": 500,
[0082] "collections": 1500
[0083] }
[0084] For the base content part, the title ("Make delicious pasta in ten minutes"), tags ("food making", "pasta", "ten-minute food", "easy recipe"), and audio information (description of background music and production sound) are integrated together. The possible representation in a similar data structure is as follows:
[0085] {
[0086] "Title": "Make delicious pasta in ten minutes",
[0087] "Tags": ["food making", "pasta", "ten-minute food", "easy recipe"],
[0088] "Audio": {
[0089] "Background music": "soft and lively music",
[0090] "Production sound": ["chopping sound", "boiling sound", "stirring sauce sound"]
[0091] }
[0092] }
[0093] Finally, the creator ("Food Kitchen Girl") in the background information, the head competitor video (title "Super Authentic Italian Pasta Making Tutorial", creator "Italian Food Master") and its creator information are integrated:
[0094] {
[0095] "Creator": "Food Kitchen Girl",
[0096] "Head competitor video": [
[0097] {
[0098] "Title": "Super Authentic Italian Pasta Making Tutorial",
[0099] "Creator": "Italian Food Master"
[0100] } ]
[0102] }
[0103] Then, the server will combine these three parts (effect data, base content, and background information) into a complete basic information package, which will serve as the basis for subsequent analysis and processing of the target short video. For example, the entire basic information package is represented in a larger JSON object as follows:
[0104] {
[0105] "effect_data": {
[0106] "on_shelf_time": "2023-03-15",
[0107] "play_count": 80000,
[0108] "like_count": 3000,
[0109] "comment_count": 500,
[0110] "collection_count": 1500
[0111] },
[0112] "base_content": {
[0113] "title": "Make Delicious Italian Pasta in Ten Minutes",
[0114] "tags": ["food making", "Italian pasta", "ten-minute dish", "simple recipe"],
[0115] "audio": {
[0116] "background_music": "soft and lively music",
[0117] "production_sounds": ["chopping vegetables", "boiling pasta", "stirring sauce"]
[0118] }
[0119] },
[0120] "background_info": {
[0121] "creator": "Foodie Little Chef",
[0122] "top competitor videos": [
[0123] {
[0124] "title": "Authentic Italian Pasta Making Tutorial",
[0125] "creator": "Italian Food Connoisseur"
[0126] }
[0128]
[0129] }
[0130] This basic information package fully describes the various features and related background of the target short video, providing comprehensive data support for subsequent calls to the influencer large language model for deep information extraction and ultimately obtaining optimization suggestions.
[0131] In an embodiment of the present invention, based on the basic information, calling a preset internet celebrity language model to extract information from the target short video to obtain depth information can be implemented through the following examples.
[0132] The target short video is transcribed by calling a preset tool set through the internet celebrity language model, and the audio is transcribed into lines;
[0133] Analyze the title, the tags, and the lines using the internet celebrity language model to obtain the theme of the target short video;
[0134] The deep information is obtained from the preset internet celebrity knowledge graph through the internet celebrity language model based on the effect data, the basic content, the background information, the lines and the theme.
[0135] In this embodiment of the present invention, after obtaining the basic information of the target short video, the server begins to call the preset KOL_LLM language model. KOL_LLM is fine-tuned based on the Llama3-70B 1M context window version and has rich functions and pre-trained knowledge.
[0136] After receiving the audio portion of the basic information sent by the server, KOL_LLM invokes its pre-set toolset to begin transcribing the audio. For example, in the aforementioned food production short video, the audio includes soft background music and various sounds from the production process. KOL_LLM's transcription tool first pre-processes the audio, separating the background music from the production sounds. Since background music primarily creates atmosphere and does not directly contribute to the content, the tool focuses on identifying the vocal portion of the production sounds.
[0137] For example, when making pasta, the creator might explain as they go: "First, we prepare fresh pasta, boil water, and add a pinch of salt to make the noodles more chewy." The transcription tool will accurately convert these sounds into text lines: "First, prepare fresh pasta, boil water, and add a pinch of salt to make the noodles chewy." While the noodles are cooking, the creator says: "Now we can see the noodles tumbling in the pot. Let's turn down the heat and let it cook slowly." The transcription tool will also transcribe it as: "Now you see the noodles tumbling in the pot. Turn down the heat and let it cook slowly."
[0138] As the production process progresses, the creator also introduces the making of the sauce: "Next we make a delicious tomato sauce, put the chopped tomatoes in a pot, add some olive oil, minced garlic and basil leaves, and slowly simmer to fully release the flavor of the tomatoes." The transcription tool also accurately transcribes it as: "Next make tomato sauce, chop tomatoes into a pot, add olive oil, garlic, basil, slow simmer to release tomato flavor."
[0139] Throughout the entire transcription process, KOL_LLM's transcription tool will accurately identify the human voice content in the audio as much as possible and convert it into complete and coherent lines, which will provide an important basis for subsequent analysis of the theme of the video and obtaining deep information from the knowledge graph.
[0140] After obtaining the transcribed lines, KOL_LLM begins to analyze the title, tags and lines of the target short video to determine the theme.
[0141] For the title "Make delicious Italian pasta in ten minutes", it explicitly mentions the time limit for making Italian pasta (ten minutes) and the dish (Italian pasta). The tags "food making" "Italian pasta" "ten-minute food" "simple recipe" further emphasize that this is a video about food making, especially Italian pasta making, and emphasize the simplicity and time efficiency of the production.
[0142] Looking at the transcribed lines, such as "First, prepare fresh Italian pasta, boil water with a small spoonful of salt to make the noodles firm. Now see the noodles rolling in the pot, turn down the heat and slowly cook. Next make tomato sauce, chop tomatoes into a pot, add olive oil, garlic, basil, slow simmer to release tomato flavor." These lines describe the process of making Italian pasta in detail.
[0143] Through deep analysis of the semantic information in the title, tags and lines, KOL_LLM identifies key elements such as "Italian pasta" as the core dish, "making process" as the main action, "ten minutes" as the time feature, and "simple" as the difficulty of making. By integrating this information, KOL_LLM determines the theme of this short video as "Simple ten-minute Italian pasta making process". This theme accurately summarizes the core content of the video, covering the dish, the simplicity of making and the time element, providing a clearer direction for subsequent deep information retrieval from the knowledge graph.
[0144] After determining the theme, KOL_LLM performs information retrieval in the preset KOL_Graph based on the previously obtained effect data, basic content, background information, lines and theme. KOL_Graph is a four-layer structure graph based on tens of millions of short videos, including large category clusters, subcategory clusters, attitude clusters and theme clusters.
[0145] First, the large category cluster matching is performed. Since the topic of the target short video is "simple ten-minute Italian pasta making process", it belongs to the category of food making. KOL_LLM matches this topic with the large category clusters in KOL_Graph. In KOL_Graph, there may be a large category cluster named "food" which contains various sub-category content related to food. This "food" large category cluster covers different types of food making, food culture, food recommendations, and other knowledge entities. KOL_LLM discovers that "simple ten-minute Italian pasta making process" has a high degree of relevance to the "food" large category cluster through semantic analysis, and determines to further search under this large category cluster.
[0146] After determining the "food" large category cluster, KOL_LLM continues to perform sub-category cluster matching. Under the "food" large category cluster, there may be multiple sub-category clusters such as "Western cuisine making", "Chinese cuisine making", "dessert making", etc. Since the target short video is about Italian pasta making, which belongs to the category of Western cuisine, KOL_LLM will match the topic with these sub-category clusters. Through semantic analysis, it will find the "Western cuisine making" sub-category cluster, which contains more specific knowledge entities, relationships, and text blocks related to Western cuisine making. For example, this sub-category cluster may contain information about different types of Italian pasta (such as macaroni, spiral noodles, etc.), as well as common ingredients, seasonings, and kitchen utensils used in Western cuisine making.
[0147] After finding the "Western cuisine making" sub-category cluster, KOL_LLM then performs attitude cluster matching. Attitude clusters are closer to user scenarios, which gather opinions, views, and other entities. In the "Western cuisine making" attitude cluster, there may be different attitude clusters such as "pursuit of efficient Western cuisine making", "inheritance of traditional Western cuisine making", "creative Western cuisine making", etc. Since the target short video emphasizes "ten minutes" of making time, reflecting an efficient making method, KOL_LLM will match the topic with the "pursuit of efficient Western cuisine making" attitude cluster. This attitude cluster may contain some tips and experience sharing on how to make delicious Western cuisine in a short time, as well as feedback and evaluation of the audience on efficient Western cuisine making, and other related knowledge entities, relationships, and text blocks.
[0148] After completing the attitude cluster matching, the KOL_LLM performs topic cluster matching. From the attitude cluster of "pursuing efficient Western food making", the KOL_LLM will find the most suitable topic cluster. The topic cluster-bottom is a cluster composed of entities with a certain degree of relevance. In this case, there may be a topic cluster named "efficient making and innovation of Italian noodles". This topic cluster contains various entities related to efficient making of Italian noodles, such as specific Italian noodle brands, food material suppliers suitable for quick making, and relevant information such as audience comparison of different Italian noodle making videos. Through semantic analysis, the KOL_LLM determines that this topic cluster has a high degree of relevance to the topic of the target short video.
[0149] After completing the above matching, the KOL_LLM obtains target entities and target text blocks from the "efficient making and innovation of Italian noodles" topic cluster according to the pre-set relevance threshold. For example, the target entities may include a specific Italian noodle brand "XX brand Italian noodles", because this brand may have unique advantages in efficient making of Italian noodles, such as easy to cook, good taste, etc.; and may also include a special seasoning "quick-flavored Italian seasoning", which can add rich flavor to Italian noodles in a short time.
[0150] The target text blocks may contain some user evaluations of "XX brand Italian noodles" in quick making, such as "XX brand Italian noodles are really convenient, according to the method in the video, ten minutes can make delicious Italian noodles", or some sharing of using techniques of "quick-flavored Italian seasoning", such as "when making ten-minute Italian noodles, add a small amount of this quick-flavored Italian seasoning to make the noodles taste perfect".
[0151] These target entities and target text blocks constitute the depth information of the target short video, which will provide more rich and detailed basis for subsequent generation of optimization suggestions.
[0152] In the embodiment of the present application, the preset KOL knowledge graph is constructed by the following method, which can be implemented by the following examples.
[0153] The video itself information and account information of a plurality of stored videos stored in the preset KOL big data service are obtained through the KOL big language model;
[0154] The video itself information and account information of each stored video are structured to obtain the text blocks corresponding to each stored video, and the text blocks include the video itself information and the account information;
[0155] The plurality of text blocks are input into the KOL large language model for multi-round extraction to obtain entities, relationships, and covariates corresponding to each of the plurality of text blocks;
[0156] Multi-layer cluster extraction is performed according to the entities, relationships, and covariates to obtain the preset KOL knowledge graph, and the preset KOL knowledge graph includes a four-layer graph structure constructed from top to bottom of a large category cluster, a subcategory cluster, an attitude cluster, and a topic cluster.
[0157] In an embodiment of the present application, an exemplary server first starts a process of constructing a preset KOL knowledge graph (KOL_Graph). It calls a KOL large language model (KOL_LLM) to obtain relevant information of a plurality of stored videos stored in a preset KOL big data service (KOL_BigData).
[0158] For each stored video, KOL_BigData stores rich information. Taking a fitness short video as an example, the video itself information may include the title of the video as "Efficient home fitness exercise, fast fat burning", the tags as "fitness exercise", "home fitness", "fat burning exercise", the audio content containing light and brisk exercise background music and the coach's guidance password, the video's play count of 50,000 times, the like count of 2,000 times, the comment number of 800, the collection count of 1,000 times, the shelf time of April 10, 2023, etc.
[0159] In terms of account information, the creator account of this fitness short video is named "Fitness Expert Xia A", the number of subscribed users of the account is 100,000, the account's attention field is fitness, healthy lifestyle, etc., the account's activity is high, and the account publishes 3-4 fitness-related videos on average every week.
[0160] KOL_LLM obtains the video itself information and account information of each stored video in sequence. This process is like an information collector who methodically extracts the relevant information of each video from the huge information base of KOL_BigData to prepare for subsequent processing. This process involves massive video data, which may be hundreds of thousands or even millions of videos. KOL_LLM sequentially performs information acquisition operations on each video to ensure that no important information is missed.
[0161] After obtaining the video itself information and account information of each video, the server begins to structure the information.
[0162] Still taking the fitness short video as an example, for the video itself information, the server will organize it into a structured text block according to a certain format. For example:
[0163] "Title: Efficient home workout, fast fat burning; Tags: workout, home fitness, fat burning exercise; Audio: light and fast sports background music + coach's orders; Effect data: 50,000 plays, 2,000 likes, 800 comments, 1,000 collections, shelf time: 2023-04-10."
[0164] For account information, it will be arranged in a similar form:
[0165] "Account name: Fitness Expert Xia A; Number of subscribed users: 100,000; Areas of interest: fitness, healthy lifestyle; Account activity: publishes 3-4 fitness-related videos per week."
[0166] Then combine the two parts of information into a complete text block:
[0167] "Title: Efficient home workout, fast fat burning; Tags: workout, home fitness, fat burning exercise; Audio: light and fast sports background music + coach's orders; Effect data: 50,000 plays, 2,000 likes, 800 comments, 1,000 collections, shelf time: 2023-04-10; Account name: Fitness Expert Xia A; Number of subscribed users: 100,000; Areas of interest: fitness, healthy lifestyle; Account activity: publishes 3-4 fitness-related videos per week."
[0168] This text block contains complete video and account information, and is presented in a clear and orderly structure. For each video in KOL_BigData, the server will perform such structured processing to convert it into the corresponding text block. This process needs to strictly follow the predetermined structure rules to ensure the uniformity of each text block format, facilitating subsequent processing and analysis.
[0169] The server inputs a large number of text blocks (assuming 100,000 processed text blocks) into KOL_LLM for the first round of basic extraction.
[0170] For each text block, KOL_LLM will mine entities, relationships, and covariates (entity attributes). Taking the previous fitness short video text block as an example, KOL_LLM may identify entities such as "home workout" (as the core content entity of the video), "Fitness Expert Xia A" (creator entity), "fat burning exercise" (concept entity related to the workout), and "fitness enthusiasts" (implicit potential audience entity watching the video).
[0171] In terms of relationships, it identifies that "Fitness Enthusiast A" has a "creator-work" relationship with "Family Fitness Routine", "Family Fitness Routine" has a "belongs to" relationship with "Fat Burning Workout", and "Fitness Enthusiast" has an "audience-work" relationship with "Family Fitness Routine", etc.
[0172] In terms of covariants (entity attributes), for the entity "Family Fitness Routine", attributes may include "efficient" (describing the effectiveness of the fitness routine) and "suitable for family environment" (describing the applicable scenario of the fitness routine); for the entity "Fitness Enthusiast A", attributes include "number of subscribers - 100000" and "focus area - fitness, healthy lifestyle".
[0173] This process is a parallel processing of large text blocks, and KOL_LLM quickly scans the information in each text block, using its pre-trained knowledge and algorithms to accurately extract entities, relationships, and covariants. This process may consume some computing resources, but since KOL_LLM is optimized, it can complete the first round of basic extraction of a large number of text blocks within a reasonable time.
[0174] After the first round of basic extraction is completed, the server sends a signal to KOL_LLM to start the second round of extraction.
[0175] In the second round of extraction, KOL_LLM will conduct more in-depth mining. For example, for the entity "Family Fitness Routine" identified in the first round, it will further analyze the more complex relationships it may have with other entities. It may find that "Family Fitness Routine" has a "complementary use" relationship with certain fitness equipment (such as yoga mats and dumbbells), although these fitness equipment are not mentioned directly in the original text block. KOL_LLM can mine this potential relationship through its knowledge reserve and semantic analysis capabilities.
[0176] For the covariants (attributes) of entities, more in-depth mining will be conducted. For example, for the entity "Fitness Enthusiast", it may mine deeper attributes such as "fitness goal - weight loss, body shaping", which may be inferred from analyzing the comments in the video or based on general knowledge of the fitness field.
[0177] The second round of extraction is a deep mining process based on the first round, aiming to discover more hidden entities, relationships, and more detailed covariants, further enriching the information content corresponding to each text block.
[0178] After the first two rounds of extraction, the server starts the third round of merging operation. In this round, KOL_LLM will perform de-duplication and merging of all entity names identified in the previous two rounds.
[0179] For example, the entity "fitness exercises" may be mentioned multiple times in different text blocks, but may have different expressions, such as "family fitness exercises" and "efficient fitness exercises". KOL_LLM will merge these entities with different expressions into the entity "fitness exercises" and integrate its related relationships and covariates.
[0180] The same is true for commutative relationships. There may be multiple similar relationships expressed, such as "the creator produced the work" and "the work was created by the creator". KOL_LLM will merge and unify these relationships.
[0181] The third round of merging ultimately generates a complete, accurate, and non-redundant combination of entities, relationships, and covariates for each text segment. This makes the information in each text segment more refined and easier to understand, providing a high-quality data foundation for subsequent knowledge graph construction.
[0182] The server begins extracting multi-layer clusters based on the entities, relations, and covariates corresponding to each text block. First, it builds large-scale clusters.
[0183] Among all entities, entities with broad commonalities are grouped into a broad cluster. For example, entities such as "aerobics," "yoga classes," and "aerobics classes" are grouped into the broad cluster "fitness classes," as they all fall into the fitness category. Entities such as "fitness expert Xiao A," "fitness coach Xiao B," and "fitness influencer Xiao C" are grouped into the broad cluster "fitness creators," as they all create content in the fitness field.
[0184] The construction of these large clusters is based on the high-level classification of entities, aiming to build the top-level structure of the knowledge graph to facilitate the subsequent organization and retrieval of information under more specific classifications.
[0185] After building the large cluster, the server continues to build the sub-cluster.
[0186] Taking the "Fitness Classes" cluster as an example, we constructed subclusters within it. We considered "Aerobics" as a separate subcluster, as it has unique characteristics and audiences within fitness classes. We also considered "Yoga Classes" as a subcluster, as it has a different movement system, practice objectives, and audience than aerobics. We further subdivided "Aerobics Classes" into smaller subclusters based on different exercise types (such as running classes and swimming classes).
[0187] For the "fitness creator" category cluster, sub-category clusters are divided according to the professional field or creation style of the creator. For example, creators who focus on strength training teaching are classified as "strength training creator" sub-category cluster, and creators who are good at yoga teaching are classified as "yoga creator" sub-category cluster, etc.
[0188] The construction of sub-category clusters further refines the knowledge graph based on the large category, enabling more precise positioning and organization of related knowledge entities, relationships, and covariants.
[0189] Next, attitude clusters are constructed. Under the "fitness exercise" sub-category cluster, there may be different attitude clusters.
[0190] For example, there is an "attitude cluster that promotes the weight loss effect of fitness exercise", which includes entities, relationships, and covariants related to videos that emphasize the significant effect of fitness exercise on weight loss. For example, some videos emphasize that fitness exercise can quickly burn fat, and there are many successful student cases sharing these information, which will be classified into this attitude cluster.
[0191] There is also an "attitude cluster that focuses on the interest of fitness exercise", which includes video-related content that highlights the interesting and easy-to-keep characteristics of fitness exercise. For example, some fitness exercise videos incorporate popular music and have interesting movement designs, and the entities, relationships, and covariants involved in these videos will be classified into this attitude cluster.
[0192] The construction of attitude clusters is more closely related to users' actual needs and views of things, and can directly reflect the distribution of knowledge content under different attitudes.
[0193] Finally, the theme cluster-bottom layer is constructed. Under the "attitude cluster that promotes the weight loss effect of fitness exercise", the theme cluster is constructed.
[0194] For example, there is a theme cluster called "efficient action combination of fitness exercise for weight loss", which includes entities, relationships, and covariants related to videos that specifically introduce which action combinations in fitness exercise can efficiently lose weight. It may include specific fitness exercise action names (such as deep squat jump, high leg lift, etc.), the relationship between these actions and weight loss effect (such as how many calories are consumed per minute by deep squat jump, etc.), and the demonstration and explanation of these action combinations in different videos, etc.
[0195] Through such a top-down step-by-step construction, the preset KOL graph (KOL_Graph) with a four-layer structure of large category cluster, sub-category cluster, attitude cluster, and theme cluster is finally obtained. This knowledge graph provides a rich knowledge reserve and powerful retrieval tool for subsequent short video analysis, optimization, and other operations.
[0196] In the embodiments of the present application, the deep information is obtained from the preset KOL knowledge graph based on the effect data, the basic content, the background information, the dialogue and the theme by the KOL large language model. The implementation can be performed through the following examples.
[0197] Through the KOL large language model, the target entity and the target text block are obtained according to the preset relevance threshold by performing large category cluster matching, subcategory cluster matching, attitude cluster matching and theme cluster matching in the preset KOL knowledge graph based on the effect data, the basic content, the background information, the dialogue and the theme.
[0198] The target entity and the target text block are taken as the deep information.
[0199] In the embodiments of the present application, for example, it is assumed that we are processing a travel short video. The server has obtained the effect data (such as 150,000 times of play, 8,000 times of likes, 2,000 comments and 3,000 times of collections, and the shelf time is July 5, 2023) of the short video, the basic content (the title is “Explore the mysterious Bali Island”, the tags are “Bali Island”, “tourism” and “tropical style”, the audio contains sea waves, local ethnic music and guide's explanation), the background information (the creator is a user named “Tourist Explorer”, the head competitor video is other popular videos about Bali Island tourism, and the creators are “Bali Island Tourism Expert” and the like), the dialogue (obtained by transcription, such as “Now we are at Kuta Beach in Bali Island. The sand here is delicate and soft, and the sea waves are very suitable for surfers”) and the theme (determined as “Bali Island tourism experience”).
[0200] The server calls the KOL large language model (KOL_LLM) and inputs these information into the preset KOL knowledge graph (KOL_Graph) to perform large category cluster matching. The large category cluster in KOL_Graph includes a plurality of broad categories, such as “tourism”, “food”, “fitness” and the like. KOL_LLM performs semantic analysis on the input information. Since the theme is “Bali Island tourism experience”, it is clear that it belongs to the large category cluster of “tourism”. Therefore, KOL_LLM determines to perform further matching operation under the large category cluster of “tourism”. This process is like first finding the correct large classification area in a huge knowledge warehouse, which lays a foundation for subsequent more accurate search.
[0201] After determining the large category cluster of “tourism”, KOL_LLM starts to perform subcategory cluster matching. Under the large category cluster of “tourism”, there are many subcategory clusters, such as “domestic tourism”, “overseas tourism”, “island tourism”, “mountain tourism” and the like.
[0202] KOL_LLM analyzes the short video's relevant information. Since the topic is "Bali travel experience", and Bali is an island, it will focus on the "island tourism" sub-cluster. This sub-cluster contains various entities, relationships, and text blocks related to island tourism. For example, different island tourist attractions, suitable seasons for island tourism, and island tourism activities (such as diving, fishing, etc.). Through this matching, KOL_LLM further narrows down the search range from the many possibilities under the large cluster to a more specific sub-cluster, in order to obtain more relevant and in-depth information.
[0203] After finding the "island tourism" sub-cluster, KOL_LLM then performs attitude cluster matching. In the "island tourism" attitude cluster, there may be different attitude clusters such as "enjoying the tranquility and beauty of the island", "enthusiastic about island adventure activities", and "pursuing luxury vacation experiences on the island".
[0204] By analyzing the dialogue in the short video, such as mentioning enjoying the sunshine and beautiful sea view on the beach in Bali, and not involving too many adventure activities, KOL_LLM will match it with the "enjoying the tranquility and beauty of the island" attitude cluster. This attitude cluster may contain tourists' love for the peaceful atmosphere of the island, praise for the beautiful scenery, and some recommended places and activities suitable for enjoying the tranquility and beauty of the island, etc. related knowledge entities, relationships and text blocks. Through this matching, KOL_LLM is closer to the emotions and experiences conveyed by the short video, and thus more accurately locates the knowledge content related to the video content.
[0205] After completing the attitude cluster matching, KOL_LLM performs theme cluster matching. From the "enjoying the tranquility and beauty of the island" attitude cluster, KOL_LLM finds the most suitable theme cluster.
[0206] For example, there may be a "Bali beach tranquility and beauty" theme cluster. This theme cluster contains information about the characteristics of various beaches in Bali (such as Kuta Beach, Nusa Dua Beach, etc.), the specific manifestations of the tranquil atmosphere on the beach (such as the sunset scenery in the evening, the tranquil sea surface in the morning, etc.), the relaxation activities that can be done on the beach (such as beach yoga, beach walking, etc.), and tourists' evaluations of the tranquil beauty of these beaches. Related information. KOL_LLM determines that this theme cluster has a high degree of relevance to the theme of the short video "Bali travel experience" through comprehensive analysis of the short video.
[0207] After completing all the cluster matching described above, KOL_LLM obtains the target entity and target text block from the "Bali beach tranquility and beauty" theme cluster according to the pre-set relevance threshold.
[0208] Assuming the preset relevance threshold is 0.8, KOL_LLM will calculate and evaluate the relevance of each entity and text chunk in the topic cluster to the short video. Target entities may include the specific location entity "Kuta Beach" because it was mentioned in the short video and it is one of the representative beaches of Bali; and the activity entity "beach yoga" because it belongs to the relaxing activities that can be done on the beach.
[0209] The target text chunk may contain a tourist's review of Kuta Beach: "Kuta Beach is one of the most enchanting beaches in Bali, with its fine sand and gentle waves that caress the shore, making one feel incredibly relaxed, especially during the evening when watching the beautiful sunset, it's like time has stood still." This text chunk is highly relevant to the theme of the short video and meets the preset relevance threshold.
[0210] In this way, KOL_LLM filters out target entities and target text chunks closely related to the short video from the preset KOL knowledge graph, which will serve as the depth information of the short video and provide detailed and targeted basis for subsequent optimization suggestions.
[0211] After receiving the target entities (such as "Kuta Beach" and "beach yoga") and target text chunks (such as the tourist's review of Kuta Beach) filtered by KOL_LLM, the server stores them as depth information of the target short video for subsequent processing.
[0212] These depth information will be integrated into a special data structure, for example, represented in JSON format:
[0213] {
[0214] "depth information": {
[0215] "target entities": ["Kuta Beach", "beach yoga"],
[0216] "target text chunks": ["Kuta Beach is one of the most enchanting beaches in Bali, with its fine sand and gentle waves that caress the shore, making one feel incredibly relaxed, especially during the evening when watching the beautiful sunset, it's like time has stood still."]
[0217] }
[0218] }
[0219] This depth information data structure will be used together with the basic information obtained previously for subsequent operations, such as input into a KOL large language model to generate optimization suggestions for the target short video. By dividing the target entity and target text into blocks as depth information, the server can use these richer and more targeted information to deeply understand the characteristics and advantages of the short video, thereby providing more valuable references for optimizing the short video.
[0220] In the embodiments of the present application, the preset KOL large language model can be obtained by the following examples.
[0221] Obtain a basic large language model and an advanced large language model, the model parameter quantity of the advanced large language model being greater than that of the basic large language model;
[0222] Through the advanced large language model, a sample short video is analyzed, self-verified, and self-selected extraction strategy is performed based on a self-discovery mechanism, and knowledge extraction data is generated through prompt engineering;
[0223] The basic large language model is fine-tuned based on the knowledge extraction data in a LoRA manner to obtain the preset KOL large language model.
[0224] In the embodiments of the present application, the server needs to obtain a basic large language model and an advanced large language model in the process of constructing a preset KOL large language model (KOL_LLM).
[0225] The server selects a suitable basic large language model from the model repository. For example, Llama3-70B is selected as the basic large language model. This model has certain pre-training knowledge and ability to handle various natural language processing tasks. Its structure and parameters are pre-trained by a large amount of text data, containing rich semantic information and language patterns. Llama3-70B is selected as the basic model because it performs well in text generation, multi-task support, reasoning, and general performance, and has the advantage of being open source, easy to deploy in a server environment.
[0226] At the same time, the server obtains an advanced large language model, which has a larger model parameter quantity than the basic large language model. For example, a model with a larger parameter quantity is selected, such as a large language model optimized for knowledge extraction tasks, which has a parameter quantity of 100 billion parameters (here is just an example, the actual model is determined according to the specific situation). This advanced large language model has stronger ability in handling complex semantic analysis and knowledge mining tasks, and due to its larger parameter quantity, it can accommodate more knowledge and language patterns, thereby providing a more in-depth and comprehensive understanding when analyzing short video related content.
[0227] The server selects a batch of sample short videos from a pre-set short video sample library for knowledge extraction data generation. This sample short video library contains various types of short videos, covering tourism, food, fitness, entertainment and other fields. For example, a sample short video of the tourism category is selected, titled "Explore the Ancient Kyoto", with tags "Kyoto", "Japan Tourism", "Historical Culture", and audio containing the noise of Kyoto streets, the bell of the temple and the guide's explanation. The video content shows the ancient buildings of Kyoto and traditional tea ceremony performances.
[0228] The server inputs the relevant information of this tourism sample short video into the advanced large language model. The advanced large language model starts analyzing the short video based on the self-discovery mechanism.
[0229] For the title "Explore the Ancient Kyoto" in the video, the model analyzes through the self-discovery mechanism that "Kyoto" is an important entity and is associated with the attribute "ancient", which may imply that the historical and cultural value of Kyoto is a core content of the video. For the tags "Kyoto", "Japan Tourism", "Historical Culture", the model further understands that this video is mainly about tourism in Kyoto, Japan, and emphasizes its historical and cultural features.
[0230] In analyzing the audio content, the model identifies that the temple bell is a symbolic element of Kyoto culture, and the guide's explanation may contain important information about Kyoto's historical buildings and cultural traditions. For the ancient buildings and tea ceremony performances shown in the video content, the model discovers that these are key elements that reflect the historical and cultural value of Kyoto. Through this analysis based on the self-discovery mechanism, the model can automatically mine the potential semantic information in the short video without the need to define explicit rules or patterns in advance.
[0231] Based on the analysis, the advanced large language model begins to perform self-verification. For example, for the conclusion that "the historical and cultural value of Kyoto is the core content" identified earlier, the model will again check whether each element in the video supports this conclusion.
[0232] The model will check whether the architectural style in the video, the details of the tea ceremony performance and the content of the guide's explanation are all closely related to the historical and cultural value of Kyoto. If it finds that some elements are contradictory or irrelevant, the model will adjust its analysis results. For example, if the guide's explanation suddenly mentions the modern commercial development of Kyoto, which deviates from the previously identified core content of historical and cultural value, the model will further analyze the relationship between this part and the whole, and may determine that it is a contrast or supplementary explanation, rather than a change in the core content, thereby verifying the previous conclusion that historical and cultural value is the core content.
[0233] After analysis and self-validation, the advanced large language model adopts a self-selected extraction strategy to determine the knowledge content to be extracted.
[0234] For this tourism sample short video, the model self-selects entities, relationships, and attributes directly related to Kyoto's history and culture as knowledge content. For example, it selects "Kinkaku-ji" (as a famous historical building entity in Kyoto), "Tea Ceremony" (as a traditional art form entity in Kyoto), "Kinkaku-ji - represents Kyoto's historical and cultural heritage" (relationship), "Tea Ceremony - embodies Kyoto's traditional culture" (relationship), etc. The model automatically selects these knowledge elements that best represent the core meaning of the video according to its judgment of their importance to the short video content, while ignoring some secondary or less relevant content, such as occasional appearances of passers-by in the video.
[0235] After self-selecting knowledge content, the advanced large language model generates knowledge extraction data through prompt engineering.
[0236] The server provides a series of prompts for the model to guide the generation of knowledge extraction data. For example, for the previously selected "Kinkaku-ji" entity, the prompt might be "Please describe the status of Kinkaku-ji in Kyoto's history and culture, its architectural features, and its relationship with other scenic spots." The model generates detailed knowledge extraction data about Kinkaku-ji based on this prompt, such as "Kinkaku-ji is one of the most representative historical and cultural heritage in Kyoto, built in the 14th century, with its architectural features of golden appearance, blending traditional Japanese architectural style. Compared with other scenic spots in Kyoto, Kinkaku-ji is a must-visit place for tourists, and the surrounding garden landscape also complements the architecture, together showing the charm of Kyoto's ancient culture."
[0237] Similarly, for the "Tea Ceremony" entity, the prompt might be "Please explain the significance of the Tea Ceremony in Kyoto's traditional culture, its process, and its connection with local residents' lives", and the model generates corresponding knowledge extraction data, such as "The Tea Ceremony holds a high status in Kyoto's traditional culture, it is an art form of self-cultivation. The process of the Tea Ceremony includes the preparation of tea utensils, the making of green tea, and the tasting of tea. In Kyoto, the Tea Ceremony is closely linked to the lives of local residents, many families still maintain traditional Tea Ceremony rituals, it is not only a beverage culture, but also a way to inherit family culture and social etiquette."
[0238] In this way, for each self-selected knowledge content, the advanced large language model generates detailed knowledge extraction data according to the prompt, which will be used to fine-tune the basic large language model.
[0239] Before fine-tuning, the server first prepares the relevant parameters and settings required for Low-Rank Adaptation (LoRA). LoRA is an effective model fine-tuning method that adjusts the parameters of the model by adding a low-rank matrix to the original model, allowing it to quickly adapt to specific tasks while maintaining the basic capabilities of the model.
[0240] The server determines the range of parameters to be adjusted, for example, selecting certain layers or specific parameter subsets in the base large language model (Llama3-70B) for fine-tuning. At the same time, the rank parameter of LoRA is set, which determines the size of the low-rank matrix and thus affects the fine-tuning effect and computational cost. For example, setting the rank to 8 is a value determined based on experiments and experience that can guarantee the fine-tuning effect while controlling the computational cost.
[0241] The server applies the knowledge extraction data generated through the advanced large language model and prompt engineering to the fine-tuning process of the base large language model.
[0242] Taking the knowledge extraction data about "Kinkaku-ji" and "Tea Ceremony" generated from the previous tourism sample short video as an example, the base large language model (Llama3-70B) will adjust its internal parameters based on these knowledge data during the fine-tuning process.
[0243] For knowledge data related to "Kinkaku-ji", the model will adjust parameters related to architecture, historical culture, tourist attractions, etc. For example, if the model's understanding of the "Kinkaku-ji" entity is not accurate or deep enough, through fine-tuning, it will better understand the importance of Kinkaku-ji in Kyoto tourism and historical culture, as well as its relationship with other related entities (such as tourists, garden landscapes, etc.).
[0244] For knowledge data about "Tea Ceremony", the model will adjust parameters related to traditional culture, artistic form, social etiquette, etc. For example, the model will more accurately understand the significance of tea ceremony in the lives of local residents in Kyoto, as well as the relationship between each step in the tea ceremony process and cultural heritage.
[0245] During the entire fine-tuning process, due to the use of LoRA, the base large language model will only make small adjustments to the selected parameters. For example, the amount of adjusted parameters is less than 0.005%. This small adjustment can preserve the basic capabilities of the model, such as text generation, multi-task processing, etc., while increasing the understanding and processing capabilities of short video content.
[0246] After LoRA fine-tuning based on knowledge extraction data, the base large language model is transformed into the preset KOL large language model (KOL_LLM).
[0247] This KOL_LLM has the ability to optimize for short video content when handling tasks related to short videos. For example, when analyzing a new travel-related short video, the KOL_LLM can more accurately identify the tourist attraction entities in the video, understand the historical and cultural significance behind the attractions, analyze the relationship between tourists and attractions, etc. It can better utilize the knowledge learned from sample short videos to analyze and process different types of short videos more deeply, thereby providing more effective support for short video optimization suggestions, content analysis, and other tasks.
[0248] In the embodiments of the present application, the input of the basic information and the depth information into the KOL_LLM to obtain the optimization suggestions for the target short video can be implemented through the following examples.
[0249] The basic information and the depth information are input into the KOL_LLM to obtain optimization suggestions for the target short video, including new titles, new tags, and new lines.
[0250] In the embodiments of the present application, the server inputs the basic information and the depth information of the target short video into the KOL_LLM (KOL_LLM) after obtaining them.
[0251] Taking the previously mentioned travel-related short video as an example, the basic information includes effect data (15 million views, 8000 likes, 2000 comments, 3000 collections, and a shelf time of July 5, 2023), basic content (the original title is "Exploring the Mysterious Bali Island", the original tags include "Bali Island", "Tourism", and "Tropical Atmosphere", the audio contains sea waves, local ethnic music, and tour guide explanations), and background information (the creator is a user named "Tourist Explorer", the head competitor video is another popular video about Bali tourism, and the creators are "Bali Tourism Expert" and others).
[0252] The depth information may include target entities (such as "Kuta Beach" and "Beach Yoga") and target text blocks (such as "Kuta Beach is one of the most enchanting beaches in Bali, with its fine sand and gentle sea waves that gently caress the shore, making you feel incredibly relaxed, especially during the sunset, it's like time has stood still.").
[0253] The server organizes this basic information and depth information according to a specific data format, such as constructing a JSON object containing all the information, and then sends this object to the KOL_LLM. After receiving this data, the KOL_LLM begins to analyze and process it in order to generate optimization suggestions for the target short video.
[0254] After receiving the basic information and depth information of the target short video, KOL_LLM first analyzes how to optimize the title.
[0255] For the original title "Exploring the Mysterious Bali Island", KOL_LLM will optimize it based on the target entities and text blocks in the depth information. For example, Kuta Beach in the depth information is an important scenic spot in Bali Island, and the description of Kuta Beach in the target text block is very attractive. KOL_LLM may generate a new title: "Bali Island Kuta Beach: Enchanting Tropical World". This new title highlights the Kuta Beach feature of Bali Island more specifically, and uses the attractive expression "Enchanting Tropical World" to better attract the attention of the audience.
[0256] During the process of generating a new title, KOL_LLM will also consider the effect data in the basic information. If the data such as the number of plays and likes shows that the original title may not have enough appeal, KOL_LLM will focus on adding more attractive elements in the new title. For example, if the analysis finds that the click rate of the original title is low, it may be because the expression "Exploring the Mysterious" is too vague, while the new title directly mentions "Kuta Beach", which can make the audience more clearly understand the content focus of the video, thereby improving the click rate.
[0257] Next, KOL_LLM begins to consider the generation of new tags.
[0258] The original tags are "Bali Island", "Tourism", and "Tropical Style". KOL_LLM supplements and optimizes them according to the content in the depth information. Since the depth information mentions Kuta Beach and beach yoga, KOL_LLM may add the tags "Kuta Beach" and "Beach Yoga". These two tags more specifically describe the content elements in the video, which helps to more accurately target the target audience in the search and recommendation system.
[0259] At the same time, KOL_LLM will also analyze the background information in the basic information. If it finds that the creator's style or the target audience's preferences are related to certain tags, it will also make corresponding adjustments. For example, if the creator often publishes beach-related travel videos and the subscribed users are interested in beach activities, KOL_LLM may further optimize the tags and adjust "Beach Yoga" to "Bali Island Beach Yoga" to make it more targeted.
[0260] Moreover, KOL_LLM also evaluates the effectiveness of tags based on the effect data. If a certain tag corresponds to a low search volume, but the video content is related to it, KOL_LLM may modify or replace this tag. For example, if the "tropical style" tag can summarize the overall atmosphere of Bali, but it is not precise enough in search, KOL_LLM may replace it with "Bali beach style", so that the video is more likely to be recommended when searching for Bali beach-related content.
[0261] For the generation of new scripts, KOL_LLM considers multiple aspects of basic information and depth information.
[0262] First, KOL_LLM analyzes the original script (obtained by transcription, such as "Now we are in Kuta Beach, Bali, where the sand is soft and delicate, and the waves are perfect for surfers"). If the original script is not vivid enough or the information is incomplete, KOL_LLM will optimize it.
[0263] According to the target text block in the depth information (such as "Kuta Beach is one of the most charming beaches in Bali, where the sand is delicate and the waves gently hit the shore, making people feel incredibly relaxed, especially in the evening, watching the beautiful sunset, it's like time has stopped"), KOL_LLM may suggest adding a description of the Kuta Beach sunset scene to the original script. The optimized script may be: "Now we are in Kuta Beach, Bali, where the sand is soft and delicate, and the waves are perfect for surfers. In the evening, Kuta Beach is even more charming, with the beautiful sunset coloring the entire beach orange-red, and the waves gently hitting the shore, making people feel incredibly relaxed, as if time has stopped."
[0264] KOL_LLM also considers the audio content in the basic information. If there are some unique elements in the audio, such as local ethnic music, KOL_LLM may suggest adding a description of the integration of this music and the beach atmosphere in the script, for example: "Now we are in Kuta Beach, Bali, accompanied by the local ethnic music, where the sand is soft and delicate, and the waves are perfect for surfers. In the evening, Kuta Beach is even more charming, with the beautiful sunset coloring the entire beach orange-red, and the waves gently hitting the shore, making people feel incredibly relaxed, as if time has stopped."
[0265] In addition, KOL_LLM can determine which parts of the original script may need to be optimized based on the effect data. If the analysis finds that the audience is more interested in the part of the video about local culture (for example, through the distribution of comment and like numbers in this part of the content), but the original script has less description in this regard, KOL_LLM may suggest adding a description of the connection between the local culture of Bali Island (such as local religious rituals, traditional customs, etc.) and Kuta Beach in the script to make the script more rich and attractive.
[0266] Through the generation of the above optimization suggestions for new titles, new tags, and new scripts, KOL_LLM provides a comprehensive optimization direction for the target short video, which helps to improve the attractiveness, searchability, and viewing experience of the short video, and thus has the potential to improve various effect data of the short video, such as play count, like count, comment count, and collection count, etc.
[0267] In order to more clearly describe the scheme provided by the embodiments of the present application, a more detailed implementation manner is provided below.
[0268] Please refer to Figure 2 , Figure 2 The system framework diagram of the short video optimization based on knowledge graph provided by the embodiments of the present application.
[0269] 1) Log in to the platform and open the chat window of the robot expert;
[0270] 2) Input the online link or design scheme of the short video, such as the short video link on the TikTok platform in the input box below;
[0271] 3) After sending, the robot expert starts running;
[0272] 4) Access KOL_BigData to obtain the basic information of the video:
[0273] Obtain effect data: shelf time, play count, like count, comment count, and collection count;
[0274] Obtain content: title, tag, and audio;
[0275] Obtain background information: creator, head competitor video, and creator;
[0276] If the video data is not stored in the storage server, call the real-time grabbing service to obtain and store;
[0277] 5) Access Graph RAG service to obtain more information of the video:
[0278] Obtain script: convert audio to text through the transcription function of KOL_LLM tool set;
[0279] Get topic / scene: Analyze the title, tags, and dialogues of the video content through KOL_LLM to obtain the topic of the video content.
[0280] Get deep content:
[0281] Through KOL_LLM, based on the information, dialogues, and topics / scene obtained in 4), identify a batch of related entities and relationships in KOL_Graph.
[0282] Use them to repeatedly extract more related entities, relationships, and cluster reports from the knowledge graph. According to the relevance, the original text blocks of the video content will also be extracted.
[0283] Sort and filter the search results according to relevance, taking into account the context window size of KOL_LLM.
[0284] Return the final search results to KOL_LLM, such as: tags, topics, items, creators' associated hot events, people, other videos, and play counts, likes, and other effect data, hot event-related characters, locations, etc. This kind of in-depth, detailed, and rich relationship information.
[0285] Output optimization suggestions and new content design:
[0286] 6) Return the information of 4) and 5) to KOL_LLM;
[0287] The large model understands, extracts, and summarizes optimization suggestions for the video from the information, and designs complete new titles, tags, and dialogues.
[0288] 7) Output optimization suggestions and new solutions to the chat window.
[0289] Core service introduction:
[0290] KOL Big Data Service (KOL_BigData)
[0291] 1) Video library: data storage of billions of videos, which can return related information of videos in real time;
[0292] 2) Crawler service: video crawler that can crawl related information of videos in real time;
[0293] KOL Large Language Model (KOL_LLM)
[0294] Based on the Llama3-70B 1M context window version, fine-tune the model with larger parameters, and use the invented Prompt framework to extract the required data for fine-tuning.
[0295] 1) Reason for choosing Llama3-70B: Open-source model, with 70 billion parameters, leading in text generation, multi-task support, reasoning, and general performance, easy to deploy;
[0296] 2) Reason for choosing 1M context window:
[0297] Establishing a popular knowledge graph: Entity and relationship extraction requires as much video content as possible to be imported to ensure accurate and efficient extraction and merging of entities and relationships, avoiding the incompleteness of local extraction and the inefficiency of more rounds of entity merging;
[0298] Knowledge graph retrieval: This stage returns a large amount of cluster, entity, relationship, and other knowledge data, which needs to be sufficient in one analysis to make the analysis and optimization results more comprehensive and accurate;
[0299] 3) Fine-tuning process:
[0300] Objective: Improve the efficiency, completeness, and accuracy of knowledge (entity, relationship, covariant-attribute of entity) extraction from short video content
[0301] Fine-tuning data production - use Gpt-4o model with more parameters:
[0302] Production method: Use prompt engineering, and focus on the knowledge extraction framework for short video content:
[0303] Prompt framework idea: Fully tap into the self-discovery mechanism of large models, use self-analysis, self-verification, and self-selection extraction strategies based on short video content to achieve accurate and comprehensive knowledge extraction;
[0304] Prompt framework implementation, please refer to Figure 3 :
[0305] Generated data: Select 50,000 videos of various types from the past 3 years, about 100MB of data, and generate corresponding knowledge;
[0306] Data effect: Compared with the extraction effect of non-self-discovery Prompt, the total number of knowledge is increased by 20%, and it basically covers the knowledge extracted by it, which is equivalent to stimulating 20% of the potential of the large model;
[0307] Fine-tuning:
[0308] Fine-tune Llama3-70B using LoRA;
[0309] Use 50,000 video contents and corresponding generated data set to fine-tune the model;
[0310] Effect: The change in parameters is less than 0.005%, preserving the basic capabilities of the model while increasing its understanding and processing capabilities for short video content. The fine-tuned model extracts video knowledge faster than Gpt-4o and performs better than the pre-fine-tuned model and Gpt-4o.
[0311] Cost: As an open-source, self-built model, the cost is much lower than Gpt-4o, meeting online service standards.
[0312] KOL Graph
[0313] 1) Knowledge graph establishment - based on tens of millions of short videos
[0314] Text chunking - building basic source data:
[0315] Get information for each video from KOL_BigData, which is a basic information unit.
[0316] Pass into KOL_LLM, the model will automatically call the toolset to get the video's lines.
[0317] The model organizes each video information into two categories: Video and Channel, which are video information and account information.
[0318] According to the specified structured format, convert to text chunks as the bottom layer of data, and each chunk is a video's information.
[0319] Entity multi-round extraction - build basic knowledge structure:
[0320] First round of basic extraction: pass as many text chunks as possible into KOL_LLM, close to half the context window, and extract entities, relationships, and covariants (entity attributes), such as: extract video, author, tags, games, books, and other entities, relationships, and entity attributes from video information, and multiple times of passing in and extracting entities from all text chunks.
[0321] Second round of extraction: give the model a clear signal to extract again, which will dig deeper.
[0322] Third round of merging: pass the existing entity names to KOL_LLM for deduplication and merging, generating the final combination of entities, relationships, and covariants. Please refer to Figure 4 .
[0323] Multi-layer cluster extraction - build knowledge in different dimensions:
[0324] There are four layers in total, each clustered by the relevance of the underlying layers or entities. This improves information usability, density, and regularity, significantly improving the quality and efficiency of search results. Each cluster's data includes a cluster report (title, summary, score, score explanation, key insights) and elements. The report is designed to better express the core information of the cluster, making search directions and results more precise, while the elements provide the next step in search directions and results.
[0325] Level 3 - Major category clusters - top level: clusters composed of related sub-category clusters, which facilitates retrieval direction through larger categories;
[0326] Level 2-Subcategory cluster: A cluster composed of related attitude clusters, which facilitates finding more detailed directions through small categories;
[0327] Level 1 - Attitude Cluster: This cluster, closest to the user scenario, aggregates opinions and views on entities, making the details of these opinions more comprehensive and accurate, making it easier to directly locate content that matches the search.
[0328] Level 0 - Theme Cluster - Bottom Layer: Clusters composed of entities with a certain degree of relevance. For example, the game Black God x cluster consists of entities such as Game-Black God x, R&D Team-Game Sience, Book-Journey to the West, and Video-Best Game Recommendations, and their corresponding relationships. This helps to increase knowledge density and makes it easier to find related entities through common nouns and themes. Please refer to Figure 5 ;
[0329] 2) Knowledge retrieval
[0330] Query understanding and optimization: KOL_LLM optimizes search queries to highlight query intent, related topics, and entities. For example, if a user provides a link to a video review of game xxx, the large model will focus on extracting: game information, review methods, dimensions, and results, influencer information and attitudes, and then retrieve more information from the knowledge graph.
[0331] Cluster positioning:
[0332] Cluster matching: Use the extracted key points to perform semantic matching on clusters. For example, a game review video is matched to the three clusters of games, reviews, and gaming influencers.
[0333] Subcategory cluster matching: Use key points to continue semantic matching in subcategory clusters within the hit major cluster, for example, matching to subcategory clusters such as "domestic ARPG games," "domestic gaming influencers," and "game reviews."
[0334] Attitude cluster matching: This layer of semantic matching has a probability of directly hitting and inputting a very relevant cluster of videos, which can maximize retrieval efficiency, such as matching to "game black god x good difficult" attitude cluster, and the content of the analyzed video is highly relevant;
[0335] Theme cluster matching: semantic matching of suitable theme clusters from attitude clusters
[0336] The matching of each layer will be based on the extracted points and cluster reports, and the matching results will be ranked in descending order of relevance, and a certain number of results will be returned, and the lower the layer, the more results will be returned;
[0337] Entity and text block positioning: semantic matching of suitable entities and relationships from selected theme clusters, returning a large amount of information, and most of the relevant entities and relationships in the input video will be found, and the source text block can also be positioned;
[0338] Result return: also based on relevance descending order to arrange a certain number of positioned entities and text blocks. For example: for the input game review video, game review tips, creator's subscription user preference information, and recent game public opinion can be mined through the relationship attribute of the knowledge graph, which can help analyze the good and bad of the video and optimize.
[0339] Please refer to Figure 6 , Figure 6 A short video optimization device 110 based on a knowledge graph is provided for an embodiment of the application, comprising:
[0340] An acquisition module 1101 is configured to acquire a target short video link, wherein the target short video link points to a target short video; call a preset net red big data service to acquire basic information of the target short video; based on the basic information and a preset net red knowledge graph, call a preset net red big language model to perform information extraction on the target short video to obtain deep information;
[0341] An optimization module 1102 is configured to input the basic information and the deep information into the net red big language model to obtain an optimization suggestion for the target short video.
[0342] It should be noted that the implementation principle of the short video optimization device 110 based on the knowledge graph can refer to the implementation principle of the short video optimization method based on the knowledge graph, which will not be described here. It should be understood that the division of each module of the above device is only a logical function division, and all or part of it can be integrated into a physical entity, or it can be physically separated. And these modules can all be in the form of software called by the processing element; all in the form of hardware; some modules can be implemented in the form of software called by the processing element, and some modules can be implemented in the form of hardware. For example, the short video optimization device 110 based on the knowledge graph can be a separate processing element, or it can be integrated into a chip of the above device, in addition, it can also be stored in the form of program code in the memory of the above device, and the function of the above short video optimization device 110 based on the knowledge graph is called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of the hardware in the processor element or the instruction in the form of software.
[0343] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of scheduling program code by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together to implement in the form of system on a chip (SOC).
[0344] The embodiment of the application provides a computer device 100, which comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the short video optimization device 110 based on the knowledge graph. As shown in Figure 7 Figure 7 A structural block diagram of a computer device 100 is provided for an embodiment of the present application. The computer device 100 comprises a short video optimization apparatus based on a knowledge graph 110, a memory 111, a processor 112, and a communication unit 113.
[0345] To realize the transmission or interaction of data, the memory 111, the processor 112, and the communication unit 113 are electrically connected with each other directly or indirectly. For example, the electrical connection between these elements can be realized by one or more communication buses or signal lines. The short video optimization apparatus based on a knowledge graph 110 comprises at least one software function module stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is configured to execute the short video optimization apparatus based on a knowledge graph 110 stored in the memory 111, such as the software function module and the computer program comprised by the short video optimization apparatus based on a knowledge graph 110.
[0346] An embodiment of the present application provides a readable storage medium, which comprises a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the aforementioned short video optimization apparatus based on a knowledge graph 110.
[0347] The foregoing description is made with reference to specific embodiments for purposes of illustration only. The above description is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application and to thereby enable others skilled in the art to best utilize the disclosure, as well as various embodiments with various modifications as are suited to the particular use contemplated.
Claims
1. A short video optimization method based on a knowledge graph, characterized in that, The method comprises the following steps: acquiring a target short video link pointing to a target short video; calling a preset influencer big data service to obtain effect data, basic content and background information of the target short video, and taking the effect data, the basic content and the background information as basic information of the target short video; transcribing the target short video into a script by calling a preset tool set through an influencer big language model; analyzing the title, label and script through the influencer big language model to obtain the theme of the target short video; performing big category cluster matching, subcategory cluster matching, attitude cluster matching and theme cluster matching in a preset influencer knowledge graph based on the effect data, the basic content, the background information, the script and the theme through the influencer big language model, and obtaining target entities and target text blocks according to a preset relevance threshold; taking the target entities and the target text blocks as depth information; inputting the basic information and the depth information into the influencer big language model to obtain optimization suggestions for the target short video; The preset influencer knowledge graph is constructed by the following method, comprising: obtaining video itself information and account information of a plurality of stored videos stored by the preset influencer big data service through the influencer big language model; structuring the video itself information and the account information of each stored video to obtain text blocks corresponding to each stored video, wherein the text blocks include the video itself information and the account information; inputting a plurality of text blocks into the influencer big language model for multi-round extraction to obtain entities, relationships and covariates corresponding to each plurality of text blocks; performing multi-layer cluster extraction according to the entities, the relationships and the covariates to obtain the preset influencer knowledge graph, wherein the preset influencer knowledge graph includes a four-layer graph structure constructed from top to bottom by big category clusters, subcategory clusters, attitude clusters and theme clusters.
2. The method of claim 1, wherein, The effect data includes shelf time, play count, like count, comment count and collection count of the target short video, the basic content includes title, label and audio of the target short video, and the background information includes the creator of the target short video, the head competitor video and the creator of the head competitor video.
3. The method of claim 1, wherein, The influencer big language model is obtained by the following method, comprising: obtaining a basic big language model and an advanced big language model, wherein the model parameter amount of the advanced big language model is greater than that of the basic big language model; performing analysis, self-verification and self-selection extraction strategies on sample short videos based on a self-discovery mechanism through the advanced big language model, and generating knowledge extraction data through prompt engineering; fine-tuning the basic big language model based on the knowledge extraction data in a LoRA manner to obtain the influencer big language model.
4. The method of claim 1, wherein, The method for inputting the basic information and the depth information into the influencer big language model to obtain optimization suggestions for the target short video comprises: Input the basic information and the depth information into the net red large language model to obtain optimization suggestions including a new title, a new label and a new catchphrase for the target short video.
5. A short video optimization device based on a knowledge graph, characterized in that, Comprise: An acquisition module is configured to acquire a target short video link, wherein the target short video link points to a target short video; A preset net red big data service is called to acquire effect data, basic content and background information of the target short video, and the effect data, the basic content and the background information are taken as basic information of the target short video; a preset tool set is called through a net red large language model to transcribe the target short video, and audio is transcribed into a catchphrase; a title, a label and the catchphrase are analyzed through the net red large language model to obtain a theme of the target short video; the effect data, the basic content, the background information, the catchphrase and the theme are used to perform large category cluster matching, subcategory cluster matching, attitude cluster matching and theme cluster matching in a preset net red knowledge graph through the net red large language model, and target entities and target text blocks are obtained according to a preset correlation threshold; the target entities and the target text blocks are taken as depth information; An optimization module is configured to input the basic information and the depth information into the net red large language model to obtain optimization suggestions for the target short video; The preset net red knowledge graph is constructed by the following method, comprising: Video itself information and account information of a plurality of stored videos stored by the preset net red big data service are acquired through the net red large language model; The video itself information and the account information of each stored video are structured to obtain text blocks corresponding to each stored video, wherein the text blocks include the video itself information and the account information; a plurality of text blocks are input into the net red large language model for multi-round extraction to obtain entities, relationships and covariants corresponding to each plurality of text blocks; multi-layer cluster extraction is performed according to the entities, the relationships and the covariants to obtain the preset net red knowledge graph, wherein the preset net red knowledge graph includes a four-layer graph structure constructed from top to bottom of a large category cluster, a subcategory cluster, an attitude cluster and a theme cluster.
6. A computer device, comprising: The computer device comprises a processor and a non-volatile memory storing computer instructions, and the computer instructions are executed by the processor, so that the computer device executes the method of any one of claims 1-4.
7. A readable storage medium characterized by, The readable storage medium comprises a computer program, and the computer program controls the computer device where the readable storage medium is located to execute the method of any one of claims 1-4 when running.
Citation Information
Patent Citations
Video title processing method and device, electronic equipment and readable storage medium
CN111353070A
Video title generation method and device, equipment, storage medium and program product
CN118673905A