Personalized content generation method of content sharing platform
By constructing dynamic prompt word templates and multimodal fusion models in the content sharing platform, combining user historical interaction behaviors and long-term preference distribution, the problem of difficulty in generating personalized style content in the existing technology is solved, and efficient and personalized content generation is achieved.
Patent Information
- Application Number
- CN202510641055.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing technology is difficult to accurately and quickly realize the generation of personalized style content in the content sharing platform, and does not consider the creation of user's "personality".
By constructing a dynamic prompt word template, embed content rules, text style and ‘personality’, and combining the user’s historical interaction behavior, the model input prompt word is enhanced. The multimodal fusion model is used to recall the graphic and text content, and the user demand content is generated through the intention keywords and long-term preference distribution.
The model input quality and personalized adaptability are improved, the generated content is closer to user needs, the user's creative efficiency and the relevance and diversity of content are improved, and personalized and real-time input optimization is achieved.
Smart Images

Figure CN120179913A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for generating personalized content on a content sharing platform. Background Art
[0002] In content sharing platforms such as short videos and social media, more and more creators use large models to generate shared content, greatly improving the efficiency of content creation. Specifically, in CN202311548464.1, "A Method for AIGC Content Operation Based on Large Model Capabilities", a method for AIGC content operation based on large model capabilities is disclosed, which improves the originality of videos and can customize personalized content according to the actual needs of users.
[0003] In content sharing platforms, some users create "personas" to quickly reach potential target users. Therefore, in the process of generating and processing personalized content in existing technical solutions, the "personas" created by users are not considered, making it difficult to accurately and quickly generate personalized style content for users.
[0004] To address the above technical problems, specifically, the present application provides a method for generating personalized content on a content sharing platform. Summary of the Invention
[0005] To achieve the purpose of the present invention, the present invention adopts the following technical solutions: The present application provides a method for generating personalized content on a content sharing platform, specifically including: S1 Construct a dynamic prompt template, embed content rules, text styles, and personas in the dynamic prompt template, and use the dynamic prompt template and the user's historical interaction behavior to enhance the user's input data to obtain the prompt for model input; S2 Use the prompt as the input of the multi-modal fusion model, use the multi-modal recall link of the multi-modal fusion model to recall matching graphic and text content, and add two additional recall links according to the personalized content and user interaction behavior that the user is interested in, respectively. Sort the recall results of different recall links by relevance, and obtain the N recall results with the highest degree of relevance as the instances input to the multi-modal fusion model; S3 Extract and analyze features from the user's input data within a specified time range to generate intent keywords, use the evaluation of the priority of the intent keywords to generate a quantitative result of short-term interest, integrate short-term interest and long-term interest through a weighted model, and optimize by supplementing low-frequency data and debiasing high-frequency data to generate a stable long-term preference distribution. Use the instance and the long-term preference distribution to generate the user's required content.
[0006] The beneficial effects of the present invention are as follows: 1. Dynamic prompt template: Improves the quality of model input and personalized adaptability, extracts the user's historical interaction data, dynamically embeds keywords and background information, makes the generated content closer to the user's needs, optimizes the prompt template with a multi-level structure, enhances the model's understanding ability, and realizes personalized definition of scenario adaptation and expansion. Customizes the prompt template for different application scenarios, improves the matching degree between the prompt and the user's needs, enhances the relevance and diversity of the generated content, realizes personalized and real-time input optimization, makes the model output more in line with the user's expectations, improves the user's creation efficiency, and lowers the threshold of prompt design. In multiple scenarios, it provides high-quality and customized generated content, improving user satisfaction.
[0007] 2. Solves the problem of poor multi-modal content adaptability in traditional RAG technology, expands the application scope of content generation, and enhances the model's ability to handle complex requirements, especially in scenarios that require high-quality graphic and text content (such as social media, marketing copywriting, etc.).
[0008] 3. Realizes the mining of the user's long-term intention and preference insight ability, ensures the coherence and stability of content generation, and improves the user's creation experience.
[0009] In summary, the present invention forms an efficient and accurate content generation system through a dynamic prompt template, multi-modal fusion RAG technology, and context behavior understanding, significantly improving the generation quality and personalization level, and bringing a more intelligent creation experience to users.
[0010] A further technical solution is that the dynamic prompt template includes role, background, task, personalized requirements, and examples.
[0011] A further technical solution is that the historical interaction behaviors include input data, search records, and click operation data.
[0012] A further technical solution is that the method for determining the prompt for model input is as follows: Based on the user's historical interaction behaviors, determine the user's keywords, themes, and preference tags; Dynamically fill the user's keywords, themes, and preference tags into the background and task in the dynamic prompt template to obtain an updated dynamic prompt template; Based on the updated dynamic prompt template, and in combination with the usage scenario, perform dynamic adjustment of the structure of the dynamic prompt template to obtain an adjusted dynamic prompt template, and use the output result of the adjusted dynamic prompt template as the prompt for model input.
[0013] A further technical solution lies in generating a stable long-term preference distribution, specifically including: Based on the historical interaction behaviors of the user, when the number of changed specific behavior feature vectors is greater than a preset feature vector number threshold, perform update processing on the long-term behavior preferences of the user; Use a hybrid model based on time weighting and frequency weighting to integrate short-term behavior preferences and long-term behavior preferences, and through the supplementation of low-frequency data and the debiasing optimization of high-frequency data, generate the corrected long-term behavior preferences and use them as the stable long-term preference distribution.
[0014] A further technical solution lies in performing generation processing on the demand content of the user, specifically including: Use the instance and the long-term preference distribution as the input of the model, and use the output of the model to perform generation processing on the demand content of the user.
[0015] Other features and advantages will be described in the subsequent specification. The objectives and other advantages of the present invention are achieved and obtained through the structures specifically pointed out in the specification and the drawings.
[0016] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically presents preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. Description of the Drawings
[0017] By referring to the drawings and describing its exemplary embodiments in detail, the above and other features and advantages of the present invention will become more obvious.
[0018] Figure 1 is a flowchart of a method for generating personalized content on a content sharing platform; Figure 2 is a flowchart of a method for determining the prompt words for model input; Figure 3 is a flowchart of a method for recalling matching graphic and text content; Figure 4 is a flowchart of a method for determining relevance ranking. Specific Embodiments
[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0020] Prompt: The prompt is the text input by the user into the language model.
[0021] BERT: BERT is an open-source machine learning framework. According to the nature of the task, the BERT model can capture the context information in the text and extract important features from it.
[0022] BERTopic: A text clustering tool based on BERT.
[0023] RAG: Retrieval-augmented Generation. When the model needs to generate text or answer questions, it first retrieves relevant information from a large collection of documents and then uses this retrieved information to guide the generation of text, thereby improving the quality and accuracy of the prediction.
[0024] Embedding: A technique for mapping high-dimensional data into a low-dimensional space, mainly used to convert discrete and sparse data into continuous and dense vector representations so that machine learning or deep learning models can better process and understand this data.
[0025] Cross-Encoder: A technique used in natural language processing (NLP) to calculate the similarity between two texts. It encodes the two texts simultaneously and generates a value between 0 and 1, representing the similarity of the input pair.
[0026] CLIP / BLIP: Multimodal large models.
[0027] The present invention aims to use multiple technical means and combine various user interaction behaviors. By constructing a dynamic prompt template and embedding requirements such as content rules, text styles, and persona / role directions in the prompt template, the prompts input to the model are enhanced from multiple dimensions. At the same time, relying on user behavior feedback, the real content creation requirements of the user are understood from the user's continuous input and behaviors, thereby enhancing the diversity and personalization of the content generated by the model and meeting the user's need to generate content that meets the "persona / role" requirements.
[0028] As Figure 1 shown, the present application provides a personalized content generation method for a content sharing platform, specifically including: S1 Construct a dynamic prompt template, embed content rules, text styles, and personas in the dynamic prompt template, and use the dynamic prompt template and the user's historical interaction behaviors to enhance the user's input data to obtain the prompts input to the model; Further, the dynamic prompt template includes roles, backgrounds, tasks, personalized requirements, and examples.
[0029] Specifically, the historical interaction behaviors include input data, search records, and click operation data.
[0030] It should be noted that, as Figure 2 shown, the method for determining the prompt words input into the model is as follows: Based on the historical interaction behaviors of the user, determine the user's keywords, topics, and preference tags; Dynamically fill the user's keywords, topics, and preference tags into the background and tasks in the dynamic prompt word template to obtain an updated dynamic prompt word template; Based on the updated dynamic prompt word template, and in combination with the usage scenario, perform dynamic adjustment of the structure of the dynamic prompt word template to obtain an adjusted dynamic prompt word template, and use the output result of the adjusted dynamic prompt word template as the prompt word input into the model.
[0031] . Dynamic Prompt Word Template: The prompt word is the core text content finally input into the model for inference calculation, and its quality directly affects the quality and relevance of the content generated by the model. The dynamic prompt word template provides the model with richer and more accurate input information by introducing a multi-level and customized structure to meet the personalized and diverse needs of users. Its construction method is based on the CRISPE prompt word framework and then undergoes customization. The original prompt word framework is defined as: ● Capacity and Role: The role that the model is required to play, such as a creator, a writer, etc. ● Insight (Background): Provide the model with the background of this role, such as when playing a creator, it is necessary to create a certain type of copywriting, etc. ● Statement (Task): The task that the model is required to complete ● Personality: Specify the style, manner, format, etc. of the model's return ● Experiment (Example): Provide multiple example references for the model Based on the original CRISPE framework, the dynamic prompt template introduces more dynamic information embeddings based on user behavior and context, thus enhancing personalization and scenario adaptability. In the improved dynamic prompt structure, effective information extracted from the user's historical interaction behaviors (input, search, click) is added to represent the potential direction of the user's content creation, and the model is explicitly required to refer to the provided keyword information in the Statement (task). When the user creates content, relevant information is dynamically filled into the improved prompt template according to different scenarios and the user's own attributes, achieving a personalized and diverse effect. The improvements are as follows: a. Extract effective information from user history interactions: The system analyzes the user's historical interaction behaviors (such as input content, search records, click operations, etc.) to extract keywords, themes, and preference tags. These information are used as identifiers of the user's potential creation direction and are dynamically filled into the prompt template. Natural language processing technologies such as TF-IDF and BERT are used to extract high-frequency keywords, and topic models (such as BERTopic) are combined to extract the user's main interest areas.
[0032] b. Dynamically embed keywords and background information: Dynamically fill the keywords and background information extracted from user behavior into Insight (background) and Statement (task), making the prompt closer to the user's current creation needs. For example, if the user has recently searched for "science fiction writing skills", relevant descriptions and task requirements related to the science fiction theme will be added to the prompt template.
[0033] c. Scenario adaptation: Dynamically adjust the prompt template structure according to different usage scenarios (such as copywriting, social media operation, academic research, etc.). For example, for the platform attributes, format requirements (such as Xiaohongshu style) can be specified in Personality (personalization), while in the creative writing scenario, the freedom of style can be emphasized.
[0034] d. Extended Personality definition: In addition to style and format, Personality also supports output methods customized according to the user's historical preferences, such as adding certain specific expression habits and citing certain specific resources, further enhancing the personalized output effect.
[0035] e. Explicit keyword reference: The model is explicitly required to refer to the extracted keywords in the Statement (task) and construct the prompt through natural language. For example: Original: Write an article about "outdoor wear". Dynamic prompt: Write an article about "outdoor wear" with an emphasis on "clothing styles" and "matching styles".
[0036] f. Enhancement of Examples: If examples need to be provided, the system will dynamically generate customized examples related to the user's needs instead of using general static examples. These examples can be generated based on the closest creation records in the user's historical interactions, thus being more in line with the user's habits.
[0037] S2 uses the prompt as the input of the multimodal fusion model, utilizes the multimodal recall link of the multimodal fusion model to recall the matching graphic and text content, and at the same time adds two recall links respectively according to the personalized content and user interaction behavior that the user is concerned about, sorts the recall results of different recall links according to relevance, and obtains the N recall results with the highest degree of relevance as the instances input to the multimodal fusion model; Specifically, as Figure 3 shown, the method for recalling the matching graphic and text content is: Use a deep learning model to extract features from images, texts, and videos in different types of files, and determine the low-dimensional vector representation based on the extracted features using a language model; Based on the vector retrieval library, perform a preliminary retrieval on the low-dimensional vector representation to obtain a retrieval result, and perform multimodal alignment processing on the retrieval results of different types of files to obtain the associated vectors of different types of files in different modalities; Perform weighted processing on the associated vectors of different types of files in different modalities to obtain the recall result of the file, and use the recall result to determine the matching graphic and text content in the file.
[0038] Optionally, the matching graphic and text content in the file is a file whose recall result is greater than a preset associated degree threshold.
[0039] It should be noted that adding two recall links respectively specifically includes: Based on the historical interaction behavior of the user, obtain the behavior data of the user, and use the behavior data to generate a behavior feature vector; Combined with the deep learning recommendation algorithm, use a factorization machine to capture the high-order interactions between the behavior feature vectors, and model through a deep neural network to obtain the long-term behavior preferences of the user; Use the dynamic query expansion technology to perform dynamic optimization queries according to the behavior feature quantity obtained from the user's instant interaction behavior to obtain the short-term behavior preferences of the user; Dynamically adjust in combination with context information, respectively index by the user's long-term behavior preferences and by the user's short-term behavior preferences, add two recall links, and use the recall links to determine the similarity of the feature vectors of the user's long-term behavior preferences and short-term behavior preferences. Generate a link recall link based on the similarity to obtain a recall result to ensure that the output content is more personalized and real-time.
[0040] Specifically, as Figure 4 shown, the method for determining the relevance ranking is: For the recall results of different recall links, use the BERT or Cross-Encoder model to calculate the semantic similarity of the text content between the recall results, use the CLIP model to calculate the relevance between the image and text of the recall results, and use the visual features based on ResNet and TimeSformer to judge the relevance of the video frames and time series between the recall results; Use the semantic similarity, the relevance of the text, and the relevance of the time series to determine the association degree of different recall results; Based on the association degree, perform a relevance ranking of different recall results.
[0041] Further, the N is a constant and is determined according to a preset value.
[0042] Optionally, the value range of N is between 5 and 10, and is specifically determined according to the number of times of generating and processing the current demand content of the user and the association degree of different recall results. The more times the user generates and processes the current demand content, and the lower the association degree of the recall results, the larger the value range of N.
[0043] It can be understood that the specific steps for determining N are: Determine the number of times of generating and processing the user's input data for the current demand content, and use it as the matching generation and processing times, and use the semantic similarity between different matching generation and processing times and the intent keywords of the current user's input data to determine the semantically similar generation and processing times in the matching generation and processing times; According to the preset association threshold corresponding to the semantically similar generation and processing times, determine the association degree threshold of the recall results; Based on the association degree of different recall results and the association degree threshold of the recall results, determine the value of N.
[0044] Further, the semantically similar generation and processing times are the matching generation and processing times in which the proportion of the number of times semantically similar to the intent keywords of the current user's input data is greater than the preset proportion.
[0045] Specifically, determining the value of N based on the correlation degree of different recall results and the correlation degree threshold of the recall results specifically includes: Select the solution with the least number of recall results whose sum of the correlation degrees of the recall results with the highest correlation degree is greater than the correlation degree threshold of the recall results, and use the number of recall results of the solution with the least number of recall results as the value of N.
[0046] Optionally, the specific steps for determining N are as follows: S21 Determine the semantic similarity of the intent keywords of the user's input data among different generation processing times of the current required content according to the intent keywords of the user's input data, and use the semantic similarity of the intent keywords among different generation processing times to determine the semantic deviation coefficient of the intent keywords among different generation processing times. Based on the weighted sum of the semantic deviation coefficients of different generation processing times, determine the matching deviation value of the generation result of the required content; S22 Obtain the correlation degree of different recall results, and determine the correlation coefficient of the recall results with the average value of the correlation degrees of different recall results; S23 Based on the matching deviation value of the generation result of the required content and the correlation coefficient of the recall results, determine the demand deviation amount of the recall results, and use the demand deviation amount of the recall results to determine the value of N.
[0047] Furthermore, the demand deviation amount of the recall results is determined according to the average value of the matching deviation value of the generation result of the required content and the correlation coefficient of the recall results.
[0048] It should be noted that using the demand deviation amount of the recall results to determine the value of N specifically includes: Based on the demand deviation amount of the recall results, determine the preset demand deviation interval corresponding to the demand deviation amount of the recall results, and use the preset value corresponding to the preset demand deviation interval to determine the value of N.
[0049] Optionally, the above step S21 includes the following content: S211 Obtain the generation processing times of the user for the current required content. When the generation processing times of the user for the current required content are greater than the preset times threshold, then use a specified quantity to determine the value of N. When the generation processing times of the user for the current required content are not greater than the preset times threshold, proceed to step S212; S212 determines the number of times of generating and processing the input data of the user for the current required content, and uses it as the matching generation and processing times. Then, based on the semantic similarity between different matching generation and processing times and the intention keywords of the current user's input data, it determines the semantically similar generation and processing times among the matching generation and processing times. When the semantically similar generation and processing times are greater than the preset generation and processing times threshold, it determines the value of N using a specified quantity. When the semantically similar generation and processing times are not greater than the preset generation and processing times threshold, it proceeds to step S213; S213 determines the semantic deviation coefficient between the intention keywords of different generation and processing times of the current required content based on the input data of the user, and also uses the semantic similarity of the intention keywords between different generation and processing times. When the average value of the semantic deviation coefficients between the intention keywords of different generation and processing times is less than the preset deviation coefficient threshold, it proceeds to step S214. When the average value of the semantic deviation coefficients between the intention keywords of different generation and processing times is not less than the preset deviation coefficient threshold, it proceeds to step S215; S214 When the number of times of generating and processing the current required content by the user is within the preset generation and processing times range, it determines the value of N using a specified quantity. When the number of times of generating and processing the current required content by the user is not within the preset generation and processing times range, it proceeds to step S215; S215 determines the matching deviation value of the generation result of the required content based on the weighted sum of the semantic deviation coefficients of different generation and processing times. When the matching deviation value of the generation result of the required content is greater than the preset deviation threshold, it determines the value of N using a specified quantity. When the matching deviation value of the generation result of the required content is not greater than the preset deviation threshold, it proceeds to step S22.
[0050] Optionally, the following content is included in the above step S22: S221 obtains the number of the recall results. When the number of the recall results is less than the preset number of recall results, all the recall results are used as instances input to the multi-modal fusion model. When the number of the recall results is not less than the preset number of recall results, it proceeds to step S222; S222 determines the correlation coefficient of the recall results based on the average value of the correlation degrees of different recall results. When the correlation coefficient of the recall results is less than the preset correlation coefficient threshold, it proceeds to step S223. When the correlation coefficient of the recall results is not less than the preset correlation coefficient threshold, it proceeds to step S23; S223 When the matching deviation value of the generation result of the required content is within the preset deviation value range, the value of N is determined using a specified quantity. When the matching deviation value of the generation result of the required content is not within the preset deviation value range, proceed to step S23.
[0051] Multimodal Fusion RAG (Retrieval-Augmented Generation): The examples provided to the Experiment (Example) module in the dynamic prompt template enable the model to generate content that better suits the platform style by referring to them. Since traditional RAG technology mainly focuses on text content for processing and retrieval and has poor retrieval effects on multimodal content such as images and videos, and the high-quality content on this platform mainly consists of text-image content. Therefore, a multimodal recall link is introduced in this stage to recall high-quality text-image content, and at the same time, two additional recall links are added for personalized content that users are interested in and user interaction behaviors respectively. Finally, the relevance of all recalled results is sorted, and the top K content results are obtained as instances input to the model. The technical details are as follows: a. Multimodal Index Construction: Multimodal indexing is the basis for efficient content recall. First, data preprocessing is performed on text, images, and videos. Text generates Embeddings through BERT; visual features are extracted from images and videos using networks such as ResNet, ViT, or TimeSformer. These features are stored in a vector database (such as FAISS or Milvus) for indexing. During the index construction process, the HNSW algorithm is used to create an efficient nearest neighbor index. For each modality, an independent sub-index is established to ensure optimized results for queries of each modality. Then, the index results of different modalities are unified through a weighted fusion strategy (such as linear fusion or neural network fusion). The purpose of this is to generate a comprehensive score based on the features of different modalities during retrieval, thereby obtaining the most relevant results. To ensure the real-time and accuracy of data, the index design supports dynamic updates. When new content is added or existing content is modified, the index can be updated immediately to ensure the timeliness of query results. Combining the text retrieval capabilities of ElasticSearch enables efficient comprehensive retrieval in multimodal data.
[0052] b. Multi-modal Content Retrieval: The goal of multi-modal content retrieval is to uniformly map text, image, and video content into a shared semantic space to ensure that data of different modalities can be effectively retrieved and ranked. First, deep learning models such as CLIP and BLIP are used to extract features from images and text, and images are processed by Vision Transformer (ViT) or ResNet, and videos are processed by TimeSformer for feature extraction. For text, language models such as BERT are used to generate text Embeddings. In the retrieval stage, these Embeddings are initially retrieved through an efficient vector retrieval library such as FAISS. FAISS supports the HNSW (Hierarchical Navigable Small World) algorithm to accelerate vector retrieval and improve search efficiency. The retrieval results are subjected to multi-modal alignment and fusion, and a weighted strategy is usually adopted to integrate the relevance of different modalities. For example, after jointly modeling the relevance between text and image through CLIP, a weighted score is obtained.
[0053] c. Personalized Content Retrieval: The core of personalized content retrieval is to provide content recommendations that meet the user's needs by analyzing the user's behavior data. First, the user's behavior feature vectors are generated through the user's behavior data (such as clicks, dwell time, browsing history). These features are jointly modeled with the content features through the Attention mechanism to highlight the user's personalized needs. In the retrieval stage, collaborative filtering and deep learning recommendation algorithms (such as DeepFM) are combined to enhance the model's understanding of user preferences. The DeepFM model uses the factorization machine (FM) to capture the high-order interactions between features, and at the same time models more complex user interest patterns through a deep neural network. The dynamic query expansion technology (such as the Seq2Seq model) is used to dynamically optimize the query according to the user's immediate needs and generate a recall request that better matches the current interest. Finally, the sorting and filtering of the recall results are dynamically adjusted in combination with context information (such as time, location, etc.) to ensure that the output content is more personalized and real-time. By adaptively adjusting the sorting strategy, the weight of each recall can be adjusted according to the user's historical behavior and preferences, thereby improving the user experience.
[0054] d. Relevance Sorting: Relevance sorting is to screen out the results that best meet the user's needs from the retrieved content. First, models such as BERT or Cross-Encoder are used to calculate the semantic similarity of the text content to ensure that the text results can be closely aligned with the user's needs. For image and video content, models such as CLIP are used to calculate the relevance between the image and the text, or visual features based on ResNet and TimeSformer are used to judge the relevance of video frames and time series. For video content, by analyzing the visual features of each frame and the relevance of the time series, it can be ensured that the video segment can accurately match the user's needs. For example, an RNN-based time series model can be used to capture the long-term and short-term dependencies of the video, and the sorting weights can be adjusted in combination with the user's context information. Finally, the model will select the top K results with the highest scores as the input instances of the generation model. This sorting process ensures that users get the most relevant and current-demand-matching note content, improving the accuracy and personalization of the model output quality.
[0055] S3 performs feature extraction and analysis on the user's input data within a specified time range to generate intent keywords, generates a quantitative result of short-term interest using the evaluation of the priority of the intent keywords, integrates the short-term interest and long-term interest through a weighted model, and generates a stable long-term preference distribution through the supplementation of low-frequency data and the debiasing of high-frequency data. The generation processing of the user's required content is performed using the said instance and the long-term preference distribution.
[0056] Further, the specified time range is the period during the generation processing of the user's current required content.
[0057] Specifically, the specified time range is determined according to the change situation of the instances during the generation processing of the user's current required content. Among them, during the generation processing of the user's current required content, the more the number of changes in the instances between different generation processing times and the current generation processing time, the longer the duration corresponding to the specified time range.
[0058] Specifically, the specific steps for determining the specified time range are as follows: Take the number of times of the generation processing of the user's input data in the current required content generation processing as the matching generation processing times, and obtain the deviation situation between the instances of different matching generation times and the instance of the current generation processing times; Use the deviation situation to determine the deviation quantity and the consistent quantity of the instances of different matching generation times, and combine the deviation situation of the correlation degree of different consistent quantities to determine the correlation degree deviation coefficient between different matching generation times and the current generation processing times; Determine the historical generation deviation coefficient by averaging the degree-of-association deviation coefficients between different numbers of matching generations and the current number of generation processes, and use the historical generation deviation coefficient to determine the reference number of matching generations. Based on the time period corresponding to the reference number of matching generations, determine the specified time range.
[0059] Further, using the historical generation deviation coefficient to determine the reference number of matching generations specifically includes: Based on the historical generation deviation coefficient, determine the minimum required number of matching generations corresponding to the historical generation deviation coefficient; Based on the minimum required number of matching generations corresponding to the historical generation deviation coefficient, determine the reference number of matching generations.
[0060] It should be noted that based on the time period corresponding to the reference number of matching generations, determining the specified time range specifically includes: Use the time period corresponding to the reference number of matching generations as the specified time range.
[0061] It can be understood that the specific steps for determining the specified time range are: Take the number of generation processes of the user's input data in the current required content as the number of matching generation processes, and obtain the deviation situation between different instances of the number of matching generations and the current instance of the number of generation processes; Use the deviation situation to determine the deviation quantity of different instances of the number of matching generations; Determine the sum of the deviation quantities of different instances of the number of matching generations in different time periods, and combine a preset deviation quantity threshold to determine the specified time range.
[0062] Further, the specified time range is the time period corresponding to the sum of the deviation quantities of different instances of the number of matching generations being greater than the preset deviation quantity threshold.
[0063] Optionally, the specific steps for determining the specified time range are: Take the number of generation processes of the user's input data in the current required content as the number of matching generation processes. When the number of matching generation processes is less than the preset number limit value, then use all the time periods corresponding to the number of matching generations as the specified time range; When the number of matching generation processes is not less than the preset number limit value: Obtain the deviation situation of different instances of the number of matching generations. When the deviation quantities of different instances of the number of matching generations from the current production process are all greater than the preset deviation instance quantity threshold, then use all the time periods corresponding to the number of matching generations as the specified time range; When there are matching generation times with the deviation quantity from the current production processing times not greater than the preset deviation instance quantity threshold, determine the deviation quantity and the consistent quantity of instances with different matching generation times using the deviation situation, and in combination with the deviation situation of the correlation degree of different consistent quantities, determine the correlation degree deviation coefficient between different matching generation times and the current generation processing times. When the correlation degree deviation coefficients between different matching generation times and the current generation processing times are all greater than the preset correlation deviation coefficient threshold, then take the time periods corresponding to all the matching generation times as the specified time range; When there are matching generation times with the correlation degree deviation coefficient from the current generation processing times not greater than the preset correlation deviation coefficient threshold: When the correlation degree deviation coefficients between different matching generation times and the current generation processing times are all within the preset correlation deviation coefficient interval, determine the reference matching generation times according to the preset quantity, and determine the specified time range based on the time periods corresponding to the reference matching generation times; When there are matching generation times with the correlation degree deviation coefficient from the current generation processing times not within the preset correlation deviation coefficient interval: Determine the historical generation deviation coefficient through the average value of the correlation degree deviation coefficients between different matching generation times and the current generation processing times, and use the historical generation deviation coefficient to determine the reference matching generation times, and determine the specified time range based on the time periods corresponding to the reference matching generation times.
[0064] It should be noted that generating a stable long-term preference distribution specifically includes: Based on the historical interaction behavior of the user, when the number of changes in specific behavior feature vectors is greater than the preset feature vector quantity threshold, perform the update processing of the user's long-term behavior preference; Use a hybrid model based on time weighting and frequency weighting to integrate the short-term behavior preference and the long-term behavior preference, and through the supplementation of low-frequency data and the debiasing optimization of high-frequency data, generate the corrected long-term behavior preference and use it as the stable long-term preference distribution.
[0065] Specifically, performing the generation processing of the user's demand content specifically includes: Take the instance and the long-term preference distribution as the input quantities of the model, and use the output quantity of the model to perform the generation processing of the user's demand content.
[0066] In-Context / Action Understanding: The user's input and interaction behaviors can reflect the user's interests and creative intentions. Therefore, continuously generating content that conforms to the "persona / character" requires identifying and understanding the user's long-term input and related interaction behaviors. To achieve this goal, the intention understanding process relies on the comprehensive management of session storage and the structured processing of input / interaction behaviors. The input and interaction behaviors are converted into a unified data structure and stored and retrieved according to the time dimension and feature dimension. When analyzing the user's creative intention, the system performs feature extraction and analysis based on the data within a specified time range, generates intention phrases or keywords, and evaluates their priorities. The long-term interest preferences are calculated through a weighted model, which serves as the core information source for personalized recall and persona refinement in the dynamic template. The following are the specific implementation technical details: a. Behavior Information Completion: The original interaction behaviors usually only contain basic information such as the ID of the interaction object (e.g., blogger, note) and timestamp. To ensure the comprehensiveness of context understanding, the system retrieves the complete information of the interaction object from the database through the ID field during storage, including key attributes such as persona description, creative style, associated tags, and categories, and maps them to the behavior records. According to the characteristics of the interaction object, the object is classified into behaviors using clustering algorithms or topic modeling methods to ensure that the stored data has available classification dimensions.
[0067] b. Session Memory Management: The session memory stores the user's historical input and interaction behavior data, and uses a storage component that supports hybrid retrieval and time-range retrieval at the bottom layer. The memory module provides functions for adding, deleting, modifying, and querying memories, and can quickly obtain relevant context records under keywords for a certain time period. The session memory supports quickly retrieving relevant context records within a specific time period based on keywords, and at the same time provides a memory weight decay mechanism (such as dynamically adjusting the relevance weight according to the time distance).
[0068] c. Short-Term Intention Understanding: Use the current input information and short-term context session memory to extract features from text, image, and video content through pre-trained models (such as CLIP and BLIP), and use the Cross-Attention mechanism to achieve cross-modal alignment. Combining the short-term context session memory and the feature extraction results, generate keywords and relevance priorities for the user's creative intention, and quantify the priorities using a scoring model based on Attention weights (such as the weight mechanism of the Transformer model).
[0069] d. Long-term interest preferences: When the user interaction behavior reaches a critical threshold (such as repeated browsing of specific content, frequent occurrence of specific types of creations, etc.), the update process of long-term interest preferences is triggered. The system uses a hybrid model based on time weighting and frequency weighting to integrate short-term interests with long-term interests, and generates a stable long-term preference distribution through the supplementation of low-frequency data and the debiasing optimization of high-frequency data. A topic attribution model (such as BERTopic) is used to map user behavior data to interest topics, generating an interpretable interest preference distribution map for subsequent content generation and persona improvement modules.
[0070] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0071] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0072] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for generating personalized content on a content sharing platform, characterized in that: Specifically include: Constructing a dynamic prompt word template, embedding content rules, text style, and personality in the dynamic prompt word template, and enhancing the user's input data using the dynamic prompt word template and the user's historical interaction behavior to obtain the prompt word input by the model; The prompt word is used as the input of the multimodal fusion model, and the multimodal recall link of the multimodal fusion model is used to recall the matching graphic content. At the same time, two recall links are added according to the personalized content concerned by the user and the user interaction behavior, and the recall results of different recall links are sorted by relevance, and the N recall results with the highest relevance are obtained as the instance of input to the multimodal fusion model; Feature extraction and analysis are performed based on the user's input data within a specified time range to generate intent keywords, and the quantitative results of short-term interests are generated by utilizing the priority evaluation of the intent keywords. Short-term interests and long-term interests are integrated through a weighted model, and a stable long-term preference distribution is generated through the supplementation of low-frequency data and the debiasing optimization of high-frequency data. The instance and the long-term preference distribution are utilized to generate the user's demand content.
2. The method for generating personalized content for a content sharing platform as claimed in claim 1, characterized in that: The dynamic prompt word template includes roles, backgrounds, tasks, personalized requirements, and examples.
3. The personalized content generation method of the content sharing platform as claimed in claim 1, characterized in that: The historical interaction behaviors include input data, search records, and click operation data.
4. The method for generating personalized content for a content sharing platform as claimed in claim 1, characterized in that: The method for determining the prompt word input by the model is: Based on the user's historical interaction behavior, determine the user's keywords, topics, and preference tags; Dynamically filling the user's keywords, topics, and preference tags into the background and tasks in the dynamic prompt word template to obtain an updated dynamic prompt word template; Based on the updated dynamic prompt word template, the structure of the dynamic prompt word template is dynamically adjusted in combination with the usage scenario to obtain an adjusted dynamic prompt word template, and the output result of the adjusted dynamic prompt word template is used as the prompt word input to the model.
5. The personalized content generation method of the content sharing platform as claimed in claim 1, characterized in that: The method for recalling the matching graphic content is: Use deep learning models to extract features from images, texts, and videos in different types of files, and use language models to determine low-dimensional vector representations based on the extracted features; Performing a preliminary search on the low-dimensional vector representation based on a vector search library to obtain a search result, and performing a multi-modal alignment process on the search results of different types of files to obtain association vectors of different types of files in different modalities; After weighted processing is performed on association vectors of different types of files in different modes, recall results of the files are obtained, and matching graphic and text contents in the files are determined using the recall results.
6. The method for generating personalized content for a content sharing platform as claimed in claim 5, characterized in that: The matching graphic content in the file is a file whose recall result is greater than a preset correlation degree threshold.
7. The method for generating personalized content for a content sharing platform as claimed in claim 1, characterized in that: Two recall links are added respectively, including: Based on the historical interaction behavior of the user, acquiring the behavior data of the user, and generating a behavior feature vector using the behavior data; Combined with the deep learning recommendation algorithm, the factorization machine is used to capture the high-order interactions between the behavioral feature vectors, and the long-term behavioral preferences of the user are obtained through deep neural network modeling; Using dynamic query expansion technology, dynamically optimize the query based on the behavioral feature quantity obtained from the user's real-time interactive behavior to obtain the user's short-term behavioral preference; Combined with dynamic adjustment of context information, two recall links are added with the user's long-term behavior preference as the index and the user's short-term behavior preference as the index, and the recall links are used to determine the similarities between the feature vectors with the user's long-term behavior preference as the index and the user's short-term behavior preference. Based on the similarities, a link recall link is generated to obtain the recall result, ensuring that the output content is more personalized and real-time.
8. The method for generating personalized content for a content sharing platform as claimed in claim 1, characterized in that: The value range of N is between 5 and 10, and is specifically determined based on the number of times the user generates and processes the current demand content and the degree of correlation between different recall results. The more times the user generates and processes the current demand content and the lower the degree of correlation between the recall results, the larger the value range of N.
9. The method for generating personalized content for a content sharing platform as claimed in claim 1, characterized in that: Generate a stable long-term preference distribution, including: Based on the historical interaction behavior of the user, when the number of changes in a specific behavior feature vector is greater than a preset feature vector number threshold, the user's long-term behavior preference is updated. A hybrid model based on time weighting and frequency weighting is used to integrate short-term behavioral preferences with long-term behavioral preferences. By supplementing low-frequency data and debiasing optimization of high-frequency data, the corrected long-term behavioral preferences are generated and used as a stable long-term preference distribution.
10. The method for generating personalized content for a content sharing platform according to claim 1, characterized in that: The generation process of the user's demand content specifically includes: The instance and the long-term preference distribution are used as the input of the model, and the output of the model is used to generate the user's demand content.
Citation Information
Patent Citations
Information recommendation method and device based on artificial intelligence, electronic equipment and storage medium
CN111444428A
Information recommendation method and device, computer equipment, storage medium and program product
CN114329176A
Resource recall method
CN117743673A
Comparative learning-based interpretable personalized recommendation method
CN118312671A
Creation model automatic training method based on user behavior preference data
CN118965037A
Cited By
E-commerce commodity image-text generation optimization system based on generative AI
CN120472050A
Question and answer generation method and device, electronic equipment and storage medium
CN121434340A