AIGC-Based Method and System for Automatic Generation and Push of Personalized Content
By obtaining user behavior data and using large language models to generate user portraits, and combining multimodal generation models to automatically generate content, the existing recommendation system lacks understanding of user interests, and realizes accurate personalized recommendation and cross-modal content generation.
Patent Information
- Application Number
- CN202510279598.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing recommendation systems lack a deep semantic understanding of user interests, and the multimodal generation technology has shortcomings in the combination of theme consistency and user portraits.
By obtaining the user's historical behavior data, using a large language model to generate user portrait tags and personalized prompt words, combining multiple generation models to automatically generate multimodal content, and optimizing recommendation strategies through feedback mechanisms.
It realizes accurate personalized recommendations, improves the relevance and creative ability of content, has cross-modal content generation capabilities, and continuously optimizes user portraits and recommendation strategies through a closed-loop feedback mechanism.
Smart Images

Figure CN119807538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence-generated content, and specifically to a method and system for automatically generating and pushing personalized content based on AIGC. Background Art
[0002] With the rapid development of artificial intelligence and big data technologies, personalized recommendation based on user interests and behaviors has become an important development trend in the current content industry.
[0003] Existing recommendation systems (such as Netflix's collaborative filtering algorithm and YouTube's deep neural network recommendation) mainly rely on user behavior data (clicks, views) or static tag systems (such as "action movies", "comedies"), lacking in-depth semantic understanding of user interests (reference: Koren et al, 2009).
[0004] Although multi-modal generation technologies (such as DALL·E, Midjourney) can generate graphic, text, or audio content, the theme consistency between modalities is poor (such as the music and video styles generated conflicting), and they are not dynamically combined with user portraits.
[0005] In view of the above problems, how to design an encrypted traffic classification method that takes into account both improving the classification accuracy and utilizing the structural information of the traffic itself is an urgent problem to be solved in this field. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for automatically generating and pushing personalized content based on AIGC. It is a personalized content generation and pushing system that combines AIGC technology, and realizes precise content customization and pushing through user portraits and content automatic generation technology.
[0007] To achieve the above purpose, the present invention is realized through the following technical solutions:
[0008] On the one hand, the present invention provides a method for automatically generating and pushing personalized content based on AIGC, including the following steps:
[0009] S1: Obtain data information of the user's historical viewing records, browsing behaviors, and interaction records;
[0010] S2: According to the data information obtained in step S1, analyze the data through a large language model to generate user portrait tags and personalized prompt words;
[0011] S3: Based on the user portrait tags and personalized prompt words generated in step S2, select and generate corresponding creation prompts;
[0012] S4: Based on the creation prompts generated in step S3, complete the automatic generation of multimodal content through multiple generation models;
[0013] S5: Push the automatically generated multimodal content in a personalized manner and collect feedback.
[0014] Preferably, the data information obtained about the user's historical viewing records includes, but is not limited to:
[0015] The movies, TV shows, videos, music, articles, and picture content viewed by the user;
[0016] The data information obtained about the user's browsing behavior includes, but is not limited to:
[0017] The user's search records and browsing trajectories;
[0018] The data information obtained about the user's interaction records includes, but is not limited to:
[0019] The user's click, favorite, like, and comment behavior data.
[0020] Preferably, in step S2, generating user portrait tags specifically includes:
[0021] Model the user behavior sequence using the multi-head attention mechanism based on Transformer:
[0022] ;
[0023] Among them, is the query matrix, is the key-value matrix, is the vector dimension;
[0024] By calculating the attention weights at different time steps in the behavior sequence, dynamically extract the long-term dependencies of the user's interests. Finally, the user portrait is represented as:
[0025] ;
[0026] Among them, is a three-layer fully connected network, the activation function is GELU, and it outputs the probability distribution of the user tags.
[0027] Preferably, in step S2, generating personalized prompt words specifically includes:
[0028] Use the mapping function from tags to prompt words and optimize the generation quality by combining reinforcement learning:
[0029] ;
[0030] Among them, For the semantic similarity between the prompt and the user tags, For the diversity score of the prompt, 、 are weight hyperparameters.
[0031] Preferably, in the step S3, the corresponding creation prompt is selected and generated, specifically: according to the media type, prompts for text-to-image, text-to-audio, and image-to-video are generated. By jointly training the text-to-image, text-to-audio, and image-to-video models, a cross-modal joint loss function is used:
[0032] ;
[0033] Among them, is the image-text alignment loss calculated by the CLIP model, is the cycle consistency loss, 、 are balance coefficients.
[0034] Preferably, the personalized push is specifically as follows:
[0035] The push strategy is executed using the context bandit algorithm with feedback:
[0036] ;
[0037] Among them, is the content type of the push, is the user context feature; is the weight parameter of the content type of the push, is the noise term, and the parameters are updated through online learning , maximizing the user click-through rate and viewing duration.
[0038] Preferably, the feedback collection is specifically as follows:
[0039] The interaction feedback and behavior data of the user are collected in real time for further optimizing the user portrait and recommendation algorithm. The feedback-driven model incremental update mechanism is: The loss function of the feedback-driven model based on the user feedback data is:
[0040] ;
[0041] Among them, is the actual feedback of the user on the content , a click is recorded as (1), and no click is recorded as (0); is the predicted click probability of the model.
[0042] On the other hand, a push system based on the above AIGC-based personalized content automatic generation and push method is also provided, including:
[0043] A data acquisition module, configured to: obtain data information of a user's historical viewing records, browsing behaviors, and interaction records;
[0044] A data processing module, configured to: analyze the data through a large language model according to the obtained data information, generate user portrait tags and personalized prompt words, and select and generate corresponding creation prompts;
[0045] An automatic generation module, configured to: complete the automatic generation of multimodal content based on the generated creation prompts through multiple generation models;
[0046] A personalized push and feedback module, configured to: perform personalized push on the automatically generated multimodal content and collect feedback.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] 1. Cross-modal content generation: The present invention can generate content in multiple modalities according to user needs (such as text-to-image, text-to-audio, image-to-video, etc.), has strong content creation capabilities, and can meet the content form requirements of different users.
[0049] 2. Precise personalized recommendation: Based on the accurate user portrait tags and prompt words generated by the large language model, it can help the system better understand user interests and improve the relevance and accuracy of recommended content.
[0050] 3. Closed-loop feedback mechanism: By collecting user feedback, the system can continuously optimize the user portrait and recommendation strategy, achieve self-iteration of content recommendation, and improve user satisfaction and platform stickiness.
[0051] 4. Efficient content generation and push: Through the automated content generation and push process, it greatly reduces manual intervention and creation time, and improves content production efficiency. Description of the Drawings
[0052] Figure 1 is the method flow chart of the present invention;
[0053] Figure 2 is the system structure schematic diagram of the present invention. Detailed Embodiments
[0054] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0055] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationships of various components or elements of the present invention, and do not specifically refer to any component or element in the present invention, and should not be construed as a limitation to the present invention.
[0056] Example:
[0057] As Figure 1 shown, this embodiment provides a method for automatically generating and pushing personalized content based on AIGC, including the following steps:
[0058] S1: Obtain the data information of the user's historical viewing records, browsing behaviors, and interaction records;
[0059] S2: According to the data information obtained in step S1, analyze the data through a large language model to generate user portrait tags and personalized prompt words;
[0060] S3: Based on the user portrait tags and personalized prompt words generated in step S2, select and generate corresponding creation prompts;
[0061] S4: Based on the creation prompts generated in step S3, complete the automatic generation of multi-modal content through multiple generation models;
[0062] S5: Push the automatically generated multi-modal content in a personalized manner and collect feedback.
[0063] In this embodiment, assume that user A is a science fiction enthusiast, and the historical behavior data includes watching movies such as "Interstellar" and "Inception", frequently clicking on science fiction videos, and the search keywords are "black hole theory" and "time travel". The system needs to generate and push personalized content for him.
[0064] According to step 1, obtain the historical behavior data of user A:
[0065] Data source: The browsing records, search records, and interaction behaviors (clicking, favoriting, commenting) of user A on the video platform;
[0066] Example data:
[0067] Viewing records: 10 science fiction movies (accounting for 80%) and 2 documentaries (20%);
[0068] Search keywords: "quantum physics" and "interstellar travel";
[0069] Interaction behaviors: Liking science fiction videos 15 times, and the comment keywords are "hardcore science fiction" and "visual effects";
[0070] Based on the above data information, perform preprocessing on it:
[0071] Anonymization: Desensitize the user ID and aggregate the behavioral data into feature vectors;
[0072] Time window: Extract the behavioral sequences in the last 30 days and sort them by timestamp.
[0073] Step S2 is specifically as follows:
[0074] (1) Model configuration:
[0075] Large language model: Qwen2.5-72B-Instruct;
[0076] Attention mechanism: 4-head attention, hidden layer dimension 512, sequence length 256;
[0077] (2) User profile generation:
[0078] Input: The behavioral sequence encoding of user A (dimension 512);
[0079] The output labels include:
[0080] Interest labels: "Science fiction enthusiast (confidence 0.92)", "Hardcore science (0.85)";
[0081] Emotion labels: "High immersion preference (0.88)", "Complex narrative tendency (0.78)";
[0082] Attention weights: The weight of long-term interest (science fiction) is 0.75, and the weight of short-term interest (documentary) is 0.25;
[0083] (3) Multimodal prompt generation:
[0084] Reinforcement learning strategy: PPO algorithm, reward function weights (Relevance), (Diversity);
[0085] The output prompts include:
[0086] Text-to-image: "Future city skyline and hovering spaceships, cyberpunk style";
[0087] Text-to-audio: "Tense and suspenseful electronic synthesized music, BPM = 120, timbre: futuristic";
[0088] Image-to-video: "Nebula rotating dynamic special effects, resolution 4K, frame rate 60fps".
[0089] Step S3 is specifically as follows:
[0090] (1) Text-to-Image Model (Stable Diffusion):
[0091] Input prompt: "Future city skyline with hovering airships";
[0092] Parameters: Sampling temperature 0.8, generation resolution 1024×1024, number of iteration steps 50;
[0093] Output: 4 candidate images, screened by CLIP (text-image similarity ≥ 0.85);
[0094] (2) Text-to-Audio Model (Improved Jukedeck):
[0095] Input prompt: "Tense and suspenseful electronic synthesized music";
[0096] Parameters: Audio track length 30 seconds, sound library selected as "science fiction synthesizer", reverb intensity 0.6;
[0097] Output: 3 candidate audio segments, user portrait matching degree ≥ 90%;
[0098] (3) Image-to-Video Model (Stable Video Diffusion):
[0099] Input image: Selected text-to-image result;
[0100] Parameters: Video length 10 seconds, dynamic blur intensity 0.5, key frame interval 0.2 seconds;
[0101] Output: Dynamic nebula rotation video, PSNR = 29dB (meeting the requirement of cycle consistency).
[0102] Step S4 is specifically as follows:
[0103] (1) Push strategy:
[0104] Contextual bandit algorithm: Action space = {video, music, text-image}, exploration rate ;
[0105] Real-time feature: Current context of User A (tag "science fiction enthusiast", online period "evening");
[0106] Push content: Preferred video (weight 0.6) > music (0.3) > text-image (0.1);
[0107] (2) Push example:
[0108] Video: "Future city special effects short film";
[0109] Music: "Cyberpunk electronic remix";
[0110] Text and image: "Popular science article on black hole theory";
[0111] (3) Feedback collection:
[0112] The user clicks on the video (CTR = 1), and the viewing duration is 8 minutes (exceeding the threshold of 5 minutes);
[0113] Does not click on the text and image, triggering a negative feedback signal;
[0114] (4) Model update:
[0115] Online learning: Update the weights of the user profile (confidence of science fiction label +0.05);
[0116] Loss function: Based on the click-through rate (actual 1 vs predicted 0.9), adjust the push strategy parameters (learning rate 0.01).
[0117] Step S5 is specifically:
[0118] (1) Cross-modal consistency verification:
[0119] Video and music alignment: CLIP similarity 0.82 (threshold 0.8);
[0120] Cyclic consistency loss: PSNR = 28.5dB (meeting the standard);
[0121] (2) Dynamic update of the user profile:
[0122] New label: "Special effects visual preference (0.78)";
[0123] Eliminated label: "Documentary interest (confidence drops to 0.15)";
[0124] (3) System performance metrics:
[0125] Generation efficiency: The time taken for single content generation is 25 minutes (4 hours for traditional methods);
[0126] User satisfaction: NPS score +45 (benchmark value +20).
[0127] As Figure 2 shown, this embodiment also provides a push system based on the above-mentioned personalized content automatic generation and push method based on AIGC, including:
[0128] A data collection module, used for: obtaining data information on the user's historical viewing records, browsing behaviors, and interaction records;
[0129] A data processing module, used for: analyzing the obtained data information through a large language model, generating user profile labels and personalized prompt words, and selecting and generating corresponding creation prompts;
[0130] An automatic generation module, configured to: based on the generated creation prompts, complete the automatic generation of multimodal content through multiple generation models;
[0131] A personalized push and feedback module, configured to: perform personalized push on the automatically generated multimodal content and collect feedback.
[0132] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. The method for automatically generating and pushing personalized content based on AIGC is characterized by: The following steps are involved: S1: Obtain data information about the user's historical viewing records, browsing behaviors, and interaction records; S2: Analyze the data using a large language model based on the data information obtained in step S1 to generate user portrait labels and personalized prompt words; The generating of user portrait labels is specifically as follows: The Transformer-based multi-head attention mechanism is used to model user behavior sequences: ; in, is the query matrix, is the key-value matrix, is the vector dimension; By calculating the attention weights of different time steps in the behavior sequence, the long-term dependencies of user interests are dynamically extracted, and the final user portrait is expressed as: ; in, It is a three-layer fully connected network with GELU as the activation function, and outputs the probability distribution of user labels; The generating of personalized prompt words is specifically as follows: Use the mapping function from label to prompt word and combine it with reinforcement learning to optimize the generation quality: ; in, is the semantic similarity between the prompt word and the user tag, Score the diversity of the prompt words, , is the weight hyperparameter; S3: Based on the user portrait tag and personalized prompt words generated in step S2, select and generate corresponding creation prompts; S4: Based on the creation prompt generated in step S3, the multimodal content is automatically generated through multiple generation models; S5: Personalize the push of automatically generated multimodal content and collect feedback.
2. The method for automatically generating and pushing personalized content based on AIGC according to claim 1, characterized in that: The data information of obtaining the user's historical viewing records includes but is not limited to: TV series, movies, videos, music, articles, and pictures watched by users; The data information of the user's browsing behavior obtained includes but is not limited to: User's search history and browsing history; The data information of obtaining the user's interaction records includes but is not limited to: User click, favorite, like, and comment behavior data.
3. The method for automatically generating and pushing personalized content based on AIGC according to claim 1, characterized in that: In step S3, corresponding creation prompts are selected and generated, specifically: prompts of text-to-image, text-to-sound, and image-to-video are generated according to the media type, and the cross-modal joint loss function is used by jointly training the text-to-image, text-to-sound, and image-to-video models: ; in, The image-text alignment loss calculated for the CLIP model, is the cycle consistency loss, , is the balance coefficient.
4. The method for automatically generating and pushing personalized content based on AIGC according to claim 1, characterized in that: The personalized push is specifically as follows: Execute push strategy using contextual bandit algorithm with feedback: ; in, The type of content to be pushed. is the user context feature; The weight parameter for the type of content being pushed. is the noise term, and the parameters are updated through online learning , maximize user click-through rate and viewing time.
5. The method for automatically generating and pushing personalized content based on AIGC according to claim 4 is characterized in that: The feedback collection is specifically: Collect user interaction feedback and behavior data in real time to further optimize user portraits and recommendation algorithms. The feedback-driven model incremental update mechanism is: The loss function of the feedback-driven model based on user feedback data is: ; in, For users to The actual feedback is: click is recorded as (1), and no click is recorded as (0); The click probability predicted by the model.
6. A push system based on the AIGC-based personalized content automatic generation and push method as claimed in claim 1, characterized in that: include: Data collection module, used to obtain data information about users' historical viewing records, browsing behaviors and interaction records; The data processing module is used to: analyze the data through a large language model based on the acquired data information, generate user portrait tags and personalized prompt words, and select and generate corresponding creation prompts; The automatic generation module is used to: automatically generate multimodal content through multiple generation models based on the generated creation prompts; The personalized push and feedback module is used to: personalize the push of automatically generated multimodal content and collect feedback.
Citation Information
Patent Citations
Advertisement scheme automatic generation system based on AIGC
CN117350783A
Content recommendation method and system based on semantic recognition
CN119089398A