A digital media content recommendation system based on AIGC
Through the large language model and text embedding model based on AIGC, the user behavior events are interpreted in real time to generate intent vectors, which solves the problem of difficulty in capturing user dynamic intentions in the existing technology, realizes more accurate content recommendations, and improves the real-time and personalization of the user experience and recommendation system.
Patent Information
- Application Number
- CN202510678240.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing digital media content recommendation system is difficult to capture the user's dynamic intentions in real time, especially in conversation scenarios, which leads to deviations from user needs and is unable to effectively respond to short-term interest changes and cross-domain needs.
A large language model based on AIGC is used to interpret user behavior events in real time, generate dynamic intent descriptions, and convert them into intent vectors through text embedding models, and use vector index to perform approximate nearest neighbor searches to filter out candidate content that is highly consistent with user intent.
It significantly optimizes the content recommendation effect in conversation scenarios, improves the personalization of user experience and recommendations, and can more accurately respond to changes in users' immediate interests and needs.
Smart Images

Figure CN120196815B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent recommendation, and more specifically, to a digital media content recommendation system based on AIGC. Background Art
[0002] In the context of the increasing richness and diversification of current digital media content, the amount of information faced by users shows an explosive growth. How to help users efficiently and accurately obtain the information they are interested in among the vast amount of content has become the key to improving user experience and platform competitiveness. Therefore, building an intelligent and personalized digital media content recommendation solution can not only effectively alleviate the problem of information overload, but also enhance user stickiness, improve content distribution efficiency, and achieve a win-win situation for both the platform and users.
[0003] Existing digital media content recommendation systems mainly rely on traditional methods such as collaborative filtering, content-based analysis, and hybrid recommendation. These methods usually make recommendations by analyzing user historical behavior data or content features, but generally have certain limitations. For example, collaborative filtering is easily affected by cold start and data sparsity, and has a weak adaptability to new users or new content; content-based methods are difficult to deeply understand complex and changing user interests and can only capture surface preferences. In addition, these traditional solutions often cannot grasp the dynamic intentions of users in the current session in real time, are not sensitive enough to short-term interest changes and cross-domain requirements, resulting in a deviation between the recommended results and the actual needs.
[0004] Therefore, an optimized digital media content recommendation system based on AIGC is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a digital media content recommendation system based on AIGC. First, it uses a large language model (LLM) fine-tuned by instructions to real-time interpret a series of behavior events shown by users in the current session, so as to intelligently generate text that can accurately describe the current dynamic intentions of users, which enables the system to keenly capture the subtle and immediate interest focuses of users; then, it uses a text embedding model to convert the dynamic intention description into a vectorized representation to obtain an intention vector, and further uses the intention vector to perform an efficient approximate nearest neighbor search in a pre-constructed content library vector index, so as to quickly locate and screen out candidate content that highly matches the current intentions of users. Finally, several of the most relevant items are selected from the candidate list to obtain the final recommended list. In this way, the system can significantly optimize the content recommendation effect in the session scenario and improve the user experience by deeply utilizing the LLM's understanding of user dynamic intentions and the generation ability of AIGC.
[0006] According to one aspect of the present application, there is provided a digital media content recommendation system based on AIGC, which includes:
[0007] A behavior event collection module for collecting a series of behavior events of a target user in the current session;
[0008] A dynamic intention description module for inputting a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description;
[0009] A text embedding module for using a text embedding model to convert the dynamic intention description into an intention vector;
[0010] An intention approximate search module for performing approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list;
[0011] A content recommendation module for selecting the top N candidate contents from the candidate content list as the final content recommendation list.
[0012] Compared with the prior art, a digital media content recommendation system based on AIGC provided by the present application first uses a large language model (LLM) fine-tuned by instructions to interpret in real time a series of behavior events demonstrated by a user in the current session, so as to intelligently generate text that can accurately describe the user's current dynamic intention, which enables the system to keenly capture the user's subtle and immediate interest focus; then, a text embedding model is used to convert the dynamic intention description into a vectorized representation to obtain an intention vector, and further the intention vector is used to perform efficient approximate nearest neighbor search in the pre-constructed content library vector index, so as to quickly locate and screen out candidate contents that highly match the user's current intention. Finally, several most relevant items are selected from the candidate list to obtain the final recommendation list. In this way, the system can significantly optimize the content recommendation effect in the session scenario and improve the user experience by deeply using the LLM's understanding of the user's dynamic intention and the generation ability of AIGC. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. They are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0014] Figure 1 It is a block diagram of a digital media content recommendation system based on AIGC according to an embodiment of the present application;
[0015] Figure 2Schematic diagram of data flow of an AIGC-based digital media content recommendation system according to an embodiment of the present application;
[0016] Figure 3 Block diagram of a dynamic intent description module in an AIGC-based digital media content recommendation system according to an embodiment of the present application. Detailed implementation manners
[0017] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0018] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0019] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0020] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0021] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0022] In the technical solution of the present application, an AIGC-based digital media content recommendation system is proposed. Figure 1 Block diagram of an AIGC-based digital media content recommendation system according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of an AIGC-based digital media content recommendation system according to an embodiment of the present application. As Figure 1 and Figure 2As shown, the AIGC-based digital media content recommendation system 300 according to an embodiment of the present application includes: a behavior event collection module 310, configured to collect a series of behavior events of a target user in the current session; a dynamic intention description module 320, configured to input the series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description; a text embedding module 330, configured to use a text embedding model to convert the dynamic intention description into an intention vector; an intention approximate search module 340, configured to perform approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list; and a content recommendation module 350, configured to select the top N candidate contents from the candidate content list as the final content recommendation list.
[0023] Specifically, the behavior event collection module 310 is configured to collect a series of behavior events of a target user in the current session. It should be understood that in the current digital media environment, the interests and needs of users are often in a state of continuous change. Especially in a specific interactive session, these changes are more significant and dynamic. Therefore, in order to capture the dynamic intentions in the user's current session, enable the recommendation system to get rid of the dependence on historical static data of traditional models, and perceive the changing needs of users in real time to achieve accurate and real-time content recommendations, in the technical solution of the present application, a series of behavior events of the target user in the current session are collected. Among them, the behavior events include multi-dimensional interaction data such as the user's clicks, browsing duration, search keywords, scrolling behavior, likes, shares, comments, etc. These data comprehensively reflect the user's current focus of attention and interest trends. Through the collection and utilization of this real-time behavior data, the recommendation system can not only significantly improve the accuracy of identifying user interests, but also enhance the response ability to short-term hot content and cross-domain interest migration, thereby bringing a more personalized, timely and user-real-needs-compliant recommendation effect, greatly improving the user experience and the content distribution efficiency of the platform.
[0024] Specifically, the dynamic intention description module 320 is configured to input a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description. In a specific example of the present application, as Figure 3 shown, the dynamic intention description module 320 includes: a semantic embedding encoding unit 321, configured to perform semantic embedding encoding on each behavior event in the series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors; a time series context encoding unit 322, configured to perform behavior event time series context encoding on the time series distribution of behavior event semantic embedding encoding vectors to obtain a user behavior pattern time series implicit encoding vector; and a dynamic intention description generation unit 323, configured to obtain a dynamic intention description based on the user behavior pattern time series implicit encoding vector.
[0025] Specifically, the semantic embedding encoding unit 321 is configured to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors. It should be understood that a user may generate a large number of diverse behavior events in a single session, such as clicks, browsing, searches, likes, comments, etc. These behaviors not only reflect the user's immediate interest in the content but also contain their dynamically changing needs and preferences. Traditional recommendation systems often only perform simple statistics or tagging on these behaviors, making it difficult to capture their deep semantic information and temporal evolution relationships, resulting in a deviation between the recommended results and the user's true intentions. Therefore, in the technical solution of this application, a behavior event embedding matrix is used to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors. Through semantic embedding encoding, not only can the information carried by a single event be captured, but also the potential temporal dependencies and evolution laws between a series of events can be revealed.
[0026] Specifically, the temporal context encoding unit 322 is configured to perform behavioral event temporal context encoding on the time series distribution of the behavioral event semantic embedding encoding vectors to obtain the user behavior pattern temporal implicit encoding vectors. It should be understood that the interaction behavior between the user and the digital media platform is not an isolated event. Actions such as clicks, browsing, and pauses form a behavioral fluid with internal logic in the time dimension. The time series distribution of the behavioral event semantic embedding encoding vectors often contains dynamic interest trajectories and potential intentions. However, traditional recommendation systems are difficult to effectively capture the complex correlations of such temporal contexts. Therefore, to break through the limitations of static feature extraction and construct an encoding representation that can adaptively fuse temporal dependencies and dynamic intentions, in the technical solution of this application, behavioral event temporal context encoding is performed on the time series distribution of the behavioral event semantic embedding encoding vectors to obtain the user behavior pattern temporal implicit encoding vectors. In this process, each behavioral event is regarded as a dynamic entity, and the transmission potential energy and repulsive force coefficient between it and the current node are calculated. Essentially, a non-uniform and adaptive information propagation channel is established in the time flow. Among them, the transmission potential energy encodes the potential driving force of historical behaviors on the current intention (such as the interest reinforcement formed by repeatedly browsing similar content), while the repulsive force simulates the information attenuation caused by time distance or behavior type differences (such as the weak influence of a viewing record three days ago on the current recommendation). This mechanism enables the model to distinguish the "mainstream trend" from the "instantaneous fluctuations" in the user behavior stream. For example, it can identify the intention turn of the user who has frequently clicked on food videos but recently started searching for fitness tutorials, thus avoiding misjudging short-term behavioral noise as the core interest. In this way, the system can transform scattered behavioral events into a coherent "intention field", which not only retains the original temporal structure of the behavioral sequence but also strengthens the implicit correlations across time points through the product state synergy effect. The generated user behavior pattern temporal implicit encoding vectors provide a semantic base with high information density for the subsequent large language model to generate accurate intention descriptions, enabling the recommendation system to break through the static preference assumption of traditional collaborative filtering and respond in real time to the dynamic needs of users in specific scenarios, thereby improving the fit between the recommended content and the immediate scenario.
[0027] Specifically, first, the current node feature vector of the behavior event is extracted from the time series distribution of the behavior event semantic embedding encoding vector, and the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector are defined as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors. It should be understood that the real-time interaction between the user and the platform has significant spatio-temporal correlation. The influence of each behavior event is not evenly distributed, but shows non-linear attenuation or transition with its spatio-temporal distance and semantic correlation strength from the current moment. For example, when the user is watching a short video, the previous like and favorite behaviors may form an "information potential energy" for the continuation of interest, while the rapid swiping or exiting operation in the middle may generate a "repulsion effect" that inhibits subsequent similar recommendations. The traditional processing methods of static time windows or fixed attenuation weights are difficult to capture such complex dynamics. Therefore, in the technical solution of this application, the current node feature vector of the behavior event is extracted from the time series distribution of the behavior event semantic embedding encoding vector, and the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector are defined as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors.
[0028] By taking the latest behavior node of the current session as the information aggregation anchor point and defining the historical behavior nodes as the to-be-propagated nodes, a dynamic perception field centered on the current intention is essentially constructed in the time flow, enabling the system to break through the uniform processing of time series data by traditional recommendation models. When the user browses digital media content, their behavior sequence naturally has the characteristic of attention focus. In a specific example, the behavior closer to the current moment often carries a higher-weight intention signal, but some early key behaviors (such as the first click on content in a certain vertical field) may continuously affect subsequent decisions due to the awakening of potential interest. By dividing the time series into the current node and the to-be-propagated nodes, the current node represents the explicit interaction focus of the user at this moment, and the to-be-propagated nodes form a potential factor pool of influence. This structure provides a precise object of action for the calculation of potential energy and repulsion force in the subsequent liquid message propagation.
[0029] In a specific example of this application, the current node feature vector of the behavior event is extracted from the time series distribution of the behavior event semantic embedding encoding vector according to the following formula, and the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector are defined as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors; where the formula is:
[0030] ,
[0031] ,
[0032] ,
[0033] Among them, is the time series distribution of the semantic embedding coding vectors of the behavior event, are the semantic embedding coding vectors of the behavior event at the 1st, 2nd, and Tth time steps in the time series distribution of the semantic embedding coding vectors of the behavior event respectively, T represents the total length of the sequence, and d represents the feature dimension of each time step. is the current node feature vector of the behavior event, is the time series of the feature vectors of the nodes to be propagated of the behavior event.
[0034] Next, calculate the dynamic transmission potential energy of each behavior event node to be propagated feature vector in the time series of the behavior event node to be propagated feature vector relative to the current node feature vector of the behavior event to obtain the time series of the dynamic transmission potential energy coding vectors of the behavior event nodes. It should be understood that due to the strong temporal dependence and context sensitivity of the user's interaction behavior in the current session, simple time series modeling methods (such as RNN or LSTM) cannot fully capture the non-linear and non-steady dynamic coupling relationships between behavior events. By introducing the physically inspired dynamic transmission potential energy calculation, the system can construct an interaction field of behavior event nodes in the high-dimensional latent space, thereby mapping the discrete behavior event nodes in the time series data into continuous vector representations with energy attributes. Therefore, to construct a deep time series representation mechanism with physical field theory characteristics, by establishing a dynamic transmission potential energy mapping between behavior nodes, the discrete user behavior is transformed into the energy gradient distribution in the continuous semantic manifold. In the technical solution of this application, calculate the dynamic transmission potential energy of each behavior event node to be propagated feature vector in the time series of the behavior event node to be propagated feature vector relative to the current node feature vector of the behavior event to obtain the time series of the dynamic transmission potential energy coding vectors of the behavior event nodes.
[0035] Among them, the dynamic transmission potential encoding vector of the behavior event node, through the parameterized learning mechanism, transforms the non-linear interaction pattern between the feature vector of the node to be propagated and the feature vector of the current node into the potential energy gradient distribution in the high-dimensional space. It not only bears the temporal correlation strength between behavior events, but more importantly, reveals the implicit semantic correlation across time steps, such as complex patterns of gradual change, jump or superposition of user interests, etc., thus providing a physically interpretable kinetic basis for the subsequent simulated dynamic liquid message propagation. Through the dynamic transmission potential encoding vector of the behavior event node, the system can effectively distinguish the surface behavior pattern and the deep intention evolution. This potential field-based modeling method enables the recommendation system to break through the static preference assumption of traditional collaborative filtering, realize the dynamic modeling and accurate response to the user's real-time intention, and finally improve the comprehensive performance of the recommendation results in terms of content novelty, temporal coherence and interest matching degree.
[0036] In a specific example of this application, the following formula is used to calculate the dynamic transmission potential of each behavior event node feature vector in the time series of the behavior event node feature vector to be propagated relative to the behavior event current node feature vector to obtain the time series of the dynamic transmission potential encoding vector of the behavior event node; where, the formula is:
[0037] ,
[0038] where, and respectively represent the learnable query weight matrix and key weight matrix, is the dynamic transmission potential factor, represents the calculation of and the square difference between, is the dynamic transmission potential encoding vector of the behavior event node.
[0039] Subsequently, calculate the behavior event message propagation repulsion coefficient of each behavior event node feature vector to be propagated in the time series of behavior event node feature vectors to be propagated relative to the current behavior event node feature vector, so as to obtain the time series of behavior event message propagation repulsion coefficients. It should be understood that the information propagation between feature vectors at different time nodes in the user behavior sequence is not unconditionally unobstructed, and a dynamic modulation mechanism needs to be established to suppress the interference of irrelevant or noise signals. Although traditional time series modeling methods (such as attention mechanisms or gated networks) can achieve partial information filtering functions, their modeling of the hindrance effect of cross-time step information conduction is mostly based on empirical parameters and lacks interpretable physical constraints. Therefore, in order to quantify the negative impact weight of historical behavior events on the current intention modeling process and establish a dynamic suppression criterion for the liquid-like message propagation mechanism, in the technical solution of this application, calculate the behavior event message propagation repulsion coefficient of each behavior event node feature vector to be propagated in the time series of behavior event node feature vectors to be propagated relative to the current behavior event node feature vector, so as to obtain the time series of behavior event message propagation repulsion coefficients. By introducing the repulsion coefficient, the system can construct a dynamic impedance model of the information propagation channel based on physical field theory, thus being more in line with the path dependence and context conflict existing in the evolution of user interests in the real scenario.
[0040] Specifically, the message propagation repulsion coefficient generates a scalar conduction hindrance factor by calculating the state difference and time series distance between the node feature vector to be propagated and the current node feature vector. This repulsion coefficient not only reflects the semantic heterogeneity between behavior events (such as the interest conflict when the user switches from "science and technology news reading" to "entertainment video watching"), but also captures the time decay effect (such as the weak correlation of early behavior events with the current intention), thus providing a dynamic modulation basis based on field theory for the subsequent liquid-like message propagation process. Through the generation of the time series of the repulsion coefficient, the system can dynamically suppress the information penetration intensity of node feature vectors with high state differences (such as cross-domain content interaction) or long time spans (such as historical behaviors outside ultra-short sessions) from the current user intention. This repulsion mechanism based on physical field theory enables the recommendation system to more accurately focus on the coherent and intensive interaction patterns in the current user session, effectively filter the interference of irrelevant time series signals, and then enhance the semantic consistency of the dynamic intention description and the scenario adaptability of the recommendation results.
[0041] In a specific example of this application, calculate the behavior event message propagation repulsion coefficient of each behavior event node feature vector to be propagated in the time series of behavior event node feature vectors to be propagated relative to the current behavior event node feature vector according to the following formula to obtain the time series of behavior event message propagation repulsion coefficients, where the formula is:
[0042] ,
[0043] Among them, represents the vector multiplication, is the timestamp extraction function, represents the time difference scaling factor, represents the square of the norm, is the repulsive force coefficient for the propagation of the behavior event message.
[0044] Then, based on the time series of the repulsive force coefficient for the propagation of the behavior event message and the time series of the dynamic transmission potential energy encoding vectors of the behavior event nodes, perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of dynamically encoded vectors of the behavior event node messages. It should be understood that the information transmission between user behavior nodes has non-linear coupling characteristics, and a multi-body interaction model inspired by physical field theory is required to achieve dynamic semantic fusion across time nodes. Traditional recommendation systems use fixed time decay functions or shallow attention mechanisms, which are difficult to capture the quantization semantic transitions (such as the sudden migration of interest topics) and multi-scale time dependencies (such as the synergistic effect of high-frequency click behaviors and low-frequency in-depth browsing) hidden in the behavior sequence. Therefore, in order to establish a dynamic message generation mechanism controlled by dual physical field parameters and achieve refined modulation of the influence of historical behavior events on the current user intention, in the technical solution of this application, based on the time series of the repulsive force coefficient for the propagation of the behavior event message and the time series of the dynamic transmission potential energy encoding vectors of the behavior event nodes, perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of dynamically encoded vectors of the behavior event node messages.
[0045] That is, by non-linearly coupling the dynamic transmission potential energy encoding vectors of the behavior event nodes (representing the driving strength of information transmission between behavior events) with the repulsive force coefficient for the propagation of the behavior event message (representing the dynamic impedance of the information channel), a dynamically encoded vector of the behavior event node message with directionality and attenuation characteristics is generated. This coupling mechanism not only simulates the interactive dynamic characteristics of multi-body particles in a liquid medium (such as the balance between viscous resistance and potential energy gradient), but more importantly, transforms the cross-time step correlations (such as the continuity of interest migration and the suddenness of exploration behaviors) hidden in the user behavior sequence into a differentiable high-dimensional vector space mapping relationship, providing a physically interpretable intermediate representation for subsequent multi-body propagation collaborative renormalization.
[0046] In a specific example of this application, the following formula is used to perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of dynamically encoded vectors of the behavior event node messages; where the formula is:
[0047] ,
[0048] Among them, represents dot product by position, represents function, represents function, represents the learnable value weight matrix, is the dynamic encoding vector of the behavior event node message.
[0049] Furthermore, multi-body propagation collaborative renormalization is performed on the set of dynamic encoding vectors of behavior event node messages to obtain an optimized set of dynamic encoding vectors of behavior event node messages. It should be understood that after the dynamic message encoding vectors in the user behavior sequence are propagated in an imitation liquid state, their internal correlation presents high-dimensional manifold characteristics, including both the spatio-temporal correlation of local propagation lattice points and the implicit coupling effect across time steps. Traditional time series aggregation methods (such as simple vector superposition or average pooling) cannot effectively handle the multi-body collaborative effect generated in the non-linear propagation process. Therefore, in a preferred example of this application, multi-body propagation collaborative renormalization is performed on the set of dynamic encoding vectors of behavior event node messages to obtain an optimized set of dynamic encoding vectors of behavior event node messages.
[0050] Specifically, by introducing a global scale correction factor μ (a renormalization operator based on the mean of eigenvalues) and an auxiliary propagation key channel, conformal mapping and dynamic equilibrium of cross-time node information are achieved. In a real-time recommendation scenario, when the user behavior stream contains a mixed mode of high-frequency interactions (such as continuously clicking on similar videos) and low-frequency dwells (such as concentrating on watching tutorials), the renormalization process projects the contribution weights of each node to the homogeneous manifold space through the non-orthogonal basis projection of the product state, ensuring that short-term behavior noise does not overwhelm long-term interest features while retaining the collaborative effect of cross-modal semantics. In this way, the finally generated temporal implicit encoding vector of the user behavior pattern not only retains the mutation characteristics of short-term interests but also maintains the continuity of long-term behavior inertia, providing high-quality feature inputs with physical significance for the downstream large language model to generate accurate dynamic intention descriptions.
[0051] In the above process, for the set of dynamic encoding vectors of the behavior event node messages, due to the adaptability of the effect propagation under non-linear representation, each dynamic encoding vector of the behavior event node message essentially constitutes a vector local propagation lattice point that can be precisely decoupled. Therefore, when aggregating each local propagation lattice point, in addition to simple superposition, it is also desirable to introduce local propagation lattice point auxiliary key corrections that conform to the propagation characteristics.
[0052] That is, the vector product state is introduced to simulate the auxiliary propagation bond channels of local propagation lattice points, so as to reflect the propagation behavior in many-body effects with a product structure, enabling the synergy effect to emerge naturally:
[0053] ,
[0054] Moreover, in order to perform the collaborative renormalization under the vector product state that can be exactly decoupled, a global scaling representation correction is introduced:
[0055] ,
[0056] where is the mean value of all the eigenvalues of the set of the dynamic encoding vectors of the behavior event node messages.
[0057] Furthermore, considering the scaling transformation of the global correction under the product state, it is necessary to calculate the partial derivative of the global scaling representation correction with respect to , that is:
[0058] ,
[0059] From this, we obtain:
[0060] ,
[0061] where is the hyperparameter for linear compression .
[0062] Then, the dynamic encoding vectors of the behavior event node messages are optimized:
[0063] ,
[0064] In this way, while ensuring the covariance under the scaling transformation is maintained, the renormalization collaboration under the vector product state that can be exactly decoupled is carried out. Thus, by introducing the auxiliary propagation bond channels of local propagation lattice points, the representation of many-body effects is strengthened, and the aggregation effect of the dynamic timing information of the dynamic encoding vectors of the behavior event node messages for each time point on the dynamic encoding vectors of the behavior event node messages is improved.
[0065] Subsequently, the set of dynamically encoded vectors of the fused and optimized behavior event node messages and the current node feature vector of the behavior event are combined to obtain the temporal implicit encoding vector of the user behavior pattern. It should be understood that the real-time interaction behaviors generated by the user in the current session not only need to immediately capture the surface operation features, but also need to reveal the potential intention evolution path through the non-linear propagation of the historical behavior sequence. The set of dynamically encoded vectors processed by multi-body propagation collaborative renormalization carries the propagation collaborative effect and information screening mechanism across time steps, while the current node feature vector represents the most up-to-date state of the user interaction. Therefore, in order to break through the dimensionality limitation of static feature splicing and achieve the deep coupling of the dynamic propagation field and the real-time state field, in the technical solution of this application, the set of dynamically encoded vectors of the fused and optimized behavior event node messages and the current node feature vector of the behavior event are combined to obtain the temporal implicit encoding vector of the user behavior pattern. That is, by introducing a spatial fusion mechanism based on the product state, the system performs manifold alignment and information entanglement on the renormalized dynamically encoded vectors (containing the non-linear correlations and multi-body collaborative effects between historical behaviors) and the current node feature vectors (reflecting the fine-grained semantics of the latest interaction events), generating an implicit encoding with spatio-temporal invariance. This fusion process is not a simple vector addition, but a differential geometric operation that preserves scale transformation covariance. While retaining the independent propagation characteristics of each information flow, it constructs a cross-modal semantic resonance channel. This dual information integration mechanism effectively solves the trade-off dilemma between real-time performance and continuity in traditional recommendation systems, enabling the dynamic intention description generation module to output a recommendation strategy that not only conforms to the long-term interest evolution trend but also accurately matches the current session focus based on high-fidelity spatio-temporal features, ultimately achieving personalized content matching across time scales.
[0066] In a specific example of this application, the following formula is used to fuse and optimize the set of dynamically encoded vectors of the behavior event node messages and the current node feature vector of the behavior event to obtain the temporal implicit encoding vector of the user behavior pattern, where the formula is:
[0067] ,
[0068] ,
[0069] Wherein, is the dynamically encoded vector of the optimized behavior event node message, represents function, represents the gating weight matrix, is the temporal implicit encoding vector of the user behavior pattern.
[0070] Specifically, the dynamic intention description generation unit 323 is configured to obtain a dynamic intention description based on the user behavior pattern time-series implicit encoding vector. It should be understood that in the current environment where digital media content is increasingly rich and user interests change rapidly, simply relying on traditional behavioral data analysis is difficult to accurately capture the immediate needs and complex intentions of users. The large language model can infer the user's current specific interest points, content preferences, or the type of content being sought in a more user-friendly and contextually relevant manner based on the embedded user behavior pattern time-series information, rather than just staying at the similarity comparison in the traditional vector space. This not only helps to reveal the diverse hidden needs of users but also effectively addresses the challenges of short-term interest fluctuations and cross-domain interest migrations. Therefore, in the technical solution of this application, after embedding the user behavior pattern time-series implicit encoding vector into a preset Prompt, it is input into the large language model fine-tuned by instructions to obtain a dynamic intention description. The preset Prompt is: Based on the semantic information of the following recent interaction behavior sequence between the user and digital media, infer the user's current specific interest points, content preferences, or the type of content being searched for, and generate a concise intention description text. That is, through the structured guidance of the preset Prompt (such as clearly requiring the inference of interest points, content preferences, and search types), the generation direction of the large language model is constrained, enabling the large language model to generate intention texts that meet the scenario requirements based on understanding the user interaction behavior sequence, enhancing the consistency and controllability of the model output results. Here, the dynamic intention description can not only accurately reflect the core needs of the user in the current conversation stage but also guide the recommendation system to be more focused on the user's true intentions during content retrieval, improving the relevance and personalization of recommendations, providing a structured and semantically rich data representation for the downstream text embedding and vector retrieval modules, thus significantly improving the candidate content matching accuracy and the final recommendation quality.
[0071] Particularly, the text embedding module 330 is configured to use a text embedding model to convert the dynamic intention description into an intention vector. That is, in the technical solution of this application, to construct a semantically continuous and structured vector space, a text embedding model based on the Word2vec model is used to perform embedding encoding on the dynamic intention description to obtain an intention vector. That is, through the text embedding based on the Word2vec model, the dynamic intention description is encoded to map the user's current interest points, content preferences, or needs into a low-dimensional and semantically rich space in the form of a continuous vector, so that the recommendation system can more accurately understand and quantify the semantic connotation of the user's intentions. By converting the dynamic intention into an intention vector, the system can quickly identify and match candidate content highly relevant to the user's current needs, improving the relevance and real-time response ability of the recommendation.
[0072] Specifically, the intention approximate search module 340 is used to perform approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a list of candidate contents. Among them, the dynamic intention description generated by the large language model based on AIGC is transformed into a semantically rich intention vector through text embedding. This vector form can highly condense the user's real-time interest information and complex intentions, laying a solid foundation for content matching. In the technical solution of this application, to improve the relevance and real-time performance of recommendations and enable the system to accurately capture and respond to the latest interest points and content preferences shown by the user in the current session, by performing approximate nearest neighbor search in the vector index of the content library using the intention vector, efficient and semantically accurate content retrieval can be achieved. It is worth mentioning that approximate nearest neighbor search, as an efficient vector retrieval technology, can quickly locate the candidate content closest to the user's intention vector in the large-scale content vector space. In the specific implementation process, each piece of content in the content library has been pre-converted into a corresponding vector representation and a corresponding vector index structure has been established, such as an index based on the ANN (Approximate Nearest Neighbor) algorithm, such as HNSW, Faiss, etc. In this process, the system inputs the user's current intention vector and performs approximate search through the vector index structure, quickly extracting several candidate contents that are semantically most similar to the user's dynamic intention. In this way, the probability of matching high-relevance candidate contents is greatly improved, and the coincidence degree between the recommendation result and the user's real needs is significantly improved.
[0073] In particular, the content recommendation module 350 is configured to select the top N candidate contents from the candidate content list as the final content recommendation list. It should be understood that the candidate content list often contains a large number of contents with varying degrees of relevance. Pushing all of them directly will not only cause information overload but also may affect the user experience and the utilization efficiency of platform resources. Therefore, it is necessary to further screen and rank the candidate contents to select the top N highest-quality and most-matched contents as the final recommendation results. Thus, in the technical solution of this application, the top N candidate contents are selected from the candidate content list as the final content recommendation list. It is worth mentioning that the system will first set a default value of N in the configuration and also support dynamic adjustment according to personalized strategies. For example, the recommended quantity is appropriately increased for highly active users or member users, while the display is streamlined for new users or low-active users. In addition, the relevance score threshold output by the algorithm model can also be combined to only retain the top N with the highest scores and qualified quality, so as to ensure that the content finally presented to the user is the highest-quality and most in line with their current intention. This method can effectively reduce the interference of redundant information, improve the utilization rate of platform resources, speed up the page loading speed, and at the same time improve the accuracy of recommendation hits and the click-through conversion effect. Finally, by selecting the top N from the candidate list as the final recommendation, it not only enhances the system's adaptability to short-term interest changes and complex demand scenarios but also greatly improves user satisfaction and platform stickiness, achieving the goals of precision, efficiency, and sustainable development in the field of digital media intelligent distribution.
[0074] As described above, the AIGC-based digital media content recommendation system 300 according to the embodiments of this application can be implemented in various wireless terminals, such as a server with an AIGC-based digital media content recommendation algorithm. In a possible implementation manner, the AIGC-based digital media content recommendation system 300 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the AIGC-based digital media content recommendation system 300 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the AIGC-based digital media content recommendation system 300 can also be one of the many hardware modules of the wireless terminal.
[0075] Alternatively, in another example, the AIGC-based digital media content recommendation system 300 and the wireless terminal can also be separate devices, and the AIGC-based digital media content recommendation system 300 can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0076] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A digital media content recommendation system based on AIGC, characterized in that, Including: A behavior event collection module for collecting a series of behavior events of a target user in the current session; A dynamic intention description module for inputting a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description; A text embedding module for using a text embedding model to convert the dynamic intention description into an intention vector; An intention approximate search module for performing approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list; A content recommendation module for selecting the top N candidate contents from the candidate content list as the final content recommendation list; Among them, the dynamic intention description module includes: A semantic embedding encoding unit for performing semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors; A temporal context encoding unit for performing behavior event temporal context encoding on the time series distribution of behavior event semantic embedding encoding vectors to obtain a user behavior pattern temporal hidden encoding vector; A dynamic intention description generation unit for obtaining a dynamic intention description based on the user behavior pattern temporal hidden encoding vector; Among them, the temporal context encoding unit includes: A node extraction sub-unit for extracting the current node feature vector of the behavior event from the time series distribution of the behavior event semantic embedding encoding vectors, and defining the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vectors as the behavior event to-be-propagated node feature vectors to obtain a time series of behavior event to-be-propagated node feature vectors; An imitation dynamic liquid message propagation sub-unit for performing behavior event imitation dynamic liquid message propagation on each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors based on the behavior event liquid interaction field between each behavior event to-be-propagated node feature vector and the current node feature vector of the behavior event in the time series of behavior event to-be-propagated node feature vectors to obtain a set of behavior event node message dynamic encoding vectors; A fusion sub-unit for fusing the set of behavior event node message dynamic encoding vectors and the current node feature vector of the behavior event to obtain a user behavior pattern temporal hidden encoding vector.
2. The digital media content recommendation system based on AIGC according to claim 1, wherein, The semantic embedding encoding unit is used for: Performing semantic embedding encoding on each behavior event in a series of behavior events using a behavior event embedding matrix to obtain a time series distribution of behavior event semantic embedding encoding vectors.
3. The AIGC-based digital media content recommendation system according to claim 2, wherein The imitation dynamic liquid message propagation sub-unit is used for: Calculating the behavior event node dynamic transmission potential energy of each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors relative to the current node feature vector of the behavior event to obtain a time series of behavior event node dynamic transmission potential energy encoding vectors; Calculating the behavior event message propagation repulsion coefficient of each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors relative to the current node feature vector of the behavior event to obtain a time series of behavior event message propagation repulsion coefficients; Based on the time series of the repulsive force coefficient of behavioral event message propagation and the time series of the dynamic transmission potential encoding vector of behavioral event nodes, perform pseudo-dynamic liquid message propagation on each behavioral event node feature vector in the time series of the feature vectors of behavioral event nodes to be propagated, so as to obtain a set of dynamic encoding vectors of behavioral event node messages.
4. The AIGC-based digital media content recommendation system according to claim 3, wherein, The fusion subunit is used for: Performing multi-body propagation collaborative renormalization on the set of dynamic encoding vectors of behavioral event node messages to obtain an optimized set of dynamic encoding vectors of behavioral event node messages; Fusing the optimized set of dynamic encoding vectors of behavioral event node messages and the current behavioral event node feature vector to obtain a user behavior pattern time-series implicit encoding vector.
5. The AIGC-based digital media content recommendation system according to claim 4, wherein The dynamic intention description generation unit is used for: After embedding the user behavior pattern time-series implicit encoding vector into a preset Prompt, input it into a large language model fine-tuned by instructions to obtain a dynamic intention description.
6. The AIGC-based digital media content recommendation system according to claim 5, wherein The preset Prompt is: "Based on the semantic information of the following sequence of the user's recent interactions with digital media, infer the user's current specific interest points, content preferences, or the type of content being searched for, and generate a concise intention description text.
7. The AIGC-based digital media content recommendation system according to claim 6, wherein The text embedding module is used for: Using a text embedding model based on the Word2vec model to perform embedding encoding on the dynamic intention description to obtain an intention vector.
Citation Information
Patent Citations
Teenager algorithm code auxiliary learning system and method based on large language model
CN117235347A
Multi-modal question and answer interpretation method and related device
CN118689997A
Cited By
AIGC-based interactive large-screen real-time drawing method and system
CN121213737A
An AIGC-based interactive large-screen real-time drawing method and system
CN121213737B