Digital media content recommendation system based on AIGC

Through the AIGC-based digital media content recommendation system, using large language models and text embedding models, the user behavior events are interpreted in real time and dynamic intention descriptions are generated, which solves the limitations of the existing system in capturing user dynamic intentions and responding to short-term interest changes, and achieves a more accurate and personalized content recommendation effect.

CN120196815AActive Publication Date: 2025-06-24BAOJI VOCATIONAL TECH COLLEGE

Patent Information

Application Number
CN202510678240.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing digital media content recommendation system has limitations in capturing user dynamic intentions and responding to short-term interest changes, especially in the adaptability and real-time intention grasp of new users or new content.

Method used

A recommendation system based on AIGC is adopted to interpret user behavior events in real time through large language models, generate dynamic intent descriptions, and use text embedding models to convert intent descriptions into vector representations, perform approximate nearest neighbor searches in the content library, and filter out content that matches user intent.

Benefits of technology

It significantly optimizes the content recommendation effect in conversation scenarios, improves the personalization of user experience and recommendations, and can more accurately capture the user's subtle and instant interest focus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196815A_ABST
    Figure CN120196815A_ABST
Patent Text Reader

Abstract

The invention discloses an AIGC-based digital media content recommendation system, which comprises the following steps of: firstly, interpreting a series of behavior events displayed by a user in a current session in real time by using a large language model (LLM) subjected to instruction fine tuning so as to intelligently generate a text capable of accurately describing the current dynamic intention of the user; this enables the system to sensitively capture subtle and instant interest focuses of the user; then, converting the dynamic intention description into a vectorized representation by adopting a text embedding model to obtain an intention vector, further performing efficient approximate nearest neighbor search in a pre-constructed content library vector index by utilizing the intention vector, so as to quickly position and screen out candidate content highly matched with the current intention of the user, and finally, extracting the current intention of the user from the candidate content. And selecting a plurality of most relevant items from the candidate list to obtain a final recommendation list. Therefore, the system can significantly optimize the content recommendation effect in the session scene by deeply utilizing the understanding of the LLM on the dynamic intention of the user and the generation capability of the AIGC, and improves the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent recommendation, and more specifically, to a digital media content recommendation system based on AIGC. Background Art

[0002] In the context of the increasing richness and diversification of current digital media content, the amount of information faced by users shows an explosive growth. How to help users efficiently and accurately obtain the information they are interested in from the vast amount of content has become the key to improving user experience and platform competitiveness. Therefore, constructing an intelligent and personalized digital media content recommendation solution can not only effectively alleviate the problem of information overload, but also enhance user stickiness, improve content distribution efficiency, and achieve a win-win situation for both the platform and users.

[0003] Existing digital media content recommendation systems mainly rely on traditional methods such as collaborative filtering, content-based analysis, and hybrid recommendation. These methods usually make recommendations by analyzing user historical behavior data or content features, but generally have certain limitations. For example, collaborative filtering is easily affected by cold start and data sparsity, and has a weak adaptability to new users or new content; content-based methods are difficult to deeply understand complex and changing user interests and can only capture surface preferences. In addition, these traditional solutions often cannot grasp the dynamic intentions of users in the current session in real time, are not sensitive enough to short-term interest changes and cross-domain requirements, resulting in a deviation between the recommended results and the actual needs.

[0004] Therefore, an optimized digital media content recommendation system based on AIGC is expected. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a digital media content recommendation system based on AIGC. First, by using a large language model (LLM) fine-tuned with instructions to interpret a series of behavior events shown by users in the current session in real time, it can intelligently generate text that can accurately describe the current dynamic intentions of users, enabling the system to keenly capture the subtle and immediate interest focus of users; then, a text embedding model is used to convert the dynamic intention description into a vectorized representation to obtain an intention vector, and further use the intention vector to perform an efficient approximate nearest neighbor search in a pre-constructed content library vector index, so as to quickly locate and screen out candidate content that highly matches the current intentions of users. Finally, several of the most relevant items are selected from the candidate list to obtain the final recommended list. In this way, the system can significantly optimize the content recommendation effect in the session scenario and improve the user experience by deeply utilizing the understanding of user dynamic intentions by the LLM and the generation ability of AIGC.

[0006] According to one aspect of the present application, there is provided a digital media content recommendation system based on AIGC, which includes: A behavior event collection module for collecting a series of behavior events of a target user in the current session; A dynamic intention description module for inputting a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description; A text embedding module for using a text embedding model to convert the dynamic intention description into an intention vector; An intention approximate search module for performing approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list; A content recommendation module for selecting the top N candidate contents from the candidate content list as the final content recommendation list.

[0007] Compared with the prior art, a digital media content recommendation system based on AIGC provided by the present application first uses a large language model (LLM) fine-tuned by instructions to interpret in real time a series of behavior events shown by a user in the current session, so as to intelligently generate text that can accurately describe the current dynamic intention of the user, which enables the system to keenly capture the subtle and immediate interest focus of the user; then, a text embedding model is used to convert the dynamic intention description into a vectorized representation to obtain an intention vector, and the intention vector is further used to perform efficient approximate nearest neighbor search in the pre-constructed content library vector index, so as to quickly locate and screen out candidate contents highly consistent with the current intention of the user. Finally, several most relevant items are selected from the candidate list to obtain the final recommendation list. In this way, the system can significantly optimize the content recommendation effect in the session scenario and improve the user experience by deeply utilizing the understanding of the user's dynamic intention by the LLM and the generation ability of AIGC. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 It is a block diagram of a digital media content recommendation system based on AIGC according to an embodiment of the present application; Figure 2 It is a data flow diagram of a digital media content recommendation system based on AIGC according to an embodiment of the present application; Figure 3Block diagram of a dynamic intent description module in an AIGC-based digital media content recommendation system according to an embodiment of the present application. Detailed implementation manners

[0010] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0011] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements.

[0012] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are merely illustrative, and different aspects of the system and method can use different modules.

[0013] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in order. Instead, as needed, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0014] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0015] In the technical solution of the present application, an AIGC-based digital media content recommendation system is proposed. Figure 1 Block diagram of an AIGC-based digital media content recommendation system according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of an AIGC-based digital media content recommendation system according to an embodiment of the present application. As Figure 1 and Figure 2As shown, the AIGC-based digital media content recommendation system 300 according to an embodiment of the present application includes: a behavior event collection module 310 for collecting a series of behavior events of a target user in the current session; a dynamic intention description module 320 for inputting the series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description; a text embedding module 330 for using a text embedding model to convert the dynamic intention description into an intention vector; an intention approximate search module 340 for performing approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list; and a content recommendation module 350 for selecting the top N candidate contents from the candidate content list as the final content recommendation list.

[0016] Specifically, the behavior event collection module 310 is used to collect a series of behavior events of a target user in the current session. It should be understood that in the current digital media environment, the interests and needs of users are often in a state of continuous change, especially in a specific interactive session, these changes are more significant and dynamic. Therefore, in order to capture the dynamic intention in the user's current session, enable the recommendation system to get rid of the dependence on historical static data of traditional models, and perceive the changing needs of users in real time to achieve accurate and real-time content recommendation, in the technical solution of the present application, a series of behavior events of the target user in the current session are collected. Among them, the behavior events include multi-dimensional interaction data such as the user's clicks, browsing duration, search keywords, scrolling behavior, likes, shares, comments, etc. These data comprehensively reflect the user's current focus of attention and interest trends. Through the collection and utilization of this real-time behavior data, the recommendation system can not only significantly improve the accuracy of identifying user interests, but also enhance the response ability to short-term hot content and cross-domain interest migration, thus bringing a more personalized, timely and user-real-needs-compliant recommendation effect, and greatly improving the user experience and the content distribution efficiency of the platform.

[0017] Specifically, the dynamic intention description module 320 is used to input a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description. In a specific example of the present application, as Figure 3 shown, the dynamic intention description module 320 includes: a semantic embedding encoding unit 321 for performing semantic embedding encoding on each behavior event in the series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors; a temporal context encoding unit 322 for performing behavior event temporal context encoding on the time series distribution of behavior event semantic embedding encoding vectors to obtain a user behavior pattern temporal hidden encoding vector; and a dynamic intention description generation unit 323 for obtaining a dynamic intention description based on the user behavior pattern temporal hidden encoding vector.

[0018] Specifically, the semantic embedding encoding unit 321 is configured to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors. It should be understood that a user may generate a large number of diverse behavior events in a single session, such as clicks, browsing, searches, likes, comments, etc. These behaviors not only reflect the user's immediate interest in the content but also contain their dynamically changing needs and preferences. Traditional recommendation systems often only perform simple statistics or tagging on these behaviors, making it difficult to capture their deep semantic information and temporal evolution relationships, resulting in a deviation between the recommended results and the user's true intentions. Therefore, in the technical solution of this application, a behavior event embedding matrix is used to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors. Through semantic embedding encoding, not only can the information carried by a single event be captured, but also the potential temporal dependencies and evolution laws between a series of events can be revealed.

[0019] Specifically, the temporal context encoding unit 322 is configured to perform behavioral event temporal context encoding on the time series distribution of the behavioral event semantic embedding encoding vectors to obtain the user behavior pattern temporal implicit encoding vectors. It should be understood that the interaction behavior between the user and the digital media platform is not an isolated event. Actions such as clicking, browsing, and pausing form a behavioral fluid with internal logic in the time dimension. The time series distribution of the behavioral event semantic embedding encoding vectors often contains dynamic interest trajectories and potential intentions. However, traditional recommendation systems are difficult to effectively capture the complex correlations of such temporal contexts. Therefore, to break through the limitations of static feature extraction and construct an encoding representation that can adaptively fuse temporal dependencies and dynamic intentions, in the technical solution of this application, behavioral event temporal context encoding is performed on the time series distribution of the behavioral event semantic embedding encoding vectors to obtain the user behavior pattern temporal implicit encoding vectors. In this process, each behavioral event is regarded as a dynamic entity, and the transmission potential energy and repulsive force coefficient between it and the current node are calculated. Essentially, a non-uniform and adaptive information propagation channel is established in the time flow. Among them, the transmission potential energy encodes the potential driving force of historical behaviors on the current intention (such as the interest reinforcement formed by repeatedly browsing similar content), while the repulsive force simulates the information decay caused by time distance or behavior type differences (such as the weak influence of a viewing record three days ago on the current recommendation). This mechanism enables the model to distinguish the "mainstream trend" and "instantaneous fluctuations" in the user behavior flow. For example, it can identify the intention turn of the user who frequently clicks on food videos but recently starts searching for fitness tutorials, thus avoiding misjudging short-term behavior noise as the core interest. In this way, the system can transform scattered behavioral events into a coherent "intention field", which not only retains the original temporal structure of the behavior sequence but also strengthens the implicit correlations across time points through the product state synergy effect. The generated user behavior pattern temporal implicit encoding vectors provide a semantic base with high information density for the subsequent large language model to generate accurate intention descriptions, enabling the recommendation system to break through the static preference assumption of traditional collaborative filtering and respond in real time to the dynamic needs of users in specific scenarios, thereby improving the fit between the recommended content and the immediate scenario.

[0020] Specifically, first, extract the current node feature vector of the behavior event from the time series distribution of the behavior event semantic embedding encoding vector, and define the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors. It should be understood that the real-time interaction between the user and the platform has significant spatio-temporal correlation. The influence of each behavior event is not evenly distributed, but shows non-linear attenuation or transition with its spatio-temporal distance and semantic correlation strength from the current moment. For example, when the user watches a short video, the previous like and favorite behaviors may form an "information potential energy" for the continuation of interest, while the quick swipe or exit operation in the middle may produce a "repulsion effect" that inhibits subsequent similar recommendations. The traditional processing methods of static time windows or fixed attenuation weights are difficult to capture such complex dynamics. Therefore, in the technical solution of this application, to establish a time series information processing structure that conforms to the cognitive law, extract the current node feature vector of the behavior event from the time series distribution of the behavior event semantic embedding encoding vector, and define the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors.

[0021] By taking the latest behavior node of the current session as the information aggregation anchor point and defining the historical behavior nodes as the to-be-propagated nodes, a dynamic perception field centered on the current intention is essentially constructed in the time flow, enabling the system to break through the uniform processing of time series data by traditional recommendation models. When the user browses digital media content, their behavior sequence naturally has the characteristic of attention focus. In a specific example, the behavior closer to the current moment often carries a higher-weight intention signal, but some early key behaviors (such as the first click on content in a certain vertical field) may continuously affect subsequent decisions due to the awakening of potential interest. By dividing the time series into the current node and the to-be-propagated nodes, the current node represents the explicit interaction focus of the user at this moment, and the to-be-propagated nodes form a potential influence factor pool. This structure provides a precise object of action for the calculation of potential energy and repulsion force in the subsequent liquid message propagation.

[0022] In a specific example of this application, extract the current node feature vector of the behavior event from the time series distribution of the behavior event semantic embedding encoding vector according to the following formula, and define the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector as the behavior event to-be-propagated node feature vectors to obtain the time series of the behavior event to-be-propagated node feature vectors; where the formula is: , , , Among them, is the time series distribution of the semantic embedding coding vectors of the behavior events, are the semantic embedding coding vectors of the behavior events at the 1st, 2nd, and Tth time steps in the time series distribution of the semantic embedding coding vectors of the behavior events respectively, T represents the total length of the sequence, and d represents the feature dimension of each time step. is the current node feature vector of the behavior event, is the time series of the feature vectors of the nodes to be propagated of the behavior event.

[0023] Next, calculate the dynamic transmission potential energy of each node feature vector to be propagated of the behavior event in the time series of the node feature vectors to be propagated of the behavior event relative to the current node feature vector of the behavior event to obtain the time series of the dynamic transmission potential energy coding vectors of the behavior event nodes. It should be understood that due to the strong temporal dependence and context sensitivity of the user's interaction behavior in the current session, simple time series modeling methods (such as RNN or LSTM) cannot fully capture the non-linear and non-steady dynamic coupling relationships between behavior events. By introducing the calculation of physically inspired dynamic transmission potential energy, the system can construct an interaction field of behavior event nodes in the high-dimensional latent space, thereby mapping the discrete behavior event nodes in the time series data into continuous vector representations with energy attributes. Therefore, to construct a deep temporal representation mechanism with the characteristics of physical field theory, by establishing a dynamic transmission potential energy mapping between behavior nodes, the discrete user behavior is transformed into the energy gradient distribution in the continuous semantic manifold. In the technical solution of this application, calculate the dynamic transmission potential energy of each node feature vector to be propagated of the behavior event in the time series of the node feature vectors to be propagated of the behavior event relative to the current node feature vector of the behavior event to obtain the time series of the dynamic transmission potential energy coding vectors of the behavior event nodes.

[0024] Among them, the dynamic transmission potential energy coding vector of the behavior event node transforms the non-linear interaction mode between the node feature vector to be propagated and the current node feature vector into the potential energy gradient distribution in the high-dimensional space through a parameterized learning mechanism. It not only bears the temporal correlation intensity between behavior events, but more importantly, reveals the implicit semantic associations across time steps, such as complex patterns of gradual change, jump, or superposition of user interests, etc., thereby providing a physically interpretable kinetic basis for subsequent simulated dynamic liquid message propagation. Through the dynamic transmission potential energy coding vector of the behavior event node, the system can effectively distinguish between surface behavior patterns and deep intention evolution. This potential field-based modeling method enables the recommendation system to break through the static preference assumption of traditional collaborative filtering, realize dynamic modeling and accurate response to the user's real-time intention, and ultimately improve the comprehensive performance of the recommendation results in terms of content novelty, temporal coherence, and interest matching degree.

[0025] In a specific example of the present application, the dynamic transmission potential energy of each behavior event to-be-propagated node feature vector in the time series of the behavior event to-be-propagated node feature vector relative to the current node feature vector of the behavior event is calculated according to the following formula to obtain the time series of the behavior event node dynamic transmission potential energy coding vector; wherein, the formula is: , wherein, and respectively represent a learnable query weight matrix and a key weight matrix, is the dynamic transmission potential energy factor, represents calculating and the squared difference between them, is the behavior event node dynamic transmission potential energy coding vector.

[0026] Subsequently, the behavior event message propagation repulsion coefficient of each behavior event to-be-propagated node feature vector in the time series of the behavior event to-be-propagated node feature vector relative to the current node feature vector of the behavior event is calculated to obtain the time series of the behavior event message propagation repulsion coefficient. It should be understood that the information propagation between different time node feature vectors in the user behavior sequence is not unconditionally unobstructed, and a dynamic modulation mechanism needs to be established to suppress the interference of irrelevant or noise signals. Although traditional time series modeling methods (such as attention mechanisms or gated networks) can achieve partial information filtering functions, their modeling of the hindrance effect of cross-time step information conduction is mostly based on empirical parameters and lacks interpretable physical constraints. Therefore, in order to quantify the negative impact weight of historical behavior events on the current intention modeling process and establish a dynamic suppression criterion for the liquid-like message propagation mechanism, in the technical solution of the present application, the behavior event message propagation repulsion coefficient of each behavior event to-be-propagated node feature vector in the time series of the behavior event to-be-propagated node feature vector relative to the current node feature vector of the behavior event is calculated to obtain the time series of the behavior event message propagation repulsion coefficient. By introducing the repulsion coefficient, the system can construct a dynamic impedance model of the information propagation channel based on the physical field theory, thus being more in line with the path dependence and context conflict existing in the evolution of user interests in the real scenario.

[0027] Specifically, the message propagation repulsion coefficient generates a scalarized conduction hindrance factor by calculating the state difference and temporal distance between the eigenvectors of the nodes to be propagated and the eigenvectors of the current node. This repulsion coefficient not only reflects the semantic heterogeneity between behavioral events (such as the interest conflict when a user switches from "reading technology news" to "watching entertainment videos"), but also captures the time decay effect (such as the weak correlation of early behavioral events with the current intention), thereby providing a dynamic modulation basis based on field theory for the subsequent liquid message propagation process. Through the generation of the time series of the repulsion coefficient, the system can dynamically suppress the information penetration intensity of the eigenvectors of the nodes with high state differences (such as cross-domain content interaction) or long time spans (such as historical behaviors outside short-term conversations) from the current user intention. This repulsion mechanism based on physical field theory enables the recommendation system to more precisely focus on the coherent and dense interaction patterns in the current user session, effectively filter the interference of irrelevant temporal signals, and then enhance the semantic consistency of dynamic intention description and the scenario adaptability of the recommendation results.

[0028] In a specific example of the present application, the following formula is used to calculate the message propagation repulsion coefficient of each behavioral event to-be-propagated node eigenvector in the time series of the behavioral event to-be-propagated node eigenvector relative to the behavioral event current node eigenvector to obtain the time series of the behavioral event message propagation repulsion coefficient, where the formula is: , where, represents vector multiplication, is a timestamp extraction function, represents a time difference scaling factor, represents the square of the norm, is the message propagation repulsion coefficient of the behavioral event.

[0029] Then, based on the time series of the behavior event message propagation repulsion coefficient and the time series of the behavior event node dynamic transmission potential energy encoding vector, perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of behavior event node message dynamic encoding vectors. It should be understood that the information transmission between user behavior nodes has non-linear coupling characteristics, and a multi-body interaction model inspired by physical field theory is required to achieve dynamic semantic fusion across time nodes. Traditional recommendation systems use fixed time decay functions or shallow attention mechanisms, which are difficult to capture the quantized semantic transitions (such as sudden migrations of interest topics) and multi-scale time dependencies (such as the synergistic effect of high-frequency click behaviors and low-frequency in-depth browsing) implicit in behavior sequences. Therefore, in order to establish a dynamic message generation mechanism controlled by dual physical field parameters and achieve refined modulation of the influence of historical behavior events on the current user's intention, in the technical solution of this application, based on the time series of the behavior event message propagation repulsion coefficient and the time series of the behavior event node dynamic transmission potential energy encoding vector, perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of behavior event node message dynamic encoding vectors.

[0030] That is, by non-linearly coupling the behavior event node dynamic transmission potential energy encoding vector (representing the driving strength of information transmission between behavior events) and the behavior event message propagation repulsion coefficient (representing the dynamic impedance of the information channel), a behavior event node message dynamic encoding vector with directionality and attenuation characteristics is generated. This coupling mechanism not only simulates the interaction dynamics characteristics of multi-body particles in a liquid medium (such as the balance between viscous resistance and potential energy gradient), but more importantly, it transforms the cross-time-step correlations (such as the continuity of interest migration and the suddenness of exploration behaviors) implicit in the user behavior sequence into a differentiable high-dimensional vector space mapping relationship, providing a physically interpretable intermediate representation for subsequent multi-body propagation co-renormalization.

[0031] In a specific example of this application, the following formula is used to perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated to obtain a set of behavior event node message dynamic encoding vectors; where the formula is: , where, denotes element-wise multiplication by position, denotes function, denotes function, denotes a learnable value weight matrix, is the behavior event node message dynamic encoding vector.

[0032] Furthermore, multi-body propagation collaborative renormalization is performed on the set of dynamic coding vectors of behavioral event node messages to obtain an optimized set of dynamic coding vectors of behavioral event node messages. It should be understood that after the dynamic message coding vectors in the user behavior sequence are propagated in an imitation liquid state, their internal correlation presents high-dimensional manifold characteristics, including both the spatio-temporal correlation of local propagation lattice points and the implicit coupling effect across time steps. Traditional time-series aggregation methods (such as simple vector superposition or average pooling) cannot effectively handle the multi-body collaborative effect generated in the non-linear propagation process. Therefore, in a preferred example of this application, in order to establish a physically interpretable multi-body collaborative representation reconstruction framework, multi-body propagation collaborative renormalization is performed on the set of dynamic coding vectors of behavioral event node messages to obtain an optimized set of dynamic coding vectors of behavioral event node messages.

[0033] Specifically, by introducing a global scale correction factor μ (a renormalization operator based on the mean eigenvalue) and an auxiliary propagation bond channel, conformal mapping and dynamic equilibrium of cross-time node information are achieved. In a real-time recommendation scenario, when the user behavior stream contains a mixed mode of high-frequency interactions (such as continuously clicking on similar videos) and low-frequency dwells (such as concentrating on watching tutorials), the renormalization process adaptively adjusts the contribution weights of each node to the homogeneous manifold space through the non-orthogonal basis projection of the product state, ensuring that short-term behavior noise does not overwhelm long-term interest features while retaining the collaborative effect of cross-modal semantics. In this way, the finally generated sequential implicit coding vectors of the user behavior pattern not only retain the mutation characteristics of short-term interests but also maintain the continuity of long-term behavior inertia, providing high-quality feature inputs with physical significance for the downstream large language model to generate accurate dynamic intention descriptions.

[0034] In the above process, for the set of dynamic coding vectors of the behavioral event node messages, due to the adaptability of the effect propagation under non-linear representation, each dynamic coding vector of the behavioral event node message essentially constitutes a vector local propagation lattice point that can be precisely decoupled. Therefore, when aggregating each local propagation lattice point, in addition to simple superposition, it is also desirable to introduce a local propagation lattice point auxiliary bond correction that conforms to the propagation characteristics.

[0035] That is, a vector product state is introduced to simulate the auxiliary propagation bond channel of the local propagation lattice point, so as to reflect the propagation behavior in the multi-body effect with a product structure, enabling the collaborative effect to emerge naturally: , Moreover, in order to perform collaborative renormalization under the vector product state that can be precisely decoupled, a global scale representation correction is introduced: , where It is the mean value of all the eigenvalues of the set of the dynamic encoding vectors of the behavior event node messages.

[0036] And, considering the global correction For the scale transformation in the product state, it is necessary to perform the partial derivative calculation with respect to That is: , Thus, we get: , where is the hyperparameter for linear compression .

[0037] Then, optimize the dynamic encoding vectors of the behavior event node messages: , In this way, while ensuring the covariance preservation under the scale transformation, the renormalization cooperation in the vector product state with exact decoupling is carried out. Thus, by introducing the auxiliary propagation bond channels of the local propagation lattice points, the representation of the many-body effect is strengthened, and the aggregation effect of the dynamic timing information of the dynamic encoding vectors of the behavior event node messages for each time point on the dynamic encoding vectors of the behavior event node messages is improved.

[0038] Subsequently, the set of dynamically encoded vectors of the fused and optimized behavioral event node messages and the current node feature vector of the behavioral event are combined to obtain the sequential implicit encoding vector of the user behavior pattern. It should be understood that the real-time interaction behaviors generated by the user in the current session not only need to immediately capture their surface operation features but also need to reveal the potential intention evolution path through the non-linear propagation of the historical behavior sequence. The set of dynamically encoded vectors processed by multi-body propagation collaborative renormalization carries the propagation collaborative effect and information screening mechanism across time steps, while the current node feature vector represents the most forefront state of the user interaction. Therefore, to break through the dimensionality limitation of static feature splicing and achieve the deep coupling of the dynamic propagation field and the real-time state field, in the technical solution of this application, the set of dynamically encoded vectors of the fused and optimized behavioral event node messages and the current node feature vector of the behavioral event are combined to obtain the sequential implicit encoding vector of the user behavior pattern. That is, by introducing a spatial fusion mechanism based on the product state, the system aligns the manifold and entangles the information of the renormalized dynamically encoded vector (containing the non-linear correlation and multi-body collaborative effect between historical behaviors) and the current node feature vector (reflecting the fine-grained semantics of the latest interaction event), generating an implicit encoding with spatio-temporal invariance. This fusion process is not a simple vector superposition but a differential geometric operation that preserves scale transformation covariance, constructing a cross-modal semantic resonance channel while retaining the independent propagation characteristics of each information flow. This dual information integration mechanism effectively solves the trade-off dilemma between real-time and continuity in traditional recommendation systems, enabling the dynamic intention description generation module to output a recommendation strategy that not only conforms to the long-term interest evolution trend but also precisely matches the current session focus based on high-fidelity spatio-temporal features, ultimately achieving personalized content matching across time scales.

[0039] In a specific example of this application, the set of dynamically encoded vectors of the fused and optimized behavioral event node messages and the current node feature vector of the behavioral event are combined according to the following formula to obtain the sequential implicit encoding vector of the user behavior pattern, where the formula is: , , where, is the dynamically encoded vector of the optimized behavioral event node message, represents function, represents the gating weight matrix, is the sequential implicit encoding vector of the user behavior pattern.

[0040] Specifically, the dynamic intention description generation unit 323 is configured to obtain a dynamic intention description based on the implicit encoding vector of the user behavior pattern time series. It should be understood that in the current environment where digital media content is increasingly rich and user interests change rapidly, simply relying on traditional behavioral data analysis is difficult to accurately capture the immediate needs and complex intentions of users. The large language model can infer the user's current specific interest points, content preferences, or the type of content being sought in a more user-friendly and contextually relevant manner based on the embedded user behavior pattern time series information, rather than just staying at the similarity comparison in the traditional vector space. This not only helps to reveal the diverse hidden needs of users but also effectively addresses the challenges of short-term interest fluctuations and cross-domain interest migrations. Therefore, in the technical solution of this application, after embedding the implicit encoding vector of the user behavior pattern time series into a preset Prompt, it is input into the large language model fine-tuned by instructions to obtain a dynamic intention description. The preset Prompt is: Based on the semantic information of the following recent interaction behavior sequence between the user and digital media, infer the user's current specific interest points, content preferences, or the type of content being searched for, and generate a concise intention description text. That is, through the structured guidance of the preset Prompt (such as clearly requiring the inference of interest points, content preferences, and search types), the generation direction of the large language model is constrained, enabling the large language model to generate intention texts that meet the scenario requirements on the basis of understanding the user interaction behavior sequence, enhancing the consistency and controllability of the model output results. Here, the dynamic intention description can not only accurately reflect the core needs of the user in the current conversation stage but also guide the recommendation system to be more focused on the user's true intentions during content retrieval, improving the relevance and personalization of recommendations, providing a structured and semantically rich data representation for the downstream text embedding and vector retrieval modules, thereby significantly improving the candidate content matching accuracy and the final recommendation quality.

[0041] Particularly, the text embedding module 330 is configured to convert the dynamic intention description into an intention vector using a text embedding model. That is, in the technical solution of this application, a text embedding model based on the Word2vec model is used to perform embedding encoding on the dynamic intention description to obtain an intention vector in order to construct a semantically continuous and structured vector space. That is, through the text embedding based on the Word2vec model, the dynamic intention description is encoded to map the user's current interest points, content preferences, or requirements into a low-dimensional and semantically rich space in the form of a continuous vector, so that the recommendation system can more accurately understand and quantify the semantic connotation of the user's intentions. By converting the dynamic intention into an intention vector, the system can quickly identify and match candidate content highly relevant to the user's current needs, improving the relevance and real-time response ability of the recommendation.

[0042] Specifically, the intention approximate search module 340 is used to perform approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a list of candidate contents. Among them, the dynamic intention description generated by the large language model based on AIGC is converted into a semantically rich intention vector through text embedding. This vector form can highly condense the user's real-time interest information and complex intentions, laying a solid foundation for content matching. In the technical solution of this application, to improve the relevance and real-time performance of recommendations and enable the system to accurately capture and respond to the latest interest points and content preferences shown by the user in the current session, by performing approximate nearest neighbor search in the vector index of the content library using the intention vector, efficient and semantically accurate content retrieval can be achieved. It is worth mentioning that approximate nearest neighbor search, as an efficient vector retrieval technology, can quickly locate the candidate content closest to the user's intention vector in the large-scale content vector space. In the specific implementation process, each piece of content in the content library has been pre-converted into a corresponding vector representation and a corresponding vector index structure has been established, such as an index based on the ANN (Approximate Nearest Neighbor) algorithm, such as HNSW, Faiss, etc. During this process, the system inputs the user's current intention vector and performs approximate search through the vector index structure to quickly extract several candidate contents that are semantically most similar to the user's dynamic intention. In this way, the probability of matching high-correlation candidate contents is greatly improved, and the coincidence degree between the recommended result and the user's real needs is significantly improved.

[0043] Specifically, the content recommendation module 350 is configured to select the top N candidate contents from the candidate content list as the final content recommendation list. It should be understood that the candidate content list often contains a large number of contents with varying degrees of relevance. Pushing all of them directly will not only cause information overload but also may affect the user experience and the utilization efficiency of platform resources. Therefore, it is necessary to further screen and rank the candidate contents to select the top N highest-quality and most matching contents as the final recommendation result. Thus, in the technical solution of this application, the top N candidate contents are selected from the candidate content list as the final content recommendation list. It is worth mentioning that the system will first set a default N value in the configuration and also support dynamic adjustment according to personalized strategies. For example, the recommended quantity can be appropriately increased for highly active users or member users, while the display can be streamlined for new users or low-active users. In addition, the relevance score threshold output by the algorithm model can be combined to only retain the top N contents with the highest scores and meeting the quality standards, so as to ensure that the contents finally presented to the user are the highest-quality and most in line with their current intentions. This method can effectively reduce the interference of redundant information, improve the utilization rate of platform resources, speed up the page loading speed, and at the same time improve the accuracy of recommendation hits and the click-through conversion effect. Finally, by selecting the top N from the candidate list as the final recommendation, it not only enhances the system's adaptability to short-term interest changes and complex demand scenarios but also greatly improves user satisfaction and platform stickiness, achieving the goals of precision, efficiency, and sustainable development in the field of digital media intelligent distribution.

[0044] As described above, the AIGC-based digital media content recommendation system 300 according to the embodiments of this application can be implemented in various wireless terminals, such as a server with an AIGC-based digital media content recommendation algorithm. In a possible implementation manner, the AIGC-based digital media content recommendation system 300 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the AIGC-based digital media content recommendation system 300 can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the AIGC-based digital media content recommendation system 300 can also be one of the many hardware modules of the wireless terminal.

[0045] Alternatively, in another example, the AIGC-based digital media content recommendation system 300 and the wireless terminal can also be separate devices, and the AIGC-based digital media content recommendation system 300 can be connected to the wireless terminal through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.

[0046] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A digital media content recommendation system based on AIGC, characterized in that, Including: A behavior event collection module, used to collect a series of behavior events of the target user in the current session; A dynamic intention description module, used to input a series of behavior events into a large language model fine-tuned by instructions to obtain a dynamic intention description; A text embedding module, used to use a text embedding model to convert the dynamic intention description into an intention vector; An intention approximate search module, used to perform approximate nearest neighbor search in the vector index of the content library using the intention vector to obtain a candidate content list; A content recommendation module, used to select the top N candidate contents from the candidate content list as the final content recommendation list.

2. The digital media content recommendation system based on AIGC according to claim 1, characterized in that The dynamic intention description module includes: A semantic embedding encoding unit, used to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors; A time series context encoding unit, used to perform behavior event time series context encoding on the time series distribution of behavior event semantic embedding encoding vectors to obtain a user behavior pattern time series implicit encoding vector; A dynamic intention description generation unit, used to obtain a dynamic intention description based on the user behavior pattern time series implicit encoding vector.

3. The AIGC-based digital media content recommendation system according to claim 2, wherein The semantic embedding encoding unit is used for: Using a behavior event embedding matrix to perform semantic embedding encoding on each behavior event in a series of behavior events to obtain a time series distribution of behavior event semantic embedding encoding vectors.

4. The AIGC-based digital media content recommendation system according to claim 3, wherein, The time series context encoding unit includes: A node extraction sub-unit, used to extract the current node feature vector of the behavior event from the time series distribution of the behavior event semantic embedding encoding vector, and define the other node feature vectors of the behavior event in the time series distribution of the behavior event semantic embedding encoding vector as the behavior event to-be-propagated node feature vector to obtain a time series of behavior event to-be-propagated node feature vectors; An imitation dynamic liquid message propagation sub-unit, used to perform behavior event imitation dynamic liquid message propagation on each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors based on the behavior event liquid interaction field between each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors and the current node feature vector of the behavior event to obtain a set of behavior event node message dynamic encoding vectors; A fusion sub-unit, used to fuse the set of behavior event node message dynamic encoding vectors and the current node feature vector of the behavior event to obtain a user behavior pattern time series implicit encoding vector.

5. The AIGC-based digital media content recommendation system according to claim 4, wherein, The imitation dynamic liquid message propagation sub-unit is used for: Calculating the behavior event node dynamic transmission potential energy of each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors relative to the current node feature vector of the behavior event to obtain a time series of behavior event node dynamic transmission potential energy encoding vectors; Calculating the behavior event message propagation repulsion coefficient of each behavior event to-be-propagated node feature vector in the time series of behavior event to-be-propagated node feature vectors relative to the current node feature vector of the behavior event to obtain a time series of behavior event message propagation repulsion coefficients; Based on the time series of the repulsive force coefficient of behavior event message propagation and the time series of the dynamic transmission potential energy coding vectors of behavior event node, perform pseudo-dynamic liquid message propagation on each behavior event node feature vector in the time series of the behavior event node feature vectors to be propagated, so as to obtain a set of behavior event node message dynamic coding vectors.

6. The AIGC-based digital media content recommendation system according to claim 5, wherein The fusion subunit is used for: Performing multi-body propagation collaborative renormalization on the set of behavior event node message dynamic coding vectors to obtain an optimized set of behavior event node message dynamic coding vectors; Fusing the optimized set of behavior event node message dynamic coding vectors and the current behavior event node feature vectors to obtain the user behavior pattern temporal implicit coding vectors.

7. The AIGC-based digital media content recommendation system according to claim 6, wherein The dynamic intention description generation unit is used for: After embedding the user behavior pattern temporal implicit coding vectors into the preset Prompt, input them into the large language model fine-tuned by instructions to obtain the dynamic intention description.

8. The AIGC-based digital media content recommendation system according to claim 7, wherein The preset Prompt is: According to the semantic information of the following sequence of the user's recent interaction behaviors with digital media, infer the user's current specific interest points, content preferences or the type of content being searched for, and generate a concise intention description text.

9. The AIGC-based digital media content recommendation system according to claim 8, wherein The text embedding module is used for: Using a text embedding model based on the Word2vec model to perform embedding coding on the dynamic intention description to obtain the intention vector.

Citation Information

Patent Citations

  • Teenager algorithm code auxiliary learning system and method based on large language model

    CN117235347A

  • Fine adjustment method, system and equipment based on large language model and medium

    CN117290480A

  • Multi-modal question and answer interpretation method and related device

    CN118689997A

  • User preference analysis method based on memory enhanced multi-modal large model

    CN119884981A

  • System for providing korean sentence rephraser service

    US20250147996A1

Cited By

  • Electric energy quality comprehensive monitoring method and system based on real-time waveform analysis

    CN120831528A

  • Catering recommendation method and system applied to digital multimedia online platform

    CN121743559A

  • Dining recommendation method and system applied to digital multimedia online platform

    CN121743559B