A Controllable Multimodal Personalized Recommendation System and Method Based on LLM
By constructing a controllable multimodal personalized recommendation system based on LLM, the problems of dependence on historical behavioral data and insufficient utilization of multimodal information in existing technologies are solved. It achieves a deep understanding of natural language queries and complex intents, enhances the controllability and interpretability of the recommendation process, and optimizes long-term user value and short-term experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA NORMAL UNIV
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing recommendation systems rely excessively on historical behavior data, lack deep semantic understanding and dynamic strategy fusion capabilities, making it difficult to optimize long-term user value. Furthermore, the insufficient utilization of multimodal information leads to cold start dilemmas and a lack of controllability and interpretability in the recommendation process.
A controllable multimodal personalized recommendation system based on LLM is constructed, including modules for data collection, preprocessing, initial understanding and selection, fusion, rapid screening, scoring and recommendation, and optimization. The system uses LLM to parse user requests and generate semantic representations, constructs a unified multimodal representation network, uses a lightweight ranking network for rapid screening, and performs real-time evaluation and parameter correction based on comprehensive utility scores to achieve closed-loop adaptive optimization throughout the entire process.
It significantly improves the system's ability to deeply understand natural language queries and complex intents, achieves unified representation of multimodal data, enhances the controllability and interpretability of the recommendation process, optimizes long-term user value and short-term experience, and ensures the real-time performance and accuracy of recommendations.
Smart Images

Figure CN122089441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data processing technology, specifically to a controllable multimodal personalized recommendation system and method based on LLM. Background Technology
[0002] With the rapid development of the internet, artificial intelligence, and mobile devices, personalized recommendation systems have become a core engine for e-commerce, short video, and news platforms to improve user experience and business conversion. Faced with increasingly complex application scenarios and diversified needs, recommendation systems are facing greater challenges in terms of real-time performance, accuracy, and intelligence.
[0003] Existing recommendation systems generally adopt a two-stage "recall-ranking" architecture: first, collaborative filtering and vector retrieval techniques are used to quickly filter the candidate set, and then deep neural networks are used to accurately score and rank them. This method relies on large-scale historical behavioral data, performs robustly in scenarios with sufficient data, and has been widely used in mainstream recommendation platforms.
[0004] However, with the enrichment of user expression methods and the increase in scenario complexity, this technical approach has gradually exposed multiple limitations: First, the system's semantic understanding of natural language input, complex intents, and contextual information remains insufficient, and it faces a severe cold start dilemma in scenarios with new users or new content; second, existing models mostly focus on short-term metrics such as click-through rate and exposure rate, making it difficult to effectively optimize user retention and long-term stickiness; third, the recommendation process lacks sufficient controllability and interpretability, making it difficult to flexibly integrate into business strategies and compliance rules; and finally, the mining and utilization of multimodal information such as images, voice, and video are still insufficient, failing to fully leverage the value of heterogeneous data.
[0005] Chinese Patent Publication No. CN102663627A discloses a personalized recommendation method, which includes steps such as collecting basic data; processing and classifying the basic data, analyzing the buyer's interests and preferences based on the buyer's historical operation behavior records; filtering and organizing seller and product information according to pre-set rules; deriving a set of products to be recommended based on a pre-set matching algorithm, according to the buyer's interest and preference data and the seller and product information; deduplicating, sorting, and weighting all products to be recommended to obtain N products with the highest matching degree as the recommendation results, and displaying the final recommendation results to the buyer. Although this invention can build a recommendation framework based on pre-set rules and historical behavior statistics, it still cannot get rid of its high dependence on historical behavior data, cannot effectively understand natural language queries and potential user intentions, and cannot flexibly integrate dynamic business strategies or multimodal information. It is difficult to meet the requirements of intelligence and generalization capabilities in modern complex recommendation scenarios. How to design a unified recommendation framework that combines deep semantic understanding, cold start mitigation, and flexible injection of business rules to break through the excessive dependence on historical behavior data and optimize long-term user value has become a key technical problem that urgently needs to be solved. Summary of the Invention
[0006] This invention provides a controllable multimodal personalized recommendation system and method based on LLM, which overcomes the technical limitations of existing recommendation systems, such as over-reliance on historical behavior data, lack of deep semantic understanding and dynamic strategy fusion capabilities, and difficulty in optimizing long-term user value.
[0007] In a first aspect, the present invention provides a controllable multimodal personalized recommendation system based on LLM, comprising, The data collection module is used to collect user behavior data, content data, and contextual information in real time. A preprocessing module, connected to the acquisition module, is used to preprocess the acquired data and to construct several user behavior sequences based on user ID and time order. The initial selection module, which is connected to the preprocessing module, is used to obtain user requests from the user behavior data and generate semantic vectors, and to determine the initial recommendation set based on the semantic vectors. A fusion module, which is connected to the preprocessing module and the initial understanding selection module, is used to map the semantic vector and the user behavior sequence into a user interaction vector based on a pre-trained fusion model. The quick screening module, which is connected to the understanding and initial selection module and the preprocessing module, is used to construct quick screening vectors based on user historical features, and to output a quick screening set based on the pre-trained quick screening model to initially screen the initial recommendation set. The rating and recommendation module is connected to the quick screening module and the fusion module. It is used to construct several rating vectors based on the quick screening set and the user interaction vectors, determine short-term ratings and long-term ratings based on the rating vectors, and comprehensively evaluate and output a new recommended solution based on the short-term ratings and long-term ratings. An optimization module, connected to the acquisition module, is used to determine a comprehensive utility score based on the user retention rate ratio, click-through rate ratio, preset click-through rate, and short-term loss penalty parameters within the current period; to determine whether the recommendation effect is qualified based on the comprehensive utility score; and to output instructions to correct the LLM parameters, quick screening parameters, and new scheme determination parameters when the effect is unqualified. The scheduling module, which is connected to the understanding and initial selection module, the rapid screening module, the recommendation module, and the optimization module, is used to correct the corresponding parameters based on the instructions from the optimization module.
[0008] Furthermore, the controllable multimodal personalized recommendation system based on LLM also includes an explanation module, which is connected to the rating recommendation module, and is used to generate natural language recommendation reasons for the new solution based on LLM; The fusion module is also used to construct a fusion model based on a unified multimodal representation network, and to train the fusion model based on the relationship between user interaction vectors in historical traffic data and the semantic vectors and the user behavior sequences. The rapid screening module is also used to construct a rapid screening model based on a lightweight ranking network, and to train the rapid screening model based on the relationship between the historical initial screening vector - initial recommendation set vector and the rapid screening set.
[0009] Furthermore, the optimization module is also used to determine the comprehensive utility score based on the user retention rate ratio, the click-through rate ratio, the preset click-through rate, and the short-term loss penalty parameter within the current period, and to determine whether the recommendation effect within the current period is qualified based on the comprehensive utility score; The optimization module is also used to maintain and continuously monitor system parameters when the system is deemed qualified, or to correct the temperature value of the LLM in the initial selection module based on the ratio of the preset comprehensive utility score to the comprehensive utility score when the system is deemed unqualified. The short-term loss penalty parameter refers to the penalty level imposed by the system on short-term losses. The user retention rate ratio refers to the ratio of the user retention rate when the new scheme is adopted in the current period to the user retention rate of the old baseline. The click-through rate ratio refers to the ratio of the click-through rate when using the new scheme during the current period to the click-through rate of the old baseline.
[0010] Furthermore, the optimization module is also used to determine a utility ratio based on the ratio of the preset comprehensive utility score to the comprehensive utility score, and to reduce the temperature value based on the utility ratio, wherein the reduction in temperature value is proportional to the utility ratio.
[0011] Furthermore, the optimization module is also used to determine the temperature difference based on the difference between the temperature values before and after the correction, and to increase the kernel sampling value of the LLM in the initial selection module based on the temperature difference, wherein the increase in the kernel sampling value is proportional to the temperature difference.
[0012] Furthermore, the optimization module is also used to determine the kernel sampling difference based on the difference between the kernel sampling values before and after the correction, and to increase the Top K value of the rapid screening model based on the kernel sampling difference, wherein the increase in the Top K value is proportional to the kernel sampling difference.
[0013] Furthermore, the optimization module is also used to determine whether the recommendation effect is qualified based on the comprehensive utility score in the first period after correction, and to maintain the system parameters and continue monitoring when the result is qualified, or to correct the exploration coefficient in the confidence upper bound algorithm based on the ratio of the preset long-term solution ratio to the long-term solution ratio in the first period after correction when the result is unqualified. The long-term solution refers to the solution whose long-term score is greater than the short-term score, and the proportion of long-term solutions refers to the proportion of long-term solutions among all new solutions output in this period.
[0014] Furthermore, the optimization module is also used to determine the long-term ratio based on the ratio of the preset long-term solution ratio to the ratio of the long-term solution ratio in the first period after correction, and to increase the exploration coefficient in the confidence upper bound algorithm based on the long-term ratio, wherein the increase in the exploration coefficient is proportional to the long-term ratio.
[0015] Furthermore, the optimization module is also used to determine the difference in exploration coefficients based on the difference in exploration coefficients before and after the correction, and to increase the attempt threshold in the confidence upper bound algorithm based on the difference in exploration coefficients, wherein the increase in the attempt threshold is proportional to the difference in exploration coefficients.
[0016] Secondly, this invention also provides a controllable multimodal personalized recommendation method based on LLM, including: Real-time collection of user behavior data, content data, and contextual information; The preprocessed data is used to construct several user behavior sequences based on user ID and time order. The user requests in the user behavior data are obtained, and several semantic expressions are generated by parsing the user requests based on LLM, and each semantic expression is converted into a semantic vector. The initial recommendation set is obtained by searching the content database based on the semantic vector. The pre-trained fusion model maps the semantic vector and the user behavior sequence to the same semantic space and outputs a user interaction vector. Based on the user ID, location, user historical click list, and user long-term interest tags, a quick screening vector is constructed, and a quick screening set is output by using the quick screening vector to perform initial screening on the initial recommendation set based on the pre-trained quick screening model. Based on the rapid screening set and the user interaction vector, several scoring vectors are constructed, and the scoring vectors are scored in the short term based on the fine ranking model, scored in the long term based on LLM, and a new recommended solution is output by comprehensively evaluating the short-term and long-term scores using the confidence upper bound algorithm. The comprehensive utility score is determined based on the user retention rate ratio, click-through rate ratio, and short-term loss penalty parameters of the new solution relative to the old baseline within this period. The comprehensive utility score is used to determine whether the recommendation effect is satisfactory. If it is unsatisfactory, the LLM parameters, quick screening parameters, and new solution determination parameters are adjusted.
[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention discloses a controllable multimodal personalized recommendation system based on LLM. By constructing a complete system architecture including a data acquisition module, a preprocessing module, an understanding and initial selection module, a fusion module, a rapid screening module, a rating and recommendation module, an optimization module, and a scheduling module, it effectively overcomes the cold start dilemma and insufficient semantic understanding problems caused by the high dependence of existing recommendation systems on historical behavioral data. By utilizing LLM to parse user requests, generate semantic expressions, and convert them into semantic vectors for initial selection and recall, it significantly improves the system's ability to deeply understand natural language queries, complex intents, and contextual information. Simultaneously, the fusion module realizes the integration of multimodal data... The unified representation solves the technical bottleneck of insufficient mining of heterogeneous data in traditional solutions. More importantly, the optimization module evaluates the recommendation effect in real time based on the comprehensive utility score, and dynamically adjusts the LLM parameters, quick screening parameters and new solution determination parameters through the scheduling module when the recommendation is deemed unqualified. This enables the entire system to form a closed-loop adaptive optimization mechanism from data collection, semantic understanding, candidate screening to scoring decision. This not only enhances the controllability of the recommendation process, but also quantifies the long-term user value and short-term experience constraints through the utility score, preventing the optimization process from damaging the core user experience. Thus, while improving the accuracy of recommendations, it also enables the flexible injection of business strategies and continuous optimization of long-term user stickiness.
[0018] Furthermore, by adding an explanation module to generate natural language recommendation reasons, the interpretability and transparency of the recommendation system are significantly enhanced, enabling users to understand the decision-making basis and thus improving user trust and satisfaction. Simultaneously, the fusion module uses a unified multimodal representation network to construct a fusion model, achieving effective mapping between semantic vectors and user behavior sequences in the same semantic space. This solves the technical challenges of isolated processing and insufficient fusion of multimodal information in traditional recommendation systems, improving the comprehensiveness and accuracy of user interest modeling. In addition, the rapid screening module uses a lightweight ranking network to construct a rapid screening model, achieving rapid filtering of the initial recommendation set while ensuring screening efficiency. This reduces the computational pressure on the downstream scoring module and ensures the real-time nature of the recommendation response through reasonable candidate set compression, thus achieving a good balance between overall system performance and computational resource consumption.
[0019] Furthermore, by constructing a comprehensive utility score calculation system based on user retention ratio, click-through rate ratio, preset click-through rate, and short-term loss penalty parameters, a unified and clear decision criterion is provided for the recommendation system. This integrates the conflicting goals of maximizing long-term user value with short-term experience constraints into a single quantitative indicator. Furthermore, by embedding a business risk tolerance within the preset click-through rate, it prevents excessive damage to the core user experience during the optimization process. This comprehensive utility score can be calculated periodically and continuously monitored, promptly identifying performance degradation caused by strategy drift, changes in data distribution, or improper algorithm exploration, providing an objective basis for adjusting system parameters. Simultaneously, by triggering a correction mechanism for the LLM temperature value when the recommendation effect is deemed unsatisfactory, a technical path of constraining semantic generation quality from the source is achieved, ensuring that the system can respond quickly and perform targeted optimizations when performance deteriorates, rather than blindly adjusting irrelevant modules.
[0020] Furthermore, by reducing the temperature value of LLM based on the utility ratio and making the reduction proportional to the utility ratio, a dynamic match between the intensity of parameter correction and the severity of recommendation performance degradation is achieved. When a large utility ratio indicates that the system is far from the optimal state, the temperature can be significantly reduced to significantly suppress the randomness and creativity of LLM, forcing the generation of more deterministic, conservative, and accurate semantic expressions, thereby quickly correcting semantic parsing biases. When the performance deviation is small, only the temperature needs to be fine-tuned to avoid the loss of semantic diversity caused by excessive constraints. This adaptive adjustment mechanism can both make aggressive corrections in severely degraded scenarios to quickly restore system performance and maintain the richness of semantic expressions in the case of slight fluctuations, avoiding the blindness of manual experience-based parameter tuning, and enabling the system to have quantitatively driven self-optimization capabilities and faster convergence speed.
[0021] Furthermore, a dynamic complementary balance strategy among LLM generation parameters was constructed by increasing the kernel sampling value proportionally to the temperature difference after reducing the temperature value. While reducing the temperature can suppress randomness and improve accuracy, it can also lead to a sharper output probability distribution, causing the model to over-concentrate on high-probability candidate words and fall into pattern collapse or semantic homogenization. In this case, increasing the kernel sampling value can effectively broaden the candidate word selection space and compensate for the diversity loss under the low temperature setting. When the temperature is significantly reduced, the kernel sampling value is correspondingly significantly increased to ensure that sufficient candidate diversity is still maintained on the basis of forced determinism. When the temperature is fine-tuned, it is moderately increased to avoid overcorrection. This two-way balance mechanism can quickly correct recommendation bias through rapid temperature reduction and prevent candidate semantic homogenization caused by excessive conservatism through dynamic compensation of kernel sampling value. Thus, it maintains a dynamic balance between accuracy and diversity during system optimization and enhances the robustness and convergence stability of the recommendation system in the face of performance fluctuations.
[0022] Furthermore, by increasing the Top K value of the rapid screening model proportionally to the kernel sampling difference, a parameter coordination and dynamic adaptation mechanism between upstream semantic understanding and downstream candidate screening was established. Since the cascading adjustment of decreasing temperature and increasing kernel sampling value has substantially broadened the semantic distribution space of the initial recommendation set, if the rapid screening module maintains the original Top K value, high-quality content may be excessively truncated in the initial screening stage. Therefore, the kernel sampling difference, as a quantitative indicator reflecting the degree of adjustment of the LLM generation strategy towards diversity, can guide the rapid screening threshold to be accurately matched. When the kernel sampling difference is large, indicating a high degree of semantic divergence in the upstream initial selection set, the Top K value is significantly increased to retain more candidates for the downstream fine ranking stage, ensuring that the diversity gain is not negated by the rapid screening stage. When the difference is small, it is moderately relaxed to avoid excessive expansion of the candidate pool. This design not only prevents the loss of high-quality candidates due to mismatch between upstream and downstream parameters, but also ensures the system's fine-grained response capability when facing different adjustment ranges, thereby maximizing the overall coverage accuracy and diversity balance of the recommendation system while maintaining computational efficiency.
[0023] Furthermore, after correcting the Top K value of the rapid screening module based on the kernel sampling difference, the recommendation effect is judged again based on the comprehensive utility score in the first cycle after correction. This constructs a full-link hierarchical diagnosis and repair mechanism from upstream semantic understanding, midstream candidate screening to downstream scoring decision. If it is judged as unqualified again, it indicates that the root cause of the problem is not the defects in candidate set generation or initial screening, but the decision-making balance between short-term benefits and long-term value in the downstream scoring recommendation module is biased. At this time, by correcting the exploration coefficient in the confidence upper bound algorithm based on the ratio of the preset long-term solution ratio to the actual long-term solution ratio, the exploration-utilization imbalance problem in the final decision layer can be accurately located and corrected. This phased verification mechanism avoids blindly continuing to adjust upstream parameters when the candidate quality has been optimized, ensuring that the system can make precise interventions according to the fault level. It not only prevents ineffective adjustments of irrelevant modules, but also ensures that the system takes into account the dual optimization goals of short-term click-through rate and long-term user retention during the repair process by quantifying the dynamic coupling of the long-term solution ratio and the exploration coefficient, thereby improving the sustainability of the recommendation strategy and the user lifetime value.
[0024] Furthermore, by increasing the exploration coefficient in the confidence upper bound algorithm proportionally to the long-term proportion ratio, quantitative adaptive correction of the recommendation decision layer is achieved. This long-term proportion ratio essentially reflects the degree of lack of exploration of long-term value solutions or the degree of deviation of the current recommendation strategy from short-term benefit bias. The larger the ratio, the more the system focuses on short-term scores for immediate feedback while neglecting the mining of candidate content with high long-term user retention value. Increasing the exploration coefficient based on this ratio can significantly increase the system's sampling probability of uncertain long-term solutions when the proportion of long-term solutions is significantly lower than expected. This forces the recommendation strategy to break out of the local trap of short-term optimization and actively discover potential high long-term value content. The increase is proportional to the ratio, ensuring a dynamic match between the adjustment intensity and the severity of the imbalance. This prevents exploration rigidity caused by fixed coefficients and ensures that the system can accurately compensate for different degrees of underestimation of long-term value. Thus, while maintaining the short-term click-through rate baseline, the aggressiveness of exploration is dynamically adjusted according to the long-term value gap, achieving a dynamic balance between maximizing user lifetime value and immediate business goals.
[0025] Furthermore, a dual constraint mechanism for exploration quality is constructed by proportionally increasing the attempt threshold in the confidence upper bound algorithm based on the difference in exploration coefficients. While simply increasing the exploration coefficient can enhance the tendency to explore potential long-term solutions, without a corresponding entry threshold, low-confidence, low-quality candidate solutions may excessively enter the actual recommendation process due to exploration incentives, thus harming the user experience. In this case, the difference in exploration coefficients, as a quantitative indicator reflecting the extent to which the system shifts from a conservative utilization strategy to an aggressive exploration strategy, can guide the dynamic adjustment of the attempt threshold. When the exploration coefficient increases significantly, the attempt threshold is simultaneously increased significantly. The trial threshold is used to build a higher quality firewall, preventing low-quality content from easily infiltrating the final recommendation due to a surge in aggressive exploration. When the exploration coefficient is fine-tuned, the threshold can be appropriately increased. This design achieves dynamic coordination and risk hedging between exploration incentives and quality control. It expands the space for discovering potential high-value solutions by increasing the exploration coefficient, and ensures the controllability of the exploration process and the stability of recommendation quality by linking the trial threshold. It avoids an out-of-control state of "exploration for the sake of exploration", so that the system can maintain a bottom line guarantee of short-term recommendation quality while pursuing the maximization of long-term user value. Attached Figure Description
[0026] Figure 1 This is a block diagram of a controllable multimodal personalized recommendation system based on LLM in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the workflow of the controllable multimodal personalized recommendation system based on LLM in this embodiment of the invention. Figure 3 This is a flowchart in an embodiment of the present invention for determining whether the system recommendation effect is qualified within the current period based on the comprehensive utility score; Figure 4 This is a flowchart illustrating the correction of the temperature value based on the utility ratio in an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0028] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0029] Firstly, please refer to Figure 1 The diagram shown is a block diagram of a controllable multimodal personalized recommendation system based on LLM in an embodiment of the present invention. The controllable multimodal personalized recommendation system based on LLM in this embodiment includes: The data collection module is used to collect user behavior data, content data, and contextual information in real time. A preprocessing module, connected to the acquisition module, is used to preprocess the acquired data and to construct several user behavior sequences based on user ID and time order. The initial selection module, which is connected to the preprocessing module, is used to obtain user requests from the user behavior data and generate semantic vectors, and to determine the initial recommendation set based on the semantic vectors. A fusion module, which is connected to the preprocessing module and the initial understanding selection module, is used to map the semantic vector and the user behavior sequence into a user interaction vector based on a pre-trained fusion model. The quick screening module, which is connected to the understanding and initial selection module and the preprocessing module, is used to construct quick screening vectors based on user historical features, and to output a quick screening set based on the pre-trained quick screening model to initially screen the initial recommendation set. The rating and recommendation module is connected to the quick screening module and the fusion module. It is used to construct several rating vectors based on the quick screening set and the user interaction vectors, determine short-term ratings and long-term ratings based on the rating vectors, and comprehensively evaluate and output a new recommended solution based on the short-term ratings and long-term ratings. An optimization module, connected to the acquisition module, is used to determine a comprehensive utility score based on the user retention rate ratio, click-through rate ratio, preset click-through rate, and short-term loss penalty parameters within the current period; to determine whether the recommendation effect is qualified based on the comprehensive utility score; and to output instructions to correct the LLM parameters, quick screening parameters, and new scheme determination parameters when the effect is unqualified. The scheduling module, which is connected to the understanding and initial selection module, the rapid screening module, the recommendation module, and the optimization module, is used to correct the corresponding parameters based on the instructions from the optimization module.
[0030] The user behavior data refers to the behavioral records generated during the interaction between the user and the recommendation system. User behavior data includes, but is not limited to, clicks, likes, favorites, purchases, shares, comments, browsing time, dwell time, scrolling depth, bounce rate, revisit frequency, etc., which will not be elaborated here. The content data refers to the content entities that can be recommended in the recommendation system and their related descriptive information. The content data includes, but is not limited to, the original files or feature vectors, categories, prices, publication times, authors, brands, resolutions, etc. of product titles, video descriptions, news summaries, tags, keywords, text descriptions, images, videos, and audio content. This will not be elaborated further. The context information refers to environmental variables and state information related to the user's current interaction scenario. Context information includes, but is not limited to, current time, date, holidays, season, user's geographical location (GPS, IP location), city, business district, interaction scenario, device type (mobile phone / tablet / PC), operating system, screen resolution, network status, current page / interface, entry source (search / recommendation / advertisement), session status, whether the user is a new user, whether the user is logged in, current activity level, etc., which will not be elaborated further. The methods for collecting user behavior data, content data, and context information are not limited. Technical personnel can obtain them based on server logs, client-side tracking, content management system (CMS) synchronization, system APIs, etc. Among them, the methods for collecting interaction scenarios are not limited. For example, technical personnel can use LLM analysis to generate a "scenario description" based on the user's search terms in the current session, the title / description of the clicked item, and historical interests. The preprocessing module includes methods for preprocessing the collected data, including but not limited to deduplication, normalization, and feature extraction, to provide high-quality input signals for subsequent decision-making. These methods will not be elaborated further.
[0031] The user behavior sequence refers to an ordered set of a series of interactive behaviors (such as clicking, browsing, purchasing, etc.) completed by the same user in a continuous time period, arranged in chronological order.
[0032] The specific functions of the initial selection module include: obtaining user requests from the user behavior data, generating several semantic expressions based on LLM parsing of user requests, converting each semantic expression into a semantic vector, and obtaining an initial recommendation set by retrieving content from the content database based on the semantic vector; When the initial selection module generates several semantic expressions based on LLM parsing of user requests, it is necessary to set LLM parameters to ensure that the generated semantic expressions are accurate, diverse, and stable, including: The temperature value is used to control the randomness of the output value. To ensure the stability and diversity of semantic expression in this invention, the temperature value TT is set to a range of 0.1-0.5. The kernel sampling value (KSV) is used to control the cumulative probability range of candidate words during generation. Here, the range of the kernel sampling value is set to 0.8-0.95 to ensure the fluency and relevance of the generated text. The number of generated expressions is used to generate multiple independent semantic expressions. There is no limit to the value of the number of generated expressions, and technical personnel can set it according to their needs. The output length controls the maximum length of the generated semantic expression. In principle, the range of the output length is not limited. For example, it can be set between 10 and 50 word units to ensure the conciseness of the semantic expression. When the initial selection module for understanding converts each semantic expression into a semantic vector, the conversion method is not limited. The content database refers to a database pre-built through offline processing, which stores the multimodal feature representation vectors and associated metadata of all content to be recommended on the platform. This database is used by the understanding and initial selection module to perform real-time vector retrieval to obtain an initial recommendation set that matches the user's semantic vector. The method of building the content database is not limited. When the fusion module maps the semantic vector and the user behavior sequence into a user interaction vector based on the pre-trained fusion model, it needs to map the semantic vector and the user behavior sequence to the same semantic space output based on the pre-trained fusion model to obtain the user interaction vector. The user historical features in the fast screening module include user ID, location, user historical click list and user long-term interest tags. The fast screening module uses fast screening vectors to perform initial screening on the initial recommendation set based on the pre-trained fast screening model and outputs a fast screening set. The user history click list in the rapid screening module refers to a sequence of unique identifiers of all content items actively interacted with by the target user (including but not limited to clicking, playing, and purchasing) within a preset time window, arranged in chronological order. This list dynamically reflects the user's recent explicit interests and behavioral patterns. The preset time window is not limited here, and the method for constructing the user history click list includes, but is not limited to, directly generating it through real-time or near real-time processing of user behavior logs obtained by the acquisition module. In the rapid screening module, the user long-term interest tag refers to a semantic classification identifier that represents the stable preference tendency of the target user, extracted through machine learning models or statistical analysis based on the full amount of behavioral data of the target user spanning a long historical period (such as several months). The method of constructing the user long-term interest tag is not limited. Technical personnel can obtain the user's behavior sequence and related content metadata across historical periods through the preprocessing module, and then use topic models (such as LDA), clustering algorithms, or pre-trained classification models to semantically classify the content that the user has interacted with, and convert the mined interest categories into readable tags, and finally store them in the user profile database. This will not be elaborated further. The specific functions of the rating and recommendation module include: constructing several rating vectors based on the quick screening set and the user interaction vectors; scoring the rating vectors in the short term based on the fine ranking model; scoring the rating vectors in the long term based on LLM; and using the confidence upper bound algorithm to comprehensively evaluate the short-term and long-term scores and output a new recommended solution. The process by which the rating recommendation module constructs several rating vectors based on the quick screening set and the user interaction vector includes: for each candidate content in the quick screening set, obtaining its corresponding content feature representation; concatenating the user interaction vector with the content feature representation of the candidate content in a predetermined dimension order to form a comprehensive feature vector, which is the rating vector for the candidate content; The fine-ranking model is used to accurately predict the probability of a user making a short-term interaction (such as clicking) with each candidate content and outputs it as a standardized "short-term score". The fine-ranking model is existing technology and will not be described in detail here. The rating and recommendation module, based on LLM (Local Mode Model) for long-term rating of rating vectors, first decodes the features of each rating vector. Then, it inputs the decoded features into an LLM model with provided prompts for inference. Finally, the LLM model outputs the rating for each decoded feature. The prompts for the LLM model are not limited in principle; for example: You are a senior recommendation system strategy analyst, skilled at evaluating the positive impact of recommended content on the long-term value of users. Based on the provided user profiles, product information, and matching analysis, please assign a long-term value score. [Assessment Context] 1. User Profile: Long-term interest tags: [Enter the user's long-term interest tags here, such as "outdoor sports", "digital technology"] Recently Focused On: [Enter the categories the user has recently searched or clicked on here; optional] 2. Candidate Content: Title: [Product / Content Title] Key attributes: [Core selling points or attributes, such as "lightweight" or "4K resolution"] 3. Matching Analysis: Deep Match Summary: [Enter a description of the matching relationship generated by decoding the "rating vector" here, such as "The key attributes of this product are highly consistent with the user's long-term interests and can expand their expertise in that area of interest."] [Assessment Task] Please evaluate this recommendation solely from the perspective of "long-term user value," specifically considering the following three dimensions: 1. Interest Deepening: Can this recommendation help users delve deeper into or expand their long-term areas of interest? 2. Trust Building: Will this recommendation make users feel that the platform understands them better, thereby increasing their long-term stickiness to the platform? 3. Retention Contribution: Does this increase the likelihood that users will continue to use the platform due to recommendations? Output format requirements You must strictly adhere to the following format when outputting, and must not include any additional explanations, apologies, or thought processes: Long-term rating: <An integer between 1 and 10, with 10 being the highest score> Core Reason: <In short, summarize the most important reason> When using the confidence upper bound algorithm to comprehensively evaluate short-term and long-term scores and output a new recommended solution, the UCB algorithm is first initialized before the system goes live. This involves setting the number of arms and initial weights to define different combinations of score weights as different arms. The number of arms and initial weights are not limited in principle. For example, the number of arms can be defined as 3, where arm 1 has a short-term weight of 0.9 and a long-term weight of 0.1, arm 2 has a short-term weight of 0.5 and a long-term weight of 0.5, and arm 3 has a short-term weight of 0.1 and a long-term weight of 0.9. Then, all initial values are initialized. The values of the initial values are not limited in principle. For example, they can be set to optimistic initial values (e.g., 5) to encourage early exploration. The initial number of attempts is set to 0, and the exploration coefficient EI is set to a value of 1-3 to balance the effective search space for exploration and utilization. Set the trial threshold ALA to 100-500 to force it to have an extremely high "exploration bonus", ensuring that it can get enough exposure opportunities, get through the cold start phase with sparse data, get a fair evaluation opportunity, and avoid good strategies being buried due to bad luck in the early stages.
[0033] Please see Figure 2 The diagram shown illustrates the workflow of a controllable multimodal personalized recommendation system based on LLM in an embodiment of the present invention. The workflow of the controllable multimodal personalized recommendation method based on LLM in this embodiment includes: S1: The data collection module collects user behavior data, content data, and contextual information in real time; S2: The preprocessing module preprocesses the collected data and, based on user ID and time sequence, constructs several user behavior sequences from the preprocessed data; S3: Understand the user requests obtained from the user behavior data by the initial selection module, and generate several semantic expressions based on LLM parsing of user requests, and convert each semantic expression into a semantic vector, and retrieve the content database based on the semantic vector to obtain the initial recommendation set; S4: The fusion module maps the semantic vector and the user behavior sequence to the same semantic space based on the pre-trained fusion model and outputs a user interaction vector. S5: The quick screening module constructs a quick screening vector based on the user ID, location, user historical click list and user long-term interest tags, and uses the quick screening vector to perform initial screening on the initial recommendation set based on the pre-trained quick screening model, outputting a quick screening set. S6: The rating and recommendation module constructs several rating vectors based on the quick screening set and the user interaction vectors, and scores the rating vectors in the short term based on the fine ranking model, scores the rating vectors in the long term based on LLM, and uses the confidence upper bound algorithm to comprehensively evaluate the short-term and long-term scores and output a new recommended solution. S7: The explanation module generates natural language recommendation reasons for the new scheme based on LLM; S8: The optimization module determines the comprehensive utility score based on the user retention rate ratio, click-through rate ratio, preset click-through rate and short-term loss penalty parameters within the current period, and judges whether the recommendation effect is qualified based on the comprehensive utility score, and outputs instructions to correct LLM parameters, quick screening parameters and new solution determination parameters when the effect is not qualified. S9: The scheduling module corrects the corresponding parameters based on the instructions from the optimization module.
[0034] Specifically, the controllable multimodal personalized recommendation system based on LLM also includes an explanation module, which is connected to the rating recommendation module, and is used to generate natural language recommendation reasons for the new solution based on LLM; The fusion module is also used to construct a fusion model based on a unified multimodal representation network, and to train the fusion model based on the relationship between user interaction vectors in historical traffic data and the semantic vectors and the user behavior sequences. The rapid screening module is also used to construct a rapid screening model based on a lightweight ranking network, and to train the rapid screening model based on the relationship between the historical initial screening vector - initial recommendation set vector and the rapid screening set.
[0035] The unified multimodal representation network includes, but is not limited to, any one of CLIP, ALBEF, BLIP, Florence, and ONE-PEACE, which will not be elaborated here. The historical traffic data refers to the collection of all online service logs accumulated during one or more statistical periods of the system's past operation, including but not limited to the following data with interrelationships: user interaction vectors, semantic vectors, user behavior sequences, etc., which will not be elaborated here. The lightweight ranking network includes, but is not limited to, any one of logistic regression, shallow multilayer perceptron, factorization machine, and width learning system, which will not be elaborated here. Specifically, the process of building a rapid screening model based on a shallow multilayer perceptron includes: Input layer: The input dimensions are not limited in principle, and technical personnel can set them according to the actual situation. For example, user ID embedding (64 dimensions) + user location embedding (16 dimensions) + 10 categories of user's historical clicks embedding (10*8=80 dimensions) + 5 long-term interest tags of user embedding (5*8=40 dimensions) + basic product features (56 dimensions) = 256 dimensions. Therefore, the dimensions are set to 256. Hidden layer First hidden layer The dimension is set to 128, compressing the original 256-dimensional features to 128 dimensions, and the activation function is set to ReLU; Second hidden layer By setting the dimension to 64, the 128-dimensional features are further condensed to extract a more advanced abstract pattern that is strongly correlated with clicks, and the activation function is set to ReLU. Output layer Set the dimension to 1 and the activation function to Sigmoid; The fast screening model is trained based on the relationship between the historical initial screening vector - initial selection recommendation set vector and the fast screening set. During training, the optimizer is set to Adam, the initial learning rate is 0.001, and 512 samples are input into the model for training. After training, the model is exported, and the Top K value is set. The range of the Top K value is not limited in principle. Technicians can set the Top K value according to their needs, for example, from 100 to 300, to achieve a balance between efficiency and accuracy.
[0036] Specifically, the optimization module is also used to determine the comprehensive utility score based on the user retention rate ratio, the click-through rate ratio, the preset click-through rate, and the short-term loss penalty parameter within the current period, and to determine whether the recommendation effect within the current period is qualified based on the comprehensive utility score; The optimization module is also used to maintain and continuously monitor system parameters when the system is deemed qualified, or to correct the temperature value of the LLM in the initial selection module based on the ratio of the preset comprehensive utility score to the comprehensive utility score when the system is deemed unqualified.
[0037] Among them, the comprehensive utility component The formula is:
[0038] Among them, the user retention rate ratio Click-through rate ratio , This refers to the user retention rate during the current period when the new solution is adopted. This refers to the user retention rate of the historical baseline plan within the same period. The specific value is not limited; technical personnel can set it according to the specific circumstances of the platform. The preset click-through rate (CTR) refers to the minimum CTR that can be expected when using the new strategy within the current period. The specific value is not limited in principle; technical personnel can set it according to actual requirements. This refers to the short-term loss penalty coefficient. The specific value of the short-term loss penalty coefficient is not limited; technical personnel can set it according to the specific circumstances of the platform. Click-through rate when using the new solution during this period. This refers to the click-through rate of the historical baseline plan within the same period. The specific value is not limited; technical personnel can set it according to the specific circumstances of the platform. The internal strategies, external user preferences, and content ecosystem of a recommendation system are all dynamically changing. Regularly calculating the comprehensive utility score allows for continuous monitoring of the system's overall performance, enabling timely detection of performance degradation caused by strategy drift, changes in data distribution, or improper algorithm exploration. It also provides an objective basis for adjusting system parameters. The comprehensive utility score is a composite indicator used to quantify the overall effectiveness of a new recommendation scheme. Its calculation relies on two core inputs: the user retention rate ratio, representing long-term user value, and the click-through rate ratio, representing short-term interaction experience. The physical meaning of this score is: under strict constraints on the short-term experience decline, the risk-adjusted net long-term value gain that the new scheme can bring. Specifically, when the short-term... When the decline in user experience does not exceed the preset click-through rate, the overall utility score directly equals the long-term value improvement. However, when the short-term decline in user experience exceeds the preset click-through rate, a penalty term will be deducted from the long-term gains. This penalty term is quadratically related to the magnitude of the deviation from the threshold, thus achieving non-linear sensitive control over different degrees of short-term losses. The overall utility score unifies long-term goals and short-term constraints, providing a single and clear decision criterion for conflicting long-term and short-term objectives. At the same time, the preset click-through rate and penalty coefficient embed a business risk tolerance, preventing the optimization process from harming the core user experience. Finally, the existence of overall utility can quickly evaluate the new solution recommendation method within the current cycle, thus providing a basis for the system's adaptive adjustment.
[0039] Please see Figure 3 As shown, this is a flowchart illustrating the process of determining whether the system recommendation effect within the current period is satisfactory based on the comprehensive utility score in an embodiment of the present invention. The process of determining whether the system recommendation effect within the current period is satisfactory based on the comprehensive utility score in an embodiment of the present invention includes: The optimization module is based on the user retention rate ratio (URRR), the click-through rate ratio (CTR), and the preset click-through rate within the current period. The short-term loss penalty parameter determines the overall utility score. The formula is:
[0040] The comprehensive utility segment With setting preset comprehensive utility score In the comparison, when the overall utility score S is below 1, it indicates that the effect of adopting the new recommended solution is not as good as the baseline. However, in a mature recommendation system, if the overall utility score can be raised to 1.2 (equivalent to a 20% increase in user retention rate) in the short term, considering both short-term losses and long-term retention, it indicates that it has achieved good results. Therefore, a preset overall utility score is set.
[0041] If the comprehensive utility score If the value is less than or equal to the preset comprehensive utility score S1, the recommendation effect in this period is deemed unqualified, and the temperature value of LLM in the understanding initial selection module is corrected based on the ratio of the preset comprehensive utility score S1 to the comprehensive utility score S. If the comprehensive utility score If the overall utility score is greater than the preset comprehensive utility score S1, the recommended effect is deemed satisfactory for this period, and the system parameters are maintained and continuously monitored.
[0042] Specifically, the optimization module is also used to determine a utility ratio based on the ratio of a preset comprehensive utility score to a comprehensive utility score, and to reduce the temperature value based on the utility ratio, wherein the reduction in temperature value is proportional to the utility ratio.
[0043] The downstream fusion, rapid screening, and scoring stages of the recommendation chain all rely on understanding the semantic vectors generated by the initial selection module as input. If the overall utility score does not meet expectations, the primary suspect is that the LLM's semantic parsing of the user request has become overly divergent or semantically drifted, causing the initial recall candidate set to deviate from the user's true intent. Therefore, it is necessary to constrain the generation behavior of the LLM from the source. The utility ratio mentioned here refers to the ratio between the preset overall utility score and the actual overall utility score. Its value directly reflects the degree of deviation or performance gap between the current recommendation performance and the expected target. The larger the ratio, the worse the actual recommendation effect and the further the system deviates from the optimal state. Based on this utility ratio, the temperature value is reduced, and the reduction magnitude is proportional to the utility ratio. This helps to achieve dynamic matching between the parameter correction intensity and the severity of the problem. It can both make aggressive corrections for severely degraded scenarios to quickly restore system performance and maintain gentle adjustments during slight fluctuations to preserve the richness of semantic expression. At the same time, it avoids the blindness of manual experience-based parameter tuning, enabling the system to have quantitatively driven self-optimization capabilities and faster convergence speed.
[0044] Please see Figure 4 The diagram shows a flowchart illustrating the process of correcting the temperature value based on the utility ratio in an embodiment of the present invention. The process of correcting the temperature value based on the utility ratio in this embodiment includes: The optimization module determines the utility ratio SR based on the preset ratio of the comprehensive utility score to the total utility score. The utility ratio SR is compared with the set first preset utility ratio SR1 and second preset utility ratio SR2. When the comprehensive utility score is less than 1 and greater than or equal to 0.95, it indicates that the currently recommended new solution may be slightly lower than the historical baseline. When it reaches 0.85, it is considered that there may be a significant deviation. Therefore, the first preset utility ratio SR1 is set to (1, 1.25] and the second preset utility ratio SR2 is set to (1.25, 1.4]. If the utility ratio SR is less than or equal to the first preset utility ratio SR1, then the temperature value is corrected using the first creative threshold β1, and the corrected temperature value TT' = TT × β1, wherein the first creative threshold β1 is set to 0.95; If the utility ratio SR is greater than the first preset utility ratio SR1 and less than or equal to the second preset utility ratio SR2, then the temperature value is corrected using the second creative threshold β2. The corrected temperature value TT' = TT × β2, where the second creative threshold β2 is set to 0.89. If the utility ratio SR is greater than the second preset utility ratio SR2, then the temperature value is corrected using the third creative threshold β3. The corrected temperature value TT' = TT × β3, where the third creative threshold β3 is set to 0.8.
[0045] Specifically, the optimization module is also used to determine the temperature difference based on the difference between the temperature values before and after the correction, and to increase the kernel sampling value of the LLM in the initial selection module based on the temperature difference, wherein the increase in the kernel sampling value is proportional to the temperature difference.
[0046] While lowering the temperature value can suppress the randomness of LLM generation and improve the accuracy of semantic parsing, it also leads to a sharper output probability distribution. This can cause the model to over-concentrate on high-probability candidate words, resulting in pattern collapse or semantic homogenization. Therefore, it is necessary to broaden the candidate word selection space by increasing the kernel sampling value to compensate for the loss of diversity under low-temperature settings. The temperature difference reflects the degree to which the system strengthens the constraints on the randomness of LLM or the adjustment range from the original state to the deterministic state. A larger difference indicates a more severe compression of the semantic generation space and a higher risk of diversity loss. Based on this, the kernel sampling value is increased proportionally to the temperature difference. The underlying mechanism is... A dynamic complementary balance strategy is constructed: when the temperature is significantly reduced, the kernel sampling value is correspondingly significantly increased to ensure sufficient candidate diversity is maintained on the basis of forced determinism; when the temperature is finely adjusted, the kernel sampling value is also moderately increased to avoid overcorrection. This design realizes adaptive coupling and bidirectional balance between LLM generation parameters. It can quickly correct recommendation bias and improve semantic accuracy by rapidly reducing the temperature, and prevent homogenization of candidate semantics caused by excessive conservatism by dynamically compensating for the kernel sampling value. Thus, it maintains a dynamic balance between accuracy and diversity during the system optimization process, and enhances the robustness and convergence stability of the recommendation system in the face of performance fluctuations.
[0047] Specifically, the process by which the optimization module increases the kernel sampling value of the LLM in the initial selection module based on the temperature difference includes: The optimization module determines the temperature difference TTD based on the difference between the temperature values before and after the correction; The temperature difference TTD is compared with the first preset temperature difference TTD1 and the second preset temperature difference TTD2. For the temperature value, a decrease of about 0.05 may have a significant impact on the comprehensive utility score. Therefore, the first preset temperature difference TTD1 is set to [0.005, 0.05] and the second preset temperature difference TTD2 is set to (0.05, 0.1]. If the temperature difference TTD is less than or equal to the first preset temperature difference TTD1, then the kernel sample value KSV is corrected using the first kernel sampling correction threshold μ1. The corrected kernel sample value KSV' = KSV × μ1, where the first kernel sampling correction threshold μ1 is set to 1.02. If the temperature difference TTD is greater than the first preset temperature difference TTD1 and less than or equal to the second preset temperature difference TTD2, then the kernel sample value KSV is corrected using the second kernel sampling correction threshold μ2. The corrected kernel sample value KSV' = KSV × μ2, where the second kernel sampling correction threshold μ2 is set to 1.06. If the temperature difference TTD is greater than the second preset temperature difference TTD2, then the third kernel sampling correction threshold μ3 is used to correct the kernel sampling value KSV. The corrected kernel sampling value KSV' = KSV × μ3, where the third kernel sampling correction threshold μ3 is set to 1.12. Specifically, the optimization module is also used to determine the kernel sampling difference based on the difference between the kernel sampling values before and after the correction, and to increase the Top K value of the rapid screening model based on the kernel sampling difference, wherein the increase in the Top K value is proportional to the kernel sampling difference.
[0048] The decrease in temperature value improves the accuracy of semantic parsing, while the increase in kernel sampling value compensates for the loss of diversity under low temperature settings, significantly widening the semantic distribution space of the initial recommendation set. At this point, the rapid screening module, as a downstream filtering stage, may suffer from excessive truncation of high-quality content in the initial screening stage if it maintains the original Top K value due to the expansion of the candidate pool. Therefore, the rapid screening threshold must be dynamically adjusted to adapt to upstream changes. Here, the kernel sampling difference directly reflects the adjustment magnitude of the LLM generation strategy towards diversity. A larger difference indicates a higher degree of semantic divergence and a wider candidate coverage in the upstream initial recommendation set. When the kernel sampling difference is large, it means that the diversity of the initial set has been greatly improved. In this case, the Top K value needs to be significantly increased accordingly to retain more candidates for the downstream fine-tuning stage, ensuring that the diversity gain is not negated by the rapid screening stage. When the kernel sampling difference is small, only the Top K value needs to be moderately relaxed. The K value avoids the waste of computational resources caused by the expansion of the candidate pool. This design realizes the parameter coordination and dynamic adaptation of semantic understanding and candidate selection in the recommendation link. By using the kernel sampling difference as a quantitative indicator, the changes in the generation characteristics of the upstream LLM are accurately transmitted to the fast screening layer. This not only prevents the loss of high-quality candidates due to the mismatch between upstream and downstream parameters, but also ensures the system's fine-grained response capability when facing different adjustment ranges. Thus, while maintaining computational efficiency, it maximizes the balance between the overall coverage accuracy and diversity of the recommendation system.
[0049] Specifically, the process by which the optimization module increases the Top K value of the rapid screening model based on the kernel sampling difference includes: The optimization module determines the kernel sampling difference (KSVD) based on the difference between the kernel sampling values before and after correction. The kernel sampling difference KSVD is compared with the first preset kernel sampling difference KSVD1 and the second preset kernel sampling difference KSVD2. The specific values of the first preset kernel sampling difference KSVD1 and the second preset kernel sampling difference KSVD2 are not limited in principle. Technicians can set them based on historical experience with experimental data. For example, here the first preset kernel sampling difference KSVD1 is set to [0.01, 0.05] and the second preset kernel sampling difference KSVD2 is set to (0.05, 0.1] based on experimental data. If the kernel sampling difference KSVD is less than or equal to the first preset kernel sampling difference KSVD1, then the Top K value TK is corrected using the first TOPK correction threshold λ1, and the corrected Top K value TK' = λ1 × TK, where the first TOPK correction threshold λ1 is set to 1.03. If the kernel sampling difference KSVD is greater than the first preset kernel sampling difference KSVD1 and less than or equal to the second preset kernel sampling difference KSVD2, then the Top K value TK is corrected using the second TOPK correction threshold λ2. The corrected Top K value TK' = λ2 × TK, where the second TOPK correction threshold λ2 is set to 1.07. If the kernel sampling difference KSVD is greater than the second preset kernel sampling difference KSVD2, then the Top K value TK is corrected using the third TOPK correction threshold λ3. The corrected Top K value TK' = λ3 × TK, where the third TOPK correction threshold λ3 is set to 1.14.
[0050] Specifically, the optimization module is also used to determine whether the recommendation effect is qualified based on the comprehensive utility score in the first period after correction, and to maintain the system parameters and continue monitoring when the result is qualified, or to correct the exploration coefficient in the confidence upper bound algorithm based on the ratio of the preset long-term solution ratio to the long-term solution ratio in the first period after correction when the result is unqualified. The long-term solution refers to the solution whose long-term score is greater than the short-term score, and the proportion of long-term solutions refers to the proportion of long-term solutions among all new solutions output in this period.
[0051] Previous parameter adjustments only targeted the upstream initial screening and rapid filtering stages. If the recommendation effect still falls short after the adjustments, it indicates that the root cause is not a defect in candidate set generation or initial filtering, but rather a deviation in the downstream scoring and recommendation module's decision-making balance between short-term gains and long-term value. Adjusting the exploration coefficient in the confidence upper bound algorithm directly controls the system's exploration intensity for uncertain long-term solutions. When the actual proportion of long-term solutions is lower than the preset target, it reflects that the system is overly inclined towards short-term scoring based on immediate feedback, thus inhibiting the discovery of high-long-term value candidates. The exploration coefficient is adjusted based on the ratio of the preset proportion of long-term solutions to the actual proportion of long-term solutions. In scenarios where recommendation results fall short of expectations and upstream corrections are ineffective, this design can accurately locate and correct the exploration-utilization imbalance in the final decision-making layer, preventing continuous degradation due to ineffective parameter adjustments in the early stages. This design establishes a layered diagnosis and repair mechanism across the entire chain, from upstream semantic understanding and midstream candidate selection to downstream scoring decisions. It not only achieves precise, step-by-step location of fault points, preventing blind adjustments to irrelevant modules, but also ensures that the system balances the dual optimization goals of short-term click-through rate and long-term user retention during the repair process by dynamically coupling the proportion of long-term solutions with the exploration coefficient, thereby improving the sustainability of the recommendation strategy and the user lifetime value.
[0052] Specifically, the process by which the optimization module determines whether the recommendation effect is satisfactory based on the comprehensive utility score within the corrected first period includes: The optimization module obtains the comprehensive utility score S' in the first corrected period; If the corrected comprehensive utility score S' is less than or equal to the preset comprehensive utility score S1, the recommendation effect in this period is deemed unqualified, and the exploration coefficient in the confidence upper bound algorithm is corrected based on the ratio of the preset long-term solution ratio to the corrected long-term solution ratio in the first period. If the comprehensive utility score S' is greater than the preset comprehensive utility score S1, the recommendation effect is deemed satisfactory in this period, and the system parameters are maintained and continuously monitored.
[0053] Specifically, the optimization module is also used to determine the long-term ratio based on the ratio of the preset long-term solution ratio to the ratio of the long-term solution ratio in the first period after correction, and to increase the exploration coefficient in the confidence upper bound algorithm based on the long-term ratio, wherein the increase in the exploration coefficient is proportional to the long-term ratio.
[0054] The long-term effectiveness ratio refers to the ratio of the preset long-term effectiveness ratio to the actual long-term effectiveness ratio in the first period after correction. Essentially, it reflects the degree to which the current recommendation strategy lacks exploration of long-term value solutions or deviates from short-term gain bias. A larger ratio indicates that the actual proportion of long-term effectiveness solutions in the recommendation results is lower than the expected target, and the system is overly focused on short-term scores for immediate feedback while neglecting to mine high-value candidate content for long-term user retention. Therefore, the exploration coefficient in the confidence upper bound algorithm should be increased proportionally to the long-term effectiveness ratio. When the ratio increases, it means the system's exploration gap for long-term value is widening. At this point, the exploration coefficient needs to be significantly increased to enhance the sampling probability of uncertain long-term solutions in the UCB algorithm, forcing the recommendation strategy to escape the local trap of short-term optimism and actively discover potential high-long-term value content. Furthermore, the increase should be proportional to the ratio to ensure the adjustment strength and balance of imbalances. Dynamic matching avoids excessive or insufficient corrections. This design achieves quantitative adaptive correction of the recommendation decision layer. By precisely coupling the long-term proportion ratio, a business indicator, with the algorithm's exploration parameters, the system maintains the short-term click-through rate baseline while dynamically adjusting the aggressiveness of exploration based on the actual size of the long-term value gap. This prevents exploration rigidity caused by fixed coefficients and ensures that the system can accurately compensate for different degrees of long-term value underestimation, thereby achieving a dynamic balance between maximizing user lifetime value and immediate business goals.
[0055] Specifically, the optimization module determines the long-term ratio (LTR) based on the ratio of the preset long-term solution ratio to the modified long-term solution ratio in the first cycle. Since 60-70% of users typically only seek immediate gratification, while only 30-40% of users are willing to accept "delayed gratification," the preset long-term solution ratio is set to a range of 40%-60% to avoid overwhelming users with too much long-term content. The long-term ratio LTR is compared with the first preset long-term ratio LTR1 and the second preset long-term ratio LTR2. When the proportion of the long-term solution in the first cycle after correction is only 35%, it may be biased towards short-term orientation. When it reaches 30%, it means that it may be seriously biased towards short-term orientation. Therefore, the first preset long-term ratio LTR1 is set to (1, 1.4] and the second preset long-term ratio LTR2 is set to (1.4, 1.8]. If the long-acting ratio LTR is less than or equal to the first preset long-acting ratio LTR1, then the exploration coefficient EI is corrected using the first exploration coefficient correction threshold θ1, and the corrected exploration coefficient EI' = EI × θ1, where the first exploration coefficient correction threshold θ1 is set to 1.02. If the long-acting ratio LTR is greater than the first preset long-acting ratio LTR1 and less than or equal to the second preset long-acting ratio LTR2, then the exploration coefficient EI is corrected using the second exploration coefficient correction threshold θ2. The corrected exploration coefficient EI' = EI × θ2, where the second exploration coefficient correction threshold θ2 is set to 1.05. If the long-term ratio LTR is greater than the second preset long-term ratio LTR2, then the exploration coefficient EI is corrected using the third exploration coefficient correction threshold θ3. The corrected exploration coefficient EI' = EI × θ3, where the third exploration coefficient correction threshold θ3 is set to 1.1.
[0056] Specifically, the optimization module is also used to determine the difference in exploration coefficients based on the difference in exploration coefficients before and after the correction, and to increase the attempt threshold in the confidence upper bound algorithm based on the difference in exploration coefficients, wherein the increase in the attempt threshold is proportional to the difference in exploration coefficients.
[0057] Simply increasing the exploration coefficient can enhance the system's tendency to explore potential long-term solutions. However, without corresponding entry thresholds, low-confidence, low-quality candidate solutions may excessively enter the actual recommendation process due to exploration incentives, thereby harming the user experience. In this case, raising the trial threshold can impose a reverse constraint on exploration behavior, ensuring that only candidate solutions with sufficient potential value are given the opportunity to be explored, avoiding the degradation of recommendation quality caused by blind exploration. Here, the difference in exploration coefficients directly reflects the magnitude or intensity of the system's shift from a conservative utilization strategy to an aggressive exploration strategy. A larger difference indicates a more significant increase in the system's openness to uncertain long-term solutions. At this point, potential exploration risks and quality control are crucial. The higher the demand, the more the exploration coefficient increases. When the exploration coefficient increases significantly, the attempt threshold must be increased simultaneously to build a higher quality firewall and prevent low-quality content from easily infiltrating the final recommendation due to the surge in exploration aggressiveness. When the exploration coefficient is finely adjusted, the threshold can be increased appropriately. This design achieves dynamic coordination and risk hedging between exploration incentives and quality control. It expands the space for discovering potential high-value solutions by increasing the exploration coefficient, and ensures the controllability of the exploration process and the stability of recommendation quality by linking the increase of the attempt threshold. It avoids the out-of-control state of "exploring for the sake of exploration", so that the system can maintain a bottom line guarantee of short-term recommendation quality while pursuing the maximization of long-term user value.
[0058] Specifically, the process by which the optimization module increases the attempt threshold in the confidence upper bound algorithm based on the difference in exploration coefficients includes: The optimization module determines the exploration coefficient difference EID based on the difference between the exploration coefficients before and after the correction; The exploration coefficient difference EID is compared with the first preset exploration coefficient difference EID and the second preset exploration coefficient difference EID2. Based on the experimental results, the first preset exploration coefficient difference EID is set to [0.02, 0.1] and the second preset exploration coefficient difference EID2 is set to (0.1, 0.2) to obtain a higher adjustment effect. If the exploration coefficient difference EID is less than or equal to the first preset exploration coefficient difference EID1, then the first attempt correction threshold З1 is used to correct the attempt threshold ALA, and the corrected attempt threshold ALA' = ALA × З1, wherein the first attempt correction threshold З1 is set to 1.01; If the exploration coefficient difference EID is greater than the first preset exploration coefficient difference EID1 and less than or equal to the second preset exploration coefficient difference EID2, then the second attempt correction threshold З2 is used to correct the attempt threshold ALA. The corrected attempt threshold ALA' = ALA × З2, where the second attempt correction threshold З2 is set to 1.03. If the exploration coefficient difference EID is greater than the second preset exploration coefficient difference EID2, then the third attempt correction threshold З3 is used to correct the attempt threshold ALA. The corrected attempt threshold ALA' = ALA × З3, where the third attempt correction threshold З3 is set to 1.08.
[0059] Secondly, this invention also provides a controllable multimodal personalized recommendation method based on LLM, including: Real-time collection of user behavior data, content data, and contextual information; The preprocessed data is used to construct several user behavior sequences based on user ID and time order. The user requests in the user behavior data are obtained, and several semantic expressions are generated by parsing the user requests based on LLM, and each semantic expression is converted into a semantic vector. The initial recommendation set is obtained by searching the content database based on the semantic vector. The pre-trained fusion model maps the semantic vector and the user behavior sequence to the same semantic space and outputs a user interaction vector. Based on the user ID, location, user historical click list, and user long-term interest tags, a quick screening vector is constructed, and a quick screening set is output by using the quick screening vector to perform initial screening on the initial recommendation set based on the pre-trained quick screening model. Based on the rapid screening set and the user interaction vector, several scoring vectors are constructed, and the scoring vectors are scored in the short term based on the fine ranking model, scored in the long term based on LLM, and a new recommended solution is output by comprehensively evaluating the short-term and long-term scores using the confidence upper bound algorithm. The comprehensive utility score is determined based on the user retention rate ratio, click-through rate ratio, and short-term loss penalty parameters of the new solution relative to the old baseline within this period. The comprehensive utility score is used to determine whether the recommendation effect is satisfactory. If it is unsatisfactory, the LLM parameters, quick screening parameters, and new solution determination parameters are adjusted.
[0060] In summary, this invention provides a controllable multimodal personalized recommendation system and method based on LLM, aiming to solve the technical problems of existing recommendation systems, such as the cold start dilemma, insufficient semantic understanding ability, lack of long-term value optimization, and insufficient utilization of multimodal information. The system achieves closed-loop control of the entire process from user request parsing to final recommendation output by constructing a complete architecture including modules for collection, preprocessing, understanding and initial selection, fusion, rapid screening, rating and recommendation, optimization, and scheduling. First, LLM is used to deeply parse user requests to generate semantic vectors and perform initial selection and recall. User behavior sequences and semantic information are fused through a unified multimodal representation network. After filtering by a lightweight rapid screening model, a rating and recommendation module based on a confidence upper bound algorithm is used to comprehensively evaluate short-term and long-term value and output a new solution. At the same time, an interpretation module provides interpretability support. The core innovation of the system lies in the comprehensive utility score evaluation system constructed by the optimization module. This system, based on user retention ratio, click-through rate ratio, preset click-through rate, and short-term loss penalty parameters, achieves a unified quantification of long-term user value and short-term experience constraints. When the recommendation effect is unsatisfactory, the system initiates a cascaded adaptive correction mechanism from upstream to downstream. It sequentially adjusts the LLM temperature value based on the utility ratio to improve semantic accuracy, adjusts the kernel sampling value based on the temperature difference to compensate for diversity loss, adjusts the Top K value of the rapid screening model based on the kernel sampling difference to adapt to changes in the candidate set, and adjusts the exploration coefficient of the confidence upper bound algorithm based on the long-term proportion ratio when upstream correction is ineffective. Finally, it adjusts the attempt threshold based on the exploration coefficient difference to ensure exploration quality, thus forming a parameter optimization chain that dynamically balances accuracy, diversity, and exploration intensity. This technical solution not only significantly enhances the system's ability to understand natural language queries and complex intents, breaking through the limitations of historical behavior data dependence, but also effectively improves long-term user retention while ensuring short-term user experience through controllable adjustment and adaptive optimization across the entire chain. This achieves a comprehensive improvement in the recommendation system in terms of accuracy, real-time performance, interpretability, and commercial value.
[0061] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A controllable multimodal personalized recommendation system based on LLM, characterized in that, include, The data collection module is used to collect user behavior data, content data, and contextual information in real time. A preprocessing module, connected to the acquisition module, is used to preprocess the acquired data and to construct several user behavior sequences based on user ID and time order. The initial selection module, which is connected to the preprocessing module, is used to obtain user requests from the user behavior data and generate semantic vectors, and to determine the initial recommendation set based on the semantic vectors. A fusion module, which is connected to the preprocessing module and the initial understanding selection module, is used to map the semantic vector and the user behavior sequence into a user interaction vector based on a pre-trained fusion model. The quick screening module, which is connected to the understanding and initial selection module and the preprocessing module, is used to construct quick screening vectors based on user historical features, and to output a quick screening set based on the pre-trained quick screening model to initially screen the initial recommendation set. The rating and recommendation module is connected to the quick screening module and the fusion module. It is used to construct several rating vectors based on the quick screening set and the user interaction vectors, determine short-term ratings and long-term ratings based on the rating vectors, and comprehensively evaluate and output a new recommended solution based on the short-term ratings and long-term ratings. An optimization module, connected to the acquisition module, is used to determine a comprehensive utility score based on the user retention rate ratio, click-through rate ratio, preset click-through rate, and short-term loss penalty parameters within the current period; to determine whether the recommendation effect is qualified based on the comprehensive utility score; and to output instructions to correct the LLM parameters, quick screening parameters, and new scheme determination parameters when the effect is unqualified. The scheduling module, which is connected to the understanding and initial selection module, the rapid screening module, the recommendation module, and the optimization module, is used to correct the corresponding parameters based on the instructions from the optimization module.
2. The controllable multimodal personalized recommendation system based on LLM according to claim 1, characterized in that, The controllable multimodal personalized recommendation system based on LLM also includes an explanation module, which is connected to the rating recommendation module, and is used to generate natural language recommendation reasons for the new solution based on LLM. The fusion module is also used to construct a fusion model based on a unified multimodal representation network, and to train the fusion model based on the relationship between user interaction vectors in historical traffic data and the semantic vectors and the user behavior sequences. The rapid screening module is also used to construct a rapid screening model based on a lightweight ranking network, and to train the rapid screening model based on the relationship between the historical initial screening vector - initial recommendation set vector and the rapid screening set.
3. The controllable multimodal personalized recommendation system based on LLM according to claim 2, characterized in that, The optimization module is also used to determine the comprehensive utility score based on the user retention rate ratio, the click-through rate ratio, the preset click-through rate, and the short-term loss penalty parameter within the current period, and to determine whether the recommendation effect within the current period is qualified based on the comprehensive utility score. The optimization module is also used to maintain and continuously monitor system parameters when the system is deemed qualified, or to correct the temperature value of the LLM in the initial selection module based on the ratio of the preset comprehensive utility score to the comprehensive utility score when the system is deemed unqualified. The short-term loss penalty parameter refers to the penalty level imposed by the system on short-term losses. The user retention rate ratio refers to the ratio of the user retention rate when the new scheme is adopted in the current period to the user retention rate of the old baseline. The click-through rate ratio refers to the ratio of the click-through rate when using the new scheme during the current period to the click-through rate of the old baseline.
4. The controllable multimodal personalized recommendation system based on LLM according to claim 3, characterized in that, The optimization module is also used to determine a utility ratio based on the ratio of the preset comprehensive utility score to the comprehensive utility score, and to reduce the temperature value based on the utility ratio, wherein the reduction in temperature value is proportional to the utility ratio.
5. The controllable multimodal personalized recommendation system based on LLM according to claim 4, characterized in that, The optimization module is also used to determine the temperature difference based on the difference between the temperature values before and after the correction, and to increase the kernel sampling value of the LLM in the initial selection module based on the temperature difference, wherein the increase in the kernel sampling value is proportional to the temperature difference.
6. The controllable multimodal personalized recommendation system based on LLM according to claim 5, characterized in that, The optimization module is also used to determine the kernel sampling difference based on the difference between the kernel sampling values before and after the correction, and to increase the Top K value of the rapid screening model based on the kernel sampling difference, wherein the increase in the Top K value is proportional to the kernel sampling difference.
7. The controllable multimodal personalized recommendation system based on LLM according to claim 6, characterized in that, The optimization module is also used to determine whether the recommendation effect is qualified based on the comprehensive utility score in the first period after correction, and to maintain the system parameters and continue monitoring when the result is qualified, or to correct the exploration coefficient in the confidence upper bound algorithm based on the ratio of the preset long-term solution ratio to the long-term solution ratio in the first period after correction when the result is unqualified. The long-term solution refers to the solution whose long-term score is greater than the short-term score, and the proportion of long-term solutions refers to the proportion of long-term solutions among all new solutions output in this period.
8. The controllable multimodal personalized recommendation system based on LLM according to claim 7, characterized in that, The optimization module is also used to determine the long-term ratio based on the ratio of the preset long-term solution ratio to the ratio of the long-term solution ratio in the first period after correction, and to increase the exploration coefficient in the confidence upper bound algorithm based on the long-term ratio, wherein the increase in the exploration coefficient is proportional to the long-term ratio.
9. The controllable multimodal personalized recommendation system based on LLM according to claim 8, characterized in that, The optimization module is also used to determine the difference in exploration coefficients based on the difference in exploration coefficients before and after the correction, and to increase the attempt threshold in the confidence upper bound algorithm based on the difference in exploration coefficients, wherein the increase in the attempt threshold is proportional to the difference in exploration coefficients.
10. A controllable multimodal personalized recommendation method based on LLM as described in any one of claims 1-9, characterized in that, include, Real-time collection of user behavior data, content data, and contextual information; The preprocessed data is used to construct several user behavior sequences based on user ID and time order. The user requests in the user behavior data are obtained, and several semantic expressions are generated by parsing the user requests based on LLM, and each semantic expression is converted into a semantic vector. The initial recommendation set is obtained by searching the content database based on the semantic vector. The pre-trained fusion model maps the semantic vector and the user behavior sequence to the same semantic space and outputs a user interaction vector. Based on the user ID, location, user historical click list, and user long-term interest tags, a quick screening vector is constructed, and a quick screening set is output by using the quick screening vector to perform initial screening on the initial recommendation set based on the pre-trained quick screening model. Based on the rapid screening set and the user interaction vector, several scoring vectors are constructed, and the scoring vectors are scored in the short term based on the fine ranking model, scored in the long term based on LLM, and a new recommended solution is output by comprehensively evaluating the short-term and long-term scores using the confidence upper bound algorithm. The comprehensive utility score is determined based on the user retention rate ratio, click-through rate ratio, and short-term loss penalty parameters of the new solution relative to the old baseline within this period. The comprehensive utility score is used to determine whether the recommendation effect is satisfactory. If it is unsatisfactory, the LLM parameters, quick screening parameters, and new solution determination parameters are adjusted.