Audio and video recommendation method, system and device and storage medium
By constructing a user interest evolution model, real-time analysis of user short-term and long-term interest label deviation, dynamically adjust recommendation weights, and introducing cross-domain content, it solves the problem of information cocoon in traditional audio and video recommendation systems, and achieves more accurate and diversified recommendation effects.
Patent Information
- Application Number
- CN202510335088.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-08-01
AI Technical Summary
Traditional personalized audio and video recommendation systems are prone to information cocoon phenomenon, and users cannot access new things, reducing information diversity.
By constructing a user interest evolution model, real-time analysis of the deviation between users' short-term and long-term interest labels, dynamically adjust recommendation weights, introduce cross-domain candidate content, generate diversified recommendation lists, and avoid information cocoons.
Effectively avoid the information cocoon effect, provide more diverse recommended content, improve user satisfaction and experience, and ensure that the recommended content meets the current interests of users.
Smart Images

Figure CN120407820A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio - video recommendation, and particularly relates to an audio - video recommendation method, system, device, and storage medium. Background Art
[0002] Traditional personalized audio - video push methods and systems mainly mine and analyze behavioral data such as users' browsing history, search history, purchase behavior, etc. Through these behavioral data, user interest preferences can be modeled, and corresponding audio - videos can be recommended to users based on these models. Although personalized content can be provided based on user preferences, it will also cause a phenomenon called "information cocoon". When the recommendation system over - emphasizes personalized recommendation, it will overly filter information that does not match the user's historical preferences, making the user in a closed environment with only the information they are interested in, thus restricting the user's access to and understanding of new things and reducing the diversity of information.
[0003] Based on this, the present invention provides an audio - video recommendation method, system, device, and storage medium to solve the above problems. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides an audio - video recommendation method, system, device, and storage medium to solve the problems in the prior art.
[0005] One embodiment of the present invention provides an audio - video recommendation method, including the following steps:
[0006] Collect user historical behavior data and audio - video content feature data;
[0007] Construct a user interest evolution model based on the user historical behavior data and audio - video content feature data, and use the user interest evolution model to analyze the deviation degree between the user's short - term interest tags and long - term interest tags in real - time;
[0008] Based on the analysis result of the deviation degree, output a user interest representation vector to obtain the current interest tags;
[0009] Generate and output a corresponding recommendation list according to the current interest tags;
[0010] Obtain the current recommended content in real - time. When it is detected that the topic similarity of the content recommended continuously for n times exceeds a preset threshold, perform the following steps:
[0011] Reduce the recommendation weight of the current interest tags;
[0012] Input cross - domain candidate content according to the preset recommendation rules;
[0013] Generate and output a dynamically adjusted recommendation list based on the cross - domain candidate content.
[0014] By adopting the above solution, the deviation degree between the user's short-term interest tags and long-term interest tags is analyzed in real time using the user interest evolution model to dynamically capture the changing trend of the user's interests. Based on the deviation degree analysis, the user's current interest tags are obtained, and then a corresponding recommendation list is generated and pushed according to the user's current interest tags to ensure that the currently recommended content is of interest to the user. When calculating the deviation degree, the user's short-term interest tags and long-term interest tags are used for calculation, so that the changes in the user's interests can be discovered, and more diverse recommended content can be introduced, effectively avoiding the information cocoon effect. If only the user's long-term interest tags are concerned, it may cause the user to fall into the information cocoon and only be exposed to the content they are already familiar with. At the same time, when recommending content to the user, when it is detected that the topic similarity of the content recommended continuously for n times exceeds the preset threshold, in order to prevent the user from falling into the information cocoon, the recommendation weight of the current interest tags is reduced, and then according to the preset recommendation rules, cross-domain candidate content is input, and a dynamically adjusted recommendation list is generated and output. This method can prevent the user from falling into the information cocoon, provide more diverse recommended content, thereby improving the user's satisfaction and experience. At the same time, the input cross-domain content is selected from the content related to the user's interest tags to avoid random pushing, which makes the user feel abrupt and results in a low acceptance rate by the user.
[0015] In one embodiment, in the step of constructing the user interest evolution model based on the user's historical behavior data and the audio-visual content feature data and using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time, the construction of the user interest evolution model includes:
[0016] Constructing a two-dimensional user portrait, including a short-term interest portrait and a long-term interest portrait; the short-term interest portrait is the behavior sequence within the first preset time, modeled using an LSTM network; the long-term interest portrait is the behavior pattern within the second preset time, modeled using matrix factorization;
[0017] Designing an interest decay factor to dynamically adjust the interest weight, and its expression is:
[0018] w(t) = e -γt γ ∈ [0.05, 0.2];
[0019] where t is the number of days since the behavior occurred.
[0020] By adopting the above solution, a two-dimensional user profile is constructed, including a short-term interest profile for a first preset time, such as the recent seven days, and a long-term interest profile for a second preset time, such as the recent year, which can provide more accurate and personalized recommendation services. The short-term interest profile uses an LSTM network to capture the user's immediate behavior patterns, while the long-term interest profile uses matrix factorization technology to identify the user's stable preferences. By designing an interest decay factor to dynamically adjust the interest weights, it is ensured that the recommendation system can respond promptly to changes in user interests while retaining the user's long-term interest preferences, thereby enhancing the user experience and engagement and optimizing the content recommendation efficiency.
[0021] In one embodiment, in the step of constructing a user interest evolution model based on user historical behavior data and audio-visual content feature data and using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time, an anti-bias mechanism is added, which specifically includes the following steps:
[0022] De-identify the user sensitive attributes in the user historical behavior data to generate generalized category tags;
[0023] When training the user interest evolution model, a loss function including a fairness constraint term is adopted, and its expression is:
[0024]
[0025] where, E k is the content exposure times of the k-th user group, is the overall average exposure times, τ is the fairness difference threshold, and λ is the weight adjustment coefficient.
[0026] By adopting the above solution, in audio-visual recommendations, there may be recommendation biases, such as biases caused by region, gender, etc., resulting in unequal exposure of user groups, reduction of content diversity, formation of information cocoons, etc.; therefore, by adding an anti-bias mechanism, the recommended content is fair and there is no recommendation deviation caused by bias; by de-identifying the user sensitive attributes, that is, the attributes that are prone to recommendation biases, such as gender, region, etc., to generate generalized category tags, while protecting user privacy and maintaining data availability, when training the user interest evolution model, a loss function including a fairness constraint term is adopted, and the fairness constraint term encourages the model to balance the exposure times of each user group and avoid over-recommending or ignoring a certain group. By reducing over-exposure, the model is more motivated to recommend diverse audio-visual content to meet the needs of different users.
[0027] In one embodiment, the loss function of the fairness constraint term is implemented through adversarial training, which specifically includes the following steps:
[0028] (a) Construct a sensitive attribute prediction model, where the input of the sensitive attribute prediction model is the user interest representation vector generated by the user interest evolution model, and the output is the predicted probability distribution of the user's sensitive attributes;
[0029] (b) During the process of training the user interest evolution model, the following operations are synchronously performed:
[0030] (i) Minimize the prediction error between the recommendation result and the user's actual behavior through the recommendation accuracy loss function;
[0031] (ii) Maximize the prediction error of the sensitive attribute prediction model for the user's sensitive attributes through the adversarial loss function;
[0032] (c) Backpropagate the loss gradient of step (b) through the gradient reversal layer, so that the association strength between the user interest representation vector and the user's sensitive attributes is lower than the preset threshold.
[0033] By adopting the above scheme, a sensitive attribute prediction model is constructed to enhance the fairness of recommendations, and its prediction error is maximized during the training process. Adversarial training can effectively reduce the dependence of the user interest evolution model on the user's sensitive attributes (such as gender, age, region, etc.). This makes the user interest representation vector generated by the user interest evolution model more neutral, reduces the bias towards specific groups, and thus improves the fairness of the recommendation system. In adversarial training, the user interest evolution model optimizes its prediction accuracy through the recommendation accuracy loss function to ensure that the recommendation result is consistent with the user's actual behavior. At the same time, the adversarial model maximizes the sensitive attribute prediction error, prompting the user interest evolution model to avoid unfairness caused by relying on sensitive attributes while improving accuracy. This dual optimization ensures the balance between the accuracy and fairness of the recommendation system. Through adversarial training, when generating the user interest representation vector, the user interest evolution model reduces its dependence on sensitive attributes, making the recommendation result more interpretable. Users can more easily understand how the recommendation system makes recommendations based on their interests and behaviors without being affected by unfair factors.
[0034] In one embodiment, in the step of outputting the user interest representation vector based on the deviation analysis result to obtain the current interest label, the following specific steps are included:
[0035] Obtain the deviation analysis result;
[0036] When the deviation between the user's short-term interest label and the long-term interest label is less than the preset threshold range, obtain several interest labels that co-occur in the user's short-term interest label and the long-term interest label and whose occurrence frequency is greater than the preset frequency threshold, and output the user interest representation vector according to the obtained several interest labels to obtain the current interest label;
[0037] When the deviation between the user's short-term interest tags and long-term interest tags is greater than the preset threshold range, obtain several interest tags with a frequency of occurrence greater than the preset frequency threshold in the user's short-term interest tags, and output a user interest representation vector based on the obtained several interest tags to obtain the current interest tags.
[0038] By adopting the above solution, according to the analysis result of the deviation degree, the smaller the deviation degree, the closer the user's short-term interest is to the long-term interest, and the greater the deviation degree, the greater the difference between the user's short-term interest and long-term interest; when the deviation degree is less than the preset threshold range, several interest tags that co-occur in the user's short-term interest tags and long-term interest tags and have a frequency of occurrence greater than the preset frequency threshold are used as the user interest representation vector for output to obtain the corresponding recommendation list. Conversely, several interest tags with a frequency of occurrence greater than the preset frequency threshold in the user's short-term interest tags are used as the user interest representation vector for output to obtain the corresponding recommendation list, so as to ensure that the recommended content is acceptable to the user and ensure that the user will be interested in the currently recommended content.
[0039] In one embodiment, in the step of inputting cross-domain candidate content according to the preset recommendation rules, the following specific steps are included:
[0040] Use the BERT model to extract semantic vectors from the video content feature data, and construct a content semantic graph based on the semantic vectors and the knowledge graph;
[0041] Construct inter-domain associations through the content semantic graph to obtain cross-domain candidate content, where the cross-domain candidate content includes core domain content, adjacent domain content, and cross-border domain content;
[0042] When the recommendation weight of the current interest tags decreases, input cross-domain candidate content according to the preset recommendation rules, and the recommendation rules are to input core domain content, adjacent domain content, and cross-border domain content in sequence.
[0043] By adopting the above solution, when breaking the cocoon with cross-domain candidate content injected after detecting that the user may be trapped in an information cocoon, it can avoid making the user feel abrupt. Build associations between fields through the content semantic graph, such as the association link of "science and technology → history of science and technology → history", to obtain cross-domain candidate content. The core domain content is the content with a relatively high similarity to the user's interest tags, the adjacent domain content is the content related to but slightly different from the core domain content, and the cross-border domain content is the content seemingly unrelated to the core domain content but can attract the user through a certain connection. By recommending corresponding content according to the preset recommendation rules, it is ensured that the injected content will not make the user feel abrupt and can be smoothly passed; for example, when recommending new content to the user, instead of suddenly recommending something completely irrelevant (such as suddenly recommending beauty makeup videos to science and technology enthusiasts), it is through "progressive transition" to let the user naturally come into contact with the content of the new field.
[0044] In one embodiment, in the step of generating and outputting a dynamically adjusted recommendation list based on the cross-domain candidate content, the dynamically adjusted recommendation list includes recommended content generated by the current interest tags, core domain content, adjacent domain content, and cross-border domain content, and the recommendation rule of the dynamically adjusted recommendation list is:
[0045] The recommended content generated by the current interest tags, core domain content, adjacent domain content, and cross-border domain content are recommended in sequence and in a cycle;
[0046] Among them, when transitioning from the content generated by the current interest tags to the core domain content, calculate the potential connection features between the core domain content and the current interest tags, and generate a transition prompt based on the calculation result;
[0047] Among them, when transitioning from the core domain content to the adjacent domain content, calculate the potential connection features between the core domain content and the adjacent domain content, and generate a transition prompt based on the calculation result;
[0048] Among them, when transitioning from the adjacent domain content to the cross-border domain content, calculate the potential connection features between the adjacent domain content and the cross-border domain content, and generate a transition prompt based on the calculation result.
[0049] By adopting the above solution, it can enable users to come into contact with different and diverse recommended content, avoid users being trapped in an information cocoon, provide more diverse recommended content, thereby improving user satisfaction and experience. At the same time, when recommending cross-domain candidate content, calculate the potential connection features between the current domain content and the next domain content, and generate a transition prompt as the reason for recommendation. For example, "Based on the science and technology content you like, it is recommended that you learn about the history of science and technology", ensuring that users will not feel that the newly recommended content is completely irrelevant, but can understand the reason for the recommendation, thereby improving user satisfaction and experience.
[0050] This application also relates to an audio-visual recommendation system, including:
[0051] A data acquisition module, configured to acquire user historical behavior data and audio-visual content feature data;
[0052] A construction and analysis module, configured to construct a user interest evolution model according to the user historical behavior data and the audio-visual content feature data, and use the user interest evolution model to analyze the deviation degree between the short-term interest label and the long-term interest label of the user in real time;
[0053] A first output module, configured to output a user interest representation vector based on the analysis result of the deviation degree to obtain a current interest label;
[0054] A second output module, configured to generate and output a corresponding recommendation list according to the current interest label;
[0055] A detection module, configured to obtain the current recommended content in real time. When it is detected that the topic similarity of the continuously recommended content exceeds a preset threshold for n times, the following steps are executed:
[0056] An adjustment sub-module, configured to reduce the recommendation weight of the current interest label;
[0057] An input sub-module, configured to input cross-domain candidate content according to a preset recommendation rule;
[0058] A dynamic adjustment module, configured to generate and output a dynamically adjusted recommendation list based on the cross-domain candidate content.
[0059] By adopting the above solution, the deviation degree between the user's short-term interest tags and long-term interest tags is analyzed in real time using the user interest evolution model to dynamically capture the changing trend of the user's interests. Based on the deviation degree analysis, the user's current interest tags are obtained, and then a corresponding recommendation list is generated and pushed according to the user's current interest tags to ensure that the content currently recommended is of interest to the user. When calculating the deviation degree, the user's short-term interest tags and long-term interest tags are used for calculation, so as to discover the changes in the user's interests and introduce more diverse recommended content, which can effectively avoid the information cocoon effect. If only the user's long-term interest tags are concerned, it may lead to the user being trapped in the information cocoon and only exposed to the content they are already familiar with. At the same time, when recommending content to the user, when it is detected that the topic similarity of the content recommended continuously for n times exceeds the preset threshold, in order to prevent the user from being trapped in the information cocoon, the recommendation weight of the current interest tags is reduced, and then according to the preset recommendation rules, cross-domain candidate content is input, and a dynamically adjusted recommendation list is generated and output. This method can prevent the user from being trapped in the information cocoon, provide more diverse recommended content, thereby improving the user's satisfaction and experience. At the same time, the input cross-domain content is selected from the content related to the user's interest tags to avoid random pushing, which makes the user feel abrupt and leads to the problem of low user acceptance.
[0060] This application also relates to a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned audio and video recommendation method are implemented.
[0061] This application also relates to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned audio and video recommendation method are implemented.
[0062] The above-mentioned audio and video recommendation method, system, device, and storage medium provided by the above embodiments have the following beneficial effects:
[0063] 1. By using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time to dynamically capture the changing trend of the user's interests, based on the deviation degree analysis, the user's current interest tags are obtained, and then a corresponding recommendation list is generated and pushed according to the user's current interest tags to ensure that the content currently recommended is of interest to the user. When calculating the deviation degree, the user's short-term interest tags and long-term interest tags are used for calculation, so as to discover the changes in the user's interests and introduce more diverse recommended content, which can effectively avoid the information cocoon effect. If only the user's long-term interest tags are concerned, it may lead to the user being trapped in the information cocoon and only exposed to the content they are already familiar with..
[0064] 2. In one embodiment, when recommending content to users, when it is detected that the topic similarity of the continuously recommended content exceeds a preset threshold for n consecutive times, to prevent users from being trapped in an information cocoon, the recommendation weight of the current interest tag is reduced. Then, according to the preset recommendation rules, cross-domain candidate content is input, and a dynamically adjusted recommendation list is generated and output. This method can prevent users from being trapped in an information cocoon, provide more diverse recommended content, thereby improving user satisfaction and experience. At the same time, the input cross-domain content is selected from content related to the user's interest tags to avoid random pushing, which may make users feel abrupt and lead to low user acceptance.
[0065] 3. In one embodiment, a sensitive attribute prediction model is constructed to enhance the fairness of recommendations, and during the training process, its prediction error is maximized. Adversarial training can effectively reduce the dependence of the user interest evolution model on user sensitive attributes (such as gender, age, region, etc.). This makes the user interest representation vector generated by the user interest evolution model more neutral, reduces bias towards specific groups, and thus improves the fairness of the recommendation system. In adversarial training, the user interest evolution model optimizes its prediction accuracy through a recommendation accuracy loss function to ensure that the recommendation results are consistent with the actual behavior of users. At the same time, the adversarial model maximizes the sensitive attribute prediction error, prompting the user interest evolution model to avoid unfairness caused by relying on sensitive attributes while improving accuracy. This dual optimization ensures the balance between the accuracy and fairness of the recommendation system. Through adversarial training, when generating the user interest representation vector, the user interest evolution model reduces its dependence on sensitive attributes, making the recommendation results more interpretable. Users can more easily understand how the recommendation system makes recommendations based on their interests and behaviors without being affected by unfair factors.
[0066] 4. In one embodiment, to avoid making users feel abrupt when injecting cross-domain candidate content to break the cocoon after detecting that users may be trapped in an information cocoon. By constructing an inter-domain association through a content semantic graph, such as the association link of "technology → history of technology → history", cross-domain candidate content is obtained. The core domain content is the content with a relatively high similarity to the user's interest tags, the adjacent domain content is the content related to the core domain content but slightly different, and the cross-border domain content is the content seemingly unrelated to the core domain content but can attract users through some connection. By recommending corresponding content according to the preset recommendation rules, it is ensured that the injected content will not make users feel abrupt and can be smoothly passed; for example, when recommending new content to users, instead of suddenly recommending something completely irrelevant (such as suddenly recommending beauty makeup videos to technology enthusiasts), through "progressive transition", users can naturally come into contact with content in new fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0068] Figure 1 It is a flowchart of an audio - video recommendation method provided by an embodiment of the present invention;
[0069] Figure 2 It is a principle block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0071] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0072] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, if "and / or" or "and / or" appears throughout the text, its meaning includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or the solution where A and B are satisfied simultaneously. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0073] One embodiment of the present invention provides an audio - video recommendation method, including the following steps:
[0074] S10. Collect user historical behavior data and audio - video content feature data;
[0075] S20. Construct a user interest evolution model based on the user's historical behavior data and the audio-visual content feature data, and use the user interest evolution model to analyze the deviation degree between the short-term interest tags and the long-term interest tags of the user in real time;
[0076] S30. Based on the analysis result of the deviation degree, output a user interest representation vector to obtain the current interest tags;
[0077] S40. Generate and output a corresponding recommendation list according to the current interest tags;
[0078] S50. Obtain the current recommended content in real time. When it is detected that the topic similarity of the content recommended for n consecutive times exceeds a preset threshold, perform the following steps:
[0079] S51. Reduce the recommendation weight of the current interest tags;
[0080] S52. Input cross-domain candidate content according to the preset recommendation rules;
[0081] S60. Generate and output a dynamically adjusted recommendation list based on the cross-domain candidate content.
[0082] As described in step S10 above, the collected user historical behavior data includes the user's viewing records, search records, like records, etc. Collecting the user historical behavior data and the audio-visual content feature data provides a data basis for constructing the user interest evolution model.
[0083] As described in step S20 above, construct a user interest evolution model using the collected user historical behavior data and the audio-visual content feature data, and use the constructed user interest evolution model to analyze the deviation degree between the short-term interest tags and the long-term interest tags in the user historical behavior data in real time. The purpose of the deviation degree analysis is to better understand the dynamic changes of the user's interests, provide more accurate, diverse and timely recommended content, so as to improve the user experience and the effectiveness of the recommendation system.
[0084] As described in the above steps S30 - S40, the analysis results of the deviation degree can be used to adjust the recommendation weights, enabling the recommendation system to more flexibly respond to changes in user interests. For example, when the short - term interests of a user deviate significantly from the long - term interests, the weights of long - term interest tags can be reduced, and the weights of short - term interest tags can be increased to provide recommendations that better meet the user's current needs. According to the analysis results of the deviation degree, the user interest evolution model outputs the corresponding user interest representation vector to obtain the current interest tags, and generates the corresponding recommendation list based on the obtained interest tags; at the same time, the recommendation weights are dynamically adjusted through the deviation degree to avoid over - emphasizing short - term interests. For example, among the long - term interest tags of a user, there are sports, technology, history, music, food, beauty, entertainment, home life, etc., and among the short - term interest tags are sports, technology, music, food; therefore, in the generated recommendation list, there are not only sports, technology, music, food, but also beauty, entertainment, home life, etc., but the proportion of the content of tags such as beauty, entertainment, home life, etc. will not be so high.
[0085] As described in the above steps S50 - S60, the current recommended content is obtained in real - time, and the similarity of the recommended content is detected. For example, when it is detected that the similarity of 20 consecutive recommended contents exceeds 80%, such as 20 consecutive recommended contents are all contents with tags such as sports, technology, music, food, etc., at this time, to prevent the user from falling into an information cocoon, the recommendation weights of the current interest tags are reduced, that is, the recommendation weights of tags such as sports, technology, music, food, etc. are reduced, so that the proportion of the recommendation of these related contents is reduced, and then according to the preset recommendation rules, cross - domain candidate contents are input, and a dynamically adjusted recommendation list is generated and output. This method can prevent the user from falling into an information cocoon, provide more diverse recommended content, thereby improving user satisfaction and experience. The input cross - domain content is selected from the content related to the user interest tags to avoid random pushing, which may make the user feel abrupt and lead to a low acceptance rate by the user.
[0086] In one embodiment, in step S20, the construction of the user interest evolution model includes:
[0087] S21. Construct a two - dimensional user portrait, including a short - term interest portrait and a long - term interest portrait; the short - term interest portrait is the behavior sequence within the first preset time, modeled using the LSTM network; the long - term interest portrait is the behavior pattern within the second preset time, modeled using matrix factorization;
[0088] S22. Design an interest decay factor to dynamically adjust the interest weights, and its expression is:
[0089] w(t)=e -γt γ ∈ [0.05, 0.2];
[0090] where t is the number of days since the behavior occurred.
[0091] In this embodiment, the short-term interest profile is modeled by an LSTM network based on the click, favorite, and share behavior sequences of the user within the last 7 days (the first preset time); the long-term interest profile is modeled by matrix factorization based on the viewing duration and content type preferences of the user within the last year (the second preset time); the long short-term memory network (LSTM) is used to model the user's behavior sequence to capture the user's immediate interests and behavior patterns; the matrix factorization technique is used to model the user's behavior patterns to identify the user's stable interests and preferences. By designing an interest decay factor to dynamically adjust the interest weights, it is ensured that the recommendation system can respond in a timely manner to changes in the user's interests while retaining the user's long-term interest preferences, thereby improving the user experience and engagement and optimizing the content recommendation efficiency.
[0092] The values 0.05 and 0.2 are used to define the decay rate parameter γ in the interest decay factor. γ = 0.05: indicates that the interest decays slowly. As time goes by, the influence of the user's early behavior on the current recommendation will gradually weaken, but the weakening speed is relatively slow. This means that the user's long-term interests still have a certain influence in the recommendation. γ = 0.2: indicates that the interest decays quickly. As time goes by, the influence of the user's early behavior on the current recommendation will weaken rapidly. This means that the user's recent interests have a greater weight in the recommendation. By adjusting the value of γ, the speed at which the user's interests decay over time can be controlled, thus balancing the influence of short-term and long-term interests in the recommendation system.
[0093] In one of the embodiments, in step S20, an anti-bias mechanism is added, which specifically includes the following steps:
[0094] S201. De-identify the user-sensitive attributes in the user's historical behavior data to generate generalized category labels;
[0095] S202. When training the user interest evolution model, use a loss function that includes a fairness constraint term, and its expression is:
[0096]
[0097] where E k is the number of content exposures of the k-th user group, is the overall average exposure number, τ is the fairness difference threshold, and λ is the weight adjustment coefficient.
[0098] In this embodiment, by adding a fairness constraint term to the loss function, the model will minimize the difference in the number of exposure times between different user groups during the training process, ensuring that each group can obtain relatively fair content recommendations. This fairness constraint term encourages the user interest evolution model to make the ratio of the exposure times of each user group to the average exposure times not exceed the threshold τ. If the exposure times of a certain group are too many, exceeding the average value plus the threshold, then this difference will be included in the loss function, thereby prompting the model to adjust and reduce the over-exposure of this group.
[0099] For example, to ensure the fairness of the recommendation results for different user groups (such as different ages, genders, regions, etc.), a loss function containing a fairness constraint term is introduced when training the user interest evolution model.
[0100] (1) Define sensitive attributes and exposure times:
[0101] Sensitive attribute: Age (youth, middle age, old age).
[0102] Exposure times: The number of times each age group receives recommended audio and video within a certain period of time.
[0103] (2) Set the threshold τ:
[0104] Assume the threshold τ = 0.15, indicating that the ratio of the exposure times of each group to the average exposure times should not exceed 15%;
[0105] (3) Calculate the average exposure times and the exposure times of each group:
[0106] Total exposure times: 2000 times.
[0107] Average exposure times: 2000 / 3 ≈ 666.67 times.
[0108] Assume the exposure times of the youth group are 800 times, the exposure times of the middle-aged group are 700 times, and the exposure times of the elderly group are 500 times.
[0109] (4) Calculate the ratio of the exposure times of each group to the average exposure times:
[0110] Youth group: 800 / 666.67 ≈ 1.2 (exceeding the threshold τ = 1.15).
[0111] Middle-aged group: 700 / 666.67 ≈ 1.05 (not exceeding the threshold τ = 1.15).
[0112] Elderly group: 500 / 666.67 ≈ 0.75 (not exceeding the threshold τ = 0.85).
[0113] (5) Calculate the loss of the fairness constraint term:
[0114] Differences among the youth group: 1.2 - 1.15 = 0.05.
[0115] Differences among the middle-aged group: 1.05 - 1.15 = -0.1 (not exceeding the threshold, not counted as loss).
[0116] Differences among the elderly group: 0.75 - 0.85 = -0.1 (not exceeding the threshold, not counted as loss).
[0117] Loss of the fairness constraint term: 0.05 2 = 0.0025.
[0118] (6) Therefore, during the model training process, the model will adjust its recommendation strategy according to the fairness constraint term in the loss function, reduce the overexposure of the youth group, and increase the exposure times of the elderly group to achieve the fairness goal.
[0119] Through the above example, it can be seen how the loss function of the fairness constraint term encourages the model to make the ratio of the exposure times of each user group to the average exposure times not exceed the threshold τ. If the exposure times of a certain group are too many and exceed the average value plus the threshold, then this difference will be included in the loss function, thus prompting the model to adjust and reduce the overexposure of this group.
[0120] By adding an anti-bias mechanism, the recommended content is made fair and there is no recommendation bias caused by bias.
[0121] In one embodiment, in step S202, the loss function of the fairness constraint term is implemented through adversarial training, specifically including the following steps:
[0122] (a) Construct a sensitive attribute prediction model. The input of the sensitive attribute prediction model is the user interest representation vector generated by the user interest evolution model, and the output is the predicted probability distribution of the user's sensitive attributes;
[0123] As described in this step, the user interest evolution model generates a vector containing the user's interest in various audio-visual contents, and then constructs a sensitive attribute prediction model, such as a simple fully connected neural network. This model receives the user interest vector as input and outputs the predicted probability distribution of sensitive attributes such as the user's gender and age.
[0124] (b) During the process of training the user interest evolution model, the following operations are synchronously executed:
[0125] (i) Minimize the prediction error between the recommendation result and the user's actual behavior through the recommendation accuracy loss function; 8]
[0126] As described in this step, a recommended precision loss function (such as the cross-entropy loss function) is used to measure the difference between the recommended result and the actual click behavior of the user, and the model parameters are updated through the backpropagation algorithm to minimize this difference.
[0127] (ii) Maximize the prediction error of the sensitive attribute prediction model for the user's sensitive attributes through the adversarial loss function;
[0128] As described in this step, an adversarial loss function (such as the cross-entropy loss function) is used to measure the difference between the prediction result of the sensitive attribute prediction model and the actual sensitive attributes of the user, and the model parameters are updated through the backpropagation algorithm to maximize this difference.
[0129] (c) Backpropagate the loss gradient of step (b) through the gradient reversal layer, so that the association strength between the user interest representation vector and the user's sensitive attributes is lower than the preset threshold.
[0130] As described in this step, the gradient reversal layer is used to backpropagate the loss gradient of step (b), so that the association strength between the user interest vector generated by the user interest evolution model and the user's sensitive attributes is reduced, thereby improving the fairness of the recommendation system.
[0131] In this embodiment, during the fairness training phase, the recommendation system synchronously runs the user interest evolution model (main model) and the sensitive attribute prediction model (adversarial model). The user interest evolution model updates the parameters through the gradient reversal layer, so that the sensitive attribute prediction model cannot accurately predict the user's gender (test accuracy 49.3%), while maintaining the recommendation accuracy.
[0132] In one of the embodiments, in step S30, the following steps are specifically included:
[0133] S31. Obtain the analysis result of the deviation degree;
[0134] S32a. When the deviation degree between the user's short-term interest tags and long-term interest tags is less than the preset threshold range, obtain several interest tags that co-occur and have a frequency greater than the preset frequency threshold in the user's short-term interest tags and long-term interest tags, and output a user interest representation vector according to the obtained several interest tags to obtain the current interest tags;
[0135] S32b. When the deviation degree between the user's short-term interest tags and long-term interest tags is greater than the preset threshold range, obtain several interest tags that have a frequency greater than the preset frequency threshold in the user's short-term interest tags, and output a user interest representation vector according to the obtained several interest tags to obtain the current interest tags.
[0136] In this embodiment, according to the analysis result of the deviation degree, the smaller the deviation degree is, the closer the user's short-term interest is to the long-term interest. The larger the deviation degree is, the greater the difference between the user's short-term interest and the long-term interest is. For example, the preset threshold range of the deviation degree is 0.4-0.7, and the preset frequency threshold is 70%. When the analysis result of the deviation degree is less than the preset threshold range, several interest tags that co-occur in the user's short-term interest tags and long-term interest tags and have a frequency greater than the preset frequency threshold are used as the user interest representation vector for output to obtain the corresponding recommendation list. Otherwise, several interest tags with a frequency greater than the preset frequency threshold in the user's short-term interest tags are used as the user interest representation vector for output to obtain the corresponding recommendation list, so as to ensure that the recommended content is acceptable to the user and ensure that the user will be interested in the currently recommended content.
[0137] In one of the embodiments, in step S52, it specifically includes the following steps:
[0138] S521. Use the BERT model to extract the semantic vectors in the video content feature data, and construct a content semantic graph based on the semantic vectors and the knowledge graph;
[0139] S522. Construct an inter-domain association through the content semantic graph to obtain cross-domain candidate content, where the cross-domain candidate content includes core domain content, adjacent domain content, and cross-border domain content;
[0140] S523. When the recommendation weight of the current interest tag decreases, input the cross-domain candidate content according to the preset recommendation rule, and the recommendation rule is to input the core domain content, adjacent domain content, and cross-border domain content in sequence.
[0141] In this embodiment, the BERT model is responsible for converting text information (such as video titles, introductions, subtitles) into mathematical vectors (a set of numbers). For example, converting "Smartphone Review" into a vector similar to [0.2, 0.7, -0.3...]. These vectors can represent the semantic meaning of the content (similar content will have similar vectors). Then, based on the obtained semantic vectors and the industry knowledge graph, a content semantic graph is constructed. The industry knowledge graph is used to add professional associations. An inter-domain association (such as the association link of "Technology → History of Technology → History") is established through the content semantic graph to obtain the corresponding cross-domain candidate content.
[0142] Cross - domain candidate content includes core domain content, adjacent domain content, and cross - border domain content; core domain content is content with a relatively high similarity to the user's interest tags, adjacent domain content is content related to but slightly different from the core domain content, and cross - border domain content is content that seemingly has no relation to the core domain content but can attract users through a certain connection; the similarity between the core domain content and the current interest tag content is greater than 70%, the similarity between the adjacent domain content and the current interest tag content is between 40% and 70%, and the similarity between the cross - border domain content and the current interest tag content is less than 40%.
[0143] After obtaining the cross - domain candidate content, when injecting the cross - domain candidate content for breaking the cocoon, corresponding content is recommended according to the preset recommendation rules to ensure that the injected content does not make the user feel abrupt and can be smoothly passed; for example, when recommending new content to users, instead of suddenly recommending completely irrelevant things (such as suddenly recommending beauty makeup videos to technology enthusiasts), through "progressive transition", users can naturally come into contact with content in new fields.
[0144] In one embodiment, in step S60, the dynamically adjusted recommendation list includes recommended content generated by the current interest tag, core domain content, adjacent domain content, and cross - border domain content, and the recommendation rule of the dynamically adjusted recommendation list is as follows:
[0145] The recommended content generated by the current interest tag, core domain content, adjacent domain content, and cross - border domain content are recommended in sequence and in a loop;
[0146] Among them, when transitioning from the content generated by the current interest tag to the core domain content, calculate the potential connection features between the core domain content and the current interest tag, and generate a transition prompt based on the calculation result;
[0147] Among them, when transitioning from the core domain content to the adjacent domain content, calculate the potential connection features between the core domain content and the adjacent domain content, and generate a transition prompt based on the calculation result;
[0148] Among them, when transitioning from the adjacent domain content to the cross - border domain content, calculate the potential connection features between the adjacent domain content and the cross - border domain content, and generate a transition prompt based on the calculation result.
[0149] In this embodiment, when the recommendation list injects cross - domain candidate content for dynamic adjustment, the recommended content generated by the current interest tag, core domain content, adjacent domain content, and cross - border domain content are recommended in sequence and in a loop, which can enable users to come into contact with different and diverse recommended content, avoid users falling into the information cocoon, provide more diverse recommended content, and thus improve user satisfaction and experience.
[0150] When recommending cross - domain candidate content, calculate the potential connection features between the current domain content and the next domain content, that is, find the common ground between the cross - domain content and the current interests, design the recommendation reasons based on the common ground, and generate transition prompts; for example, "Based on the technology content you like, it is recommended that you learn about the history of technology", ensuring that users do not feel that the newly recommended content is completely irrelevant, but can understand the reasons for the recommendation, thereby improving user satisfaction and experience.
[0151] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0152] In one of the embodiments, an audio - video recommendation system is provided. This audio - video recommendation system corresponds to the audio - video recommendation method in the above - mentioned embodiment. This audio - video recommendation system includes:
[0153] A data acquisition module, used to acquire user historical behavior data and audio - video content feature data;
[0154] A construction and analysis module, used to construct a user interest evolution model based on user historical behavior data and audio - video content feature data, and use the user interest evolution model to analyze the deviation degree between the short - term interest tags and long - term interest tags of users in real time;
[0155] A first output module, used to output a user interest representation vector based on the analysis result of the deviation degree to obtain the current interest tags;
[0156] A second output module, used to generate and output a corresponding recommendation list according to the current interest tags;
[0157] A detection module, used to obtain the current recommended content in real time. When it is detected that the topic similarity of the content recommended continuously for n times exceeds a preset threshold, the following steps are executed:
[0158] An adjustment sub - module, used to reduce the recommendation weight of the current interest tags;
[0159] An input sub - module, used to input cross - domain candidate content according to preset recommendation rules;
[0160] A dynamic adjustment module, used to generate and output a dynamically adjusted recommendation list based on the cross - domain candidate content.
[0161] For the specific limitations of an audio - video recommendation system, reference can be made to the limitations of an audio - video recommendation method in the foregoing text, which will not be elaborated here. Each module in the above - mentioned audio - video recommendation system can be implemented in whole or in part by software, hardware, or a combination thereof. The above - mentioned modules can be embedded in the processor of a computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above - mentioned modules.
[0162] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 2 shown. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non - volatile storage medium and an internal memory. The non - volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non - volatile storage medium. The database of the computer device is used to store user data, data processing, and analysis of interest tags, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio - video recommendation method.
[0163] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements an audio - video recommendation method.
[0164] In one embodiment, a computer - readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements an audio - video recommendation method.
[0165] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0167] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An audio-video recommendation method, characterized in that, It includes the following steps: Collect user historical behavior data and audio-visual content feature data; Construct a user interest evolution model based on the user historical behavior data and the audio-visual content feature data, and use the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time; Based on the analysis result of the deviation degree, output a user interest representation vector to obtain the current interest tags; Generate and output a corresponding recommendation list according to the current interest tags; Obtain the current recommended content in real time. When it is detected that the topic similarity of the content recommended continuously for n times exceeds a preset threshold, perform the following steps: Reduce the recommendation weight of the current interest tags; Input cross-domain candidate content according to the preset recommendation rules; Generate and output a dynamically adjusted recommendation list based on the cross-domain candidate content.
2. The audio-video recommendation method according to claim 1, wherein In the step of constructing a user interest evolution model based on the user historical behavior data and the audio-visual content feature data, and using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time, the construction of the user interest evolution model includes: Construct a two-dimensional user portrait, including a short-term interest portrait and a long-term interest portrait; the short-term interest portrait is the behavior sequence within the first preset time, modeled by an LSTM network; the long-term interest portrait is the behavior pattern within the second preset time, modeled by matrix factorization; Design an interest decay factor to dynamically adjust the interest weight, and its expression is: w(t) = e -γt γ ∈ [0.05, 0.2]; where t is the number of days since the behavior occurred.
3. The audio and video recommendation method according to claim 1, characterized in that In the step of constructing a user interest evolution model based on the user historical behavior data and the audio-visual content feature data, and using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time, an anti-bias mechanism is added, which specifically includes the following steps: Perform de-identification processing on the user sensitive attributes in the user historical behavior data to generate generalized category tags; When training the user interest evolution model, use a loss function containing a fairness constraint term, and its expression is: Among them, E k is the content exposure times of the k-th user group, is the overall average exposure times, τ is the fairness difference threshold, and λ is the weight adjustment coefficient.
4. The audio-visual recommendation method according to claim 3, wherein The loss function of the fairness constraint term is realized through adversarial training, which specifically includes the following steps: (a) Construct a sensitive attribute prediction model. The input of the sensitive attribute prediction model is the user interest representation vector generated by the user interest evolution model, and the output is the predicted probability distribution of the user sensitive attributes; (b) During the process of training the user interest evolution model, synchronously perform the following operations: (i) Minimize the prediction error between the recommendation result and the user's actual behavior through the recommendation accuracy loss function; (ii) Maximize the prediction error of the sensitive attribute prediction model for the user sensitive attributes through the adversarial loss function; (c) Backpropagate the loss gradient in step (b) through a gradient reversal layer, so that the association strength between the user interest representation vector and the user sensitive attributes is lower than a preset threshold.
5. The audio-visual recommendation method according to claim 1, wherein In the step of outputting a user interest representation vector based on the analysis result of the deviation degree to obtain the current interest tags, it specifically includes the following steps: Obtain the analysis result of the deviation degree; When the deviation degree between the user's short-term interest tags and long-term interest tags is less than the preset threshold range, obtain several interest tags that co-occur in the user's short-term interest tags and long-term interest tags and whose occurrence frequencies are greater than the preset frequency threshold, and output a user interest representation vector based on the obtained several interest tags to obtain the current interest tags; When the deviation degree between the user's short-term interest tags and long-term interest tags is greater than the preset threshold range, obtain several interest tags whose occurrence frequencies are greater than the preset frequency threshold in the user's short-term interest tags, and output a user interest representation vector based on the obtained several interest tags to obtain the current interest tags.
6. The audio-video recommendation method according to claim 5, characterized in that In the step of inputting cross-domain candidate content according to the preset recommendation rules, the following steps are specifically included: Use the BERT model to extract semantic vectors from the video content feature data, and construct a content semantic graph based on the semantic vectors and the knowledge graph; Construct an inter-domain association through the content semantic graph to obtain cross-domain candidate content, and the cross-domain candidate content includes core domain content, adjacent domain content, and cross-border domain content; When the recommendation weight of the current interest tags decreases, input cross-domain candidate content according to the preset recommendation rules, and the recommendation rules are to input core domain content, adjacent domain content, and cross-border domain content in sequence.
7. The audio and video recommendation method according to claim 6, wherein In the step of generating and outputting a dynamically adjusted recommendation list based on the cross-domain candidate content, the dynamically adjusted recommendation list includes recommended content generated by the current interest tags, core domain content, adjacent domain content, and cross-border domain content, and the recommendation rules of the dynamically adjusted recommendation list are: The recommended content generated by the current interest tags, core domain content, adjacent domain content, and cross-border domain content are recommended in sequence and cyclically; Among them, when transitioning from the content generated by the current interest tags to the core domain content, calculate the potential connection features between the core domain content and the current interest tags, and generate a transition prompt based on the calculation result; Among them, when transitioning from the core domain content to the adjacent domain content, calculate the potential connection features between the core domain content and the adjacent domain content, and generate a transition prompt based on the calculation result; Among them, when transitioning from the adjacent domain content to the cross-border domain content, calculate the potential connection features between the adjacent domain content and the cross-border domain content, and generate a transition prompt based on the calculation result.
8. An audio - video recommendation system for implementing the steps of an audio - video recommendation method according to any one of claims 1 - 7, characterized in that, Include: A data collection module for collecting user historical behavior data and audio-visual content feature data; A construction and analysis module for constructing a user interest evolution model based on the user historical behavior data and audio-visual content feature data, and using the user interest evolution model to analyze the deviation degree between the user's short-term interest tags and long-term interest tags in real time; A first output module for outputting a user interest representation vector based on the analysis result of the deviation degree to obtain the current interest tags; A second output module for generating and outputting a corresponding recommendation list according to the current interest tags; A detection module for obtaining the current recommended content in real time. When it is detected that the topic similarity of the content recommended continuously for n times exceeds the preset threshold, the following steps are executed: An adjustment sub-module for reducing the recommendation weight of the current interest tags; An input sub-module, configured to input cross-domain candidate content according to a preset recommendation rule; A dynamic adjustment module, configured to generate and output a dynamically adjusted recommendation list based on the cross-domain candidate content.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of an audio-video recommendation method according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of an audio-video recommendation method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Intelligent pushing method and system based on dynamic interest attenuation analysis
CN120711235A
Intelligent push method and system based on dynamic interest decay analysis
CN120711235B
Personalized article recommendation method and device, electronic equipment and storage medium
CN121434501A
Personalized item recommendation method and apparatus, electronic device, and storage medium
CN121434501B