Personalized video recommendation method and system based on big data

Through multi-threshold user group integration and optimized genetic algorithm, the problem of insufficient adaptability of personalized video recommendation system to preference differences and user behavior noise is solved, more efficient user preference capture and dynamic adaptation are achieved, and the accuracy and adaptability of the recommendation system are improved.

CN120769084AActive Publication Date: 2025-10-10CHINA UNICOM VIDEO TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511283904.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing personalized video recommendation systems cannot simultaneously take into account preference differences within and across content channels. The recommendation system has high operation and maintenance costs, is easily misled by user behavior noise, has low recommendation accuracy, and is insufficiently sensitive to user preference mutation patterns and has poor adaptability, resulting in poor recommendation results.

Method used

Through multi-threshold preliminary user group integration, the candidate user group structure is enriched, the optimal user preference segmentation of each content channel is adaptively retained, the optimized genetic algorithm is introduced to optimize the user segmentation of user viewing behavior data, and a dynamic adjustment mechanism of crossover probability and neighborhood mutation probability is designed to enhance the adaptability to dynamic changes in user preferences.

Benefits of technology

It improves the accuracy and adaptability of personalized video recommendations, enhances the ability to capture users' real preferences and the sensitivity to potential preference changes, and improves the overall effect of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769084A_ABST
    Figure CN120769084A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized video recommendation method and system based on big data. The method comprises the steps of data acquisition, preliminary user group complete set division, preliminary user group refinement, final user subdivision group division and personalized video recommendation. The invention belongs to the technical field of user recommendation, and particularly relates to a personalized video recommendation method and system based on big data, according to the scheme, through multi-threshold preliminary user group integration, a candidate user group structure is enriched, the optimal user preference subdivision degree of each content channel is adaptively reserved, and the capturing ability for the real preference of a user is enhanced; a migration path of user preference is highlighted by refining the interior of a preliminary user group of user watching behavior data; an optimization genetic algorithm is introduced to optimize user subdivision group division of user watching behavior data, and a multi-objective optimization strategy is adopted to improve sensitivity to potential preference changes; designing a mechanism for dynamically adjusting crossover probability and neighborhood mutation probability; and the personalized video recommendation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of user recommendation technology, and specifically to a method and system for personalized video recommendation based on big data. Background Art

[0002] Personalized video recommendation methods are based on user viewing behavior data. By analyzing user preference similarities to segment and optimize groups, they accurately push video content that matches the shared preferences of different user groups. However, typical personalized video recommendation systems fail to account for preference differences within and across content channels. They also have high operational costs and are easily misled by user behavior noise, resulting in low recommendation accuracy. They also lack sensitivity to sudden changes in user preferences and are poorly adaptable to real-time changes in user preferences, leading to poor personalized video recommendation results. Summary of the Invention

[0003] To address the above-mentioned issues and overcome the shortcomings of the prior art, the present invention provides a method and system for personalized video recommendation based on big data. This method addresses the problems of general personalized video recommendation systems, such as their inability to simultaneously account for preference differences within and across content channels, high operational and maintenance costs, and susceptibility to misleading user behavior noise, resulting in low recommendation accuracy. This solution integrates preliminary user groups using multiple thresholds to enrich the structure of candidate user groups, adaptively retaining the optimal user preference segmentation for each content channel, and enhancing the ability to capture users' true preferences. It also refines the preliminary user groups based on user viewing behavior data to highlight the migration paths of user preferences, thereby improving the accuracy of personalized video recommendations. Furthermore, this method addresses the problem of general personalized video recommendation systems being insufficiently sensitive to user preference mutation patterns and poorly adaptable to real-time changes in user preferences, resulting in poor personalized video recommendation performance. This solution introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, employing a multi-objective optimization strategy to improve sensitivity to potential preference changes. Furthermore, a mechanism for dynamically adjusting crossover probability and neighborhood mutation probability is designed to automatically adjust the exploration intensity of user group segmentation, enhancing the overall recommendation system's adaptability to dynamic changes in user preferences and thereby improving the effectiveness of personalized video recommendations.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides a personalized video recommendation method based on big data, which includes the following steps:

[0005] Step S1: data collection;

[0006] Step S2: preliminary division of the entire user group;

[0007] Step S3: preliminary user group refinement;

[0008] Step S4: final user segmentation;

[0009] Step S5: Personalized video recommendation.

[0010] Furthermore, in step S1, the data collection is to collect user video viewing data; and perform feature coding and standardization processing on the user video viewing data.

[0011] Furthermore, in step S2, the preliminary user group full set division is performed by obtaining the user viewing behavior similarity matrix R by the Jaccard similarity of the video viewing data of two users, which is expressed as: ; and get the distance matrix D, expressed as: ; Perform indirect association operation on user preference on R to obtain user preference equivalent association matrix ; Take the threshold set ;in, is the Jaccard similarity between the u-th user video viewing data and the g-th user video viewing data; n is the total number of data; 、 and is the threshold used to divide the strength of user preference association; for each threshold Will As a preliminary user group segmentation; It is based on the threshold The obtained preliminary user groups are divided; the preliminary user group sets under all thresholds are combined to obtain the complete set of preliminary user groups.

[0012] Furthermore, in step S3, the preliminary user group refinement is to initially treat each user's video viewing data as a user segment within each preliminary user group; gradually merge the two user segments with the smallest distance; until the maximum number of merges is reached or the merged distance is lower than the merge threshold; obtain candidate user segments; and calculate the average correlation for each candidate user segment, which is expressed as: ;in, is the i-th candidate user segment The average correlation degree of g and k is the index of user video viewing data; is the Jaccard similarity between the video viewing data of the gth user and the video viewing data of the kth user; and based on the average correlation, the candidate user segments are divided into high correlation, medium correlation and low correlation; for the candidate user segments, if all user video viewing data belong to the same content channel and the candidate user segment is highly correlated, then it is retained. If all user video viewing data do not belong to the same content channel and the candidate user segment is low-correlated, retain The candidate user segments below; the rest The candidate user segments below.

[0013] Furthermore, in step S4, the final user segmentation is performed by screening all candidate user segment groups through a genetic algorithm, specifically including the following steps:

[0014] Step S41: Cluster coding; for the candidate user segment group set ;in, and The first and second user segments; let each candidate user segment correspond to a binary vector, the binary vector elements Expressed as: ; Chromosome length Boolean vector of Select clusters as the final partition; among them, and are the first and second Boolean vectors respectively. elements; Indicates the selection of candidate user segments ; Randomly generate N chromosomes;

[0015] Step S42: Multi-objective fitness evaluation; define the multi-objective functions to minimize the similarity between clusters of user video viewing data. , maximize the intra-cluster density of user video viewing data and minimize the number of clusters of user video viewing data ; are expressed as follows: ; ; ;in, is the jth candidate user segment; is the Jaccard similarity between the video viewing data of the kth user and the video viewing data of the lth user; X is the set of candidate user segments selected in the current solution;

[0016] Step S43: Crossover operation: Use a multi-objective solution ranking algorithm to assign an optimal solution level to each solution in the population; for solutions within the same level, calculate the distribution density based on the normalized distance of each target in that level; obtain the ranking; and use a roulette wheel to select and retain N parents. For each parent team, choose whether to perform a crossover operation based on the crossover probability; the crossover operation is as follows: randomly select a cluster subset S, copy all gene positions corresponding to S in parent 1 to offspring 1, and fill the remaining positions in the order of parent 2;

[0017] Step S44: Neighborhood mutation operation: For each offspring, a neighborhood mechanism is introduced to select a gene position based on the neighborhood mutation probability, and a right shift operation is performed to generate a neighborhood solution; the optimal solution is selected from all non-taboo neighborhoods and the current chromosome is updated; the used moves are recorded in the search taboo record, and the number of iterations is set.

[0018] Step S45: Conflict handling: For each new offspring, if there exists g such that , then there is a conflict. For each conflict point, keep one and close the other based on fitness. If there exists g such that , then the coverage point is missing, and the missing point is remedied to construct the submatrix ; Solve the linear equations to obtain the auxiliary Boolean vector y, select the cluster completion in the least incremental way, let when ;in, is the i-th element of the Boolean vector; is the i-th element of the auxiliary Boolean vector; select the one with the highest diversity in the population according to the Hamming distance chromosomes, and run a neighborhood search again;

[0019] Step S46: Update crossover mutation probability; calculate intermediate quantity: ; ; After fuzzy calculation is clarified, we get and ,renew: ; ;in, and are the minimum fitness value and average fitness value of all individuals in the t-th generation population respectively; is the average fitness value of all individuals in the t-1 generation population; and are the crossover probabilities of the tth generation and the t-1th generation respectively; and are the neighborhood mutation probabilities of the tth and t-1th generations, respectively; and are the increments of the crossover probability and neighborhood mutation probability of the tth generation, respectively; and is the intermediate quantity;

[0020] Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the final selected candidate user segmentation group as the final user segmentation group.

[0021] Furthermore, in step S5, the personalized video recommendation is based on the end user segmentation result, and directly recommends highly relevant videos watched by users in the cluster to users in the cluster, thereby achieving real-time similar preference matching in cluster units.

[0022] The personalized video recommendation system based on big data provided by the present invention includes a data acquisition module, a preliminary user group full set division module, a preliminary user group refinement module, a final user sub-group division module and a personalized video recommendation module;

[0023] The data acquisition module collects user video viewing data;

[0024] The preliminary user group full set division module divides the user video viewing data into preliminary user groups to obtain a user preference equivalent association matrix, and divides the user video viewing data into connected areas based on different thresholds to form preliminary user group full sets of user video viewing data;

[0025] The preliminary user group refinement module further refines the preliminary user group of user video viewing data into candidate user segments with high, medium and low correlation by merging user segments;

[0026] The final user segmentation module screens candidate user segments of user video viewing data using a genetic algorithm by designing a crossover operation and a neighborhood mutation operation to obtain the final user segmentation;

[0027] The personalized video recommendation module performs personalized video recommendation based on the end user segmentation.

[0028] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0029] (1) In view of the problems that general personalized video recommendation systems cannot take into account the preference differences within and across content channels at the same time, the operation and maintenance costs of the recommendation system are high, and it is easily misled by user behavior noise, resulting in low recommendation accuracy, this solution integrates multiple threshold preliminary user groups, enriches the candidate user group structure, adaptively retains the optimal user preference segmentation of each content channel, and enhances the ability to capture users' true preferences; through the refinement of the preliminary user group within the user viewing behavior data, it highlights the migration path of user preferences; and thus improves the accuracy of personalized video recommendations.

[0030] (2) In view of the problem that general personalized video recommendation systems are not sensitive enough to the sudden change pattern of user preferences and have poor adaptability to the real-time changes in user preferences, resulting in poor personalized video recommendation results, this solution introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, and adopts a multi-objective optimization strategy to improve the sensitivity to potential preference changes; a dynamic adjustment mechanism of crossover probability and neighborhood mutation probability is designed to automatically adjust the exploration intensity of user group segmentation, thereby enhancing the adaptability of the entire recommendation system to dynamic changes in user preferences; and thus improving the personalized video recommendation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A flowchart of the method for personalized video recommendation based on big data provided by the present invention;

[0032] Figure 2 This is a schematic diagram of the big data-based personalized video recommendation system provided by the present invention.

[0033] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0035] In the description of the present invention, it should be understood that terms such as "up", "down", "front", "back", "left", "right", "top", "bottom", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the system or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0036] Example 1, see Figure 1 The present invention provides a personalized video recommendation method based on big data, which includes the following steps:

[0037] Step S1: Data collection: Collect user video viewing data;

[0038] Step S2: preliminary division of the entire user group;

[0039] Step S3: Refine the preliminary user group; by merging user segments, further refine the preliminary user group of user video viewing data into candidate user segments with high, medium, and low correlation;

[0040] Step S4: final user segmentation: by designing crossover operations and neighborhood mutation operations, using genetic algorithms to screen candidate user segments of user video viewing data, and obtain final user segmentation;

[0041] Step S5: Personalized video recommendation: Personalized video recommendation is performed based on the end user segmentation.

[0042] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, data collection is to collect user video viewing data; the user video viewing data includes the user viewing time window, single video viewing time, viewing frequency, viewing cycle, number of video types watched by the user, cumulative viewing days and video category ratio; and the user video viewing data is feature coded and standardized.

[0043] Example 3, see Figure 1 This embodiment is based on the above embodiment. In step S2, the preliminary user group full set division is performed. In personalized video recommendation, the video viewing preferences of different users may be correlated. These correlations can be discovered through the preliminary user group full set division. The user viewing behavior similarity matrix R is obtained by the Jaccard similarity of the video viewing data of two users, which is expressed as: ; and get the distance matrix D, expressed as: ; Perform indirect association operation on user preference on R to obtain user preference equivalent association matrix ; Take the threshold set ;in, is the Jaccard similarity between the u-th user video viewing data and the g-th user video viewing data; n is the total number of data; 、 and It is the threshold used to divide the strength of user preference association, which corresponds to the low threshold for screening the loosely associated user group, the medium threshold for screening the moderately associated user group, and the high threshold for screening the strictly associated user group; for each threshold Will As a preliminary user group segmentation; It is based on the threshold The preliminary user group division is obtained; the preliminary user group sets under all thresholds are combined to obtain the full set of preliminary user groups; through the preliminary user group division, user groups with similar viewing preference patterns can be preliminarily identified, providing a basis for further refined analysis.

[0044] Example 4, see Figure 1This embodiment is based on the above embodiment. In step S3, the preliminary user group is refined. In personalized video recommendation, the correlation distribution between user video viewing data and preference feedback data of different users is quite different. Due to the different content tones and user group characteristics of different content channels, the correlation of user viewing data in the same content channel is often high. Users of the same education channel are more inclined to watch similar courses, while the correlation between different content channels may be very low. The viewing preferences of film and television users and game live broadcast users are significantly different. In order to avoid the recommendation bias and waste of computing resources caused by a single granularity, for each preliminary user group, each user's video viewing data is initially regarded as a user segment group. The two user segments with the smallest distance are gradually merged until the maximum number of merges is reached or the merged distance is lower than the merge threshold. Candidate user segments are obtained. The average correlation is calculated for each candidate user segment group, which is expressed as: ;in, is the i-th candidate user segment The average correlation degree of g and k is the index of user video viewing data; is the Jaccard similarity between the video viewing data of the gth user and the video viewing data of the kth user; and based on the average correlation, the candidate user segments are divided into high correlation, medium correlation and low correlation; for the candidate user segments, if all user video viewing data belong to the same content channel and the candidate user segment is highly correlated, then it is retained. The candidate user segments under the , fine division; if all user video viewing data do not belong to the same content channel and the candidate user segments are low correlation, retain The candidate user segments below should be avoided to avoid being too detailed; the rest Candidate user segments under the following conditions; for channels with highly relevant content, user preference segmentation is carried out to accurately capture user segmentation preferences; and emphasis is placed on the migration characteristics of user preferences and the impact of niche interests.

[0045] By performing the above operations, we can address the problems of general personalized video recommendation systems that cannot simultaneously take into account preference differences within and across content channels, have high operation and maintenance costs, and are easily misled by user behavior noise, resulting in low recommendation accuracy. This solution integrates multi-threshold preliminary user groups, enriches the candidate user group structure, adaptively retains the optimal user preference segmentation of each content channel, and enhances the ability to capture users' true preferences. By refining the preliminary user groups within the user viewing behavior data, it highlights the migration path of user preferences, thereby improving the accuracy of personalized video recommendations.

[0046] Example 5, see Figure 1This embodiment is based on the above embodiment. In step S4, the final user segmentation is performed by screening all candidate user segmentation groups through a genetic algorithm, specifically including the following steps:

[0047] Step S41: Cluster coding; for the candidate user segment group set ;in, and The first and second user segments; let each candidate user segment correspond to a binary vector, the binary vector elements Expressed as: ; Chromosome length Boolean vector of Select clusters as the final partition; among them, and are the first and second Boolean vectors respectively. elements; Indicates the selection of candidate user segments ; Randomly generate N chromosomes;

[0048] Step S42: Multi-objective fitness evaluation; define the multi-objective functions to minimize the similarity between clusters of user video viewing data. , maximize the intra-cluster density of user video viewing data and minimize the number of clusters of user video viewing data ; are expressed as follows: If there is a high degree of similarity between different user clusters, it means that the group boundaries are fuzzy, making it difficult to push differentiated content to different clusters. We encourage low coupling between clusters to ensure the uniqueness of the recommendation strategy. The viewing behavior of users within the same user cluster should be as similar as possible, so that they can push video content with the same preferences to users within the cluster. The higher the cohesion, the better the recommendation consistency among users within the cluster. ;in, is the jth candidate user segment; is the Jaccard similarity between the video viewing data of the kth user and the video viewing data of the lth user; X is the set of candidate user segmentation groups selected in the current solution; over-segmented clusters will increase the computational cost and strategy complexity of the recommendation system; therefore, while ensuring the accuracy of recommendations, the number of user clusters must be controlled;

[0049] Step S43: Crossover operation; use a multi-objective solution ranking algorithm to assign an optimal solution level to each solution in the population; for solutions within the same level, calculate the distribution density based on the normalized distance of each target in that level; obtain the ranking; and use a roulette wheel to select N parents to retain. For each parent, decide whether to perform a crossover operation based on the crossover probability; the crossover operation is as follows: randomly select a cluster subset S, copy all gene positions corresponding to S in parent 1 to offspring 1, and fill the remaining positions in the order of parent 2; the same process is repeated for offspring 2; the initial crossover probability is set in the range [0.6, 0.95];

[0050] Step S44: Neighborhood mutation operation: For each offspring, a neighborhood mechanism is introduced to select gene positions based on neighborhood mutation probability, and a right shift operation is performed to generate a neighborhood solution; the optimal solution is selected from all non-taboo neighborhoods and the current chromosome is updated; the used moves are recorded in the search taboo record, and the number of iterations is set; the initial neighborhood mutation probability is set in the range of [0.01, 0.05];

[0051] Step S45: Conflict handling: For each new offspring, if there exists g such that , then there is a conflict. For each conflict point, keep one and close the other based on fitness. If there exists g such that , then the coverage point is missing, and the missing point is remedied to construct the submatrix ; Solve the linear equations to obtain the auxiliary Boolean vector y, select the cluster completion in the least incremental way, let when ;in, is the i-th element of the Boolean vector; is the i-th element of the auxiliary Boolean vector; select the one with the highest diversity in the population according to the Hamming distance chromosomes, and run a neighborhood search again;

[0052] Step S46: Update crossover mutation probability; calculate intermediate quantity: ; ; After fuzzy calculation is clarified, we get and ,renew: ; ;in, and are the minimum fitness value and average fitness value of all individuals in the t-th generation population respectively; is the average fitness value of all individuals in the t-1 generation population; and are the crossover probabilities of the tth generation and the t-1th generation respectively; and are the neighborhood mutation probabilities of the tth and t-1th generations, respectively; and are the increments of the crossover probability and neighborhood mutation probability of the tth generation, respectively; and is the intermediate quantity;

[0053] Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the final selected candidate user segmentation group as the final user segmentation group.

[0054] By performing the above operations, we can address the problem that general personalized video recommendation systems are insufficiently sensitive to user preference mutation patterns and have poor adaptability to real-time changes in user preferences, resulting in poor personalized video recommendation results. This solution introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, and adopts a multi-objective optimization strategy to improve sensitivity to potential preference changes. It also designs a mechanism for dynamically adjusting crossover probability and neighborhood mutation probability, automatically adjusts the exploration intensity of user group segmentation, and enhances the adaptability of the entire recommendation system to dynamic changes in user preferences, thereby improving the effectiveness of personalized video recommendations.

[0055] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S5, personalized video recommendation is based on the end user segmentation result, and highly relevant videos watched by users in the same cluster are directly recommended to users in the cluster, so as to achieve real-time similar preference matching based on clusters; for each end user segment, the video collection of all users in the cluster is extracted, and after merging and removing duplicates, a cluster-related video pool is obtained; the video relevance within the user group of each video in the video pool is calculated. , expressed as: ;according to Sort in descending order, take the first N R videos as candidate recommendation videos; for each user in the cluster, the videos that the user has watched are removed from the candidate recommendation videos to avoid repeated recommendations; finally, the top K videos after sorting are pushed to the user.

[0056] Example 7, see Figure 2 This embodiment is based on the above embodiment. The present invention provides a personalized video recommendation system based on big data, which includes a data acquisition module, a preliminary user group full set division module, a preliminary user group refinement module, a final user sub-group division module and a personalized video recommendation module;

[0057] The data acquisition module collects user video viewing data;

[0058] The preliminary user group full set division module divides the user video viewing data into preliminary user groups to obtain a user preference equivalent association matrix, and divides the user video viewing data into connected areas based on different thresholds to form preliminary user group full sets of user video viewing data;

[0059] The preliminary user group refinement module further refines the preliminary user group of the user video watching data into candidate user subgroups of high, medium and low correlation degrees by merging user subgroups;

[0060] The final user subgroup division module obtains final user subgroup division by designing cross operation and neighborhood mutation operation, and screening the candidate user subgroups of the user video watching data by using genetic algorithm.

[0061] The personalized video recommendation module performs personalized video recommendation based on the final user subgroup division.

[0062] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between such entities or operations. Moreover, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not expressly listed or inherent to such process, method, article or apparatus.

[0063] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made thereto without departing from the principles and spirit of the present application.

[0064] The above describes the present application and its embodiments, which are not restrictive, and the embodiments shown in the drawings are only one of the embodiments of the present application, and the actual structure is not limited thereto. In general, if a person skilled in the art is inspired by it, without departing from the principles and spirit of the present application, without creative design, similar structure and embodiments of the technical solution can be designed, which should belong to the protection scope of the present application.

Claims

1. A personalized video recommendation method based on big data, characterized by: The method comprises the following steps: Step S1: Data collection: Collect user video viewing data; Step S2: preliminary division of the entire user group; Step S3: Refine the preliminary user group: by merging user segments, refine the preliminary user group of user video viewing data into candidate user segments; Step S4: final user segmentation: by designing crossover operations and neighborhood mutation operations, using genetic algorithms to screen candidate user segments of user video viewing data, and obtain final user segmentation; Step S5: Personalized video recommendation: Personalized video recommendation is performed based on the end user segmentation.

2. The method for personalized video recommendation based on big data according to claim 1, characterized in that: In step S2, the preliminary user group full set division is to obtain the user viewing behavior similarity matrix R by the Jaccard similarity of the video viewing data of two users, which is expressed as: ; and get the distance matrix D, expressed as: ; Perform indirect association operation on user preference on R to obtain user preference equivalent association matrix ; Take the threshold set ;in, is the Jaccard similarity between the u-th user video viewing data and the g-th user video viewing data; n is the total number of data; 、 and is the threshold used to divide the strength of user preference association; for each threshold Will As a preliminary user group segmentation; It is based on the threshold The obtained preliminary user groups are divided; the preliminary user group sets under all thresholds are combined to obtain the complete set of preliminary user groups.

3. The method for personalized video recommendation based on big data according to claim 2, characterized in that: In step S3, the preliminary user group refinement is to initially treat each user's video viewing data as a user segment within each preliminary user group; and gradually merge the two user segments with the smallest distance; Until the maximum number of merges is reached, or the merge distance is lower than the merge threshold; Get candidate user segments; calculate the average relevance for each candidate user segment, expressed as: ;in, is the i-th candidate user segment The average correlation degree of g and k is the index of user video viewing data; is the Jaccard similarity between the video viewing data of the gth user and the video viewing data of the kth user; and based on the average correlation, the candidate user segments are divided into high correlation, medium correlation and low correlation; for the candidate user segments, if all user video viewing data belong to the same content channel and the candidate user segment is highly correlated, then it is retained. If all user video viewing data do not belong to the same content channel and the candidate user segment is low-correlated, retain The candidate user segments below; the rest The candidate user segments below.

4. The method for personalized video recommendation based on big data according to claim 3, characterized in that: In step S4, the final user segmentation is performed by screening all candidate user segment groups using a genetic algorithm, which specifically includes the following steps: Step S41: Cluster coding; for the candidate user segment group set ;in, and The first and second User segments; let each candidate user segment correspond to a binary vector, the binary vector elements Expressed as: ; Chromosome length Boolean vector of Select clusters as the final partition; among them, and are the first and second Boolean vectors respectively. elements; Indicates the selection of candidate user segments ; Randomly generate N chromosomes; Step S42: Multi-objective fitness evaluation; define the multi-objective functions to minimize the similarity between clusters of user video viewing data. , maximize the intra-cluster density of user video viewing data and minimize the number of clusters of user video viewing data ; are expressed as follows: ; ; ;in, is the jth candidate user segment; is the Jaccard similarity between the video viewing data of the kth user and the video viewing data of the lth user; X is the set of candidate user segments selected in the current solution; Step S43: Crossover operation: Use a multi-objective solution ranking algorithm to assign an optimal solution level to each solution in the population; for solutions within the same level, calculate the distribution density based on the normalized distance of each target in that level; obtain the ranking; and use a roulette wheel to select and retain N parents. For each parent team, choose whether to perform a crossover operation based on the crossover probability; the crossover operation is as follows: randomly select a cluster subset S, copy all gene positions corresponding to S in parent 1 to offspring 1, and fill the remaining positions in the order of parent 2; Step S44: Neighborhood mutation operation: For each offspring, a neighborhood mechanism is introduced to select a gene position based on the neighborhood mutation probability, and a right shift operation is performed to generate a neighborhood solution; the optimal solution is selected from all non-taboo neighborhoods and the current chromosome is updated; the used moves are recorded in the search taboo record, and the number of iterations is set. Step S45: conflict handling; Step S46: crossover mutation probability update; Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the final selected candidate user segmentation group as the final user segmentation group.

5. The method for personalized video recommendation based on big data according to claim 4, characterized in that: In step S4, the conflict handling is performed on each new offspring, if there is g such that , then there is a conflict. For each conflict point, keep one and close the other based on fitness. If there exists g such that , then the coverage point is missing, and the missing point is remedied to construct the submatrix ; Solve the linear equations to obtain the auxiliary Boolean vector y, select the cluster completion in the least incremental way, and let when ;in, is the i-th element of the Boolean vector; is the i-th element of the auxiliary Boolean vector; Select the population with the highest diversity according to the Hamming distance chromosomes and run another neighborhood search.

6. The method for personalized video recommendation based on big data according to claim 5, characterized in that: In step S4, the crossover probability update is to calculate the intermediate quantity: ; ; After fuzzy calculation is clarified, we get and ,renew: ; ;in, and are the minimum fitness value and average fitness value of all individuals in the t-th generation population respectively; is the average fitness value of all individuals in the t-1 generation population; and are the crossover probabilities of the tth generation and the t-1th generation respectively; and are the neighborhood mutation probabilities of the tth and t-1th generations, respectively; and are the increments of the crossover probability and neighborhood mutation probability of the tth generation, respectively; and It is an intermediate amount.

7. The method for personalized video recommendation based on big data according to claim 6, characterized in that: In step S5, the personalized video recommendation is based on the end user segmentation result, and directly recommends highly relevant videos watched by users in the cluster to users in the cluster, thereby achieving real-time similar preference matching in cluster units.

8. A personalized video recommendation system based on big data, for implementing the personalized video recommendation method based on big data as described in any one of claims 1 to 7, characterized in that: It includes data collection module, preliminary user group full set division module, preliminary user group refinement module, final user segmentation module and personalized video recommendation module; The data acquisition module collects user video viewing data; The preliminary user group full set division module divides the user video viewing data into preliminary user groups to obtain a user preference equivalent association matrix, and divides the user video viewing data into connected areas based on different thresholds to form preliminary user group full sets of user video viewing data; The preliminary user group refinement module further refines the preliminary user group of user video viewing data into candidate user segments with high, medium and low correlation by merging user segments; The final user segmentation module screens candidate user segments of user video viewing data using a genetic algorithm by designing a crossover operation and a neighborhood mutation operation to obtain the final user segmentation; The personalized video recommendation module performs personalized video recommendation based on the end user segmentation.

Citation Information

Patent Citations

  • Video recommendation method and device based on user group behavioral analysis

    CN103731738A

  • Video content recommending method, device and system

    CN105915949A

  • Evaluation optimization method and system based on user joint preference

    CN117056613A

  • Power marketing recommendation method and system based on big data

    CN118967211A

  • E-commerce personalized recommendation service method based on genetic fuzzy clustering

    CN119090588A