Personalized Video Recommendation Methods and Systems Based on Big Data

By integrating a preliminary user group with multiple thresholds and optimizing the genetic algorithm, the problem of insufficient adaptability to preference differences and user behavior noise in personalized video recommendation systems is solved, achieving higher recommendation accuracy and dynamic adaptability.

CN120769084BActive Publication Date: 2025-12-02CHINA UNICOM VIDEO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511283904.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-12-02
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing personalized video recommendation systems cannot simultaneously take into account the differences in preferences within and across content channels. These systems have high maintenance costs, are easily misled by user behavior noise, have low recommendation accuracy, and are not sensitive enough to sudden changes in user preferences, resulting in poor adaptability and unsatisfactory recommendation performance.

Method used

By integrating preliminary user groups with multiple thresholds, the structure of candidate user groups is enriched. The optimal user preference segmentation of each content channel is adaptively preserved. An optimized genetic algorithm is introduced to optimize the user segmentation of user viewing behavior data. A mechanism for dynamically adjusting crossover probability and neighborhood mutation probability is designed to enhance the adaptability to dynamic changes in user preferences.

Benefits of technology

It improves the accuracy and adaptability of personalized video recommendations, enhances the ability to capture users' true preferences and the sensitivity to potential changes in preferences, and improves the dynamic adaptability of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769084B_ABST
    Figure CN120769084B_ABST
Patent Text Reader

Abstract

This invention discloses a personalized video recommendation method and system based on big data. The method includes data collection, initial user group segmentation, initial user group refinement, final user segmentation, and personalized video recommendation. This invention belongs to the field of user recommendation technology, specifically referring to a personalized video recommendation method and system based on big data. This solution enriches the candidate user group structure through multi-threshold initial user group integration, adaptively retains the optimal user preference segmentation of each content channel, and enhances the ability to capture users' true preferences. It highlights the migration path of user preferences by refining the initial user group within user viewing behavior data; it introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, and adopts a multi-objective optimization strategy to improve sensitivity to potential preference changes; it designs a mechanism for dynamically adjusting crossover probability and neighborhood mutation probability; thereby improving the personalized video recommendation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of user recommendation technology, specifically to a personalized video recommendation method and system based on big data. Background Technology

[0002] Personalized video recommendation methods are based on user viewing behavior data. By analyzing user preference similarities, they segment and optimize user groups, accurately pushing video content that matches their shared preferences to different user groups. However, typical personalized video recommendation systems suffer from several drawbacks: they cannot simultaneously account for preference differences within and across content channels; they have high maintenance costs; they are easily misled by user behavior noise, leading to low recommendation accuracy; and they generally lack sensitivity to sudden changes in user preferences and are poorly adaptable to real-time changes in user preferences, resulting in unsatisfactory personalized video recommendation performance. Summary of the Invention

[0003] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides a personalized video recommendation method and system based on big data. Addressing the problems of general personalized video recommendation systems, such as their inability to simultaneously consider preference differences within and across content channels, high system maintenance costs, susceptibility to user behavior noise leading to low recommendation accuracy, this solution employs multi-threshold initial user group integration to enrich the candidate user group structure, adaptively retaining the optimal user preference segmentation for each content channel, and enhancing the ability to capture true user preferences. Furthermore, by refining the initial user group segmentation based on user viewing behavior data, it highlights the migration path of user preferences, thereby improving the accuracy of personalized video recommendations. To address the issues of insufficient sensitivity to sudden changes in user preferences and poor adaptability to real-time user preference changes in general personalized video recommendation systems, resulting in poor personalized video recommendation performance, this solution introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, employs a multi-objective optimization strategy to improve sensitivity to potential preference changes, and designs a mechanism for dynamically adjusting crossover probability and neighborhood mutation probability to automatically adjust the exploration intensity of user group segmentation, enhancing the overall recommendation system's adaptability to dynamic changes in user preferences, thereby improving the performance of personalized video recommendations.

[0004] The technical solution adopted by this invention is as follows: The personalized video recommendation method based on big data provided by this invention includes the following steps:

[0005] Step S1: Data Acquisition;

[0006] Step S2: Preliminary segmentation of the entire user base;

[0007] Step S3: Initial user group refinement;

[0008] Step S4: Final user segmentation;

[0009] Step S5: Personalized video recommendations.

[0010] Furthermore, in step S1, the data acquisition involves collecting user video viewing data and performing feature encoding and standardization processing on the user video viewing data.

[0011] Further, in step S2, the initial user group segmentation is achieved by obtaining a user viewing behavior similarity matrix R based on the Jaccard similarity of the video viewing data of two users, expressed as: And the distance matrix D is obtained, which is represented as: Perform indirect correlation operations on user preferences on R to obtain the user preference equivalence matrix. Take the threshold set ;in, is the Jaccard similarity between the video viewing data of user u and user g; n is the total number of data points; , and It is a threshold used to classify the strength of association between user preferences; for each threshold Will As an initial user group segmentation; Based on threshold The initial user group segmentation is obtained; the initial user group sets under all thresholds are merged to obtain the complete initial user group set.

[0012] Further, in step S3, the initial user group refinement involves, within each initial user group, initially treating each user's video viewing data as a separate user segment; gradually merging the two user segments with the smallest distance; until the maximum number of merges is reached, or the merge distance is lower than the merge threshold; obtaining candidate user segments; and calculating the average correlation degree for each candidate user segment, expressed as: ;in, It is the i-th candidate user segment The average correlation coefficient; g and k are indexes of user video viewing data; The Jaccard similarity is calculated between the video viewing data of the g-th user and the video viewing data of the k-th user. Based on the average correlation, candidate user segments are divided into high-correlation, medium-correlation, and low-correlation groups. For each candidate user segment, if all user video viewing data belongs to the same content channel and the candidate user segment is highly correlated, then the segment is retained. The following candidate user segments; if all user video viewing data do not belong to the same content channel and the candidate user segments have low correlation, retain them. The following are candidate user segments; the rest are selected. The candidate user segments below.

[0013] Further, in step S4, the final user segmentation is achieved by screening all candidate user segments using a genetic algorithm, specifically including the following steps:

[0014] Step S41: Cluster coding; for the candidate user segment set ;in, and They are the 1st and the 2nd respectively Each candidate user segment is assigned a binary vector, and the elements of the binary vector are... Represented as: Chromosome length Boolean vector Clusters were selected as the final partitions; among them, and These are the first and second Boolean vectors, respectively. One element; Indicates the selection of candidate user segments Randomly generate N chromosomes;

[0015] Step S42: Multi-objective fitness evaluation; Define the multi-objective functions as minimizing the inter-cluster similarity of user video viewing data. Maximize the intra-cluster density of user video viewing data and minimizing the number of clusters in user video viewing data ; are represented as follows: ; ; ;in, It is the j-th candidate user segment; X is the Jaccard similarity between the video viewing data of the k-th user and the video viewing data of the l-th user; X is the set of candidate user segments selected in the current solution.

[0016] Step S43: Crossover operation; Use a multi-objective scheme ranking algorithm to assign the optimal scheme level to each solution in the population; For solutions within the same level, calculate the distribution density based on the normalized distance of each objective in that level; Obtain the ranking; Use roulette wheel selection to retain N parent generations; For each parent generation, select whether to perform crossover operation based on the crossover probability; The crossover operation is as follows: Randomly select a subset S, copy all gene positions corresponding to S in parent generation 1 to child generation 1, and fill the remaining positions according to the order of parent generation 2;

[0017] Step S44: Neighborhood mutation operation; For each offspring, introduce a neighborhood mechanism, select the gene position with the neighborhood mutation probability, perform a right shift operation to generate a neighborhood solution; Select the optimal solution from all non-taboo neighborhoods and update the current chromosome; Record the used moves in the search taboo record and iterate until the set number of times;

[0018] Step S45: Conflict resolution; For each new offspring, if there exists g such that... Then, for each conflict point, based on fitness, keep one and close the other; if there exists g such that If there are missing covered points, perform missing point repair and construct a submatrix. Solving the system of linear equations yields an auxiliary Boolean vector y. Cluster completion is then selected using the least increment method, letting... when ;in, It is the i-th element of the Boolean vector; It is the i-th element of the auxiliary Boolean vector; the population with the highest diversity is selected according to the Hamming distance. For each chromosome, run the neighborhood search again;

[0019] Step S46: Update crossover and mutation probabilities; calculate intermediate quantities: ; After fuzzy computation is clarified, the result is... and ,renew: ; ;in, and These are the minimum fitness value and the average fitness value of all individuals in the t-th generation population, respectively. It is the average fitness value of all individuals in the (t-1)th generation population; and These are the crossover probabilities for generation t and generation (t-1), respectively. and These are the neighborhood mutation probabilities for generation t and generation (t-1), respectively. and These are the increments of the crossover probability and the neighborhood mutation probability in the t-th generation, respectively; and It is an intermediate quantity;

[0020] Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the finally selected candidate user subgroups as the final user subgroup segmentation.

[0021] Furthermore, in step S5, the personalized video recommendation is based on the final user segmentation results, and directly recommends highly relevant videos that users in the same cluster watch together, thereby achieving real-time similarity preference matching on a cluster-by-cluster basis.

[0022] The personalized video recommendation system based on big data provided by this invention includes a data acquisition module, a preliminary user group full set segmentation module, a preliminary user group refinement module, a final user segmentation module, and a personalized video recommendation module.

[0023] The data acquisition module collects user video viewing data;

[0024] The preliminary user group full set segmentation module segments the user video viewing data into a preliminary user group full set, obtains the user preference equivalent correlation matrix, and divides the user video viewing data into connection regions based on different thresholds to form a preliminary user group full set of user video viewing data.

[0025] The preliminary user group refinement module further refines the preliminary user group based on user video viewing data into candidate user segments with high, medium, and low relevance by merging user subgroups.

[0026] The final user segmentation module uses a genetic algorithm to filter candidate user segments from user video viewing data by designing crossover and neighborhood mutation operations, thereby obtaining the final user segmentation.

[0027] The personalized video recommendation module makes personalized video recommendations based on the segmentation of end-user groups.

[0028] The beneficial effects achieved by the present invention using the above solution are as follows:

[0029] (1) To address the problems of general personalized video recommendation systems being unable to simultaneously consider the differences in preferences within and across content channels, resulting in high system operation and maintenance costs and susceptibility to user behavior noise leading to low recommendation accuracy, this solution enriches the candidate user group structure by integrating a preliminary user group with multiple thresholds, adaptively retaining the optimal user preference segmentation of each content channel, and enhancing the ability to capture users' true preferences; by refining the preliminary user group within the user viewing behavior data, the migration path of user preferences is highlighted; thereby improving the accuracy of personalized video recommendations.

[0030] (2) To address the problem that general personalized video recommendation systems are not sensitive enough to sudden changes in user preferences and have poor adaptability to real-time changes in user preferences, resulting in poor personalized video recommendation performance, this solution introduces an optimized genetic algorithm to optimize the user segmentation of user viewing behavior data, adopts a multi-objective optimization strategy to improve the sensitivity to potential changes in preferences, and designs a mechanism to dynamically adjust the crossover probability and neighborhood mutation probability to automatically adjust the exploration intensity of user group segmentation, thereby enhancing the overall recommendation system's adaptability to dynamic changes in user preferences and thus improving the personalized video recommendation effect. Attached Figure Description

[0031] Figure 1 A flowchart illustrating the personalized video recommendation method based on big data provided by this invention;

[0032] Figure 2 This is a schematic diagram of the personalized video recommendation system based on big data provided by the present invention.

[0033] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0034] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0035] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the system or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0036] Example 1, see Figure 1 The present invention provides a personalized video recommendation method based on big data, which includes the following steps:

[0037] Step S1: Data Acquisition; Collect user video viewing data;

[0038] Step S2: Preliminary segmentation of the entire user base;

[0039] Step S3: Initial User Group Refinement; By merging user segments, the initial user group based on user video viewing data is further refined into candidate user segments with high, medium, and low relevance.

[0040] Step S4: Final user segmentation; By designing crossover and neighborhood mutation operations, a genetic algorithm is used to screen candidate user segments based on user video viewing data to obtain the final user segmentation.

[0041] Step S5: Personalized video recommendation; Personalized video recommendations are made based on end-user segmentation.

[0042] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, data collection involves collecting user video viewing data. The user video viewing data includes user viewing time window, single video viewing duration, viewing frequency, viewing cycle, number of video types viewed by the user, cumulative viewing days, and video category ratio. The user video viewing data is then subjected to feature encoding and standardization processing.

[0043] Example 3, see Figure 1 This embodiment is based on the above embodiment. In step S2, the preliminary user group segmentation is performed because, in personalized video recommendation, there may be correlations in the video viewing preferences of different users. These correlations can be discovered through the preliminary user group segmentation. The user viewing behavior similarity matrix R is obtained by using the Jaccard similarity of the video viewing data of two users, as follows: And the distance matrix D is obtained, which is represented as: Perform indirect correlation operations on user preferences on R to obtain the user preference equivalence matrix. Take the threshold set ;in, is the Jaccard similarity between the video viewing data of user u and user g; n is the total number of data points; , and These are thresholds used to classify the strength of user preference associations: a low threshold for filtering users with loose associations, a medium threshold for filtering users with moderate associations, and a high threshold for filtering users with strict associations; for each threshold... Will As an initial user group segmentation; Based on threshold The initial user group segmentation is obtained; the initial user group sets under all thresholds are merged to obtain the initial user group complete set; through the initial user group segmentation, user groups with similar viewing preference patterns can be initially identified, providing a basis for further refined analysis.

[0044] Example 4, see Figure 1This embodiment is based on the above embodiment. In step S3, the initial user group refinement is carried out in personalized video recommendation, where the correlation distribution between user video viewing data and preference feedback data varies greatly among different users. Due to the different content tone and user group characteristics of different content channels, the correlation of user viewing data within the same content channel is often high. Users of the same educational channel tend to watch similar courses, while the correlation between different content channels may be very low, and the viewing preferences of film and television users and game live streaming users differ significantly. To avoid the recommendation bias and waste of computing resources caused by a single granularity, for each initial user group, each user's video viewing data is initially treated as a separate user segment. The two user segments with the smallest distance are gradually merged until the maximum number of merges is reached, or the merging distance is lower than the merging threshold, to obtain candidate user segments. The average correlation is calculated for each candidate user segment, expressed as: ;in, It is the i-th candidate user segment The average correlation coefficient; g and k are indexes of user video viewing data; The Jaccard similarity is calculated between the video viewing data of the g-th user and the video viewing data of the k-th user. Based on the average correlation, candidate user segments are divided into high-correlation, medium-correlation, and low-correlation groups. For each candidate user segment, if all user video viewing data belongs to the same content channel and the candidate user segment is highly correlated, then the segment is retained. The candidate user segments are further refined; if all users' video viewing data do not belong to the same content channel and the candidate user segments have low correlation, they are retained. The candidate user segments should be further segmented to avoid over-segmentation; the rest should be selected. The candidate user segments are identified; for channels with highly relevant content, user preferences are segmented to accurately capture user segment preferences; the migration characteristics of user preferences and the influence of niche interests are highlighted.

[0045] By performing the above operations, this solution addresses the problems of general personalized video recommendation systems, such as the inability to simultaneously consider preference differences within and across content channels, high system maintenance costs, susceptibility to user behavior noise, and low recommendation accuracy. It enriches the candidate user group structure through multi-threshold initial user group integration, adaptively retains the optimal user preference segmentation for each content channel, and enhances the ability to capture users' true preferences. Furthermore, by refining the initial user group segmentation based on user viewing behavior data, it highlights the migration path of user preferences, thereby improving the accuracy of personalized video recommendations.

[0046] Example 5, see Figure 1This embodiment is based on the above embodiment. In step S4, the final user segmentation is achieved by screening all candidate user segments using a genetic algorithm, specifically including the following steps:

[0047] Step S41: Cluster coding; for the candidate user segment set ;in, and They are the 1st and the 2nd respectively Each candidate user segment is assigned a binary vector, and the elements of the binary vector are... Represented as: Chromosome length Boolean vector Clusters were selected as the final partitions; among them, and These are the first and second Boolean vectors, respectively. One element; Indicates the selection of candidate user segments Randomly generate N chromosomes;

[0048] Step S42: Multi-objective fitness evaluation; Define the multi-objective functions as minimizing the inter-cluster similarity of user video viewing data. Maximize the intra-cluster density of user video viewing data and minimizing the number of clusters in user video viewing data ; are represented as follows: If there is a high degree of similarity between different user clusters, it indicates that the grouping boundaries are blurred, making it difficult to push differentiated content to different clusters; low coupling between clusters is encouraged to ensure the uniqueness of the recommendation strategy. It is required that the viewing behavior of users within the same user cluster be as similar as possible, so as to facilitate the recommendation of video content with the same preferences to users within the cluster; the higher the cohesion, the better the recommendation consistency among users within the cluster. ;in, It is the j-th candidate user segment; X is the Jaccard similarity between the video viewing data of the k-th user and the video viewing data of the l-th user; X is the set of candidate user subgroups selected in the current solution; excessively subdivided clusters will increase the computational cost and strategy complexity of the recommendation system; therefore, while ensuring the accuracy of the recommendation, the number of user clusters needs to be controlled.

[0049] Step S43: Crossover operation; Use a multi-objective scheme ranking algorithm to assign the optimal scheme level to each solution in the population; For solutions within the same level, calculate the distribution density based on the normalized distance of each objective in that level; Obtain the ranking; Use roulette wheel selection to retain N parent generations. For each parent generation, select whether to perform crossover operation based on the crossover probability; The crossover operation is as follows: Randomly select a subset S, copy all gene positions corresponding to S in parent generation 1 to child generation 1, and fill the remaining positions according to the order of parent generation 2; The same applies to child generation 2; The initial crossover probability is set within the range of [0.6, 0.95];

[0050] Step S44: Neighborhood mutation operation; For each offspring, introduce a neighborhood mechanism, select the gene position with the neighborhood mutation probability, perform a right shift operation to generate a neighborhood solution; Select the optimal solution from all non-taboo neighborhoods and update the current chromosome; Record the used moves in the search taboo record and iterate until the set number of times; The initial neighborhood mutation probability is set within the range of [0.01, 0.05];

[0051] Step S45: Conflict resolution; For each new offspring, if there exists g such that... Then, for each conflict point, based on fitness, keep one and close the other; if there exists g such that If there are missing covered points, perform missing point repair and construct a submatrix. Solving the system of linear equations yields an auxiliary Boolean vector y. Cluster completion is then selected using the least increment method, letting... when ;in, It is the i-th element of the Boolean vector; It is the i-th element of the auxiliary Boolean vector; the population with the highest diversity is selected according to the Hamming distance. For each chromosome, run the neighborhood search again;

[0052] Step S46: Update crossover and mutation probabilities; calculate intermediate quantities: ; After fuzzy computation is clarified, the result is... and ,renew: ; ;in, and These are the minimum fitness value and the average fitness value of all individuals in the t-th generation population, respectively. It is the average fitness value of all individuals in the (t-1)th generation population; and These are the crossover probabilities for generation t and generation (t-1), respectively. and These are the neighborhood mutation probabilities for generation t and generation (t-1), respectively. and These are the increments of the crossover probability and the neighborhood mutation probability in the t-th generation, respectively; and It is an intermediate quantity;

[0053] Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the finally selected candidate user subgroups as the final user subgroup segmentation.

[0054] By performing the above operations, this solution addresses the problem of general personalized video recommendation systems lacking sensitivity to sudden changes in user preferences and exhibiting poor adaptability to real-time changes in user preferences, resulting in unsatisfactory personalized video recommendation performance. It introduces an optimized genetic algorithm to refine user segmentation based on viewing behavior data, employs a multi-objective optimization strategy to improve sensitivity to potential preference changes, and designs a mechanism for dynamically adjusting crossover and neighborhood mutation probabilities to automatically adjust the exploration intensity of user segmentation, thereby enhancing the overall recommendation system's adaptability to dynamic changes in user preferences and ultimately improving the effectiveness of personalized video recommendations.

[0055] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S5, personalized video recommendation is based on the final user segmentation results, directly recommending highly relevant videos that users in the same cluster have watched together, realizing real-time similarity preference matching on a cluster-by-cluster basis; for each final user segmentation, the set of videos watched by all users in the cluster is extracted, merged and deduplicated to obtain a cluster-related video pool; the video relevance within the user group of each video in the video pool is calculated. , represented as: ;according to Sort in descending order and take the top N. R The first K videos are selected as candidate recommendation videos; for each user in the cluster, videos that they have already watched are removed from the candidate recommendation videos to avoid duplicate recommendations; finally, the top K videos after sorting are pushed to the user.

[0056] Example 7, see Figure 2 Based on the above embodiments, the personalized video recommendation system based on big data provided by the present invention includes a data acquisition module, a preliminary user group full set segmentation module, a preliminary user group refinement module, a final user segmentation module, and a personalized video recommendation module.

[0057] The data acquisition module collects user video viewing data;

[0058] The preliminary user group full set segmentation module segments the user video viewing data into a preliminary user group full set, obtains the user preference equivalent correlation matrix, and divides the user video viewing data into connection regions based on different thresholds to form a preliminary user group full set of user video viewing data.

[0059] The preliminary user group refinement module further refines the preliminary user group based on user video viewing data into candidate user segments with high, medium, and low relevance by merging user subgroups.

[0060] The final user segmentation module uses a genetic algorithm to filter candidate user segments from user video viewing data by designing crossover and neighborhood mutation operations, thereby obtaining the final user segmentation.

[0061] The personalized video recommendation module makes personalized video recommendations based on the segmentation of end-user groups.

[0062] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0063] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0064] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A personalized video recommendation method based on big data, characterized by: The method includes the following steps: Step S1: Data Acquisition; Collect user video viewing data; Step S2: Preliminary segmentation of the entire user base; Step S3: Initial User Group Refinement; By merging user segments, the initial user group based on user video viewing data is refined into candidate user segments; Step S4: Final user segmentation; By designing crossover and neighborhood mutation operations, a genetic algorithm is used to screen candidate user segments based on user video viewing data to obtain the final user segmentation. Step S5: Personalized video recommendation; Personalized video recommendations are made based on end-user segmentation. In step S2, the initial user group partitioning is achieved by obtaining a user viewing behavior similarity matrix R based on the Jaccard similarity of video viewing data between two users, expressed as: And the distance matrix D is obtained, which is represented as: Perform indirect correlation operations on user preferences on R to obtain the user preference equivalence matrix. Take the threshold set ;in, is the Jaccard similarity between the video viewing data of the u-th user and the video viewing data of the g-th user; n is the total number of data points; , and It is a threshold used to classify the strength of association between user preferences; for each threshold Will As an initial user group segmentation; Based on threshold The initial user group segmentation is obtained; the initial user group sets under all thresholds are merged to obtain the complete initial user group set; In step S3, the initial user group refinement involves treating each user's video viewing data as a separate user segment within each initial user group; gradually merging the two user segments with the smallest distance; until the maximum number of merges is reached, or the merge distance is below the merge threshold; resulting in candidate user segments; and calculating the average correlation degree for each candidate user segment, expressed as: ;in, It is the i-th candidate user segment. The average correlation coefficient; g and k are indexes of user video viewing data; The Jaccard similarity is calculated between the video viewing data of the g-th user and the video viewing data of the k-th user. Based on the average correlation, candidate user segments are divided into high-correlation, medium-correlation, and low-correlation groups. For each candidate user segment, if all user video viewing data belongs to the same content channel and the candidate user segment is highly correlated, then the segment is retained. The following candidate user segments; if all user video viewing data do not belong to the same content channel and the candidate user segments have low correlation, retain them. The following are candidate user segments; the rest are selected. The following candidate user segments; In step S4, the final user segmentation is achieved by screening all candidate user segments using a genetic algorithm, specifically including the following steps: Step S41: Cluster coding; for the candidate user segment set ;in, and They are the 1st and the 2nd respectively Each candidate user segment is assigned a binary vector, and the elements of the binary vector are... Represented as: Chromosome length Boolean vector Clusters were selected as the final partitions; among them, and These are the first and second Boolean vectors, respectively. One element; Indicates the selection of candidate user segments Randomly generate N chromosomes; Step S42: Multi-objective fitness evaluation; Define the multi-objective functions as minimizing the inter-cluster similarity of user video viewing data. Maximize the intra-cluster density of user video viewing data and minimizing the number of clusters in user video viewing data ; are represented as follows: ; ; ;in, It is the j-th candidate user segment; X is the Jaccard similarity between the video viewing data of the k-th user and the video viewing data of the l-th user; X is the set of candidate user segments selected in the current solution. Step S43: Crossover operation; Use a multi-objective scheme ranking algorithm to assign the optimal scheme level to each solution in the population; For solutions within the same level, calculate the distribution density based on the normalized distance of each objective in that level; Obtain the ranking; Use roulette wheel selection to retain N parent generations; For each parent generation, select whether to perform crossover operation based on the crossover probability; The crossover operation is as follows: Randomly select a subset S, copy all gene positions corresponding to S in parent generation 1 to child generation 1, and fill the remaining positions according to the order of parent generation 2; Step S44: Neighborhood mutation operation; For each offspring, introduce a neighborhood mechanism, select the gene position with the neighborhood mutation probability, perform a right shift operation to generate a neighborhood solution; Select the optimal solution from all non-taboo neighborhoods and update the current chromosome; Record the used moves in the search taboo record and iterate until the set number of times; Step S45: Conflict resolution; Step S46: Update crossover and mutation probabilities; Step S47: Determine; when the maximum number of iterations is reached, output the optimal user group segmentation scheme; obtain the finally selected candidate user subgroups as the final user subgroup segmentation.

2. The personalized video recommendation method based on big data according to claim 1, characterized in that: In step S4, the conflict handling is performed on each new offspring if there exists a g such that Then, for each conflict point, based on fitness, keep one and close the other; if there exists g such that If there are missing covered points, perform missing point repair and construct a submatrix. ; Solving the system of linear equations yields an auxiliary Boolean vector y. Cluster completion is then selected using the least incremental approach, letting... when ;in, It is the i-th element of the Boolean vector; It is the i-th element of the auxiliary Boolean vector; Select the population with the highest diversity based on Hamming distance. For each chromosome, run the neighborhood search again.

3. The personalized video recommendation method based on big data according to claim 2, characterized in that: In step S4, the crossover mutation probability update involves calculating intermediate quantities: ; After fuzzy computation is clarified, the result is... and ,renew: ; ;in, and These are the minimum fitness value and the average fitness value of all individuals in the t-th generation population, respectively. It is the average fitness value of all individuals in the (t-1)th generation population; and These are the crossover probabilities for generation t and generation (t-1), respectively. and These are the neighborhood mutation probabilities for generation t and generation (t-1), respectively. and These are the increments of the crossover probability and the neighborhood mutation probability in the t-th generation, respectively; and It is an intermediate quantity.

4. The personalized video recommendation method based on big data according to claim 3, characterized in that: In step S5, the personalized video recommendation is based on the final user segmentation results, and directly recommends highly relevant videos that users in the same cluster watch together, thereby achieving real-time similarity preference matching on a cluster-by-cluster basis.

5. A personalized video recommendation system based on big data, used to implement the personalized video recommendation method based on big data as described in any one of claims 1-4, characterized in that: It includes a data acquisition module, a preliminary user group segmentation module, a preliminary user group refinement module, a final user segmentation module, and a personalized video recommendation module; The data acquisition module collects user video viewing data; The preliminary user group full set segmentation module segments the user video viewing data into a preliminary user group full set, obtains the user preference equivalent correlation matrix, and divides the user video viewing data into connection regions based on different thresholds to form a preliminary user group full set of user video viewing data. The preliminary user group refinement module further refines the preliminary user group based on user video viewing data into candidate user segments with high, medium, and low relevance by merging user subgroups. The final user segmentation module uses a genetic algorithm to filter candidate user segments from user video viewing data by designing crossover and neighborhood mutation operations, thereby obtaining the final user segmentation. The personalized video recommendation module makes personalized video recommendations based on the segmentation of end-user groups.

Citation Information

Patent Citations

  • Video content recommending method, device and system

    CN105915949A

  • E-commerce personalized recommendation service method based on genetic fuzzy clustering

    CN119090588A