User Preference Mining Method, Device, Storage Medium and Computer Equipment

By constructing content-tagged and user-viewing matrices and applying modified TF-IDF with time decay, IPTV systems can accurately profile user preferences, enhancing personalized content delivery.

CN115866344BActive Publication Date: 2025-07-15E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211501370.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-07-15
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing IPTV operations cannot accurately explore users, resulting in the inability to achieve refined operations for each user.

Method used

By constructing the program tag matrix and user viewing program matrix, combining the TF-IDF algorithm, time decay algorithm and duration influence factor, the user's preference value for content tags is determined, and then the user's preference tag is determined.

Benefits of technology

It improves the accuracy of user preference mining and improves the reach rate of content recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115866344B_ABST
    Figure CN115866344B_ABST
Patent Text Reader

Abstract

A user preference mining method provided by this application includes: obtaining program data, matching corresponding content tags to each program in the program data, and constructing a program tag matrix based on each content tag; obtaining the viewing data of users, and constructing a user viewing program matrix according to the user information and viewing program information in the viewing data; obtaining a target matrix based on the program tag matrix and the user viewing program matrix, and determining target users; based on the target matrix, determining the IDF values of each content tag, as well as the TF values, time decay coefficients, and duration influence factors of the target users on each content tag; determining the preference values of the target users for each content tag based on each IDF value, TF value, time decay coefficient, and duration influence factor, and determining the preference tags corresponding to the target users. By means of an improved TF-IDF algorithm combined with a time decay algorithm and a duration influence factor, the accuracy of user preference mining can be improved, and the reach rate of content recommendation can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and particularly to a method, apparatus, storage medium, and computer device for mining user preferences. Background Art

[0002] In the era of continuous development of information technology, in order to provide users with a better program viewing experience, IPTV has emerged as the times require. IPTV, namely Internet Protocol Television, is a technology that integrates Internet, multimedia, and communication technologies and uses broadband networks as a medium to provide various interactive services including digital television to home users. IPTV operation can further meet the personalized needs of users to enjoy video programs.

[0003] Existing IPTV operations only simply group IPTV users or use user preference mining methods in other industries, and cannot accurately mine the preferences of IPTV users, thus unable to achieve refined operation for each IPTV user. Summary of the Invention

[0004] The purpose of this application aims to solve at least one of the above technical defects, especially the technical defect that it is impossible to achieve refined operation for each IPTV user in the prior art.

[0005] In a first aspect, this application provides a method for mining user preferences, and the method includes:

[0006] Obtain program data, match each program in the program data with a corresponding content label, and construct a program label matrix according to each content label;

[0007] Obtain the viewing data of users, and construct a user viewing program matrix according to the user information and viewing program information in the viewing data;

[0008] Obtain a target matrix based on the program label matrix and the user viewing program matrix, and determine target users according to the target matrix;

[0009] Based on the target matrix, determine the inverse document frequency index IDF value of each content label, and determine the term frequency index TF value, time decay coefficient, and duration influence factor of the target user on each content label;

[0010] Determine the preference value of the target user for each content label based on each IDF value, the TF value, the time decay coefficient, and the duration influence factor;

[0011] Determine the preference label corresponding to the target user according to each preference value.

[0012] In one embodiment, the viewing program information in the user viewing program matrix includes the program viewed by the user, the viewing duration of the program, the total duration of the program, and the viewing time of the program each time. The step of obtaining the target matrix based on the program tag matrix and the user viewing program matrix includes:

[0013] According to the program tag matrix, match the corresponding content tags for the programs viewed by the user in the user viewing program matrix to obtain the target matrix.

[0014] In one embodiment, the step of determining the IDF value of each content tag based on the target matrix includes:

[0015] Based on the target matrix, determine the number of users corresponding to each content tag and the total number of users in the target matrix;

[0016] Determine the IDF value of each content tag according to the number of users corresponding to each content tag and the total number of users.

[0017] In one embodiment, the step of determining the TF value of the target user on each content tag based on the target matrix includes:

[0018] Based on the target matrix, for each content tag, determine the total number of first viewing times and the total first viewing duration of the program corresponding to the content tag by the target user, and the total first program duration of the program corresponding to the content tag;

[0019] For each content tag, determine a first value according to the total number of first viewing times, the total first viewing duration, and the total first program duration;

[0020] Based on the target matrix, determine the total number of second viewing times and the total second viewing duration of the programs by the target user, and the total second program duration of the programs corresponding to the target user;

[0021] Determine a second value according to the total number of second viewing times, the total second viewing duration, and the total second program duration;

[0022] For each content tag, determine the TF value according to the first value and the second value.

[0023] In one embodiment, the step of determining the time decay coefficient of the target user on each content tag based on the target matrix includes:

[0024] Based on the target matrix, for each of the content tags, determine the time difference between the first viewing time and the second viewing time of the program corresponding to the content tag for the target user;

[0025] For each of the content tags, determine the time decay coefficient according to a preset constant and the time difference.

[0026] In one embodiment, the step of determining the duration influence factor of the target user on each of the content tags based on the target matrix includes:

[0027] Based on the target matrix, for each of the content tags, determine the total target viewing duration of the program corresponding to the content tag for the target user;

[0028] For each of the content tags, determine the duration influence factor according to a preset value and the total target viewing duration.

[0029] In one embodiment, the step of determining the preference tag corresponding to the target user according to each of the preference values includes:

[0030] Sort the content tags corresponding to each of the preference values according to the numerical magnitudes of the preference values to obtain a sorting result;

[0031] Filter the content tags in the sorting result according to a preset selection number;

[0032] Use the filtered content tags as the preference tags of the target user.

[0033] In a second aspect, the present application provides a device for mining user preferences, including:

[0034] A program tag matrix construction module, configured to obtain program data, match a corresponding content tag for each program in the program data, and construct a program tag matrix according to each content tag;

[0035] A user viewing program matrix construction module, configured to obtain user viewing data, and construct a user viewing program matrix according to the user information and viewing program information in the viewing data;

[0036] A target matrix acquisition module, configured to obtain a target matrix based on the program tag matrix and the user viewing program matrix, and determine a target user according to the target matrix;

[0037] A value determination module, configured to determine the inverse document frequency index IDF value of each content tag based on the target matrix, and determine the term frequency index TF value, time decay coefficient, and duration influence factor of the target user on each content tag;

[0038] A preference value determination module, configured to determine the preference values of the target user for each of the content tags based on each of the IDF values, the TF values, the time decay coefficient, and the duration influence factor;

[0039] A preference tag determination module, configured to determine the preference tags corresponding to the target user according to each of the preference values.

[0040] In a third aspect, the present application provides a storage medium, characterized in that: computer-readable instructions are stored in the storage medium, and when the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the user preference mining method in any of the above embodiments.

[0041] In a fourth aspect, the present application provides a computer device, characterized by including: one or more processors, and a memory;

[0042] Computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the steps of the user preference mining method in any of the above embodiments are executed.

[0043] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0044] A user preference mining method provided by the present application includes: obtaining program data, matching each program in the program data with a corresponding content tag, and constructing a program tag matrix according to each content tag; obtaining user viewing data, and constructing a user viewing program matrix according to the user information and viewing program information in the viewing data; obtaining a target matrix based on the program tag matrix and the user viewing program matrix, and determining a target user; determining the IDF value of each content tag, and the TF value, time decay coefficient, and duration influence factor of the target user on each content tag based on the target matrix; determining the preference values of the target user for each content tag based on each IDF value, TF value, time decay coefficient, and duration influence factor, and determining the preference tags corresponding to the target user. By improving the TF-IDF algorithm and combining the time decay algorithm and the duration influence factor, the accuracy of user preference mining can be improved, and the reach rate of content recommendation can be enhanced. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 A flowchart of a user preference mining method provided by an embodiment of the present application;

[0047] Figure 2 A schematic diagram of the implementation process of a user preference mining method provided by an embodiment of the present application;

[0048] Figure 3 A schematic diagram of the structure of a user preference mining device provided by an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the internal structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0051] In one embodiment, the present application provides a user preference mining method. The following embodiments will be described by taking the application of this method to a server as an example. It can be understood that the server executing the user preference mining method can be a single server or a server cluster composed of multiple servers. The present application does not make specific limitations in this regard. As Figure 1 shown, the method specifically includes the following steps:

[0052] S101: Obtain program data, match corresponding content tags to each program in the program data, and construct a program tag matrix according to each content tag.

[0053] In this step, program data is obtained, and type tagging processing is performed on the program data, that is, corresponding content tags are matched to each program in the program data, so that a program tag matrix can be obtained.

[0054] Among them, each program can match one or more content tags, and the content tags corresponding to different programs can be the same. For example, the content tags matched by program 1 are tag 1, the content tags matched by program 2 are tag 2 and tag 3, and the content tags matched by program 3 are tag 1 and tag 3.

[0055] S102: Obtain the viewing data of the user, and construct a user viewing program matrix according to the user information and viewing program information in the viewing data.

[0056] In this step, obtain the viewing data of the user, clean the viewing data, filter out invalid data, and then construct a user viewing program matrix based on the user information and viewing program information in the viewing data.

[0057] Among them, the user viewing program matrix can be an m*n matrix. In the user viewing program matrix, the user information corresponds to the viewing program information, that is, the programs viewed by the user and the corresponding program information can be obtained from this matrix.

[0058] S103: Obtain a target matrix based on the program tag matrix and the user viewing program matrix, and determine the target user according to the target matrix.

[0059] In this step, the program tag matrix and the user viewing program matrix are associated to obtain a target matrix, and the target user can be determined from the user information in the target matrix.

[0060] Among them, since each program does not contain all content tags, the obtained target matrix is a sparse matrix. The target user is any one of the users in the target matrix.

[0061] S104: Based on the target matrix, determine the inverse document frequency index IDF value of each content tag, and determine the term frequency index TF value, time decay coefficient, and duration influence factor of the target user on each content tag.

[0062] In this step, according to the user information and program information in the target matrix, calculate the IDF value of each content tag, and at the same time calculate the TF value, time decay coefficient, and duration influence factor of the target user on each content tag.

[0063] Among them, the inverse document frequency index IDF value refers to a measure of the general importance of a content tag, the term frequency index TF value refers to the frequency of a content tag appearing in this matrix, the time decay factor characterizes the degree of decay of the user's interest in the content tag over time, and the duration influence factor characterizes the degree of change of the user's interest in the content tag with the viewing duration.

[0064] S105: Determine the preference value of the target user for each content tag based on each IDF value, TF value, time decay coefficient, and duration influence factor.

[0065] In this step, the TF-IDF algorithm integrates the time decay coefficient and the duration influence factor to calculate the preference value of the target user for each content tag.

[0066] Further, in one embodiment, the product of the TF value, IDF value, time decay factor, and duration impact factor can be calculated, and the product value can be used as the preference value of the target user for the content label.

[0067] S106: Determine the preference labels corresponding to the target user according to each of the preference values.

[0068] In this step, after obtaining the preference values of the target user for each content label, based on the numerical magnitudes of each preference value, multiple content labels are selected as the preference labels of the target user according to the selection rules.

[0069] Further, for each user in the target matrix, after determining the preference values of the user for each content label, a user preference matrix can be obtained. In this matrix, according to the numerical magnitudes of each preference value, the preference labels of each user can be determined.

[0070] A user preference mining method provided by this application includes: obtaining program data, matching corresponding content labels to each program in the program data, and constructing a program label matrix according to each content label; obtaining the viewing data of users, and constructing a user viewing program matrix according to the user information and viewing program information in the viewing data; obtaining a target matrix based on the program label matrix and the user viewing program matrix, and determining the target user; determining the IDF value of each content label based on the target matrix, as well as the TF value, time decay coefficient, and duration impact factor of the target user for each content label; determining the preference value of the target user for each content label based on each IDF value, TF value, time decay coefficient, and duration impact factor, and determining the preference labels corresponding to the target user. By improving the TF-IDF algorithm and combining the time decay algorithm and the duration impact factor, the accuracy of user preference mining can be improved, and the reach rate of content recommendation can be enhanced.

[0071] In one embodiment, the viewing program information in the user viewing program matrix includes the program viewed by the user, the viewing duration of the program, the total duration of the program, and the viewing time of the program each time. The step of obtaining the target matrix based on the program label matrix and the user viewing program matrix includes:

[0072] According to the program label matrix, match the corresponding content labels to the programs viewed by the user in the user viewing program matrix to obtain the target matrix.

[0073] Specifically, the program label matrix includes the watched programs and the content labels of the watched programs. The user watched program matrix includes user information, the programs watched by the user, the viewing duration of the programs, the total duration of the programs, and the viewing time of each program. Therefore, based on the watched programs, the program label matrix and the user watched program matrix are associated, and the programs watched by the user in the user watched program matrix are matched with the corresponding content labels to obtain the target matrix.

[0074] For example, if the user watched program matrix A is an m*5 matrix, where A1 represents the primary key of the user, A2 represents the program watched by the user, A3 represents the viewing duration of the program, A4 represents the total duration of the program, and A5 represents the viewing time of each program. The target matrix P obtained by associating the program label matrix and the user watched program matrix A based on the watched programs is an m*n matrix, where P1 represents the primary key of the user, P2 represents the program watched by the user, P3 to P(n - 3) represent the corresponding content labels of the program, P(n - 2) represents the viewing duration of the program, P(n - 1) represents the total duration of the program, and Pn represents the viewing time of each program.

[0075] It can be understood that since there is a situation where the program content watched by most users is relatively concentrated, therefore, tagging the programs can, to a certain extent, solve the problem of data sparsity, thereby improving the accuracy of mining user preferences.

[0076] In one embodiment, the step of determining the IDF value of each of the content labels based on the target matrix includes:

[0077] Based on the target matrix, determine the number of users corresponding to each content label and the total number of users in the target matrix;

[0078] Determine the IDF value of each content label according to the number of users corresponding to each content label and the total number of users.

[0079] Specifically, determine the total number of users and the number of users corresponding to each content label from the target matrix, and calculate and determine the IDF value of each content label through a formula.

[0080] Further, for each content label, use the sum of the number of users corresponding to the content label and a preset value as the first sum, use the sum of the total number of users and the preset value as the second sum, use the quotient of the first sum and the second sum as the true number, take the logarithm of the true number with base 10 as the IDF value of the content label. Among them, the preset value can be 1.

[0081] For example, if calculating the IDF value for content label j, the formula is specifically:

[0082]

[0083] In one embodiment, the step of determining the TF value of the target user on each of the content tags based on the target matrix includes:

[0084] Based on the target matrix, for each of the content tags, determine the first total viewing times and the first total viewing duration of the programs corresponding to the content tag by the target user, and the first total program duration of the programs corresponding to the content tag;

[0085] For each of the content tags, determine a first value according to the first total viewing times, the first total viewing duration, and the first total program duration;

[0086] Based on the target matrix, determine the second total viewing times and the second total viewing duration of the target user on each of the programs, and the second total program duration of the programs corresponding to the target user;

[0087] Determine a second value according to the second total viewing times, the second total viewing duration, and the second total program duration;

[0088] For each of the content tags, determine the TF value according to the first value and the second value.

[0089] Specifically, based on the target matrix, for each content tag, determine the first total viewing times and the first total viewing duration of the programs corresponding to the content tag by the target user, and the first total program duration of the programs corresponding to the content tag. Multiply the ratio of the first total viewing duration to the first total program duration by the first total viewing times to obtain the first value. Based on the target matrix, determine the second total viewing times and the second total viewing duration of the target user on each of the programs, and the second total program duration of the programs corresponding to the target user. Multiply the ratio of the second total viewing duration to the second total program duration by the second total viewing times to obtain the second value. Take the ratio of the first value to the second value to obtain the TF value.

[0090] For example, if calculating the TF value for content tag j, the formula is specifically:

[0091]

[0092] It can be understood that, in combination with the actual situation, the problems faced when only considering the number of times alone include that the number of viewing times of the programs corresponding to a certain content tag is large, but the viewing duration each time is very short, or the number of viewing times of the programs corresponding to a certain content tag is small, but the viewing duration each time is very long. By improving the calculation method of the TF value, the problem of being unable to accurately mine the user's preferences can be effectively solved.

[0093] In one embodiment, the step of determining the time decay coefficient of the target user on each content tag based on the target matrix includes:

[0094] Based on the target matrix, for each content tag, determine the time difference between the first viewing time and the second viewing time of the program corresponding to the content tag by the target user;

[0095] For each content tag, determine the time decay coefficient according to a preset constant and the time difference.

[0096] Specifically, based on the target matrix, for each content tag, determine the time difference between the first viewing time and the second viewing time of the program corresponding to the content tag by the target user, take the negative value of the product of the time difference and the preset constant as the exponent, and take e as the base to perform a power operation to obtain the time decay coefficient.

[0097] Wherein, the first viewing time is the most recent time when the target user watched the program corresponding to the content tag, and the second viewing time is the previous viewing time of the first viewing time.

[0098] For example, if calculating the time decay coefficient for content tag j, the formula is specifically:

[0099] α(j) = e -factor*Δt

[0100] Wherein, factor is the preset constant, and Δt is the time difference between the first viewing time and the second viewing time.

[0101] In one embodiment, the step of determining the duration influence factor of the target user on each content tag based on the target matrix includes:

[0102] Based on the target matrix, for each content tag, determine the target total viewing duration of the program corresponding to the content tag by the target user;

[0103] For each content tag, determine the duration influence factor according to a preset value and the target total viewing duration.

[0104] Specifically, based on the target matrix, for each content tag, determine the target total viewing duration of the program corresponding to the content tag by the target user, take the logarithm of the target total viewing duration with a preset value as the base as the duration influence factor. Wherein, the preset value can be 1.2

[0105] For example, if calculating the duration influence factor for content tag j, the formula is specifically:

[0106] β(j) = log 1.2 sum(j)

[0107] In one embodiment, the step of determining the preference tags corresponding to the target user according to each of the preference values includes:

[0108] Sort the content tags corresponding to each of the preference values according to the numerical magnitudes of the preference values to obtain a sorting result;

[0109] Filter the content tags in the sorting result according to a preset selection number;

[0110] Use the filtered content tags as the preference tags of the target user.

[0111] Specifically, sort the content tags corresponding to each preference value from largest to smallest according to the numerical magnitudes of the preference values to obtain a sorting result, filter the content tags in the sorting result according to a preset selection number, select the first one or more content tags in the sorting result, and use the filtered content tags as the preference tags of the target user.

[0112] Among them, the preset selection number is a positive integer greater than or equal to 1, and the specific value can be determined according to the actual situation.

[0113] It can be understood that the content in the target matrix is automatically updated, and the preference values of the user on each content tag can be updated in real time, so as to ensure the real-time performance and accuracy of the data. Then, the preference tags of the user are determined according to the preference values, further accurately mining the preferences of the user, reducing manual operation, and improving the reach rate of content recommendation.

[0114] In an example, as Figure 2 shown, the user preference mining method provided by this application may include the following steps:

[0115] S201: Obtain a program data set, match corresponding content tags for each program in the program data set, and construct a program tag matrix T according to each content tag;

[0116] S202: Obtain a user viewing data set, and construct a user viewing program matrix A according to the user information and viewing program information in the viewing data;

[0117] S203: Obtain a target matrix P based on the program tag matrix T and the user viewing program matrix A;

[0118] S204: Based on the target matrix P, calculate the TF value, IDF value, time decay factor α, and duration preference coefficient β;

[0119] S205: Determine the user tag preference value according to the TF value, IDF value, time decay factor α, and duration preference coefficient β;

[0120] S206: Select the tags with higher preference values as the user's preference tags.

[0121] The user preference mining device provided by the embodiments of the present application will be described below. The user preference mining device described below can be correspondingly referred to the user preference mining method described above.

[0122] In one embodiment, the present application provides a user preference mining device. As Figure 3 shown, the device specifically includes a program tag matrix construction module 301, a user viewing program matrix construction module 302, a target matrix acquisition module 303, a value determination module 304, a preference value determination module 305, and a preference tag determination module 306, where:

[0123] The program tag matrix construction module 301 is configured to obtain program data, match corresponding content tags for each program in the program data, and construct a program tag matrix according to each content tag;

[0124] The user viewing program matrix construction module 302 is configured to obtain user viewing data, and construct a user viewing program matrix according to the user information and viewing program information in the viewing data;

[0125] The target matrix acquisition module 303 is configured to obtain a target matrix based on the program tag matrix and the user viewing program matrix, and determine a target user according to the target matrix;

[0126] The value determination module 304 is configured to determine the inverse document frequency index IDF value of each content tag based on the target matrix, and determine the term frequency index TF value, time decay coefficient, and duration impact factor of the target user on each content tag;

[0127] The preference value determination module 305 is configured to determine the preference value of the target user for each content tag based on each IDF value, the TF value, the time decay coefficient, and the duration impact factor;

[0128] The preference tag determination module 306 is configured to determine the preference tags corresponding to the target user according to each preference value.

[0129] In one embodiment, the viewing program information in the user viewing program matrix includes the program viewed by the user, the viewing duration of the program, the total duration of the program, and the viewing time of the program each time. The target matrix acquisition module 303 includes a content tag matching unit, where:

[0130] A content label matching unit, configured to match corresponding content labels for the programs watched by the user in the user program watching matrix according to the program label matrix, so as to obtain the target matrix.

[0131] In one embodiment, the numerical value determination module 304 includes a user number determination unit and an IDF value determination unit, where:

[0132] The user number determination unit is configured to determine, based on the target matrix, the number of users corresponding to each content label, and the total number of users in the target matrix;

[0133] The IDF value determination unit is configured to determine the IDF value of each content label according to the number of users corresponding to each content label and the total number of users.

[0134] In one embodiment, the numerical value determination module 304 includes a first viewing numerical value determination unit, a first numerical value determination unit, a second viewing numerical value determination unit, a second numerical value determination unit, and a TF value determination unit, where:

[0135] The first viewing numerical value determination unit is configured to determine, based on the target matrix, for each content label, the total number of first viewings and the total first viewing duration of the program corresponding to the content label by the target user, and the total first program duration of the program corresponding to the content label;

[0136] The first numerical value determination unit is configured to determine, for each content label, a first numerical value according to the total number of first viewings, the total first viewing duration, and the total first program duration;

[0137] The second viewing numerical value determination unit is configured to determine, based on the target matrix, the total number of second viewings and the total second viewing duration of the target user for each program, and the total second program duration of the program corresponding to the target user;

[0138] The second numerical value determination unit is configured to determine a second numerical value according to the total number of second viewings, the total second viewing duration, and the total second program duration;

[0139] The TF value determination unit is configured to determine the TF value for each content label according to the first numerical value and the second numerical value.

[0140] In one embodiment, the numerical value determination module 304 includes a time difference determination unit and a time decay coefficient determination unit, where:

[0141] The time difference determination unit is configured to determine, based on the target matrix, for each content label, the time difference between the first viewing time and the second viewing time of the program corresponding to the content label by the target user;

[0142] A time decay coefficient determination unit, configured to determine the time decay coefficient for each of the content tags according to a preset constant and the time difference.

[0143] In one embodiment, the numerical value determination module 304 includes a target total viewing duration determination unit and a duration influence factor determination unit, where:

[0144] The target total viewing duration determination unit is configured to determine, for each of the content tags, the target total viewing duration of the target user for the program corresponding to the content tag based on the target matrix;

[0145] The duration influence factor determination unit is configured to determine, for each of the content tags, the duration influence factor according to a preset value and the target total viewing duration.

[0146] In one embodiment, the preference tag determination module 306 includes a content tag sorting unit, a content tag screening unit, and a preference tag determination unit, where:

[0147] The content tag sorting unit is configured to sort the content tags corresponding to the respective preference values according to the numerical magnitudes of the respective preference values to obtain a sorting result;

[0148] The content tag screening unit is configured to screen the content tags in the sorting result according to a preset selection number;

[0149] The preference tag determination unit is configured to use the screened content tags as the preference tags of the target user.

[0150] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the user preference mining method according to any one of the above embodiments.

[0151] In one embodiment, the present application further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the user preference mining method according to any one of the above embodiments.

[0152] Schematically, as Figure 4 shown, Figure 4 is an internal structural schematic diagram of a computer device provided by an embodiment of the present application. The computer device 400 can be provided as a server. Referring to Figure 4, the computer device 400 includes a processing component 402, which further includes one or more processors, and memory resources represented by a memory 401 for storing instructions executable by the processing component 402, such as application programs. The application programs stored in the memory 401 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 402 is configured to execute instructions to perform the user preference mining method of any of the above embodiments.

[0153] The computer device 400 may further include a power component 403 configured to perform power management of the computer device 400, a wired or wireless network interface 404 configured to connect the computer device 400 to a network, and an input / output (I / O) interface 405. The computer device 400 may operate based on an operating system stored in the memory 401, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM, or the like.

[0154] Those skilled in the art can understand that Figure 4 the structure shown in

[0155] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0156] The user preference mining device provided by this application can be applied to the IPTV operation field. For example, it can be applied to public Internet platforms such as Internet TVs to provide certain personalized needs for users to enjoy video programs. When using this application, an indication sign can be set up in a specific application scenario to remind users that when using the user preference mining device of this application, the viewing data of users will be collected. Furthermore, with the authorization of users, the viewing data of users can be legally collected to implement the user preference mining device of this application.

[0157] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0158] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0159] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A user preference mining method, characterized in that, The method includes: Obtaining program data, matching corresponding content tags to each program in the program data, and constructing a program tag matrix according to each of the content tags; Obtaining the viewing data of the user, and constructing a user viewing program matrix according to the user information and viewing program information in the viewing data, where the viewing program information includes the program viewed by the user, the viewing duration of the program, the total duration of the program, and the viewing time of the program each time; According to the program tag matrix, matching corresponding content tags to the programs viewed by the user in the user viewing program matrix to obtain a target matrix, and determining a target user according to the target matrix; Based on the target matrix, determining the inverse document frequency index IDF value of each of the content tags, and determining the term frequency index TF value, time decay coefficient, and duration influence factor of the target user on each of the content tags; Among them, the step of determining the TF value of the target user on each of the content tags based on the target matrix includes: Based on the target matrix, for each of the content tags, determining the first total viewing times and the first total viewing duration of the target user for the program corresponding to the content tag, and the first total duration of the program corresponding to the content tag; For each of the content tags, determining a first value according to the first total viewing times, the first total viewing duration, and the first total duration of the program; Based on the target matrix, determining the second total viewing times and the second total viewing duration of the target user for each of the programs, and the second total duration of the program corresponding to the target user; Determining a second value according to the second total viewing times, the second total viewing duration, and the second total duration of the program; For each of the content tags, determining the TF value according to the first value and the second value; Determining the preference value of the target user for each of the content tags based on each of the IDF values, the TF values, the time decay coefficient, and the duration influence factor; Determining the preference tags corresponding to the target user according to each of the preference values.

2. The user preference mining method according to claim 1, wherein The step of determining the IDF value of each of the content tags based on the target matrix includes: Based on the target matrix, determining the number of users corresponding to each of the content tags, and the total number of users in the target matrix; Determining the IDF value of each of the content tags according to the number of users corresponding to each of the content tags and the total number of users.

3. The user preference mining method according to claim 1, wherein The step of determining the time decay coefficient of the target user on each of the content tags based on the target matrix includes: Based on the target matrix, for each of the content tags, determining the time difference between the first viewing time and the second viewing time of the target user for the program corresponding to the content tag; For each of the content tags, determining the time decay coefficient according to a preset constant and the time difference.

4. The user preference mining method according to claim 1, wherein The step of determining the duration influence factor of the target user on each of the content tags based on the target matrix includes: Based on the target matrix, for each of the content tags, determine the total target viewing duration of the program corresponding to the content tag for the target user; For each of the content tags, determine the duration influence factor according to a preset value and the total target viewing duration.

5. The user preference mining method according to claim 1, wherein The step of determining the preference tags corresponding to the target user according to each of the preference values includes: Sort the content tags corresponding to each of the preference values according to the numerical magnitudes of the preference values to obtain a sorting result; Screen the content tags in the sorting result according to a preset selection number; Use the screened content tags as the preference tags of the target user.

6. A user preference mining device, characterized in that, including: A program tag matrix construction module, configured to obtain program data, match corresponding content tags to each program in the program data, and construct a program tag matrix according to each of the content tags; A user viewing program matrix construction module, configured to obtain the viewing data of the user, and construct a user viewing program matrix according to the user information and the viewing program information in the viewing data, where the viewing program information includes the program viewed by the user, the viewing duration of the program, the total duration of the program, and the viewing time of the program each time; A target matrix obtaining module, configured to match corresponding content tags to the programs viewed by the user in the user viewing program matrix according to the program tag matrix to obtain a target matrix, and determine a target user according to the target matrix; A numerical value determination module, configured to determine the inverse document frequency index IDF value of each of the content tags based on the target matrix, and determine the term frequency index TF value, the time decay coefficient, and the duration influence factor of the target user on each of the content tags; wherein, the numerical value determination module includes: A first viewing numerical value determination unit, configured to determine, based on the target matrix, for each of the content tags, the total first viewing times and the total first viewing duration of the program corresponding to the content tag for the target user, and the total first program duration of the program corresponding to the content tag; A first numerical value determination unit, configured to determine a first numerical value for each of the content tags according to the total first viewing times, the total first viewing duration, and the total first program duration; A second viewing numerical value determination unit, configured to determine, based on the target matrix, the total second viewing times and the total second viewing duration of the target user for each of the programs, and the total second program duration of the programs corresponding to the target user; A second numerical value determination unit, configured to determine a second numerical value according to the total second viewing times, the total second viewing duration, and the total second program duration; A TF value determination unit, configured to determine the TF value for each of the content tags according to the first numerical value and the second numerical value; A preference value determination module, configured to determine the preference value of the target user for each of the content tags based on each of the IDF values, the TF values, the time decay coefficient, and the duration influence factor; A preference tag determination module, configured to determine the preference tags corresponding to the target user according to each of the preference values.

7. A storage medium, characterized in that: The storage medium stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to execute the steps of the user preference mining method according to any one of claims 1 to 5.

8. A computer device, characterized in that, Comprising: One or more processors, and a memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, execute the steps of the user preference mining method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • User content portrait determination method, access object recommendation method and related devices

    CN110209875A

  • User portrait generation method and device based on applet game, equipment and medium

    CN113780415A