Family structure member identification method, device, electronic device and storage medium

By analyzing multi-dimensional features and behaviors based on the media library and combining the TextRank and Embedding algorithms to enrich media tags, the problem of inaccurate identification of family members is solved, achieving more efficient and accurate identification of family users.

CN116541585BActive Publication Date: 2025-09-26CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210087423.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-09-26
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing technology is not accurate in identifying family members and relies on expert experience, resulting in low recognition efficiency and poor accuracy.

Method used

Based on the media library of family structure members, family structure members are identified by extracting multi-dimensional features such as viewing loyalty, viewing intensity and usage ability characteristics, combining media tags and preset thresholds, and using TextRank and Embedding algorithms to enrich the media tag system, build a family structure media library, and combine on-demand and live broadcast behaviors to identify family structures.

Benefits of technology

The accuracy and efficiency of family structure member identification are improved, the inaccuracy of expert experience identification is avoided, and more accurate family user identification is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116541585B_ABST
    Figure CN116541585B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device, and storage medium for identifying family structure members. The method includes: determining a first parameter of each family user among multiple family users to be identified based on a media library of family structure members; the first parameter represents the ratio of viewing time of a specific media asset associated with the media library to total viewing time; the media library includes family structure members and media asset tags; identifying family structure members of a first family user based on the first parameter and a preset threshold; the first family user is at least one family user among the multiple family users; extracting multi-dimensional features based on the media library, and identifying family structure members of other family users to be identified based on the multi-dimensional features combined with the family structure members of the first family user; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for identifying family members. Background Art

[0002] In the related technologies, the identification of family members of home broadband users is based on expert experience, which has the problem of inaccurate identification. Summary of the Invention

[0003] To solve related technical problems, embodiments of the present application provide a method, device, electronic device, and storage medium for identifying family structure members.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present application provides a method for identifying family members, including:

[0006] Determining a first parameter of each of a plurality of family users to be identified based on a media library of family members; the first parameter representing a ratio of viewing time of a specific media asset associated with the media library to total viewing time; the media library including family members and media asset tags;

[0007] identifying family members of a first family user according to the first parameter and a preset threshold; the first family user being at least one family user among the plurality of family users;

[0008] Multi-dimensional features are extracted based on the media library, and family structure members of other family users to be identified are identified based on the multi-dimensional features combined with the family structure members of the first family user; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user.

[0009] In the above solution, the method further includes:

[0010] Extracting label data corresponding to the media data watched by each family user;

[0011] The media resource library is updated using the family structure members corresponding to each family user and the tag data.

[0012] In the above solution, the step of extracting the tag data corresponding to the media data watched by each family user includes:

[0013] Obtaining a brief introduction to the media data watched by each household user;

[0014] Key information is extracted from the introduction, and the tag data is determined according to the key information.

[0015] In the above solution, the step of extracting the tag data corresponding to the media data watched by each family user includes:

[0016] Obtaining attribute information of the media data watched by each family user;

[0017] Performing cluster analysis on the attribute information to obtain words that represent common attributes of the media asset data;

[0018] The words are used as the label data.

[0019] In the above solution, the method further includes:

[0020] Obtaining a storage method of the family structure members and the media asset tags;

[0021] The media resource library is constructed according to the data stored in the storage method.

[0022] In the above solution, the multi-dimensional features include at least: viewing loyalty features, viewing intensity features, and usage capability features; and the multi-dimensional features extracted based on the media resource library include:

[0023] Obtaining on-demand behavior data of the family members based on the media resource library;

[0024] The viewing loyalty feature, the viewing intensity feature, and the usage capability feature are extracted according to the on-demand behavior data.

[0025] In the above solution, extracting the viewing loyalty feature based on the on-demand behavior data includes:

[0026] Obtaining, based on the on-demand behavior data, a unit length of viewing time and an average unit length of interval time of the on-demand viewing of the family members within a first preset period;

[0027] Determining a maximum on-demand viewing time unit length and a maximum on-demand viewing interval time unit length among the family members;

[0028] The viewing loyalty feature is determined according to the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length.

[0029] In the above solution, determining the viewing loyalty feature according to the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length includes:

[0030] Determining a first ratio of the viewing time unit length to the maximum on-demand viewing time unit length and a second ratio of the average interval time unit length to the maximum on-demand interval time unit length;

[0031] Obtaining a first weight coefficient for the first ratio and a second weight coefficient for the second ratio;

[0032] The viewing loyalty feature is determined based on the first ratio, the second ratio, the first weight coefficient, and the second weight coefficient.

[0033] In the above solution, extracting the viewing intensity feature based on the on-demand behavior data includes:

[0034] Obtaining the frequency and average duration of playback of the program requested by the family members within a second preset period based on the on-demand behavior data;

[0035] Determine the maximum on-demand playback frequency and maximum on-demand playback duration among the family members;

[0036] The viewing intensity feature is determined according to the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration.

[0037] In the above solution, determining the viewing intensity feature based on the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration includes:

[0038] Determining a third ratio of the playback frequency to the maximum on-demand playback frequency and a fourth ratio of the average playback duration to the maximum on-demand playback duration;

[0039] obtaining a third weight coefficient for the third ratio and a fourth weight coefficient for the fourth ratio;

[0040] The viewing intensity feature is determined based on the third ratio, the fourth ratio, the third weight coefficient, and the fourth weight coefficient.

[0041] In the above solution, extracting the usage capability feature based on the on-demand behavior data includes:

[0042] Determining the number of collections, searches, and payments for on-demand playback by the family member within a third preset period based on the on-demand playback behavior data;

[0043] Obtaining a fifth weight coefficient of the number of collections, a sixth weight coefficient of the number of searches, and a seventh weight coefficient of the number of payments;

[0044] The usage capability feature is determined based on the number of collections, the number of searches, the number of payments, the fifth weight coefficient, the sixth weight coefficient, and the seventh weight coefficient.

[0045] In the above solution, the step of identifying the family members of other family users to be identified based on the multi-dimensional features in combination with the family members of the first family user includes:

[0046] performing clustering processing on the other family users to be identified according to the multi-dimensional features to obtain clustering results of the other family users to be identified;

[0047] Determining the similarity between the family users of each category in the clustering result and the first family user;

[0048] identifying the family structure members of the second family user in the category corresponding to the maximum similarity among the similarities as the same family structure members as the first family user;

[0049] At least one family user to be identified among the other family users to be identified except the second family user is identified based on the family structure members of the second family user.

[0050] In the above solution, the step of identifying at least one of the other to-be-identified family users except the second family user according to the family structure members of the second family user includes:

[0051] Determining, from among the other to-be-identified family users, a third family user whose live broadcast behavior is associated with the on-demand behavior of the second family user;

[0052] identifying the family structure members of the third family user as the same family structure members as the second family user;

[0053] Constructing a live broadcast behavior network graph of the third family user using a preset algorithm;

[0054] At least one to-be-identified household user other than the second household user and the third household user among the other to-be-identified household users is predicted based on the network diagram.

[0055] The present application also provides a device for identifying family members, including:

[0056] a determining unit configured to determine a first parameter of each of a plurality of family users to be identified based on a media library of family members, wherein the first parameter represents a ratio of viewing time of a specific media asset associated with the media library to total viewing time; the media library including family members and media asset tags;

[0057] a first identification unit, configured to identify family members of a first family user according to the first parameter and a preset threshold; the first family user being at least one family user among the plurality of family users;

[0058] The second identification unit is used to extract multi-dimensional features based on the media library, and identify the family structure members of other family users to be identified based on the multi-dimensional features combined with the family structure members of the first family user; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user.

[0059] An embodiment of the present application further provides an electronic device, including:

[0060] a memory for storing executable instructions;

[0061] The processor is configured to implement any step of the above-described method when executing the executable instructions stored in the memory.

[0062] An embodiment of the present application also provides a computer-readable storage medium storing executable instructions for implementing any step of the above-described method when executed by a processor.

[0063] The embodiments of the present application provide a method, device, electronic device, and storage medium for identifying family structure members, wherein the method comprises: determining a first parameter of each family user among a plurality of family users to be identified based on a media library of family structure members; the first parameter represents the ratio of viewing time of a specific media associated with the media library to total viewing time; the media library includes family structure members and media tags; identifying a family structure member of a first family user based on the first parameter and a preset threshold; the first family user is at least one family user among the plurality of family users; extracting multi-dimensional features based on the media library, and extracting a family structure member based on the multi-dimensional features combined with the family structure of the first family user. Family structure members identify family structure members of other family users to be identified; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user. The solution of the embodiment of the present application first accurately identifies the family structure members of the first family user through a media library based on family structure members; then extracts multi-dimensional features based on the media library, and identifies the family structure members of other family users to be identified based on the multi-dimensional features and the family structure members of the first family user, thereby avoiding the inaccuracy of expert experience identification, greatly improving the accuracy of family user family structure member identification, and also improving the identification efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of a family structure member identification method process is provided for an embodiment of the present application;

[0065] Figure 2 This is a schematic diagram showing the distribution of the viewing time P of a media asset and the total viewing time L of a media asset using a box plot analysis in an embodiment of the present application;

[0066] Figure 3 Schematic diagram of the S-shaped function of the embodiment of the present application;

[0067] Figure 4 This is a schematic diagram of clustering categories in an embodiment of the present application;

[0068] Figure 5 This is a schematic diagram of a family structure obtained based on user on-demand behavior data in an embodiment of the present application;

[0069] Figure 6 This is a network diagram of live broadcast behavior in an embodiment of the present application;

[0070] Figure 7 This is a schematic diagram of the processing flow for identifying family members in an embodiment of the present application;

[0071] Figure 8 This is a box plot analysis of the P / L ratio distribution diagram;

[0072] Figure 9 This is a schematic diagram of user clustering in this application;

[0073] Figure 10 This is a schematic diagram of a family structure member identification device for this application;

[0074] Figure 11 This is a schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0075] The present application will be described in further detail below with reference to the accompanying drawings and embodiments.

[0076] In related technologies, the method of extracting plot description labels from user-watched video data (including basic information such as actors, regions, years, and categories) is not intelligent enough. It is more based on business experience and subjective judgment, and the workload is large. The label description of the video content is incomplete.

[0077] On the other hand, in related technologies, since there is no real labeled data on family structure in most cases, some solutions choose to manually observe the user's historical viewing data with the naked eye and manually label the user. This is not only very labor-intensive, but also will inevitably lead to large human errors in the later stage of manual screening.

[0078] Based on this, this paper proposes using the TextRank algorithm to extract keywords from video descriptions to supplement the video's tag descriptions, making the media asset's plot descriptions more objective and comprehensive. Subsequently, mapping the video's tag descriptions to each family user facilitates the establishment of a family structure identity tag model.

[0079] At the same time, faced with massive amounts of media data such as movies, TV series, and variety shows, we proposed using the Embedding algorithm to perform cluster analysis on media data, mapping the basic information of the video data into a high-dimensional space. Then, similar video data will be in a similar position in the transformed semantic space. Then, a clustering algorithm is used to more accurately divide the media data converted into dense vectors into different categories.

[0080] To accurately and efficiently identify partial family structures and thus improve the accuracy of subsequent semi-supervised learning, a method for constructing family member media data based on the granularity of media asset tags was proposed. Media assets were categorized into those for children and those for the elderly. Indicators were designed for three different dimensions: viewing loyalty, viewing intensity, and usage. By observing the characteristics of different family structure groups within and across indicators, the discriminability and usability of the extracted features were improved when applying the algorithm.

[0081] Taking into account the existence of on-demand and live broadcasting in home broadband scenarios, a family structure identification scheme is designed for live broadcast and on-demand scenarios. At the same time, the live broadcast and on-demand users are correlated with each other for label inference, and the behavioral time series characteristics of live broadcast users are mined based on the real-time characteristics of the live broadcast user viewing scenario.

[0082] The present application provides a method for identifying family members, which is applied to an electronic device. The functions implemented by the method can be implemented by a processor in the electronic device calling program code. Of course, the program code can also be stored in a computer storage medium. Therefore, the electronic device includes at least a processor and a storage medium. As an example, the electronic device can be a mobile phone, a computer, a terminal, an information transceiver, a tablet device, a personal digital assistant, etc.

[0083] Figure 1 A schematic diagram of a family member identification method process is provided for an embodiment of the present application; as shown in FIG1 , the method includes:

[0084] Step 101: Determine a first parameter of each of a plurality of family users to be identified based on a media library of family members; the first parameter represents a ratio of viewing time of a specific media asset associated with the media library to total viewing time; the media library includes family members and media asset tags;

[0085] Step 102: Identify family members of a first family user based on the first parameter and a preset threshold; the first family user is at least one family user among the multiple family users;

[0086] Step 103: Extract multi-dimensional features based on the media library, and identify family members of other family users to be identified based on the multi-dimensional features combined with the family members of the first family user; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user.

[0087] In step 101, the media library includes family members and media labels; both the family members and the media labels can be determined based on actual circumstances and are not limited here. As an example, the family members can be children, teenagers (i.e., K12), middle-aged and young people, the elderly, etc. In actual applications, the family members can also be divided into more detailed categories based on the user's viewing behavior and viewing content. For example, the user's educational level can be divided into literary youth, the number of times the user stays up late to watch dramas can be divided into otaku, and the content of the user's viewing can be divided into star chasers, etc.; the media labels can be labels corresponding to the media that the family members like to watch, for example, for the elderly: opera, elderly care, square dancing, etc.; for children: early education, enlightenment education, cartoon animation, etc.

[0088] The first parameter represents the ratio of viewing time of a specific media asset associated with the media asset library to total viewing time. The specific media asset can be determined based on actual circumstances and is not limited here. As an example, the specific media asset can be media for seniors, media for children, etc. For ease of understanding, viewing time can be denoted as P, total viewing time can be denoted as L, and the ratio can be denoted as P / L.

[0089] In one embodiment, the method further comprises:

[0090] Extracting label data corresponding to the media data watched by each family user;

[0091] The media resource library is updated using the family structure members corresponding to each family user and the tag data.

[0092] The media resource data may be determined according to actual conditions and is not limited here. As an example, the media resource data may be movies, TV series, variety shows, and the like.

[0093] In one embodiment, extracting the tag data corresponding to the media data watched by the plurality of to-be-identified family users further includes:

[0094] Obtaining a brief introduction to the media asset data;

[0095] Key information is extracted from the introduction, and the tag data is determined according to the key information.

[0096] This embodiment addresses the situation where each household user has a description of the media asset being viewed. Key information can be extracted from the description using a preset algorithm based on the description. This preset algorithm can be a keyword extraction (TextRank) algorithm. The algorithm process involves segmenting the given media asset description text into sentences, T = [S1, S2, ..., Sm]. Each sentence is then tokenized to obtain Si = [pi1, pi2, ..., pin], and a word graph is constructed. The window size is set to k, with [p1, p2, ..., pk] [p2, p3, ..., pk+1] representing each window. If two words appear simultaneously in a window, an edge is considered to exist between the corresponding word nodes. Based on this concept and formula, the weights of each node in the word graph are iteratively propagated until convergence. Finally, the node weights are ranked to identify the most important words as candidate keywords. If several extracted keywords are adjacent in the text, they form a key phrase. This method yields an objective tag description of the media asset data.

[0097] In one embodiment, extracting the tag data corresponding to the media data watched by each family user includes:

[0098] Obtaining a brief introduction to the media data watched by each household user;

[0099] Key information is extracted from the introduction, and the tag data is determined according to the key information.

[0100] The attribute information may be determined according to actual conditions and is not limited here. As an example, the attribute information may be the author, organization, and name corresponding to the media asset data.

[0101] This embodiment addresses the situation where the media data viewed by each household lacks a description. Instead, the data is categorized by author and organization, allowing the media to be labeled as songs, movies, and so on. The remaining media is then segmented by word, converted into word vectors using Word2Vector, and mapped to a high-dimensional space using an embedding algorithm. This places similar data in close proximity in the transformed semantic space. Clustering is then performed to select the most representative words as labels for the media assets.

[0102] In this embodiment, the labels manually added by operators to the media data are supplemented in a more comprehensive and objective manner. The TextRank algorithm and the Embedding algorithm are used to conduct more abstract learning on the media text data, which greatly enriches the media label system, thereby making the family structure media library more accurate.

[0103] In one embodiment, the method further comprises:

[0104] Obtaining a storage method of the family structure members and the media asset tags;

[0105] The media resource library is constructed according to the data stored in the storage method.

[0106] The storage method can be determined according to actual conditions and is not limited here. As an example, the storage method can use a distributed (Hadoop) technology architecture for standard structured storage. Specifically, the family structure members and the media asset tags can be stored in a one-to-one correspondence.

[0107] In actual applications, based on the family structure member attribute tag, the media asset identifier (ID) and name and other specific information are mapped one by one to build a media asset library of the family structure members.

[0108] In step 102, the preset threshold value can be determined according to actual conditions and is not limited here. As an example, the preset threshold value can be obtained using box plot analysis.

[0109] In practice, we can calculate the average monthly viewing time P of the media library for each family member and the total viewing time L of each media library. We can then use a box plot to analyze the P / L ratio distribution across all users and identify the threshold by observing the percentiles. Based on this analysis, we can filter users whose viewing time for a specific media library exceeds the threshold and identify them as users of the corresponding family structure.

[0110] In step 103, the multi-dimensional features can be determined based on actual conditions and are not limited here. As an example, the multi-dimensional features may include at least: viewing loyalty features, viewing intensity features, and usage capability features; the viewing loyalty features represent the different stickiness of different user groups for traditional media (e.g., television) and non-traditional media (e.g., mobile phones, tablets, and computers); the viewing intensity features represent the different viewing durations and viewing frequencies of different users; and the usage capability features represent the extent to which users use collection, search, payment, and other features in media on-demand.

[0111] In one embodiment, the multi-dimensional features include at least: viewing loyalty features, viewing intensity features, and usage capability features; and extracting the multi-dimensional features based on the media asset library includes:

[0112] Obtaining on-demand behavior data of the family members based on the media resource library;

[0113] The viewing loyalty feature, the viewing intensity feature, and the usage capability feature are extracted according to the on-demand behavior data.

[0114] The on-demand behavior data can be determined according to actual conditions and is not limited here. As an example, the on-demand behavior data can be television on-demand behavior data, computer on-demand behavior data, etc.

[0115] In one embodiment, extracting the viewing loyalty feature based on the on-demand behavior data includes:

[0116] Obtaining, based on the on-demand behavior data, a unit length of viewing time and an average unit length of interval time of the on-demand viewing of the family members within a first preset period;

[0117] Determining a maximum on-demand viewing time unit length and a maximum on-demand viewing interval time unit length among the family members;

[0118] The viewing loyalty feature is determined according to the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length.

[0119] Among them, the first preset period, the unit length, and the interval time unit length can all be determined according to actual conditions and are not limited here. As an example, the first preset period can be the last three months of the family structure members; the unit length can be days; and the interval time unit length can be interval time days. For ease of understanding, the viewing loyalty feature can be recorded as PlayLoyalty, the viewing time unit length can be recorded as PlayDays, the maximum on-demand viewing time unit length can be recorded as MaxPlayDays, the average interval time unit length can be recorded as PlayIntervalDays, and the maximum on-demand interval time unit length is MaxPlayIntervalDays.

[0120] In one embodiment, determining the viewing loyalty feature based on the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length includes:

[0121] Determining a first ratio of the viewing time unit length to the maximum on-demand viewing time unit length and a second ratio of the average interval time unit length to the maximum on-demand interval time unit length;

[0122] Obtaining a first weight coefficient for the first ratio and a second weight coefficient for the second ratio;

[0123] The viewing loyalty feature is determined based on the first ratio, the second ratio, the first weight coefficient, and the second weight coefficient.

[0124] The first weight coefficient and the second weight coefficient can be determined according to actual conditions and are not limited here. As an example, the first weight coefficient can be 0.4 and the second weight coefficient can be 0.6. For ease of understanding, the first weight coefficient can be denoted as a1 and the second weight coefficient can be denoted as b1.

[0125] In practice, television plays a different role in the lives of different user groups. For young families, they have a wider range of media options, such as mobile phones, tablets, and computers, and their personal entertainment lives are more diverse. Therefore, traditional television media is less popular with this group. For older adults, however, their lifestyles and locations are more fixed, and they use fewer media to access various types of information, so they still rely heavily on traditional television. To address this situation, a user viewing loyalty feature is constructed, as shown in Formula (1).

[0126]

[0127] PlayLoyalty represents viewing loyalty, PlayDays represents the number of days a user has watched on-demand content over the past three months, MaxPlayDays represents the maximum number of days a user has watched on-demand content over the past three months, PlayIntervalDays represents the average number of days between on-demand viewings over the past three months, and MaxPlayIntervalDays represents the maximum number of days a user has watched on-demand content over the past three months. a and b are used to balance the weights of these two dimensions.

[0128] The Sigmoid function is a common S-shaped function in biology, also known as the S-shaped growth curve. In information science, due to its monotonic increasing properties and the monotonic increasing properties of its inverse function, the Sigmoid function is often used as an activation function in neural networks, mapping variables to a value between 0 and 1. Substituting the viewing loyalty value into the Sigmoid function, the viewing loyalty is implicitly mapped to the range of 0 and 1. A larger calculated value indicates a higher viewing loyalty.

[0129] In one embodiment, extracting the viewing intensity feature based on the on-demand behavior data includes:

[0130] Obtaining the frequency and average duration of playback of the program requested by the family members within a second preset period based on the on-demand behavior data;

[0131] Determine the maximum on-demand playback frequency and maximum on-demand playback duration among the family members;

[0132] The viewing intensity feature is determined according to the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration.

[0133] The second preset period can be determined based on actual circumstances and is not limited here. As an example, the second preset period can be the last three months for the family members. For ease of understanding, the viewing intensity feature can be recorded as PlayIntensity, the playback frequency can be recorded as PlayFrequency, the average playback duration can be recorded as PlayTime, and the maximum point playback frequency can be recorded as MaxPlayTime.

[0134] In one embodiment, determining the viewing intensity feature based on the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration includes:

[0135] Determining a third ratio of the playback frequency to the maximum on-demand playback frequency and a fourth ratio of the average playback duration to the maximum on-demand playback duration;

[0136] obtaining a third weight coefficient for the third ratio and a fourth weight coefficient for the fourth ratio;

[0137] The viewing intensity feature is determined based on the third ratio, the fourth ratio, the third weight coefficient, and the fourth weight coefficient.

[0138] The third weight coefficient and the fourth weight coefficient can be determined based on actual conditions and are not limited here. As an example, the third weight coefficient can be 0.5, and the fourth weight coefficient can be 0.5. For ease of understanding, the first weight coefficient can be denoted as a2, and the second weight coefficient can be denoted as b2.

[0139] In practical applications, user viewing intensity features are constructed based on viewing duration and frequency. For example, for users whose families include K-12 students, they typically control the duration and frequency of TV viewing to ensure that it does not affect their studies. However, for users with young children, who do not have as much academic pressure, viewing duration and frequency may differ from other family structures.

[0140]

[0141] PlayIntensity represents playback intensity, PlayFrequency represents the user's playback frequency over the past three months, MaxPlayFrequency represents the maximum playback frequency among all users over the past three months, PlayTime represents the average playback duration, and MaxPlayTime represents the maximum playback duration among all users over the past three months. Similarly, we apply PlayIntensity to a Sigmoid function and map it between 0 and 1. Larger values ​​closer to 1 indicate a greater viewing intensity, and vice versa.

[0142] In one embodiment, extracting the usage capability feature based on the on-demand behavior data includes:

[0143] Determining the number of collections, searches, and payments for on-demand playback by the family member within a third preset period based on the on-demand playback behavior data;

[0144] Obtaining a fifth weight coefficient of the number of collections, a sixth weight coefficient of the number of searches, and a seventh weight coefficient of the number of payments;

[0145] The usage capability feature is determined based on the number of collections, the number of searches, the number of payments, the fifth weight coefficient, the sixth weight coefficient, and the seventh weight coefficient.

[0146] Among them, the third preset period can be determined according to actual conditions and is not limited here. As an example, the third preset period can be the last three months of the family structure members; the fifth weight coefficient, the sixth weight coefficient and the seventh weight coefficient can all be determined according to actual conditions and are not limited here. As an example, the fifth weight coefficient can be 0.4, the sixth weight coefficient can be 0.3, and the seventh weight coefficient can be 0.3. For ease of understanding, the fifth weight coefficient can be recorded as a3, the sixth weight coefficient can be recorded as b3, the seventh weight coefficient can be recorded as c, the usability feature can be recorded as useability, the number of collections can be recorded as collecttimes, the number of searches can be recorded as searchtimes, the number of payments can be recorded as paytimes, and the maximum point-to-play frequency can be recorded as MaxPlayTime.

[0147] In actual applications, the user capability characteristics of the elderly are relatively simple compared to the younger family user group. Their learning and usage abilities are weaker than those of the elderly. They rarely use functions such as collection, search, and payment in on-demand TV. Based on this, we build user capability characteristics.

[0148] usability=log e (a3·collecttimes+b3·searchtimes+c·paytimes) (3)

[0149] Useability represents the user's ability to use the platform, collecttimes represents the total number of historical collections, searchtimes represents the total number of historical searches, and paytimes represents the total number of historical payments. Substituting the calculated value into a sigmoid function, the useability is mapped to a value between 0 and 1. A larger value indicates a stronger user's ability to use the on-demand platform, and vice versa.

[0150] In one embodiment, identifying family members of other family users to be identified based on the multi-dimensional features in combination with the family members of the first family user includes:

[0151] performing clustering processing on the other family users to be identified according to the multi-dimensional features to obtain clustering results of the other family users to be identified;

[0152] Determining the similarity between the family users of each category in the clustering result and the first family user;

[0153] identifying the family structure members of the second family user in the category corresponding to the maximum similarity among the similarities as the same family structure members as the first family user;

[0154] At least one family user to be identified among the other family users to be identified except the second family user is identified based on the family structure members of the second family user.

[0155] The clustering process may include a first clustering process and a second clustering process; both the first clustering process and the second clustering process may be determined based on actual conditions and are not limited herein. As an example, the first clustering process may be performed using a k-means clustering algorithm; and the second clustering process may be performed using a density-based spatial clustering of application with noise (DBSCAN) algorithm.

[0156] In practical applications, the K-means algorithm can be used for coarse user clustering. Selecting a smaller K value, depending on the data, allows users to be roughly clustered into broad categories. Easily distinguishable users are then segmented to avoid users with unclear boundaries being grouped into different categories. Within each broad category, the DBSCAN algorithm is then used for secondary fine-grained clustering, clustering users into fine-grained categories based on user density and spatial distribution. This minimizes the distance between users in the same category and maximizes the distance between different categories.

[0157] Assume that the clustering results are K1, K2, K3, K4..., and each category of users contains three features. Calculate the similarity between each category and the first family user. Based on the K-nearest neighbor method, the family structure of each clustering result that is most similar to the first family user is the family structure of the current category.

[0158] In this embodiment, three main data features are designed based on the behavioral characteristics of different family structures, and the similarity calculation method is applied to greatly improve the accuracy of family structure recognition.

[0159] In one embodiment, the identifying of at least one to-be-identified family user other than the second family user among the other to-be-identified family users according to the family structure members of the second family user includes:

[0160] Determining, from among the other to-be-identified family users, a third family user whose live broadcast behavior is associated with the on-demand behavior of the second family user;

[0161] identifying the family structure members of the third family user as the same family structure members as the second family user;

[0162] Constructing a live broadcast behavior network graph of the third family user using a preset algorithm;

[0163] At least one to-be-identified household user other than the second household user and the third household user among the other to-be-identified household users is predicted based on the network diagram.

[0164] The preset algorithm may be determined according to actual conditions and is not limited here. As an example, the preset algorithm may be a label propagation algorithm (LabelPropagation).

[0165] In actual applications, based on the user's on-demand behavior data, the family structure results of families with young children, families with elderly people, and families with K12 students were obtained. Many family users not only have daily video on-demand behavior, but also have live viewing behavior. Through the overlap of on-demand behavior and live broadcast behavior, part of the family structure in the live broadcast behavior was obtained. In the user watching live broadcast scenario, if two users have watched the same TV program in the past three months, it is considered that there is similarity between the two users and they are connected with an edge. The longer the two users watch the same program, the greater the edge weight. At the same time, the more the same programs are watched, the greater the edge weight of the two users. Construct a user program viewing network based on the live viewing behavior history of each user. Use the LabelPropagation label propagation algorithm, such as Figure 6 As shown in the figure, there are network connections between users who have watched the same program. The larger circles represent some users with known family structures. The model is trained based on the existing partial family structure labels to infer the family structures of the remaining users, and finally the family structure results of all users are obtained.

[0166] The embodiment of the present application constructs a live broadcast user viewing program network based on user on-demand and live broadcast viewing behavior data and an associated label propagation algorithm, thereby improving model recognition efficiency.

[0167] For ease of understanding, a method for identifying family members is exemplified herein, specifically a method for identifying the family structure of a home broadband user. The specific implementation process is as follows:

[0168] Step 1: Based on a distributed Hadoop system architecture, users' daily viewing history is stored as standard structured data. Because user viewing behavior is granular on a daily basis, this information, including fields such as user ID, viewing date, media source, media source tags, media source start and end times, media source channel, media source author, and organization, is stored in a static database (Hive) partitioned by day.

[0169] Step 2: Typically, the media data that users watch on demand includes tags like "action" and "comedy." These tags are assigned by operators or maintenance personnel, which lacks objective and comprehensive descriptions of the media and often contains incomplete information. Therefore, we first need to develop an objective and comprehensive media tagging system.

[0170] (1) Based on the introduction of media asset data, the TextRank algorithm is applied. The algorithm process is as follows: the given media asset introduction text is divided into sentences, T = [S1, S2, ..., Sm], and then each sentence is segmented to obtain Si = [pi1, pi2, ..., pin], and a word graph is constructed. The window size is set to k, [p1, p2, ..., pk][p2, p3, ..., pk+1], etc. are all windows. If two words appear at the same time in a window, it is considered that there is an edge between the corresponding word nodes. According to this idea and formula, the weights of each node in the word graph are iteratively propagated until convergence. Finally, the node weights are sorted to obtain the most important words as candidate keywords. If several extracted keywords are adjacent in the text, they form a key phrase. This method can obtain an objective tag description of the media asset data.

[0171] (2) For media assets without video descriptions, they can be classified by the author and the organization where the data is located, and the media assets can be labeled as songs, movies, etc. The remaining media assets are segmented by the media asset names, converted into word vectors using Word2Vector, and the media asset segmentation results are mapped to a high-dimensional space using the Embedding algorithm. Then, similar data will be in a similar position in the transformed semantic space, and then the most representative words will be selected through clustering as the labels of the media assets.

[0172] Step 3: By combining the above methods with existing data, we developed a more complete and objective pool of media asset tags. These tags are abstract representations of media asset data, significantly reducing the number of tags required compared to the vast amount of media asset data. Leveraging expert experience, we selected tags with distinct family structure attributes, such as those for seniors: opera, elderly care, and square dancing, and for children: early childhood education, enlightenment education, and cartoons. Based on these family structure attribute tags, we mapped specific information, such as media asset IDs and names, to construct a family structure media library.

[0173] (1) Count the average monthly viewing time P of the media library of the corresponding family structure members and the total viewing time L of the media, use the box plot to analyze the distribution of the P / L ratio of all users, and find the threshold by observing the quantile line of the ratio. Figure 2 As shown, Figure 2This embodiment of the present application uses a box plot to analyze the distribution of the viewing time P of media resources and the total viewing time L of media resources. Based on the results of the analysis, users whose viewing time of specific media resources exceeds a threshold are filtered to be users with corresponding family structures.

[0174] (2) Through the above method, we have obtained some family structure users, such as family members including children, family members including elderly people, family members including K12 students, etc.

[0175] Step 4: Construct features.

[0176] 1) User Viewing Loyalty Characteristics: The proportion of television in the personal lives of different user groups varies. For young families, there are many media options available, such as mobile phones, tablets, and computers, and their personal entertainment lives are relatively rich. Traditional television media is not very popular with this group. For older people, however, their lifestyles and regions are relatively fixed, and they use fewer media to receive various types of information, so they still rely heavily on traditional television. To address this situation, a user viewing loyalty characteristic is constructed. Refer to the previous formula (1).

[0177] The Sigmoid function is a common S-shaped function in biology, also known as the S-shaped growth curve. In information science, due to its monotonic and inverse monotonic properties, the Sigmoid function is often used as an activation function in neural networks to map variables between 0 and 1. Substituting the viewing loyalty value into the Sigmoid function, the viewing loyalty is implicitly mapped to the range of 0 and 1. The larger the calculated value, the higher the user's viewing loyalty. The S-shaped function can be expressed as like Figure 3 As shown, Figure 3 Schematic diagram of the S-shaped function of the embodiment of the present application.

[0178] 2) User viewing intensity characteristics: User viewing intensity characteristics are constructed based on the user's viewing time and frequency. For example, for users whose families include K-12 students, the user usually controls the duration and frequency of TV viewing to avoid affecting the students' studies. However, for users in families with young children, the children do not have much academic pressure, so the viewing time and frequency differ from other family structures. Refer to the previous formula (2).

[0179] 3) User Usage Capability Characteristics: Compared to young family users, the elderly have a more limited TV on-demand behavior and are less capable of learning and using TV on-demand features. They rarely use functions such as favorites, search, and payment. Based on this, we construct a user usage capability characteristic. Refer to the previous formula (3).

[0180] Useability represents the user's ability to use the platform, collecttimes represents the total number of historical collections, searchtimes represents the total number of historical searches, and paytimes represents the total number of historical payments. Substituting the calculated value into a sigmoid function, the useability is mapped to a value between 0 and 1. A larger value indicates a stronger user's ability to use the on-demand platform, and vice versa.

[0181] Step 5: Extract user loyalty features, viewing intensity features, and usage ability features from the full amount of user TV on-demand behavior data, filter out some family structure users obtained in the second step, and use the K-means algorithm to coarsely cluster users. The K value is selected to be smaller, depending on the data situation, so that users can be roughly clustered into large categories. Divide users who are easier to distinguish to avoid users with blurred interfaces being assigned to other categories. Then, in each large category, use DBSCAN for secondary fine clustering, clustering into fine-grained categories based on user density and spatial distribution, such as Figure 4 As shown, Figure 4 This is a schematic diagram of clustering categories in an embodiment of the present application; the distance between the same categories is minimized as much as possible, and the distance between different categories is maximized.

[0182] Assume that the clustering results are K1, K2, K3, K4..., and each category of users contains three features. Calculate the similarity between each category and the partial family structure users obtained in the second step. Based on the K-nearest neighbor method, the family structure of the current category is the one that is most similar to the known partial family structure users in each clustering result.

[0183] Through the overlap between on-demand and live streaming behaviors, we can derive some family structures in live streaming behaviors. Figure 5 As shown, Figure 5 This is a schematic diagram of the family structure obtained based on the user's on-demand behavior data in the embodiment of this application; in the scenario of users watching live broadcasts, if two users have watched the same TV program in the past three months, it is considered that there is similarity between the two users and they are connected by an edge. The longer the two users watch the same program, the greater the edge weight. At the same time, the more the same programs are watched, the greater the edge weight of the two users. A user viewing program network is constructed based on the live viewing behavior history of each user. Use LabelPropagation, such as Figure 6 As shown, Figure 6 This is a diagram of the live broadcast behavior network in an embodiment of the present application; there is a network connection between users who have watched the same program, where the larger circles represent some users with known family structures. The model is trained based on the existing partial family structure labels to infer the family structures of the remaining users, and finally the family structure results of all users are obtained. Figure 7This is a schematic diagram of the process flow for identifying family members in an embodiment of the present application. Figure 7 shown.

[0184] In this application, the TextRank algorithm is used to extract keywords from video descriptions to supplement the video's tag descriptions, making the media asset's plot descriptions more objective and comprehensive. Subsequently, when the video's tag descriptions are mapped to each family user, this is more conducive to modeling family structure and identity tags.

[0185] At the same time, faced with massive amounts of media data such as movies, TV series, and variety shows, we proposed using the Embedding algorithm to perform cluster analysis on media data, mapping the basic information of the video data into a high-dimensional space. Then, similar video data will be in a similar position in the transformed semantic space. Then, a clustering algorithm is used to more accurately divide the media data converted into dense vectors into different categories.

[0186] To accurately and efficiently identify partial family structures and thus improve the accuracy of subsequent semi-supervised learning, a method for constructing family member media data based on the granularity of media asset tags was proposed. Media assets were categorized into those for children and those for the elderly. Indicators were designed for three different dimensions: viewing loyalty, viewing intensity, and usage. By observing the characteristics of different family structure groups within and across indicators, the discriminability and usability of the extracted features were improved when applying the algorithm.

[0187] Taking into account the existence of on-demand and live broadcasting in home broadband scenarios, a family structure identification scheme is designed for live broadcast and on-demand scenarios. At the same time, the live broadcast and on-demand users are correlated with each other for label inference, and the behavioral time series characteristics of live broadcast users are mined based on the real-time characteristics of the live broadcast user viewing scenario.

[0188] In order to further understand the present application, an actual application scenario is given as an example.

[0189] Assume that the current family structure to be identified is that there are elderly users in the family.

[0190] Step 1: Apply the TextRank algorithm to all video descriptions in the user on-demand behavior data. This algorithm segments the given media asset description into sentences and words. A sliding window size of k is set, with [p1,p2,...,pk][p2,p3,...,pk+1] representing the sliding window. If two words appear simultaneously within a window, an edge is considered to exist between the corresponding word nodes. This method generates keywords for the media asset data, which serve as objective tags for the current video data. The choice of K depends on the data. A larger K value results in a larger sliding window, yielding more keywords, while a smaller K value yields fewer keywords.

[0191] Step 2: If a media resource does not have a description, labels can be extracted based on the author and institution of the data. For the remaining media resources, the names of the media resources are segmented and converted into word vectors using the Word2Vector algorithm. Then, the most representative words are clustered and selected as the media resource labels.

[0192] Step 3: This yields a pool of media tags. Based on pre-defined business experience, we filter out tags with obvious elderly-related attributes, such as opera, elderly care, and square dancing. This ultimately creates a library of media resources targeted at the elderly. This data is processed into standard structured data and stored in a Hive static database.

[0193] Step 4: Count the average monthly viewing time P of elderly media and the total viewing time L of media, and use a box plot to analyze the P / L ratio distribution of all users. Figure 8 The box plot analysis shows the distribution of P / L, as shown below: Figure 8 As shown in the figure, we identify the threshold by observing the quantiles of the percentages. We found that users whose viewing time for elderly media exceeds 15% are in the critical zone. These users are those whose families include elderly people.

[0194] Step 5: Extract multi-dimensional features.

[0195] (1) User viewing loyalty features: Statistics are collected for all users, including the number of days they watched on-demand content in the past three months, the maximum number of days they watched on-demand content, the average interval between users’ on-demand content in the past three months, and the maximum interval between users’ on-demand content. Here, we consider the number of days watched and the interval between users to be equally important, so we set a = b = 0.5 and apply the calculated viewing loyalty value to the Sigmoid function, which maps the viewing loyalty value to the range of 0 and 1.

[0196] (2) User viewing intensity characteristics: Statistics are collected for the frequency of on-demand playback of all users in the past three months, the maximum playback frequency among all users in the past three months, the playback duration in the past three months, and the maximum playback duration among all users in the past three months. Substitute the calculated results into the viewing intensity formula and finally substitute them into the Sigmoid function to map them to the range 0 and 1.

[0197] (3) User Usage Capacity Characteristics: Count each user's total number of historical collections, searches, and payments. Substitute the calculated values ​​into the Usage Capacity Characteristics calculation formula and feed them into the Sigmoid function. Usage capacity is mapped to the range of 0 and 1.

[0198] Step 6: Eliminate some users with elderly people in the fourth step, and use the K-means algorithm to roughly cluster the remaining users, such as Figure 9 As shown, Figure 9This is a diagram of user clustering in this application. It was observed that clustering works best when K = 7. Users are first clustered into large categories. Then, within each large category, the DBSCAN algorithm is used for secondary clustering, clustering into fine-grained categories based on user density and spatial distribution. Assuming the clustering results are K1, K2, K3, K4…, and each category of users contains three features, the K-nearest neighbor method is used to calculate the similarity between each category and the users with elderly families obtained in step 4. The similarities are sorted from highest to lowest, and the users with the highest similarity are those with elderly families.

[0199] Step 7: Through the above method, users with elderly people at home are obtained based on the user's on-demand behavior data. Partially identified users with elderly people at home are screened out from users who have both video on demand behavior and live broadcast viewing behavior. Based on the fact that two users have watched the same TV program in the past three months, it is considered that there is similarity between the two users. The longer the two users watch the same program, the greater the similarity. At the same time, the more the same programs they watch, the greater the similarity between the two users. The user viewing behavior network is constructed. Using the LabelPropagation label propagation algorithm, the family structure of the remaining users is inferred based on the existing partial family structure label training model, and finally the users with elderly people at home among all users are obtained.

[0200] In order to implement the method of the embodiment of the present application, the embodiment of the present application further provides a family member identification device 1000, which is set on an electronic device. Figure 10 This is a schematic diagram of a family structure member identification device for this application, such as Figure 10 Shown, including:

[0201] A determining unit 1001 is configured to determine a first parameter of each of a plurality of family users to be identified based on a media library of family members, wherein the first parameter represents a ratio of viewing time of a specific media asset associated with the media library to total viewing time; the media library includes family members and media asset tags;

[0202] A first identification unit 1002 is configured to identify family members of a first family user according to the first parameter and a preset threshold; the first family user is at least one family user among the plurality of family users;

[0203] The second identification unit 1003 is used to extract multi-dimensional features based on the media library, and identify family structure members of other family users to be identified based on the multi-dimensional features combined with the family structure members of the first family user; the other family users to be identified are at least one family user to be identified among the multiple family users to be identified except the first family user.

[0204] Here, in one embodiment, the device further includes an updating unit for extracting label data corresponding to the media data watched by each family user; and updating the media library using the family structure members corresponding to each family user and the label data.

[0205] Here, in one embodiment, the updating unit is further configured to obtain a brief introduction of the media data watched by each family user;

[0206] Key information is extracted from the introduction, and the tag data is determined according to the key information.

[0207] Here, in one embodiment, the updating unit is further configured to obtain attribute information of the media data watched by each household user; perform cluster analysis on the attribute information to obtain words representing common attributes of the media data; and use the words as the label data.

[0208] In one embodiment, the device further includes a construction unit configured to obtain a storage method of the family structure members and the media asset tags; and construct the media asset library according to the data stored in the storage method.

[0209] In one embodiment, the multi-dimensional features include at least: viewing loyalty features, viewing intensity features and usage ability features; the second identification unit 1003 is also used to obtain the on-demand behavior data of the family structure members based on the media library; and extract the viewing loyalty features, the viewing intensity features and the usage ability features based on the on-demand behavior data.

[0210] Here, in one embodiment, the second identification unit 1003 is further used to obtain the viewing time unit length and the average interval time unit length of the on-demand viewing of the family structure members within the first preset period based on the on-demand behavior data; determine the maximum on-demand viewing time unit length and the maximum on-demand interval time unit length among the family structure members; and determine the viewing loyalty feature based on the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length and the maximum on-demand interval time unit length.

[0211] Here, in one embodiment, the second identification unit 1003 is further used to determine a first ratio of the viewing time unit length to the maximum on-demand viewing time unit length and a second ratio of the average interval time unit length to the maximum on-demand interval time unit length; obtain a first weight coefficient of the first ratio and a second weight coefficient of the second ratio; and determine the viewing loyalty feature based on the first ratio, the second ratio, the first weight coefficient and the second weight coefficient.

[0212] Here, in one embodiment, the second identification unit 1003 is further used to obtain the frequency of on-demand playback and the average playback duration of the family structure members within a second preset period based on the on-demand behavior data; determine the maximum on-demand playback frequency and the maximum on-demand playback duration among the family structure members; and determine the viewing intensity characteristics based on the playback frequency, the maximum on-demand playback frequency, the average playback duration and the maximum on-demand playback duration.

[0213] Here, in one embodiment, the second identification unit 1003 is further used to determine a third ratio of the playback frequency to the maximum on-demand playback frequency and a fourth ratio of the average playback time to the maximum on-demand playback time; obtain a third weight coefficient of the third ratio and a fourth weight coefficient of the fourth ratio; and determine the viewing intensity feature based on the third ratio, the fourth ratio, the third weight coefficient and the fourth weight coefficient.

[0214] Here, in one embodiment, the second identification unit 1003 is further used to determine the number of collections, searches and payments of the family structure members in the on-demand video playback within a third preset period based on the on-demand behavior data; obtain the fifth weight coefficient of the number of collections, the sixth weight coefficient of the number of searches and the seventh weight coefficient of the number of payments; and determine the usage capability characteristics based on the number of collections, the number of searches, the number of payments, the fifth weight coefficient, the sixth weight coefficient and the seventh weight coefficient.

[0215] In one embodiment, the second identification unit 1003 is further used to cluster the other family users to be identified based on the multi-dimensional features to obtain clustering results of the other family users to be identified; determine the similarity between the family users of each category in the clustering results and the first family user; identify the family structure members of the second family user in the category corresponding to the maximum similarity in the similarities as having the same family structure members as the first family user; and identify at least one family user to be identified among the other family users to be identified except the second family user based on the family structure members of the second family user.

[0216] In one embodiment, the second identification unit 1003 is further used to determine a third family user whose live broadcast behavior is associated with the on-demand behavior of the second family user among the other family users to be identified; identify the family structure members of the third family user as having the same family structure members as the second family user; construct a live broadcast behavior network graph of the third family user using a preset algorithm; and predict at least one family user to be identified among the other family users to be identified, except the second family user and the third family user, based on the network graph.

[0217] It should be noted that the above-described embodiments of the family member identification device, when performing family member identification, illustrate the division of the aforementioned program modules only as an example. In actual applications, the aforementioned processing can be assigned to different program modules as needed, i.e., the internal structure of the device can be divided into different program modules to perform all or part of the aforementioned processing. Furthermore, the family member identification device and the family member identification method embodiments provided in the above-described embodiments share the same concept. The specific implementation process is detailed in the method embodiments and will not be further elaborated here.

[0218] Based on the hardware implementation of the above-mentioned program module, an embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the family structure member identification method provided in the above-mentioned embodiment are implemented.

[0219] Correspondingly, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the family structure member identification method provided in the above embodiment are implemented.

[0220] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0221] It should be noted that Figure 11 FIG. 1 is a schematic diagram of a hardware entity structure of an electronic device in an embodiment of the present application, such as Figure 11 As shown, the hardware entity of the electronic device 1100 includes: a processor 1101 and a memory 1103 . Optionally, the electronic device 1100 may further include a communication interface 1102 .

[0222] It is understood that the memory 1103 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk or a magnetic tape. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 1103 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.

[0223] The methods disclosed in the above embodiments of the present application can be applied to the processor 1101 or implemented by the processor 1101. The processor 1101 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1101 or by instructions in the form of software. The above processor 1101 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1101 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 1103. The processor 1101 reads the information in the memory 1103 and completes the steps of the above method in combination with its hardware.

[0224] In an exemplary embodiment, the device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0225] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0226] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0227] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0228] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0229] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0230] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for identifying family members, characterized in that: include: Determining a first parameter of each family user among a plurality of family users to be identified based on a media resource library of family members; The first parameter represents the ratio of viewing time of a specific media asset associated with the media asset library to total viewing time; the media asset library includes family structure members and media asset tags; Identify family members of a first family user according to the first parameter and a preset threshold; the first family user is at least one family user among the multiple family users to be identified; extracting multi-dimensional features based on the media asset library, identifying family members of other to-be-identified family users according to the multi-dimensional features combined with the family members of the first family user; the other to-be-identified family users being at least one to-be-identified family user among the plurality of to-be-identified family users other than the first family user; The step of identifying family members of other to-be-identified family users based on the multi-dimensional features in combination with the family members of the first family user includes: performing clustering processing on the other family users to be identified according to the multi-dimensional features to obtain clustering results of the other family users to be identified; Determining the similarity between the family users of each category in the clustering result and the first family user; identifying the family structure members of the second family user in the category corresponding to the maximum similarity among the similarities as the same family structure members as the first family user; Determining, from among the other to-be-identified family users, a third family user whose live broadcast behavior is associated with the on-demand behavior of the second family user; identifying the family structure members of the third family user as the same family structure members as the second family user; Constructing a live broadcast behavior network graph of the third family user using a preset algorithm; At least one to-be-identified household user other than the second household user and the third household user among the other to-be-identified household users is predicted based on the network diagram.

2. The method according to claim 1, characterized in that The method further comprises: Extracting label data corresponding to the media data watched by each family user; The media resource library is updated using the family structure members corresponding to each family user and the tag data.

3. The method according to claim 2, characterized in that The extracting of label data corresponding to the media data watched by each family user includes: Obtaining a brief introduction to the media data watched by each household user; Key information is extracted from the introduction, and the tag data is determined according to the key information.

4. The method according to claim 2, characterized in that The extracting of label data corresponding to the media data watched by each family user includes: Obtaining attribute information of the media data watched by each family user; Performing cluster analysis on the attribute information to obtain words that represent common attributes of the media asset data; The words are used as the label data.

5. The method according to claim 1, wherein The method further comprises: Obtaining a storage method of the family structure members and the media asset tags; The media resource library is constructed according to the data stored in the storage method.

6. The method according to claim 1, characterized in that The multi-dimensional features include at least: viewing loyalty features, viewing intensity features, and usage capability features; and the multi-dimensional features extracted based on the media resource library include: Obtaining on-demand behavior data of the family members based on the media resource library; The viewing loyalty feature, the viewing intensity feature, and the usage capability feature are extracted according to the on-demand behavior data.

7. The method according to claim 6, characterized in that The extracting the viewing loyalty feature according to the on-demand behavior data includes: Obtaining, based on the on-demand behavior data, a unit length of viewing time and an average unit length of interval time of the on-demand viewing of the family members within a first preset period; Determining a maximum on-demand viewing time unit length and a maximum on-demand viewing interval time unit length among the family members; The viewing loyalty feature is determined according to the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length.

8. The method according to claim 7, characterized in that The determining the viewing loyalty feature according to the viewing time unit length, the maximum on-demand viewing time unit length, the average interval time unit length, and the maximum on-demand interval time unit length includes: Determining a first ratio of the viewing time unit length to the maximum on-demand viewing time unit length and a second ratio of the average interval time unit length to the maximum on-demand interval time unit length; Obtaining a first weight coefficient for the first ratio and a second weight coefficient for the second ratio; The viewing loyalty feature is determined based on the first ratio, the second ratio, the first weight coefficient, and the second weight coefficient.

9. The method according to claim 6, characterized in that The extracting the viewing intensity feature according to the on-demand behavior data includes: Obtaining the frequency and average duration of playback of the program requested by the family members within a second preset period based on the on-demand behavior data; Determine the maximum on-demand playback frequency and maximum on-demand playback duration among the family members; The viewing intensity feature is determined according to the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration.

10. The method according to claim 9, characterized in that The determining the viewing intensity feature according to the playback frequency, the maximum on-demand playback frequency, the average playback duration, and the maximum on-demand playback duration includes: Determining a third ratio of the playback frequency to the maximum on-demand playback frequency and a fourth ratio of the average playback duration to the maximum on-demand playback duration; obtaining a third weight coefficient for the third ratio and a fourth weight coefficient for the fourth ratio; The viewing intensity feature is determined based on the third ratio, the fourth ratio, the third weight coefficient, and the fourth weight coefficient.

11. The method according to claim 6, characterized in that The extracting the usage capability feature according to the on-demand behavior data includes: Determining the number of collections, searches, and payments for on-demand playback by the family member within a third preset period based on the on-demand playback behavior data; Obtaining a fifth weight coefficient of the number of collections, a sixth weight coefficient of the number of searches, and a seventh weight coefficient of the number of payments; The usage capability feature is determined based on the number of collections, the number of searches, the number of payments, the fifth weight coefficient, the sixth weight coefficient, and the seventh weight coefficient.

12. A family member identification device, characterized in that: include: a determining unit, configured to determine a first parameter of each family user among a plurality of family users to be identified based on a media resource library of family members; The first parameter represents the ratio of viewing time of a specific media asset associated with the media asset library to total viewing time; the media asset library includes family structure members and media asset tags; a first identification unit, configured to identify family members of a first family user according to the first parameter and a preset threshold; the first family user being at least one family user among the plurality of family users to be identified; a second identification unit configured to extract multi-dimensional features based on the media asset library, and identify family members of other to-be-identified family users based on the multi-dimensional features in combination with the family members of the first family user; the other to-be-identified family users being at least one to-be-identified family user among the plurality of to-be-identified family users other than the first family user; The second identification unit is further configured to cluster the other family users to be identified based on the multi-dimensional features to obtain clustering results of the other family users to be identified; determine similarity between each category of family users in the clustering results and the first family user; identify family members of a second family user in a category corresponding to a maximum similarity among the similarities as having the same family members as the first family user; determine, among the other family users to be identified, a third family user whose live broadcast behavior is associated with the on-demand behavior of the second family user; and identify family members of the third family user as having the same family members as the second family user; A preset algorithm is used to construct a live broadcast behavior network diagram of the third family user; and at least one family user to be identified other than the second family user and the third family user is predicted based on the network diagram.

13. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the family structure member identification method according to any one of claims 1 to 11 when executing the executable instructions stored in the memory.

14. A computer-readable storage medium, characterized in that Executable instructions are stored for implementing the family structure member identification method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Family member identification method and device, electronic equipment and readable storage medium

    CN113065058A

  • Systems and methods for generating and implementing knowledge graphs for knowledge representation and analysis

    US10496678B1