User portrait generation method and generation system based on behavior data

Through time domain scrolling update of user behavior data and combining LSTM-Bayesian module and converter model, the problem of neglecting the timing characteristics of user portraits in the prior art is solved, and a high accuracy and timeliness mass user portrait is achieved, which improves the personalization of integrated media content recommendations.

CN119048147BActive Publication Date: 2025-05-13BEIJING TONGFANG LEGENDSILICON TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411541271.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-13
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

The prior art ignores the timing characteristics of user behavior data in mass user portraits, resulting in a decrease in accuracy and timeliness of user portraits, and is unable to respond to changes in user behavior in real time.

Method used

The user behavior sequence data is continuously updated through time domain scrolling, and combined with the long and short-term memory network-Bayesian module and converter model, the user behavior interest tag is accurately output to form a group user portrait with depth and breadth.

Benefits of technology

It improves the accuracy and timeliness of user portraits, can respond to user behavior changes in real time or near real time, and enhances the accuracy and personalization of integrated media content recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048147B_ABST
    Figure CN119048147B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating a user portrait based on behavior data, including obtaining the first behavior data of the user and the user identity data, and updating the first behavior data in real time, and recording the timestamp of the first behavior data. Determine whether to generate second user behavior data, and update the user behavior data sequence. Input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tags, convert the interest tags into the input format of the converter model, feature extract the user interest tag sequence, and obtain the global feature vector of the user interest tags. Generate the user portrait based on the global feature vector of the user interest tags and the user identity data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a method and system for generating a user portrait based on behavior data. Background Art

[0002] In the current converged media platforms, the user's behavior data such as clicking, browsing, forwarding and commenting on content contains the user's interests and preferences. In particular, establishing a group user portrait for the converged media platform based on user groups can help the platform better understand and segment user groups, thereby providing more accurate and personalized content recommendations and services. This is of great significance for enhancing user stickiness and improving user experience and satisfaction. Through user portraits, the converged media platform can gain insight into user needs and behavior patterns, optimize content creation and dissemination strategies, achieve precision marketing, and improve operational efficiency. At the same time, this will also help the platform to tap into new business value and promote the in-depth development of media convergence.

[0003] However, existing technologies often ignore the temporal characteristics of user behavior data in group user portraits and fail to effectively integrate user behavior sequences, thus affecting the accuracy of user portraits and the depth of group analysis. This neglect may have the following adverse effects: reduced portrait accuracy, failure to integrate user behavior sequences will result in the inability to capture the development trends and pattern changes of user behavior, thus affecting the accuracy and timeliness of user portraits: portraits lacking temporal analysis cannot provide users with personalized services customized according to their behavior evolution, which may lead to a decline in user experience, and cannot respond to changes in user behavior in real time or near real time, which may cause the platform to lag in responding to user needs.

[0004] The current cutting-edge directions in the field include using deep learning technology to process time series data, and extracting the global features and relationships of user interest tags through large-scale parallel computing, but the main challenges include: how to dynamically update user behavior sequences and how to accurately and efficiently output typical representative behavior interest tags of user groups. The problem has never been solved. Specifically, the amount of user behavior data is huge and needs to be updated in real time, which places extremely high demands on the performance of data storage and processing systems. Dynamically updating behavior sequences requires algorithms to be able to adapt to changes in data streams, which means that the algorithms need to be flexible and robust enough; pattern change recognition: user behavior patterns may change over time. Identifying these changes and updating user sequences accordingly is a complex learning process. Extracting representative interest tags from them requires in-depth semantic understanding and pattern recognition. Summary of the invention

[0005] In view of the above-mentioned defects, the technical problem solved by the present invention is to provide a method for generating user portraits based on behavioral data, continuously update user behavior sequence data in a time domain rolling manner, and combine the long short-term memory network-Bayesian module and the converter model to accurately output user behavior interest tags and form a group user portrait with depth and breadth, so as to improve the accuracy and personalization of integrated media content recommendations, and solve the problems of insufficient utilization of user behavior data and untimely updating of user portraits in the prior art.

[0006] A first aspect of the present invention provides a method for generating a user portrait based on behavior data, comprising:

[0007] S101: Acquire the first behavior data of the user and the user identity data, update the first behavior data in real time, and record the timestamp of the first behavior data.

[0008] S102: Determine whether second user behavior data is generated, and update the user behavior data sequence.

[0009] S103: Input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tags.

[0010] S104: Convert the interest tags into an input format of a converter model, feature extract the user interest tag sequence, and obtain a global feature vector of the user interest tags.

[0011] S105: Generate the user portrait according to the global feature vector of the user interest tag and the user identity data.

[0012] According to an embodiment of the present invention, in step S102, determining whether second user behavior data is generated and updating the user behavior data includes: S201: setting a fixed sampling time period , and record the fixed sampling time period The first user behavior data within.

[0013] S202: Determine whether second user behavior data is generated. If the second user behavior data is generated, the non-fixed sampling time period in the first user behavior data is The user behavior data is deleted and the user behavior data sequence is updated, including:

[0014] Record the timestamp of the second user behavior data, and determine whether the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period .

[0015] If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period , record the second user behavior data.

[0016] If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is not less than the fixed sampling time period , delete the first user behavior data.

[0017] According to an embodiment of the present invention, in step S103, the updated user behavior data sequence is input into the long short-term memory network-Bayesian module to obtain the user's interest tags, including: using the Bayesian formula to predict the user's interest tags, as shown in formula (1).

[0018] Formula (1)

[0019] in, represents the probability that a user has a specific label under a given behavior sequence, It represents the probability that a user with a specific label will generate this behavior data sequence. is the prior probability, which indicates the inherent probability that the user has this label, is the marginal probability of the user behavior data sequence.

[0020] According to an embodiment of the present invention, in step S104, converting the interest tag into an input format of a converter model, feature extracting the user interest tag sequence, and obtaining a global feature vector of the user interest tag includes:

[0021] The interest tags are embedded in the vector , forming a sequence, expressed as formula (2):

[0022] Formula (2)

[0023] Among them, e1,e2,...,e n represents the vector into which each interest is embedded, Represents the sequence of interest tags formed after being embedded in the vector.

[0024] Then in the embedding vector Add positional encoding , as shown in formula (3):

[0025] Formula (3)

[0026] in, Represents the sequence of interest tags formed by adding position information to the embedded vector.

[0027] The position code , as shown in formula (4):

[0028] Formula (4)

[0029] in, is the position index, is the vector dimension index, is the embedding vector Dimension.

[0030] Through the self-attention multi-head attention mechanism, the global feature vector of the user interest sequence is extracted.

[0031] A second aspect of the present invention provides a system for generating a user portrait based on behavior data, the system comprising:

[0032] The acquisition module is used to acquire the first behavior data of the user and the user identity data, and to update the first behavior data in real time and record the timestamp of the first behavior data.

[0033] The updating module is used to determine whether the second user behavior data is generated and to update the user behavior data sequence.

[0034] The first computing module is used to input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tags.

[0035] The second computing module is used to convert the interest tags into an input format of a converter model, feature extract the user interest tag sequence, and obtain a global feature vector of the user interest tags.

[0036] The third computing module is used to generate the user portrait according to the global feature vector of the user interest tag and the user identity data.

[0037] The third aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory and a processor; the memory is used to store computer instructions; the processor is used to run the computer instructions stored in the memory to implement the above method for generating user portraits based on behavioral data.

[0038] A fourth aspect of the present invention provides a storage medium, comprising: a readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; the computer instructions are used to implement the above-mentioned method for generating a user portrait based on behavioral data.

[0039] The beneficial effects provided by the present invention are as follows: user behavior sequence data is continuously updated in a time domain rolling manner, thereby avoiding computational redundancy, and at the same time being able to effectively capture the development trend and pattern changes of user behavior, improve the accuracy and timeliness of user portraits, and achieve real-time or near real-time response to user behavior changes; increase the timeliness and accuracy of detection, and provide a guarantee for more accurate subsequent judgment of user preferences through user behavior; and provide a high-quality data foundation for subsequent interest tag prediction, construction of group user portraits, and user portrait analysis.

[0040] At the same time, combined with the long short-term memory network-Bayesian module and the transformer model, user behavior interest tags are accurately output, and a group user portrait with depth and breadth is formed. The transformer is used to output typical representative behavior interest tags of user groups in parallel, which can meet the requirements of large data volume and real-time update of group user behavior analysis, and has sufficient flexibility and robustness to improve the accuracy and personalization of integrated media content recommendations. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0042] Figure 1 A flow chart of a method for generating a user portrait based on behavior data disclosed in an embodiment of the present invention;

[0043] Figure 2 A method flow chart of step S102 in the method for generating a user portrait based on behavior data disclosed in an embodiment of the present invention;

[0044] Figure 3 A block diagram of a system for generating a user portrait based on behavioral data disclosed in an embodiment of the present invention.

[0045] The above drawings show clear embodiments of the present disclosure, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0046] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0047] The method for generating a user portrait based on behavior data disclosed in the present invention is as follows: Figure 1 shown.

[0048] S101: Collect first behavior data of the user, update the first behavior data in real time, and record a timestamp of the first behavior data.

[0049] S102: Determine whether second user behavior data is generated, and update the user behavior data.

[0050] Set a fixed time window , in a fixed time window The user’s behavior data is collected and expressed as formula (1). Formula (1)

[0051] in Representing the user The behavior record.

[0052] When the user is in a fixed time window When new user behavior data is generated within a fixed time window, the behavior data is appended to the end of the current user behavior sequence, and the old data outside the fixed time window is removed.

[0053] Specific logical process, such as Figure 2 S201: Setting a fixed sampling time period , and record the fixed sampling time period The first user behavior data within a period; define a time window , which represents the validity period of user behavior data.

[0054] The size of the time window depends on the specific application scenario and is usually determined based on the frequency and change rate of user behavior.

[0055] Furthermore, initialize the behavior sequence, for each user , at time When it is 0, an empty user behavior sequence is initialized, as shown in formula (2).

[0056] Formula (2)

[0057] That is to say, the fixed time window here can be understood as the sampling time, and the fixed time window is the fixed time The sampling time.

[0058] S202: Determine whether second user behavior data is generated, and if the second user behavior data is generated, replace the non-fixed sampling time period in the first user behavior data Deleting user behavior data and updating the user behavior data include:

[0059] Record the timestamp of the second user behavior data, and determine whether the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period .

[0060] If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period , record the second user behavior data.

[0061] If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is not less than the fixed sampling time period , delete the first user behavior data.

[0062] According to the timestamp and fixed sampling time period of the second user behavior data , delete the non-fixed sampling time period in the first user behavior data User behavior data.

[0063] Specifically, when the user In time Generate new behaviors When the data of this behavior is encoded into a feature vector , and add it to the current user behavior sequence At the same time, check the earliest behavior record in the sequence Timestamp Is it less than If yes, it means that the behavior record has exceeded the time window and should be removed from the sequence. The mathematical description of the update is shown in formula (3).

[0064] Formula (3)

[0065] in, It is user behavior The time of occurrence, is the updated user behavior sequence, It is behavior The timestamp of the

[0066] In order to adapt to the dynamic changes of time series data, each behavior record It can be expressed as a vector, as shown in formula (4).

[0067] Formula (4)

[0068] in Indicates The behavior in For each user, we continuously track their behavior sequence and ensure that all behavior data in the sequence are within the time window. Therefore, formula (3) can also be expressed as formula (5).

[0069] Formula (5)

[0070] in, Is a user Updated behavior sequence, is the new behavior feature vector, is the earliest behavioral feature vector.

[0071] Assume user In time Behavior , the eigenvector corresponding to this behavior is recorded. At this time, the user's behavior sequence is shown in formula (6):

[0072] Formula (6)

[0073] in, is the most recent behavior record. We check the first behavior record in the sequence If the timestamp is less than ,Right now , it is removed from the sequence. Then, the new behavior record Add to the end of the sequence to complete the dynamic update of the behavior sequence.

[0074] Specifically, for example, the time when the second user behavior data is generated is 8:15 am, and the fixed sampling time period is If the time limit is 30 minutes, the user behavior data of the first user behavior before 7:30 am is the data to be deleted.

[0075] By user For example, in the most recent time window Within, the user produced the following sequence of behaviors: , each behavior are encoded into corresponding feature vectors When new behavior When it occurs, we first determine whether it is within the time window If so, then its feature vector Add to the current sequence of behaviors and remove the first behavior outside the time window The eigenvector of In this way, the user behavior sequence data is dynamically updated.

[0076] It avoids computational redundancy, increases the timeliness and accuracy of detection, and provides a guarantee for more accurate judgment of user preferences through user behavior in the future. It ensures the timeliness and dynamics of user behavior sequence data, and provides a basis for subsequent interest tag prediction and the construction of group user portraits.

[0077] In this way, we ensure the real-time and continuity of user behavior sequence data, providing a high-quality data foundation for subsequent interest tag prediction and user portrait analysis.

[0078] After the user behavior sequence data is dynamically updated, in order to predict the user's behavior interest tags, we use the long short-term memory network (LSTM) combined with the Bayesian reasoning method. The LSTM network can capture the long-term dependencies in the time series data, and the Bayesian method can model the uncertainty of user interests on this basis.

[0079] Step S103: Input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag. LTSM-Bayesian module processing: Input the updated user behavior sequence data into the LTSM-Bayesian module and express it using a mathematical formula such as formula (7).

[0080] Formula (7)

[0081] in, represents the probability that a user has a specific label under a given behavior sequence, It represents the probability that a user with a specific label will generate this behavior data sequence. is the prior probability, which indicates the inherent probability that the user has this label, is the marginal probability of the user behavior data sequence.

[0082] Output user behavior interest tags. The LTSM-Bayes module processes user behavior sequence data and interest tag prediction.

[0083] The specific processing steps are as follows: First, update the user behavior sequence data Input into the LSTM network. The LSTM network learns the long-term dependencies in the sequence data through its special gate structure (forget gate, input gate and output gate). For the The mathematical description of LSTM can be simplified to formula (8).

[0084] Formula (8)

[0085] Based on the features extracted by LSTM, Bayesian reasoning is applied to predict user interest tags. The Bayesian formula is expressed as formula (7).

[0086] Prediction of interest tags,By calculating the above probabilities, we can assign a set of interest tags to each user. In practice, the tag with the largest posterior probability is usually selected as the user's interest tag:

[0087] Formula (9)

[0088] Assume that after being processed by the LSTM network, the user Behavior sequence data is converted into a sequence of feature vectors Next, we use Bayesian inference to predict user interest tags.

[0089] Calculate the conditional probability: For each possible interest tag ,calculate , which is usually estimated through statistics on the training dataset.

[0090] Calculate the prior probability: Similarly, It can be estimated from the training data or set based on domain knowledge.

[0091] Normalization: Since the denominator Since is a constant for all labels, we can omit it and directly compare the sizes of the molecules to determine the most likely label, as shown in Equation (10).

[0092] Formula (10)

[0093] Through the above steps, we provide users The most likely interest tag is predicted*. Such predictions can provide a basis for applications such as personalized recommendations and advertising targeting.

[0094] Through the above detailed description, readers can clearly understand how to use the LSTM-Bayesian module to process user behavior sequence data and predict user behavior interest tags.

[0095] In order to deeply explore the global features and complex relationships between user behavior interest tags, the Transformer model is used for processing. The Transformer model, with its self-attention mechanism and parallel computing capabilities, shows excellent performance in processing sequence data.

[0096] Step S104: converting the interest tags into an input format of a converter model, feature extracting the user interest tag sequence, and obtaining a global feature vector of the user interest tags.

[0097] Transformer model feature extraction: Using the transformer model (transformer model), the global features and complex relationships of user behavior interest tags are extracted in high parallelism. The mathematical formula is described as follows:

[0098] Formula (11)

[0099] The following are the specific steps for extracting user behavior interest tag features using the transformer model:

[0100] Preparation of interest tag sequence. First, the user interest tag sequence output by the LTSM-Bayes module is converted into an input format suitable for the transformer model. is represented as an embedding vector , forming a sequence, as expressed in formula (12).

[0101] Formula (12)

[0102] In order to enable the model to understand the positional relationship of elements in the sequence, we embed each Adding positional encoding , and obtain a vector sequence with position information, as expressed in formula (13).

[0103] Formula (13)

[0104] The mathematical formula of position encoding can be expressed as formula (14):

[0105] Formula (14)

[0106] in, is the position index, is the vector dimension index, is the dimension of the embedding vector.

[0107] Self-attention mechanism: The core of the Transformer model is the self-attention mechanism, which allows the model to take into account the information of other elements in the sequence when processing each element in the sequence. The mathematical description of the self-attention mechanism is formula (15):

[0108] Formula (15)

[0109] in, They are query, key and value matrices respectively, which are obtained by linear transformation of the input sequence.

[0110] Multi-head attention. In order to improve the expressiveness of the model, we adopt a multi-head attention mechanism, passing the input sequence through multiple self-attention layers at the same time, and then concatenating the results, as shown in Equation (16).

[0111] Formula (16)

[0112] in, It's the number of heads. is the output weight matrix.

[0113] Through the multi-head attention mechanism, we obtain the global feature representation of the user behavior interest tag sequence. These features represent the complex relationship between user interest tags and can be used for subsequent user profile construction.

[0114] The mathematical formula for feature extraction is formula (17):

[0115] Formula (17)

[0116] By user Take the interest tag sequence as an example, first, each tag Convert to embedding vector , and add position code Subsequently, these embedding vectors with position information are fed into the transformer model for processing. Through self-attention and multi-head attention mechanisms, the model is able to capture the global features and relationships between interest tags.

[0117] For example, for the tag sequence , the corresponding embedding vector sequence is After position encoding, we get Through the transformer model, these vectors are transformed into a global feature vector that integrates the information of all labels in the sequence.

[0118] In this way, we can not only obtain a comprehensive representation of user interests, but also reveal the intrinsic connections between different interest tags, providing strong support for building more refined user portraits.

[0119] Step S105: Generate the user portrait based on the global feature vector of the user interest tag and the user identity data. Aggregate similar user interest tags in the full set feature vector to identify the common interest tags of the users.

[0120] The user portrait is generated according to the common interest tags and user identity data.

[0121] Group user portrait generation: Based on the extracted feature vectors, a group user portrait is constructed to optimize the converged media content recommendation algorithm.

[0122] Group user portrait generation method: After the transformer model extracts the global feature vectors of user behavior interest tags, we further apply these feature vectors to construct group user portraits. Group user portraits are designed to reflect the common interests and behavior patterns of a group of users on the converged media platform, which is of great significance for optimizing content recommendation algorithms.

[0123] The following are the detailed steps for generating group user portraits: Feature vector aggregation: For user groups with similar interest tags, their feature vectors are aggregated. The aggregation method can be simple averaging or weighted averaging. The weight can be determined based on factors such as user activity and behavior frequency. is a set of individual user feature vectors, and the aggregated group feature vector It can be expressed as formula (18):

[0124] Formula (18)

[0125] If weighted average is used, the formula becomes formula (19):

[0126] Formula (19)

[0127] in, It is The weight of each user. Group interest tag identification. Using the aggregated group feature vector, the common interest tags of the group are identified through the clustering algorithm (K-means). These tags will serve as the core component of the group user portrait.

[0128] Portrait construction. Combine the identified group interest tags with other statistical information of the user group (such as age, gender, geographic location, etc.) to construct a multi-dimensional group user portrait.

[0129] Taking a group of users with similar interest tags as an example, firstly, their feature vectors are weighted averaged to obtain the group feature vector . Assume that we use the user's activity as the weight , then the feature vector of each user is The contribution to the group feature vector will be weighted according to its activity.

[0130] Subsequently, the K-means clustering algorithm is applied to cluster the group feature vectors to identify the group’s common interest tags, which represent the group’s common preferences in consuming converged media content.

[0131] Finally, combined with other statistical information of the user group, such as age distribution, gender ratio, geographical distribution, etc., we can construct a comprehensive group user portrait. This portrait not only includes the group's interest preferences, but also reflects the group's basic demographic characteristics.

[0132] Suppose we select a group of users who are interested in technology content and extract the feature vectors of their interest tags through the transformer model. We perform weighted average on these feature vectors to obtain a group feature vector representing the interests of the group. Using the K-means algorithm, we find that this group is mainly concentrated on two interest tags: electronic products and artificial intelligence.

[0133] Combined with other information about the user group, such as most users are males aged 30-40, mainly distributed in first-tier cities, we built a group user portrait that includes this information. This portrait will help the integrated media platform more accurately recommend content that meets the interests of this group, improve user satisfaction and user stickiness of the platform.

[0134] Through the above methods, we not only optimized the content recommendation algorithm, but also provided users with a more personalized service experience through group user portraits.

[0135] The second aspect of the present invention provides a user profile generation system 300 based on behavior data, such as Figure 3 As shown. Includes:

[0136] The acquisition module 301 is used to acquire the first behavior data of the user and the user identity data, and to update the first behavior data in real time and record the timestamp of the first behavior data.

[0137] The updating module 302 is used to determine whether the second user behavior data is generated and update the user behavior data sequence.

[0138] The first calculation module 303 is used to input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag.

[0139] The second calculation module 304 is used to convert the interest tags into an input format of a converter model, feature extract the user interest tag sequence, and obtain a global feature vector of the user interest tags.

[0140] The third calculation module 305 is used to generate the user portrait according to the global feature vector of the user interest tag and the user identity data.

[0141] The third aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory and a processor; the memory is used to store computer instructions; the processor is used to run the computer instructions stored in the memory to implement the above method for generating user portraits based on behavioral data.

[0142] A fourth aspect of the present invention provides a storage medium, comprising: a readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; the computer instructions are used to implement the above-mentioned method for generating a user portrait based on behavioral data.

[0143] The beneficial effects provided by the present invention: The beneficial effects provided by the present invention: Continuously updating user behavior sequence data in a time domain rolling manner avoids calculation redundancy, increases the timeliness and accuracy of detection, and provides a guarantee for more accurate judgment of user preferences through user behavior in the subsequent process; provides a high-quality data foundation for subsequent interest tag prediction, construction of group user portraits, and user portrait analysis.

[0144] At the same time, combined with the long short-term memory network-Bayesian module and the converter model, user behavior interest tags can be accurately output, and a group user portrait with depth and breadth can be formed to improve the accuracy and personalization of integrated media content recommendations.

[0145] Obviously, the above specific implementation cases are only examples for illustrating the application of the present method, rather than limiting the implementation methods. For those skilled in the art, other different forms of changes and modifications can be made on the basis of the above description to study other related issues. Therefore, the protection scope of the present invention should be the protection scope of the claims.

[0146] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.

[0147] The electronic device and other embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the embodiments of the present invention have been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

[0150] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0151] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating a user profile based on behavioral data, characterized in that: The method comprises: S101: Acquire first behavior data and user identity data of a user, update the first behavior data in real time, and record a timestamp of the first behavior data; S102: Determine whether second user behavior data is generated, and update the user behavior data sequence; S103: Input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag; S104: converting the interest tags into an input format of a converter model, feature extracting a user interest tag sequence, and obtaining a global feature vector of the user interest tag; S105: Generate the user portrait according to the global feature vector of the user interest tag and the user identity data; Extract the global feature vector of group users and build group user portraits to optimize the converged media content recommendation algorithm, including: For user groups with similar interest tags, their feature vectors are aggregated; Using the aggregated group feature vector, the common interest tags of the group are identified through the clustering algorithm K-means; Combine the identified group interest tags with other statistical information of the user group to build a multi-dimensional group user portrait. Other statistical information includes age, gender, and geographic location. In the step S102, it is determined whether the second user behavior data is generated, and the user behavior data sequence is updated, including: S201: Setting a fixed sampling time period , and record the fixed sampling time period First user behavior data within; S202: Determine whether second user behavior data is generated. If the second user behavior data is generated, delete the user behavior data of the non-fixed sampling time period in the first user behavior data, and update the user behavior data sequence, including: Record the timestamp of the second user behavior data, and determine whether the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period , recording the second user behavior data; If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is not less than the fixed sampling time period , deleting the first user behavior data; in the step S103, inputting the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag, including: The updated user behavior sequence data is input into the LSTM network. The LSTM network learns the long-term dependencies in the sequence data through the forget gate, input gate, and output gate. For the The hidden state corresponding to each behavior, the mathematical description of LSTM can be simplified to formula (5): Formula (5) Based on the features extracted by LSTM, Bayesian reasoning is applied to predict user interest tags. The Bayesian formula is expressed as follows: Formula (1) in, represents the probability that a user has a specific label under a given behavior sequence, It represents the probability that a user with a specific label will generate this behavior data sequence. is the prior probability, which indicates the inherent probability that the user has this label, is the marginal probability of the user behavior data sequence; by calculating the above probability, the tag with the maximum posterior probability is selected as the user's interest tag: Formula (6) After being processed by the LSTM network, the user Behavior sequence data is converted into a sequence of feature vectors , using Bayesian reasoning to predict user interest tags; Calculate the conditional probability: For each possible interest tag ,calculate , estimated by statistics on the training dataset; Calculate the prior probability: Estimated from training data, or set based on domain knowledge; Normalization: The denominator is a constant for all labels, and the numerator is directly compared to determine the most likely label, as shown in formula (7): Formula (7).

2. The method according to claim 1, characterized in that In step S104, the interest tags are converted into an input format of a converter model, and feature extraction of the user interest tag sequence is performed to obtain a global feature vector of the user interest tag, including: The interest tags are embedded in the vector , forming a sequence, expressed as formula (2): Formula (2) in, represents the vector where each interest is embedded, Represents the sequence of interest tags formed after being embedded in the vector; Then in the embedding vector Add positional encoding , as shown in formula (3): Formula (3) in, Represents the sequence of interest tags formed by adding position information to the embedded vector; The position code , as shown in formula (4): Formula (4) in, is the position index, is the vector dimension index, is the embedding vector Dimension.

3. The method according to claim 2, characterized in that The step S105 generates the user portrait according to the global feature vector of the user interest tag and the user identity data, including: Aggregating similar user interest tags in the full set of feature vectors to identify common interest tags of the users; The user portrait is generated according to the common interest tags and user identity data.

4. A user profile generation system based on behavioral data, characterized in that: The system comprises: An acquisition module, used to acquire the user's first behavior data and user identity data, and to update the first behavior data in real time and record a timestamp of the first behavior data; An updating module, used to determine whether to generate second user behavior data and update the user behavior data sequence; The first computing module is used to input the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag; A second computing module, configured to convert the interest tags into an input format of a converter model, feature extract the user interest tag sequence, and obtain a global feature vector of the user interest tag; The third calculation module is used to generate the user portrait according to the global feature vector of the user interest tag and the user identity data; extract the global feature vector of the group user, construct the group user portrait, and optimize the converged media content recommendation algorithm, including: For user groups with similar interest tags, their feature vectors are aggregated; Using the aggregated group feature vector, the common interest tags of the group are identified through the clustering algorithm K-means; Combine the identified group interest tags with other statistical information of the user group to build a multi-dimensional group user portrait. Other statistical information includes age, gender, and geographic location. The determining whether to generate second user behavior data and updating the user behavior data sequence includes: S201: Set fixed sampling time period , and record the fixed sampling time period S202: determining whether second user behavior data is generated, and if the second user behavior data is generated, deleting the user behavior data of the non-fixed sampling time period in the first user behavior data, and updating the user behavior data sequence, including: Record the timestamp of the second user behavior data, and determine whether the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period ; If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is less than the fixed sampling time period , recording the second user behavior data; If the difference between the timestamp of the second user behavior data and the timestamp of the first user behavior data is not less than the fixed sampling time period , deleting the first user behavior data; in the third calculation module, inputting the updated user behavior data sequence into the long short-term memory network-Bayesian module to obtain the user's interest tag, including: The updated user behavior sequence data is input into the LSTM network. The LSTM network learns the long-term dependencies in the sequence data through the forget gate, input gate, and output gate. For the The hidden state corresponding to each behavior, the mathematical description of LSTM can be simplified to formula (5): Formula (5) Based on the features extracted by LSTM, Bayesian reasoning is applied to predict user interest tags. The Bayesian formula is expressed as follows: Formula (1) in, represents the probability that a user has a specific label under a given behavior sequence, It represents the probability that a user with a specific label will generate this behavior data sequence. is the prior probability, which indicates the inherent probability that the user has this label, is the marginal probability of the user behavior data sequence; By calculating the above probabilities, the tag with the largest posterior probability is selected as the user's interest tag: Formula (6) After being processed by the LSTM network, the user Behavior sequence data is converted into a sequence of feature vectors , using Bayesian reasoning to predict user interest tags; Calculate the conditional probability: For each possible interest tag ,calculate , estimated by statistics on the training dataset; Calculate the prior probability: Estimated from training data, or set based on domain knowledge; Normalization: The denominator is a constant for all labels, and the numerator is directly compared to determine the most likely label, as shown in formula (7): Formula (7).

5. A smart device, characterized in that: include: transmitter, receiver, memory and processor; The memory is used to store computer instructions; The processor is configured to execute the computer instructions stored in the memory to implement any one of claims 1 to 3.

6. A storage medium, characterized in that: include: A readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; The computer instructions are used to implement any one of claims 1 to 3.

Citation Information

Patent Citations

  • Consumer behavior portrait tool based on big data

    CN109615432A

  • User portrait updating method and system, network equipment and storage medium

    CN111523026A

  • Electronic commerce user portrait construction method based on big data

    CN118485464A

  • Click rate prediction method based on multi-gradient interest context network

    CN118552261A