Knowledge recommendation method fusing text tag and user behavior and related equipment

By preprocessing and subject mining of user interaction knowledge text, generating target text tags, building user knowledge scoring matrix and introducing virtual users, the low accuracy and cold start problems of knowledge recommendation in professional fields are solved, and the adaptability and coverage of the recommendation system are improved.

CN120508647AActive Publication Date: 2025-08-19CENT SOUTH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510984796.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-19
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

The existing knowledge recommendation methods have difficulties in obtaining tags, poor user interest modeling capabilities and sparse data in the professional field, resulting in low accuracy of knowledge recommendation.

Method used

By obtaining the knowledge text and behavior records of the target user's interaction with the electronic device, pre-processing and subject mining, generating initial text tags, selecting target text tags, building a user knowledge scoring matrix, and introducing virtual users for personalized recommendations.

Benefits of technology

It improves the accuracy of knowledge recommendation, overcomes the problem of sparse user data, avoids cold starts, and enhances the adaptability and coverage of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508647A_ABST
    Figure CN120508647A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent recommendation, and provides a knowledge recommendation method fusing text tags and user behaviors and related equipment. The method comprises the following steps: preprocessing each knowledge text to obtain a plurality of initial text tags of each knowledge text; performing subject mining on the knowledge text to obtain subject terms of the knowledge text, and generating text representation of each initial text label of the knowledge text based on the subject terms; based on the text representation of all the initial text tags of each knowledge text, selecting a target text tag of each knowledge text from all the initial text tags of each knowledge text; according to the target text labels of all the knowledge texts, all the behavior records of the target user and the virtual user, constructing a user knowledge score matrix; and performing personalized knowledge recommendation for the target user according to the user knowledge score matrix. According to the method, the accuracy of knowledge recommendation for the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent knowledge recommendation technology, and in particular to a knowledge recommendation method and related equipment that integrates text tags and user behaviors. Background Art

[0002] With the widespread adoption of internet information services, users face challenges with inefficient information screening and inaccurate knowledge acquisition in a massive data environment. User knowledge needs are not only dynamic and ever-changing, but also highly personalized. This is especially true in specialized vertical fields like healthcare and fintech, where knowledge texts often contain a large number of domain-specific terms and complex semantic relationships. User needs are highly specialized and real-time. Building a precise, domain-specific knowledge recommendation system has become crucial for improving the effectiveness of information services in these verticals.

[0003] Knowledge recommendation methods have been widely applied in various fields. Traditional recommendation methods primarily rely on collaborative filtering and content-based models, but these methods often perform poorly when faced with diverse and dynamic knowledge service demands. In recent years, the introduction of deep learning technology has driven rapid development in the recommendation field. Through precise user behavior modeling and deep semantic analysis, the accuracy and diversity of knowledge recommendations have been significantly improved. However, existing research primarily focuses on feature extraction and matching mechanisms for general scenarios, while fine-grained semantic mining and dynamic interest modeling for domain-specific knowledge recommendation remain significantly insufficient.

[0004] In text knowledge recommendation, text tags, as semantic connectors between domain knowledge and user profiles, have the advantages of structured representation and interpretability. High-quality text tags can not only present users' potential interest preferences in a fine-grained form, but also effectively reflect the similarity between knowledge content, providing a basis for accurately matching user needs. However, text recommendation for professional fields still faces many challenges: (1) Difficulty in obtaining tags: Domain knowledge texts usually contain a large number of professional terms, are long and semantically complex, and existing methods have the problem of insufficient topic level recognition in long text processing, resulting in low efficiency in building a tag system. At the same time, the inherent long-tail distribution characteristics of text tags make it difficult to effectively utilize low-frequency professional tags. (2) Poor user interest modeling capabilities: Users' interest preferences are highly domain-related and dynamic. Traditional static modeling methods often lack in-depth understanding of users' specific needs and personalized services, making it difficult to adapt to changes in demand in professional scenarios, resulting in poor recommendation effects. (3) Data sparsity and cold start dilemma: The behavioral data density of user groups is significantly lower than that of general scenarios. The recommendation quality of traditional collaborative filtering methods drops sharply under limited data conditions. This leads to the problem of low accuracy in knowledge recommendation for users. Summary of the Invention

[0005] This application provides a knowledge recommendation method and related equipment that integrates text tags and user behaviors, which can solve the problem of low accuracy of knowledge recommendation to users.

[0006] In a first aspect, an embodiment of the present application provides a knowledge recommendation method that integrates text tags and user behavior. The knowledge recommendation method includes: Acquire multiple knowledge texts that the target user interacts with the electronic device, as well as a behavior record of the target user interacting with each knowledge text with the electronic device; Preprocess each knowledge text to obtain multiple initial text labels for each knowledge text; the initial text labels are used to describe the semantics in the knowledge text; For each knowledge text, perform subject mining on the knowledge text to obtain the subject words of the knowledge text, and generate the text representation of each initial text label of the knowledge text based on the subject words; Based on the text representation of all initial text labels of each knowledge text, a target text label of each knowledge text is selected from all initial text labels of each knowledge text; the matching degree between the knowledge text and the target text label is greater than the matching degree between the knowledge text and each other initial text label; Introducing virtual users, constructing a user knowledge scoring matrix based on the target text labels of all knowledge texts, all target users' behavior records, and virtual users; the elements in the user knowledge scoring matrix are the scores of the target users for each initial text label; Personalized knowledge recommendation is performed for target users based on the user knowledge scoring matrix.

[0007] Optionally, perform subject mining on the knowledge text to obtain the key words of the knowledge text, including: By formula: ; ; ; ; Keywords of computing knowledge texts ; in, represents the latent variable, and represents the variational parameter, represents the parameter distribution, 、 、 represents a multi-layer perceptron computing network, The feature vector representing the knowledge text, represents the latent variable mean vector, represents the latent variable variance vector, Represents the activation function.

[0008] Optionally, a text representation of each initial text tag of the knowledge text is generated based on the subject words, including: Perform context sensing on the knowledge text to obtain the vector representation of each word in the knowledge text; The text representation of each initial text tag is calculated based on all vector representations and knowledge text keywords.

[0009] Optionally, a text representation for each initial text tag is calculated based on all vector representations and knowledge text keywords, including: By formula: ; Calculate the Text representation of the initial text labels ; in, , Indicates the number of initial text labels corresponding to the knowledge text, 、 represents the weight matrix, 、 represents the bias term, represents the transpose operation, Indicates the Topic-guided representation of initial text labels: ; ; ; in, represents the normalized coefficient guided by the subject word, Indicates the The vector representation of the subject words of the knowledge text, Indicates the The number of knowledge texts corresponding to the initial text labels, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, represents the attention parameter, represents a constant, 、 Indicates parameters, Represents the subject term of the knowledge text.

[0010] Optionally, based on the text representation of all initial text labels of each knowledge text, a target text label of each knowledge text is selected from all initial text labels of each knowledge text, including: For each knowledge text, perform the following steps: Calculating the matching degree between the knowledge text and each initial text label according to the text representation of each initial text label of the knowledge text; All matching degrees are sorted from large to small, and the initial text labels corresponding to the first multiple matching degrees are used as the target text labels of the knowledge text.

[0011] Optionally, based on the text representation of each initial text label of the knowledge text, a matching degree between the knowledge text and each initial text label is calculated, including: By formula: ; Computational Knowledge Text and The matching degree between the initial text labels ; in, Indicates the The text representation of the initial text labels, Represents knowledge text, , Indicates the number of initial text labels corresponding to the knowledge text, represents a constant, represents the weight vector, Represents a transpose operation.

[0012] Optionally, a user knowledge scoring matrix is constructed based on the target text labels of all knowledge texts, all behavior records of target users, and virtual users, including: A user interaction scoring matrix is constructed based on all the behavior records of the target user and the virtual user. The elements in the user interaction scoring matrix are the interaction scores between the target user and each knowledge text, or the interaction scores between the virtual user and each knowledge text. A text label matrix is constructed based on the target text labels of all knowledge texts; the elements in the text label matrix are scored for the matching between each knowledge text and each initial text label; A user knowledge scoring matrix is constructed based on the user interaction scoring matrix and the text label matrix; the elements in the user knowledge scoring matrix are the scores of the target user for each initial text label.

[0013] Optionally, the behavior record includes the user's browsing time, collection behavior, and like behavior of the knowledge text; Build a user interaction scoring matrix based on all target user behavior records, including: By formula: ; Calculating users With knowledge text Interaction score between ; in, Represents a user No. The number of days between the browsing behavior and the current time, , Represents a user The total number of browsing behaviors, represents the time decay rate constant, Represents a user With knowledge text The interaction weight of: ; in, 、 、 Both represent weights, Represents a user Browse knowledge text duration, Represents a user Knowledge Text Performed a like action. Represents a user No knowledge text Perform a like action, Represents a user Knowledge Text Collected behavior, Represents a user No knowledge text Collecting behavior, , The total number of numbers representing knowledge texts; Construct a text label matrix based on the target text labels of all knowledge texts, including: By formula: ; Computational Knowledge Text With initial text label Match score between ; in, , Represents the set of all initial text labels corresponding to all knowledge text labels.

[0014] Optionally, a user knowledge scoring matrix is constructed based on the user interaction scoring matrix and the text label matrix, including: By formula: ; Calculate the target user's initial text label Rating ; in, The total number of numbers representing knowledge texts, Indicates target users Knowledge Text The interaction score, Represents a user Knowledge Text The interaction score, Indicates the number of users, , represents the user label matrix, , , Represents the transposed matrix of the text label matrix Corresponding knowledge text With initial text label The matching score between the elements, represents the user interaction rating matrix, Represents a text label matrix.

[0015] In a second aspect, an embodiment of the present application provides a knowledge recommendation device that integrates text tags and user behavior, including: An acquisition module, configured to acquire multiple knowledge texts that a target user interacts with an electronic device, and a behavior record of the target user interacting with each knowledge text with the electronic device; The preprocessing module is used to preprocess each knowledge text to obtain multiple initial text labels for each knowledge text; the initial text labels are used to describe the semantics in the knowledge text; The subject mining module is used to perform subject mining on each knowledge text, obtain the subject words of the knowledge text, and generate the text representation of each initial text label of the knowledge text based on the subject words; a selection module for selecting a target text label for each knowledge text from all initial text labels of each knowledge text based on the text representation of all initial text labels of each knowledge text; a matching degree between the knowledge text and the target text label is greater than a matching degree between the knowledge text and each other initial text label; The construction module is used to introduce virtual users and construct a user knowledge scoring matrix based on the target text labels of all knowledge texts, all behavior records of the target users, and the virtual users. The elements in the user knowledge scoring matrix are the scores of the target users for each initial text label. The personalized recommendation module is used to make personalized knowledge recommendations for target users based on the user knowledge scoring matrix.

[0016] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned knowledge recommendation method integrating text tags and user behaviors when executing the above-mentioned computer program.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned knowledge recommendation method that integrates text tags and user behaviors.

[0018] The above solution of the present application has the following beneficial effects: In some embodiments of the present application, by obtaining multiple knowledge texts that the target user interacts with the electronic device, and the behavior records of the target user when interacting with the electronic device for each knowledge text, each knowledge text is preprocessed to obtain multiple initial text tags for each knowledge text, and then for each knowledge text, the knowledge text is subject mined to obtain the subject words of the knowledge text, and the text representation of each initial text tag of the knowledge text is generated based on the subject words. Then, based on the text representation of all initial text tags of each knowledge text, the target text tag of each knowledge text is selected from all initial text tags of each knowledge text, and then a virtual user is introduced. According to the target text tags of all knowledge texts, all behavior records of the target user and the virtual user, a user knowledge scoring matrix is constructed, and finally, personalized knowledge recommendation is performed for the target user based on the user knowledge scoring matrix. Among them, the user knowledge scoring matrix is constructed based on the behavior records of the user interaction and the target text tags of the knowledge text, taking into account the text semantics of the user interaction and the knowledge text, effectively analyzing the text tags of the user's interest, and performing personalized knowledge recommendation for the user based on the user knowledge scoring matrix, thereby improving the accuracy of the knowledge recommendation for the user.

[0019] In addition, introducing virtual users to construct the user knowledge rating matrix can overcome the problem of sparse user data, avoid the cold start of knowledge recommendation, and improve the quality of recommendation.

[0020] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 A flowchart of a knowledge recommendation method that integrates text tags and user behavior provided in one embodiment of the present application; Figure 2 A schematic diagram of the structure of a knowledge recommendation device that integrates text tags and user behaviors, provided in one embodiment of the present application; Figure 3 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0023] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0024] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0025] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0027] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0028] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0029] In response to the problem of low accuracy of existing knowledge recommendation for users, the embodiment of the present application provides a knowledge recommendation method that integrates text tags and user behavior. The knowledge recommendation method obtains multiple knowledge texts that the target user interacts with the electronic device, as well as the target user's behavior record when interacting with the electronic device for each knowledge text, and then preprocesses each knowledge text to obtain multiple initial text tags for each knowledge text. Then, for each knowledge text, the knowledge text is subject mined to obtain the subject words of the knowledge text, and the text representation of each initial text tag of the knowledge text is generated based on the subject words. Then, based on the text representation of all initial text tags of each knowledge text, the target text tag of each knowledge text is selected from all initial text tags of each knowledge text. Then, a virtual user is introduced, and a user knowledge scoring matrix is constructed based on the target text tags of all knowledge texts, all behavior records of the target user, and the virtual user. Finally, personalized knowledge recommendation is performed for the target user based on the user knowledge scoring matrix. Among them, the user knowledge scoring matrix is constructed based on the behavior record of the user interaction and the target text tag of the knowledge text, taking into account the text semantics of the user interaction and the knowledge text, effectively analyzing the text tags of the user's interest, and performing personalized knowledge recommendation for the user based on the user knowledge scoring matrix, thereby improving the accuracy of knowledge recommendation for the user.

[0030] Next, an exemplary description is given of the knowledge recommendation method provided in this application that integrates text tags and user behaviors.

[0031] like Figure 1 As shown, the knowledge recommendation method provided by this application that integrates text tags and user behavior includes the following steps: Step 11: Acquire multiple knowledge texts that the target user interacts with the electronic device, as well as a behavior record of the target user interacting with each knowledge text with the electronic device.

[0032] The aforementioned knowledge texts are electronic texts that target users browse when interacting with electronic devices, such as news on disease prevention and control, articles on health and wellness, etc. The aforementioned behavioral records include the user's browsing time, collection behavior, and like behavior of the knowledge texts.

[0033] In some embodiments of the present application, a log recording script or software may be used to obtain the knowledge text of the target user interaction and the corresponding behavior record.

[0034] For example, the data format of the behavior record is in vector format: ; in, Indicates target users Knowledge Text The vector form of Indicates target users Knowledge Text The time of interaction, Indicates target users Knowledge Text The duration of browsing, Indicates target users Knowledge Text Performed a like action. Indicates target users No knowledge text Perform a like action, Indicates target users Knowledge Text Collected behavior, Indicates target users No knowledge text Collecting behavior.

[0035] Step 12: pre-process each knowledge text to obtain multiple initial text labels for each knowledge text.

[0036] The above initial text tags are used to describe the semantics in the knowledge text.

[0037] For example, a support vector machine or other method can be used to preprocess the knowledge text to obtain multiple initial text labels for the knowledge text. A comprehensive frequency statistics is performed on the initial text labels of all knowledge texts, and the initial text label with the highest frequency among all the initial text labels of the knowledge text is used as the standard label of the knowledge text. For example, if the initial text labels of knowledge text A are "epidemic prevention" and "disinfection", and they appear 8 and 6 times respectively among the initial text labels of all knowledge texts, then "epidemic prevention" is used as the standard label of knowledge text A.

[0038] Step 13: For each knowledge text, subject mining is performed on the knowledge text to obtain the subject words of the knowledge text, and a text representation of each initial text label of the knowledge text is generated based on the subject words.

[0039] In some embodiments of the present application, the steps of performing subject mining on the knowledge text to obtain the subject words of the knowledge text and generating a text representation of each initial text tag of the knowledge text based on the subject words include:

[0040] The first step is to conduct subject mining on the knowledge text to obtain the key words of the knowledge text.

[0041] By formula: ; ; ; ; Keywords of computing knowledge texts .

[0042] in, represents the latent variable, and represents the variational parameter, represents the parameter distribution, , 、 、 represents a multi-layer perceptron computing network, To calculate the computing network of keywords, is the computational network for calculating the mean vector of latent variables, is the computational network for calculating the latent variable variance vector, The feature vector representing the knowledge text (which can be obtained by calculating the knowledge text using algorithms such as the bag-of-words model, the word frequency-inverse document frequency algorithm, and the variational autoencoder. For example, the knowledge text is input into the bag-of-words model for calculation, the word frequency-inverse document frequency algorithm is used to calculate the output data of the bag-of-words model, and then the variational autoencoder is used to calculate the output data of the word frequency-inverse document frequency algorithm to obtain the feature vector of the knowledge text). represents the latent variable mean vector, represents the latent variable variance vector, Represents the activation function.

[0043] The second step is to perform context sensing on the knowledge text and obtain the vector representation of each word in the knowledge text.

[0044] For example, a bidirectional encoder representation model with a sliding window mechanism (BERT) and a bidirectional gated recurrent unit (BiGRU) can be used to contextualize knowledge text and generate vector representations of vocabulary. BERT generates a textual semantic representation vector for the knowledge text, while the BiGRU extracts all vocabulary in the knowledge text and calculates a vector representation for each vocabulary word based on the textual semantic representation vector.

[0045] In the third step, the text representation of each initial text tag is calculated based on all vector representations and knowledge text keywords.

[0046] Specifically, through the formula: ; Calculate the Text representation of the initial text labels .

[0047] in, , Indicates the number of initial text labels corresponding to the knowledge text, 、 represents the weight matrix, 、 represents the bias term, represents the transpose operation, Indicates the Topic-guided representation of initial text labels: ; ; ; in, represents the normalized coefficient guided by the subject word, Indicates the The vector representation of the subject words of the knowledge text, Indicates the The number of knowledge texts corresponding to the initial text labels, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, represents the attention parameter, represents a constant, 、 Indicates parameters, Represents the subject term of the knowledge text.

[0048] Step 14 : Based on the text representation of all the initial text labels of each knowledge text, a target text label of each knowledge text is selected from all the initial text labels of each knowledge text.

[0049] The matching degree between the above knowledge text and the target text label is greater than the matching degree between the knowledge text and each other initial text label.

[0050] In some embodiments of the present application, the step of selecting a target text label for each knowledge text from all initial text labels of each knowledge text based on the text representation of all initial text labels of each knowledge text includes: For each knowledge text, perform the following steps: In the first step, the matching degree between the knowledge text and each initial text label is calculated based on the text representation of each initial text label of the knowledge text.

[0051] Specifically, through the formula: ; Computational Knowledge Text and The matching degree between the initial text labels .

[0052] in, Indicates the The text representation of the initial text labels, Represents knowledge text, , Indicates the number of initial text labels corresponding to the knowledge text, represents a constant, represents the learnable weight vector, Represents a transpose operation.

[0053] In the second step, all matching degrees are sorted from large to small, and the initial text labels corresponding to the first multiple matching degrees are used as the target text labels of the knowledge text.

[0054] For example, in order to improve the accuracy of the target text label, it is necessary to train the parameters of BERT and BiGRU for generating the vector representation of the vocabulary in step 13, as well as the formula for calculating the topic words, the formula for calculating the text representation, and the formula for calculating the matching degree between the knowledge text and the initial text label. Specifically, sample knowledge texts that have a true match with the initial text labels are used as training data. BERT is first trained, and BERT is used to generate a text semantic representation vector for each sample knowledge text. The fully connected network is then used to classify the text semantic representation vector to obtain a predicted category label. The standard labels of all sample knowledge texts are obtained according to the process of step 12. A classification loss function is constructed based on all standard labels and all predicted category labels. With the purpose of minimizing the value of the classification loss function, the Adam optimizer and the back-propagation algorithm are used to adjust the parameters of BERT. Then, the parameters of BiGRU and the three formulas are trained, and the text semantic representation vector generated by the trained BERT is input into BiGRU to generate a vector representation of the vocabulary. According to steps 13 and 14, the target text label of each sample knowledge text is obtained, and a ranking loss function and a reconstruction loss function are constructed based on all target text labels. The final loss function is constructed based on the ranking loss function and the reconstruction loss function. With the purpose of minimizing the final loss function, the Adam optimizer and the back-propagation algorithm are used to train the parameters of BiGRU and the three formulas.

[0055] The above classification loss function is: ; in, represents the number of sample knowledge texts, represents the value of the classification loss function, represents the number of standard labels, represents the true label, Indicates that the sample knowledge text belongs to The probability of a standard label.

[0056] The above ranking loss function is: ; in, represents the value of the ranking loss function, represents the number of sample knowledge texts, The number of target text labels representing the sample knowledge text, Indicates the Sample knowledge texts, Indicates the Sample knowledge text and The true matching degree between the target text labels (which can be calculated by using the trained bag-of-words model, support vector machine and other models to calculate the sample knowledge text and the target text label), Indicates the Sample knowledge text and The matching degree between the target text labels.

[0057] The above reconstruction loss function is: ; in, represents the value of the reconstruction loss function, represents the latent variable, represents the process of inferring the model, represents the process of generating the model, represents the text reconstructed according to the latent variables, represents the prior distribution of the latent variable z, represents the KL divergence that measures the difference between two distributions, Representation based on distribution expectations, Feature vector representing the sample knowledge text.

[0058] The final loss function mentioned above is: ; in, represents the balancing hyperparameter.

[0059] Step 15: Introduce virtual users and construct a user knowledge scoring matrix based on the target text labels of all knowledge texts, all behavior records of target users, and virtual users.

[0060] The elements in the above user knowledge rating matrix are the target user's ratings for each initial text tag, and multiple elements correspond one-to-one to multiple ratings. The above virtual users are virtual users who have browsed the knowledge texts, and it is assumed that the virtual users have browsed each knowledge text.

[0061] In some embodiments of the present application, the step of constructing a user knowledge scoring matrix based on the target text labels of all knowledge texts, all behavior records of target users, and virtual users includes: The first step is to build a user interaction rating matrix based on all the behavior records of the target user and the virtual user.

[0062] The elements in the above user interaction score matrix are the interaction scores between the target user and each knowledge text, or the interaction scores between the virtual user and each knowledge text. The interaction score is used to describe the interest level between the target user or virtual user and the knowledge text. Multiple elements correspond to multiple interaction scores. For example, if there are 3 virtual users and 3 knowledge texts, then there are 3 elements in the user interaction score matrix. elements.

[0063] By formula: ; Calculating users With knowledge text Interaction score between .

[0064] in, Represents a user No. The number of days between the browsing behavior and the current time, , Represents a user The total number of browsing behaviors, The browsing behavior is the The behavior of opening the device to browse text, and interacting with at least one text in this browsing behavior, represents the time decay rate constant, Represents a user With knowledge text The interaction weight of: ; in, 、 、 Both represent weights, Represents a user Browse knowledge text duration, Represents a user Knowledge Text Performed a like action. Represents a user No knowledge text Perform a like action, Represents a user Knowledge Text Collected behavior, Represents a user No knowledge text Collecting behavior, , Indicates the total number of knowledge text numbers. In this formula, the superscript Indicates the Browsing behavior, that is, users Browse knowledge text Browsing behavior, such as user Browse knowledge text When the user opens the device to browse text for the second time, Through this superscript, the user's interactive behavior on the knowledge text can be expressed while expressing the relative time the user browses the knowledge text.

[0065] In the second step, a text label matrix is constructed based on the target text labels of all knowledge texts.

[0066] The elements in the above text label matrix are the matching scores between each knowledge text and each initial text label.

[0067] Specifically, through the formula: ; Computational Knowledge Text With initial text label Match score between .

[0068] in, , Represents the set of all initial text labels corresponding to all knowledge text labels.

[0069] The third step is to construct a user knowledge scoring matrix based on the user interaction scoring matrix and the text label matrix.

[0070] The elements in the user knowledge rating matrix are the target user's ratings for each initial text tag. This rating expresses the target user's interest in the initial text tag.

[0071] Specifically, through the formula: By formula: ; Calculate the target user's initial text label Rating ; in, The total number of numbers representing knowledge texts, Indicates target users Knowledge Text The score of the interaction, in which the user is considered For target users, Represents a user Knowledge Text The interaction score, Indicates the number of users (that is, the number of virtual users, is the sum of the number of virtual users and target users), , represents the user label matrix, , , Represents the transposed matrix of the text label matrix Corresponding knowledge text With initial text label The matching score between the elements, represents the user interaction rating matrix, Represents a text label matrix.

[0072] Step 16: Perform personalized knowledge recommendations for target users based on the user knowledge scoring matrix.

[0073] Specifically, according to the score between the target user and each initial text label in the user knowledge score matrix, the initial text labels corresponding to the highest scores are used as recommended labels, and knowledge texts with recommended labels are recommended to the target user.

[0074] It should be noted that for new users with no browsing history, a basic profile of the new user is constructed, and based on the clustering algorithm, the new user and existing users are clustered according to the basic profile, and knowledge recommendations are made to the new user based on other users who belong to the same cluster as the new user. For example, the tag with the highest score among other users is used as the recommended tag for the new user, and the knowledge text including the recommended tag is recommended to the new user.

[0075] For new users who need to browse disease treatment and health-related knowledge, the basic portrait includes the new user's age, gender, medical history, inherent ability characteristics, and monitoring indicators including heart rate and blood sugar.

[0076] Then process the data in the basic portrait and convert it into a data vector form that is convenient for subsequent analysis. The data can be processed using methods such as one-hot encoding. For data with high dimensions after encoding (such as disease history), principal component analysis (PCA) can be used to reduce the dimensionality of the encoded data. The core principle of the PCA method is to convert the original high-dimensional data into a set of new representations with linear independence in each dimension through linear transformation, thereby effectively reducing the data dimension while retaining the main features of the data. Assume that the user's original disease history data matrix is Udisease, and the dimension is (M is the number of samples, is the feature dimension, here is the total number of disease types), first calculate the covariance matrix of the data . Perform eigenvalue decomposition on the covariance matrix C and obtain the eigenvalue and the corresponding eigenvector . Sort by eigenvalue size, select the eigenvectors corresponding to the first d largest eigenvalues, and form the projection matrix Through projection transformation , project the original data U into the low-dimensional space to obtain the reduced-dimensional data , whose dimensions are .

[0077] The K-means clustering algorithm can be used for clustering, using the processed basic profiles and the user profiles of existing users as data points for the K-means clustering algorithm. The elbow method is used to optimize the optimal number of clusters, K. The core idea is to measure the closeness of samples within the same cluster and the degree of difference between samples in different clusters during the clustering process, thereby finding an appropriate K value that strikes a balance between clustering accuracy and model complexity. The following normalization formula is used to map the eigenvalues in the basic profile to the range [0, 1]: ,in is the first The first sample (i.e. new user or existing user) eigenvalues (i.e., the data in the basic portrait after numbering and dimension reduction), Indicates the first The first sample eigenvalues, and are the first The minimum and maximum values of the features. Secondly, the normalized data samples are divided into several clusters. In cluster analysis, the within-cluster sum of squares (WCSS) is used as an indicator to measure the tightness of the clustering results. Its calculation formula is: , where K is the number of clusters, Indicates the clusters, u is the number of clusters in the normalized data set. The sample points in It is a cluster As the value of K increases, the number of samples in each cluster decreases, and the sum of the squared distances between sample points and the cluster center (WCSS) gradually decreases because more clusters provide a better fit to the data. However, too many clusters increase model complexity and can easily lead to overfitting. Finally, we calculate the WCSS for different K values and plot a curve of WCSS versus K. We find the inflection point in the curve. The K value corresponding to this inflection point is the optimal balance between reducing WCSS and controlling model complexity, which serves as the input parameter for the subsequent K-means clustering algorithm.

[0078] After completing data dimensionality reduction and determining the optimal K value, the normalized basic profile and the user profile of existing users are used as data points to perform cluster analysis using the K-means clustering algorithm. Specifically, randomly select data points as initial cluster centers For each data point in the dataset , using Euclidean distance Calculate the distance between it and each cluster center, where is the data dimension, Represents a data point No. The value of the dimension, Represents the cluster center No. dimension) and Assigned to the cluster with the closest cluster center Recalculate the center of each cluster, that is, take the mean of all data points in the cluster as the new cluster center, and the calculation formula is ,in Represents a cluster The number of data points in the cluster is quantified. The steps of data point assignment and cluster center update are repeated until a stopping condition is met. This stopping condition can be when the cluster center no longer changes, when the change in cluster center is less than a preset threshold, or when a preset number of iterations is reached. Through continuous iterative optimization, a stable clustering result is eventually achieved.

[0079] It is worth mentioning that a user knowledge scoring matrix is constructed based on the behavioral records of user interactions and the target text tags of knowledge texts. It takes into account the text semantics of user interactions and knowledge texts, effectively analyzes the text tags that users are interested in, and makes personalized knowledge recommendations to users based on the user knowledge scoring matrix, thereby improving the accuracy of knowledge recommendations to users.

[0080] In addition, introducing virtual users to construct the user knowledge rating matrix can overcome the problem of sparse user data, avoid the cold start of knowledge recommendation, and improve the quality of recommendation.

[0081] The method of the present application also has the following advantages: By deeply mining the fine-grained features of text tags, we can more accurately identify users’ interests and preferences. Based on this, the recommendation system can provide users with knowledge content that meets their needs, improving the adaptability and coverage of the recommendation system for complex knowledge content.

[0082] The introduction of the BERT pre-training model and the adoption of a two-stage training method fully utilize the correlation between head labels and tail labels, effectively resolving the common problems of long length and complex content in Chinese knowledge texts, alleviating the long-tail and imbalanced data distribution problems in text label classification, and overcoming the shortcomings of low efficiency and information omissions in traditional methods when processing long texts.

[0083] It effectively integrates topic guidance and multi-layer attention mechanism, which can efficiently mine the potential topic information in knowledge text, enable the model to adaptively capture the topic features related to different tags in the text, establish label-aware context associations in the encoding stage, and improve the accuracy and reliability of label acquisition.

[0084] The introduction of virtual users and their appropriate weighting alleviates the knowledge cold-start problem faced by traditional recommendation systems, allowing new knowledge to be smoothly incorporated. Leveraging cluster analysis of user profile characteristics, health knowledge recommendations are recommended to new users based on behavioral data from similar groups, effectively addressing the cold-start issue for new users. Furthermore, the integration of historical user interaction records with a health knowledge tag matrix effectively addresses the issue of sparse user group data and improves the real-time and accuracy of recommendations.

[0085] The following is an exemplary description of the knowledge recommendation device that integrates text tags and user behaviors provided by this application.

[0086] like Figure 2 As shown, an embodiment of the present application provides a knowledge recommendation device that integrates text tags and user behaviors. The knowledge recommendation device 200 that integrates text tags and user behaviors includes: An acquisition module 201 is configured to acquire multiple knowledge texts that a target user interacts with an electronic device, and a behavior record of the target user interacting with each knowledge text with the electronic device; The preprocessing module 202 is used to preprocess each knowledge text to obtain multiple initial text labels for each knowledge text; the initial text labels are used to describe the semantics in the knowledge text; The subject mining module 203 is used to perform subject mining on each knowledge text to obtain the subject words of the knowledge text and generate a text representation of each initial text label of the knowledge text based on the subject words; a selection module 204 for selecting a target text label for each knowledge text from all the initial text labels of each knowledge text based on the text representation of all the initial text labels of each knowledge text; a matching degree between the knowledge text and the target text label is greater than a matching degree between the knowledge text and each other initial text label; Construction module 205 is used to introduce a virtual user and construct a user knowledge scoring matrix based on the target text labels of all knowledge texts, all behavior records of the target user, and the virtual user; the elements in the user knowledge scoring matrix are the scores of the target user for each initial text label; The personalized recommendation module 206 is used to make personalized knowledge recommendations for target users based on the user knowledge scoring matrix.

[0087] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0088] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0089] like Figure 3 As shown, an embodiment of the present application provides a terminal device. The terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above-mentioned method embodiments when executing the computer program D102.

[0090] Specifically, when the processor D100 executes the computer program D102, it obtains multiple knowledge texts that the target user interacts with the electronic device, as well as the behavior records of the target user when interacting with the electronic device for each knowledge text, and then preprocesses each knowledge text to obtain multiple initial text tags for each knowledge text. Then, for each knowledge text, the knowledge text is subject-mined to obtain the subject words of the knowledge text, and a text representation of each initial text tag of the knowledge text is generated based on the subject words. Then, based on the text representation of all initial text tags of each knowledge text, a target text tag of each knowledge text is selected from all initial text tags of each knowledge text. Then, a virtual user is introduced, and a user knowledge scoring matrix is constructed based on the target text tags of all knowledge texts, all behavior records of the target user, and the virtual user. Finally, personalized knowledge recommendation is performed for the target user based on the user knowledge scoring matrix. Among them, the user knowledge scoring matrix is constructed based on the behavior records of the user interaction and the target text tags of the knowledge text, taking into account the text semantics of the user interaction and the knowledge text, effectively analyzing the text tags of interest to the user, and performing personalized knowledge recommendation for the user based on the user knowledge scoring matrix, thereby improving the accuracy of knowledge recommendation for the user.

[0091] The processor D100 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0092] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device D10. Furthermore, the memory D101 may include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or is about to be output.

[0093] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0094] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0095] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the knowledge recommendation method apparatus / terminal device that integrates text tags and user behavior, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a mobile hard drive, a magnetic disk, or an optical disk.

[0096] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0097] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0098] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A knowledge recommendation method that integrates text labels and user behaviors, characterized in that: include: Acquire multiple knowledge texts that a target user interacts with an electronic device, and a behavior record of the target user interacting with each of the knowledge texts with the electronic device; Preprocessing each of the knowledge texts to obtain a plurality of initial text labels for each of the knowledge texts; the initial text labels are used to describe the semantics in the knowledge texts; For each of the knowledge texts, subject mining is performed on the knowledge text to obtain the subject words of the knowledge text, and a text representation of each initial text label of the knowledge text is generated based on the subject words; Based on the textual representation of all initial text labels of each knowledge text, a target text label of each knowledge text is selected from all initial text labels of each knowledge text; the matching degree between the knowledge text and the target text label is greater than the matching degree between the knowledge text and each other initial text label; Introducing a virtual user, constructing a user knowledge scoring matrix based on the target text labels of all knowledge texts, all behavior records of the target user, and the virtual user; the elements in the user knowledge scoring matrix are the scores of the target user for each of the initial text labels; Perform personalized knowledge recommendations for the target user based on the user knowledge scoring matrix.

2. The knowledge recommendation method according to claim 1, characterized in that: The subject mining of the knowledge text to obtain the subject words of the knowledge text includes: By formula: ; ; ; ; Keywords of Computational Knowledge Texts ; in, represents the latent variable, and represents the variational parameter, represents the parameter distribution, 、 、 represents a multi-layer perceptron computing network, The feature vector representing the knowledge text, represents the latent variable mean vector, represents the latent variable variance vector, Represents the activation function.

3. The knowledge recommendation method according to claim 1, characterized in that: Generating a text representation of each initial text tag of the knowledge text based on the subject word includes: Performing context sensing on the knowledge text to obtain a vector representation of each word in the knowledge text; A text representation of each initial text tag is calculated based on all vector representations and the knowledge text keywords.

4. The knowledge recommendation method according to claim 3, characterized in that: The calculating of the text representation of each initial text tag based on all vector representations and the knowledge text keywords includes: By formula: ; Calculate the Text representation of the initial text labels ; in, , Indicates the number of initial text labels corresponding to the knowledge text, 、 represents the weight matrix, 、 represents the bias term, represents the transpose operation, Indicates the Topic-guided representation of initial text labels: ; ; ; in, represents the normalization coefficient guided by the subject word, Indicates the The vector representation of the subject words of the knowledge text, Indicates the The number of knowledge texts corresponding to the initial text labels, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, Indicates the The initial text label corresponds to the The normalized attention weight of the subject words of the knowledge text, represents the attention parameter, represents a constant, 、 Indicates parameters, Represents the subject term of the knowledge text.

5. The knowledge recommendation method according to claim 1, characterized in that: The text representation based on all initial text labels of each knowledge text, selecting a target text label for each knowledge text from all initial text labels of each knowledge text, includes: For each of the knowledge texts, perform the following steps: Calculating a matching degree between the knowledge text and each of the initial text labels according to a text representation of each of the initial text labels of the knowledge text; All matching degrees are sorted from large to small, and the initial text labels corresponding to the first multiple matching degrees are used as the target text labels of the knowledge text.

6. The knowledge recommendation method according to claim 5, characterized in that: The calculating the matching degree between the knowledge text and each of the initial text labels according to the text representation of each of the initial text labels of the knowledge text comprises: By formula: ; Computing Knowledge Text and The matching degree between the initial text labels ; in, Indicates the The text representation of the initial text labels, represents the knowledge text, , Indicates the number of initial text labels corresponding to the knowledge text, represents a constant, represents the weight vector, Represents a transpose operation.

7. The knowledge recommendation method according to claim 1, characterized in that: The user knowledge scoring matrix is constructed based on the target text labels of all knowledge texts, all behavior records of the target users, and virtual users, including: A user interaction scoring matrix is constructed based on all behavioral records of the target user and the virtual user; the elements in the user interaction scoring matrix are interaction scores between the target user and each knowledge text, or interaction scores between the virtual user and each knowledge text; Constructing a text label matrix based on the target text labels of all knowledge texts; the elements in the text label matrix are scores for the matching between each knowledge text and each initial text label; A user knowledge scoring matrix is constructed based on the user interaction scoring matrix and the text label matrix; the elements in the user knowledge scoring matrix are the scores of the target user for each initial text label.

8. The knowledge recommendation method according to claim 7, characterized in that: The behavior record includes the user's browsing time, collection behavior and like behavior of the knowledge text; The step of constructing a user interaction scoring matrix based on all behavior records of the target user includes: By formula: ; Calculating users With knowledge text Interaction score between ; in, Represents a user No. The number of days between the browsing behavior and the current time, , Represents a user The total number of browsing behaviors, represents the time decay rate constant, Represents a user With knowledge text The interaction weight of: ; in, 、 、 Both represent weights, Represents a user Browse knowledge text duration, Represents a user Knowledge Text Performed a like action. Represents a user No knowledge text Perform a like action, Represents a user Knowledge Text Collecting behavior, Represents a user No knowledge text Collecting behavior, , The total number of numbers representing knowledge texts; The text label matrix is constructed according to the target text labels of all knowledge texts, including: By formula: ; Computational Knowledge Text With initial text label Match score between ; in, , Represents the set of all initial text labels corresponding to all knowledge text labels.

9. The knowledge recommendation method according to claim 8, characterized in that: The constructing of a user knowledge scoring matrix based on the user interaction scoring matrix and the text label matrix includes: By formula: ; Calculate the target user's initial text label Rating ; in, The total number of numbers representing knowledge texts, Indicates target users Knowledge Text The interaction score, Represents a user Knowledge Text The interaction score, Indicates the number of users, , represents the user label matrix, , , Represents the transposed matrix of the text label matrix Corresponding knowledge text With initial text label The matching score between the elements, represents the user interaction rating matrix, Represents a text label matrix.

10. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the knowledge recommendation method integrating text tags and user behaviors as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Association recommendation method and device based on label knowledge graph, and computer equipment

    CN112307337A

  • Label identification method and device, computer equipment, storage medium and program product

    CN113627447A

  • Tag recommendation method of online question and answer platform based on knowledge graph and tag association

    CN113672693A

  • Tag recommendation method fusing text similarity and collaborative filtering

    CN113722443A

  • Synchronization and tagging of image and text data

    US11093690B1