Intelligent anti-cancer science popularization content recommendation method and system based on patient big data

By building a user portrait and content recommendation engine, combining the patient's treatment stage and psychological state, the accurate and personalized recommendation of anti-cancer popular science content is achieved, solving the problems of insufficient accuracy of recommended content and neglecting psychological needs in the existing technology, and improving the scientificity and credibility of recommended content.

CN119988738APending Publication Date: 2025-05-13XIAMEN COBBLESTONE NETWORK TECH CO LTD

Patent Information

Application Number
CN202510114273.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the recommendation of anti-cancer popular science content in the prior art, the recommended content is insufficiently accurate, which ignores the psychological and emotional needs of the patients, and the content classification and labeling are insufficient, resulting in the recommended content that does not match the actual needs of the patients.

Method used

By collecting static data and behavioral data of patients, building user portraits, combining the treatment stage and psychological state of the patients, using natural language processing technology and machine learning algorithms to attach tags to popular science content, building a content recommendation engine to dynamically match user portraits and tags, generating a personalized content recommendation list, and introducing a dynamic feedback mechanism to adjust user portraits and recommendation strategies.

Benefits of technology

It has achieved accurate and personalized recommendations for popular science content against cancer, met the health knowledge needs of patients, provided psychological support, improved the scientificity and credibility of recommended content, and continuously optimized the recommendation effect through dynamic feedback mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988738A_ABST
    Figure CN119988738A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent anti-cancer science popularization content recommendation method and system based on patient big data, and the method comprises the following steps: collecting patient data, including static data and behavior data, the behavior data including search history, click record, browsing duration, activity participation record, and reading content; extracting user features from the patient data, and generating a user portrait; anti-cancer popular science contents in the platform are collected, and labels are added to the popular science anti-cancer contents by using a natural language processing technology and a machine learning algorithm; constructing a content recommendation engine for dynamically matching the user portrait with the tag; according to the matching result of the content recommendation engine, generating a personalized content recommendation list in combination with the current anti-cancer treatment stage of the patient; and transmitting the recommended content by using multiple channels according to the content recommendation list. According to the method, the user portrait is constructed through the multi-dimensional data of the patient, so that the pertinence and adaptability of recommendation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease course management and content recommendation, and in particular to an intelligent anti-cancer science popularization content recommendation method and system based on patient big data. Background Art

[0002] Getting news and information is a habit of people in modern society. With the development of computer technology and the continuous expansion of the scale of Internet users, more and more people use the Internet to obtain various information they need. In this context, it becomes extremely important to filter out the most valuable information from the massive amount of information and recommend the news and information that users are most interested in. With the rapid development of artificial intelligence (AI), big data and natural language processing (NLP) technology, intelligent recommendation systems have been widely used in e-commerce, social media, education and other fields to provide users with personalized content recommendations.

[0003] The Chinese invention patent application with publication number CN117972214A discloses a real-time movie recommendation processing method and device based on a knowledge graph, the method comprising the following steps: obtaining user operation behavior data, and constructing a user portrait graph based on the collected user operation behavior data; based on the constructed user portrait graph, associating the user portrait graph with media resource data to construct a joint knowledge graph for the user; using a knowledge graph embedding method to convert the entity and relationship parameters of the joint knowledge graph into vectors for subsequent entity and relationship retrieval as data preparation; based on the user operation behavior data, using a real-time recommendation algorithm to construct real-time recommendation pool data for each user; based on the construction of the user portrait graph, combining the real-time recommendation pool data for each user, and filtering the content that the user has seen and is not interested in, outputting a recommendation prediction result that updates the recommended content in real time according to the user operation.

[0004] However, in the field of medical health, especially in the recommendation of popular science content for cancer patients, existing technologies still have many shortcomings. Cancer is a disease that seriously threatens human health. Its treatment cycle is long and staged. Different patients have significantly different needs for anti-cancer knowledge at different stages of treatment. Scientific and effective anti-cancer popular science content can help patients understand the disease, choose appropriate treatment plans, and provide psychological support. However, most of the current content recommendation methods are based on traditional advertising recommendation models, and their application in the field of anti-cancer popular science has the following problems: The accuracy of recommended content is insufficient. Conventional advertising recommendation systems usually make recommendations based on user consumption habits, browsing history and other behavioral data. They lack in-depth analysis of medical and health data and patient-specific needs, resulting in a mismatch between recommended content and patients' actual needs. For example, the recommendation system may recommend content that is irrelevant to the patient's condition or of low quality, making it difficult to provide patients with scientific disease management advice.

[0005] Ignoring psychological and emotional needs: During cancer treatment, patients generally face tremendous psychological pressure and emotional fluctuations. Psychological support and emotional care are crucial to the patient's recovery. Traditional recommendation systems mainly focus on the user's consumption behavior, ignoring the user's emotional characteristics and psychological state, and it is difficult to provide recommendations with psychological support functions.

[0006] Content classification and labeling are not sophisticated enough. Anti-cancer science popularization content has strong professional and phased characteristics. Different types of content (such as cancer type, treatment stage, psychological support) need to be accurately classified and labeled according to the specific needs of patients. However, existing recommendation systems usually use coarse-grained classification in content processing, resulting in insufficient pertinence and practicality of content recommendations. Summary of the invention

[0007] The present invention provides an intelligent anti-cancer popular science content recommendation method and system based on patient big data, aiming to solve the problems of insufficient accuracy and low adaptability of recommended content caused by the single dimension of user portrait when the existing technology is applied in the field of anti-cancer popular science content recommendation.

[0008] To solve the above technical problems, the content recommendation method proposed in the present invention comprises the following steps: Collect patient data, including static data and behavioral data, including search history, click records, browsing time, activity records, and reading content; Extract user features from patient data and generate user profiles; Collect anti-cancer science content on the platform and use natural language processing technology and machine learning algorithms to add tags to the anti-cancer science content; Build a content recommendation engine to dynamically match user portraits and tags; Generate a personalized content recommendation list based on the matching results of the content recommendation engine and the patient's current anti-cancer treatment stage; Use multiple channels to deliver recommended content based on the content recommendation list.

[0009] Preferably, the method further includes dynamic feedback on the recommended content, specifically: Collect patient feedback data on recommended content, including click-through rate, reading time, and feedback score; Based on patients' feedback data, analyze the dynamic changes of interest points and adjust the user portrait and recommendation strategy for the corresponding patients.

[0010] Preferably, the method further performs desensitization processing on the patient data and extracts user features from the desensitized patient data, and the desensitization processing methods include masking, data generalization, data segmentation, data perturbation and data anonymization.

[0011] Preferably, the user feature extraction includes static data feature extraction and behavioral data feature extraction, the static data feature extraction is directly stored as user features according to the field name, and the behavioral data feature extraction includes search history extraction, reading record extraction and emotional feature extraction; Search history extraction first obtains the search keywords input by the patient and performs preprocessing, which includes word segmentation, removal of stop words, and part-of-speech tagging; extracts keyword topics using a latent Dirichlet distribution model; and calculates interest intensity based on the frequency of use or timestamp of the keywords; Reading record extraction first presets tags for platform articles, then counts the frequency of tags of articles read by patients, and weights them based on reading time and stay ratio, and uses weighted frequency to construct patient reading preferences; Emotional feature extraction extracts emotional states by analyzing patients' comments, search content, or question-and-answer interaction records. First, a sentiment analysis model is used to classify the corresponding text into positive, neutral, negative, or specific emotions and assign weights.

[0012] Preferably, the generation of the label includes content classification and label refinement; Content classification: The anti-cancer science popularization content on the platform is classified according to preset classification dimensions, including cancer type, treatment stage, and psychological support. The classification methods include manual classification and text clustering algorithm classification; Label refinement includes the following steps: Extract keywords and use natural language processing technology to extract the core keywords of each article; Semantic matching: semantically match the extracted core keywords with the standard tag vocabulary to generate fine-grained tags.

[0013] Preferably, the method for the content recommendation engine to generate a content recommendation list is: Receive the patient's user profile and label data; Use the candidate set generation algorithm to preliminarily filter out content that matches the user profile and generate a candidate set; Construct a multi-level weighted scoring algorithm to calculate the similarity, stage relevance, and feedback weight of user portraits and tag data, sort the candidate set content, comprehensively calculate the matching score, and output the candidate list based on the matching score; adjusting the content scores to optimize the candidate list based on the patient's mental state extracted from the patient's participation activity record; The optimized candidate list is further optimized based on group behavior data and hot content, and the final recommendation list is output.

[0014] Preferably, the scoring formula of the multi-level weighted scoring algorithm is:

[0015] In the formula, is the comprehensive score calculated by the multi-level weighted scoring algorithm. is a weight parameter that can be set manually, obtained through a multi-armed bandit model, or dynamically adjusted based on A / B testing.

[0016] Preferably, the method for obtaining weight parameters of the multi-armed bandit model is: Initialize weight parameters to equal values; After each recommendation, the weight parameters are updated according to the feedback. The update algorithm is:

[0017] In the formula, is the updated weight parameter, is the initial weight parameter for this update, is the set learning rate, For the corresponding Recommended feedback score, determined by experts based on current The generated recommendations are scored, or the weight parameters are collected. Get user feedback when , corresponding to Update algorithm for The update is repeated until the set stop condition is reached, and the latest weight parameter is used as the final parameter for calculating the comprehensive score.

[0018] Preferably, the similarity is calculated as follows: Use the pre-trained language model to convert user interest tags and tags into two semantic vectors, and calculate the cosine similarity of the two semantic vectors as the similarity involved in the comprehensive score calculation; The calculation method of the stage correlation is as follows: constructing a correlation matrix according to the co-occurrence probability of the treatment stages, obtaining correlation score values ​​from the key matrix based on the treatment stage data in the patient static data and the treatment stage label of the content; The calculation method of the feedback weight is:

[0019]

[0020]

[0021]

[0022] In the formula, is the preset weight, , The number of time periods divided, The time difference between when the recommended content is clicked and now. is the time decay coefficient.

[0023] Another aspect of the present invention is to provide an intelligent anti-cancer science content recommendation system based on patient big data. The system is used to implement the above-mentioned content recommendation method, including: Data collection module, used to collect patient data and desensitize the data, including static data and behavioral data; User portrait generation module, used to extract user features from patient data and generate user portraits; The content processing module is used to collect anti-cancer science popularization content on the platform and classify and label the content using natural language processing technology and machine learning algorithms; Content recommendation engine, which receives user portrait and tag data and generates recommendation lists through candidate set generation, multi-level weight scoring, psychological state optimization and group behavior optimization steps; Dynamic feedback module, which is used to collect patients' feedback data on recommended content, including click-through rate, reading time, and feedback score. Based on the feedback data, it analyzes the dynamic changes of interest points and adjusts the user portrait and recommendation strategy for the corresponding patients. The multi-channel push module is used to deliver personalized recommendation content to patients through multiple channels based on the recommendation list.

[0024] Compared with the prior art, the present invention has the following technical effects: 1. The content recommendation method proposed in the present invention constructs a user portrait through the patient's static data and behavioral data, and generates personalized recommendation content based on the patient's treatment stage and psychological state, which can accurately meet the patient's health knowledge needs, rather than simply recommending goods or services based on consumption behavior characteristics.

[0025] 2. The content recommendation method proposed in the present invention introduces a dynamic feedback mechanism to analyze patients' feedback data and interest changes on recommended content in real time, dynamically adjust user portraits and recommendation strategies, and continuously optimize recommendation effects.

[0026] 3. The content recommendation method proposed in the present invention strictly classifies and labels popular science content in a refined manner to ensure the scientificity and credibility of the recommended content. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the process of the recommended method of the present invention. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.

[0029] Embodiment 1 This embodiment is an intelligent anti-cancer science content recommendation method based on patient big data. Figure 1 As shown, the following steps are included: Collect patient data, including static data and behavioral data, including search history, click records, browsing time, activity records, and reading content; Extract user features from patient data and generate user profiles; Collect anti-cancer science content on the platform and use natural language processing technology and machine learning algorithms to add tags to the anti-cancer science content; Build a content recommendation engine to dynamically match user portraits and tags; Generate a personalized content recommendation list based on the matching results of the content recommendation engine and the patient's current anti-cancer treatment stage; Use multiple channels to deliver recommended content based on the content recommendation list.

[0030] The method also includes dynamic feedback on the recommended content, specifically: Collect patient feedback data on recommended content, including click-through rate, reading time, and feedback score; Based on patients' feedback data, analyze the dynamic changes of interest points and adjust the user portrait and recommendation strategy for the corresponding patients.

[0031] The method also performs desensitization processing on the patient data and extracts user features from the desensitized patient data. The desensitization processing method includes mask processing, data generalization, data segmentation, data perturbation and data anonymization.

[0032] The user feature extraction includes static data feature extraction and behavioral data feature extraction. The static data feature extraction is directly stored as user features according to the field name. The behavioral data feature extraction includes search history extraction, reading record extraction and emotional feature extraction. Search history extraction first obtains the search keywords input by the patient and performs preprocessing, which includes word segmentation, removal of stop words, and part-of-speech tagging; extracts keyword topics using a latent Dirichlet distribution model; and calculates interest intensity based on the frequency of use or timestamp of the keywords; Reading record extraction first presets tags for platform articles, then counts the frequency of tags of articles read by patients, and weights them based on reading time and stay ratio, and uses weighted frequency to construct patient reading preferences; Emotional feature extraction extracts emotional states by analyzing patients' comments, search content, or question-and-answer interaction records. First, a sentiment analysis model is used to classify the corresponding text into positive, neutral, negative, or specific emotions and assign weights.

[0033] The generation of the labels includes content classification and label refinement; Content classification: The anti-cancer science popularization content on the platform is classified according to preset classification dimensions, including cancer type, treatment stage, and psychological support. The classification methods include manual classification and text clustering algorithm classification; Label refinement includes the following steps: Extract keywords and use natural language processing technology to extract the core keywords of each article; Semantic matching: semantically match the extracted core keywords with the standard tag vocabulary to generate fine-grained tags.

[0034] The method for the content recommendation engine to generate a content recommendation list is as follows: Receive the patient's user profile and label data; Use the candidate set generation algorithm to preliminarily filter out content that matches the user profile and generate a candidate set; Construct a multi-level weighted scoring algorithm to calculate the similarity, stage relevance, and feedback weight of user portraits and tag data, sort the candidate set content, comprehensively calculate the matching score, and output the candidate list based on the matching score; adjusting the content scores to optimize the candidate list based on the patient's mental state extracted from the patient's participation activity record; The optimized candidate list is further optimized based on group behavior data and hot content, and the final recommendation list is output.

[0035] The scoring formula of the multi-level weighted scoring algorithm is:

[0036] In the formula, is the comprehensive score calculated by the multi-level weighted scoring algorithm. is a weight parameter that can be set manually, obtained through a multi-armed bandit model, or dynamically adjusted based on A / B testing.

[0037] The method for obtaining weight parameters of the multi-armed bandit model is: Initialize weight parameters to equal values; After each recommendation, the weight parameters are updated according to the feedback. The update algorithm is:

[0038] In the formula, is the updated weight parameter, is the initial weight parameter for this update, is the set learning rate, For the corresponding Recommended feedback score, determined by experts based on current The generated recommendations are scored, or the weight parameters are collected. Get user feedback when , corresponding to Update algorithm for The update is repeated until the set stop condition is reached, and the latest weight parameter is used as the final parameter for calculating the comprehensive score.

[0039] The similarity is calculated as follows: Use the pre-trained language model to convert user interest tags and tags into two semantic vectors, and calculate the cosine similarity of the two semantic vectors as the similarity involved in the comprehensive score calculation; The calculation method of the stage correlation is as follows: constructing a correlation matrix according to the co-occurrence probability of the treatment stages, obtaining correlation score values ​​from the key matrix based on the treatment stage data in the patient static data and the treatment stage label of the content; The calculation method of the feedback weight is:

[0040]

[0041]

[0042]

[0043] In the formula, is the preset weight, , The number of time periods divided, The time difference between when the recommended content is clicked and now. is the time decay coefficient.

[0044] Embodiment 2 This embodiment is an intelligent anti-cancer science content recommendation system based on patient big data. The system is used to implement the content recommendation method described in the first embodiment, including: Data collection module, used to collect patient data and desensitize the data, including static data and behavioral data; User portrait generation module, used to extract user features from patient data and generate user portraits; The content processing module is used to collect anti-cancer science popularization content on the platform and classify and label the content using natural language processing technology and machine learning algorithms; Content recommendation engine, which receives user portrait and tag data and generates recommendation lists through candidate set generation, multi-level weight scoring, psychological state optimization and group behavior optimization steps; Dynamic feedback module, which is used to collect patients' feedback data on recommended content, including click-through rate, reading time, and feedback score. Based on the feedback data, it analyzes the dynamic changes of interest points and adjusts the user portrait and recommendation strategy for the corresponding patients. The multi-channel push module is used to deliver personalized recommendation content to patients through multiple channels based on the recommendation list.

[0045] The above is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, which all fall within the protection scope of the present invention.

Claims

1. An intelligent anti-cancer science content recommendation method based on patient big data, characterized in that: The following steps are involved: Collect patient data, including static data and behavioral data, including search history, click records, browsing time, activity records, and reading content; Extract user features from patient data and generate user profiles; Collect anti-cancer science content on the platform and use natural language processing technology and machine learning algorithms to add tags to the anti-cancer science content; Build a content recommendation engine to dynamically match user portraits and tags; Generate a personalized content recommendation list based on the matching results of the content recommendation engine and the patient's current anti-cancer treatment stage; Use multiple channels to deliver recommended content based on the content recommendation list.

2. The intelligent anti-cancer science content recommendation method based on patient big data according to claim 1 is characterized in that: The method also includes dynamic feedback on the recommended content, specifically: Collect patient feedback data on recommended content, including click-through rate, reading time, and feedback rating; Based on patients' feedback data, analyze the dynamic changes of interest points and adjust the user portrait and recommendation strategy for the corresponding patients.

3. The intelligent anti-cancer science content recommendation method based on patient big data according to claim 1 is characterized in that: The method also performs desensitization processing on the patient data and extracts user features from the desensitized patient data. The desensitization processing method includes mask processing, data generalization, data segmentation, data perturbation and data anonymization.

4. The intelligent anti-cancer science popularization content recommendation method based on patient big data according to claim 1 is characterized in that: The user feature extraction includes static data feature extraction and behavioral data feature extraction. The static data feature extraction is directly stored as user features according to the field name. The behavioral data feature extraction includes search history extraction, reading record extraction and emotional feature extraction. Search history extraction first obtains the search keywords entered by the patient and performs preprocessing, which includes word segmentation, stop word removal, and part-of-speech tagging; Extract keyword topics using latent Dirichlet distribution model; calculate interest intensity based on keyword usage frequency or timestamp; Reading record extraction first presets tags for platform articles, then counts the frequency of tags of articles read by patients, and weights them based on reading time and stay ratio, and uses weighted frequency to construct patient reading preferences; Emotional feature extraction extracts emotional states by analyzing patients' comments, search content, or question-and-answer interaction records. First, a sentiment analysis model is used to classify the corresponding text into positive, neutral, negative, or specific emotions and assign weights.

5. The intelligent anti-cancer science popularization content recommendation method based on patient big data according to claim 1 is characterized in that: The generation of the labels includes content classification and label refinement; Content classification: The anti-cancer science popularization content on the platform is classified according to preset classification dimensions, including cancer type, treatment stage, and psychological support. The classification methods include manual classification and text clustering algorithm classification; Label refinement includes the following steps: Extract keywords and use natural language processing technology to extract the core keywords of each article; Semantic matching: semantically match the extracted core keywords with the standard tag vocabulary to generate fine-grained tags.

6. The intelligent anti-cancer science popularization content recommendation method based on patient big data according to claim 1 is characterized in that: The method for the content recommendation engine to generate a content recommendation list is as follows: Receive the patient's user profile and label data; Use the candidate set generation algorithm to preliminarily filter out content that matches the user profile and generate a candidate set; Construct a multi-level weighted scoring algorithm to calculate the similarity, stage relevance, and feedback weight of user portraits and tag data, sort the candidate set content, comprehensively calculate the matching score, and output the candidate list based on the matching score; adjusting the content scores to optimize the candidate list based on the patient's mental state extracted from the patient's participation activity record; The optimized candidate list is further optimized based on group behavior data and hot content, and the final recommendation list is output.

7. The intelligent anti-cancer science popularization content recommendation method based on patient big data according to claim 6 is characterized in that: The scoring formula of the multi-level weighted scoring algorithm is: In the formula, is the comprehensive score calculated by the multi-level weighted scoring algorithm. is a weight parameter that can be set manually, obtained through a multi-armed bandit model, or dynamically adjusted based on A / B testing.

8. The intelligent anti-cancer science content recommendation method based on patient big data according to claim 7 is characterized in that: The method for obtaining weight parameters of the multi-armed bandit model is: Initialize weight parameters to equal values; After each recommendation, the weight parameters are updated according to the feedback. The update algorithm is: In the formula, is the updated weight parameter, is the initial weight parameter for this update, is the set learning rate, For the corresponding Recommended feedback score, determined by experts based on current The generated recommendations are scored, or the weight parameters are collected. Get user feedback when , corresponding to Update algorithm for The update is repeated until the set stop condition is reached, and the latest weight parameter is used as the final parameter for calculating the comprehensive score.

9. The intelligent anti-cancer science content recommendation method based on patient big data according to claim 7 is characterized in that: The similarity is calculated as follows: Use the pre-trained language model to convert user interest tags and tags into two semantic vectors, and calculate the cosine similarity of the two semantic vectors as the similarity involved in the comprehensive score calculation; The calculation method of the stage correlation is as follows: constructing a correlation matrix according to the co-occurrence probability of the treatment stages, obtaining correlation score values ​​from the key matrix based on the treatment stage data in the patient static data and the treatment stage label of the content; The calculation method of the feedback weight is: In the formula, is the preset weight, , The number of time periods divided, The time difference between when the recommended content is clicked and now. is the time decay coefficient.

10. An intelligent anti-cancer science content recommendation system based on patient big data, characterized in that: The system is used to implement the method according to any one of claims 1 to 9, comprising: Data collection module, used to collect patient data and desensitize the data, including static data and behavioral data; User portrait generation module, used to extract user features from patient data and generate user portraits; The content processing module is used to collect anti-cancer science popularization content on the platform and classify and label the content using natural language processing technology and machine learning algorithms; Content recommendation engine, which receives user portrait and tag data, and generates recommendation lists through candidate set generation, multi-level weight scoring, psychological state optimization, and group behavior optimization steps; Dynamic feedback module, which is used to collect patients' feedback data on recommended content, including click-through rate, reading time, and feedback score. Based on the feedback data, it analyzes the dynamic changes of interest points and adjusts the user portrait and recommendation strategy for the corresponding patients. The multi-channel push module is used to deliver personalized recommendation content to patients through multiple channels based on the recommendation list.

Citation Information

Patent Citations

  • Real-time film recommendation processing method and device based on knowledge graph

    CN117972214A

Cited By

  • Same patient friend matching method and system based on pathological characteristics and treatment cycle

    CN121565502A

  • Matching method and system for same-disease patients based on pathological characteristics and treatment cycles

    CN121565502B