A large-scale language model-enhanced medical information recommendation method

By using large-scale language models in the medical information recommendation system to extract user interests, and combining heterogeneous hypergraphs and neural network technology to alleviate the sparseness of collaborative data, the problems of increasing model complexity and increasing data noise in the existing technology are solved, and more accurate user interest mining and recommendation effects are achieved.

CN119046449BActive Publication Date: 2025-05-02ZHEJIANG NARI DIGITAL HEALTH TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411050005.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-05-02
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

When the existing medical information recommendation method handles long user behavior sequences, the model complexity increases and data noise increases, and it cannot effectively alleviate the problem of collaborative data sparse caused by the sparseness of user interaction behavior.

Method used

A large-scale language model is used to extract important information from the medical information viewed by users, reduce the complexity of the model, and by constructing heterogeneous hypergraphs and designing heterogeneous hypergraph neural networks, the user's collaborative data and category data of medical information are used to alleviate the sparseness of collaborative information.

Benefits of technology

It realizes that without increasing the complexity of the model, fully tap user interests and effectively alleviates the problem of collaborative data sparse caused by the sparseness of user interaction behavior, and improves the accuracy of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046449B_ABST
    Figure CN119046449B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical information recommendation method enhanced by a large-scale language model, which recommends information of interest to users based on the medical information sequences that users have watched historically. The forward propagation part of the present invention mainly consists of three parts. The first part is to connect the titles of medical information watched by users to obtain long title text as the input of the large-scale language model, and design the prompt instructions of the large-scale language model so that the model outputs a short summary text; then design a title encoder to process the summary text to obtain the user's interest under a semantic view; the second part is to construct a heterogeneous hypergraph containing two relationships based on the user's collaborative data and the category data of medical information, and at the same time design a heterogeneous hypergraph neural network to obtain the vector representation of users and information under the collaborative view; the third part is to combine the user vectors under the semantic view and the collaborative view to calculate the probability of the user browsing the candidate information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet service technology, and in particular relates to a medical information recommendation method enhanced by a large-scale language model. Background Art

[0002] Medical information recommendation methods refer to mining user interests and recommending the next medical information they like. Existing recommendation methods usually use convolutional neural networks or attention mechanisms to model the medical information sequences that users have recently viewed, thereby obtaining user interests. However, when users view a large amount of medical information, using the information in a long list of medical information as input not only increases the complexity of the model, but also introduces more noise. Therefore, existing recommendation methods only use a small amount of medical information that users have recently viewed as model input. In addition, in order to alleviate the problem of sparse user behavior, existing recommendation methods introduce collaborative information to update user representations, such as constructing an item-user bipartite graph. Unpopular information and users with less behavior have fewer edges in this graph, so they still cannot be well modeled.

[0003] In view of the fact that the existing recommendation methods only use a small amount of medical information that the user has recently viewed as model input to reduce the complexity of the model, the present invention uses a large-scale language model to extract important information from the medical information that the user has viewed. Large-scale language models have been proven to generate summary information, that is, to extract important information from long texts and generate short texts. The present invention uses a large-scale language model to extract important information from a large amount of medical information that the user has viewed. Commonly used open source large-scale language models include LLaMA and ChatGLM. The application scenario of the present invention is the medical field. The general large-scale language model cannot understand some Chinese diseases and TCM professional terms well. Therefore, the large-scale language model used in the present invention is ChatGLM-Med in the medical field. ChatGLM-Med is a ChatGLM fine-tuning model based on Chinese medical knowledge. The present invention uses the user's interactive information as the input of the large-scale language model to fully explore the user's interests; and uses the characteristics of extracting important information from the large-scale language model to reduce the noise of the user's long interactive information.

[0004] In view of the sparsity of user interaction behavior, the present invention further introduces the category of medical information into the recommendation system. The expansion of the relationship from "user-information-user" to "user-category-user" has alleviated the sparsity of the data to a great extent. In the original "user-information-user" collaborative information, users are only similar to users who have clicked on the same information. After the introduction of the "user-category-user" collaborative information, users are similar to users who have clicked on the same category. At this time, the number of similar user sets is expanded, so the sparsity of collaborative data is alleviated. Summary of the invention

[0005] The technical difficulty that this method aims to solve is that when all medical information viewed by users is used as the input of the recommendation system in order to fully explore user interests, the model complexity increases and the data noise increases. At the same time, this method alleviates the problem of collaborative data sparsity in the recommendation method. When the user's interactive behavior is small, this method can still obtain a relatively accurate user vector representation.

[0006] The present invention proposes a medical information recommendation method enhanced by a large-scale language model. The titles of medical information viewed by users are connected to obtain a long title text as the input of the large-scale language model. The prompt instructions of the large-scale language model are designed so that the model outputs a short summary text. Then, a title encoder is designed to process the summary text to obtain the user's interest under the semantic view. User u i The medical information sequence is represented as Subscript Is a sequence The length, v j Is user u i The jth information browsed. Any medical information v j The title is represented as Will Connect the titles of medical information in the system to get the long text of the titles of the information that the user has browsed. The formula is: in It is a text connection symbol. Input into a large-scale language model and design prompt instructions for generating summaries s Make large-scale language models output short summary text LLM refers to Large Scale Language Model. Designing a title encoder to get the user’s interest in a semantic view

[0007] Based on the user's collaborative data and the category data of medical information, a heterogeneous hypergraph containing two relationships is constructed, and a heterogeneous hypergraph neural network is designed to obtain the vector representation of users and information under the collaborative view. The method of constructing a heterogeneous hypergraph is to allow the information nodes viewed by the same user to be connected by the same user hyperedge, and similarly, the information nodes belonging to the same category are connected by the same category hyperedge. j The vector representation is Encode via header encoder s Get, that is in It's information j Next, the heterogeneous hypergraph neural network is used to update the vectors of users, information, and categories in the hypergraph.j The initial vector is The heterogeneous hypergraph neural network consists of a multi-layer information transmission process, and the number of layers is L. The heterogeneous hypergraph neural network is used for the l-th layer node v i Vector and edge j Vector The updated calculation formula is as follows:

[0008]

[0009] Among them, B(e i ) is the edge e i The information node set contained in the set B(e i ) in any information node V τ The vector representation of at the l-1 layer is E(v j ) indicates that it contains information node v j The set of hyperedges, the set E(v j ) in any side e τ The vector representation at the l-1th layer is The formula is:

[0010]

[0011] Among them, |B(e i The operator |·| in )| represents the number of sets, and σ is the activation function. node→edge The formula is as follows:

[0012]

[0013] Among them, type(e τ )∈{c,u}, when the edge type is user, the matrix mapping It is W u , when the edge type is category, the mapping matrix It is W c .

[0014] Combining the user vectors in the semantic view and the collaborative view, the probability of the user browsing the candidate information is calculated. The title text of the candidate information vτ is Its vector representation is The user vector is fused with the function f merge get: The probability is calculated using the dot product formula of the vector and normalized using the softmax activation function:

[0015]

[0016] Among them, the superscript is the transpose symbol, represents probability. The loss function is:

[0017]

[0018] Among them, y τ Representative Informationv τ One-hot encoding. Functions are optimized using an optimizer.

[0019] The beneficial technical effects of the present invention are as follows:

[0020] (1) This method uses a large-scale medical language model to extract representations from a semantic perspective from a long sequence of user behaviors, making full use of user interaction data to mine user interests without increasing the complexity of the model.

[0021] (2) This method proposes to construct a heterogeneous hypergraph from the "user-information-category" collaborative data to alleviate the sparsity of collaborative information, and designs a heterogeneous hypergraph neural network to model the heterogeneity in the heterogeneous hypergraph structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flowchart of a large-scale language model-enhanced medical information recommendation method of the present invention;

[0023] Figure 2 This is a model framework diagram of a large-scale language model enhanced medical information recommendation method of the present invention. DETAILED DESCRIPTION

[0024] In order to further understand the present invention, a large-scale language model enhanced medical information recommendation method provided by the present invention is specifically described below in combination with a specific implementation method, but the present invention is not limited to this. Non-essential improvements and adjustments made by technical personnel in this field under the core guiding ideology of the present invention still fall within the scope of protection of the present invention.

[0025] The problem definition of this method is to recommend the next interesting information to the user based on the sequence of medical information that the user has watched in the past. The forward propagation part of a large-scale language model-enhanced medical information recommendation method mainly consists of three parts. The first part is to connect the titles of medical information viewed by the user to obtain the long title text as the input of the large-scale language model, and design the prompt instructions of the large-scale language model so that the model outputs a short summary text; then design a title encoder to process the summary text to obtain the user's interest under the semantic view; the second part is to construct a heterogeneous hypergraph containing two relationships based on the user's collaborative data and the category data of medical information, and at the same time design a heterogeneous hypergraph neural network to obtain the vector representation of users and information under the collaborative view; the third part is to combine the user vectors under the semantic view and the collaborative view to calculate the probability of the user browsing the candidate information.

[0026] like Figure 1 As shown, according to an embodiment of the present invention in a medical information recommendation scenario, the method includes the following steps:

[0027] S100, connect the titles of medical information viewed by the user to obtain the long title text as the input of the large-scale language model, design the prompt instruction of the large-scale language model so that the model outputs a short summary text; then design the title encoder to process the summary text to obtain the user's interest under the semantic view. i The medical information sequence is represented as Subscript Is a sequence The length, v j Is user u i The jth information browsed. Any medical information v j The title is represented as Will Connect the titles of medical information in the system to get the long text of the titles of the information that the user has browsed. The formula is: in It is a text connection symbol. The prompt instruction of large-scale language model s The design is "This is a medical information that the user has browsed ``` long title text```, generate a summary, limited to 50 words". The three backticks ``` here are commonly used delimiters in large-scale language models. Technicians can try to modify the prompt instruction repeatedly according to different scenarios. Because large-scale language models tend to generate longer answers, it is recommended that technicians limit the number of words in the generated summary in the prompt instruction. The short summary text output by the large-scale language model is represented as Right now The LLM here stands for Large-Scale Language Model. This application scenario is Chinese medical information, so the open source ChatGLM-Med large-scale language model in the medical field is used. Technicians can choose other large-scale language models according to their own application scenarios, such as the closed-source ChatGPT, open-source LLaMA and ChatGLM, etc.

[0028] Get the streamlined news headline and use the BERT model and a two-layer neural network as the title encoder Encode l Get the user's interest pui under the semantic view. The specific formula is:

[0029]

[0030] Among them, W1 and W2 are model matrix parameters, b1 is the model vector parameter, and b2 is the model scalar parameter. W1, W2, b1, and b2 are updated during model training. is the output of the BERT model, that is, the BERT model for text The encoding vector of the BERT model is frozen and not updated during training. The parameters of the BERT model are initialized using the existing checkpoint. The process of reducing the long text of the news title to short text by the large-scale language model can reduce the noise of the original long text. In addition, the refined short text is used as the input of the BERT model, which reduces the complexity of the model.

[0031] S200, based on the user's collaborative data and the category data of medical information, constructs a heterogeneous hypergraph containing two relationships, and designs a heterogeneous hypergraph neural network to obtain the vector representation of users and information under the collaborative view. The category data of medical information refers to the sub-classification of the information, such as: Internet medical care, telemedicine, Internet diagnosis and treatment management, smart medical cloud imaging platform, medical big data, and electronic medical records. A hypergraph is a type of graph structure in which one edge can contain any number of nodes. Figure 2 An example graph of a heterogeneous hypergraph constructed based on the collaborative data of users u1 and u2 is given. Information nodes viewed by the same user are connected by the same user hyperedge. Similarly, information nodes belonging to the same category are connected by the same category hyperedge. Because the hypergraph contains two types of hyperedges, one is the user hyperedge and the other is the category hyperedge, the hypergraph is a heterogeneous hypergraph. In this heterogeneous hypergraph, information belonging to the same category is connected, and information viewed by the same user is also connected. Compared with the traditional user-information bipartite graph, this heterogeneous hypergraph has more category hyperedges, so even if it is unpopular information, it can still establish a connection relationship with other information. The vectors of information nodes in the hypergraph still use the BERT model and a two-layer neural network as the title encoder Encode s To get any informationj Vector representation of The input to the title encoder is the title text The specific formula is:

[0032]

[0033] Among them, W4, W3, b3 and b4 are model parameters. is the output of the BERT model, that is, the BERT model for text Encode vector. Note that Encode s and Encode l Although the structures are the same, the parameters of the two layers of neural networks are not shared. Next, the heterogeneous hypergraph neural network is used to update the vectors of users, information, and categories in the hypergraph, where any information v j The initial vector is The heterogeneous hypergraph neural network is the same as the ordinary graph neural network, which consists of a multi-layer information transmission process, and the number of layers is L. In this embodiment, the number of layers L = 2. Technicians can set L in the range of 1 to 4 according to the scenario to find the optimal value. From the heterogeneous hypergraph structure, it can be seen that there are three types of connections between information nodes in the hypergraph: only belonging to the same category, only belonging to the same user, and belonging to the same category and the same user. The structures of LightGCN without weights and GCN with weights have been proven to be effective. The computational complexity of LightGCN is lower than that of GCN, but it cannot be applied to heterogeneous graphs. The heterogeneous hypergraph neural network proposed in the present invention performs a multi-layered computation on the node v of the lth layer. i Vector and edge j Vector The updated calculation formula is as follows:

[0034]

[0035] Among them, B(e i ) is the edge e i The information node set contained in the set B(e i ) in any information node v τ The vector representation of at the l-1 layer is E(v j ) indicates that it contains information node v j The set of hyperedges, the set E(v j ) in any side e τ The vector representation at the l-1th layer is AGGR node→edge The role of is to pass the information of the information node in the heterogeneous hypergraph to the edge and update the edge vector. The formula is:

[0036]

[0037] Among them, |B(e i The operator |·| in )| represents the number of sets, and σ is the sigmoid activation function. edge→node的 The function is to pass the edge information of the heterogeneous hypergraph to the information node and update the node vector. Because there are two types of edges in the heterogeneous hypergraph, representing collaborative information and category attribute information respectively, matrix mapping is used Where type(e τ )∈{c,u}. That is, when the edge type is user, the mapping matrix is ​​W u , when the edge type is category, the mapping matrix is ​​W c AGGR node→edge The formula is as follows:

[0038]

[0039] It can be seen that AGGR node→edge Without introducing weights, AGGR node→edge The weight is introduced. The advantage of this setting is that it can model the heterogeneity of the hypergraph and balance the complexity of the model. After L iterations, the information v under the collaborative view that integrates the user collaborative information j The vector is User i The vector is After the category data of medical information is introduced into the collaborative hypergraph, users with sparse behaviors and information that is only viewed by a small number of users can also be linked to other similar consultations in the graph. Therefore, the sparsity of interaction data is greatly alleviated.

[0040] S300, combining the user vectors in the semantic view and the collaborative view, and calculating the probability that the user browses the candidate information. τ The title text is Its vector representation is The user vector is represented as The probability is calculated using the dot product formula of the vector and normalized using the softmax activation function:

[0041]

[0042] Among them, the superscript is the transpose symbol, represents probability. The loss function is:

[0043]

[0044] Among them, y τ Representative Informationvτ One-hot encoding. The function is optimized using the Adam optimizer.

[0045] The above description of the embodiments is to facilitate the understanding and application of the present invention by those skilled in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made to the present invention by those skilled in the art based on the disclosure of the present invention should be within the scope of protection of the present invention.

Claims

1. A large-scale language model-enhanced medical information recommendation method, characterized by: Connect the titles of medical information viewed by users to obtain long title text as the input of large-scale language model, design prompt instructions of large-scale language model so that the model outputs short summary text; then design title encoder to process the summary text to obtain user’s interest under semantic view; user u i The medical information sequence is represented as Subscript Is a sequence The length, v j Is user u i The jth information browsed; any medical information v j The title is represented as Will Connect the titles of medical information in the system to get the long text of the titles of the information that the user has browsed. The formula is: in Is a text connection symbol; Input into a large-scale language model and design prompt instructions for generating summaries s Make large-scale language models output short summary text LLM stands for Large Scale Language Model; Design title encoder Encode l Get the user's interest in the semantic view Based on the user's collaborative data and the category data of medical information, a heterogeneous hypergraph containing two kinds of relationships is constructed, and a heterogeneous hypergraph neural network is designed to obtain the vector representation of users and information under the collaborative view. The method of constructing a heterogeneous hypergraph is to connect the information nodes viewed by the same user with the same user hyperedge, and similarly, the information nodes belonging to the same category are connected with the same category hyperedge. Any information v in the hypergraph j The vector representation is Encode via header encoder s Get, that is in It's information j Next, the heterogeneous hypergraph neural network is used to update the vectors of users, information, and categories in the hypergraph. j The initial vector is The heterogeneous hypergraph neural network consists of a multi-layer information transmission process, the number of layers is L; the heterogeneous hypergraph neural network for the l-th layer node v i Vector and edge j Vector The updated calculation formula is as follows: Among them, B(e i ) is the edge e i The information node set contained in the set B(e i ) in any information node v τ The vector representation of at the l-1 layer is E(v j ) indicates that it contains information node v j The set of hyperedges, the set E(v j ) in any side e τ The vector representation at the l-1th layer is The formula is: Among them, |B(e i )|, the operator |·| represents the number of sets, σ ​​is the activation function; AGGR node→edge The formula is as follows: Among them, type(e τ )∈{c,u}, when the edge type is user, the matrix mapping It is W u , when the edge type is category, the mapping matrix It is W c ; Combine the user vectors in the semantic view and the collaborative view to calculate the probability of the user browsing the candidate information; the candidate information v τ The title text is Its vector representation is The user vector is fused with the function f merge get: The probability is calculated using the dot product formula of the vector and normalized using the softmax activation function: The superscript T is the transposition symbol. represents probability; the loss function is: Among them, y τ Representative Informationv τ One-hot encoding; Functions are optimized using an optimizer.

2. The method for recommending medical information using a large-scale language model according to claim 1, characterized in that: The large-scale language model prompt instruction s Designed as "This is a medical information that the user has browsed ``` long text of the title```, generate a summary, limited to 50 words." 3. The method for recommending medical information using a large-scale language model according to claim 1, characterized in that: The large-scale language model is the open source ChatGLM-Med large-scale language model in the medical field.

4. The method for recommending medical information using a large-scale language model according to claim 1, wherein: The title encoder Encode l It consists of a BERT model and a two-layer neural network. The specific formula is: Among them, W1 and W2 are model matrix parameters, b1 is the model vector parameter, and b2 is the model scalar parameter; W1, W2, b1, and b2 are updated during model training; the parameters of the BERT model are frozen and not updated during training. The input is The output vector of the BERT model when .

5. The method for recommending medical information using a large-scale language model according to claim 1, characterized in that: The title encoder Encode s It consists of a BERT model and a two-layer neural network. The specific formula is: Among them, W4, W3, b3 and b4 are model parameters, The input is is the output vector of the BERT model.

6. The method for recommending medical information using a large-scale language model according to claim 1, characterized in that: The fusion function f merge The specific calculation formula is:

7. The method for recommending medical information using a large-scale language model according to claim 1, characterized in that: The optimizer is the Adam optimizer.

Citation Information

Patent Citations

  • Long text matching method based on graph convolution

    CN116304749A

  • Recommendation method, device and equipment

    CN117493703A