Intelligent analysis method and system for medical literature data based on natural language processing

By building an intelligent medical literature data analysis system based on natural language processing, using BP neural network and Jaccard coefficients for analysis, the inefficiency and error-prone problems in medical literature management are solved, and efficient, real-time and intelligent analysis results are achieved.

CN119474365BActive Publication Date: 2025-08-15BEIJING YAOYUN DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510066377.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-08-15
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing medical literature management has problems of inefficiency and error-proneness. The existing technology analyzes the combination of subject words by obtaining the method of combining topic words.

Method used

Using a natural language processing method, we collect and process medical literature data, build a sample set and build a literature topic analysis model, use BP neural network for training, and combine Jaccard coefficients for real-time analysis and matching to improve the accuracy and intelligence of the analysis.

Benefits of technology

It realizes efficient, real-time and intelligent analysis of medical literature data, improves the accuracy and reliability of the analysis, and reduces manual errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474365B_ABST
    Figure CN119474365B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data analysis technology, and discloses a method and system for intelligent analysis of medical literature data based on natural language processing. The method collects historical medical literature data, processes the collected historical medical literature data, and constructs a sample set based on the processed historical medical literature data. Furthermore, a literature topic analysis model is constructed based on the processed historical medical literature data in the sample set, and the topics of the processed historical medical literature data are determined. Furthermore, based on the determined topics of the processed historical medical literature data, the topics of the processed historical medical literature data are analyzed and trained using an intelligent analysis method. After the training is completed, medical literature data is collected in real time, and analysis and matching is performed based on the topics of the trained historical medical literature data using natural language processing, thereby improving the accuracy of the medical literature data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and specifically to a method and system for intelligent analysis of medical literature data based on natural language processing. Background Art

[0002] With the continuous increase of medical literature data, the difficulty of manual inspection and testing of medical literature is also increasing. Medical literature management has pain points such as low efficiency and prone to errors. There is an urgent need for an intelligent medical literature analysis system to further improve the recognition rate of medical critical values and reduce manual errors.

[0003] The existing patent application CN118916459A discloses a method that acquires subject terms from documents by collecting and traversing combinations, and records a word combination consisting of subject terms as the first target combination. All first target combinations in the document are then marked. After marking, a second target combination consisting of subject terms different from the first target combination is acquired and traversed. The distribution of all subject terms that constitute the target combination is statistically analyzed within each paragraph of the document, resulting in a distribution vector for the subject terms within each paragraph. This vector is then represented by a reference vector group to obtain a comparison vector for the target combination. Several documents containing the target combination are then obtained and compared with reference documents. The comparison vectors of the reference documents are then compared with the comparison vectors of the several documents to determine the differences in subject terms across the different documents. However, since the document data is analyzed solely by acquiring and combining subject terms, the analysis lacks intelligence and is inaccurate. Summary of the Invention

[0004] (1) Technical problems solved

[0005] In response to the shortcomings of the existing technology, the present invention provides a method and system for intelligent analysis of medical literature data based on natural language processing, which has the advantages of high efficiency, real-time, intelligence, and accuracy, and solves the pain points of low efficiency and proneness to errors in medical literature management.

[0006] (2) Technical solution

[0007] In order to solve the above-mentioned technical problems of low efficiency and easy errors in medical literature management, the present invention provides the following technical solutions:

[0008] The present invention discloses a method for intelligent analysis of medical literature data based on natural language processing, which specifically comprises the following steps:

[0009] S1. Collecting historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data;

[0010] S2. Based on the processed historical medical literature data in the sample set, a literature theme analysis model is constructed, and the themes of the processed historical medical literature data are determined based on the constructed literature theme analysis model;

[0011] S3. Analyze and train the processed historical medical literature data themes using an intelligent analysis method to obtain a trained historical medical literature data theme set;

[0012] S4. Collect medical literature data in real time, determine the theme of the medical literature data collected in real time based on the constructed literature theme analysis model, and at the same time, analyze and match the trained historical medical literature data theme set through natural language processing to obtain the analyzed historical medical literature data theme.

[0013] The present invention collects historical medical literature data, processes the collected historical medical literature data, and constructs a sample set based on the processed historical medical literature data. At the same time, based on the processed historical medical literature data in the sample set, a document theme analysis model is constructed, and the theme of the processed historical medical literature data is determined. At the same time, based on the determined theme of the processed historical medical literature data, the theme of the processed historical medical literature data is analyzed and trained through an intelligent analysis method. After the training is completed, the medical literature data is collected in real time and analyzed and matched based on the theme of the trained historical medical literature data through natural language processing, thereby improving the accuracy of the medical literature data analysis.

[0014] Preferably, the collecting of historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data comprises the following steps:

[0015] S11, collecting historical medical literature data, filtering the collected historical medical literature data, and obtaining processed historical medical literature data;

[0016] S12. constructing a sample set based on the processed historical medical literature data;

[0017] Building a sample set ;

[0018] in, represents the sample set, represents the first set of processed historical medical literature data in the sample set, The nth group of processed historical medical literature data in the sample set.

[0019] Preferably, filtering the collected historical medical literature data to obtain the processed historical medical literature data comprises the following steps:

[0020] Create a length of Array, select A hash function is used to calculate the name and time of each collected historical medical literature data, and the calculation results are stored in an array;

[0021] in, represents the number of hash functions, Indicates the amount of historical medical literature data collected;

[0022] In the process of calculating the collected historical medical literature data using the hash function, when the calculation results of two sets of historical medical literature data are consistent, the names and times in the two sets of historical medical literature data are compared;

[0023] When the comparison results are consistent, the two sets of historical medical literature data are set as duplicate historical medical literature data;

[0024] When two sets of historical medical literature data are duplicate historical medical literature data, one set of historical medical literature data is deleted to achieve filtering processing;

[0025] After the calculation is completed, the historical medical literature data stored in the array during the calculation process is summarized to obtain the processed historical medical literature data.

[0026] The present invention ensures the reliability of the historical medical literature data by filtering the collected historical medical literature data and constructing a sample set based on the processed historical medical literature data.

[0027] Preferably, the process of constructing a document theme analysis model based on the processed historical medical document data in the sample set, and determining the theme of the processed historical medical document data based on the constructed document theme analysis model comprises the following steps:

[0028] S21. Randomly select a set of processed historical medical literature data from the sample set, extract all the word pairs, and construct a word pair set ;

[0029] S22. For vocabulary pair sets For each word pair in , calculate the probability that the word pair becomes the subject of the processed historical medical literature data;

[0030] ;

[0031] in, Represented in the vocabulary pair set Chinese vocabulary pairs Become the first The probability of the topic of historical medical literature data after group processing, Indicates in Word pairs in historical medical literature data after group processing Number of occurrences;

[0032] S23. Based on the calculated probability of each word pair becoming a topic, a document topic analysis model is constructed to determine the topic of the processed historical medical document data.

[0033] Preferably, the process of constructing a document topic analysis model based on the calculated probability of each word pair becoming a topic and determining the topic of the processed historical medical document data comprises the following steps:

[0034] The formula for calculating word pair weights is as follows:

[0035] ;

[0036] in, Representing word pairs The weight of represents the probability parameter, Represents positional parameters, Representing word pairs Location in historical medical literature data, including title, subheadings, text, and abstract;

[0037] Combine the word pair probability calculation formula and the word pair weight calculation formula, and set the combined word pair probability calculation formula and word pair weight calculation formula as the literature topic analysis model;

[0038] The weight threshold of the word pairs in the literature topic analysis model is set, the word pairs with weights greater than or equal to the weight threshold are summarized, and the topics of the processed historical medical literature data are determined based on the summarized word pairs.

[0039] The present invention selects a set of processed historical medical literature data, extracts all word pairs, calculates the probability of the word pairs becoming topics, constructs a document topic analysis model based on the calculated probability, determines the topics of the processed historical medical literature data, and improves the accuracy of the medical literature data.

[0040] Preferably, the analyzing and training of the processed historical medical literature data themes by the intelligent analysis method to obtain a trained historical medical literature data theme set comprises the following steps:

[0041] S31, selecting some of the processed historical medical literature data topics from the determined processed historical medical literature data topics to construct a training set;

[0042] S32, inputting the themes of the historical medical literature data processed in the training set into the BP neural network, and determining the BP neural network model through iterative training;

[0043] S33. Based on the determined BP neural network model, all processed historical medical literature data topics are trained, and a trained historical medical literature data topic set is output.

[0044] Preferably, the step of inputting the processed historical medical literature data into the BP neural network and determining the BP neural network model through iterative training comprises the following steps:

[0045] The themes of the historical medical literature data processed in the training set are input into the BP neural network. The structure of the BP neural network includes: input layer, hidden layer and output layer;

[0046] Set the input set of the input layer to ,in Indicates the input The total number of historical medical literature data themes processed in the training set is ;

[0047] The output set of the output layer is ,in Indicates the output The total number of historical medical literature data topics after training is output as ;

[0048] Assume that the hidden layer contains q neurons, v is the weight from the input layer to the hidden layer, and w is the weight from the hidden layer to the output layer;

[0049] Training is performed through the forward propagation process of the BP neural network;

[0050] During the forward propagation process, the calculation formula for transferring data from the input layer to the hidden layer is:

[0051] ;

[0052] in, represents the input of the hth hidden layer neuron, represents the weight of the i-th input to the h-th hidden layer neuron, and m represents the number of input data in the input layer; represents the bias value of the hth hidden layer neuron;

[0053] The calculation formula for data transmission from the hidden layer to the output layer is:

[0054] ;

[0055] in, represents the input of the j-th output layer neuron, represents the weight from the hth hidden layer neuron to the jth output layer neuron, and q represents the number of hidden layer neurons; represents the output of the hth hidden layer neuron; represents the bias value of the j-th output layer neuron;

[0056] After receiving the parameter data transmitted by the hidden layer, the output layer directly outputs the parameter data received from the hidden layer;

[0057] Right now ;

[0058] Calculate the error between the output layer and the expected value;

[0059] Set the error threshold between the output layer and the expected value. When the error between the output layer and the expected value is greater than or equal to the set error threshold, iterative calculation is performed by adjusting the weights from the hidden layer to the output layer and the weights from the input layer to the hidden layer in sequence until the BP neural network model is obtained.

[0060] The iterative calculation formula is as follows:

[0061] ;

[0062] in, Represents the weight adjustment value, represents the learning rate, Indicates error, Represents the output of the output layer.

[0063] The present invention improves the intelligence of medical literature data analysis by selecting the topics of some historical medical literature data from the topics of the processed historical medical literature data to construct a training set, and training the topics of all the processed historical medical literature data through the forward propagation process of the BP neural network.

[0064] Preferably, the real-time collection of medical literature data, determining the subject of the real-time collected medical literature data based on the constructed literature subject analysis model, and simultaneously analyzing and matching the obtained trained historical medical literature data subject set through natural language processing to obtain the analyzed historical medical literature data subject include the following steps:

[0065] Collect medical literature data in real time, determine the theme of the collected medical literature data based on the constructed literature theme analysis model, and calculate the similarity between the theme of the collected medical literature data and the theme set of the trained historical medical literature data through the Jaccard coefficient;

[0066] The calculation formula is:

[0067] ;

[0068] Among them, KT is the real-time collected medical literature data theme set, KN is the trained historical medical literature data theme set, Represents the similarity between set KT and set KN;

[0069] A Jaccard similarity threshold is set, and analysis and matching are performed through similarity matching. When the similarity between the real-time collected medical literature data topic and the trained historical medical literature data topic is greater than or equal to the threshold, it is determined that the real-time collected medical literature data topic is related to the trained historical medical literature data topic. If the similarity is lower than the threshold, it is determined that the real-time collected medical literature data topic is not related to the trained historical medical literature data topic.

[0070] The present invention determines the topics of real-time collected medical literature data based on a constructed literature topic analysis model, and calculates the similarity between the topics of real-time collected medical literature data and the topics of trained historical medical literature data through the Jaccard coefficient, thereby improving the real-time performance of medical literature data analysis.

[0071] The present invention also discloses a medical literature data intelligent analysis system based on natural language processing, which is used to implement a medical literature data intelligent analysis method based on natural language processing. The system includes: a data acquisition module, a data processing module, an analysis and training module, and a natural language processing module;

[0072] The data acquisition module is used to receive medical literature data in real time;

[0073] The data processing module is used to process the received medical literature data and transmit the processed medical literature data to the analysis and training module;

[0074] The analysis and training module is used to analyze the received processed medical literature data;

[0075] The natural language processing module is used to judge real-time medical literature data by calculating relevance.

[0076] (3) Beneficial effects

[0077] Compared with the existing technology, the present invention provides a method and system for intelligent analysis of medical literature data based on natural language processing, which has the following beneficial effects:

[0078] 1. The invention collects historical medical literature data, processes the collected historical medical literature data, and constructs a sample set based on the processed historical medical literature data. At the same time, based on the processed historical medical literature data in the sample set, a document theme analysis model is constructed, and the theme of the processed historical medical literature data is determined. At the same time, based on the determined theme of the processed historical medical literature data, the theme of the processed historical medical literature data is analyzed and trained through an intelligent analysis method. After the training is completed, the medical literature data is collected in real time and analyzed and matched based on the theme of the trained historical medical literature data through natural language processing, thereby improving the accuracy of the medical literature data analysis.

[0079] 2. The invention ensures the reliability of historical medical literature data by filtering the collected historical medical literature data and constructing a sample set based on the processed historical medical literature data.

[0080] 3. The invention selects a set of processed historical medical literature data, extracts all word pairs, and calculates the probability of the word pairs becoming topics. Based on the calculated probability, a document topic analysis model is constructed to determine the topics of the processed historical medical literature data, thereby improving the efficiency of medical literature data.

[0081] 4. The invention improves the intelligence of medical literature data analysis by selecting the topics of some historical medical literature data from the topics of the processed historical medical literature data to construct a training set, and training the topics of all processed historical medical literature data through the forward propagation process of the BP neural network.

[0082] 5. This invention determines the topics of medical literature data collected in real time based on a constructed literature topic analysis model, and calculates the similarity between the topics of medical literature data collected in real time and the topics of historical medical literature data after training through the Jaccard coefficient, thereby improving the real-time performance of medical literature data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 This is a schematic diagram of the structure of the medical literature data intelligent analysis process of the present invention. DETAILED DESCRIPTION

[0084] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0085] Example 1

[0086] See also Figure 1 This embodiment discloses a method for intelligent analysis of medical literature data based on natural language processing, which specifically includes the following steps:

[0087] S1. Collecting historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data;

[0088] S2. Based on the processed historical medical literature data in the sample set, a literature theme analysis model is constructed, and the themes of the processed historical medical literature data are determined based on the constructed literature theme analysis model;

[0089] S3. Analyze and train the processed historical medical literature data themes using an intelligent analysis method to obtain a trained historical medical literature data theme set;

[0090] S4. Collect medical literature data in real time, determine the theme of the medical literature data collected in real time based on the constructed literature theme analysis model, and analyze and match the trained historical medical literature data theme set through natural language processing to obtain the analyzed theme of the historical medical literature data;

[0091] Further, see Figure 1 , collecting historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data includes the following steps:

[0092] S11, collecting historical medical literature data, filtering the collected historical medical literature data, and obtaining processed historical medical literature data;

[0093] Create a length of Array, select A hash function is used to calculate the name and time of each collected historical medical literature data, and the calculation results are stored in an array;

[0094] in, represents the number of hash functions, Indicates the amount of historical medical literature data collected;

[0095] In the process of calculating the collected historical medical literature data using the hash function, when the calculation results of two sets of historical medical literature data are consistent, the names and times in the two sets of historical medical literature data are compared;

[0096] When the comparison results are consistent, the two sets of historical medical literature data are set as duplicate historical medical literature data;

[0097] When two sets of historical medical literature data are duplicate historical medical literature data, one set of historical medical literature data is deleted to achieve filtering processing;

[0098] After the calculation is completed, the historical medical literature data stored in the array during the calculation process is summarized to obtain the processed historical medical literature data;

[0099] S12. constructing a sample set based on the processed historical medical literature data;

[0100] Building a sample set ;

[0101] in, represents the sample set, represents the first set of processed historical medical literature data in the sample set, The nth group of processed historical medical literature data in the sample set;

[0102] Furthermore, the historical medical literature data after centralized sample processing includes diverse data such as literature content, literature citation information, and time information;

[0103] Further, see Figure 1 , based on the processed historical medical literature data in the sample set, constructing a literature topic analysis model, and determining the theme of the processed historical medical literature data based on the constructed literature topic analysis model includes the following steps:

[0104] S21. Randomly select a set of processed historical medical literature data from the sample set, extract all the word pairs, and construct a word pair set ;

[0105] S22. For vocabulary pair sets For each word pair in , calculate the probability that the word pair becomes the subject of the processed historical medical literature data;

[0106] The formula for calculating word pair probability is as follows:

[0107] ;

[0108] in, Represented in the vocabulary pair set Chinese vocabulary pairs Become the first The probability of the topic of historical medical literature data after group processing, Indicates in Word pairs in historical medical literature data after group processing Number of occurrences;

[0109] S23, based on the calculated probability of each word pair becoming a topic, constructing a literature topic analysis model to determine the topic of the processed historical medical literature data;

[0110] The formula for calculating word pair weights is as follows:

[0111] ;

[0112] in, Representing word pairs The weight of represents the probability parameter, Represents positional parameters, Representing word pairs The location in historical medical literature data, including title, subtitle, text, abstract, etc.;

[0113] Furthermore, the word pair probability calculation formula and the word pair weight calculation formula are combined, and the combined word pair probability calculation formula and word pair weight calculation formula are set as the literature topic analysis model;

[0114] Setting a weight threshold for word pairs in the literature topic analysis model, aggregating word pairs with weights greater than or equal to the weight threshold, and determining the themes of the processed historical medical literature data based on the aggregated word pairs;

[0115] Further, see Figure 1 , analyzing and training the processed historical medical literature data topics through intelligent analysis methods to obtain the trained historical medical literature data topic set includes the following steps:

[0116] S31, selecting some of the processed historical medical literature data topics from the determined processed historical medical literature data topics to construct a training set;

[0117] S32, inputting the themes of the historical medical literature data processed in the training set into the BP neural network, and determining the BP neural network model through iterative training;

[0118] The themes of the historical medical literature data processed in the training set are input into the BP neural network. The structure of the BP neural network includes: input layer, hidden layer and output layer;

[0119] Set the input set of the input layer to ,in Indicates the input The total number of historical medical literature data themes processed in the training set is ;

[0120] The output set of the output layer is ,in Indicates the output The total number of historical medical literature data topics after training is output as ;

[0121] Assume that the hidden layer contains q neurons, v is the weight from the input layer to the hidden layer, and w is the weight from the hidden layer to the output layer;

[0122] Training is performed through the forward propagation process of the BP neural network;

[0123] During the forward propagation process, the calculation formula for transferring data from the input layer to the hidden layer is:

[0124] ;

[0125] in, represents the input of the hth hidden layer neuron, represents the weight of the i-th input to the h-th hidden layer neuron, and m represents the number of input data in the input layer; represents the bias value of the hth hidden layer neuron;

[0126] The calculation formula for data transmission from the hidden layer to the output layer is:

[0127] ;

[0128] in, represents the input of the j-th output layer neuron, represents the weight from the hth hidden layer neuron to the jth output layer neuron, and q represents the number of hidden layer neurons; represents the output of the hth hidden layer neuron; represents the bias value of the j-th output layer neuron;

[0129] After receiving the parameter data transmitted by the hidden layer, the output layer directly outputs the parameter data received from the hidden layer;

[0130] Right now ;

[0131] Calculate the error between the output layer and the expected value;

[0132] Set the error threshold between the output layer and the expected value. When the error between the output layer and the expected value is greater than or equal to the set error threshold, iterative calculation is performed by adjusting the weights from the hidden layer to the output layer and the weights from the input layer to the hidden layer in sequence until the BP neural network model is obtained.

[0133] The iterative calculation formula is as follows:

[0134] ;

[0135] in, Represents the weight adjustment value, represents the learning rate, Indicates error, Represents the output of the output layer;

[0136] S33, based on the determined BP neural network model, training all processed historical medical literature data themes, and outputting a trained historical medical literature data theme set;

[0137] Further, see Figure 1 , collect medical literature data in real time, determine the theme of the medical literature data collected in real time based on the constructed literature theme analysis model, and analyze and match the theme set of historical medical literature data after training through natural language processing. The analyzed theme of the historical medical literature data includes the following steps:

[0138] Collect medical literature data in real time, determine the theme of the collected medical literature data based on the constructed literature theme analysis model, and calculate the similarity between the theme of the collected medical literature data and the theme set of the trained historical medical literature data through the Jaccard coefficient;

[0139] The calculation formula is:

[0140] ;

[0141] Among them, KT is the real-time collected medical literature data theme set, KN is the trained historical medical literature data theme set, Represents the similarity between set KT and set KN;

[0142] A Jaccard similarity threshold is set, and analysis and matching are performed using a similarity matching method. When the similarity between the real-time collected medical literature data topic and the trained historical medical literature data topic is greater than or equal to the threshold, it is determined that the real-time collected medical literature data topic is related to the trained historical medical literature data topic. If the similarity is lower than the threshold, it is determined that the real-time collected medical literature data topic is not related to the trained historical medical literature data topic.

[0143] Furthermore, relevant medical literature data were combined and submitted to manual review;

[0144] Example 2

[0145] See also Figure 1 , this embodiment also discloses a medical literature data intelligent analysis system based on natural language processing, which is used to implement a medical literature data intelligent analysis method based on natural language processing. The system includes: a data acquisition module, a data processing module, an analysis and training module, and a natural language processing module;

[0146] The data acquisition module is used to receive medical literature data in real time;

[0147] The data processing module is used to process the received medical literature data and transmit the processed medical literature data to the analysis and training module;

[0148] The analysis and training module is used to analyze the received processed medical literature data;

[0149] The natural language processing module is used to judge real-time medical literature data by calculating relevance.

[0150] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligent analysis of medical literature data based on natural language processing, characterized in that: The following steps are involved: S1. Collecting historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data; S2. Based on the processed historical medical literature data in the sample set, a literature theme analysis model is constructed, and the themes of the processed historical medical literature data are determined based on the constructed literature theme analysis model; S3, analyzing and training the processed historical medical literature data themes through an intelligent analysis method to obtain a trained historical medical literature data theme set; including S31, selecting some processed historical medical literature data themes from the determined processed historical medical literature data themes to construct a training set; S32, inputting the processed historical medical literature data themes in the training set into a BP neural network, and determining a BP neural network model through an iterative training method; S33, training all processed historical medical literature data themes based on the determined BP neural network model, and outputting a trained historical medical literature data theme set; S4. Collect medical literature data in real time, determine the theme of the medical literature data collected in real time based on the constructed literature theme analysis model, and analyze and match the trained historical medical literature data theme set through natural language processing to obtain the analyzed theme of the historical medical literature data; The method includes real-time collection of medical literature data, determining the theme of the real-time collected medical literature data based on the constructed literature theme analysis model, and calculating the similarity between the theme of the real-time collected medical literature data and the theme of the trained historical medical literature data through the Jaccard coefficient; A Jaccard similarity threshold is set, and analysis and matching are performed through similarity matching. When the similarity between the real-time collected medical literature data topic and the trained historical medical literature data topic is greater than or equal to the threshold, it is determined that the real-time collected medical literature data topic is related to the trained historical medical literature data topic. If the similarity is lower than the threshold, it is determined that the real-time collected medical literature data topic is not related to the trained historical medical literature data topic.

2. The method for intelligent analysis of medical literature data based on natural language processing according to claim 1, characterized in that: Collecting historical medical literature data, processing the collected historical medical literature data to obtain processed historical medical literature data, and constructing a sample set based on the obtained processed historical medical literature data includes the following steps: S11, collecting historical medical literature data, filtering the collected historical medical literature data, and obtaining processed historical medical literature data; S12. constructing a sample set based on the processed historical medical literature data; Construct sample set N={N1,N2,...,N n }; Among them, N represents the sample set, N1 represents the first set of processed historical medical literature data in the sample set, and N n The nth group of processed historical medical literature data in the sample set.

3. The method for intelligent analysis of medical literature data based on natural language processing according to claim 2, characterized in that: Filtering the collected historical medical literature data to obtain the processed historical medical literature data includes the following steps: Create an array of length l and select A hash function is used to collect each historical medical Calculate the name and time of the drug literature data and save the calculation results in an array; Where k represents the number of hash functions, and z represents the amount of historical medical literature data collected; In the process of calculating the collected historical medical literature data using the hash function, when the calculation results of two sets of historical medical literature data are consistent, the names and times in the two sets of historical medical literature data are compared; When the comparison results are consistent, the two sets of historical medical literature data are set as duplicate historical medical literature data; When two sets of historical medical literature data are duplicate historical medical literature data, one set of historical medical literature data is deleted to achieve filtering processing; After the calculation is completed, the historical medical literature data stored in the array during the calculation process is summarized to obtain the processed historical medical literature data.

4. The method for intelligent analysis of medical literature data based on natural language processing according to claim 3, characterized in that: Based on the processed historical medical literature data in the sample set, a literature topic analysis model is constructed, and the topics of the processed historical medical literature data are determined based on the constructed literature topic analysis model, including the following steps: S21. Randomly select a set of processed historical medical literature data from the sample set, extract all word pairs, and construct a word pair set B; S22. For each word pair in word pair set B, calculate the probability that the word pair becomes a topic of the processed historical medical literature data; Where P(b|g) represents the probability that word pair b in word pair set B becomes the topic of the g-th group of processed historical medical literature data, and g(b) represents the number of times word pair b appears in the g-th group of processed historical medical literature data. S23. Based on the calculated probability of each word pair becoming a topic, a document topic analysis model is constructed to determine the topic of the processed historical medical document data.

5. The method for intelligent analysis of medical literature data based on natural language processing according to claim 4, characterized in that: Based on the calculated probability of each word pair becoming a topic, a literature topic analysis model is constructed to determine the topic of the processed historical medical literature data. The following steps are included: The formula for calculating word pair weights is as follows: W b =a×P(b|g)+c×local b ; Among them, W b Represents the weight of vocabulary pair b, a represents the probability parameter, c represents the position parameter, local b Indicates the position of word pair b in historical medical literature data, including title, subtitle, text and abstract; Combine the word pair probability calculation formula and the word pair weight calculation formula, and set the combined word pair probability calculation formula and word pair weight calculation formula as the literature topic analysis model; The weight threshold of the word pairs in the literature topic analysis model is set, the word pairs with weights greater than or equal to the weight threshold are summarized, and the topics of the processed historical medical literature data are determined based on the summarized word pairs.

6. The method for intelligent analysis of medical literature data based on natural language processing according to claim 5, characterized in that: The themes of the historical medical literature data processed in the training set are input into the BP neural network, and the BP neural network model is determined by iterative training, which includes the following steps: The themes of the historical medical literature data processed in the training set are input into the BP neural network. The structure of the BP neural network includes an input layer, a hidden layer and an output layer. Set the input set of the input layer to X = [x1, x2, ..., x i ,...,x t ], where x i represents the historical medical literature data theme after processing in the input training set i, and the total number of historical medical literature data theme groups after processing in the input training set is t; The output set of the output layer is Y = [y1,y2,...,y j ,...,y s ], where y j represents the outputted j-th group of trained historical medical literature data topics, and the total number of trained historical medical literature data topics outputted is s; Assume that the hidden layer contains q neurons, and w is the weight from the hidden layer to the output layer; Training is performed through the forward propagation process of the BP neural network; During the forward propagation process, the calculation formula for transferring data from the input layer to the hidden layer is: Among them, α h represents the input of the hth hidden layer neuron, v ih represents the weight of the i-th input to the h-th hidden layer neuron, m represents the number of input data in the input layer; d h represents the bias value of the hth hidden layer neuron; The calculation formula for data transmission from the hidden layer to the output layer is: Among them, β j represents the input of the j-th output layer neuron, w hj represents the weight from the hth hidden layer neuron to the jth output layer neuron, q represents the number of hidden layer neurons; a h represents the output of the hth hidden layer neuron; d j represents the bias value of the j-th output layer neuron; After receiving the parameter data transmitted by the hidden layer, the output layer directly outputs the parameter data received from the hidden layer; Right now Calculate the error between the output layer and the expected value; Set the error threshold between the output layer and the expected value. When the error between the output layer and the expected value is greater than or equal to the set error threshold, iterative calculation is performed by adjusting the weights from the hidden layer to the output layer and the weights from the input layer to the hidden layer in sequence until the BP neural network model is obtained. The iterative calculation formula is as follows: Δw=(r)Ey; Among them, Δw represents the weight adjustment value, r represents the learning rate, E represents the error, and y represents the output of the output layer.

7. The method for intelligent analysis of medical literature data based on natural language processing according to claim 6, characterized in that: The similarity between the real-time collected medical literature data topics and the trained historical medical literature data topics is calculated by the Jaccard coefficient; The calculation formula is: Among them, KT is the real-time collected medical literature data theme set, KN is the trained historical medical literature data theme set, and J(KN, KT) represents the similarity between set KT and set KN.

8. A system for implementing the method for intelligent analysis of medical literature data based on natural language processing according to claim 7, characterized in that: include: Data acquisition module, data processing module, analysis and training module, and natural language processing module; The data acquisition module is used to receive medical literature data in real time; The data processing module is used to process the received medical literature data and transmit the processed medical literature data to the analysis and training module; The analysis and training module is used to analyze the received processed medical literature data; The natural language processing module is used to judge real-time medical literature data by calculating relevance.

Citation Information

Patent Citations

  • Duplicated data deletion method for medical big data

    CN114722013A

  • Theme-based medical literature retrieval method and system, storage medium and terminal

    CN115658851A