Biomedical data classification method and device, equipment and storage medium

Optimizing biomedical text data classification through the R-BERT model and triple loss strategy solves the problem of insufficient classification accuracy and robustness in the prior art, and achieves more efficient biomedical data classification.

CN120561306AInactive Publication Date: 2025-08-29DALIAN UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510696526.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the accuracy and robustness of biomedical data classification are low, especially when processing complex biomedical text data, it is difficult to effectively distinguish different entity relationships.

Method used

The R-BERT model is used to preprocess the biomedical text data, construct triple data pairs, and convert the function into hidden state sequences to determine the entity pair features and text features, and optimize the classification results using triple loss strategies and decision rules.

Benefits of technology

It improves the accuracy and robustness of biomedical data classification, especially in drug-drug interaction and protein-protein interaction tasks, and improves the generalization performance of classification models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561306A_ABST
    Figure CN120561306A_ABST
Patent Text Reader

Abstract

The invention discloses a biomedical data classification method and device, equipment and a storage medium. The method comprises the following steps: preprocessing biomedical text data to obtain a plurality of text data units; determining a triple data pair according to each text data unit; the triple data pairs are input into an R-BERT model, the R-BERT model converts the triple data pairs into a hidden state sequence, and the hidden state sequence comprises a plurality of entity pairs; determining entity pair features corresponding to each entity pair and text features corresponding to each text data unit; determining a vector of each text data unit according to the plurality of entity pair features and the text features; carrying out the conversion of the vector through a # imgabs0 # function, and obtaining a classification probability; and performing decision selection on the classification probability through a # imgabs1 # function to obtain a classification result label of the biomedical text data. According to the invention, the accuracy and robustness of biomedical data classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data classification technology, and in particular to a biomedical data classification method, device, equipment and storage medium. Background Art

[0002] Biomedical literature is rich in valuable information, making it imperative to extract semantic relationships between biomedical entities from unstructured text. Biomedical relationship extraction (BioRE) aims to extract semantic relationships between biomedical entities, helping researchers quickly determine whether entities interact. Current research on biomedical entity interactions primarily focuses on protein-protein interactions (PPIs) and drug-drug interactions (DDIs). In recent years, deep learning has been widely used to address the problem of biomedical relationship extraction. CNN-based models achieve high accuracy in DDI extraction. Integrating attention mechanisms into CNN-based relationship extraction has been proposed. Recurrent neural network (RNN)-based models also perform very well in relationship extraction tasks, with pre-trained models, in particular, reaching a state-of-the-art. However, biomedical text data is complex, characterized by numerous specialized terms, complex sentence structures, and information redundancy. Furthermore, different entity relationship pairs may share common semantic information, making them difficult to distinguish. Existing methods for classifying biomedical data suffer from low accuracy and robustness. Summary of the Invention

[0003] Based on this, it is necessary to propose a biomedical data classification method, device, computer equipment and storage medium to address the above problems.

[0004] A biomedical data classification method, comprising: Acquiring biomedical text data to be classified, and preprocessing the biomedical text data to obtain a plurality of text data units; determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; Determining an entity pair feature corresponding to each entity pair and a text feature corresponding to each of the text data units; Determine a vector for each text data unit according to a plurality of entity pair features and the text features; pass The function transforms the vector to obtain the classification probability; pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0005] In one embodiment, preprocessing the biomedical text data to obtain a plurality of text data units includes: A CLS tag is added before each biomedical text data to obtain labeled text data, and each labeled text data is segmented by a tokenization function to obtain multiple text data units.

[0006] In one embodiment, determining a triple data pair according to each of the text data units comprises: When there are positive text instances and negative text instances of the same origin in the text data unit: Acquire a first text data unit from the plurality of text data units, and determine a first negative text instance corresponding to the first text data unit, wherein the first text data unit is a text data unit marked as a positive text instance; The first text data unit, the first negative text instance, and the second text data unit constitute the triple data pair, wherein the second text data unit is a text data unit randomly selected from a plurality of the text data units and marked as a positive text instance; and Acquire a third text data unit from the plurality of text data units, and determine a first positive text instance corresponding to the third text data unit, wherein the third text data unit is a text data unit marked as a negative text instance; The third text data unit, the first positive text instance and the fourth text data unit constitute the triple data pair; the fourth text data unit is a text data unit randomly selected from the plurality of text data units and marked as a negative text instance.

[0007] In one embodiment, determining a triple data pair according to each of the text data units comprises: When there are no positive text instances and negative text instances of the same origin in the text data unit: Obtain a fifth text data unit from the plurality of text data units, randomly match a first text data having the same text instance label as the fifth text data unit in the Faiss library based on the BGE model, and randomly match a second text data having an opposite text instance label as the fifth text data unit in the Faiss library based on the BGE model; the text instance label is a positive text instance or a negative text instance; The fifth text data unit, the first text data, and the second text data constitute the triple data pair.

[0008] In one embodiment, determining the entity pair feature corresponding to each entity pair and the text feature corresponding to each text data unit is achieved by the following expression: in, is the text feature, is the first weight matrix, is the first bias function, is the hidden state sequence, ; First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the second weight matrix, First Entity At the beginning sequence of the hidden state layer, For the second entity At the beginning sequence of the hidden state layer, is the second bias function, ={ ...... } , ={ ...... } .

[0009] In one embodiment, determining the vector of each text data unit according to the plurality of entity pair features and the text features is implemented by the following expression: Among them, e is a vector, is the text feature, First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the third weight matrix, is the third bias function.

[0010] In one embodiment, the The function transforms the vector to obtain the classification probability; The function makes a decision on the classification probability to obtain the classification result label of the biomedical text data through the following expression: in, is the classification probability, e is a vector, is the classification result label.

[0011] A biomedical data classification device, comprising: A preprocessing module, configured to obtain biomedical text data to be classified and preprocess the biomedical text data to obtain a plurality of text data units; a first determining module, configured to determine a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; A second determination module is configured to input the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; A third determining module is used to determine an entity pair feature corresponding to each entity pair and a text feature corresponding to each text data unit; a fourth determining module, configured to determine a vector of each of the text data units based on the plurality of entity pair features and the text features; Conversion module, used to The function transforms the vector to obtain the classification probability; Decision selection module, used to The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0012] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Acquiring biomedical text data to be classified, and preprocessing the biomedical text data to obtain a plurality of text data units; determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; Determining an entity pair feature corresponding to each entity pair and a text feature corresponding to each of the text data units; Determine a vector for each text data unit according to a plurality of entity pair features and the text features; pass The function transforms the vector to obtain the classification probability; pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0013] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the following steps: Acquiring biomedical text data to be classified, and preprocessing the biomedical text data to obtain a plurality of text data units; determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; Determining an entity pair feature corresponding to each entity pair and a text feature corresponding to each of the text data units; Determine a vector for each text data unit according to a plurality of entity pair features and the text features; pass The function transforms the vector to obtain the classification probability; pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0014] The present application determines a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; inputs the triple data pair into the R-BERT model, and the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; determines an entity pair feature corresponding to each entity pair and a text feature corresponding to each of the text data units; determines a vector for each of the text data units according to the plurality of entity pair features and the text features; and The function transforms the vector to obtain the classification probability; The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data, thereby improving the accuracy and robustness of biomedical data classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] in: Figure 1 A diagram illustrating an application environment of a biomedical data classification method according to an embodiment; Figure 2 A flowchart of biomedical data classification in one embodiment; Figure 3 is a structural block diagram of a biomedical data classification device in one embodiment; Figure 4 A schematic diagram of a triple loss strategy in one embodiment; Figure 5 This is a schematic diagram of the characteristic distribution of an example in an embodiment; Figure 6 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] Biomedical literature is rich in valuable information, making it imperative to extract semantic relationships between biomedical entities from unstructured text. Biomedical relationship extraction (BioRE) aims to extract semantic relationships between biomedical entities, helping researchers quickly determine whether entities interact. Current research on biomedical entity interactions primarily includes protein-protein interactions (PPIs) and drug-drug interactions (DDIs). In recent years, deep learning has been widely used by researchers to address the problem of biomedical relationship extraction. CNN-based models can achieve high accuracy in DDI extraction tasks. There are proposals to integrate attention mechanisms into CNN-based relationship extraction. Recurrent neural network (RNN)-based models also perform very well in relationship extraction tasks, with pre-trained models, in particular, reaching a relatively advanced level. However, biomedical text data is complex, characterized by numerous specialized terms, complex sentence structures, and information redundancy. Furthermore, different entity relationship pairs may share consistent semantic information, making them difficult to distinguish. Existing methods for classifying biomedical data suffer from low accuracy and robustness. To address these technical issues, this application provides a method for biomedical data classification.

[0019] Figure 1 FIG. 1 is an application environment diagram of a biomedical data classification method in an embodiment. Figure 1 , the biomedical data classification method is applied to a biomedical data classification system. The biomedical data classification system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster composed of multiple servers. The terminal 110 is used to obtain the biomedical text data to be classified, and the server 120 is used to preprocess the biomedical text data to obtain multiple text data units; determine a triple data pair according to each of the text data units, and the number of the triple data pairs is the same as the number of the text data units; input the triple data pair into the R-BERT model, and the R-BERT model converts the triple data pair into a hidden state sequence, and the hidden sequence contains multiple entity pairs; determine the entity pair features corresponding to each entity pair and the text features corresponding to each of the text data units; determine the vector of each of the text data units according to the multiple entity pair features and the text features; through The function transforms the vector to obtain the classification probability; The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0020] A biomedical data classification method, such as Figure 2 As shown, the method includes: S10: Acquire biomedical text data to be classified, and preprocess the biomedical text data to obtain a plurality of text data units; S20: determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; S30: Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; S40: Determine entity pair features corresponding to each entity pair and text features corresponding to each text data unit; S50: determining a vector of each text data unit according to a plurality of entity pair features and the text features; S60: Pass The function transforms the vector to obtain the classification probability; S70: Pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0021] In one embodiment, the preprocessing of the biomedical text data to obtain a plurality of text data units in step S10 includes: A CLS tag is added before each biomedical text data to obtain labeled text data, and each labeled text data is segmented by a tokenization function to obtain multiple text data units.

[0022] In one embodiment, determining a triple data pair according to each text data unit in step S20 includes: When there are positive text instances and negative text instances of the same origin in the text data unit: S201: Acquire a first text data unit from the plurality of text data units, and determine a first negative text instance corresponding to the first text data unit, wherein the first text data unit is a text data unit marked as a positive text instance; S202: The first text data unit, the first negative text instance, and the second text data unit constitute the triple data pair, wherein the second text data unit is a text data unit randomly selected from the plurality of text data units and marked as a positive text instance; and S203: Acquire a third text data unit from the plurality of text data units, and determine a first positive text instance corresponding to the third text data unit, wherein the third text data unit is a text data unit marked as a negative text instance; S204: The third text data unit, the first positive text instance and the fourth text data unit constitute the triple data pair; the fourth text data unit is a text data unit randomly selected from the plurality of text data units and marked as a negative text instance.

[0023] In another embodiment, determining a triple data pair according to each text data unit in step S20 includes: like Figure 4 As shown, when there are no positive text instances and negative text instances of the same source in the text data unit: S205: Obtain a fifth text data unit from the plurality of text data units, randomly match a first text data having the same text instance label as the fifth text data unit in the Faiss library based on the BGE model, and randomly match a second text data having an opposite text instance label as the fifth text data unit in the Faiss library based on the BGE model; the text instance label is a positive text instance or a negative text instance; S206: The fifth text data unit, the first text data, and the second text data constitute the triple data pair.

[0024] In one embodiment, the determination of the entity pair features corresponding to each entity pair and the text features corresponding to each text data unit in step S40 is implemented by the following expression: (1) (2) (3) in, is the text feature, is the first weight matrix, is the first bias function, is the hidden state sequence, ; First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the second weight matrix, First Entity At the beginning sequence of the hidden state layer, For the second entity At the beginning sequence of the hidden state layer, is the second bias function, ={ ...... } , ={ ...... } .

[0025] In this embodiment, after obtaining the text features and two entity features, they are then sent to the fully connected layer for processing. The fully connected layer plays a key role in feature transformation and classification in the present invention. Specifically, the obtained sentence-level text features and two entity features are first processed. and Activate the operation, Used to prevent overfitting, some input elements are randomly discarded with a certain probability. The hyperbolic tangent activation function is used to introduce nonlinearity. The present invention links the text feature and the two entity features and obtains the vector through a fully connected layer. . use Function obtains a classification probability , and then passed to the cross entropy loss function for loss calculation, where They represent three instances based on index triple matching respectively, and the whole process is given in the following embodiments.

[0026] In one embodiment, the step S50 of determining the vector of each text data unit according to the plurality of entity pair features and the text features is implemented by the following expression: (4) Among them, e is a vector, is the text feature, First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the third weight matrix, is the third bias function.

[0027] In one embodiment, for the step S60, The function converts the vector to obtain the classification probability through the following expression: (5) For the pass in step S70 The function makes a decision on the classification probability to obtain the classification result label of the biomedical text data through the following expression: (6) in, is the classification probability, e is a vector, is the classification result label. Indicates that it returns the value of the independent variable corresponding to the maximum value of a function. This represents the final classification label for the text data. This label is the classification result that determines whether the entity pairs in the text data are related. For each group of predicted text data, all classification results within that group are concatenated sequentially. After all groups are predicted, the concatenated classification results are further merged to form a complete result list, which is then saved as a text file in txt format.

[0028] The biomedical data classification method of this application can be viewed as inputting the biomedical text data to be classified into a data classification model to obtain a classification result label, and the data classification model performs steps S10-S70. During the establishment process, the data classification model needs to be optimized to improve its classification ability. The following is an optimization method for the data classification model: The three public biomedical text datasets used in this experiment all suffer from the homology-different-class problem, where different entity pairs often share the same semantic information, making it difficult for classifiers to capture the distance between text data with different labels. Therefore, the present invention uses the three pairs of entity features contained in the triple text data to calculate the triple loss. The triple loss can be used to improve the model's feature representation capabilities, allowing the model to better distinguish samples from different categories.

[0029] Specifically, first, for each sentence instance in the triple data pair, there is a pair of entity-level vectors and , and the stacking belongs to the stacking feature Furthermore, in order to capture relevant information in entity pairs, the stacking feature Perform one-dimensional convolution to obtain convolution features . Represents a set of entity features. Represents a set of stacked entity features. Finally, the triplet loss is calculated using the three entity-level convolutional feature vectors. The entire process is shown in the following formula: (7) (8) (9) (10) in, represents the true label, Indicates that the predicted sample belongs to The probability of the categories, represents the cross entropy loss, is a stacking feature. is the convolution feature, First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, It's a triple loss.

[0030] Finally, the triple loss and cross entropy loss are combined to obtain the total loss value, as shown in formula (11). This allows the data classification model to learn better feature representations while learning the classification task, thereby improving the generalization performance of the model.

[0031] (11) in, is the total loss value, is a hyperparameter that can be set manually.

[0032] Continuously optimize the loss function during training The data classification model continuously performs predictions on the test or validation set data. The data classification model processes the text data batch by batch, obtains the predicted pairwise number, and then calculates the classification result label using the formula shown below.

[0033] There are three indicators for evaluating model performance, namely P, R, and F1 value.

[0034] Specifically, in order to evaluate the proposed method, the evaluation indicators of the present invention include: precision (P), recall (R) and F1 score, TP: true positive, FP: false positive, TN: true negative, FN: false negative.

[0035] Precision (P): The proportion of samples that are actually true among all samples predicted to be true. It is an important indicator for evaluating the quality of the model in predicting positive samples. The formula is shown in (12).

[0036] (12) Recall (R): The proportion of samples predicted to be true among all actually true samples. It is an important evaluation indicator for evaluating the model's ability to find positive samples. The formula is shown in (13).

[0037] (13) F1 score: It is expressed as the harmonic mean of precision and recall, and attempts to achieve a balance between precision and recall. When precision and recall are equally important and the dataset is severely unbalanced, the F1 score is a very necessary evaluation indicator. The formula is shown in (14).

[0038] (14) Experimental results and analysis of DDI2013 dataset Table 1 Experimental results on the DDI2013 dataset Method P R F1 SCNN 77.5 76.9 77.2 MCCNN - - 79.0 FBK-irst 79.4 80.6 80.0 Tree-LSTM 83.6 84.0 83.0 ATT-BiLSTM - - 84.0 R-BERT 87.12 88.46 87.79 TL-BERT 89.1 88.5 88.8 Ours 88.2 91.2↑ 89.7↑ To compare the performance of these methods, we first conducted a comparison on the DDI2013 dataset using SCNN, MCCNN, FBK-irst, a tree-based LSTM, a bidirectional long short-term memory network with attention, R-BERT, and BERT with triplet loss. The results are shown in Table 1. On the DDI2013 dataset, our method significantly outperformed the comparison methods in all evaluation metrics, including recall and F1 score, and was only slightly behind TL-BERT in precision. This demonstrates the advanced nature of our proposed framework. In particular, compared to the TL-BERT model, we improved the triplet data matching method based on the triplet loss, adding key rule decision making and semantic information indexing strategies. Since the F1 score is the most important evaluation metric for biomedical relationship extraction, ITM-BERT achieved an F1 score of 89.7, 0.9 percentage points higher than TL-BERT and significantly higher than R-BERT. Furthermore, compared to models such as SCNN, MCCNN, Tree-LSTM, and ATT-BLSTM, ITM-BERT utilizes pre-trained strategies as its fundamental components, while other methods do not utilize any pre-training mechanisms. This demonstrates the importance of pre-training mechanisms in biomedical relationship classification tasks. Overall, the proposed method performs well on drug-drug relationship classification tasks.

[0039] PPI dataset experimental results and analysis Table 2 Experimental results on the PPI dataset In order to further compare the advantages and disadvantages of the methods horizontally, the present invention uses Deep neural, McDepCNN, sdpCNN, R-BERT, BERT based on triple loss and the method of the present invention for comparison on the two datasets BioInfer and Aimed, and the results are shown in Table 2. First of all, it is worth noting that on the AImed dataset, the model of the present invention maintains the best performance in all three evaluation indicators. Compared with other methods such as TL-BERT, it has achieved comprehensive improvements. The F1 index is nearly 6 percentage points higher than the TL-BERT model, achieving a substantial performance leap. In the BioInfer dataset, it is observed that our invention also has improved performance over TL-BERT. In particular, sdpCNNHua performs significantly better in BioInfer than in the AImed dataset. Considering that the imbalance of positive and negative samples in the AImed dataset is more serious, we infer that sdpCNN is more suitable for datasets with relatively balanced samples, while our invention method performs better on unbalanced datasets.

[0040] DDI2013, BioInfer, and AImed are typical examples of imbalanced datasets, where the number of positive samples is far less than the number of negative samples. However, the present invention significantly alleviates this imbalance between positive and negative samples by constructing a triplet loss. More importantly, the present invention optimizes triplet data pairs and triplet data construction rules by adding decision rules and indexes. This approach aims to reduce the uncertainty of random selection caused by the lack of homologous matching examples, thereby improving the convergence of the triplet loss algorithm.

[0041] Further experimental analysis of the impact of different indexing strategies is as follows: To compare the impact of different index ranges on performance, we compared the effects of different index range strategies on triple matching data in three datasets. In the triple matching, we randomly selected instances with the same label as the anchor instance, matched instances with opposite categories and high semantic similarity, and verified the impact of the index range on the experimental results by expanding or narrowing the index range. The index range for the opposite category was set to the top 5, top 10, and top 15, respectively. The results are shown in Table 3.

[0042] Table 3 Comparison results in different indicator ranges in DDI2013 Index P R F1 Top 5 87.7 87.4 87.6 Top 15 87.0 89.0 80.0 Top 10 88.2 91.2 89.7 Table 4 Comparison results in different index ranges in BioInfer Index P R F1 Top 5 73.38 75.16 74.46 Top 10 74.39 75.78 75.08 Top 15 76.36 74.22↑ 75.28 Table 5 Comparison results in different indicator ranges in AImed Index P R F1 Top 10 81.82 68.48 74.56 Top 15 73.00 79.35 76.04 Top 5 82.50 71.74 76.44 It can be seen that the method of the present invention has improvements in most range dimensions compared with the comparative experiments, which shows that the index-based triple matching strategy is meaningful.

[0043] At the same time, it should be noted that when constructing triple data, the indexing method is different from obtaining the same label as the anchor instance and the most semantically dissimilar instance, but using a global random method, because there will be some noisy data in the dataset itself. For example, for some extremely non-standard or chaotic data, its sentence-level features will be scattered across all instances, which can easily match these noisy data, causing the data cluster center to be disturbed by the noisy data, affecting the final classification effect.

[0044] Further, the feature distribution experiment analysis is as follows: In order to further verify the role of the index-based triple loss strategy of the present invention in distinguishing the feature distribution of samples, we conducted a visualization study on the feature distribution of samples under different strategies. First, the three rows of images come from different selection strategies of the three data sets, where the first picture in each row represents the feature distribution of the optimal performance of the experiment. Through the influence of different feature distribution graphs on the same data set, it can be observed and concluded that the optimal indexing scheme can reduce the discreteness of features between the same category, and the difference in feature distances between different categories is more obvious. The experimental results prove the effectiveness of the selection strategy of this method. The index-based triple loss strategy adopted on the three data sets affects the feature distribution of the instances. The ITM-BERT algorithm proposed in the present invention can significantly improve the overall performance of the framework, as shown in the following examples. Figure 5 shown.

[0045] This application also provides a biomedical data classification device, such as Figure 3 As shown, the device includes: A preprocessing module 10 is used to obtain biomedical text data to be classified and preprocess the biomedical text data to obtain multiple text data units; A first determining module 20 is configured to determine a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; A second determination module 30 is configured to input the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden state sequence includes a plurality of entity pairs; A third determining module 40 is configured to determine entity pair features corresponding to each entity pair and text features corresponding to each text data unit; A fourth determining module 50 is configured to determine a vector of each text data unit based on the plurality of entity pair features and the text features; The conversion module 60 is used to The function transforms the vector to obtain the classification probability; Decision selection module 70, for The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0046] Figure 6 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 6 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement the biomedical data classification method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement the biomedical data classification method. It will be understood by those skilled in the art that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0047] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: S10: Acquire biomedical text data to be classified, and preprocess the biomedical text data to obtain a plurality of text data units; S20: determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; S30: Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; S40: Determine entity pair features corresponding to each entity pair and text features corresponding to each text data unit; S50: determining a vector of each text data unit according to a plurality of entity pair features and the text features; S60: Pass The function transforms the vector to obtain the classification probability; S70: Pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0048] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the processor performs the following steps: S10: Acquire biomedical text data to be classified, and preprocess the biomedical text data to obtain a plurality of text data units; S20: determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; S30: Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; S40: Determine entity pair features corresponding to each entity pair and text features corresponding to each text data unit; S50: determining a vector of each text data unit according to a plurality of entity pair features and the text features; S60: Pass The function transforms the vector to obtain the classification probability; S70: Pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

[0049] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0050] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0051] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A biomedical data classification method, characterized in that: The method comprises: Acquiring biomedical text data to be classified, and preprocessing the biomedical text data to obtain a plurality of text data units; determining a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; Inputting the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; Determining an entity pair feature corresponding to each entity pair and a text feature corresponding to each of the text data units; Determine a vector for each text data unit according to a plurality of entity pair features and the text features; pass The function transforms the vector to obtain the classification probability; pass The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

2. The biomedical data classification method according to claim 1, characterized in that: The preprocessing of the biomedical text data to obtain a plurality of text data units comprises: A CLS tag is added before each biomedical text data to obtain labeled text data, and each labeled text data is segmented by a tokenization function to obtain multiple text data units.

3. The biomedical data classification method according to claim 1, characterized in that: Determining a triple data pair according to each of the text data units comprises: When there are positive text instances and negative text instances of the same origin in the text data unit: Acquire a first text data unit from the plurality of text data units, and determine a first negative text instance corresponding to the first text data unit, wherein the first text data unit is a text data unit marked as a positive text instance; The first text data unit, the first negative text instance, and the second text data unit constitute the triple data pair, wherein the second text data unit is a text data unit randomly selected from a plurality of the text data units and marked as a positive text instance; and Acquire a third text data unit from the plurality of text data units, and determine a first positive text instance corresponding to the third text data unit, wherein the third text data unit is a text data unit marked as a negative text instance; The third text data unit, the first positive text instance and the fourth text data unit constitute the triple data pair; the fourth text data unit is a text data unit randomly selected from the plurality of text data units and marked as a negative text instance.

4. The biomedical data classification method according to claim 1, characterized in that: Determining a triple data pair according to each of the text data units comprises: When there are no positive text instances and negative text instances of the same origin in the text data unit: Obtain a fifth text data unit from the plurality of text data units, randomly match a first text data having the same text instance label as the fifth text data unit in the Faiss library based on the BGE model, and randomly match a second text data having an opposite text instance label as the fifth text data unit in the Faiss library based on the BGE model; the text instance label is a positive text instance or a negative text instance; The fifth text data unit, the first text data, and the second text data constitute the triple data pair.

5. The biomedical data classification method according to claim 1, characterized in that: The determination of the entity pair features corresponding to each entity pair and the text features corresponding to each text data unit is achieved by the following expression: in, is the text feature, is the first weight matrix, is the first bias function, is the hidden state sequence, ; First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the second weight matrix, First Entity At the beginning sequence of the hidden state layer, For the second entity At the beginning sequence of the hidden state layer, is the second bias function, ={ ...... } , ={ ...... } .

6. The biomedical data classification method according to claim 5, characterized in that: The vector of each text data unit is determined according to the plurality of entity pair features and the text features by the following expression: Among them, e is a vector, is the text feature, First Entity The corresponding entity pair features, For the second entity The corresponding entity pair features, is the third weight matrix, is the third bias function.

7. The biomedical data classification method according to claim 6, characterized in that: Said through The function transforms the vector to obtain the classification probability; The function makes a decision on the classification probability to obtain the classification result label of the biomedical text data through the following expression: in, is the classification probability, e is a vector, is the classification result label.

8. A biomedical data classification device, characterized in that: The device comprises: A preprocessing module, configured to obtain biomedical text data to be classified and preprocess the biomedical text data to obtain a plurality of text data units; a first determining module, configured to determine a triple data pair according to each of the text data units, wherein the number of the triple data pairs is the same as the number of the text data units; A second determination module is configured to input the triple data pair into an R-BERT model, wherein the R-BERT model converts the triple data pair into a hidden state sequence, wherein the hidden sequence includes a plurality of entity pairs; A third determining module is used to determine an entity pair feature corresponding to each entity pair and a text feature corresponding to each text data unit; The fourth determining module is used to determine the entity pair features and the text features according to the entity pair features and the text features. Determining a vector for each of the text data units; Conversion module, used to The function transforms the vector to obtain the classification probability; Decision selection module, used to The function makes a decision selection on the classification probability to obtain a classification result label of the biomedical text data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Entity relationship mining method based on biomedical literature

    CN111428036A

  • Medical entity relationship extraction method and device

    CN112599211A

  • Biomedical entity relation extraction method based on ternary loss training strategy

    CN114238561A

  • Children medical text data classification method

    CN118733772A

  • Entity relation mining method based on biomedical literature

    US20230007965A1