Text representation and language model training method and device, equipment and storage medium

By comparing the BERT language model and TF-IDF technology after fine-tuning of learning technology, Chinese medical texts are characterized and processed, solving the problem of aligning medical terms and distinguishing the importance of words in the prior art, and achieving a more accurate and distinctive medical text representation.

CN120146039APending Publication Date: 2025-06-13THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411480272.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing text representation techniques are difficult to align medical terms with the same meaning in Chinese medical texts, and cannot effectively distinguish the importance of words, resulting in the learned text representations not being strongly distinguishable.

Method used

The BERT language model fine-tuned by comparative learning technology is used to characterize sentences and words in the text to be processed, and the representation aggregation strategy of TF-IDF weighted average word vectors is used to improve the accuracy of medical text representation.

Benefits of technology

By comparing the BERT language model fine-tuned by learning technology, we can effectively align words with the same meaning in the text, improve the ability to represent medical texts, and make artificial intelligence technology more effective in intelligent medical services. At the same time, combined with TF-IDF technology, the distinction and accuracy of text representation are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146039A_ABST
    Figure CN120146039A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, in particular to a text representation and language model training method and device, equipment and a storage medium. The method comprises the steps that sentences in a to-be-processed text and words in the sentences are acquired; sentences in the to-be-processed text and words in the sentences are input into the trained BERT language model, first sentence-level representation of the sentences in the to-be-processed text and representation of the words in the sentences output by the BERT language model are obtained, and the trained BERT language model can align the words with the same meaning during text representation; calculating TF-IDF weights of words in the sentences; obtaining a second sentence level representation of the sentence based on the TF-IDF weight of the word in the sentence and the representation of the word in the sentence; and combining the first sentence-level representation and the second sentence-level representation of the sentence to obtain a text representation of the sentence in the to-be-processed text. According to the technical scheme, the text representation capability can be improved, and the Chinese medical text is mainly represented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of text processing, and in particular, to a text representation and language model training method, apparatus, device, and storage medium. Background Art

[0002] Chinese medical texts are full of professional terms and abbreviations, usually including disease names, drug names, treatment procedures, etc. In addition, Chinese medical texts often contain complex medical information, such as symptom descriptions, diagnosis results, treatment plans, etc. This information may be written by nurses, doctors, or other medical professionals, with variable structures and rich and complex contents. This results in Chinese medical texts usually being unstructured and free-form texts. The text representation technology is a text preprocessing process that converts unstructured text content into structured data that can be understood and processed by a computer, and is a key link in natural language processing. This conversion transforms unstructured text content into structured data, enabling the text data to be applied to machine learning models for tasks such as classification and prediction. The quality of text representation will directly affect the accuracy and applicability of subsequent text information mining. The text representation of Chinese medical texts has important application values in scenarios such as medical information extraction, medical question-answering systems, medical text anomaly detection, and clinical decision support systems. By accurately representing and processing Chinese medical texts, the utilization rate and accuracy of medical information can be improved, providing effective auxiliary diagnosis and treatment suggestions for physicians and customized more accurate and personalized treatment plans for patients, so as to improve the problems of shortage of medical resources and uneven distribution of medical resources.

[0003] Currently, the commonly used text representation scheme is the word embedding technology. A deep word vector model can be constructed by using deep learning technology and large-scale text corpora. The deep word vector model can map words to a dense vector space, and these vectors can capture the semantic relationships between words. Words with similar semantics usually have similar vector representations. However, compared with general Chinese text data, Chinese medical texts have obvious inconsistencies in expression and word usage. Different hospitals and different doctors may have different descriptions for the same disease or drug. For example, diabetes and hyperglycemia, penicillin and benzylpenicillin, etc. Existing text representation technologies cannot align these medical terms to obtain similar representations. Moreover, existing text representation technologies do not distinguish the importance of words, which easily leads to the learned text representations not having strong discriminability. Summary of the Invention

[0004] To solve the problems in the related art, embodiments of the present disclosure provide a text representation and language model training method, apparatus, device, and storage medium.

[0005] In a first aspect, an embodiment of the present disclosure provides a text representation method, including:

[0006] Preprocess the text to be processed to obtain sentences in the text to be processed and words in the sentences;

[0007] Input the sentences in the text to be processed and the words in the sentences into a trained BERT language model, execute the trained BERT language model, and obtain a first sentence-level representation of the sentences in the text to be processed and a representation of the words in the sentences output by the trained BERT language model. The trained BERT language model is a model obtained by fine-tuning a pre-trained BERT language model based on contrastive learning technology, and the trained BERT language model can align words with the same meaning during text representation;

[0008] Calculate the term frequency-inverse document frequency TF-IDF weights of the words in the sentence;

[0009] Based on the TF-IDF weights of the words in the sentence and the representation of the words in the sentence, obtain a second sentence-level representation of the sentence;

[0010] Merge the first sentence-level representation and the second sentence-level representation of the sentence to obtain a text representation of the sentence in the text to be processed.

[0011] In a possible implementation manner, the method further includes:

[0012] Obtain a plurality of word pairs, each word pair including at least two words with the same meaning;

[0013] Use the plurality of word pairs to fine-tune the model parameters of the pre-trained BERT language model so that the similarity between the representations of the two words in the same word pair is greater until the adjustment termination condition is reached, and obtain a trained BERT language model.

[0014] In a possible implementation manner, the method further includes:

[0015] Sample a plurality of samples from the training set, where each sample is composed of two partially masked sentences in the text to be processed spliced together;

[0016] Input the plurality of samples into the BERT language model to be trained for masked word prediction training and next sentence prediction training, and calculate the sum of the losses of the masked word prediction and the next sentence prediction as the model loss;

[0017] Based on the model loss, use gradient descent to update the model parameters of the BERT language model to be trained;

[0018] Repeat the above steps until the model converges or meets the termination condition. The training ends, and a pre-trained BERT language model is obtained.

[0019] In a possible implementation, calculating the TF-IDF weights of the words in the sentence includes:

[0020] Calculate the TF-IDF weight of the j-th word in sentence s according to the following formula i where the TF-IDF weight of the j-th word in sentence s is

[0021]

[0022] where tf i,j represents the frequency of the j-th word in sentence s i idf j represents the total number of sentences in the text to be processed that contain the j-th word, N is the total number of sentences in the text to be processed, 1 ≤ idf j ≤ N, and m is the total number of words in sentence s i where m is the total number of words in sentence s.

[0023] In a possible implementation, the text to be processed includes Chinese medical texts, and the first sentence-level representation is the CLS Token.

[0024] In a second aspect, an embodiment of the present disclosure provides a method for training a language model, including:

[0025] Obtain a plurality of word pairs, each word pair including at least two words with the same meaning;

[0026] Fine-tune the model parameters of the pre-trained BERT language model using the plurality of word pairs so that the similarity between the representations of the two words in the same word pair is greater until the adjustment termination condition is reached, and a trained BERT language model is obtained.

[0027] In a third aspect, an embodiment of the present disclosure provides a text representation device, including:

[0028] A preprocessing module configured to preprocess the text to be processed to obtain the sentences in the text to be processed and the words in the sentences;

[0029] A model processing module, configured to input sentences in the text to be processed and words in the sentences into a trained BERT language model, execute the trained BERT language model, and obtain a first sentence-level representation of the sentences in the text to be processed and representations of the words in the sentences output by the trained BERT language model. The trained BERT language model is a model obtained by fine-tuning a pre-trained BERT language model based on contrastive learning techniques, and the trained BERT language model can align words with the same meaning during text representation;

[0030] A weight calculation module, configured to calculate the term frequency-inverse document frequency TF-IDF weights of the words in the sentence;

[0031] A weighting module, configured to obtain a second sentence-level representation of the sentence based on the TF-IDF weights of the words in the sentence and the representations of the words in the sentence;

[0032] A representation merging module, configured to merge the first sentence-level representation and the second sentence-level representation of the sentence to obtain a text representation of the sentence in the text to be processed.

[0033] In a fourth aspect, an embodiment of the present disclosure provides a language model training device, including:

[0034] A word pair acquisition module, configured to acquire a plurality of word pairs, each word pair including at least two words with the same meaning;

[0035] A fine-tuning module, configured to fine-tune the model parameters of the pre-trained BERT language model using the plurality of word pairs, so that the similarity of the representations of the two words in the same word pair is greater, until an adjustment termination condition is reached, to obtain a trained BERT language model.

[0036] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer instructions, and wherein the one or more computer instructions are executed by the processor to implement the method according to any one of the first aspect.

[0037] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the method according to any one of the first aspect is implemented.

[0038] According to the technical solution provided by the embodiments of the present disclosure, by using the BERT language model fine-tuned through contrastive learning technology, text representation is performed on the sentences in the text to be processed and the words in the sentences, which can align the words with the same meaning in the text, further improve the medical text representation ability, enable the artificial intelligence technology to be more effectively applied to intelligent medical services, and provide more useful and effective medical services for physicians and patients; at the same time, the BERT language model combines the representation aggregation strategy of TF-IDF weighted average word vectors, which not only utilizes the powerful text representation ability of the BERT language model but also makes up for the possible insufficient representation of key vocabulary in the text, providing a more accurate text representation for subsequent medical text information mining tasks.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In combination with the drawings, through the following detailed description of non-limiting embodiments, other features, objects, and advantages of the present disclosure will become more obvious. In the drawings:

[0041] Figure 1 The flowchart of a text representation method provided by an embodiment of the present disclosure is shown.

[0042] Figure 2 The flowchart of the training process of a pre-trained BERT language model provided by an embodiment of the present disclosure is shown.

[0043] Figure 3 The flowchart of a language model training method provided by an embodiment of the present disclosure is shown.

[0044] Figure 4 The structural block diagram of a text representation device provided by an embodiment of the present disclosure is shown.

[0045] Figure 5 The structural block diagram of a language model training provided by an embodiment of the present disclosure is shown.

[0046] Figure 6 The structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0047] Figure 7 The structural schematic diagram of a computer system suitable for implementing the method of the embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In the following, the exemplary embodiments of the present disclosure will be described in detail with reference to the drawings, so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts irrelevant to the description of the exemplary embodiments are omitted in the drawings.

[0049] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0050] In addition, it should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0051] Figure 1 The flowchart of a text representation method provided by an embodiment of the present disclosure is shown. As Figure 1 shown, the text representation method includes the following steps S101 - S105:

[0052] In step S101, the text to be processed is preprocessed to obtain the sentences in the text to be processed and the words in the sentences.

[0053] In step S102, the sentences in the text to be processed and the words in the sentences are input into the trained BERT language model, and the trained BERT language model is executed to obtain the first sentence-level representation of the sentences in the text to be processed output by the trained BERT language model and the representations of the words in the sentences.

[0054] Among them, the trained BERT language model is a model obtained by fine-tuning the pre-trained BERT language model based on contrastive learning technology, and the trained BERT language model can align words with the same meaning during text representation.

[0055] In step S103, the term frequency-inverse document frequency TF-IDF weights of the words in the sentence are calculated.

[0056] In step S104, based on the TF-IDF weights of the words in the sentence and the representations of the words in the sentence, a second sentence-level representation of the sentence is obtained.

[0057] In step S105, the first sentence-level representation and the second sentence-level representation of the sentence are combined to obtain the text representation of the sentence in the text to be processed.

[0058] In a possible implementation manner, this text representation method is applicable to devices such as computers, computing devices, servers, server clusters, etc. that can perform text representation.

[0059] In a possible implementation, this text representation method mainly performs text representation processing on the text to be processed to obtain the text representations of each sentence in the text to be processed. The text to be processed here can be any kind of text, such as Chinese medical texts, academic papers, media articles, and so on. Especially for Chinese medical texts with complex content and variable structures, it has better representation effects.

[0060] In a possible implementation, after obtaining the text to be processed, the original text cannot be directly subjected to representation learning. It is necessary to preprocess the text to be processed to segment the sentences and words in the text to be processed. By way of example, taking the text to be processed as a Chinese medical text, the text to be processed can be first segmented into sentences, and the text to be processed is segmented based on specific delimiters (such as full stops, semicolons, question marks, etc.) in the text to be processed to obtain a set S = {s 1 , s 2 , …, s n} of sentences or paragraphs. Since there are no obvious delimiters between words in Chinese, it is necessary to use word segmentation technology to segment each sentence into individual words. The purpose of word segmentation is to segment a continuous text sequence into a meaningful word sequence, so as to facilitate subsequent text analysis and processing. Existing natural language processing technologies (Natural Language Processing, NLP) provide efficient and feasible text word segmentation technologies (for example, jieba segmentation, Snow NLP, etc.). The specific word segmentation technology is not limited in this disclosure example, and the optimal word segmentation technology can be selected according to the actual scenario requirements. Of course, for texts in other languages, other preprocessing technologies can be used to segment the text to be processed into multiple sentences and segment the sentences into words, which is not limited here.

[0061] In a possible implementation, the pre-trained BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representation based on Transformer) is a pre-trained language representation model based on the Transformer architecture. It has learned rich language patterns through pre-training using a large amount of text data, and thus has achieved remarkable results in various natural language processing tasks. The core feature of the pre-trained BERT language model lies in its bidirectional training mechanism, which enables the model to capture the context information of words, thereby better understanding the meaning of language. The pre-training of BERT includes two main tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP).

[0062] Among them, in the task of masked language model (MLM), the model can randomly mask some words in the input sentence and try to predict the original content of these masked words, which helps the model learn the meaning of words in a specific context. In the next sentence prediction (NSP) task, the model receives a pair of sentences and judges whether the second sentence is a reasonable continuation of the first sentence. The purpose of this task is to enable the language model to understand the semantic relevance between sentences. Many downstream text mining tasks are based on this level of language understanding ability. Especially for Chinese medical texts, subsequent processing and mining technologies highly rely on the text representation's ability to learn the semantic relationships between sentences.

[0063] In a possible implementation, to better perform text representation on the text to be processed, the trained BERT language model used in this disclosure is a model obtained by fine-tuning the model parameters in the pre-trained BERT language model using contrastive learning technology (Contrastive Learning). Contrastive learning technology is a self-supervised learning method that learns the representation of data by comparing the similarities or differences between different samples. The core idea of this method is that if two samples are similar, their representations should be close; if two samples are dissimilar, their representations should be far apart. Thus, after fine-tuning the model using contrastive learning technology, when the fine-tuned BERT language model performs text representation on diabetes and hyperglycemia, penicillin and penicillin, which have the same meaning, it will output similar representations.

[0064] In a possible implementation, the input of the trained BERT language model is the sentence in the text to be processed and the words in the sentence, and the output is the first sentence-level representation of the sentence and the representations of the words in the sentence. For example, when the sentence s in the text to be processed i and the m words in sentence s i are input into the trained BERT language model, after executing the trained BERT language model, the representations {t i , t 1 , …, t 2 , …, t m} of the m words in sentence s i and the first sentence-level representation c can be obtained. Preferably, the first sentence-level representation is the CLS Token; this CLS Token can be directly output by the trained BERT language model without other processing. This CLS (Classify, sentence-level marker) Token can represent the semantic association between sentences and can perform better sentence representation.

[0065] In a possible implementation, TF-IDF (Term Frequency-Inverse Document Frequency) is a weighting technique. It reflects the importance of a word for the text to be processed. The TF-IDF weight can increase as the frequency of the word in the sentence (Term Frequency, TF) increases, but at the same time, it will decrease as the frequency of the word in the document to be processed (Document Frequency, DF) increases. This means that TF-IDF tends to filter out common words and retain important words. The TF-IDF weight of the words in the sentence can be calculated based on this TF-IDF technique.

[0066] In a possible implementation, the representation of the words in the sentence can be weighted and calculated based on the TF-IDF weight of the words in the sentence to obtain the second sentence-level representation of the sentence. For some downstream tasks, the effective representation of a specific part (word) in the sentence may be more useful than the aggregated representation of the entire sentence. Therefore, in this implementation, the TF-IDF technique is used to perform a weighted calculation on the word vectors learned by the language model to obtain another sentence-level text representation, that is, the second sentence-level representation. Most of the information in this second sentence-level representation comes from the words with high TF-IDF values, highlighting the role of important words in the text representation. Then, the first sentence-level representation and the second sentence-level representation of the sentence are merged to obtain the text representation of the sentence in the text to be processed. In this way, the text representations of each sentence in the text to be processed can be obtained, that is, the text representation of the text to be processed.

[0067] Exemplarily, the first sentence-level representation c i of the i-th sentence s i in the text to be processed and the second sentence-level representation v i can be merged to obtain the text representation e i of the sentence in the text to be processed:

[0068]

[0069] where c i is the first sentence-level representation of the sentence s i output by the trained BERT language model, t j is the representation of the j-th word in the sentence s i output by the trained BERT language model, is the TF-IDF weight of the j-th word in the sentence s i , and the second sentence-level representation v i is obtained using For t j obtained after weighted and summation operations, where m is the number of words in the i-th sentence s i Here, represents the vector concatenation operator.

[0070] This embodiment uses the BERT language model fine-tuned by contrastive learning technology to perform text representation on the sentences in the text to be processed and the words in the sentences, which can align the words with the same meaning in the text, further improve the medical text representation ability, enable the artificial intelligence technology to be more effectively applied to intelligent medical services, and provide more useful and effective medical services for physicians and patients; at the same time, the BERT language model combines the representation aggregation strategy of TF-IDF weighted average word vectors, which not only utilizes the powerful text representation ability of the BERT language model but also makes up for the possible deficiency of keyword vocabulary in text representation, providing a more accurate text representation for subsequent medical text information mining tasks.

[0071] In a possible implementation manner, the method further includes:

[0072] Obtain a plurality of word pairs, each word pair including at least two words with the same meaning;

[0073] Use the model parameters of the pre-trained BERT language model pre-trained with the plurality of word pairs for fine-tuning, so that the similarity of the representations of the two words in the same word pair is greater until the adjustment termination condition is reached, and a trained BERT language model is obtained.

[0074] In this implementation manner, in order to align the words with the same meaning, after obtaining the pre-trained BERT language model, the contrastive learning technology can be used to fine-tune the pre-trained BERT language model to further improve the semantic representation ability of the model.

[0075] In this implementation manner, a plurality of word pairs can be used to fine-tune the pre-trained BERT language model. Each word pair includes at least two words with the same meaning but different wordings. When the text to be processed is a Chinese medical text, the word pair is different medical terms with the same meaning. For example, diabetes and hyperglycemia are a pair of word pairs, and penicillin and benzylpenicillin are a pair of word pairs, and so on.

[0076] In this implementation manner, l word pairs can be obtained, where l is a relatively large value. Each word pair includes at least two words with the same meaning. These l word pairs can be denoted as These l word pairs can be used as a training set to fine-tune the model parameters of the pre-trained BERT language model. For any word pair t i, sample two words each time As the input of the pre-trained BERT language model, after executing the model, obtain the vector representations of these two words output by the pre-trained BERT language model Calculate the similarity between them. In this embodiment, the dot product of vectors can be used as the similarity measure between two representation vectors. When adjusting the model parameters, it is necessary to make the sum of the similarities of the vector representations of the two words in the same word pair among these l word pairs large enough until the adjustment termination condition is met. The adjustment termination condition can be that the epoch (number of iterations) reaches one or two, and then the adjusted parameter θ can be obtained ** , and obtain the trained language model , for example, the following formula can be used to fine-tune the model parameter θ of the pre-trained BERT language model * :

[0077]

[0078] In a possible embodiment, the method further includes:

[0079] Sample multiple samples from the training set, and each sample is composed of two partially masked sentences spliced from the text to be processed;

[0080] Input the multiple samples into the BERT language model to be trained for masked word prediction training and next sentence prediction training, and calculate the sum of the losses of masked word prediction and next sentence prediction as the model loss;

[0081] Based on the model loss, use gradient descent to update the model parameters of the BERT language model to be trained;

[0082] Repeat the above steps until the model converges or meets the termination condition, and the training ends to obtain the trained pre-trained BERT language model.

[0083] In this embodiment, first, a predetermined proportion (such as 10%) of the words in each sentence of the text to be processed need to be masked, for example, replaced with [MASK] tags to obtain the masked sentence set after masking Then, sample two masked sentences from the masked sentence set and splice them. For example, sample and for splicing, and the splicing method is "[CLS] [SEP] [SEP]", where [CLS] and [SEP] are identifiers; splice them pairwise in this way, and a second predetermined proportion (such as 50% or 49%) of It needs to be the next sentence in the text to be processed, and the remaining proportion of the concatenated data can be any combination of sentences; after concatenation, the training set X is obtained. For the next sentence, the remaining proportion of the concatenated data can be any combination of sentences; after concatenation, the training set X is obtained.

[0084] In this embodiment, Figure 2 A flowchart showing the training process of a pre-trained BERT language model provided by an embodiment of the present disclosure is shown. As Figure 2 shown, multiple samples can be sampled from the training set X, for example, bs samples {x 1 , x 2 , …, x bs}, and the multiple samples are input into the BERT language model f to be trained θ . The BERT language model to be trained can use these bs samples to perform two training subtasks: masked word prediction and next sentence prediction. The masked word prediction task is to let the BERT language model f to be trained θ predict the replaced vocabulary based on the surrounding words; the purpose of doing this is to enable the model to not only learn the semantic order of the words in the sample during training, but also learn to infer the word meaning according to the context. This training task can use cross entropy as the loss function, denoted as In the next sentence prediction task, it is necessary to train the model by judging whether two concatenated sentences actually appear continuously to improve the model's semantic understanding ability at the sentence level, and can provide effective text representations for downstream sentence-level text mining tasks. This task can also use cross entropy as the loss function, denoted as

[0085] In this embodiment, in order to enable the model to simultaneously learn the semantic understanding ability and context association ability at the word level and sentence level, this embodiment simultaneously optimizes the above-mentioned masked word prediction task and next sentence prediction task. Therefore, the loss function of the model loss is as follows:

[0086]

[0087] In this embodiment, based on the model loss, the model parameters θ of the BERT language model f to be trained can be updated using gradient descent θ ; then, the sampling of samples, the calculation of the loss function, and the update of the model parameters can be repeated until the model converges or meets the termination condition (such as the number of updates of the model parameters reaches a preset threshold, etc.), and the training ends, obtaining the trained pre-trained BERT language model .

[0088] In a possible implementation, calculating the TF-IDF weights of the words in the sentence includes:

[0089] Calculating the TF-IDF weight of the j-th word in sentence s i according to the following formula

[0090]

[0091] where tf i,j represents the frequency of occurrence of the j-th word in sentence s i idf j represents the total number of sentences in the text to be processed in which the j-th word appears, N is the total number of sentences in the text to be processed, 1 ≤ idf j ≤ N, and m is the total number of words in sentence s i .

[0092] The present disclosure also provides a method for training a language model. Figure 3 The flowchart showing a method for training a language model provided by an embodiment of the present disclosure is as Figure 3 shown. The method may include the following steps:

[0093] In step S301, a plurality of word pairs are obtained, and each word pair includes at least two words with the same meaning.

[0094] In step S302, the model parameters of the pre-trained BERT language model are fine-tuned using the plurality of word pairs so that the similarity of the representations of the two words in the same word pair is greater until the adjustment termination condition is reached, and a trained BERT language model is obtained.

[0095] In a possible implementation, this language model training is applicable to devices such as computers, computing devices, servers, and server clusters that can perform language model training.

[0096] In a possible implementation, the pre-trained BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language representation model based on the Transformer architecture. By using a large amount of text data for pre-training, it has learned rich language patterns, thus achieving remarkable results in various natural language processing tasks. The core feature of the pre-trained BERT language model lies in its bidirectional training mechanism, which enables the model to capture the context information of words, thereby better understanding the meaning of language. The pre-training of BERT includes two main tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP).

[0097] Among them, in the Masked Language Model (MLM) task, the model can randomly mask some words in the input sentence and try to predict the original content of these masked words, which helps the model learn the meaning of words in a specific context. In the Next Sentence Prediction (NSP) task, the model receives a pair of sentences and judges whether the second sentence is a reasonable continuation of the first sentence, which helps the model understand the relationship between sentences.

[0098] In a possible implementation, in order to better perform text representation on the text to be processed, the trained BERT language model used in this disclosure is a model obtained by fine-tuning the model parameters in the pre-trained BERT language model using contrastive learning technology. Contrastive learning technology is a self-supervised learning method that learns the representation of data by comparing the similarities or differences between different samples. The core idea of this method is that if two samples are similar, their representations should be close; if two samples are dissimilar, their representations should be far apart. Thus, after fine-tuning the model using contrastive learning technology, when the fine-tuned BERT language model performs text representation on diabetes and hyperglycemia, penicillin and benzylpenicillin with the same meaning, it will output similar representations.

[0099] In a possible implementation, each word pair includes at least two words with different expressions but the same meaning. When the text to be processed is Chinese medical text, the word pair is different medical terms with the same meaning. For example, diabetes and hyperglycemia are a word pair, penicillin and benzylpenicillin are a word pair, and so on.

[0100] In a possible implementation, l word pairs can be obtained, where l is a relatively large value. Each word pair includes at least two words with the same meaning. These l word pairs can be denoted as These l word pairs can be used as a training set to fine-tune the model parameters of the pre-trained BERT language model. For any word pair t i , every time two words are sampled as the input of the pre-trained BERT language model. After executing the model, the vector representations of these two words output by the pre-trained BERT language model are obtained Calculate the similarity between them. In this implementation, the dot product of vectors can be used as the similarity metric between two representation vectors. When adjusting the model parameters, it is necessary to make the sum of the similarities of the vector representations of the two words in the same word pair among these l word pairs large enough until the adjustment termination condition is met. The adjustment termination condition can be that the epoch (number of iterations) reaches one or two, and then the adjusted parameter θ can be obtained ** , and a trained language model is obtained . Exemplarily, the following formula can be used to fine-tune the model parameter θ of the pre-trained BERT language model * :

[0101]

[0102] Through the contrastive learning technique, the BERT language model fine-tuned in this implementation can obtain similar text representations for words with the same meaning when using this fine-tuned BERT language model for text representation. In this way, the medical text representation ability can be further improved, enabling the artificial intelligence technology to be more effectively applied to intelligent medical services and providing more useful and effective medical services for physicians and patients.

[0103] The present disclosure also provides a text representation device. Figure 4 FIG. shows a structural block diagram of a text representation device provided by an embodiment of the present disclosure. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 4 shown, the text representation device includes:

[0104] A preprocessing module 401, configured to preprocess the text to be processed to obtain the sentences in the text to be processed and the words in the sentences;

[0105] The model processing module 402 is configured to input the sentences in the text to be processed and the words in the sentences into a trained BERT language model, execute the trained BERT language model, and obtain the first sentence-level representation of the sentences in the text to be processed and the representations of the words in the sentences output by the trained BERT language model. The trained BERT language model is a model obtained by fine-tuning a pre-trained BERT language model based on contrastive learning techniques. The trained BERT language model can align words with the same meaning during text representation;

[0106] The weight calculation module 403 is configured to calculate the term frequency-inverse document frequency (TF-IDF) weights of the words in the sentences;

[0107] The weighting module 404 is configured to obtain a second sentence-level representation of the sentences based on the TF-IDF weights of the words in the sentences and the representations of the words in the sentences;

[0108] The representation merging module 405 is configured to merge the first sentence-level representation and the second sentence-level representation of the sentences to obtain the text representation of the sentences in the text to be processed.

[0109] In a possible implementation manner, the apparatus further includes:

[0110] The model fine-tuning module is configured to obtain multiple word pairs, each word pair including at least two words with the same meaning; use the model parameters of the pre-trained BERT language model with the multiple word pairs for fine-tuning, so that the similarity of the representations of the two words in the same word pair is greater until the adjustment termination condition is reached, and a trained BERT language model is obtained.

[0111] In a possible implementation manner, the apparatus further includes:

[0112] The pre-training module is configured to: sample multiple samples from the training set, where each sample is composed of two partially masked sentences spliced from the text to be processed; input the multiple samples into the BERT language model to be trained for masked word prediction training and next sentence prediction training, and calculate the sum of the losses of the masked word prediction and the next sentence prediction as the model loss; based on the model loss, use gradient descent to update the model parameters of the BERT language model to be trained; repeat the above steps until the model converges or meets the termination condition, and the training ends, obtaining a pre-trained BERT language model.

[0113] In a possible implementation manner, the weight calculation module is configured to:

[0114] Calculate the sentence s according to the following formula iThe TF-IDF weight of the j-th word in

[0115]

[0116] wherein, tf i,j represents the frequency of occurrence of the j-th word in the sentence s i and idf j represents the total number of sentences in the text to be processed in which the j-th word appears, N is the total number of sentences in the text to be processed, 1 ≤ idf j ≤ N, and m is the total number of words in the sentence s i .

[0117] In a possible implementation manner, the text to be processed includes Chinese medical texts.

[0118] An embodiment of the present disclosure also provides a language model training device, Figure 5 as shown in the structural block diagram of a language model training provided by an embodiment of the present disclosure. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. As Figure 5 shown, the language model training includes:

[0119] A word pair acquisition module 501, configured to acquire a plurality of word pairs, each word pair including at least two words with the same meaning;

[0120] A fine-tuning module 502, configured to fine-tune the model parameters of the pre-trained BERT language model using the plurality of word pairs, so that the similarity of the representations of the two words in the same word pair is greater, until the adjustment termination condition is reached, and a trained BERT language model is obtained.

[0121] The technical terms and technical features mentioned in the implementation manner of this device are the same as or similar to those mentioned in the above method implementation manner. For the explanations and descriptions of the technical terms and technical features involved in this device, reference can be made to the explanations and descriptions of the above method implementation manner, and details are not described herein again.

[0122] The present disclosure also discloses an electronic device, Figure 6 as shown in the structural block diagram of the electronic device according to an embodiment of the present disclosure.

[0123] As Figure 6 shown, the electronic device 600 includes a memory 601 and a processor 602. Among them, the memory 601 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor 602 to implement the method according to the embodiment of the present disclosure.

[0124] Figure 7A schematic structural diagram of a computer system suitable for implementing the method of the embodiments of the present disclosure is shown.

[0125] As Figure 7 shown, the computer system 700 includes a processing unit 701, which can perform various processes in the above embodiments according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the computer system 700 are also stored. The processing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0126] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read from it can be installed into the storage section 708 as required. Among them, the processing unit 701 can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0127] Specifically, according to the embodiments of the present disclosure, the above-described method can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes computer instructions that implement the above-described method steps when executed by a processor. In such an embodiment, the computer program product can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711.

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0129] The units or modules involved in the embodiments described in the present disclosure can be implemented in software or in programmable hardware. The described units or modules can also be provided in a processor, and the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0130] As another aspect, the present disclosure also provides a computer-readable storage medium, which can be the computer-readable storage medium included in the electronic device or computer system in the above embodiments; or it can exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the methods described in the present disclosure.

[0131] The above description is only the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

Claims

1. A text representation method, characterized in that: include: Preprocessing the text to be processed to obtain sentences in the text to be processed and words in the sentences; Inputting sentences in the text to be processed and words in the sentences into a trained BERT language model, executing the trained BERT language model, and obtaining first sentence-level representations of the sentences in the text to be processed and representations of the words in the sentences output by the trained BERT language model, wherein the trained BERT language model is a model obtained by fine-tuning a pre-trained BERT language model based on contrastive learning technology, and the trained BERT language model can align words with the same meaning when representing text; Calculate the term frequency-inverse document frequency TF-IDF weight of the words in the sentence; Obtaining a second sentence-level representation of the sentence based on the TF-IDF weights of the words in the sentence and the representations of the words in the sentence; The first sentence-level representation and the second sentence-level representation of the sentence are combined to obtain a text representation of the sentence in the text to be processed.

2. The method according to claim 1, characterized in that: The method further comprises: Acquire a plurality of word pairs, each word pair including at least two words with the same meaning; The model parameters of the pre-trained BERT language model are fine-tuned using the multiple words until the similarities of the representations of two words in the same word pair reach an adjustment termination condition, thereby obtaining a trained BERT language model.

3. The method according to claim 1, characterized in that The method further comprises: Sampling a plurality of samples from a training set, wherein the samples are formed by concatenating two partially masked sentences in the text to be processed; Inputting the multiple samples into the BERT language model to be trained to perform mask word prediction training and next sentence prediction training, and calculating the sum of the losses of mask word prediction and next sentence prediction as the model loss; Based on the model loss, updating the model parameters of the BERT language model to be trained using gradient descent; Repeat the above steps until the model converges or meets the termination condition, the training ends, and the pre-trained BERT language model is obtained.

4. The method according to claim 1, characterized in that The calculating the term frequency-inverse document frequency TF-IDF weight of the words in the sentence includes: Calculate the sentence s according to the following formula i The TF-IDF weight of the jth word in Among them, tf i,j Represents the jth word in sentence s i The frequency of occurrence, idf j represents the total number of sentences in which the jth word appears in the text to be processed, N is the total number of sentences in the text to be processed, 1≤idf j ≤N, m is the sentence s i The total number of words in .

5. The method according to claim 1, characterized in that The text to be processed includes Chinese medical text, and the first sentence-level representation is CLS Token.

6. A language model training method, characterized in that: include: Acquire a plurality of word pairs, each word pair including at least two words with the same meaning; The model parameters of the pre-trained BERT language model are fine-tuned using the multiple words so that the representations of two words in the same word pair are more similar, until the adjustment termination condition is reached, thereby obtaining a trained BERT language model.

7. A text representation device, characterized in that: include: A preprocessing module is configured to preprocess the text to be processed to obtain sentences in the text to be processed and words in the sentences; A model processing module is configured to input the sentences in the text to be processed and the words in the sentences into a trained BERT language model, execute the trained BERT language model, and obtain the first sentence-level representation of the sentences in the text to be processed and the representation of the words in the sentences output by the trained BERT language model, wherein the trained BERT language model is a model obtained by fine-tuning the pre-trained BERT language model based on contrastive learning technology, and the trained BERT language model can align words with the same meaning when representing text; A weight calculation module, configured to calculate the term frequency-inverse document frequency TF-IDF weight of the words in the sentence; A weighting module configured to obtain a second sentence-level representation of the sentence based on the TF-IDF weights of the words in the sentence and the representations of the words in the sentence; The representation merging module is configured to merge the first sentence-level representation and the second sentence-level representation of the sentence to obtain a text representation of the sentence in the text to be processed.

8. A language model training device, characterized in that: include: A word pair acquisition module is configured to acquire a plurality of word pairs, each word pair including at least two words with the same meaning; The fine-tuning module is configured to use the multiple words to fine-tune the model parameters of the pre-trained BERT language model so that the similarity of the representations of two words in the same word pair is greater, until the adjustment termination condition is reached, thereby obtaining a trained BERT language model.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 6.

10. A readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the method described in any one of claims 1 to 6 is implemented.