A method for mining professional terminology based on contrastive learning

Through comparative learning methods, the problem of low efficiency in extracting keywords in professional literature in the existing technology is solved, and more efficient professional term recognition and relationship recognition are achieved, which improves the accuracy of literature search and understanding.

CN115794998BActive Publication Date: 2025-05-16ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211632497.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-05-16
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently explore keywords in professional literature, resulting in inefficient searching and understanding professional literature.

Method used

Using a method based on contrast learning, new phrases are mined through information entropy algorithms and mutual information algorithms, combined with word segmentation thesaurus and professional knowledge bases to filter words, used sentence vector learning model and classification model to perform domain classification, and construct a tree-type term relationship tree.

Benefits of technology

It improves the accuracy and efficiency of professional terms, enhances the classification and relationship recognition capabilities of the model in downstream tasks, and improves the effect of literature retrieval and understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115794998B_ABST
    Figure CN115794998B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for mining professional terminology based on contrastive learning, which belongs to the technical field of machine learning, including forming a term list based on a corpus of a professional field; performing field classification on the terms in the term list; and constructing a tree-type term relationship tree of the professional field based on the term list. The present invention uses a BERT pre-training model to train word vectors, and uses contrastive learning to train sentence vectors. Through pre-training, the model's ability to perform classification and relationship recognition in downstream tasks can be greatly enhanced, so that the model can achieve the maximum effect. At the same time, the professional term word vector and the overall text vector are considered, and features are extracted after mutual fusion, which has better predictability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a method for mining professional domain terms based on contrastive learning. Background Art

[0002] Professional terms are regarded as the description and summary of important subject information in professional field literature. They are the smallest unit of summarizing text and are also considered the smallest abstract of professional field literature. They can be effectively used to understand, organize and retrieve the content of papers. For example, in academic publications, the keywords at the beginning of the article mention the words that best represent its content. Readers can use professional terms to decide whether to read various information systems, and can also use professional terms to easily complete tasks such as paper classification and quick retrieval.

[0003] The data of professional literature is huge, which makes it difficult to access literature related to the topic when searching for any topic on the Internet. If there are some words to represent the main features of the document, such as the content and theme, it will be easier to retrieve relevant documents. According to the development history of keyword extraction technology, it can be subdivided into the keyword extraction stage and the keyword generation stage. Among them, the keyword extraction stage refers to the selection of words that can express the theme of the professional field from the original text as professional terms, and the professional terms must appear in the document; the professional term generation stage refers to the selection of words that best fit the theme of the document from the vocabulary or original text as the professional terms of the text, regardless of whether the professional terms appear in the document.

[0004] Professional terminology can help people quickly understand the main idea of ​​professional literature and grasp the main thread of the paper. Professional terminology extraction technology is an important method to extract a number of practical and representative words or phrases in professional literature in order to quickly obtain the professional theme. It has important applications in literature retrieval, automatic summarization, text clustering and text classification.

[0005] Compared with other documents, scientific and technological documents have some unique characteristics. Among them, the abstract part is the most critical and core part of scientific and technological documents. The abstract part needs to contain all the essential technical features that reflect the novelty, creativity and practicality of the paper to illustrate the scope of the paper research. Although the abstract part does not have explicit structured information, the abstract has certain writing rules. By mining the abstract of the paper, useful information can be provided for professional document keyword extraction. Accordingly, the present invention performs keyword extraction and processing based on professional field documents to summarize its related fields and the technologies used. Summary of the invention

[0006] The purpose of the present invention is to provide a professional field term mining method based on contrastive learning to accurately mine professional field terms.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A method for mining professional domain terms based on contrastive learning, the method comprising:

[0009] Step 1: Form a term list based on the corpus of the professional field;

[0010] Step 1-1, using information entropy algorithm and mutual information algorithm to mine new phrases in the corpus;

[0011] Step 1-2: Add the mined new phrases to the word segmentation lexicon, use the word segmentation lexicon to segment all sentences in the corpus, extract the keywords in each segmented word, remove the duplicates of the extracted keywords and merge them to form a keyword list;

[0012] Step 1-3, filter out non-professional words in the keyword list, and perform term matching to obtain a term list;

[0013] Step 2: Classify the terms in the term list into different fields;

[0014] Step 2-1, take the entry corresponding to each term in the term list;

[0015] Step 2-2, divide the entry into sentences, and input each sentence into a sentence vector learning model based on contrastive learning, and output the sentence vector corresponding to each sentence in the entry;

[0016] Step 2-3: concatenate the word vector trained by the BERT model corresponding to the term and the sentence vector corresponding to each clause in the term's entry, and input them into the classification model to obtain the domain classification result of the term;

[0017] Step 3: construct a tree-type term relationship tree in the professional field based on the term list;

[0018] Step 3-1, clustering the terms in the term list according to the field classification results;

[0019] Step 3-2, concatenate the term with the sentence vector where the term is located, and then concatenate it with the paragraph vector where the term is located to obtain the feature vector of the term; the sentence vector where the term is located is trained by the contrastive learning model based on the abstract part of the professional literature where the term is located, and the paragraph vector where the term is located is trained by the BERT model based on the abstract part of the professional literature where the term is located;

[0020] Step 3-3, input the feature vectors of terms belonging to the same category into the relation recognition model based on contrastive learning in pairs, and obtain the relationship between the two terms in each pair;

[0021] Step 3-4: Based on the relationship between two terms, construct a tree-type term relationship tree for each professional field.

[0022] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.

[0023] Preferably, the method of mining new phrases in the corpus using an information entropy algorithm and a mutual information algorithm comprises:

[0024] Step 1-1-1, segment all sentences in the corpus into single characters, each character as a candidate segment, to form a candidate segment set CS;

[0025] Step 1-1-2, taking two adjacent candidate segments in the candidate segment set CS as a candidate phrase cp, and calculating the left information entropy LE(cp) and the right information entropy RE(cp) of the candidate phrase cp;

[0026] Step 1-1-3, for the candidate phrase cp, calculating the internal mutual information MI(cp) of the candidate phrase cp;

[0027] Step 1-1-4, set the score threshold δ and the probability threshold λ, and for each candidate phrase cp, calculate the score S(cp) as S(cp)=MI(cp)+min(LE(cp),RE(cp)). If the probability of the candidate phrase cp in the corpus is greater than λ and the score S(cp) is greater than δ, then confirm that the candidate phrase cp is a new phrase;

[0028] Step 1-1-5: Combine the two candidate segments confirmed as new phrases as a new candidate segment, update the candidate segment set CS, and jump to step 1-1-2 to start the next round of iteration until no new phrases appear.

[0029] Preferably, the calculation formulas of the left information entropy LE(cp) and the right information entropy RE(cp) are as follows:

[0030]

[0031]

[0032] Where L(cp) refers to the set of all left candidate segments of the candidate phrase cp in the corpus, R(cp) refers to the set of all right candidate segments of the candidate phrase cp in the corpus, and p(x) refers to the probability of occurrence of the candidate segment x in the corpus;

[0033] The calculation formula of the internal mutual information MI(cp) is as follows:

[0034]

[0035] Where x and y are two candidate segments that constitute the candidate phrase cp, p(y) refers to the probability of candidate segment y appearing in the corpus, and p(x,y) refers to the probability of candidate phrase cp appearing in the corpus.

[0036] Preferably, the filtering of non-professional words in the keyword list and performing term matching to obtain a term list includes:

[0037] Step 1-3-1, obtain the part of speech of each keyword in the keyword list based on the part-of-speech tagging algorithm, and filter out keywords that are not nouns or verbs;

[0038] Step 1-3-2, match each keyword in the keyword list filtered in step 1-3-1 to an entry in the Internet knowledge base, discard keywords that do not match an entry, and finally form a term list, in which each term in the term list corresponds to an entry.

[0039] Preferably, the classification model includes a long short-term memory network, a fully connected layer and a softmax layer connected in sequence, and the vector input to the long short-term memory network passes through the fully connected layer and then the softmax layer outputs the domain classification result of the term.

[0040] Preferably, the concatenation formula of the feature vectors of the terms is:

[0041] H″t=W0[concat(H0,H′ t )]+b0

[0042] H′ t =concat(S t ,H t )

[0043] In the formula, S t is the tth term in the term list, H t For term S t The clause vector where it is located, concat(,) is the vector concatenation function, H′ t For term S t The clause vector H where the term is located t The new vector after concatenation, H0 is the term S t The text vector where it is located, b0 is the bias vector, W0 is the weight, H″t is the term S t The corresponding feature vector.

[0044] Preferably, the training process of the sentence vector learning model and the classification model is as follows:

[0045] Taking a term list formed based on the training data as a training term list;

[0046] For each term r in the training term list i , construct classification samples s i =(doc(r i ),l i ), where doc(r i ) is the term r i Matching terms, l i For the term r i Type;

[0047] The training of the sentence vector model includes:

[0048] Input layer: doc(r i ) Term sentence, randomly select two clauses S k1 and S k2 , correspond the two clauses to a 300×1 small matrix SD k1 and SD k2 , and mark whether the two clauses are adjacent clauses, record the [CLS] flag, and convert the small matrix SD k Input to the input layer of the sentence vector model:

[0049] Coding layer: for each small matrix SD k Use Encoder to encode it, connect it to the dropout layer to prevent overfitting, and output the 300×1 small matrix SV corresponding to each sentence k As sentence vectors for clauses;

[0050] Interaction layer: Determine whether a group of sentences is a positive sample or a negative sample based on the [CLS] tag, so that the sentence vector has a high similarity with its positive sample and a low similarity with its negative sample. Map the sentence vector distribution output by the encoding layer to the MLP layer for representation, and train the sentence vector model parameters at the MLP layer according to the loss formula of contrastive learning;

[0051] The training of the classification model includes:

[0052] After the sentence vector model training is completed, the sentence vector model is used to output the sentence vector of each sentence;

[0053] The word vector trained by the BERT model corresponding to the term and the sentence vector output by the sentence vector model for each sentence are concatenated and input into the classification model to predict the domain classification result of the term;

[0054] Based on the domain classification results of the predicted terms and the actual types of the terms i The loss function will be calculated to train the classification model parameters.

[0055] The present invention provides a method for mining professional terminology based on contrastive learning. The method uses the BERT pre-training model to train word vectors and uses contrastive learning to train sentence vectors. Through pre-training, the model's ability to perform classification and relationship recognition in downstream tasks can be greatly enhanced, so that the model can achieve the maximum effect. At the same time, the professional terminology word vector and the overall text vector are considered, and the features are extracted after mutual fusion, which has better predictability. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flowchart of the professional field terminology mining method based on contrastive learning of the present invention.

[0057] Figure 2 A schematic diagram of a contrastive learning model provided by the present invention;

[0058] Figure 3 Flowchart for term relationship identification of the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0061] Reference Figure 1 This embodiment provides a method for mining domain professional terms based on contrastive learning, comprising the following steps:

[0062] (1) Extraction and screening of terms: The abstract corpus of the paper is preprocessed, and after processing, the mutual information algorithm and information entropy algorithm are used to extract phrases and keywords, and then filtered using the professional knowledge base. Since the abstract part of professional documents is representative content, in order to reduce the processing pressure, this embodiment takes the abstract part of professional documents as the content in the corpus.

[0063] (1-1) Use information entropy algorithm and mutual information algorithm to mine new phrases in the corpus.

[0064] (1-1-1) Segment all sentences in corpus D into single characters, and take each character as a candidate segment to form a candidate segment set CS.

[0065] (1-1-2) Take two adjacent candidate segments in the set CS as a candidate phrase cp, calculate the left information entropy LE(cp) of the candidate phrase cp according to formula (1), and calculate the right information entropy RE(cp) of the candidate phrase cp according to formula (2). Wherein, L(cp) refers to the set of all left candidate segments of the candidate phrase cp in the corpus D, R(cp) refers to the set of all right candidate segments of the candidate phrase cp in the corpus D, and p(x) refers to the probability of occurrence of the candidate segment x in the corpus D. In this embodiment, the probability of occurrence can be expressed by the ratio of the number of target candidate segments to the total number of candidate segments.

[0066]

[0067]

[0068] (1-1-3) For the candidate phrase cp, the internal mutual information MI(cp) of the candidate phrase cp is calculated according to formula (3). In which, x and y are two candidate segments constituting the candidate phrase cp, p(y) refers to the probability of candidate segment y appearing in the corpus, and p(x,y) refers to the probability of candidate segments x and y appearing together in the corpus D.

[0069]

[0070] (1-1-4) Set the score threshold δ and probability threshold λ, and for each candidate phrase cp, calculate the score S(cp) according to formula (4). If the occurrence probability p(x, y) of the candidate phrase cp is greater than λ and the score S(cp) is greater than δ, the candidate phrase cp is considered to be a new phrase.

[0071] S(cp)=MI(cp)+min(LE(cp),RE(cp))(4)

[0072] (1-1-5) The segment combination confirmed as the new phrase is used as the new candidate segment, the candidate segment set CS is updated, and the process jumps to step (1-1-2) to start the next round of iteration until no new phrase can be found.

[0073] (1-2) Keyword list establishment: First, add the new phrases extracted in step (1-1) to the word segmentation lexicon, and perform word segmentation on all sentences in the corpus D. Then, use the TF-IDF algorithm to calculate the keywords in each word segmentation sentence, remove duplicate keywords and merge them to obtain the keyword list KS.

[0074] (1-3) Terminology list establishment: Filter out non-professional words in the keyword list KS. The specific steps are as follows:

[0075] (1-3-1) Based on the part-of-speech tagging algorithm, the part of speech of each keyword in the keyword list KS is obtained, and keywords that are not nouns or verbs are filtered out.

[0076] (1-3-2) Each keyword in the keyword list KS filtered in step (1-3-1) is matched with entries in an Internet knowledge base (such as Baidu Encyclopedia), and keywords that cannot be matched with entries are filtered out, and finally a term list RS is formed, in which each term in the term list RS corresponds to an entry.

[0077] (2) Pre-training and domain classification of terms: First, semantic expansion is performed on candidate terms to expand the corpus available for training. Then, pre-training is performed using a contrastive learning method. Finally, the obtained word vectors are imported into a classification model to obtain the domain classification of the terms.

[0078] (2-1) Take the entry corresponding to each term in the term list.

[0079] (2-2) Divide the entry into sentences, and input each sentence into the sentence vector learning model based on contrastive learning, and output the sentence vector corresponding to each sentence in the entry.

[0080] (2-3) The word vector trained by the BERT model corresponding to the term and the sentence vector corresponding to each clause in the term's corresponding entry are concatenated and input into the classification model to obtain the domain classification result of the term.

[0081] The sentence vector learning model in this embodiment is a contrastive learning model, and the classification model includes a long short-term memory network, a fully connected layer, and a softmax layer connected in sequence. The vector input to the long short-term memory network passes through the fully connected layer and then the softmax layer outputs the domain classification result of the term.

[0082] The training process of the sentence vector learning model and the classification model is as follows:

[0083] (a) Forming a training term list RS' based on the open source dataset.

[0084] (b) Term classification sample construction and expansion: For each term r in the training term list RS' i , construct classification samples s i =(doc(r i ),l i ). Among them, doc(r i ) is the term r i The entry description text in the Internet knowledge base, l i For r i The type of i ranges from 1 to I, where I is the total number of terms in the training term list RS'.

[0085] (c) Sentence embedding learning based on contrastive learning: It specifically includes the following modules:

[0086] Input layer: doc(r i ) get K clauses, randomly select two clauses J k1 and J k2 , correspond the two clauses to a 300×1 small matrix SD k1 and SD k2 , and mark whether the two clauses are adjacent clauses, record the [CLS] flag, and convert the small matrix SD k Input to the input layer of the sentence vector model. Wherein, k1, k2, and k are all between 1 (inclusive) and K (inclusive).

[0087] Coding layer: for each small matrix SD k Use Encoder to encode it, connect it to the dropout layer to prevent overfitting, and output the 300×1 small matrix SV corresponding to each sentence k As sentence vectors for clauses.

[0088] Interaction layer: According to the [CLS] tag, determine whether a group of sentences is a positive sample or a negative sample, so that the similarity between the sentence vector and its positive sample is high, and the similarity between the sentence vector and its negative sample is small. Figure 2 This is an example of contrastive learning. The sentence vector distribution output by the encoding layer is mapped to the MLP layer for representation, and the sentence vector model parameters are trained at the MLP layer according to the loss formula (loss function formula) of contrastive learning. After training, the sentence vector corresponding to the sentence is output.

[0089] (d) Term classification: First, the word vectors trained by the BERT model and the doc(r i ) The sentence vectors of each sentence in the paragraph are concatenated and input into the long short-term memory network (LSTM), and then into a fully connected layer. The output of the fully connected layer is then input into a softmax layer. Finally, the softmax layer is based on doc(r i ) Output term r i According to the predicted domain classification results and the actual type of the term i The loss function will be calculated to train the classification model parameters.

[0090] (3) Relationship extraction and identification of terms: Based on contrastive learning, the term text is processed and analyzed to extract the feature vector of the text, which is then fused with the word vector of the term to finally identify the relationship between the terms and construct a tree-like term relationship tree in the professional field. Figure 3 shown.

[0091] (3-1) Define the relationship between terms: The relationship between two terms is defined as ["subordinate", "equal", "irrelevant"], and the terms in the term list are clustered according to the domain classification results.

[0092] (3-2) Obtain model input: Substitute the term S t (For example, for the vector {A i ,…,A j}) and the clause vector H where the term is located t Concatenate to get a new vector H' that integrates term features and sentence features t , and then the new vector H' t The matrix calculation is performed with the text vector H0 obtained by BERT model training. The specific calculation formula is as follows, W0 is the weight, and b0 is the bias vector. t The final feature vector H of the fused text segment and term t . Vector H' t It has both the global features of the text and the features of the two terms, so when obtaining relational data, it can better grasp the global features and correctly propose the relationship between the two terms.

[0093] H″t=W0[concat(H0,H′ t )]+b0 (5)

[0094] H′ t =concat(S t ,H t ) (6)

[0095] In order to distinguish the representation of terms in the term list RS' during training from the representation of terms in the term list RS during actual application, this embodiment uses S t represents the tth term in the term list RS. And the clause vector H where the term is located t The contrastive learning model is trained based on the abstract part of the professional literature where the term is located, and the paragraph vector H0 where the term is located is trained by the BERT model based on the specific paragraph of the abstract part of the professional literature where the term is located.

[0096] (3-3) Relationship identification model construction: Use the contrastive learning model SimCSE to identify the relationship between terms. For two terms S1 and S2 belonging to the same category, the corresponding H”1 and H”2 are input into the contrastive learning model.

[0097] (3-4) Based on the relationship between two terms, a tree-type term relationship tree is constructed for each professional field.

[0098] When training the relationship recognition model, the feature vectors of the two terms are input into the encoder and then distributed to the MLP layer, and the model parameters are trained according to the loss function of contrastive learning.

[0099] After the MLP layer of the comparison learning model, the softmax layer is connected to map to the specific relationship matrix to obtain the relationship between the two terms. The loss function used by the softmax layer is the cross entropy loss function, which performs very well when dealing with classification problems.

[0100] In the present embodiment, in relation identification, based on the contrastive learning model, the sentence features (sentence vectors, paragraph vectors) and entity features (terms themselves) are fused, and the obtained fused feature vector can have the features of both the sentence and the two entities. Such fusion processing enhances the model's ability to process feature vectors and improves the generalization of the model, so that it performs better when performing relation extraction.

[0101] It should be noted that, in this embodiment, contrastive learning models and BERT models are applied in many places, and the models used in different stages are independent models. For example, the contrastive learning models in the sentence vector learning model and the relationship recognition model are two independent models, which are trained and used separately, but the two contrastive learning models can be models of the same structure or models of different structures.

[0102] This embodiment uses the semantic expansion method to improve the problem of poor model performance caused by insufficient training corpus. By introducing contrastive learning to solve the text pre-training problem, it will have better performance in the downstream keyword extraction task.

[0103] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0104] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A method for mining professional domain terms based on contrastive learning, characterized in that: The professional domain term mining method based on contrastive learning includes: Step 1: Form a term list based on the corpus of the professional field; Step 1-1, using information entropy algorithm and mutual information algorithm to mine new phrases in the corpus; Step 1-2: Add the mined new phrases to the word segmentation lexicon, use the word segmentation lexicon to segment all sentences in the corpus, extract the keywords in each segmented word, remove the duplicates of the extracted keywords and merge them to form a keyword list; Step 1-3, filter out non-professional words in the keyword list, and perform term matching to obtain a term list; Step 2: Classify the terms in the term list into different fields; Step 2-1, take the entry corresponding to each term in the term list; Step 2-2, divide the entry into sentences, and input each sentence into a sentence vector learning model based on contrastive learning, and output the sentence vector corresponding to each sentence in the entry; Step 2-3: concatenate the word vector trained by the BERT model corresponding to the term and the sentence vector corresponding to each clause in the term's entry, and input them into the classification model to obtain the domain classification result of the term; Step 3: construct a tree-type term relationship tree in the professional field based on the term list; Step 3-1, clustering the terms in the term list according to the field classification results; Step 3-2, concatenate the term with the sentence vector where the term is located, and then concatenate it with the paragraph vector where the term is located to obtain the feature vector of the term; the sentence vector where the term is located is trained by the contrastive learning model based on the abstract part of the professional literature where the term is located, and the paragraph vector where the term is located is trained by the BERT model based on the abstract part of the professional literature where the term is located; Step 3-3, input the feature vectors of terms belonging to the same category into the relation recognition model based on contrastive learning in pairs, and obtain the relationship between the two terms in each pair; Step 3-4: Based on the relationship between two terms, construct a tree-type term relationship tree for each professional field.

2. The method for mining professional domain terms based on contrastive learning according to claim 1, characterized in that: The method of mining new phrases in the corpus using the information entropy algorithm and the mutual information algorithm includes: Step 1-1-1, segment all sentences in the corpus into single characters, each character as a candidate segment, to form a candidate segment set CS; Step 1-1-2, taking two adjacent candidate segments in the candidate segment set CS as a candidate phrase cp, and calculating the left information entropy LE(cp) and the right information entropy RE(cp) of the candidate phrase cp; Step 1-1-3, for the candidate phrase cp, calculating the internal mutual information MI(cp) of the candidate phrase cp; Step 1-1-4, set the score threshold δ and the probability threshold λ, and for each candidate phrase cp, calculate the score S(cp) as S(cp)=MI(cp)+min(LE(cp),RE(cp)). If the probability of the candidate phrase cp in the corpus is greater than λ and the score S(cp) is greater than δ, then confirm that the candidate phrase cp is a new phrase; Step 1-1-5: Combine the two candidate segments confirmed as new phrases as a new candidate segment, update the candidate segment set CS, and jump to step 1-1-2 to start the next round of iteration until no new phrases appear.

3. The method for mining professional domain terms based on contrastive learning according to claim 2, characterized in that: The calculation formulas of the left information entropy LE(cp) and the right information entropy RE(cp) are as follows: Where L(cp) refers to the set of all left candidate segments of the candidate phrase cp in the corpus, R(cp) refers to the set of all right candidate segments of the candidate phrase cp in the corpus, and p(x) refers to the probability of occurrence of the candidate segment x in the corpus; The calculation formula of the internal mutual information MI(cp) is as follows: Where x and y are two candidate segments that constitute the candidate phrase cp, p(y) refers to the probability of candidate segment y appearing in the corpus, and p(x,y) refers to the probability of candidate phrase cp appearing in the corpus.

4. The method for mining professional domain terms based on contrastive learning according to claim 1, characterized in that: The filtering of non-professional words in the keyword list and performing term matching to obtain a term list includes: Step 1-3-1, obtain the part of speech of each keyword in the keyword list based on the part-of-speech tagging algorithm, and filter out keywords that are not nouns or verbs; Step 1-3-2, match each keyword in the keyword list filtered in step 1-3-1 to an entry in the Internet knowledge base, discard keywords that do not match an entry, and finally form a term list, in which each term in the term list corresponds to an entry.

5. The method for mining professional domain terms based on contrastive learning according to claim 1, characterized in that: The classification model includes a long short-term memory network, a fully connected layer and a softmax layer connected in sequence. The vector input to the long short-term memory network passes through the fully connected layer and then the softmax layer outputs the domain classification result of the term.

6. The method for mining professional domain terms based on contrastive learning according to claim 1, characterized in that: The concatenation formula of the feature vectors of the terms is: H″t=W0[concat(H0,H′ t )]+b0 H′ t =concat(S t ,H t ) In the formula, S t is the tth term in the term list, H t For term S t The clause vector where it is located, concat(,) is the vector concatenation function, H′ t For term S t The clause vector H where the term is located t The new vector after concatenation, H0 is the term S t The text vector where it is located, b0 is the bias vector, W0 is the weight, H″t is the term S t The corresponding feature vector.

7. The method for mining professional domain terms based on contrastive learning according to claim 1, characterized in that: The training process of the sentence vector learning model and the classification model is as follows: Taking a term list formed based on the training data as a training term list; For each term r in the training term list i , construct classification samples s i =(doc(r i ),l i ), where doc(r i ) is the term r i Matching terms, l i For the term r i Type; The training of the sentence vector model includes: Input layer: doc(r i ) Term sentence, randomly select two clauses S k1 and S k2 , correspond the two clauses to a 300×1 small matrix SD k1 and SD k2 , and mark whether the two clauses are adjacent clauses, record the [CLS] flag, and convert the small matrix SD k Input to the input layer of the sentence vector model: Coding layer: for each small matrix SD k Use Encoder to encode it, connect it to the dropout layer to prevent overfitting, and output a small matrix SV of 300×1 corresponding to each sentence k As sentence vectors for clauses; Interaction layer: Determine whether a group of sentences is a positive sample or a negative sample based on the [CLS] tag, so that the sentence vector has a high similarity with its positive sample and a low similarity with its negative sample. Map the sentence vector distribution output by the encoding layer to the MLP layer for representation, and train the sentence vector model parameters at the MLP layer according to the loss formula of contrastive learning; The training of the classification model includes: After the sentence vector model training is completed, the sentence vector model is used to output the sentence vector of each sentence; The word vector trained by the BERT model corresponding to the term and the sentence vector output by the sentence vector model for each sentence are concatenated and input into the classification model to predict the domain classification result of the term; Based on the domain classification results of the predicted terms and the actual types of the terms i The loss function will be calculated to train the classification model parameters.

Citation Information

Patent Citations

  • Marketing text-based data processing method and device

    CN114357157A

  • Knowledge injection method of pre-training language model and corresponding interaction system

    CN114936287A