Vertical field professional vocabulary mining method and device and storage medium

By calculating word segmentation and conditional probability of vertical field text, combined with Bayesian mean and semantic conduction, the accuracy and efficiency of professional vocabulary mining in the existing technology are solved, and efficient noise word filtering and professional vocabulary screening are achieved.

CN119940351APending Publication Date: 2025-05-06FAW VOLKSWAGEN AUTOMOTIVE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311443796.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When the prior art explores professional vocabulary from vertical field texts, there are problems such as inaccurate labeling, conflict of labeling, high labor consumption, and difficulty in filtering noise words.

Method used

By dividing text fragments of vertical domain text, generating original words, and dividing molecular word combinations according to predefined word combination methods, calculating conditional probability and information entropy, generating vocabulary lists, and filtering through Bayesian mean and semantic conduction, finally iteratively determine professional vocabulary.

Benefits of technology

It effectively alleviates the problem of noise word filtering and improves the accuracy and efficiency of professional vocabulary mining. Especially in the relatively rare vertical fields, it avoids the problem of word segmentation boundary errors and excessive word segmentation efforts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940351A_ABST
    Figure CN119940351A_ABST
Patent Text Reader

Abstract

The invention provides a vertical field professional vocabulary mining method and device and a storage medium. The method comprises the steps of obtaining vertical field original words according to a vertical field text and dividing the vertical field original words into sub-word combinations; calculating the conditional probability of the sub-word combination; obtaining a final adjacent word information entropy of the original word in the vertical field; generating a first word list according to the words of which the conditional probabilities and the final adjacent word information entropies are higher than corresponding thresholds; performing word segmentation on the universal text by utilizing the first word list; calculating a first word frequency and a second word frequency of words in the first word list, and calculating a Bayesian mean value of the first word frequency and the second word frequency; generating a second word list according to the words of which the Bayesian mean values are greater than the threshold value; performing word segmentation on the vertical field text by using words in the second word list to obtain vertical field second words, and vectorizing the vertical field second words; and iteratively calculating the similarity between the seed word vector and other second word vectors, and retaining words with the similarity greater than a threshold value, thereby obtaining the vertical field specialized vocabularies. The problem of noise word filtering is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of natural language processing, and more specifically, to a method, device and storage medium for mining professional vocabulary in a vertical field. Background Art

[0002] For any technical direction in the field of natural language processing, such as text classification, sequence labeling, semantic error correction, relationship extraction, etc., if a large number of professional vocabulary in the vertical field can be obtained and combined with the model, the technical indicators can be greatly improved.

[0003] Text resources in vertical fields can be easily obtained through crawlers, but due to the limitations of human resources and personnel expertise, organizing professional vocabulary in the field from vertical domain texts leads to problems such as inaccurate annotations, annotation conflicts, and high manpower consumption.

[0004] In the existing hot word mining system, there is a solution to mine the hot words of the day from user logs by calculating conditional probability and information entropy. However, the words mined from the vertical domain text by calculating conditional probability and information entropy are mixed with many noise words that are not closely related to the vertical domain or even irrelevant.

[0005] In some papers, it is proposed to use multiple open source word segmenters to segment vertical texts respectively, then vote on the results, and finally perform semantic clustering to mine vertical professional vocabulary. However, for some relatively uncommon vertical fields, the effects of several open source word segmenters on the market are limited, and there are often problems with word segmentation boundary errors and excessive word segmentation strength. Summary of the invention

[0006] In order to solve the above-mentioned problems in the prior art, in the first aspect, an embodiment of the present invention provides a method for mining professional vocabulary in a vertical field, the method comprising: dividing a vertical field text into text segments to generate original words in a vertical field; dividing the original words in the vertical field into a combination of multiple sub-words according to a predefined word combination method; calculating the conditional probability of the sub-word combination of the original words in the vertical field; calculating the left neighbor word information entropy and the right neighbor word information entropy of the original words in the vertical field text; taking the smaller value of the left neighbor word information entropy and the right neighbor word information entropy as the final neighbor word information entropy of the original words in the vertical field; determining the words in the original words in the vertical field whose conditional probability is greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold, and generating a first word list; using the first word list to segment the general text to generate original words in the general text; calculating the first word frequency of the words in the first word list in the vertical field text and the first word frequency in the general text according to the original words in the vertical field and the first word frequency in the general text; a second word frequency in a text; calculating the Bayesian mean of words in the first word list according to the first word frequency and the second word frequency; determining words in the first word list whose Bayesian mean is greater than the Bayesian mean threshold to generate a second word list; using the words in the second word list to segment the vertical field text to obtain second words in the vertical field; vectorizing the second words in the vertical field to obtain vectors of the second words in the vertical field; selecting seed words from the second words in the vertical field; calculating the similarity between the vector of the seed word and the vectors of other second words in the vertical field; determining the words in the second words in the vertical field whose similarity is greater than the first similarity threshold as current reserved words; calculating the similarity between the vector of the seed word and the vector of the words in the current reserved words; determining the words in the current reserved words whose similarity is greater than the current similarity threshold as next-level reserved words; iteratively performing the above steps of determining the current reserved words and calculating the similarity between the vector of the seed word and the vector of the current reserved words, and taking the words finally retained as professional vocabulary in the vertical field.

[0007] In some embodiments, calculating the conditional probability of the sub-word combination of the original word in the vertical field includes: when the original word in the vertical field has multiple sub-word combinations according to a predefined word combination method, calculating the conditional probability of each combination separately, and taking the minimum value of the multiple conditional probabilities as the conditional probability of the sub-word combination of the word.

[0008] In some embodiments, calculating the conditional probability of the sub-word combination of the original word in the vertical field includes: dividing the joint probability of multiple sub-words in the sub-word combination by the product of the probability of each sub-word as the conditional probability of the sub-word combination of the original word in the vertical field.

[0009] In some implementations, the current similarity threshold of the current iteration is higher than the similarity threshold of the previous iteration.

[0010] In some implementations, selecting a seed word from the second words in the vertical field includes: displaying the second words in the vertical field; receiving a user's selection of the second words in the vertical field, and using the word selected by the user as the seed word.

[0011] In some implementations, selecting a seed word from the second words in the vertical field includes: inputting the second words in the vertical field into a professional vocabulary recognition model; and determining the seed word according to a recognition result of the professional vocabulary recognition model.

[0012] In some implementations, iteratively performing the step of determining the current reserved words and the next level reserved words includes: performing a predetermined number of iterations.

[0013] In some embodiments, iteratively performing the above step of determining the current reserved words and the next level reserved words includes: when the number of words in the currently generated reserved words whose similarity with the seed words is greater than the final similarity threshold reaches a predefined word number threshold, terminating the iteration and using the currently generated reserved words as the vertical field professional vocabulary.

[0014] In a second aspect, an embodiment of the present invention proposes a vertical field professional vocabulary mining device, the device comprising: a vertical field original word generation module, configured to divide the vertical field text into text segments to generate vertical field original words; a sub-word combination division module, configured to divide the vertical field original words into a combination of multiple sub-words according to a predefined word combination method; a conditional probability calculation module, configured to calculate the conditional probability of the sub-word combination of the vertical field original words; a left and right neighbor word information entropy calculation module, configured to calculate the left neighbor word information entropy and the right neighbor word information entropy of the vertical field original words in the vertical field text; a final neighbor word information entropy determination module, configured to take the left neighbor word information entropy as ... neighbor word information entropy determination module, configured to take the left neighbor word information entropy as the conditional probability of the sub-word combination of the vertical field original words; a left neighbor word information entropy determination module, configured to take the left neighbor word information en The smaller value of the information entropy of the right neighbor word and the information entropy of the right neighbor word is used as the final neighbor word information entropy of the original word in the vertical field; a first word list generation module is configured to determine the words whose conditional probability is greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold in the original words in the vertical field, and generate a first word list; a general text original word generation module is configured to use the first word list to segment the general text and generate general text original words; a word frequency calculation module is configured to calculate the first word frequency of the words in the first word list in the vertical field text and the second word frequency in the general text according to the original words in the vertical field and the original words in the general text; a Bayesian average value calculation module is configured to calculate the first word frequency of the words in the first word list in the vertical field text and the second word frequency in the general text based on the original words in the vertical field and the original words in the general text; According to the first word frequency and the second word frequency, the Bayesian mean of the words in the first word list is calculated; a second word list generation module is configured to determine the words in the first word list whose Bayesian mean is greater than the Bayesian mean threshold, and generate a second word list; a vertical field second word acquisition module is configured to use the words in the second word list to segment the vertical field text and obtain the second words in the vertical field; a vertical field second word vectorization module is configured to vectorize the second words in the vertical field and obtain the second word vector in the vertical field; a seed word selection module is configured to select a seed word from the second words in the vertical field; a first similarity calculation module is configured to calculate the vector of the seed word and other similarity calculation modules. similarity of the second word vector in the vertical field; a current reserved word determination module, configured to determine the words in the second word in the vertical field whose similarity is greater than the first similarity threshold as the current reserved words; a current similarity calculation module, configured to calculate the similarity between the vector of the seed word and the vector of the words in the current reserved words; a next-level reserved word determination module, configured to determine the words in the current reserved words whose similarity is greater than the current similarity threshold as the next-level reserved words; a vertical field professional vocabulary determination module, configured to iteratively perform the above steps of determining the current reserved words and calculating the similarity between the vector of the seed word and the vector of the current reserved words, and use the words finally retained as the vertical field professional vocabulary.

[0015] In a third aspect, an embodiment of the present invention provides a storage medium storing computer-readable instructions, which, when executed by a processor, executes the method according to any of the above embodiments.

[0016] The vertical field professional vocabulary mining method, device and storage medium proposed in the embodiments of the present invention do not use dictionaries or open source word segmenters, but achieve the purpose of professional vocabulary mining through a purely statistical method and solve the problem of noise word filtering.

[0017] The implementation mode of the present invention simply calculates conditional probability and information entropy to mine words from vertical domain text, which are mixed with many noise words that have little or no relevance to the vertical domain. The solution effectively alleviates this problem by calculating the Bayesian mean and using semantic conduction.

[0018] For some relatively obscure vertical fields, the effects of several open source word segmenters (Jieba, Hanlp, LAC) are limited, and often have problems such as word segmentation boundary errors and excessive word segmentation strength. The embodiments of the present invention adopt a statistical method to directly mine information from text, effectively alleviating this problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other objects, features and advantages of the embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which:

[0020] Figure 1 A flowchart of a vertical field professional vocabulary mining method according to an embodiment of the present invention is shown;

[0021] Figure 2 A block diagram of a vertical field professional vocabulary mining device according to an embodiment of the present invention is shown.

[0022] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0023] The principle and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way.

[0024] In one aspect, an embodiment of the present invention provides a method for mining professional vocabulary in a vertical field. Figure 1 The flowchart of the vertical domain professional vocabulary mining method 100 according to the embodiment of the present invention is shown. The method 100 includes steps S101-S118.

[0025] In step S101, the vertical domain text (or vertical domain text) is divided into text segments to generate vertical domain original words. As an example, the vertical domain text can be obtained by crawling. As an example, the length of the word can be limited to a predefined length range, for example, limited to words with a length greater than 1 and less than 7.

[0026] In step S102, the original words in the vertical field are divided into a combination of multiple sub-words according to a predefined word combination method. The predefined word combination method may be, for example: (1) the word "AB" is composed of sub-words "A" and "B"; (2) the word "ABC" is composed of sub-words "A", "BC" or "AB", "C".

[0027] In step S103, the conditional probability of the sub-word combination of the original words in the vertical field is calculated.

[0028] As an embodiment of the present invention, calculating the conditional probability of a sub-word combination of an original word in a vertical field may include: when the original word in the vertical field has multiple sub-word combinations according to a predefined word combination method, calculating the conditional probability of each combination respectively, and taking the minimum value of the multiple conditional probabilities as the conditional probability of the sub-word combination of the word.

[0029] As an embodiment of the present invention, calculating the conditional probability of a sub-word combination of an original word in a vertical field may include: dividing the joint probability of multiple sub-words in the sub-word combination by the product of the probability of each sub-word as the conditional probability of the sub-word combination of the original word in the vertical field.

[0030] As an example only, taking the word "ABC" consisting of sub-words "A", "BC" or "AB", "C", calculate P(A,BC) / (P(A)*P(BC)), P(AB,C) / (P(AB)*P(C)), and retain the smaller one as the score.

[0031] In step S104, the information entropy of the left neighboring words and the information entropy of the right neighboring words of the vertical domain original word in the vertical domain text are calculated.

[0032] In step S105, the smaller value of the left neighboring word information entropy and the right neighboring word information entropy is taken as the final neighboring word information entropy of the original word in the vertical domain.

[0033] In step S106, the words whose conditional probabilities are greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold are determined in the original words of the vertical field, and a first word list is generated.

[0034] Optionally, the conditional probability threshold and the information entropy threshold may be predefined fixed values ​​or dynamically adjusted values. For example, after the conditional probability and the information entropy are obtained, they may be displayed to the user, and the user may determine the threshold based on the displayed conditional probability and the information entropy, or adjust the original threshold.

[0035] In step S107, the first vocabulary is used to segment the general text to generate original words of the general text.

[0036] In step S108, based on the original words in the vertical field and the original words in the general text, the first word frequency Ai of the words in the first vocabulary in the vertical field text and the second word frequency Bi in the general text are calculated.

[0037] In step S109, the Bayesian mean of the words in the first vocabulary is calculated according to the first word frequency and the second word frequency.

[0038] As an example, when the first word frequency is Ai and the second word frequency is Bi, let c = ((A1+B1)+(A2+B2)+…+(An+Bn)) / n, d = (A1 / (A1+B1)+A2 / (A2+B2)+…+An / (An+Bn)) / n, and calculate the Bayesian mean of each word: (Ai+c*d) / (Ai+Bi+c).

[0039] In step S110, the words in the first vocabulary whose Bayesian mean value is greater than the Bayesian mean value threshold are determined to generate a second vocabulary.

[0040] Optionally, the Bayesian mean value threshold may be a predefined fixed value or a dynamically adjusted value. For example, after obtaining the Bayesian mean value, it may be displayed to the user, and the user determines the Bayesian mean value threshold based on the displayed Bayesian mean value, or adjusts the original threshold value.

[0041] The second vocabulary obtained at this time is the vocabulary unique to this vertical domain corpus. The following steps will further screen the words in the second vocabulary.

[0042] In step S111, the vertical field text is segmented using the words in the second vocabulary to obtain the second words in the vertical field.

[0043] In step S112, the second word in the vertical domain is vectorized to obtain a vector of the second word in the vertical domain. For example, the second word in the vertical domain is vectorized using a trained word2vec model.

[0044] The following steps S113-S117 use the obtained word vectors to perform semantic transmission on the vocabulary.

[0045] In step S113, seed words are selected from the second words in the vertical field. The seed words represent words with strong professionalism in the field.

[0046] As an embodiment of the present invention, selecting a seed word from the second words in the vertical field may include: displaying the second words in the vertical field; receiving a user's selection of the second words in the vertical field, and using the word selected by the user as a seed word.

[0047] As an embodiment of the present invention, selecting a seed word from the second words in the vertical field may include: inputting the second words in the vertical field into a professional vocabulary recognition model; and determining the seed word according to a recognition result of the professional vocabulary recognition model.

[0048] In step S114, the similarity between the vector of the seed word and the vectors of the second words in other vertical fields is calculated.

[0049] In step S115, the words in the second words in the vertical field whose similarity is greater than the first similarity threshold are determined as currently reserved words.

[0050] In step S116, the similarity between the vector of the seed word and the vector of the word in the current reserved words is calculated.

[0051] In step S117, the words whose similarity among the currently reserved words is greater than the current similarity threshold are determined as the next level of reserved words, so that the next round of semantic transmission is carried out using the newly reserved words.

[0052] As an embodiment of the present invention, the current similarity threshold of the current iteration is higher than the similarity threshold of the previous iteration. By gradually increasing the threshold, words with higher similarity to the seed word can be gradually screened out.

[0053] The first similarity threshold and the current similarity threshold mentioned above may be predefined fixed values ​​or dynamically adjusted values. For example, the threshold may be set by a user, or the original threshold may be modified.

[0054] In step S118, the above steps of determining the currently reserved words and calculating the similarity between the vector of the seed word and the vector of the currently reserved words are iteratively performed, and the words finally reserved are used as vertical field professional vocabulary.

[0055] As an embodiment of the present invention, iteratively performing the above step of determining the current reserved words and the next level reserved words may include: performing a predetermined number of iterations.

[0056] As an embodiment of the present invention, iteratively performing the above-mentioned step of determining the current reserved words and the next level reserved words may include: when the number of words in the currently generated reserved words whose similarity with the seed words is greater than the final similarity threshold reaches a predefined word number threshold, terminating the iteration and using the currently generated reserved words as vertical field professional vocabulary.

[0057] The vertical field professional vocabulary mining method proposed in the embodiment of the present invention does not use a dictionary or an open source word segmenter, but achieves the purpose of professional vocabulary mining through a purely statistical method and solves the problem of noise word filtering.

[0058] On the other hand, the embodiment of the present invention proposes a vertical field professional vocabulary mining device. Figure 2 , which shows a block diagram of a vertical domain professional vocabulary mining device according to an embodiment of the present invention. The device includes modules 201-218.

[0059] The vertical field original word generation module 201 may be configured to divide the vertical field text into text segments and generate vertical field original words.

[0060] The sub-word combination division module 202 may be configured to divide the original words in the vertical field into a combination of multiple sub-words according to a predefined word combination method.

[0061] The conditional probability calculation module 203 may be configured to calculate the conditional probability of the sub-word combination of the original words in the vertical field.

[0062] The left and right neighbor word information entropy calculation module 204 may be configured to calculate the left neighbor word information entropy and the right neighbor word information entropy of the vertical domain original word in the vertical domain text.

[0063] The final neighbor word information entropy determination module 205 may be configured to take the smaller value of the left neighbor word information entropy and the right neighbor word information entropy as the final neighbor word information entropy of the original word in the vertical domain.

[0064] The first vocabulary generation module 206 can be configured to determine the words in the original words of the vertical field whose conditional probability is greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold, and generate the first vocabulary.

[0065] The general text original word generation module 207 can be configured to segment the general text using the first word list to generate the general text original words.

[0066] The word frequency calculation module 208 may be configured to calculate the first word frequency of the words in the first vocabulary in the vertical field text and the second word frequency in the general text based on the original words in the vertical field and the original words in the general text.

[0067] The Bayesian mean calculation module 209 may be configured to calculate the Bayesian mean of the words in the first vocabulary according to the first word frequency and the second word frequency.

[0068] The second vocabulary generation module 210 may be configured to determine the words in the first vocabulary whose Bayesian mean is greater than a Bayesian mean threshold, and generate a second vocabulary.

[0069] The vertical field second word acquisition module 211 may be configured to segment the vertical field text using the words in the second vocabulary to obtain the vertical field second words.

[0070] The vertical-domain second word vectorization module 212 may be configured to vectorize the vertical-domain second word to obtain the vertical-domain second word vector.

[0071] The seed word selection module 213 may be configured to select a seed word from the second words in the vertical field.

[0072] The first similarity calculation module 214 may be configured to calculate the similarity between the vector of the seed word and the vectors of the second words in other vertical fields.

[0073] The currently reserved word determination module 215 may be configured to determine the words in the second words of the vertical field whose similarity is greater than a first similarity threshold as currently reserved words.

[0074] The current similarity calculation module 216 may be configured to calculate the similarity between the vector of the seed word and the vector of the word in the current reserved words.

[0075] The next-level reserved word determination module 217 may be configured to determine the words in the current reserved words whose similarity is greater than the current similarity threshold as the next-level reserved words.

[0076] The vertical field professional vocabulary determination module 218 can be configured to iteratively perform the above steps of determining the current reserved words and calculating the similarity between the vector of the seed word and the vector of the current reserved words, and use the words finally reserved as the vertical field professional vocabulary.

[0077] It should be noted that the functions implemented by each module in the vertical field professional vocabulary mining device proposed in the embodiment of the present invention correspond one-to-one to the various steps of the vertical field professional vocabulary mining method described above. For its specific implementation methods, examples and beneficial effects, please refer to the above description of the method.

[0078] On the other hand, an embodiment of the present invention provides a storage medium storing computer-readable instructions, which, when executed by a processor, executes the vertical field professional vocabulary mining method described in any of the above embodiments.

[0079] The vertical field professional vocabulary mining method, device and storage medium proposed in the embodiments of the present invention do not use dictionaries or open source word segmenters, but achieve the purpose of professional vocabulary mining through a purely statistical method and solve the problem of noise word filtering.

[0080] The implementation mode of the present invention simply calculates conditional probability and information entropy to mine words from vertical domain text, which are mixed with many noise words that have little or no relevance to the vertical domain. The solution effectively alleviates this problem by calculating the Bayesian mean and using semantic conduction.

[0081] For some relatively obscure vertical fields, the effects of several open source word segmenters (Jieba, Hanlp, LAC) are limited, and often have problems such as word segmentation boundary errors and excessive word segmentation strength. The embodiments of the present invention adopt a statistical method to directly mine information from text, effectively alleviating this problem.

[0082] For the purpose of illustration, the foregoing description of the embodiments of the present invention has been given, which is not exhaustive nor intended to limit the present invention to the disclosed exact form. It will be appreciated by those skilled in the art that various changes can be made without departing from the scope of the present invention, and that the elements therein can be replaced with equivalents. In addition, without departing from the basic scope of the present invention, many modifications can be made so that specific situations or materials are adapted to the teachings of the present invention. Therefore, the present invention is not intended to be limited to the specific embodiments disclosed as the best mode for realizing the present invention, and the present invention will include all embodiments falling within the scope of the appended claims.

Claims

1. A method for mining professional vocabulary in a vertical field, characterized in that: The method comprises: Divide the vertical field text into text segments and generate vertical field original words; According to a predefined word combination method, the original words in the vertical field are divided into a combination of multiple sub-words; Calculating the conditional probability of the sub-word combination of the original words in the vertical field; Calculate the left neighbor word information entropy and the right neighbor word information entropy of the original word in the vertical field in the vertical field text; Taking the smaller value of the left neighbor word information entropy and the right neighbor word information entropy as the final neighbor word information entropy of the original word in the vertical domain; Determine the words in the original words of the vertical field whose conditional probability is greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold, and generate a first word list; Using the first vocabulary to segment the general text to generate original words of the general text; Calculate, based on the original words in the vertical field and the original words in the general text, a first word frequency of the words in the first word list in the vertical field text and a second word frequency in the general text; Calculating the Bayesian mean of the words in the first vocabulary according to the first word frequency and the second word frequency; Determine the words in the first vocabulary whose Bayesian mean is greater than the Bayesian mean threshold, and generate a second vocabulary; Using the words in the second vocabulary, segmenting the vertical field text to obtain the second words in the vertical field; Vectorizing the second word in the vertical field to obtain a second word vector in the vertical field; Selecting a seed word from the second words in the vertical field; Calculate the similarity between the vector of the seed word and the vector of the second word in other vertical fields; Determine the words in the second words of the vertical field whose similarity is greater than the first similarity threshold as currently reserved words; Calculating the similarity between the vector of the seed word and the vector of the word in the currently reserved words; Determine the words whose similarity among the currently reserved words is greater than the current similarity threshold as the next level of reserved words; The above steps of determining the currently reserved words and calculating the similarity between the vector of the seed word and the vector of the currently reserved words are iteratively performed, and the words finally reserved are used as vertical field professional vocabulary.

2. The method according to claim 1, characterized in that Calculating the conditional probability of the sub-word combination of the original words in the vertical field includes: When the original word in the vertical field has multiple sub-word combinations according to the predefined word combination method, the conditional probability of each combination is calculated respectively, and the minimum value of the multiple conditional probabilities is taken as the sub-word combination conditional probability of the word.

3. The method according to claim 1, characterized in that Calculating the conditional probability of the sub-word combination of the original words in the vertical field includes: The joint probability of multiple sub-words in the sub-word combination is divided by the product of the probability of each sub-word, which is taken as the conditional probability of the sub-word combination of the original words in the vertical field.

4. The method according to claim 1, characterized in that: The current similarity threshold of the current iteration is higher than the similarity threshold of the previous iteration.

5. The method according to claim 1, characterized in that Selecting a seed word from the second word in the vertical field includes: Display the second word in the vertical field; Receive a user's selection of a second word in the vertical field, and use the word selected by the user as the seed word.

6. The method according to claim 1, characterized in that Selecting a seed word from the second word in the vertical field includes: Inputting the second word in the vertical field into a professional vocabulary recognition model; The seed word is determined according to the recognition result of the professional vocabulary recognition model.

7. The method according to claim 1, characterized in that The steps of iteratively performing the above steps of determining the current reserved words and the next level reserved words include: A predetermined number of iterations are performed.

8. The method according to claim 1, characterized in that The steps of iteratively performing the above steps of determining the current reserved words and the next level reserved words include: When the number of words in the currently generated reserved words whose similarity with the seed words is greater than the final similarity threshold reaches a predefined word number threshold, the iteration is terminated and the currently generated reserved words are used as the vertical field professional vocabulary.

9. A vertical field professional vocabulary mining device, characterized in that: The device comprises: A vertical field original word generation module is configured to divide the vertical field text into text segments and generate vertical field original words; A sub-word combination division module, configured to divide the original words in the vertical field into a combination of multiple sub-words according to a predefined word combination method; A conditional probability calculation module, configured to calculate the conditional probability of the sub-word combination of the original words in the vertical field; A left and right neighbor word information entropy calculation module is configured to calculate the left neighbor word information entropy and the right neighbor word information entropy of the vertical field original word in the vertical field text; A final neighbor word information entropy determination module is configured to take the smaller value of the left neighbor word information entropy and the right neighbor word information entropy as the final neighbor word information entropy of the original word in the vertical field; A first vocabulary generation module is configured to determine the words whose conditional probability is greater than the conditional probability threshold and whose final neighbor word information entropy is greater than the information entropy threshold in the original words of the vertical field, and generate a first vocabulary; A general text original word generation module, configured to segment the general text using the first word list to generate general text original words; A word frequency calculation module, configured to calculate a first word frequency of words in the first word list in the vertical field text and a second word frequency in the general text according to the vertical field original words and the general text original words; A Bayesian mean calculation module, configured to calculate the Bayesian mean of the words in the first vocabulary according to the first word frequency and the second word frequency; A second vocabulary generation module configured to determine words in the first vocabulary whose Bayesian mean is greater than a Bayesian mean threshold, and generate a second vocabulary; A vertical field second word acquisition module, configured to use the words in the second vocabulary to segment the vertical field text to obtain the vertical field second words; A second vertical domain word vectorization module, configured to vectorize the second vertical domain word to obtain a second vertical domain word vector; A seed word selection module, configured to select a seed word from the second words in the vertical field; A first similarity calculation module, configured to calculate the similarity between the vector of the seed word and the vector of the second word in other vertical fields; a currently reserved word determination module, configured to determine the words in the second words of the vertical field whose similarity is greater than a first similarity threshold as currently reserved words; A current similarity calculation module, configured to calculate the similarity between the vector of the seed word and the vector of the word in the current reserved words; a next-level reserved word determination module, configured to determine the words in the current reserved words whose similarity is greater than the current similarity threshold as next-level reserved words; The vertical field professional vocabulary determination module is configured to iteratively perform the above steps of determining the current reserved words and calculating the similarity between the vector of the seed word and the vector of the current reserved words, and use the finally reserved words as the vertical field professional vocabulary.

10. A storage medium storing computer-readable instructions, wherein when the instructions are executed by a processor, the method according to any one of claims 1 to 8 is executed.