Professional term extraction method and device and medium
By extracting candidate terms from corpus documents in specific fields, calculating their professional scores, and screening out target terms, and building a professional term database, the problem of high recognition error rate of speech recognition system when processing professional terms is solved, and efficient and accurate professional term extraction and recognition are achieved.
Patent Information
- Application Number
- CN202510595557.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing speech recognition systems process professional terms in specific fields, the recognition error rate is high, and they cannot accurately capture professional terms in voice input, affecting the user experience.
By extracting candidate terms from corpus documents in designated fields, identifying their professional impact indicators, including significance indicators, professional confidence and importance indicators, calculating professional scores, and filtering out target terms with scores greater than the threshold, and building a professional term database.
It realizes efficient and accurate extraction of professional terms in specific fields, builds a high-quality professional vocabulary library, and improves the accuracy of the speech recognition system in recognition of professional terms.
Smart Images

Figure CN120106071A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device and medium for extracting professional terms. Background Art
[0002] With the continuous development of speech recognition technology, voice input has become an important way for users to interact with computers. At present, the speech recognition model trained by large-scale general corpus has a high recognition maturity and accuracy for daily language, and is widely used in smart driving, smart home and other aspects of people's lives.
[0003] However, in speech recognition in specific fields, especially for the recognition of professional terms in complex fields such as astronomy and medicine, due to the limited vocabulary coverage of the general corpus and the low accuracy of professional vocabulary in specific fields, it is impossible to accurately capture the professional terms in the speech input, resulting in a high recognition error rate, which seriously affects the user experience.
[0004] It can be seen that how to build a high-quality professional vocabulary library in a specific field is crucial to improving the recognition accuracy of professional terms in various fields by the speech recognition system. Summary of the invention
[0005] In view of this, one aspect of the present application provides a method for extracting professional terms, the method comprising: Extract candidate terms from corpus documents in a specified domain; Determining a professional impact index for each of the candidate terms; the professional impact index includes at least two of a significance index, a professional confidence index, and an importance index; Determining a professional score for each of the candidate terms according to the professional impact index; Target terms whose professional scores are greater than a score threshold are screened out; and a professional term library for the designated field is constructed based on the target terms.
[0006] Optionally, determining the professional impact index of each candidate term includes: Determining the significance index through a latent Dirichlet distribution model; the significance index is used to reflect the degree of relevance between the candidate term and different topics in the corpus document; Determining the professional confidence by a named entity recognition method based on a BERT pre-trained model; the professional confidence is used to reflect the professionalism of the candidate term; The importance index of the candidate term is determined by a term frequency-inverse document frequency algorithm; the importance index is used to reflect the importance of the candidate term in different corpus documents.
[0007] Optionally, the professional impact index includes the significance index, the professional confidence and the importance index; and according to the professional impact index, determining the professional score of each candidate term includes: Assigning a first weight to the significance indicator and a second weight to the importance indicator; Determining a weighted sum of the significance index and the importance index according to the first weight and the second weight; The professional score is determined according to the professional confidence and the weighted sum value; when the professional confidence is greater, the professional score is greater; when the weighted sum value is greater, the professional score is greater.
[0008] Optionally, determining the professional score of each candidate term according to the professional impact index includes: Obtain a pre-built authoritative terminology database about the specified field; Determining an indicator function value of each candidate term according to the authoritative term database; the indicator function value is used to indicate whether the candidate term belongs to the authoritative term database; assigning a third weight to the indicator function value; the sum of the first weight, the second weight and the third weight is 1; The professional score is determined according to the professional confidence, the weighted sum value, the indicator function value and the third weight; when the indicator function value is larger, the professional score is larger.
[0009] Optionally, before constructing the professional terminology library of the specified field based on the target term, the method further includes: Filtering designated terms from the target terms; the designated terms are terms with at least two conflicting professional directions represented by the professional impact indicators; the professional directions include professional and non-professional; Obtaining the professional verification result of the specified term; If the professional verification result is that the designated term is a professional term, the designated term is stored in the authoritative term database, and the third weight is increased; If the professional verification result shows that the designated term is a non-professional term, the designated term is removed.
[0010] Optionally, the professional direction represented by the professional impact indicator is conflict, including: When the significance index is less than the first significance threshold, and the importance index is greater than the first importance threshold, the professional direction is in conflict; wherein the first significance threshold is less than the first importance threshold; Alternatively, when the significance index is greater than the second significance threshold and the importance index is less than the second importance threshold, the professional direction is in conflict; wherein the second significance threshold is greater than the first significance threshold, the second significance threshold is greater than the second importance threshold, and the second importance threshold is less than the first importance threshold.
[0011] Optionally, extracting candidate terms from corpus documents in a specified field includes: Collect the corpus documents of the specified field; and extract text from the corpus documents; Preprocessing the extracted text to obtain a target text; Segmenting the target text to obtain a segmentation set; After filtering the word segmentation set, the candidate terms are obtained.
[0012] Another aspect of the present application provides a device for extracting professional terms, the device comprising: A candidate term extraction module is used to extract candidate terms from corpus documents in a specified field; An impact index determination module, used to determine the professional impact index of each candidate term; the professional impact index includes at least two of a significance index, a professional confidence index and an importance index; A score determination module, used to determine the professional score of each candidate term according to the professional impact index; The terminology library construction module is used to screen out the target terms whose professional scores are greater than the score threshold; and to construct a professional terminology library in the specified field based on the target terms.
[0013] Another aspect of the present application provides a device for extracting professional terms, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps of the method for extracting professional terms when executing the program.
[0014] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method for extracting professional terms are implemented.
[0015] The present application provides a method, device and medium for extracting professional terms, which have the following beneficial effects: extracting candidate terms based on corpus documents in a specified field, and providing an efficient and accurate automated extraction method for building professional terminology libraries in different fields. By integrating multiple professional impact indicators, the candidate terms are scored for professionalism, so as to accurately screen out professional terms that meet the requirements of the field, so as to build a high-quality professional vocabulary library in a specified field, and provide data support for application scenarios such as speech recognition systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a method for extracting professional terms provided in an embodiment of the present application; Figure 2 A schematic diagram of the principle of a method for extracting professional terms provided in an embodiment of the present application; Figure 3 A schematic diagram of a process for extracting professional terms provided in another embodiment of the present application; Figure 4 A schematic diagram of the structure of a device for extracting professional terms provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a device for extracting professional terms provided in another embodiment of the present application.
[0017] The reference numerals are as follows: 50 is a memory, 51 is a processor, 52 is a display screen, 53 is an input / output interface, 54 is a communication interface, 55 is a power supply, 56 is a communication bus, 501 is a computer program, 502 is an operating system, and 503 is data. DETAILED DESCRIPTION
[0018] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0019] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0020] Figure 1 A schematic diagram of a method for extracting professional terms provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes: S10: extract candidate terms from corpus documents in a specified field; Figure 2 A schematic diagram of the principle of a method for extracting professional terms provided in an embodiment of the present application. In a specific embodiment, Figure 2 As shown, a large amount of corpus documents in a specified field are collected, and candidate terms are extracted from the collected corpus documents.
[0021] Among them, candidate terms refer to objects in the corpus document that may be professional terms, and candidate terms can be words, phrases, or word groups, which are not limited in this application. In addition, the designated field can be different fields such as astronomy, geology, medical health, etc., which are not limited in this application.
[0022] In an optional embodiment, the designated field may be the field of Chinese Survey Space Telescope (CSST), and correspondingly, the collected corpus documents may be academic papers, research reports, etc. on CSST. This application does not limit the method of obtaining the corpus documents.
[0023] When extracting candidate terms from the corpus document, it can be based on a statistical method or a method based on preset rules, and the present application does not limit the extraction method. In an optional embodiment, in order to extract candidate terms more accurately and quickly, extraction can be performed by predicting a well-trained large language model.
[0024] S11: Determine the professional impact index of each candidate term; the professional impact index includes at least two of the significance index, professional confidence index and importance index; S12: Determine the professional score of each candidate term based on the professional impact index; It is understandable that the candidate terms initially extracted from the corpus documents may be professional terms or non-professional terms, that is, they may be non-professional terms. In order to further accurately screen out professional terms from the candidate terms, in an optional embodiment, as shown in FIG. Figure 2 As shown, the professional impact index of each candidate term is calculated, and the professionalism score of the candidate term is given according to the professional impact index.
[0025] It should be noted that, in a specific embodiment, in order to improve the accuracy of the professionalism score, multiple influencing factors affecting the professionalism of the candidate terms are comprehensively considered, that is, multiple professionalism impact indicators of the candidate terms are calculated, among which the professionalism impact indicators include at least two of the significance indicators, professional confidence and importance indicators.
[0026] The significance index is used to reflect the relevance of the candidate terms to different topics in the corpus documents, the professional confidence is used to reflect the professionalism of the candidate terms, and the importance index is used to reflect the importance of the candidate terms in different corpus documents.
[0027] S13: Filter out target terms whose professional scores are greater than the score threshold; and build a professional terminology library in the specified field based on the target terms.
[0028] It is understandable that when the professional score is higher, the corresponding candidate term is more professional, that is, the possibility that the candidate term is a professional term is higher. Therefore, further, the target term with a professional score greater than the score threshold is screened out, the target term is used as the professional term in the specified field, and a professional term library is constructed based on the target term.
[0029] In an optional embodiment, after obtaining the professional terminology library, the professional terminology library can be applied to fields such as speech recognition systems, thereby improving the recognition accuracy of professional terms by the speech input system and solving the problem of inaccurate recognition and insufficient coverage of professional terms in specific fields by the speech recognition system.
[0030] Therefore, the method for extracting professional terms provided in the embodiment of the present application extracts candidate terms based on corpus documents in a specified field, providing an efficient and accurate automated extraction method for building professional terminology libraries in different fields. By integrating multiple professional impact indicators, the candidate terms are scored for professionalism, thereby accurately screening out professional terms that meet the requirements of the field, so as to build a high-quality professional vocabulary library in a specified field and provide data support for application scenarios such as speech recognition systems.
[0031] Figure 3 A flow chart of a method for extracting professional terms provided in another embodiment of the present application, in an optional embodiment, as Figure 3 As shown, determine the professional impact indicators of each candidate term, including: S30: Determine the significance index through the latent Dirichlet distribution model; the significance index is used to reflect the degree of relevance between the candidate term and different topics in the corpus document; In order to improve the accuracy of the professional score of each candidate term so as to accurately screen out professional terms in a specified field, in an optional embodiment, three professional influencing factors, namely, significance index, professional confidence and importance index, are comprehensively considered.
[0032] In a specific embodiment, in order to measure the rarity of candidate terms in the entire corpus document, potential topics in the corpus document can be first identified, and the probability distribution of each candidate term under different topics can be calculated, thereby obtaining a value that reflects the degree of relevance of the candidate terms to different topics in the corpus document.
[0033] In an optional embodiment, when calculating the significance index of each candidate term, the Latent Dirichlet Allocation (LDA) model can be used for calculation. In a specific embodiment, an LDA model is pre-constructed, and a training data set with pre-annotated topics and vocabulary probabilities is obtained. The LDA model is trained with the training data set to learn the topic distribution in the training data set. During the training process, the performance of the LDA model under different numbers of topics is evaluated by calculating indicators such as perplexity and topic coherence, so as to optimize the LDA model and obtain an LDA model with high recognition performance in a specified field.
[0034] Furthermore, the candidate terms obtained in the above embodiment are used as inputs of the LDA model of the training number, so as to calculate the significance index of the candidate terms through the LDA model. Specifically, each candidate term is obtained The topic distribution vector , that is, the distribution of each candidate term under different topics. is the number of topic types contained in the corpus documents, Indicates The probability of candidate terms under the first topic, therefore, It means the Candidate terms in The probability distribution vector under different topics.
[0035] Determine the topic distribution vector Afterwards, according to Determine the significance index of the candidate term, where is a significant indicator. The larger the value is, the more likely it is that the candidate term belongs to the core vocabulary of a topic. For example, the corpus document includes three topics: Topic A, Topic B, and Topic C. The significance index of the first candidate term under Topic A is 0.2, the significance index under Topic B is 0.3, and the significance index under Topic C is 0.8. Therefore, the first candidate term is the core vocabulary of Topic C.
[0036] In an optional embodiment, when calculating the significance index of the candidate terms through the LDA model, the Python Gensim library can be introduced, and the LdaModel class method of the Gensim.Models module can be used to build and train the LDA model to obtain the topic distribution.
[0037] It should be noted that the LDA model is an unsupervised learning algorithm that can discover potential topic structures from a large number of corpus documents in a specified field. Therefore, the LDA model can automatically divide a large number of candidate terms into several topics and extract keywords from each topic, thereby helping to analyze the content structure and semantics of the corpus documents.
[0038] S31: Determine professional confidence through a named entity recognition method based on the BERT pre-trained model; professional confidence is used to reflect the degree of professionalism of the candidate term; As an optional embodiment, a named entity recognition method based on a BERT pre-trained model (BERT-base-NER) can be used to calculate the candidate terms to determine the professional confidence of each candidate term. The named entity recognition method (NER) can identify professional entities in the candidate terms, that is, professional subjects. For example, in the field of astronomy, professional entities such as galaxy names, astrophysical concepts, and CSST telescopes can be directly identified.
[0039] In a specific embodiment, the calculated professional confidence is used to reflect the professional degree of the candidate term, that is, the professional confidence can be used to determine whether the candidate term is a professional term in a specified field. In an optional embodiment, in order to reduce the amount of calculation for calculating the professional score based on the professional impact index, after obtaining the professional confidence, the candidate terms whose professional confidence is less than the confidence threshold can be eliminated.
[0040] It can be understood that the smaller the professional confidence, the lower the possibility of characterizing the candidate term as a professional term. Therefore, in order to improve the efficiency of selecting the target term, after obtaining the professional confidence, some candidate terms can be eliminated according to the professional confidence, thereby retaining candidate terms greater than or equal to the confidence threshold.
[0041] For example, professional confidence is , the confidence threshold is 0.5, then in a specific embodiment, retain 0.5 The candidate terms are thus obtained, and the set of professional terms with a greater probability after screening is obtained. .
[0042] It should be noted that, in a specific embodiment, after the candidate terms are processed by the named entity recognition method based on the BERT pre-training model, not only can the professional entities in a large number of candidate terms be identified, but also the professional confidence of each professional entity can be calculated. However, there may be a situation where a complete professional entity is identified as multiple entities. In this case, it is necessary to merge continuous professional entities. For example, a complete professional entity is a high-sensitivity detector, which may be identified as three entities of "high", "sensitivity" and "detector". In this case, the three entities need to be merged into one entity. In an optional embodiment, continuous professional entities can be merged through a large language model, or they can be merged through preset fixed rules. The present application does not limit the merging method.
[0043] In addition, it should be noted that in the specific embodiment where the professional impact index includes a significance index, a professional confidence and an importance index, the amount of calculation for the professional score is reduced. As an optional embodiment, the named entity recognition method based on the BERT pre-training model can be used to identify professional entities in a large number of candidate terms. In fact, this process is a verification of the candidate terms extracted in the above embodiment. When the candidate terms are extracted accurately in the above embodiment, the output professional entity is the same as the candidate term. Otherwise, the candidate terms can be corrected, and the credibility (i.e., confidence) of each candidate term as a professional term can also be calculated. Furthermore, based on the recognition results of the professional entity, multiple entities that are continuous and can be merged are merged into a whole, and entities whose professional confidence is less than the confidence threshold are eliminated.
[0044] S32: Determine the importance index of the candidate term through the term frequency-inverse document frequency algorithm; the importance index is used to reflect the importance of the candidate term in different corpus documents.
[0045] In an optional embodiment, in order to evaluate the importance of each candidate term in the corpus document and identify representative professional terms in the corpus document, the candidate term can be calculated by the term frequency-inverse document frequency (TF-IDF) algorithm to determine the importance index of the candidate term. The importance index can reflect the degree of importance of the candidate term in different corpus documents.
[0046] It is understandable that the TF-IDF algorithm is a weighted technique used in text mining and information retrieval to evaluate the importance of a word in a document set. Among them, TF (term frequency) indicates the frequency of a word in a single document, while IDF (inverse document frequency) can measure the rarity of the word in the entire document set. The TF-IDF algorithm combines term frequency and inverse document frequency to effectively identify professional terms that are highly relevant in the corpus documents of a specified field and are relatively rare in the entire corpus.
[0047] In a specific embodiment, the formula Calculate the word frequency, where is the word frequency, For the candidate terms in the corpus documents The number of times it appears in For corpus documents In addition, we can use the formula Calculate the inverse document frequency, where is the inverse document frequency, is the total number of corpus documents, To include The number of corpus documents containing candidate terms. Further, we can use the formula Calculate the importance index of the candidate term, where For the The importance index of each candidate term.
[0048] It should be noted that, in the specific embodiment, The higher the frequency of a candidate term in its own corpus document, and the lower the frequency of its occurrence in other corpus documents, the importance index is calculated. The larger the value is, the more professional the candidate term is. The document where the candidate term is located is the document in which the candidate term is located. Other corpus documents refer to the documents in all corpus documents except the own corpus document. For ease of understanding, an example will be given below.
[0049] For example, the corpus documents include document F1, document F2, and document F3. When the TF-IDF algorithm is used to calculate, the first candidate term's own corpus document is document F1, and the number of times it appears in other corpus documents refers to the total number of times it appears in document F2 and document F3. Therefore, when the frequency of the first candidate term in document F1 is higher and the frequency of its appearance in document F2 and document F3 is lower, the importance index of the first candidate term calculated is The larger the value, the more likely the first candidate term is a specialized term.
[0050] In an optional embodiment, the importance index of the candidate term is calculated by the TF-IDF algorithm When using the TfidfVectorizer class in Python's Sklearn library, combined with the Jieba word segmentation tool, you can calculate the TF-IDF of the candidate terms to obtain the importance index. .
[0051] In an optional embodiment, the professional impact index includes a significance index, a professional confidence index, and an importance index; according to the professional impact index, the professional score of each candidate term is determined, including: Assigning a first weight to the significance indicator and a second weight to the importance indicator; Determining a weighted sum value of the significance index and the importance index according to the first weight and the second weight; The professional score is determined based on the professional confidence and the weighted sum value; the greater the professional confidence, the greater the professional score; the greater the weighted sum value, the greater the professional score.
[0052] In a specific embodiment, in order to improve the accuracy of professional term extraction, the three dimensions of significance index, professional confidence and importance index are integrated to calculate the professional score of the candidate term. It can be understood that in a specific embodiment, the significance index is used to reflect the degree of relevance between the candidate term and different topics in the corpus documents. When the topic direction in all corpus documents is the same, that is, the topics are close, the significance index is The larger the value, the more likely the candidate term is a professional term. In this case, the frequency of a candidate term appearing in its own corpus documents and in other corpus documents will be relatively high. The smaller it is, the smaller the possibility that the candidate term is a professional term, that is, there is a conflict.
[0053] Therefore, in an optional embodiment, in order to avoid the above technical problems, the significance index can be and importance indicators The corresponding weights are assigned to each of them, so as to adjust the influence of the two on the professional score when calculating the professional score. Specifically, the significance index Assign the first weight , is the importance index Assign the second weight , where the first weight and the second weight are all dynamically adjustable weights that can be adjusted according to different business needs. In an optional embodiment, the first weight can be initially set is 0.6, the second weight is 0.3.
[0054] Further, according to the first weight and the second weight , for the significance index and importance indicators Perform a weighted sum calculation, and determine the professional score based on the professional confidence and the weighted sum value. Specifically, the calculation formula for the professional score is formula (1): (1) in, For the The professionalism score of each candidate term.
[0055] According to formula (1), when professional confidence The larger the professional score, The larger the weighted sum is, the The larger the professional score, The bigger.
[0056] Therefore, the method for extracting professional terms provided in the embodiment of the present application takes into account the multi-dimensional influencing factors that affect the extraction of professional terms, and on this basis comprehensively considers the influence of the significance index and the importance index on each other, and assigns different weights to them accordingly, so that the professional score can more accurately represent the professional degree of the candidate terms and improve the accuracy of the professional terms.
[0057] In order to further improve the accuracy of professional term extraction, in an optional embodiment, the professional score of each candidate term is determined according to the professional impact index, including: Get pre-built authoritative terminology lexicon for a specific field; According to the authoritative terminology lexicon, determining the indicator function value of each candidate term; the indicator function value is used to indicate whether the candidate term belongs to the authoritative terminology lexicon; Assigning a third weight to the indicator function value; the sum of the first weight, the second weight and the third weight is 1; The professional score is determined based on the professional confidence, weighted sum value, indicator function value and the third weight; the larger the indicator function value, the larger the professional score.
[0058] On the basis of the above embodiment, in order to further improve the professional score The calculation accuracy of the candidate terms is to ensure the final professional evaluation accuracy of the candidate terms. In the embodiment of the present application, Figure 2As shown, an authoritative term database is introduced. In a specific embodiment, the authoritative term database is a pre-constructed authoritative term database related to the field. In an optional embodiment, the authoritative terms can be manually selected and constructed from authoritative documents.
[0059] For example, in the field of astronomy, astronomy terms can be obtained from the National Astronomical Science Data Center through experts in the field of astronomy, thereby ensuring that each word in the authoritative terminology database is a designated professional term.
[0060] Furthermore, the indicator function value of each candidate term is determined based on the authoritative term database. , where the indicator function value Used to indicate whether the candidate term belongs to the authoritative term library. Specifically, when the candidate term When it belongs to the authoritative terminology database, is equal to 1. If it does not belong to the authoritative terminology database, then, is equal to 0.
[0061] Based on the above embodiment, combined with the significance index , professional confidence , Importance Index and indicator function value , rate professionalism Calculate, specifically, the indicator function value Assign third weight , similarly, the third weight It can be dynamically adjusted according to actual business needs to adjust the indicator function value Rating of professionalism The contribution degree of the calculation result. It should be noted that, in an optional embodiment, the first weight , second weight and the third weight The sum of is equal to 1.
[0062] Further, the professional score is calculated according to formula (2): : (2) in, For the The indicator function value of the candidate terms.
[0063] According to formula (2), in a specific embodiment, when the indicator function value The larger the professional score, The bigger.
[0064] Therefore, the method for extracting professional terms provided in the embodiment of the present application further improves the accuracy of professional term extraction by introducing an indicator function value determined based on an authoritative term library when calculating the professional score.
[0065] In an optional embodiment, before constructing a professional terminology library of a specified field based on the target term, the method further includes: From the target terms, designated terms are selected; designated terms are terms with conflicting professional directions represented by at least two professional impact indicators; professional directions include professional and non-professional; Get the professional verification results of the specified term; If the result of the professional verification is that the designated term is a professional term, the designated term is stored in the authoritative term database and the third weight is increased; If the result of the professional verification is that the specified term is a non-professional term, the specified term will be eliminated.
[0066] It is understandable that determining the significance indicator When the topic direction in all corpus documents is the same, that is, when the topics are close, the significance index The larger the value, the more likely the candidate term is a professional term. In this case, the frequency of a candidate term appearing in its own corpus documents and in other corpus documents will be relatively high. The smaller it is, the less likely it is that the candidate term is a professional term, thus resulting in a conflict.
[0067] That is, it is possible that a candidate term has It can be determined as a professional term, that is, the professional direction is professional and positive feedback. It can be determined as a non-professional term, that is, the professional direction is non-professional, which is negative feedback. That is, for the same candidate term, different professional impact indicators may have opposite results in representing whether it is a professional term. Therefore, in order to avoid the above conflict, in an optional embodiment, such as Figure 2 As shown, before constructing a professional terminology library in a specified field based on the target terminology, a conflict resolution mechanism is set. Specifically, the designated subject with conflicts is first screened out from all target terms, and the professionalism of the designated subject is further verified. In an optional embodiment, the verification can be performed manually or by a large language model, which is not limited in this application.
[0068] After verification, if the professional verification result indicates that the specified term is a professional term, the specified term can be stored as a professional term in the authoritative term database and the third weight can be increased. , in order to improve the indicator function value Contribution to the professionalism score.
[0069] It is understandable that in a specific embodiment, when a new professional term appears, the above-mentioned conflict may occur. At this time, the conflict resolution mechanism provided in the embodiment of the present application can be used to avoid the professional term being mistakenly treated as a non-professional term when the conflict occurs, and the specified term is stored in the authoritative term library, which can improve the accuracy of subsequent professional term extraction.
[0070] Of course, if the professional verification result indicates that the designated term is a non-professional term, the designated term can be eliminated. It should be noted that, in an optional embodiment, when judging whether there is a designated term with professional conflict in the target term, that is, the professional direction represented by the professional impact indicator is a conflict, it includes: When the significance index is less than the first significance threshold and the importance index is greater than the first importance threshold, the professional direction of the target term is determined to be conflicting; wherein the first significance threshold is less than the first importance threshold. And the importance index When determining the professional orientation of the target term, it is determined that the professional orientation of the target term is in conflict.
[0071] Alternatively, in another optional embodiment, when the significance index is greater than the second significance threshold and the importance index is less than the second importance threshold, the professional direction of the target term is determined to be conflicting; wherein the second significance threshold is greater than the first significance threshold, the second significance threshold is greater than the second importance threshold, and the second importance threshold is less than the first importance threshold. And the importance index When determining the professional orientation of the target term, it is determined that the professional orientation of the target term is in conflict.
[0072] Thus, further, in a specific embodiment, the target term after the conflict resolution mechanism is processed, that is, the target term after the conflict resolution mechanism and the professional score can be obtained. The set of terms greater than the scoring threshold ,in, is the scoring threshold. It should be noted that the scoring threshold It can be set according to actual business needs. For example, if the accuracy requirement of the constructed professional terminology library is higher, a larger scoring threshold can be set. On the contrary, if the accuracy requirement of the constructed professional terminology library is lower, a smaller scoring threshold can be set. .
[0073] Depend on Figure 2It can be seen that in an optional embodiment, after extracting candidate terms from the corpus document, accurate professional terms (i.e., target terms) can be obtained by performing a series of calculations and screening on the candidate terms. Specifically, professional scores are calculated based on significance indicators, professional confidence, importance indicators, and indicator function values, and conflicts are resolved based on a conflict resolution mechanism, so as to obtain target terms with professional scores greater than a score threshold, so as to build a high-quality professional terminology library in a specified field.
[0074] As an optional embodiment, extracting candidate terms from corpus documents in a specified field includes: Collect corpus documents in a specified field; and extract text from the corpus documents; Preprocess the extracted text to obtain the target text; Segment the target text to obtain a segmentation set; After filtering the word set, candidate terms are obtained.
[0075] In a specific embodiment, in order to ensure that the professional database includes a large number of professional terms in a specified field, in a specific embodiment, a large number of corpus documents in the specified field can be collected, for example, a large number of CSST-related technical documents can be collected, such as academic papers, research reports, etc. The formats of the collected corpus documents include but are not limited to PDF and HTML.
[0076] Further, the text included in the collected corpus document is extracted. In an optional embodiment, the text can be extracted by a large language model, or by a document parsing Python library, such as the PymuPDF library for parsing PDF documents and the BeautifulSoup library for parsing HTML documents. This application does not limit the text extraction method.
[0077] In order to ensure the quality of the text, the extracted text is preprocessed. Specifically, the preprocessing includes but is not limited to removing stop words, punctuation marks, and unifying vocabulary forms (for example, conversion between simplified and traditional Chinese). This improves the quality of the target text and facilitates the subsequent rapid and accurate extraction of candidate terms.
[0078] In an optional embodiment, the target text may be segmented by a large language model, or by a segmentation tool suitable for Chinese (e.g., Jieba segmentation tool) to generate a segmentation set. Further, in order to ensure the quality of the segmentation set, the segmentation set may be filtered. Specifically, the filtering process includes but is not limited to deleting pure digital words, removing Chinese and English words whose text length is less than a length threshold (e.g., the length threshold is 2), etc., and a candidate term set that may be a professional term is obtained after filtering.
[0079] In the above embodiments, the method for extracting professional terms is described in detail. The present application also provides an embodiment corresponding to a device for extracting professional terms.
[0080] Figure 4 A schematic diagram of a structure of a professional term extraction device provided in an embodiment of the present application, such as Figure 4 As shown, the device comprises: A candidate term extraction module 40 is used to extract candidate terms from corpus documents in a specified field; The impact index determination module 41 is used to determine the professional impact index of each candidate term; the professional impact index includes at least two of the significance index, professional confidence index and importance index; A score determination module 42, used to determine the professional score of each candidate term according to the professional impact index; The terminology library construction module 43 is used to screen out target terms with professional scores greater than a score threshold; and to construct a professional terminology library in a specified field based on the target terms.
[0081] In addition, the professional term extraction device provided in the embodiment of the present application also includes: A significance index determination module is used to determine the significance index through a latent Dirichlet distribution model; the significance index is used to reflect the degree of relevance between the candidate term and different topics in the corpus document; A professional confidence determination module is used to determine professional confidence through a named entity recognition method based on a BERT pre-trained model; professional confidence is used to reflect the professionalism of a candidate term; The importance index determination module is used to determine the importance index of the candidate term through the word frequency-inverse document frequency algorithm; the importance index is used to reflect the importance of the candidate term in different corpus documents.
[0082] A weight assignment module, used to assign a first weight to the significance index and a second weight to the importance index; A weighted summation module, used to determine a weighted sum value of the significance index and the importance index according to the first weight and the second weight; The first determination module is used to determine the professional score according to the professional confidence and the weighted sum value; when the professional confidence is greater, the professional score is greater; when the weighted sum value is greater, the professional score is greater.
[0083] An authoritative terminology word library acquisition module is used to acquire a pre-built authoritative terminology word library for a specified field; An indicator function value determination module is used to determine the indicator function value of each candidate term according to the authoritative term vocabulary; the indicator function value is used to indicate whether the candidate term belongs to the authoritative term vocabulary; The weight allocation module is further used to allocate a third weight to the indicator function value; the sum of the first weight, the second weight and the third weight is 1; The second determination module is used to determine the professional score according to the professional confidence, the weighted sum value, the indicator function value and the third weight; when the indicator function value is larger, the professional score is larger.
[0084] A screening module is used to screen out designated terms from target terms; designated terms are terms with conflicting professional directions represented by at least two professional impact indicators; professional directions include professional and non-professional; A verification result acquisition module is used to obtain the professional verification result of the specified term; The processing module is used to store the designated term in an authoritative term database and increase the third weight when the professional verification result shows that the designated term is a professional term; and to remove the designated term when the professional verification result shows that the designated term is a non-professional term.
[0085] The text extraction module is used to collect corpus documents in a specified field and extract text from the corpus documents; A preprocessing module is used to preprocess the extracted text to obtain the target text; The word segmentation module is used to segment the target text and obtain a word segmentation set; The filtering module is used to filter the word set to obtain candidate terms.
[0086] Figure 5 A schematic diagram of a structure of a professional term extraction device provided in another embodiment of the present application is shown in FIG. Figure 5 As shown, the device for extracting professional terms includes: a memory 50 for storing a computer program; The processor 51 is used to implement the steps of the method for extracting professional terms mentioned in the above embodiment when executing the computer program.
[0087] The device for extracting professional terms provided in this embodiment may include but is not limited to a laptop computer or a desktop computer.
[0088] Among them, the processor 51 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 51 can be implemented in at least one hardware form of a digital signal processor (Digital Signal Processor, referred to as DSP), a field programmable gate array (Field-Programmable Gate Array, referred to as FPGA), and a programmable logic array (Programmable Logic Array, referred to as PLA). The processor 51 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (Central Processing Unit, referred to as CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 51 may be integrated with a graphics processing unit (Graphics Processing Unit, referred to as GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 51 may also include an artificial intelligence (Artificial Intelligence, referred to as AI) processor, which is used to process computing operations related to machine learning.
[0089] The memory 50 may include one or more computer-readable storage media, which may be non-transitory. The memory 50 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 50 is at least used to store the following computer program 501, wherein, after the computer program is loaded and executed by the processor 51, it can implement the relevant steps of the method for extracting professional terms disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 50 may also include an operating system 502 and data 503, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 502 may include Windows, Unix, Linux, etc. Data 503 may include but is not limited to relevant data involved in the method for extracting professional terms, etc.
[0090] In some embodiments, the device for extracting professional terms may further include a display screen 52 , an input / output interface 53 , a communication interface 54 , a power supply 55 , and a communication bus 56 .
[0091] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation of the extraction device of professional terminology, and may include more or fewer components than shown in the figure.
[0092] The device for extracting professional terms provided in the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the method for extracting professional terms in the above embodiment.
[0093] It should be noted that, although the operations are depicted in a specific order in the accompanying drawings, this should not be understood as requiring these operations to be performed in the specific order shown or to be performed sequentially, or requiring all illustrated operations to be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
Claims
1. A method for extracting professional terms, characterized in that: The method comprises: Extract candidate terms from corpus documents in a specified domain; Determining a professional impact index for each of the candidate terms; the professional impact index includes at least two of a significance index, a professional confidence index, and an importance index; Determining a professional score for each of the candidate terms according to the professional impact index; Target terms whose professional scores are greater than a score threshold are screened out; and a professional term library for the designated field is constructed based on the target terms.
2. The method for extracting professional terms according to claim 1, characterized in that: The step of determining the professional impact index of each candidate term includes: Determining the significance index through a latent Dirichlet distribution model; the significance index is used to reflect the degree of relevance between the candidate term and different topics in the corpus document; Determining the professional confidence by a named entity recognition method based on a BERT pre-trained model; the professional confidence is used to reflect the professionalism of the candidate term; The importance index of the candidate term is determined by a term frequency-inverse document frequency algorithm; the importance index is used to reflect the importance of the candidate term in different corpus documents.
3. The method for extracting professional terms according to claim 1, characterized in that: The professional impact index includes the significance index, the professional confidence index and the importance index; Determine the professional score of each candidate term according to the professional impact index, including: Assigning a first weight to the significance indicator and a second weight to the importance indicator; Determining a weighted sum of the significance index and the importance index according to the first weight and the second weight; The professional score is determined according to the professional confidence and the weighted sum value; when the professional confidence is greater, the professional score is greater; when the weighted sum value is greater, the professional score is greater.
4. The method for extracting professional terms according to claim 3, characterized in that: Determine the professional score of each candidate term according to the professional impact index, including: Obtain a pre-built authoritative terminology database about the specified field; Determining an indicator function value of each candidate term according to the authoritative term database; the indicator function value is used to indicate whether the candidate term belongs to the authoritative term database; assigning a third weight to the indicator function value; the sum of the first weight, the second weight and the third weight is 1; The professional score is determined according to the professional confidence, the weighted sum value, the indicator function value and the third weight; when the indicator function value is larger, the professional score is larger.
5. The method for extracting professional terms according to claim 4, characterized in that: Before constructing the professional terminology library of the specified field based on the target term, the method further includes: Filtering designated terms from the target terms; the designated terms are terms with at least two conflicting professional directions represented by the professional impact indicators; the professional directions include professional and non-professional; Obtaining the professional verification result of the specified term; If the professional verification result is that the designated term is a professional term, the designated term is stored in the authoritative term database, and the third weight is increased; If the professional verification result shows that the designated term is a non-professional term, the designated term is removed.
6. The method for extracting professional terms according to claim 5, characterized in that: The professional direction represented by the professional impact index is conflict, including: When the significance index is less than the first significance threshold, and the importance index is greater than the first importance threshold, the professional direction is in conflict; wherein the first significance threshold is less than the first importance threshold; Alternatively, when the significance index is greater than the second significance threshold and the importance index is less than the second importance threshold, the professional direction is in conflict; wherein the second significance threshold is greater than the first significance threshold, the second significance threshold is greater than the second importance threshold, and the second importance threshold is less than the first importance threshold.
7. The method for extracting professional terms according to claim 1, characterized in that: The step of extracting candidate terms from corpus documents in a specified field includes: Collect the corpus documents of the specified field; and extract text from the corpus documents; Preprocessing the extracted text to obtain a target text; Segmenting the target text to obtain a segmentation set; After filtering the word segmentation set, the candidate terms are obtained.
8. A device for extracting professional terms, characterized in that: The device comprises: A candidate term extraction module is used to extract candidate terms from corpus documents in a specified field; An impact index determination module, used to determine the professional impact index of each candidate term; the professional impact index includes at least two of a significance index, a professional confidence index and an importance index; A score determination module, used to determine the professional score of each candidate term according to the professional impact index; The terminology library construction module is used to screen out the target terms whose professional scores are greater than the score threshold; and to construct a professional terminology library in the specified field based on the target terms.
9. A device for extracting professional terms, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method for extracting professional terms according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for extracting professional terms described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Construction method of keyword library in power industry
CN116484846A
Method and system for extracting Chinese terminologies by fusing rules and statistical characteristics
CN116702786A
Professional lexicon construction method and device, medium and program product
CN117290500A