Two-stage new word discovery method

CN119990116APending Publication Date: 2025-05-13CHINA MOBILE QUANTONG SYST INTEGRATION CO LTD +5
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411695723.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing new word discovery methods lack generality and high maintenance costs, relying on expert knowledge or a large amount of pre-screened data.

Method used

Using a two-stage new word discovery method, firstly, by extracting character segments from the corpus text and calculating their adjacency entropy, the word set with high information is selected; then, these words are input into the new word discovery model, and the target words are determined based on the probability characteristics, and combined with deep learning algorithms to improve accuracy.

Benefits of technology

Effectively filter out irrelevant character segments, reduce the amount of subsequent processing data, improve processing efficiency, ensure that the extracted words have high semantic correlation, improve the accuracy and efficiency of new word discovery, and solve the problems of poor generalization and high maintenance costs of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990116A_ABST
    Figure CN119990116A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage new word discovery method and device, equipment and a storage medium, and belongs to the field of natural language process.The method comprises the steps that a character field set containing a plurality of to-be-verified target character fields is extracted from a corpus text, and the adjacent entropy of each target character field in the character field set is determined; determining a word set of the corpus text according to the adjacent entropy of each target character field in the character field set; and inputting the words in the word set into a new word discovery model, so as to determine a first target word from the word set according to the probability feature of each word determined by the new word discovery model. Based on the method provided by the embodiment of the invention, the problems confronted based on rules and statistical methods in the prior art are overcome, the problems of relying on high-quality expert knowledge and pre-screening a large amount of data are solved, the efficiency and accuracy of new word discovery are improved, and the problems of poor universality and high maintenance cost in the existing new word discovery process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of natural language processing, and specifically relates to a two-stage new word discovery method, device, equipment and storage medium. Background Art

[0002] New words and buzzwords usually reflect the priorities and directions of the work of various departments and new trends in social and economic development. By understanding these new words, the public can better understand the priorities and future development directions of various departments, while also improving the transparency and public participation of the work of various departments and promoting communication between departments and the public.

[0003] Existing new word discovery methods usually use rule-based methods to establish a rule base, pattern base or professional word base based on the word formation features or appearance features of new words, and extract text fragments with specific pattern features from the corpus as new words according to pre-set rules. Statistical methods use statistical methods to extract and screen possible new words based on the frequency, context relevance and other features of the text.

[0004] However, existing new word discovery methods either rely on expert knowledge or on a large amount of pre-screened data, and therefore suffer from the problems of lack of versatility and high maintenance costs. Summary of the invention

[0005] The present application aims to provide a two-stage new word discovery method, device, equipment and storage medium, which at least solves the problem of lack of versatility and high maintenance cost in the new word screening process.

[0006] In a first aspect, the present application embodiment discloses a two-stage new word discovery method, comprising: Extracting a character segment set including a plurality of target character segments to be verified from the corpus text, and determining the adjacency entropy of each target character segment in the character segment set; the adjacency entropy is used to characterize the amount of information carried by the target character segment; Determine a word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set; the word set contains at least one word to be verified; The words in the word set are respectively input into the new word discovery model to determine the first target word from the word set according to the probability features of each of the words determined by the new word discovery model; the probability features are used to characterize the meaning type and / or position type of each character in the word.

[0007] In a second aspect, the present application also discloses a two-stage new word discovery device, including: An extraction module is used to extract a character segment set including a plurality of target character segments to be verified from the corpus text, and determine the adjacency entropy of each target character segment in the character segment set; the adjacency entropy is used to characterize the amount of information carried by the target character segment; An adjacency entropy module, used to determine a word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set; the word set contains at least one word to be verified; A selection module is used to input the words in the word set into a new word discovery model respectively, so as to determine a first target word from the word set according to the probability features of each of the words determined by the new word discovery model; the probability features are used to characterize the meaning type and / or position type of each character in the word.

[0008] In a third aspect, an embodiment of the present application further discloses an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0009] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] In summary, in the embodiment of the present application, by introducing the adjacency entropy indicator, insignificant character segments in the corpus are effectively filtered out, thereby reducing the amount of data for subsequent processing, improving processing efficiency, and then determining the word set of the corpus text, more accurately defining the boundaries of the character segments, and ensuring that the extracted words have high semantic relevance. Then the words in the word set are respectively input into the new word discovery model, and the first target word is determined from the word set through the probability characteristics of each word determined by the new word discovery model. In combination with the deep learning algorithm, the meaning type and / or position type of each character in the word can be represented according to the probability characteristics, realizing the transition from statistical features to semantic features, and utilizing the powerful representation ability of deep learning to improve the accuracy of new word discovery. Therefore, the method based on the embodiment of the present application overcomes the problems faced by rule-based and statistical methods in the prior art, solves the problems of relying on high-quality expert knowledge and pre-screening of large amounts of data, realizes the improvement of new word discovery efficiency and accuracy, and solves the problems of poor versatility and high maintenance cost in the existing new word discovery process. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the attached picture: Figure 1 It is a flowchart of the steps of a two-stage new word discovery method provided in an embodiment of the present application; Figure 2is a flowchart of another two-stage new word discovery method provided by an embodiment of the present application; Figure 3 It is a process of determining a word set of a corpus text under the method provided in the embodiment of the present application; Figure 4 It is a new word discovery model and its operation process provided by the embodiment of the present application; Figure 5 This is the architecture and operation process of a two-stage new word discovery system provided by the embodiment of the present application; Figure 6 This is a data capture process provided by an embodiment of the present application; Figure 7 is a block diagram of a two-stage new word discovery device provided by an embodiment of the present application; Figure 8 is a block diagram of an electronic device according to an embodiment of the present application; Fig. 9 It is a block diagram of an electronic device of another embodiment provided by the embodiments of the present application. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0013] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0014] The method of the present application is based on utilizing the basic concept of information theory - information entropy. Information entropy is a concept used to represent the uncertainty of a random variable: the greater the entropy of a random variable, the greater its uncertainty and the more information it carries.

[0015] Figure 1 This is a two-stage new word discovery method provided by this embodiment.

[0016] The method may include the following steps: Step 101: extract a character segment set including multiple target character segments to be verified from a corpus text, and determine the adjacency entropy of each target character segment in the character segment set.

[0017] Among them, the adjacency entropy is used to characterize the amount of information carried by the target character segment.

[0018] In some embodiments of the present application, multiple target character segments to be verified are extracted from the corpus text, and the adjacency entropy of these target character segments is calculated. The adjacency entropy is used to measure the amount of information carried by each target character segment to characterize the semantic relevance of the character segment. In the specific implementation process, the corpus text is first segmented by a word segmentation tool, and then the segmentation results are statistically analyzed to extract character segments with higher frequencies as candidate target character segments. Next, the left and right adjacency entropies of each target character segment are calculated to determine its information content, thereby screening out target character segments with higher semantic relevance.

[0019] For example, in the processed text, you can use a word segmentation tool (such as Jieba) to segment the text using the n-gram model. This algorithm is very effective in language models and is mainly used to calculate the probability of each element in a sequence, so as to predict the next element or generate text, and finally obtain the word segmentation result. Then, count the character segments with higher frequency in all the word segmentation results, such as "data processing" and "public information". Next, calculate the left and right adjacent entropy of each character segment, such as the left and right adjacent entropy of "public information", and select the character segments with higher adjacent entropy as the target character segments based on the calculation results. These target character segments will serve as the input of the subsequent new word discovery model to provide data support for new word discovery.

[0020] Step 102, determining the word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set.

[0021] The word set contains at least one word to be verified.

[0022] In some embodiments of the present application, the word set of the corpus text is determined based on the adjacency entropy of each target character segment in the character segment set. In the specific implementation process, by calculating the adjacency entropy of each target character segment, character segments with higher entropy values ​​are screened out. These character segments usually carry more information and have higher semantic relevance. These character segments are combined together to construct the word set of the corpus text, ensuring that the words in the word set are more representative and important.

[0023] For example, the character segments extracted from the processed text include "governance system" and "information disclosure". By calculating the adjacency entropy of these character segments, it is found that the left adjacency entropy and right adjacency entropy of "governance system" are both high, indicating that it has high information content and semantic relevance. Therefore, "governance system" is included in the word set. Similarly, after screening and calculation, other character segments with higher entropy values ​​are included in the word set to form a semantically rich and information-rich word set, providing a basis for new word discovery.

[0024] Step 103 , input the words in the word set into the new word discovery model respectively, so as to determine the first target word from the word set according to the probability feature of each word determined by the new word discovery model.

[0025] The probability feature is used to characterize the meaning type and / or position type of each character in the word.

[0026] In some embodiments of the present application, the words extracted from the word set are input into the new word discovery model, and the probability features of each word are calculated by the model to select the first target word that best meets the conditions. The probability features are used to characterize the meaning type and position type of each character, helping the model to more accurately identify new words. In the specific implementation process, the pre-trained model can be combined with the screening algorithm (such as Bi-LSTM and CRF) to perform deep semantic analysis and feature extraction on the words, so as to determine the new words that best meet the semantic features.

[0027] For example, in the new word discovery system, the extracted word set includes words such as "digital economy", "governance capacity", and "openness and transparency". These words are input into the new word discovery model, and the model combines the powerful feature extraction capabilities of the pre-trained model to analyze the semantic features and character position relationships of each word. Through calculation, the model determines that "digital economy" has a high semantic relevance and position matching, so it is selected as the first target word. In this way, the system can accurately identify new words with important significance and provide support for analysis in related fields.

[0028] In summary, in the embodiment of the present application, by introducing the adjacency entropy indicator, insignificant character segments in the corpus are effectively filtered out, thereby reducing the amount of data for subsequent processing, improving processing efficiency, and then determining the word set of the corpus text, more accurately defining the boundaries of the character segments, and ensuring that the extracted words have high semantic relevance. Then the words in the word set are respectively input into the new word discovery model, and the first target word is determined from the word set through the probability characteristics of each word determined by the new word discovery model. In combination with the deep learning algorithm, the meaning type and / or position type of each character in the word can be represented according to the probability characteristics, realizing the transition from statistical features to semantic features, and utilizing the powerful representation ability of deep learning to improve the accuracy of new word discovery. Therefore, the method based on the embodiment of the present application overcomes the problems faced by rule-based and statistical methods in the prior art, solves the problems of relying on high-quality expert knowledge and pre-screening of large amounts of data, realizes the improvement of new word discovery efficiency and accuracy, and solves the problems of poor versatility and high maintenance cost in the existing new word discovery process.

[0029] Figure 2 This is another two-stage new word discovery method provided in an embodiment of the present application.

[0030] The method may include the following steps: Step 201, obtaining web page data with text from a preset target website.

[0031] In some embodiments of the present application, web page data with text is obtained from a preset target website. In the specific implementation process, a timer can be set to regularly visit the target website and automatically capture web page data using web crawling technology. The web crawling program will identify the page structure of the target website, locate and extract the text content in the web page, and ensure that the acquired text data includes the latest public information. This process provides a data basis for subsequent text preprocessing and new word discovery.

[0032] For example, suppose the target website is the official website of a department. The system sets a timer to automatically access the official website at 1 a.m. every day. The web crawler written using the Scrapy framework will automatically access the homepage and each subpage of the website, locate the area containing public files, and crawl the text data in the page. For example, the crawler will access the "Public Documents" column, extract the contents of all newly released public documents, and save this data to the local database to provide a data source for subsequent text preprocessing.

[0033] Step 202, delete the format characters in the web page data to obtain the corpus text.

[0034] In some embodiments of the present application, format characters are deleted from the captured web page data to obtain clean corpus text. In the specific implementation process, regular expressions or other text processing tools can be used to screen and filter out irrelevant content such as Hypertext Markup Language (HTML), special characters, blank characters (such as extra line breaks, tabs, spaces, etc.), Chinese and English garbled characters, etc. in the web page data. This process is intended to clean up the noise data in the text, making the subsequent new word discovery more accurate.

[0035] For example, an announcement captured from a department's official website contains HTML tags and some formatting characters. During the cleaning process, first use regular expressions to remove all HTML tags, such as 、 , Then, all redundant whitespace characters, such as redundant line breaks, tabs, and spaces, are further screened and removed. After cleaning, the text content originally containing formatting characters is processed into clean plain text, providing a basis for subsequent corpus analysis and new word discovery.

[0036] Step 203: extract a character segment set including multiple target character segments to be verified from the corpus text, and determine the adjacency entropy of each target character segment in the character segment set.

[0037] Among them, the adjacency entropy is used to characterize the amount of information carried by the target character segment.

[0038] The method shown in this step has been explained in step 101 and will not be repeated here.

[0039] In some embodiments of the present application, in order to balance the richness of stop words and the accuracy and effectiveness of new word discovery, the contextual relationship of words is introduced. The stop word definition is:

[0040] Where Q represents the set of all terms in the corpus, a is a preset constant, and the formula means that if the number of different terms w' adjacent to the right (left) side of a word w exceeds a, then the word w is a stopword on the left. l (Stopword on the right r ), the candidate phrases with the term as the left (right) boundary will be filtered out from the candidate phrases for new word discovery.

[0041] Considering that we can use the internal cohesion of words (the degree of close connection between words is used to measure the possibility of words forming a word. The greater the internal cohesion of words, the tighter the combination of Chinese characters, and the greater the possibility that they form a word), we use the threshold to make decisions. When the internal cohesion of words is greater than the threshold, they are considered to be able to form a word. Therefore, the concept of mutual information is used to screen character segments: mutual information is usually used to measure the degree of mutual dependence between two signals, and can be used to measure the degree of internal cohesion of bigrams. The larger the mutual information, the greater the internal cohesion of the bigram, and the greater the possibility that the bigram will become a new word or part of a new word. Pointwise Mutual Information (PMI) is defined as:

[0042] Among them, p(x) and p(y) represent the probability of x and y appearing alone in the corpus; p(x,y) represents the probability of x and y appearing together in the corpus. If PMI(x,y)>0, it means that x and y are related, and the larger the value, the stronger the correlation; if PMI(x,y)<0, it means that x and y are not related; if PMI(x,y)=0, it means that x and y are independent of each other.

[0043] As a statistic to measure the internal cohesion of a phrase, mutual information has a good effect on bigrams. In this invention, in order to improve the ability of the new word discovery system, candidate phrases include not only bigrams, but also trigrams and quadrugrams. For multigrams, we introduce an improved mutual information, multivariate mutual information (MMI), which is defined as follows:

[0044] Where n is the number of tuples in the multi-tuple (n>2); N is the total number of words in the corpus; w j ,…,w k is a tuple substring, which means a continuous substring from the jth term to the kth term in the tuple; fre(w1,w2,…,w n ) represents the frequency of occurrence of a word string in the corpus; p(w j ,…,w k ) represents the probability of a multi-tuple substring appearing in the corpus; avg(w1,w2,…,w n ) represents the average probability of all substring combinations of the multi-tuple.

[0045] Optionally, in order to extract a character segment set including multiple target character segments to be verified from the corpus text, step 203 includes the following sub-steps: Sub-step 2031, extracting multiple character segments of the corpus text according to a preset character truncation length, and determining the mutual information value of each character segment.

[0046] Among them, the mutual information value of a character segment is positively correlated with the probability that the single words in the character segment just constitute a word.

[0047] In some embodiments of the present application, multiple character segments are extracted from the corpus text according to a preset character interception length, and the mutual information value of each character segment is determined. The mutual information value is used to measure the closeness between the characters within the character segment, that is, the probability of whether the character segment may constitute a word. In the specific implementation process, the corpus text is intercepted character by character through the sliding window technology, and a character segment of length n is intercepted each time. Then, the mutual information value of each character segment is calculated to determine the correlation between the characters within it.

[0048] For example, extract character segments from the corpus text of the field, and the preset character extraction length is 3 characters. Use the sliding window to extract the character segments "digital economy", "character economy", "economic zone", etc. in turn. Then, calculate the mutual information value of each character segment. Assuming that the mutual information value of "economic zone" is high, it means that the correlation between its internal characters is strong, and the probability of forming a word is high. In this way, high-probability character segments that may form new words in the corpus text can be screened out, providing basic data for subsequent new word discovery.

[0049] Sub-step 2032, determining the character segments whose mutual information values ​​are greater than a preset mutual information threshold as target character segments to obtain a character segment set.

[0050] In some embodiments of the present application, a character segment with a higher mutual information value is determined as a target character segment by utilizing a preset mutual information threshold to obtain a character segment set. In the specific implementation process, the mutual information value of each character segment can be calculated first, and then these mutual information values ​​can be compared with the preset threshold. The character segments with mutual information values ​​greater than the preset threshold are retained as target character segments to form the final character segment set. This process ensures that the selected character segments have strong internal correlation and high word formation probability, laying the foundation for subsequent new word discovery.

[0051] For example, among the multiple extracted segments, the mutual information value of "data information" is 0.8, the mutual information value of "informatization" is 0.85, and the mutual information value of "data analysis" is 0.65. Assuming that the preset mutual information threshold is 0.7, "data information" and "informatization" are determined as target segments and included in the segment set, while "data analysis" is screened out. In this way, by comparing the mutual information values ​​and screening the segments above the threshold, the system finally determines the target segments in the corpus text, providing a reliable data basis for new word discovery.

[0052] Optionally, after obtaining the character segment set through sub-steps 2031 and 2032, the method of the embodiment of the present application may further update the character segment set through the following additional sub-steps: Sub-step 2033, in the character segment set, the target character segments that are identical to the reference words in the preset reference word library are deleted to update the character segment set.

[0053] In some embodiments of the present application, after obtaining the character segment set, the target character segments in the character segment set that are identical to the reference words in the preset reference word library are deleted to update the character segment set. In the specific implementation process, all the character segments in the character segment set are compared with the words in the reference word library through the preset reference word library. If the character segment is found to be identical to the reference word, the character segment is deleted from the character segment set. This process is intended to exclude those words that already exist in the reference word library, ensuring that the content in the character segment set is more unique and has the potential to discover new words.

[0054] For example, after sub-steps 2031 and 2032, the character segment set includes character segments such as "data information", "informatization", and "big data". Assume that the preset reference word library contains the two words "informatization" and "big data". Through comparison, the system recognizes that "informatization" and "big data" in the character segment set are the same as the words in the reference word library, so these two character segments are deleted from the character segment set. The updated character segment set only retains "data information", providing more accurate data support for subsequent new word discovery.

[0055] Step 204, determining the word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set.

[0056] The word set contains at least one word to be verified.

[0057] The method shown in this step has been explained in step 102 and will not be repeated here.

[0058] Optionally, step 204 includes the following sub-steps: Sub-step 2041, determining candidate words from the target character segments according to the relationship between the adjacency entropy of each target character segment in the character segment set and a preset adjacency entropy threshold.

[0059] In some embodiments of the present application, candidate words are determined from the target character segments based on the relationship between the adjacency entropy of each target character segment in the character segment set and a preset adjacency entropy threshold. In the specific implementation process, the adjacency entropy of each target character segment is first calculated, and then these adjacency entropy values ​​are compared with the preset threshold. Character segments with adjacency entropy values ​​greater than the preset threshold are retained as candidate words. These candidate words have a higher amount of information and semantic relevance, and further screen out words with greater new word potential.

[0060] For example, the character segments extracted from the processed corpus text include "digital economy", "information management", "data disclosure", etc. By calculating the adjacency entropy of each character segment, it is found that the left adjacency entropy of "digital economy" is 1.5 and the right adjacency entropy is 1.3; the left adjacency entropy of "information management" is 0.8 and the right adjacency entropy is 0.7. Assuming that the preset adjacency entropy threshold is 1.0, "digital economy" is determined as an alternative term, while "information management" is excluded due to its low adjacency entropy value. In this way, through the screening of adjacency entropy values, alternative terms with high semantic relevance can be obtained, providing a basis for the subsequent generation of word sets.

[0061] Considering that the left and right neighbor entropy is an effective statistic for measuring the degree of freedom of the boundary, the left neighbor entropy of a candidate phrase refers to the sum of the information entropy of the candidate phrase and its adjacent left neighbor set, which is used to measure the uncertainty of the left neighbor words of the candidate phrase. The larger the left neighbor entropy, the greater the uncertainty of the left neighbor words of the candidate phrase, that is, the greater the possibility that the candidate phrase becomes the left boundary of a new word. The same is true for the right neighbor entropy. The left and right neighbor entropies can be used to determine the left and right boundaries of a word, thereby determining a new word. The calculation formulas for the left and right neighbor entropies are as follows:

[0062] Among them, Q l is the set of left adjacent words of candidate word w; p(w l |w) is a candidate word When it appears, its left adjacent word is w l The conditional probability of H l (w) is the left adjacent entropy of candidate word w; Q r is the set of right adjacent words of candidate word w; p(w r |w) is the case when candidate word w appears, its right adjacent word is w r The conditional probability of H r (w) is the right neighbor entropy of candidate word w.

[0063] Based on the above theory, optionally, the adjacency entropy includes a first adjacency entropy and a second adjacency entropy, the adjacency entropy threshold includes a first adjacency entropy threshold and a second adjacency entropy threshold, and sub-step 2041 includes the following sub-steps: Sub-step 20411, when the first adjacency entropy is less than or equal to the first adjacency entropy threshold, update the target character segment with the first updated character segment of the target character segment.

[0064] The first updated character segment is a character segment in the corpus text that has one more preceding character than the target character segment.

[0065] In some embodiments of the present application, when the first adjacency entropy (left adjacency entropy) is less than or equal to the first adjacency entropy threshold, the target character segment is updated with the first updated character segment of the target character segment. In the specific implementation process, when it is found that the first adjacency entropy of the target character segment is less than or equal to the preset first adjacency entropy threshold, the character segment is extended forward by one character to form a new character segment (i.e., the first updated character segment). This first updated character segment contains one more preceding character than the original target character segment, thereby potentially improving its semantic relevance and information content.

[0066] For example, the target character segment extracted from the corpus text is "据公开" (publicly reported), and its first adjacency entropy is 0.6, which is lower than the preset first adjacency entropy threshold of 0.7. Therefore, the target character segment is extended forward by one character to become "数据公开" (public data disclosure), and the first updated character segment "数据公开" (public data disclosure) is obtained. Through this extension operation, the semantic relevance of the target character segment is improved, meeting the requirements for further screening and processing.

[0067] Sub-step 20412, when the second adjacency entropy is less than or equal to the second adjacency entropy threshold, the target character segment is updated with the second updated character segment of the target character segment.

[0068] Among them, the second updated character segment is a character segment in the corpus text that has one more subsequent character than the target character segment.

[0069] In some embodiments of the present application, when the second adjacency entropy (right adjacency entropy) is less than or equal to the second adjacency entropy threshold, the target character segment is updated with the second updated character segment of the target character segment. In the specific implementation process, when the second adjacency entropy of the target character segment is less than or equal to the preset second adjacency entropy threshold, the character segment is extended backward by one character to form a new character segment (i.e., the second updated character segment). The second updated character segment contains one more subsequent character than the original target character segment, thereby potentially increasing its semantic relevance and information content.

[0070] For example, the target character segment extracted from the corpus text is "数字经" (digital economy), and its second adjacency entropy is 0.4, which is lower than the preset second adjacency entropy threshold of 0.5. Therefore, the target character segment is extended backward by one character to become "数字经济" (digital economy), and the second updated character segment "数字经济" (digital economy) is obtained. Through this extension operation, the semantic relevance of the target character segment is enhanced, making it suitable for further screening and processing.

[0071] Sub-step 20413, when the number of characters in the updated target character segment is greater than the preset character truncation length, the target character segment is discarded.

[0072] In some embodiments of the present application, if the number of characters in the updated target character segment is greater than the preset character interception length, the target character segment will be discarded. In the specific implementation process, when the target character segment is updated (forward or backward expansion), if its number of characters exceeds the preset interception length, it is considered that the character segment is too long, which may affect its effectiveness in new word discovery. Therefore, these character segments with too long characters will be discarded to ensure that the length of the character segments in the word set is appropriate, which is conducive to subsequent processing and analysis.

[0073] For example, the preset character interception length is 3 characters. During the update process, the target character segment "development zone" is extended forward, and the updated character segment is "economic development zone". Since the updated character segment length exceeds the preset interception length (3 characters), the character segment will be discarded. In this way, the character segment length in the character segment set is kept within a reasonable range, ensuring the effectiveness and accuracy of the subsequent new word discovery process.

[0074] Sub-step 20414, when the first adjacency entropy is greater than the first adjacency entropy threshold and the second adjacency entropy is greater than the second adjacency entropy threshold, determine the target character segment as a candidate word.

[0075] In some embodiments of the present application, when the first adjacency entropy is greater than the first adjacency entropy threshold and the second adjacency entropy is greater than the second adjacency entropy threshold, the target character segment is determined as an alternative word. In the specific implementation process, the first adjacency entropy and the second adjacency entropy of each target character segment are calculated and compared with the preset adjacency entropy thresholds respectively. When both adjacency entropy values ​​of the target character segment are greater than the corresponding thresholds, it indicates that the character segment has a higher semantic relevance and information content, and is thus determined as an alternative word.

[0076] For example, the target character segment extracted from the corpus text is "digital economy", whose first adjacent entropy is 1.5 and second adjacent entropy is 1.3, both higher than the preset first adjacent entropy threshold of 1.0 and second adjacent entropy threshold of 1.0. Therefore, "digital economy" is determined as an alternative word. Through this screening process, it is ensured that the alternative words have high semantic relevance and information content, providing a reliable basis for subsequent new word discovery.

[0077] like Figure 3 As shown, in a specific program execution scheme of the present application, the method of sub-steps 20411 to 20414 can be implemented through the process of steps S1 to S5: Step S1: Determine whether the length of the character string is exceeded. Before executing sub-steps 20411 to 20414, it is first necessary to determine whether the length of the updated target character segment exceeds the preset character interception length.

[0078] Step S2: Calculate the first adjacency entropy and the second adjacency entropy The first adjacency entropy and the second adjacency entropy of each target character segment are calculated. The adjacency entropy is used to measure the semantic relevance and information content of the character segment.

[0079] Step S3: Determine whether the first adjacency entropy is greater than the first adjacency entropy threshold: If the first adjacency entropy is greater than the first adjacency entropy threshold, the target segment has good semantic relevance, skip the update, and enter step S4.1. If the first adjacency entropy is less than or equal to the first adjacency entropy threshold, the target segment needs to be updated, and enter step S3.2: Forward update: Update the target segment with the first updated segment of the target segment (i.e., the segment with one more preceding character than the target segment).

[0080] Step S4.1: Determine whether the second adjacency entropy is greater than the second adjacency entropy threshold: If the second adjacency entropy is greater than the second adjacency entropy threshold, the target character segment has good semantic relevance, skip the update, and enter step S5. If the second adjacency entropy is less than or equal to the second adjacency entropy threshold, the target character segment needs to be updated: Enter step 4.2 Backward Update: Update the target character segment with the second updated character segment of the target character segment (i.e., the character segment with one more character after the target character segment).

[0081] Step S5: Determine the target character segment as a candidate word. When the first adjacency entropy and the second adjacency entropy of the updated target character segment are both greater than the corresponding adjacency entropy thresholds, it indicates that the character segment has high semantic relevance and information content, and it is determined as a candidate word and added to the word set.

[0082] Sub-step 2042, generating a word set of the corpus text based on all candidate words.

[0083] In some embodiments of the present application, a word set of the corpus text is generated based on all the candidate words. In the specific implementation process, all the candidate words screened in the previous sub-step are summarized and further sorted and verified. This process ensures that all the candidate words are reasonably screened and processed to form a complete and representative word set, providing a reliable data basis for subsequent new word discovery and semantic analysis.

[0084] For example, the candidate words selected from the previous sub-step include "digital economy", "data disclosure", "data analysis", etc. These candidate words are aggregated and sorted to form a word set of the corpus text. The word set can contain words of different types and levels to ensure that a wide range of semantics is covered, providing rich reference data for subsequent new word discovery. In this way, by aggregating and sorting the candidate words, a representative word set is generated, providing a solid foundation for the new word discovery process.

[0085] Step 205 , in response to a selection command for a second target word in the word set, the second target word in the word set is deleted to update the word set.

[0086] In some embodiments of the present application, a specific second target word in the word set is deleted according to the user's operation command on the word set to achieve dynamic update of the word set. In the specific implementation process, the system receives a selection command through the user interface or the background system, and updates the word set in real time, deleting the second target word specified by the user, so that the content in the word set is more accurate and meets the needs.

[0087] For example, in the new word discovery system, the user selects the second target word "public document" to be deleted through the interface. After receiving the selection command, the system immediately deletes the word "public document" from the word set. After the deletion operation is completed, the word set is immediately updated to remove the word. In this way, the vocabulary content in the word set is more in line with actual application needs, providing a clear data basis for further new word discovery and semantic analysis.

[0088] Step 206 , input the words in the word set into the new word discovery model respectively, so as to determine the first target word from the word set according to the probability feature of each word determined by the new word discovery model.

[0089] The probability feature is used to characterize the meaning type and / or position type of each character in the word.

[0090] The method shown in this step has been explained in step 103 and will not be repeated here.

[0091] Optional, such as Figure 4 As shown, the new word discovery model includes a representation layer, a feature extraction layer, a fully connected layer, and an algorithm output layer. In order to determine the first target word from the word set according to the probability feature of each word determined by the new word discovery model, step 206 specifically includes the following sub-steps: Sub-step 2061, input the word into the representation layer to obtain the representation features of the word.

[0092] In some embodiments of the present application, words are input into the representation layer to obtain representation features of the words. In the specific implementation process, each word is input into the representation layer for feature representation. The representation layer uses a pre-trained word vector model (such as word vector (Word2Vec) or global vector (GloVe)) or a language model (such as a bidirectional encoder representation converter (Bidirectional Encoder Representations from Transformers, BERT)) to encode words and generate high-dimensional representation features. The representation features are used to capture the semantic information and contextual relationships of words, providing a basis for subsequent feature extraction and model prediction.

[0093] For example, input the word "digital economy" to the representation layer. The representation layer uses the BERT model to encode the word and generate a representation feature vector containing semantic information and contextual relationships. This representation feature vector can effectively reflect the semantic characteristics of "digital economy" and provide basic data support for the subsequent feature extraction layer. In this way, the system can perform feature representation on each word, ensuring that the model can capture the rich semantic information of the word during the processing process.

[0094] Sub-step 2062, inputting the representation features of the words into the extraction layer to obtain enhanced features of the representation features.

[0095] In some embodiments of the present application, the representation features of the words are input into the extraction layer to obtain enhanced features of the representation features. In the specific implementation process, the representation features of the words are further processed and abstracted through the feature extraction layer in the deep neural network. The extraction layer is usually composed of multiple bidirectional long short-term memory networks (Bidirectional Long Short-Term Memory, Bi-LSTM), which are used to capture and strengthen the context and semantic information in the representation features of the words. The enhanced features generated by this process can better reflect the deep semantic characteristics of the words and provide high-quality input for the subsequent fully connected layer processing.

[0096] For example, the feature vector "[1.2, 0.8, 2.5, ...]" is input into the feature extraction layer. The feature extraction layer consists of two Bi-LSTM layers, which are processed to output the enhanced feature vector "[3.4, 2.1, 4.8, ...]". Through this processing, the enhanced feature vector more comprehensively represents the semantic information and contextual relationship of the word "digital economy", providing a solid foundation for further model prediction and probability feature calculation.

[0097] Sub-step 2063, inputting the enhanced features of the words into the fully connected layer to obtain the probability features of the words.

[0098] In some embodiments of the present application, the enhanced features of words are input into a fully connected layer to obtain the probability features of words. In the specific implementation process, the enhanced feature vectors generated by the feature extraction layer are input into the fully connected layer for further processing. The fully connected layer consists of multiple neurons, which perform weighted summation and activation function processing on the input enhanced features to generate the probability features of words. The probability features are used to describe the possible meanings and position types of each character in the word, providing a basis for the final output of the model.

[0099] For example, the enhanced feature vector "[3.4, 2.1, 4.8,...]" is input into the fully connected layer. The fully connected layer generates a probability feature vector "[0.7, 0.3, 0.9,...]" through weighted summation and activation function processing. This probability feature vector represents the probability distribution of the meanings and position types of each character in the word "digital economy". For example, the probability that the character "数" belongs to the initial or other positions of the vocabulary and belongs to multiple preset vocabulary type labels provides necessary semantic information support for the next step of processing of the model. In this way, through the calculation of the fully connected layer, the system can more accurately describe the meaning and position of the word, which helps to improve the accuracy of new word discovery.

[0100] In a specific embodiment of the present application, when the specific field of the new word is limited to departmental documents, the position type can be divided into start (B) or internal (I); and at the same time, the vocabulary types are divided into six entity types: organization (ORG), geographical location (LOC), term slogan (SLO), rules and regulations (PLC), theoretical document (THY), activity meeting (SPC); at the same time, when the vocabulary type does not belong to the above twelve types, that is, when the corresponding character is a single character and does not belong to any target vocabulary, the position type of the vocabulary is defined as the external type (O); therefore, each character can be characterized by a 13-dimensional probability vector to represent the probability of belonging to the corresponding type and position. That is, at this time, each component of the vector corresponds to the probability of the character belonging to the thirteen types of B-ORG, I-ORG, B-LOC, I-LOC, B-SLO, I-SLO, B-PLC, I-PLC, B-THY, I-THY, B-SPC, I-SPC, and O. For example, assuming that the component [B-ORG] of the probability vector of the character "数" in the above example is 0.3, it means that "the probability that this character is the first character of an organization type word is 0.3".

[0101] Sub-step 2064: Input the probability features of the word into the algorithm output layer to determine the matching relationship between the meaning type and / or position type of each character included in the probability features of the word, and when the matching relationship meets the matching logic requirements, determine the word as the first target word.

[0102] In some embodiments of the present application, the probability features of words are input into the algorithm output layer to determine the matching relationship between the meaning type and / or position type of each character in the word, and when the matching relationship meets the matching logic requirements, the word is determined as the first target word. In the specific implementation process, by inputting the probability features generated by the fully connected layer into the algorithm output layer, such as a Conditional Random Fields (CRF), the algorithm output layer will combine the meaning and position relationship of each character in the word to determine whether it meets the preset matching logic. If the matching logic holds, the word is determined as the first target word.

[0103] For example, the probability feature vectors of the characters "shu" and "zi" in the word "digital economy" are respectively input into the algorithm output layer. The algorithm output layer analyzes the character meaning and position type relationship in the probability features. For example, if the probability that a character "shu" is the initial character (B-THY) of a theoretical document is very high, then the second character "zi" is assumed to have the internal character (I-THY) of the same theoretical document and the internal character (I-SLO) of the term slogan. Obviously, it is more logical for the initial character of the theoretical document to be followed by the internal character of the theoretical document. Therefore, while the algorithm output layer can determine that "digital economy" is the first target word, the word "digital economy" belongs to a theoretical document word. Through this process, the system can accurately identify and determine new words, ensuring the accuracy and effectiveness of new word discovery.

[0104] Reference Figure 5 , which shows a two-stage new word discovery system designed according to the method provided in the embodiments of the present application, mainly composed of three main subsystems: Data collection and preprocessing subsystem: Responsible for scraping data from the target website and cleaning it to ensure the reliability of the corpus.

[0105] First-stage new word discovery subsystem: Screen and discover new words through preliminary statistical methods.

[0106] Second-stage new word discovery subsystem: Further optimize and confirm new words through a deep learning model.

[0107] Based on the above architecture, the following steps are specifically implemented: Step T1: Data collection and preprocessing: Use the scraping module to regularly obtain the latest public information from the target website, and use the cleaning module to process the scraped original text data, delete HTML tags and special characters, clean up noise, and ensure that a clean corpus text is obtained.

[0108] Step T2: The first stage of new word discovery: Use the data preparation module I to segment the corpus, filter out the character segments with higher frequency, and generate preliminary candidate phrases through filtering through the stop word library; then based on statistical methods, calculate the mutual information and left and right neighbor entropy values, and perform boundary expansion and screening to obtain the new word discovery results of the first stage.

[0109] Step T3: Second stage new word discovery: Use data preparation module II to build a basic vocabulary, classify and annotate new word entities, form a data set for training deep learning models, and then use new word discovery module II to extract deep semantic features of the text, and finally output new words to form the new word discovery results of the second stage.

[0110] Specifically, refer to Figure 6 , step T1 performs data collection and preprocessing through the following process: Step T1.1: First, the system sets a timer to automatically trigger the data collection process at fixed time intervals (such as daily, weekly, etc.). This ensures that the system regularly obtains the latest text data from the target website to maintain the timeliness and freshness of the data.

[0111] Step T1.2: Use a web crawler to automatically access the target website. The program identifies the structure of the web page, locates the web page area containing information, and extracts the text data in the web page. The crawled text data includes newly released public documents, announcements and other public information.

[0112] Step T1.3: Clean the captured raw text data and use regular expressions to filter out irrelevant content. Specific operations include deleting HTML tags, special characters (such as &, #, etc.) and blank characters (such as extra line breaks, tabs, spaces, etc.) to ensure clean text data.

[0113] Step T1.4: Compare the cleaned text data with the preset reference word library. Remove the words in the text data that already exist in the reference word library to ensure that the extracted words are innovative and unique in the subsequent new word discovery process. In this way, after comparison and filtering, the obtained text data is more accurate and suitable for further analysis of new word discovery.

[0114] In summary, in the embodiment of the present application, by introducing the adjacency entropy indicator, insignificant character segments in the corpus are effectively filtered out, thereby reducing the amount of data for subsequent processing, improving processing efficiency, and then determining the word set of the corpus text, more accurately defining the boundaries of the character segments, and ensuring that the extracted words have high semantic relevance. Then the words in the word set are respectively input into the new word discovery model, and the first target word is determined from the word set through the probability characteristics of each word determined by the new word discovery model. In combination with the deep learning algorithm, the meaning type and / or position type of each character in the word can be represented according to the probability characteristics, realizing the transition from statistical features to semantic features, and utilizing the powerful representation ability of deep learning to improve the accuracy of new word discovery. Therefore, the method based on the embodiment of the present application overcomes the problems faced by rule-based and statistical methods in the prior art, solves the problems of relying on high-quality expert knowledge and pre-screening of large amounts of data, realizes the improvement of new word discovery efficiency and accuracy, and solves the problems of poor versatility and high maintenance cost in the existing new word discovery process.

[0115] refer to Figure 7 , which shows a two-stage new word discovery device 30 provided in an embodiment of the present application, comprising: The extraction module 301 is used to extract a character segment set including multiple target character segments to be verified from the corpus text, and determine the adjacency entropy of each target character segment in the character segment set; the adjacency entropy is used to characterize the amount of information carried by the target character segment; The adjacency entropy module 302 is used to determine a word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set; the word set contains at least one word to be verified; The selection module 303 is used to input the words in the word set into the new word discovery model respectively, so as to determine the first target word from the word set according to the probability feature of each word determined by the new word discovery model; the probability feature is used to characterize the meaning type and / or position type of each character in the word.

[0116] Optionally, the extraction module 301 includes: The mutual information calculation submodule is used to extract multiple character segments of the corpus text according to a preset character truncation length, and determine the mutual information value of each character segment; the mutual information value of the character segment is positively correlated with the probability that the single characters in the character segment just constitute a word; The mutual information comparison submodule is used to determine the character segments whose mutual information values ​​are greater than a preset mutual information threshold as target character segments to obtain a character segment set.

[0117] Optionally, based on the mutual information calculation submodule and the mutual information comparison submodule, the two-stage new word discovery device 30 provided in the embodiment of the present application further includes: The character segment set updating module is used to delete the target character segments in the character segment set that are the same as the reference words in the preset reference word library to update the character segment set.

[0118] Optionally, the adjacency entropy module 302 includes: The candidate word submodule is used to determine the candidate words from the target character segment according to the relationship between the adjacency entropy of each target character segment and the preset adjacency entropy threshold in the character segment set; The word set submodule is used to generate a word set of the corpus text based on all the candidate words.

[0119] Optionally, the adjacency entropy includes a first adjacency entropy and a second adjacency entropy, the adjacency entropy threshold includes a first adjacency entropy threshold and a second adjacency entropy threshold, and the candidate word submodule includes: The first expansion unit is used to update the target character segment with a first update character segment of the target character segment when the first adjacency entropy is less than or equal to the first adjacency entropy threshold; the first update character segment is a character segment in the corpus text that has one more preceding character than the target character segment; The second expansion unit is used to update the target character segment with a second update character segment of the target character segment when the second adjacency entropy is less than or equal to the second adjacency entropy threshold; the second update character segment is a character segment in the corpus text that has one more character after the target character segment; A discarding unit, used to discard the target character segment when the number of characters in the updated target character segment is greater than a preset character interception length; The confirmation unit is used to determine the target character segment as a candidate word when the first adjacency entropy is greater than a first adjacency entropy threshold and the second adjacency entropy is greater than a second adjacency entropy threshold.

[0120] Optionally, the new word discovery model includes a representation layer, a feature extraction layer, a fully connected layer, and an algorithm output layer, and the selection module 303 includes: The representation layer submodule is used to input words into the representation layer to obtain the representation features of the words; A feature extraction layer submodule, used for inputting the representation features of the words into the extraction layer to obtain enhanced features of the representation features; The fully connected layer submodule is used to input the enhanced features of the words into the fully connected layer to obtain the probability features of the words; The algorithm output layer submodule is used to input the probability characteristics of the word into the algorithm output layer to determine the matching relationship between the meaning type and / or position type of each character contained in the probability characteristics of the word, and determine the word as the first target word when the matching relationship meets the matching logic requirements.

[0121] Optionally, the two-stage new word discovery device 30 provided in the embodiment of the present application further includes: A crawling module is used to obtain web page data with text from a preset target website; The cleaning module is used to remove format characters in web page data to obtain corpus text.

[0122] Optionally, the two-stage new word discovery device 30 provided in the embodiment of the present application further includes: The word set updating module is used to delete the second target word in the word set in response to a selection command for the second target word in the word set, so as to update the word set.

[0123] In summary, in the embodiment of the present application, by introducing the adjacency entropy indicator, insignificant character segments in the corpus are effectively filtered out, thereby reducing the amount of data for subsequent processing, improving processing efficiency, and then determining the word set of the corpus text, more accurately defining the boundaries of the character segments, and ensuring that the extracted words have high semantic relevance. Then the words in the word set are respectively input into the new word discovery model, and the first target word is determined from the word set through the probability characteristics of each word determined by the new word discovery model. In combination with the deep learning algorithm, the meaning type and / or position type of each character in the word can be represented according to the probability characteristics, realizing the transition from statistical features to semantic features, and utilizing the powerful representation ability of deep learning to improve the accuracy of new word discovery. Therefore, the method based on the embodiment of the present application overcomes the problems faced by rule-based and statistical methods in the prior art, solves the problems of relying on high-quality expert knowledge and pre-screening of large amounts of data, realizes the improvement of new word discovery efficiency and accuracy, and solves the problems of poor versatility and high maintenance cost in the existing new word discovery process.

[0124] Reference Figure 8 , the electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output (I / O) interface 512 , a sensor component 514 , and a communication component 516 .

[0125] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0126] The memory 504 is used to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0127] The power supply component 506 provides power to the various components of the electronic device 500. The power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 500.

[0128] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0129] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC), and when the electronic device 500 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is used to receive an external audio signal. The received audio signal can be further stored in the memory 504 or sent via the communication component 516. In some embodiments, the audio component 510 also includes a speaker for outputting audio signals.

[0130] The input / output I / O interface 512 provides an interface between the processing component 502 and the peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0131] The sensor assembly 514 includes one or more sensors for providing various aspects of status assessment for the electronic device 500. For example, the sensor assembly 514 can detect the open / closed state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500, and the sensor assembly 514 can also detect the position change of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and the temperature change of the electronic device 500. The sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0132] The communication component 516 is used to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0133] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.

[0134] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the instructions can be executed by a processor 520 of an electronic device 500 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0135] Fig. 9 FIG. 6 is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server.

[0136] Reference Fig. 9 , the electronic device 600 includes a processing component 622, which further includes one or more processors, and a memory resource represented by a memory 632, for storing instructions that can be executed by the processing component 622, such as an application. The application stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute instructions to perform the method provided in the embodiment of the present application.

[0137] The electronic device 600 may also include a power supply component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.

[0138] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0139] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A two-stage new word discovery method, characterized in that: include: Extracting a character segment set including a plurality of target character segments to be verified from the corpus text, and determining the adjacency entropy of each target character segment in the character segment set; The adjacency entropy is used to characterize the amount of information carried by the target character segment; Determine a word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set; the word set contains at least one word to be verified; Inputting the words in the word set into a new word discovery model respectively, so as to determine a first target word from the word set according to the probability feature of each of the words determined by the new word discovery model; The probability feature is used to characterize the meaning type and / or position type of each character in the word.

2. The method according to claim 1, characterized in that The step of extracting a character segment set including a plurality of target character segments to be verified from the corpus text includes: According to a preset character truncation length, a plurality of character segments of the corpus text are extracted, and a mutual information value of each character segment is determined; the mutual information value of the character segment is positively correlated with the probability that a single word in the character segment just constitutes a word; The character segment whose mutual information value is greater than a preset mutual information threshold is determined as the target character segment to obtain the character segment set.

3. The method according to claim 1, characterized in that Determining the word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set includes: Determine candidate words from the target character segments according to the relationship between the adjacency entropy of each target character segment in the character segment set and a preset adjacency entropy threshold; A word set of the corpus text is generated based on all the candidate words.

4. The method according to claim 3, characterized in that The adjacency entropy includes a first adjacency entropy and a second adjacency entropy, the adjacency entropy threshold includes a first adjacency entropy threshold and a second adjacency entropy threshold, and determining the candidate words from the target character segment according to the relationship between the adjacency entropy of each target character segment in the character segment set and a preset adjacency entropy threshold, comprises: When the first adjacency entropy is less than or equal to the first adjacency entropy threshold, the target character segment is updated with a first updated character segment of the target character segment; the first updated character segment is a character segment in the corpus text that has one more preceding character than the target character segment; When the second adjacency entropy is less than or equal to the second adjacency entropy threshold, the target character segment is updated with a second updated character segment of the target character segment; the second updated character segment is a character segment in the corpus text that has one more character after the target character segment; When the number of characters in the updated target character segment is greater than the preset character interception length, the target character segment is discarded; When the first adjacency entropy is greater than the first adjacency entropy threshold, and the second adjacency entropy is greater than the second adjacency entropy threshold, the target character segment is determined as an alternative word.

5. The method according to claim 1, characterized in that The new word discovery model comprises a representation layer, a feature extraction layer, a fully connected layer and an algorithm output layer. The first target word is determined from the word set according to the probability feature of each word determined by the new word discovery model, including: Inputting the word into the representation layer to obtain the representation features of the word; Inputting the representation features of the words into the extraction layer to obtain enhanced features of the representation features; Inputting the enhanced features of the word into the fully connected layer to obtain the probability features of the word; The probability feature of the word is input into the algorithm output layer to determine the matching relationship between the meaning type and / or position type of each character contained in the probability feature of the word, and when the matching relationship meets the matching logic requirements, the word is determined as the first target word.

6. The method according to claim 1, characterized in that The method further comprises: Get web page data with text from the preset target website; Format characters in the web page data are deleted to obtain the corpus text.

7. The method according to claim 1, characterized in that The method further comprises: In response to a selection command for a second target word in the word set, the second target word in the word set is deleted to update the word set.

8. A two-stage new word discovery device, characterized in that: include: An extraction module is used to extract a character segment set including a plurality of target character segments to be verified from the corpus text, and determine the adjacency entropy of each target character segment in the character segment set; the adjacency entropy is used to characterize the amount of information carried by the target character segment; An adjacency entropy module, used to determine a word set of the corpus text according to the adjacency entropy of each target character segment in the character segment set; the word set contains at least one word to be verified; A selection module, configured to input the words in the word set into a new word discovery model respectively, so as to determine a first target word from the word set according to the probability feature of each of the words determined by the new word discovery model; The probability feature is used to characterize the meaning type and / or position type of each character in the word.

9. An electronic device, characterized in that: include: A processor, a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method as claimed in any one of claims 1 to 7.