Military vocabulary mining method, device and equipment and storage medium
By matching the word segmentation and rule database of military corpus, combined with new word probability screening and proofreading, the problem of the inability to identify military vocabulary in the existing technology is solved, and efficient and accurate military vocabulary mining is achieved.
Patent Information
- Application Number
- CN202311714978.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art cannot accurately identify military vocabulary in the original data, resulting in the inability to effectively mine new words existing in the original data.
By segmenting the recognized military corpus, obtaining the initial words, and identifying and screening these initial words according to the rule base of the target military dictionary, obtaining candidate words, and then filtering out the target words based on the probability of new words, and proofing the target words to determine the new military vocabulary.
It has achieved accurate identification and excavation of new words in the military industry, and improved the efficiency and accuracy of military industry vocabulary mining.
Smart Images

Figure CN120163155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data mining, and particularly to a method, device, equipment and storage medium for military vocabulary mining. Background Art
[0002] In the military field, the text content often has a certain specific format, and military vocabulary has the characteristics of professionalism. Therefore, the current natural language processing (NLP) methods cannot accurately identify the military vocabulary in the original data, resulting in the inability to effectively mine the new words existing in the original data.
[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, device, equipment and storage medium for military vocabulary mining, aiming to solve the technical problem that the existing technology cannot accurately identify the military vocabulary in the original data, resulting in the inability to effectively mine the new words existing in the original data.
[0005] To achieve the above purpose, the present invention provides a method for military vocabulary mining, the method comprising the following steps:
[0006] Perform word segmentation on the military corpus to be recognized to obtain a plurality of initial words;
[0007] Identify each of the plurality of initial words according to the rule base corresponding to the target military dictionary, and screen out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results;
[0008] Obtain the new word probabilities of each of the candidate words, and screen out target words from each of the candidate words according to the new word probabilities;
[0009] Proofread each of the target words, and determine the military new words in each of the target words according to the proofreading results.
[0010] Optionally, the performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words includes:
[0011] Obtain the byte information of the military corpus to be recognized;
[0012] Determine the segmentation strategy of the military corpus to be recognized according to the byte information;
[0013] Perform word segmentation on the military corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words.
[0014] Optionally, determining the segmentation strategy of the military industrial corpus to be recognized according to the byte information includes:
[0015] Determining each byte included in the military industrial corpus to be recognized according to the byte information, and the byte sequence of the military industrial corpus to be recognized;
[0016] Determining the segmentation path according to each byte and the byte sequence;
[0017] Determining the segmentation strategy of the military industrial corpus to be recognized according to the segmentation path.
[0018] Optionally, determining the segmentation path according to each byte and the byte sequence includes:
[0019] Obtaining a target military industrial dictionary and a general word segmentation dictionary;
[0020] Determining the segmentation length of the military industrial corpus to be recognized according to the target military industrial dictionary and the general word segmentation dictionary;
[0021] Traversing each byte according to the byte sequence and the segmentation length;
[0022] Performing segmentation planning according to the traversal result to determine the segmentation path.
[0023] Optionally, segmenting the military industrial corpus to be recognized to obtain multiple initial words includes:
[0024] Obtaining a target military industrial dictionary and a general word segmentation dictionary;
[0025] Segmenting the military industrial corpus to be recognized according to the target military industrial dictionary and the general word segmentation dictionary to obtain original words;
[0026] Performing string matching on the original words according to a stop word dictionary to determine the stop words to be removed in the original words;
[0027] Removing the stop words to be removed from the original words to obtain multiple initial words.
[0028] Optionally, respectively recognizing the multiple initial words according to the rule base corresponding to the target military industrial dictionary, and screening out candidate words that conform to each rule in the rule base from the multiple initial words includes:
[0029] Obtaining the character order of the multiple initial words;
[0030] Combining the multiple initial words according to the character order to obtain multiple groups of initial phrases;
[0031] Obtaining the rule base corresponding to the target military industrial dictionary;
[0032] Identify each group of initial phrases according to the respective rules in the rule base, and determine whether each group of initial phrases matches each of the rules based on the identification results;
[0033] Screen candidate phrases that conform to the respective rules in the rule base from the multiple initial words according to the identification results, and determine candidate words based on the candidate phrases.
[0034] Optionally, obtaining the new word probabilities of each of the candidate words, and screening target words from each of the candidate words according to the new word probabilities, includes:
[0035] Obtain the feature vectors of each of the candidate words;
[0036] Input the feature vectors into a pre-constructed classification model to obtain the new word probabilities of each of the candidate words;
[0037] Determine whether the new word probabilities of each of the candidate words are greater than a preset probability threshold;
[0038] Screen target words from each of the candidate words according to the judgment results.
[0039] Optionally, obtaining the feature vectors of each of the candidate words, includes:
[0040] Traverse the military industrial corpus to be identified based on each of the candidate words;
[0041] Determine the word frequency and inverse document frequency index of each of the candidate words according to the traversal results;
[0042] Conduct a weight analysis on each of the candidate words according to the word frequency and the inverse document frequency index to obtain the semantic weights of each of the candidate words;
[0043] Determine the feature vectors of each of the candidate words according to the word features and the semantic weights of each of the candidate words.
[0044] Optionally, proofreading each of the target words, and determining the military industrial new words in each of the target words according to the proofreading results, includes:
[0045] Traverse the text corresponding to the military industrial corpus to be identified based on each of the target words;
[0046] Locate the target new words in the text according to the traversal results, and mark the located target new words;
[0047] Proofread each of the target words based on the marked text;
[0048] Determine the military industrial new words in each of the target words according to the proofreading results.
[0049] Optionally, after proofreading each of the target words and determining the new military industry words among the target words according to the proofreading results, the following steps are included:
[0050] Synchronize the new military industry words to the target military industry dictionary to complete the update of the target military industry dictionary;
[0051] Generate a proofreading positive sample set and a proofreading negative sample set according to the proofreading results;
[0052] Iteratively train the classification model according to the proofreading positive sample set and the proofreading negative sample set.
[0053] In addition, to achieve the above object, the present invention also proposes a military industry vocabulary mining device, and the military industry vocabulary mining device includes:
[0054] A word segmentation module, configured to segment the to-be-recognized military industry corpus to obtain a plurality of initial words;
[0055] A rule matching module, configured to respectively identify the plurality of initial words according to a rule base corresponding to the target military industry dictionary, and screen out candidate words that conform to each rule in the rule base from the plurality of initial words according to the identification results;
[0056] A probability identification module, configured to obtain the new word probabilities of the candidate words, and screen out target words from the candidate words according to the new word probabilities;
[0057] A new word proofreading module, configured to proofread each of the target words and determine the new military industry words among the target words according to the proofreading results.
[0058] Further, the word segmentation module is further configured to obtain the byte information of the to-be-recognized military industry corpus; determine the segmentation strategy of the to-be-recognized military industry corpus according to the byte information; segment the to-be-recognized military industry corpus according to the segmentation strategy to obtain a plurality of initial words.
[0059] Further, the word segmentation module is further configured to determine each byte included in the to-be-recognized military industry corpus and the byte sequence of the to-be-recognized military industry corpus according to the byte information; determine a segmentation path according to each byte and the byte sequence; determine the segmentation strategy of the to-be-recognized military industry corpus according to the segmentation path.
[0060] Further, the word segmentation module is further configured to obtain the target military industry dictionary and a general word segmentation dictionary; determine the segmentation length of the to-be-recognized military industry corpus according to the target military industry dictionary and the general word segmentation dictionary; traverse each byte according to the byte sequence and the segmentation length; perform a segmentation plan according to the traversal result to determine a segmentation path.
[0061] Further, the word segmentation module is further configured to obtain a target military dictionary and a general word segmentation dictionary; segment the to-be-recognized military corpus according to the target military dictionary and the general word segmentation dictionary to obtain original words; perform string matching on the original words according to a stop word dictionary to determine the stop words to be removed in the original words; remove the stop words to be removed from the original words to obtain a plurality of initial words.
[0062] Further, the rule matching module is further configured to obtain the character order of the plurality of initial words; combine the plurality of initial words according to the character order to obtain multiple groups of initial phrases; obtain a rule library corresponding to the target military dictionary; respectively identify the multiple groups of initial phrases according to each rule in the rule library, and determine whether the multiple groups of initial phrases match each of the rules according to the identification results; screen out candidate phrases that conform to each rule in the rule library from the plurality of initial words according to the identification results, and determine candidate words according to the candidate phrases.
[0063] Further, the probability recognition module is further configured to obtain the feature vectors of the candidate words; input the feature vectors into a pre-constructed classification model to obtain the new word probabilities of the candidate words; determine whether the new word probabilities of the candidate words are greater than a preset probability threshold; screen out target words from the candidate words according to the determination results.
[0064] Further, the probability recognition module is further configured to traverse the to-be-recognized military corpus based on the candidate words; determine the word frequency and inverse document frequency index of each candidate word according to the traversal results; perform weight analysis on each candidate word according to the word frequency and the inverse document frequency index to obtain the semantic weights of the candidate words; determine the feature vectors of the candidate words according to the word features and the semantic weights of the candidate words.
[0065] In addition, to achieve the above object, the present invention further provides a military vocabulary mining device, where the military vocabulary mining device includes: a memory, a processor, and a military vocabulary mining program stored on the memory and executable on the processor, and the military vocabulary mining program is configured to implement the steps of the military vocabulary mining method as described above.
[0066] In addition, to achieve the above object, the present invention further provides a storage medium, where a military vocabulary mining program is stored on the storage medium, and when the military vocabulary mining program is executed by a processor, the steps of the military vocabulary mining method as described above are implemented.
[0067] The present invention performs word segmentation on the military-industry corpus to be recognized, obtains a plurality of initial words, respectively recognizes the plurality of initial words according to the rule base corresponding to the target military-industry dictionary, and screens out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results, obtains the new-word probabilities of the candidate words, and screens out target words from the candidate words according to the new-word probabilities, proofreads the target words, and determines military-industry new words among the target words according to the proofreading results; since the present invention performs word segmentation on the military-industry corpus to be recognized and obtains a plurality of initial words, the lexical segmentation of the military-industry corpus to be recognized is completed, the plurality of initial words are respectively recognized according to the rule base corresponding to the target military-industry dictionary, and candidate words that conform to each rule in the rule base are screened out from the plurality of initial words according to the recognition results, so that rule matching for the segmented initial words is realized to retain candidate words that conform to each rule in the rule base of the target military-industry dictionary, and then screening is performed according to the new-word probabilities of the candidate words, so that invalid old words are effectively eliminated, the target words are proofread, and military-industry new words among the target words are determined according to the proofreading results, so that the screened target words are all new words, accurately mining new words in the military-industry field is realized, and the efficiency of military-industry vocabulary mining is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a schematic structural diagram of a military-industry vocabulary mining device in the hardware operating environment related to the solution of the embodiment of the present invention;
[0069] Figure 2 is a schematic flowchart of the first embodiment of the military-industry vocabulary mining method of the present invention;
[0070] Figure 3 is a schematic flowchart of the second embodiment of the military-industry vocabulary mining method of the present invention;
[0071] Figure 4 is a schematic flowchart of the third embodiment of the military-industry vocabulary mining method of the present invention;
[0072] Figure 5 is a structural block diagram of the first embodiment of the military-industry vocabulary mining device of the present invention.
[0073] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0075] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a military-industry vocabulary mining device in the hardware operating environment related to the solution of the embodiment of the present invention.
[0076] As shown Figure 1 , the military vocabulary mining device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0077] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the military vocabulary mining device, and it may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0078] As shown Figure 1 , in the memory 1005 as a storage medium, there may be included an operating system, a network communication module, a user interface module, and a military vocabulary mining program.
[0079] In Figure 1 the military vocabulary mining device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the military vocabulary mining device of the present invention may be provided in the military vocabulary mining device. The military vocabulary mining device calls the military vocabulary mining program stored in the memory 1005 through the processor 1001 and executes the military vocabulary mining method provided by the embodiments of the present invention.
[0080] The embodiments of the present invention provide a military vocabulary mining method. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of a military vocabulary mining method of the present invention.
[0081] In this embodiment, the military vocabulary mining method includes the following steps:
[0082] Step S10: Segment the military industrial corpus to be recognized to obtain multiple initial words.
[0083] It should be understood that the execution subject of the method in this embodiment can be a military industrial vocabulary mining device with data processing, network communication, and program running functions, such as a computer, or other devices or equipment that can achieve the same or similar functions. Here, the above-mentioned military industrial vocabulary mining device (hereinafter referred to as the vocabulary mining device) is taken as an example for illustration.
[0084] It should be noted that the military industrial corpus to be recognized can be language materials in the military industrial field that need to be mined for new military words. The above-mentioned initial words can be words related to the military industrial field segmented from the military industrial corpus to be recognized.
[0085] It should be understood that the vocabulary mining device obtains the original resources in the military industrial database, loads the original resources to obtain the original data, processes the original data, thereby obtaining the text content, extracts the military industrial corpus to be recognized that needs to be recognized according to the text content, and segments the military industrial corpus to be recognized to obtain multiple initial words.
[0086] Further, in order to effectively segment the military industrial corpus to be recognized, the above step S10 may include:
[0087] Obtain the target military industrial dictionary and the general segmentation dictionary;
[0088] Segment the military industrial corpus to be recognized according to the target military industrial dictionary and the general segmentation dictionary to obtain the original words;
[0089] Perform string matching on the original words according to the stop word dictionary to determine the stop words in the original words;
[0090] Remove the stop words from the original words to obtain multiple initial words.
[0091] It should be noted that the target military industrial dictionary can be a professional term dictionary registered in the vocabulary mining device in the military industrial field, that is, the target military industrial dictionary can be a database for storing military industrial vocabulary; the target military industrial dictionary can also be a dictionary that needs to register new military words. The vocabulary in the target military industrial dictionary can be obtained by the vocabulary mining device from existing databases in the military industrial professional field (such as the military encyclopedia thesaurus, etc.), or can be military industrial vocabulary mined by the vocabulary mining device in the past and saved in the dictionary. The format of the target military industrial dictionary can include the form of a triple of word, type, and weight. Among them, word is the vocabulary, type is the part of speech, and weight is the weight value. Among them, word is a required field, type may be multi-valued, different values are separated by commas, and weight is of double type.
[0092] The above-mentioned general word segmentation dictionary can be a general word segmentation dictionary, which can be from open-source word segmentation tools. The dictionary format of the general word segmentation dictionary is the same as that of the target military dictionary. After the general word segmentation dictionary is loaded, it can be stored in memory for use. The above-mentioned segmentation length can be the character length of segmentation. For example, if the segmentation length is 2, then segmentation is performed in units of two characters. The segmentation length includes one or more types. For example, during the current segmentation process, several segmentation lengths can be used for segmentation. First, segmentation is performed with a segmentation length of 2, and then segmentation is performed with a segmentation length of 3, etc. This embodiment does not impose any limitations.
[0093] It should be understood that the vocabulary mining device performs word segmentation on the military corpus to be recognized according to the target military dictionary and the general word segmentation dictionary to obtain the original words, and then filters the words to be deactivated in the original words according to the stop word dictionary, and removes the words to be deactivated from the original words to obtain a plurality of initial words.
[0094] The above-mentioned stop word dictionary can be that in information retrieval by the vocabulary mining device, in order to save storage space and improve search efficiency, certain characters or words will be automatically filtered before or after processing natural language data (or text). The stop word dictionary is constructed according to the filtering history and filtering rules.
[0095] Step S20: Identify the plurality of initial words respectively according to the rule base corresponding to the target military dictionary, and screen out candidate words that meet each rule in the rule base from the plurality of initial words according to the identification results.
[0096] It should be noted that the target military dictionary can be a professional term dictionary stored in the vocabulary mining device that registers military-related fields, that is, the target military dictionary can be a database for storing military vocabulary; the target military dictionary can also be a dictionary that needs to register new military words. The vocabulary in the target military dictionary can be obtained by the vocabulary mining device from existing databases in the military professional field (such as military encyclopedia thesaurus, etc.), or can be military vocabulary mined by the vocabulary mining device in the past and saved in the dictionary. The format of the target military dictionary can include the form of a triple of word, type, and weight. Among them, word is the vocabulary, type is the part of speech, and weight is the weight value. Among them, word is a required field, type may have multiple values, and different values are separated by commas. Weight is of double type.
[0097] The above-mentioned rule base can be a database for storing new word mining rules. One or more vocabulary mining rules are stored in the rule base. For example, the rules in the rule base can include the part-of-speech requirements for the elements of new words, the candidate word length, the number of candidate word elements, etc.
[0098] It should be understood that in order to ensure that the segmented words conform to the various rules in the rule base of the target military-industry dictionary, the vocabulary mining device in this embodiment determines the target military-industry dictionary that needs to add new military-industry words, obtains the rule base corresponding to the target military-industry dictionary, and identifies each of the initially segmented words according to the rule base, so as to judge whether each initially segmented word matches the various rules in the rule base according to the identification result, eliminate the initially segmented words that do not match the rules, retain the initially segmented words that match the various rules, and use the initially segmented words that match the various rules as candidate words.
[0099] Further, in order to ensure that the selected military-industry vocabulary conforms to the rules, step S20 may include:
[0100] Obtain the character order of the multiple initially segmented words;
[0101] Combine the multiple initially segmented words according to the character order to obtain multiple groups of initial phrases;
[0102] Obtain the rule base corresponding to the target military-industry dictionary;
[0103] Identify each group of initial phrases according to the respective rules in the rule base, and judge whether the multiple groups of initial phrases match each of the rules according to the identification result;
[0104] Screen out candidate phrases that conform to the respective rules in the rule base from the multiple initially segmented words according to the identification result, and determine candidate words according to the candidate phrases.
[0105] It should be noted that the character order may be the character arrangement order between the initially segmented words. The above initial phrases may be phrases combined based on the arrangement order of each initially segmented word. For example, adjacent initially segmented words may be combined to obtain multiple groups of phrases.
[0106] It should be understood that after the vocabulary mining device segments the military-industry corpus to be recognized, it obtains multiple initially segmented words, combines adjacent words into phrases according to the character order of each initially segmented word, and through the rule base, identifies whether the phrases can match the rules in the rule base to become potential candidate new words.
[0107] Step S30: Obtain the new-word probabilities of the respective candidate words, and screen out target words from the respective candidate words according to the new-word probabilities.
[0108] It should be noted that the new-word probability may be the probability of whether each candidate word is a new word. The candidate words with lower new-word probabilities are used as invalid old words, and the candidate words with higher new-word probabilities are used as target words.
[0109] It should be understood that in order to eliminate invalid old words and only retain new words, in this embodiment, the vocabulary mining device sets a probability threshold, then obtains the vocabulary feature information of each candidate new word, determines the new word probability of each candidate new word according to the vocabulary feature information, judges whether the new word probability of each candidate word is greater than the probability threshold, takes the words greater than the probability threshold as target words (i.e., new words), and takes the words not greater than the probability threshold as invalid old words.
[0110] In a specific implementation, the vocabulary mining device screens out 4 candidate words, namely A, B, C, and D, and obtains the new word probabilities of each candidate word as 30%, 35%, 72%, and 56% respectively. The set probability threshold is 50%. Therefore, C and D are taken as target words (i.e., new words), and A and B are invalid old words.
[0111] Step S40: Proofread each of the target words, and determine the military-industry new words among each of the target words according to the proofreading results.
[0112] It should be noted that the military-industry new words can be the words that pass the proofreading among the target words. The number of military-industry new words retained after proofreading is less than or equal to the target words. That is, after being proofread by the vocabulary mining device, the target words that pass the proofreading are all military-industry new words. That is to say, if all the target words pass the proofreading, all the target words are taken as military-industry new words.
[0113] It should be understood that in order to improve the accuracy of military-industry new words and avoid the words dug out not being new words or the vocabulary not conforming to the rules, the vocabulary mining device in this embodiment also needs to proofread each target new word, determine the words that pass the proofreading among the target words according to the proofreading results, and take the words that pass the proofreading as military-industry new words.
[0114] In a specific implementation, after the vocabulary mining device confirms the accuracy of the dug-out military-industry new words, it synchronously updates the military-industry new words to the target military-industry dictionary. After that, on the one hand, it further improves the accuracy in the word segmentation stage, and on the other hand, in a new round of recognition process, it automatically ignores the new words that have been recognized to avoid repeated calculation.
[0115] Furthermore, in order to effectively improve the accuracy of new word mining, the above step S40 may include:
[0116] Traverse the text corresponding to the military-industry corpus to be recognized according to each of the target words;
[0117] Locate the target new words in the text according to the traversal results, and mark the located target new words;
[0118] Proofread each of the target words based on the marked text;
[0119] Determine the military-industry new words among each of the target words according to the proofreading results.
[0120] It should be understood that in order to improve the accuracy of new words, the vocabulary mining device in this embodiment also needs to proofread the target words. During the recognition process of the military industrial corpus to be recognized, multiple target words are obtained, and a candidate set is constructed according to the target words. In order to know in which paragraphs and sentences of different articles these candidate new words are included, they are reviewed in the form of a list on the web page. The vocabulary mining device can see the list of recognized new words. After clicking, it can see the different article paragraphs where each candidate word appears, and the position of the candidate word is highlighted to facilitate proofreading whether the recognition of the word is accurate. The candidate words after calibration can be considered as new words.
[0121] In this embodiment, the military industrial corpus to be recognized is segmented to obtain multiple initial words. Each of the multiple initial words is recognized according to the rule base corresponding to the target military industrial dictionary, and candidate words that meet each rule in the rule base are screened out from the multiple initial words according to the recognition results. The new word probability of each candidate word is obtained, and target words are screened out from each candidate word according to the new word probability. Each of the target words is proofread, and military new words in each of the target words are determined according to the proofreading results. Since in this embodiment, the military industrial corpus to be recognized is segmented to obtain multiple initial words, the vocabulary segmentation of the military industrial corpus to be recognized is completed. Each of the multiple initial words is recognized according to the rule base corresponding to the target military industrial dictionary, and candidate words that meet each rule in the rule base are screened out from the multiple initial words according to the recognition results, so as to realize rule matching for the segmented initial vocabulary to retain candidate words that meet each rule in the rule base of the target military industrial dictionary. Then, screening is performed according to the new word probability of each candidate word, so as to effectively eliminate invalid old words. Each of the target words is proofread, and military new words in each of the target words are determined according to the proofreading results, so as to ensure that the screened target words are all new words, realize accurate mining of new words in the military industrial field, and effectively improve the efficiency of military industrial vocabulary mining.
[0122] Reference Figure 3 , Figure 3 is a schematic flowchart of the second embodiment of a method for mining military industrial vocabulary of the present invention.
[0123] Based on the above first embodiment, in this embodiment, step S10 includes:
[0124] Step S11: Obtain the byte information of the military industrial corpus to be recognized.
[0125] It should be noted that the byte information may be the relevant information of the byte text in the military industrial corpus to be recognized. For example, the byte information may include: information such as the number of bytes, byte sequence, byte sorting, each included byte, and part of speech.
[0126] It should be understood that the vocabulary mining device traverses each byte in the text of the military-industry corpus to be recognized, and obtains the byte information of the military-industry corpus to be recognized according to the traversal result.
[0127] Step S12: Determine the segmentation strategy of the military-industry corpus to be recognized according to the byte information.
[0128] It should be noted that the segmentation strategy can be a vocabulary segmentation strategy for the military-industry corpus to be recognized, and the segmentation strategy can include the segmentation length and segmentation path of each byte in the military-industry corpus to be recognized, etc.
[0129] It should be understood that the vocabulary mining device sets the segmentation length of the military-industry corpus to be recognized according to the byte information, and determines the segmentation strategy for generating the military-industry corpus to be recognized according to the segmentation length and the byte information.
[0130] In a specific implementation, the vocabulary mining device in this embodiment can perform N-gram word segmentation through multi-gram grammar, perform word segmentation processing on the text content of the military-industry corpus to be recognized, and obtain multiple words with context semantics included in the text content. There is an overlapping part between adjacent words obtained by multi-gram grammar word segmentation, so that the segmented words have context semantics; it is also possible to perform word segmentation on the military-industry corpus to be recognized according to the existing general word segmentation dictionary and the target military-industry dictionary. The above general word segmentation dictionary can be a general word segmentation dictionary, which can come from an open-source word segmentation tool. The dictionary format of the general word segmentation dictionary is the same as that of the target military-industry dictionary, and the general word segmentation dictionary can be stored in memory for use after being loaded.
[0131] Furthermore, in order to ensure the accuracy of the segmentation strategy, the above step S12 may include:
[0132] Step S121: Determine each byte included in the military-industry corpus to be recognized according to the byte information, and the byte sequence of the military-industry corpus to be recognized;
[0133] Step S122: Determine the segmentation path according to each byte and the byte sequence;
[0134] Step S123: Determine the segmentation strategy of the military-industry corpus to be recognized according to the segmentation path.
[0135] It should be noted that the byte can be each byte included in the text of the military-industry corpus to be recognized. The above byte sequence can be the arrangement order of each byte in the text. The above segmentation path can be the segmentation path of each byte in the text.
[0136] It should be understood that the vocabulary mining device traverses the text of the military industrial corpus to be recognized, determines each byte and the byte sequence contained in the text according to the traversal result, plans the path according to each byte and the byte sequence, obtains the segmentation path of the byte, and determines the segmentation strategy of the military industrial corpus to be recognized according to the segmentation path.
[0137] Further, in order to accurately plan the segmentation path, step S122 may include:
[0138] Step S1221: Obtain the target military industrial dictionary and the general word segmentation dictionary;
[0139] Step S1222: Determine the segmentation length of the military industrial corpus to be recognized according to the target military industrial dictionary and the general word segmentation dictionary;
[0140] Step S1223: Traverse each byte according to the byte sequence and the segmentation length;
[0141] Step S1224: Perform segmentation planning according to the traversal result to determine the segmentation path.
[0142] It should be noted that the target military industrial dictionary can be a professional term dictionary registered in the vocabulary mining device in the military industrial field, that is, the target military industrial dictionary can be a database for storing military industrial vocabulary; the target military industrial dictionary can also be a dictionary for registering new military industrial words. The vocabulary in the target military industrial dictionary can be obtained by the vocabulary mining device from an existing database in the military industrial professional field (such as a military encyclopedia thesaurus, etc.), or can be military industrial vocabulary mined by the vocabulary mining device in the past and saved in the dictionary. The format of the target military industrial dictionary can include the form of a triple of word, type, and weight. Among them, word is the vocabulary, type is the part of speech, and weight is the weight value. Among them, word is a required field, type may be multi-valued, and different values are separated by commas. Weight is of double type.
[0143] The above general word segmentation dictionary can be a general word segmentation dictionary, which can come from an open-source word segmentation tool. The dictionary format of the general word segmentation dictionary is the same as that of the target military industrial dictionary. After the general word segmentation dictionary is loaded, it can be saved in the memory for use. The above segmentation length can be the character length of the segmentation. For example, if the segmentation length is 2, then segmentation is performed in units of two characters. The segmentation length includes one or more types. For example, in the current segmentation process, several segmentation lengths can be used for segmentation. First, segmentation is performed with a segmentation length of 2, and then segmentation is performed with a segmentation length of 3, etc. This embodiment does not limit it.
[0144] It should be understood that the vocabulary mining device determines the segmentation length of the military-related corpus to be recognized according to the target military dictionary and the general word segmentation dictionary, and traverses each byte according to the segmentation length and the byte sequence, so as to realize the planning of the segmentation path.
[0145] For example, the vocabulary mining device determines that the segmentation lengths of the military-related corpus to be recognized according to the target military dictionary and the general word segmentation dictionary are 1, 2, and 3. The military-related corpus to be recognized contains bytes a, b, c, d, e, f, g, and h. Traverse each byte according to the byte sequence and the segmentation length, and perform segmentation planning according to the traversal result. The determined segmentation path is as follows, where n is the segmentation length:
[0146] When n = 1, the segmentation path is: [a, b, c, d, e, f, g, h];
[0147] When n = 2, the segmentation path is: [ab, bc, cd, de, ef, fg, gh];
[0148] When n = 3, the segmentation path is: [abc, bcd, cde, def, efg, fgh].
[0149] Step S13: Segment the military-related corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words.
[0150] It should be understood that the vocabulary mining device determines the byte segmentation path of the text of the military-related corpus to be recognized according to the segmentation strategy, and segments the military-related corpus to be recognized according to the segmentation path to obtain a plurality of initial words.
[0151] In this embodiment, by obtaining the byte information of the military-related corpus to be recognized, determining the segmentation strategy of the military-related corpus to be recognized according to the byte information, and segmenting the military-related corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words; since in this embodiment, the byte information of the military-related corpus to be recognized is obtained, and then the segmentation strategy is determined according to the byte information, and the vocabulary of the military-related corpus to be recognized is segmented according to the segmentation strategy, so as to obtain a plurality of initial words, thereby effectively extracting potential new words in the military-related corpus to be recognized, so as to facilitate subsequent screening and proofreading of the initial words, thereby improving the mining efficiency and accuracy of new words.
[0152] Reference Figure 4 , Figure 4 is a schematic flowchart of the third embodiment of a military vocabulary mining method of the present invention.
[0153] Based on the above first embodiment, in this embodiment, the step S30 includes:
[0154] Step S31: Obtain the feature vectors of each of the candidate words.
[0155] It should be noted that the feature vector can be the feature information of the candidate word. For example, the feature vector includes factors such as part of speech, length, preposition, postposition, term, word frequency, tf-idf weight, single-character (constituent word) word frequency, etc.
[0156] Furthermore, in order to accurately obtain the feature vector of the candidate word, the above step S31 may include:
[0157] Step S311: Traverse the to-be-recognized military industrial corpus based on each of the candidate words;
[0158] Step S312: Determine the word frequency and inverse document frequency index of each of the candidate words according to the traversal result;
[0159] Step S313: Perform weight analysis on each of the candidate words according to the word frequency and the inverse document frequency index to obtain the semantic weight of each of the candidate words;
[0160] Step S314: Determine the feature vector of each of the candidate words according to the word feature and the semantic weight of each of the candidate words.
[0161] It should be noted that the word frequency can be (TF, Term Frequency), that is, the frequency of the candidate word appearing in the to-be-recognized military industrial corpus. The above inverse document frequency index can be (Inverse Document Frequency), that is, a measure of the general importance of the candidate word in the to-be-recognized military industrial corpus. The above semantic weight can be the semantic importance degree of the candidate word in the to-be-recognized military industrial corpus.
[0162] It should be understood that in order to accurately determine whether a candidate word can become a new word, the vocabulary mining device can calculate the weight value of each candidate word in the to-be-recognized military industrial corpus through TF-IDF (term frequency–inverse document frequency) weight calculation, including TF weight and IDF weight, to measure the semantic importance degree of the candidate word.
[0163] Step S32: Input the feature vector into a pre-constructed classification model to obtain the new word probability of each of the candidate words.
[0164] It should be noted that the classification model can be a classification model pre-constructed by the vocabulary mining device with an SVM classifier.
[0165] It should be understood that the vocabulary mining device can use the SVM classifier as a classification model to extract feature vectors from the previously detected candidate words (new words), including factors such as part of speech, length, preposition, postposition, term, word frequency, tf-idf weight, single-character (constituent word) frequency, etc. A classifier is constructed to judge the probability that a candidate word is a new word or not through the classifier, and a threshold of the probability is set, and those greater than the threshold are judged as new words.
[0166] Step S33: Judge whether the new word probability of each of the candidate words is greater than a preset probability threshold.
[0167] It should be noted that the preset probability threshold can be a probability threshold preset by the vocabulary mining device for judging whether a candidate word is a new word. For example, if the preset probability threshold is 50%, and the new word probability of a candidate word is 45%, then this candidate word does not meet the new word probability requirement, and this candidate word is marked as an invalid old word.
[0168] Step S34: Screen out target words from each of the candidate words according to the judgment result.
[0169] In a specific implementation, the vocabulary mining device judges whether the new word probability of each candidate word is greater than the preset probability threshold, marks the candidate words with new word probabilities greater than the preset probability threshold as target words, and marks the candidate words with new word probabilities not greater than the preset probability threshold as invalid old words.
[0170] Furthermore, in order to improve the accuracy and performance of the classification model, after the above step S40, it further includes:
[0171] Synchronize the military new words to the target military dictionary to complete the update of the target military dictionary;
[0172] Generate a proofreading positive sample set and a proofreading negative sample set according to the proofreading result;
[0173] Iteratively train the classification model according to the proofreading positive sample set and the proofreading negative sample set.
[0174] It should be understood that after the vocabulary mining device confirms the accuracy of the new word, it automatically synchronizes the new word to the professional dictionary. After synchronization, on the one hand, it further improves the accuracy in the word segmentation stage, and on the other hand, in a new round of recognition process, it automatically ignores the already recognized new words to avoid repeated calculation. During the proofreading and recognition process, there will be two situations: the recognition result is judged to be accurate after proofreading, denoted as 1, and the recognition result is judged to be incorrect after proofreading, denoted as 0. In this way, for all the judgment results, two sets of positive samples and negative samples are formed, and the above classification model is retrained through the positive and negative samples to realize the automatic upgrade and iteration of the algorithm model and improve the system accuracy.
[0175] In this embodiment, by obtaining the feature vectors of each candidate word, inputting the feature vectors into a pre-constructed classification model to obtain the new word probabilities of each candidate word, determining whether the new word probabilities of each candidate word are greater than a preset probability threshold, and screening out target words from each candidate word according to the determination result; since in this embodiment, by inputting the feature vectors of each candidate word into a pre-constructed classification model, the new word probabilities of each candidate word are obtained, and by determining whether the new word probabilities of each candidate word are greater than a preset probability threshold, it is determined whether each candidate word is a new word, thereby effectively improving the accuracy of new word screening.
[0176] In addition, an embodiment of the present invention further provides a storage medium, on which a military vocabulary mining program is stored. When the military vocabulary mining program is executed by a processor, the steps of the military vocabulary mining method described above are implemented.
[0177] Since this storage medium adopts all the technical solutions of the above-mentioned all embodiments, it has at least all the beneficial effects brought by the technical solutions of the above-mentioned embodiments, and will not be elaborated here one by one.
[0178] Refer to Figure 5 , Figure 5 which is the structural block diagram of the first embodiment of the military vocabulary mining device of the present invention.
[0179] As Figure 5 shown, the military vocabulary mining device proposed by the embodiment of the present invention includes:
[0180] A word segmentation module 10, configured to perform word segmentation on the military corpus to be recognized to obtain a plurality of initial words;
[0181] A rule matching module 20, configured to respectively recognize the plurality of initial words according to the rule library corresponding to the target military dictionary, and screen out candidate words that conform to each rule in the rule library from the plurality of initial words according to the recognition result;
[0182] A probability recognition module 30, configured to obtain the new word probabilities of each candidate word, and screen out target words from each candidate word according to the new word probabilities;
[0183] A new word proofreading module 40, configured to proofread each target word, and determine the military new words in each target word according to the proofreading result.
[0184] Further, the word segmentation module 10 is further configured to obtain the byte information of the military corpus to be recognized; determine the segmentation strategy of the military corpus to be recognized according to the byte information; and perform word segmentation on the military corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words.
[0185] Further, the word segmentation module 10 is further configured to determine each byte included in the military industrial corpus to be recognized and the byte sequence of the military industrial corpus to be recognized according to the byte information; determine a segmentation path according to each byte and the byte sequence; and determine a segmentation strategy for the military industrial corpus to be recognized according to the segmentation path.
[0186] Further, the word segmentation module 10 is further configured to obtain a target military industrial dictionary and a general word segmentation dictionary; determine the segmentation length of the military industrial corpus to be recognized according to the target military industrial dictionary and the general word segmentation dictionary; traverse each byte according to the byte sequence and the segmentation length; and perform segmentation planning according to the traversal result to determine a segmentation path.
[0187] Further, the word segmentation module 10 is further configured to obtain a target military industrial dictionary and a general word segmentation dictionary; perform word segmentation on the military industrial corpus to be recognized according to the target military industrial dictionary and the general word segmentation dictionary to obtain original words; perform string matching on the original words according to a stop word dictionary to determine the stop words to be removed in the original words; and remove the stop words to be removed from the original words to obtain a plurality of initial words.
[0188] Further, the rule matching module 20 is further configured to obtain the character order of the plurality of initial words; combine the plurality of initial words according to the character order to obtain multiple groups of initial phrases; obtain a rule library corresponding to the target military industrial dictionary; respectively identify the multiple groups of initial phrases according to each rule in the rule library, and determine whether the multiple groups of initial phrases match each rule according to the identification result; screen out candidate phrases that conform to each rule in the rule library from the plurality of initial words according to the identification result, and determine candidate words according to the candidate phrases.
[0189] Further, the probability recognition module 30 is further configured to obtain the feature vectors of the candidate words; input the feature vectors into a pre-constructed classification model to obtain the new word probabilities of the candidate words; determine whether the new word probabilities of the candidate words are greater than a preset probability threshold; and screen out target words from the candidate words according to the judgment result.
[0190] Further, the probability recognition module 30 is further configured to traverse the military industrial corpus to be recognized based on the candidate words; determine the word frequency and inverse document frequency index of each candidate word according to the traversal result; perform weight analysis on each candidate word according to the word frequency and the inverse document frequency index to obtain the semantic weights of the candidate words; and determine the feature vectors of the candidate words according to the word features and the semantic weights of the candidate words.
[0191] Further, the new word proofreading module 40 is further configured to traverse the text corresponding to the to-be-recognized military industrial corpus according to each of the target words; locate the target new words in the text according to the traversal result, and mark the located target new words; proofread each of the target words based on the marked text; and determine the military industrial new words in each of the target words according to the proofreading result.
[0192] Further, the military industrial vocabulary mining device further includes:
[0193] A model training module 50, configured to synchronize the military industrial new words to the target military industrial dictionary to complete the update of the target military industrial dictionary; generate a proofreading positive sample set and a proofreading negative sample set according to the proofreading result; and perform iterative training on the classification model according to the proofreading positive sample set and the proofreading negative sample set.
[0194] In this embodiment, by segmenting the to-be-recognized military industrial corpus, a plurality of initial words are obtained, each of the plurality of initial words is recognized according to the rule base corresponding to the target military industrial dictionary, and candidate words that meet each rule in the rule base are screened out from the plurality of initial words according to the recognition result. The new word probability of each candidate word is obtained, and target words are screened out from each candidate word according to the new word probability. Each of the target words is proofread, and the military industrial new words in each of the target words are determined according to the proofreading result. Since in this embodiment, by segmenting the to-be-recognized military industrial corpus, a plurality of initial words are obtained, the lexical segmentation of the to-be-recognized military industrial corpus is completed. Each of the plurality of initial words is recognized according to the rule base corresponding to the target military industrial dictionary, and candidate words that meet each rule in the rule base are screened out from the plurality of initial words according to the recognition result, thereby realizing rule matching for the segmented initial vocabulary to retain candidate words that meet each rule in the rule base of the target military industrial dictionary. Then, screening is performed according to the new word probability of each candidate word, so as to effectively eliminate invalid old words. Each of the target words is proofread, and the military industrial new words in each of the target words are determined according to the proofreading result, so as to ensure that the screened target words are all new words, realizing accurate mining of new words in the military industrial field and effectively improving the efficiency of military industrial vocabulary mining.
[0195] It should be understood that the above is only an example for illustration, and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not make any limitation thereto.
[0196] It should be noted that the above-described work process is only illustrative and does not constitute a limitation to the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.
[0197] In addition, for the technical details not described in detail in this embodiment, reference may be made to the military vocabulary mining method provided in any embodiment of the present invention, which will not be elaborated here.
[0198] The present invention provides A1. A military vocabulary mining method, the military vocabulary mining method includes:
[0199] Performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words;
[0200] Identifying the plurality of initial words respectively according to the rule base corresponding to the target military dictionary, and screening out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results;
[0201] Obtaining the new word probabilities of the candidate words, and screening out target words from the candidate words according to the new word probabilities;
[0202] Proofreading each of the target words, and determining military new words in each of the target words according to the proofreading results.
[0203] A2. The military vocabulary mining method as described in A1, where performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words includes:
[0204] Obtaining the byte information of the military corpus to be recognized;
[0205] Determining the segmentation strategy of the military corpus to be recognized according to the byte information;
[0206] Performing word segmentation on the military corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words.
[0207] A3. The military vocabulary mining method as described in A2, where receiving the message information sent by the target communication program through the interaction interface and visually displaying the message information includes:
[0208] Receiving the message information sent by the target communication program through the interaction interface;
[0209] Screening the message information to obtain the unread message information in the message information;
[0210] Determining the unread message display strategy corresponding to the target communication program according to the instant messaging configuration information;
[0211] Visually displaying the unread message information according to the unread message display strategy.
[0212] A4. The military vocabulary mining method as described in A3, where determining the segmentation path according to each of the bytes and the byte sequence includes:
[0213] Obtain the target military-industry dictionary and the general word segmentation dictionary;
[0214] Determine the segmentation length of the military-industry corpus to be recognized according to the target military-industry dictionary and the general word segmentation dictionary;
[0215] Traverse each byte according to the byte sequence and the segmentation length;
[0216] Perform segmentation planning based on the traversal result to determine the segmentation path.
[0217] A5. The military-industry vocabulary mining method as described in A1, wherein segmenting the military-industry corpus to be recognized to obtain a plurality of initial words, including:
[0218] Obtain the target military-industry dictionary and the general word segmentation dictionary;
[0219] Segment the military-industry corpus to be recognized according to the target military-industry dictionary and the general word segmentation dictionary to obtain original words;
[0220] Perform string matching on the original words according to the stop word dictionary to determine the stop words in the original words;
[0221] Remove the stop words from the original words to obtain a plurality of initial words.
[0222] A6. The military-industry vocabulary mining method as described in any one of A1 to A5, wherein respectively identifying the plurality of initial words according to the rule base corresponding to the target military-industry dictionary, and screening candidate words that conform to each rule in the rule base from the plurality of initial words according to the identification result, including:
[0223] Obtain the character order of the plurality of initial words;
[0224] Combine the plurality of initial words according to the character order to obtain multiple groups of initial phrases;
[0225] Obtain the rule base corresponding to the target military-industry dictionary;
[0226] Respectively identify the multiple groups of initial phrases according to each rule in the rule base, and judge whether the multiple groups of initial phrases match each rule according to the identification result;
[0227] Screen candidate phrases that conform to each rule in the rule base from the plurality of initial words according to the identification result, and determine candidate words according to the candidate phrases.
[0228] A7. The military-industry vocabulary mining method as described in any one of A1 to A5, wherein obtain the new word probability of each candidate word, and screen target words from each candidate word according to the new word probability, including:
[0229] Obtain the feature vectors of each of the candidate words;
[0230] Input the feature vectors into a pre-constructed classification model to obtain the new word probabilities of each of the candidate words;
[0231] Determine whether the new word probabilities of each of the candidate words are greater than a preset probability threshold;
[0232] Screen out target words from each of the candidate words according to the judgment result.
[0233] A8. The military-industry vocabulary mining method as described in A7, wherein the obtaining the feature vectors of each of the candidate words includes:
[0234] Traverse the to-be-recognized military-industry corpus based on each of the candidate words;
[0235] Determine the word frequency and inverse document frequency index of each of the candidate words according to the traversal result;
[0236] Conduct weight analysis on each of the candidate words according to the word frequency and the inverse document frequency index to obtain the semantic weights of each of the candidate words;
[0237] Determine the feature vectors of each of the candidate words according to the word features of each of the candidate words and the semantic weights.
[0238] A9. The military-industry vocabulary mining method as described in A8, wherein the proofreading each of the target words and determining the military-industry new words in each of the target words according to the proofreading result includes:
[0239] Traverse the text corresponding to the to-be-recognized military-industry corpus based on each of the target words;
[0240] Locate the target new words in the text according to the traversal result, and mark the located target new words;
[0241] Proofread each of the target words based on the marked text;
[0242] Determine the military-industry new words in each of the target words according to the proofreading result.
[0243] A10. The military-industry vocabulary mining method as described in A9, after the proofreading each of the target words and determining the military-industry new words in each of the target words according to the proofreading result, includes:
[0244] Synchronize the military-industry new words to the target military-industry dictionary to complete the update of the target military-industry dictionary;
[0245] Generate a proofreading positive sample set and a proofreading negative sample set according to the proofreading result;
[0246] Iteratively train the classification model according to the corrected positive sample set and the corrected negative sample set.
[0247] The present invention also discloses B11, a military vocabulary mining device, which includes:
[0248] A word segmentation module for segmenting the to-be-recognized military corpus to obtain a plurality of initial words;
[0249] A rule matching module for respectively recognizing the plurality of initial words according to the rule base corresponding to the target military dictionary, and screening out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results;
[0250] A probability recognition module for obtaining the new word probabilities of the candidate words, and screening out target words from the candidate words according to the new word probabilities;
[0251] A new word proofreading module for proofreading each of the target words and determining the military new words among the target words according to the proofreading results.
[0252] B12. The military vocabulary mining device according to B11, wherein the word segmentation module is further configured to obtain byte information of the to-be-recognized military corpus; determine a segmentation strategy for the to-be-recognized military corpus according to the byte information; and segment the to-be-recognized military corpus according to the segmentation strategy to obtain a plurality of initial words.
[0253] B13. The military vocabulary mining device according to B12, wherein the word segmentation module is further configured to determine each byte included in the to-be-recognized military corpus and the byte sequence of the to-be-recognized military corpus according to the byte information; determine a segmentation path according to each byte and the byte sequence; and determine the segmentation strategy for the to-be-recognized military corpus according to the segmentation path.
[0254] B14. The military vocabulary mining device according to B13, wherein the word segmentation module is further configured to obtain a target military dictionary and a general word segmentation dictionary; determine a segmentation length of the to-be-recognized military corpus according to the target military dictionary and the general word segmentation dictionary; traverse each byte according to the byte sequence and the segmentation length; and perform segmentation planning according to the traversal result to determine a segmentation path.
[0255] B15. The military vocabulary mining device according to B11, wherein the word segmentation module is further configured to obtain a target military dictionary and a general word segmentation dictionary; segment the to-be-recognized military corpus according to the target military dictionary and the general word segmentation dictionary to obtain original words; perform string matching on the original words according to a stop word dictionary to determine the stop words in the original words; and remove the stop words from the original words to obtain a plurality of initial words.
[0256] B16. The military-industry vocabulary mining device according to any one of B11 to B15, wherein the rule matching module is further configured to obtain the character order of the multiple initial words; combine the multiple initial words according to the character order to obtain multiple groups of initial phrases; obtain the rule base corresponding to the target military-industry dictionary; respectively identify the multiple groups of initial phrases according to each rule in the rule base, and determine whether the multiple groups of initial phrases match each of the rules according to the identification results; screen out candidate phrases that conform to each rule in the rule base from the multiple initial words according to the identification results, and determine candidate words according to the candidate phrases.
[0257] B17. The military-industry vocabulary mining device according to any one of B11 to B15, wherein the probability identification module is further configured to obtain the feature vectors of the candidate words; input the feature vectors into a pre-constructed classification model to obtain the new word probabilities of the candidate words; determine whether the new word probabilities of the candidate words are greater than a preset probability threshold; screen out target words from the candidate words according to the judgment results.
[0258] B18. The military-industry vocabulary mining device according to B17, wherein the probability identification module is further configured to traverse the military-industry corpus to be identified based on the candidate words; determine the word frequency and inverse document frequency index of each candidate word according to the traversal results; perform weight analysis on each candidate word according to the word frequency and the inverse document frequency index to obtain the semantic weights of the candidate words; determine the feature vectors of the candidate words according to the word features and the semantic weights of the candidate words.
[0259] The present invention also discloses C19. A military-industry vocabulary mining device, the device comprising: a memory, a processor, and a military-industry vocabulary mining program stored on the memory and executable on the processor, the military-industry vocabulary mining program being configured to implement the military-industry vocabulary mining method as described above.
[0260] The present invention also discloses D20. A storage medium, on which a military-industry vocabulary mining program is stored, and when the military-industry vocabulary mining program is executed by a processor, it implements the military-industry vocabulary mining method as described above.
[0261] In addition, it should be noted that in this article, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including the element.
[0262] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0263] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM) / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0264] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for mining military industry vocabulary, characterized in that, The military vocabulary mining method described above includes: Performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words; Respectively recognizing the plurality of initial words according to the rule base corresponding to the target military dictionary, and screening out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results; Obtaining the new word probabilities of each of the candidate words, and screening out target words from each of the candidate words according to the new word probabilities; Proofreading each of the target words, and determining the military new words in each of the target words according to the proofreading results.
2. The method for mining military industry vocabulary according to claim 1, characterized in that, The performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words includes: Obtaining the byte information of the military corpus to be recognized; Determining the segmentation strategy of the military corpus to be recognized according to the byte information; Performing word segmentation on the military corpus to be recognized according to the segmentation strategy to obtain a plurality of initial words.
3. The method for mining military industry vocabulary according to claim 2, characterized in that, The determining the segmentation strategy of the military corpus to be recognized according to the byte information includes: Determining each byte included in the military corpus to be recognized according to the byte information, and the byte sequence of the military corpus to be recognized; Determining the segmentation path according to each of the bytes and the byte sequence; Determining the segmentation strategy of the military corpus to be recognized according to the segmentation path.
4. The method for mining military industry vocabulary according to claim 3, characterized in that, The determining the segmentation path according to each of the bytes and the byte sequence includes: Obtaining the target military dictionary and the general word segmentation dictionary; Determining the segmentation length of the military corpus to be recognized according to the target military dictionary and the general word segmentation dictionary; Traversing each of the bytes according to the byte sequence and the segmentation length; Performing segmentation planning according to the traversal results to determine the segmentation path.
5. The method for mining military industry vocabulary according to claim 1, characterized in that, The performing word segmentation on the military corpus to be recognized to obtain a plurality of initial words includes: Obtaining the target military dictionary and the general word segmentation dictionary; Performing word segmentation on the military corpus to be recognized according to the target military dictionary and the general word segmentation dictionary to obtain the original words; Performing string matching on the original words according to the stop word dictionary to determine the stop words to be removed in the original words; Removing the stop words to be removed from the original words to obtain a plurality of initial words.
6. The method for mining military industry vocabulary according to any one of claims 1 to 5, characterized in that, The respectively recognizing the plurality of initial words according to the rule base corresponding to the target military dictionary, and screening out candidate words that conform to each rule in the rule base from the plurality of initial words according to the recognition results includes: Obtaining the character order of the plurality of initial words; Combining the plurality of initial words according to the character order to obtain multiple groups of initial phrases; Obtaining the rule base corresponding to the target military dictionary; Respectively recognizing the multiple groups of initial phrases according to each rule in the rule base, and determining whether the multiple groups of initial phrases match each of the rules according to the recognition results; Screening out candidate phrases that conform to each rule in the rule base from the plurality of initial words according to the recognition results, and determining candidate words according to the candidate phrases.
7. The method for mining military industry vocabulary according to any one of claims 1 to 5, characterized in that, The obtaining the new word probabilities of each of the candidate words, and screening out target words from each of the candidate words according to the new word probabilities includes: Obtaining the feature vectors of each of the candidate words; Inputting the feature vectors into a pre-constructed classification model to obtain the new word probabilities of each of the candidate words; Judging whether the new word probabilities of each of the candidate words are greater than a preset probability threshold; Select target words from each of the candidate words according to the judgment result.
8. A military vocabulary mining device, characterized in that, The military vocabulary mining device includes: A word segmentation module, configured to perform word segmentation on the military corpus to be recognized to obtain a plurality of initial words; A rule matching module, configured to respectively recognize the plurality of initial words according to the rule base corresponding to the target military dictionary, and screen out candidate words that conform to each rule in the rule base from the plurality of initial words; A probability recognition module, configured to obtain the new word probability of each of the candidate words, and screen out target words from each of the candidate words according to the new word probability; A new word proofreading module, configured to proofread each of the target words, and determine military new words among each of the target words according to the proofreading result.
9. A military vocabulary mining equipment, characterized in that, The military vocabulary mining device includes: a memory, a processor, and a military vocabulary mining program stored on the memory and executable on the processor, where the military vocabulary mining program is configured to implement the military vocabulary mining method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, A military vocabulary mining program is stored on the storage medium, and when the military vocabulary mining program is executed by a processor, it implements the military vocabulary mining method according to any one of claims 1 to 7.