Unregistered word vector calculation method, system, electronic device and storage medium

By acquiring and processing prior knowledge and calculating the word vectors of unlogged words, the problem that the existing word embedding model cannot handle unlogged words is solved, and efficient unlogged words representation learning is achieved.

CN113255326BActive Publication Date: 2025-06-06BEIJING XUEZHITU NETWORK TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110539232.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-18
Publication Date
2025-06-06
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

Existing word embedding models cannot learn word embedding representations of unlogged words, and as text corpus increases, computing power and hardware costs also increase, which may lead to overflow.

Method used

By obtaining prior knowledge, including dictionary, unlabeled text corpus and unlogged words, corpus preprocessing is carried out, co-occurrence data of Chinese characters is counted, entropy data is calculated, and finally the word vector of unlogged words is calculated based on the entropy data.

Benefits of technology

The problem that the pre-trained model cannot handle the unlogged words is solved. The word-composition contribution of Chinese characters to unlogged words is measured by information entropy, and the effective word vector of the unlogged words is calculated, which reduces the calculation cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113255326B_ABST
    Figure CN113255326B_ABST
Patent Text Reader

Abstract

The present invention proposes a method, system, electronic device and storage medium for calculating word vectors of unregistered words. The method and technical solution include a prior knowledge acquisition step, which acquires a prior knowledge for pre-training, wherein the prior knowledge includes a dictionary, an unlabeled text corpus and unregistered words; a corpus preprocessing step, which preprocesses the unlabeled text corpus; a character co-occurrence statistics step, which counts the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus; a character entropy data calculation step, which calculates the entropy data of the Chinese character based on the co-occurrence data; and a word vector calculation step, which calculates the word formation contribution of the Chinese character to the unregistered word based on the entropy data, and calculates the word vector of the unregistered word based on the word formation contribution. The present invention solves the problem that the existing pre-training model cannot process unregistered words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method, system, electronic device and storage medium for calculating word vectors of unregistered words. Background Art

[0002] Word embedding models have become the most important technical means of text representation in the field of natural language processing (NLP). With the rise of deep learning (DL), since word2vec, GloVe, Elmo, Bert, GPT-1, GPT-2, and GPT3 have pushed word embedding to the peak. These word embedding models are based on large-scale text corpora collected, transforming various deep neural networks, learning word embedding, and have also achieved good responses in various fields.

[0003] However, these methods can only learn word embeddings for known words, and cannot learn word embedding representations for unknown words. Although the more text corpus collected, the lower the probability of unknown words, this method still cannot solve the problem of word embedding learning for unknown words, but only avoids it as much as possible. However, as the text corpus increases, the cost of computing power, hardware equipment, memory, time, etc. also increases, and even overflows may occur, making it impossible to train word embeddings. Summary of the invention

[0004] The embodiments of the present application provide a method, system, electronic device and storage medium for calculating word vectors of unregistered words, so as to at least solve the problem that the existing pre-trained model cannot process unregistered words.

[0005] In a first aspect, an embodiment of the present application provides a method for calculating a word vector of an unregistered word, comprising: a prior knowledge acquisition step, acquiring a prior knowledge for pre-training, wherein the prior knowledge includes a dictionary, an unlabeled text corpus and an unregistered word; a corpus preprocessing step, preprocessing the unlabeled text corpus; a character co-occurrence statistics step, counting the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus; a character entropy data calculation step, calculating the entropy data of the Chinese character based on the co-occurrence data; a word vector calculation step, calculating the word formation contribution of the Chinese character to the unregistered word based on the entropy data, and calculating the word vector of the unregistered word based on the word formation contribution.

[0006] Preferably, the character co-occurrence counting step further includes: a character occurrence counting step, counting the number of times the Chinese character appears in the unlabeled text corpus; a left and right neighbor character acquisition step, acquiring the left neighbor character and the right neighbor character that co-occur on the left and right sides of the Chinese character in the unlabeled text corpus; a co-occurrence count counting step, acquiring the number of times the Chinese character co-occurs with the left neighbor character and the right neighbor character.

[0007] Preferably, the character entropy data calculation step further includes: an information entropy calculation step, calculating the left information entropy and the right information entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0008] Preferably, the character entropy data calculation step further includes: a conditional entropy calculation step, calculating the left conditional entropy and the right conditional entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0009] Preferably, the word vector calculation step further includes: calculating the word formation contribution of the left neighboring character to the left neighboring character of the unregistered word according to the left information entropy and the left conditional entropy, and calculating the word formation contribution of the right neighboring character to the right neighboring character of the unregistered word according to the right information entropy and the right conditional entropy.

[0010] Preferably, the word vector calculation step further includes: normalizing the word formation contribution of the left neighbor character and the word formation contribution of the right neighbor character, and calculating the word formation contribution of the Chinese character to the unregistered word based on the normalized word formation contribution of the left neighbor character and the word formation contribution of the right neighbor character.

[0011] Preferably, the word vector calculation step further includes: calculating the word vector of the unregistered word based on the word formation contribution of the Chinese character to the unregistered word and the character vector of the Chinese character.

[0012] In a second aspect, an embodiment of the present application provides a system for calculating word vectors of unregistered words, which is applicable to the above-mentioned method for calculating word vectors of unregistered words, and includes: a prior knowledge acquisition module, which acquires prior knowledge for pre-training, wherein the prior knowledge includes a dictionary, an unlabeled text corpus and unregistered words; a corpus preprocessing module, which preprocesses the unlabeled text corpus; a text co-occurrence statistics module, which counts the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus; a character entropy data calculation module, which calculates the entropy data of the Chinese character based on the co-occurrence data; a word vector calculation module, which calculates the word formation contribution of the Chinese character to the unregistered word based on the entropy data, and calculates the word vector of the unregistered word based on the word formation contribution.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for calculating word vectors of unregistered words as described in the first aspect above is implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for calculating word vectors of unregistered words as described in the first aspect above.

[0015] The present invention can be applied to the field of deep learning technology. Compared with the related art, the embodiment of the present application measures the contribution of Chinese characters to the word formation of unregistered words based on information entropy, and calculates the word vector of unregistered words from the character vector, thereby solving the problem that the existing pre-trained word vector model cannot process unregistered words, that is, solving the problem that the pre-trained model word vector model cannot perform embedding representation on unregistered words. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 This is a flow chart of the method for calculating the word vector of an unregistered word of the present invention;

[0018] Figure 2 for Figure 1 A step-by-step flow chart of step S3;

[0019] Figure 3 for Figure 1 The step-by-step flow chart of step S4;

[0020] Figure 4 It is a framework diagram of the unregistered word vector calculation system of the present invention;

[0021] Figure 5 A framework diagram of an electronic device of the present invention;

[0022] In the above picture:

[0023] 1. Prior knowledge acquisition module; 2. Corpus preprocessing module; 3. Word co-occurrence statistics module; 4. Word entropy data calculation module; 5. Corpus preprocessing module; 31. Word occurrence statistics unit; 32. Left and right neighbor word acquisition unit; 33. Co-occurrence times statistics unit; 41. Information entropy calculation unit; 42. Conditional entropy calculation unit; 60. Bus; 61. Processor; 62. Memory; 63. Communication interface. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0025] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.

[0026] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0027] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a kind of", "the" and the like involved in this application do not indicate a quantitative limitation and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices.

[0028] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings:

[0029] Figure 1 This is a flow chart of the method for calculating the word vector of an unregistered word in the present invention. Figure 1 The method for calculating the word vector of an unregistered word of the present invention comprises the following steps:

[0030] S1: Acquire prior knowledge for pre-training, wherein the prior knowledge includes a dictionary, an unlabeled text corpus, and unregistered words.

[0031] In the specific implementation, prior knowledge such as pre-trained word vectors, dictionaries, unlabeled text corpora and unregistered words are obtained, and the words and their corresponding vectors are stored in a hashmap.

[0032] S2: Preprocessing the unlabeled text corpus.

[0033] In the specific implementation, the obtained unlabeled corpus is preprocessed, including segmentation, sentence segmentation, word segmentation, removal of duplicate and redundant symbols, etc.

[0034] S3: Counting the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus.

[0035] In a specific implementation, the number of times a word appears in a text corpus, the preceding word and the succeeding word with which it co-occurs, and the number of times it co-occurs are counted.

[0036] Optional, Figure 2 for Figure 1 For a step-by-step flowchart of step S3 in Figure 2 :

[0037] S31: Counting the number of times the Chinese character appears in the unlabeled text corpus;

[0038] In a specific implementation, the number of times the Chinese characters in the dictionary appear in the unlabeled text corpus is counted.

[0039] S32: Obtaining the left neighboring characters and the right neighboring characters that co-occur on the left and right sides of the Chinese character in the unlabeled text corpus;

[0040] In a specific implementation, the Chinese characters in the dictionary that co-occur on the left and right in the unlabeled text corpus are counted and recorded as left neighbor characters and right neighbor characters respectively.

[0041] This application provides a specific embodiment for further explanation:

[0042] In the sentence "Xiao Ming is going to participate in Tomorrow's Star tomorrow", the left neighbors of the Chinese character "Ming" are "Xiao, Ming, Jia", and the right neighbors of the Chinese character "Ming" are "Ming, Tian, ​​Ri".

[0043] S33: Obtain the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0044] In a specific implementation, the number of times a Chinese character in a dictionary co-occurs with its left neighboring character and its right neighboring character in an unlabeled text corpus is counted.

[0045] Please continue to see Figure 1 :

[0046] S4: Calculate the entropy data of the Chinese characters according to the co-occurrence data; Optionally, Figure 3 for Figure 1 For a step-by-step flowchart of step S4 in Figure 3 :

[0047] S41: Calculate the left information entropy and the right information entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0048] S42: Calculate the left conditional entropy and the right conditional entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0049] In a specific implementation, the left information entropy and the right information entropy of a character are calculated according to the number of times the character appears in the text corpus and the number of times the character co-occurs with its left and right neighboring characters.

[0050] In the specific implementation, the embodiment of the present application is to calculate the Chinese character w i The left information entropy and right information entropy are described as an example:

[0051] Chinese character w i The calculation method of the left information entropy is:

[0052]

[0053] In the formula, f(w i ) means that in the text corpus, word w i The set of left neighboring words. k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0054] This application provides a specific embodiment for further explanation:

[0055] In the sentence “Xiao Ming is going to participate in the Tomorrow’s Star tomorrow”, f(Ming) = {Xiao, Ming, Jia}, P(Xiao|Ming) = 0.33, P(Ming|Ming) = 0.33, P(Jia|Ming) = 0.33.

[0056] Chinese character w i The calculation method of the right information entropy is:

[0057]

[0058] In the formula, g(w i) means that in the text corpus, word w i The set of right neighboring words. Here P(w k |w i ) represents the word w in the text corpus i The next character to the right is w k probability.

[0059] This application provides a specific embodiment for further explanation:

[0060] In the sentence “Xiao Ming is going to participate in Tomorrow’s Star tomorrow”, g(Ming) = {Ming, Tian, ​​Ri}, P(Ming|Ming) = 0.33, P(Tian|Ming) = 0.33, P(Ri|Ming) = 0.33.

[0061] In a specific implementation, the conditional entropy of two adjacent characters in the unregistered word is calculated. The left conditional entropy and the right conditional entropy of the character are calculated according to the number of times the character appears in the text corpus and the number of times the character co-occurs with the left and right adjacent characters.

[0062] In the specific implementation, the embodiment of the present application is to calculate the Chinese character w i The left neighbor word w k The conditional entropy of is described as an example:

[0063] With Chinese character w i The Chinese character w for the character to the right k The calculation method of the left conditional entropy is:

[0064] H left (w i , w k )=E[-logP(w k |w i )]

[0065] =-P(w k |w i )logP(w k |w i )

[0066] In the formula, P(w k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0067] With Chinese character w i The Chinese character w for the left neighbor k The calculation method of the right conditional entropy is:

[0068] H right (w i , w k )=E[-logP(w k |wi )]

[0069] =-P(w k |w i )logP(w k |w i )

[0070] Among them, P(w k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0071] Please continue to see Figure 1 :

[0072] S5: Calculate the word-forming contribution of the Chinese character to the unregistered word according to the entropy data, and calculate the word vector of the unregistered word according to the word-forming contribution.

[0073] Optionally, the contribution of the left neighboring character to the word-forming of the unregistered word is calculated based on the left information entropy and the left conditional entropy, and the contribution of the right neighboring character to the word-forming of the right neighboring character of the unregistered word is calculated based on the right information entropy and the right conditional entropy.

[0074] Optionally, the word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character are normalized, and the word-forming contribution of the Chinese character to the unregistered word is calculated based on the normalized word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character.

[0075] Optionally, the word vector of the unregistered word is calculated based on the word-forming contribution of the Chinese character to the unregistered word and the word vector of the Chinese character.

[0076] In a specific implementation, the contribution of a character in an unregistered word to the word formation of the unregistered word is directly proportional to the conditional entropy of its adjacent left and right neighboring characters.

[0077] In the specific implementation, the embodiment of the present application is to calculate the Chinese character w i and its adjacent left neighbor w k and the right neighbor word w j Let’s take the contribution of the unregistered word t as an example to describe it:

[0078] w i With the left adjacent word w k The contribution calculation method in the process of word formation of unregistered words is:

[0079]

[0080] w i With the right neighbor word w jThe calculation method of the contribution made during the word formation of out-of-vocabulary words is as follows:

[0081]

[0082] Chinese character w i The contribution to the word formation during the word formation of out-of-vocabulary word t is:

[0083] R(w i , t) = R left (w i , w k ) + R right (w i , w j )

[0084] Normalize the contribution of the left adjacent character to word formation and the contribution of the right adjacent character to word formation:

[0085]

[0086]

[0087] The contribution calculation method of the normalized Chinese character w i during the word formation of out-of-vocabulary word t is:

[0088]

[0089] In the formula, t[n] represents the nth Chinese character in the out-of-vocabulary word t, and |t| represents the number of Chinese characters contained in the out-of-vocabulary word t.

[0090] This application provides a specific embodiment for further illustration:

[0091] In the out-of-vocabulary word "Mindray Technology", the calculation method of the contribution of the Chinese character "Tech" to word formation is as follows:

[0092] weight(Tech, Mindray Technology)

[0093] = [σ left (Tech,略) + σ right (Tech,技)] / [σ right (明,略) + σ left (明,略) + σ right (略,Tech) + σ left (Tech,略) + σ right (Tech,技) + σ left (Tech,技)]

[0094] In specific implementation, calculate the word vector of the out-of-vocabulary word according to the contribution to word formation:

[0095]

[0096] In the formula, VT(t) represents the word vector of the unregistered word t, VW(w i ) represents the Chinese character w i The word vector of .

[0097] This application provides a specific embodiment for further explanation:

[0098] The word vector calculation method for the unregistered word "Minglue Technology" is as follows:

[0099] VT(Venture Technology)=

[0100] weight(Ming, Minglue Technology)*VW(Ming)+weight(Lue, Minglue Technology)*VW(Lue)+weight(Ke, Minglue Technology)*VW(Ke)+weight(Ji, Minglue Technology)*VW(Ji)

[0101] Figure 4 For a framework diagram of the unregistered word vector calculation system according to the present invention, see Figure 4 ,include:

[0102] Prior knowledge acquisition module 1: Acquire prior knowledge for pre-training, wherein the prior knowledge includes a dictionary, an unlabeled text corpus, and unregistered words.

[0103] In the specific implementation, prior knowledge such as pre-trained word vectors, dictionaries, unlabeled text corpora and unregistered words are obtained, and the words and their corresponding vectors are stored in a hashmap.

[0104] Corpus preprocessing module 2: preprocessing the unlabeled text corpus.

[0105] In the specific implementation, the obtained unlabeled corpus is preprocessed, including segmentation, sentence segmentation, word segmentation, removal of duplicate and redundant symbols, etc.

[0106] Character co-occurrence statistics module 3: counts the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus.

[0107] In a specific implementation, the number of times a word appears in a text corpus, the preceding word and the succeeding word with which it co-occurs, and the number of times it co-occurs are counted.

[0108] Optionally, the corpus preprocessing module 2 further includes:

[0109] Character occurrence counting unit 31: counting the number of times the Chinese character appears in the unlabeled text corpus;

[0110] In a specific implementation, the number of times the Chinese characters in the dictionary appear in the unlabeled text corpus is counted.

[0111] The left and right neighbor character acquisition unit 32 is used to acquire the left and right neighbor characters of the Chinese character that co-occur on the left and right sides in the unlabeled text corpus;

[0112] In a specific implementation, the Chinese characters in the dictionary that co-occur on the left and right in the unlabeled text corpus are counted and recorded as left neighbor characters and right neighbor characters respectively.

[0113] The co-occurrence number counting unit 33 is used to obtain the number of co-occurrences of the Chinese character with the left neighboring character and the right neighboring character.

[0114] In a specific implementation, the number of times a Chinese character in a dictionary co-occurs with its left neighboring character and its right neighboring character in an unlabeled text corpus is counted.

[0115] Character entropy data calculation module 4: calculates the entropy data of the Chinese characters according to the co-occurrence data; optionally, the character entropy data calculation module 4 also includes:

[0116] The information entropy calculation unit 41 calculates the left information entropy and the right information entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0117] The conditional entropy calculation unit 42 calculates the left conditional entropy and the right conditional entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character.

[0118] In a specific implementation, the left information entropy and the right information entropy of a character are calculated according to the number of times the character appears in the text corpus and the number of times the character co-occurs with its left and right neighboring characters.

[0119] In the specific implementation, the embodiment of the present application calculates the Chinese character w i The left information entropy and right information entropy are described as an example:

[0120] Chinese character w i The calculation method of the left information entropy is:

[0121]

[0122] In the formula, f(w i ) means that in the text corpus, word w i The set of left neighboring words. k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0123] Chinese character w i The calculation method of the right information entropy is:

[0124]

[0125] In the formula, g(w i ) means that in the text corpus, word w i The set of right neighboring words. Here P(w k |w i ) represents the word w in the text corpus i The next character to the right is w k probability.

[0126] In a specific implementation, the conditional entropy of two adjacent characters in the unregistered word is calculated. The left conditional entropy and the right conditional entropy of the character are calculated according to the number of times the character appears in the text corpus and the number of times the character co-occurs with the left and right adjacent characters.

[0127] In the specific implementation, the embodiment of the present application calculates the Chinese character w i The left neighbor word w k The conditional entropy of is described as an example:

[0128] With Chinese character w i The Chinese character w for the character to the right k The calculation method of the left conditional entropy is:

[0129] H left (w i , w k )=E[-logP(w k |w i )]

[0130] =-P(w k |w i )logP(u k |w i )

[0131] In the formula, P(w k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0132] With Chinese character w i The Chinese character w for the left neighbor k The calculation method of the right conditional entropy is:

[0133] H right (w i , w k )=E[-logP(w k |w i )]

[0134] =-P(w k |wi )logP(w k |w i )

[0135] Among them, P(w k |w i ) represents the word w in the text corpus i The left neighbor is w k probability.

[0136] Corpus preprocessing module 5: Calculate the word-forming contribution of the Chinese character to the unregistered word according to the entropy data, and calculate the word vector of the unregistered word according to the word-forming contribution.

[0137] Optionally, the contribution of the left neighboring character to the word-forming of the unregistered word is calculated based on the left information entropy and the left conditional entropy, and the contribution of the right neighboring character to the word-forming of the right neighboring character of the unregistered word is calculated based on the right information entropy and the right conditional entropy.

[0138] Optionally, the word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character are normalized, and the word-forming contribution of the Chinese character to the unregistered word is calculated based on the normalized word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character.

[0139] Optionally, the word vector of the unregistered word is calculated based on the word-forming contribution of the Chinese character to the unregistered word and the word vector of the Chinese character.

[0140] In a specific implementation, the contribution of a character in an unregistered word to the word formation of the unregistered word is directly proportional to the conditional entropy of its adjacent left and right neighboring characters.

[0141] In the specific implementation, the embodiment of the present application calculates the Chinese character w i and its adjacent left neighbor w k and the right neighbor word w j Let’s take the contribution of the unregistered word t as an example to describe it:

[0142] w i With the left adjacent word w k The contribution calculation method in the process of word formation of unregistered words is:

[0143]

[0144] w i With the right neighbor word w j The contribution calculation method in the process of word formation of unregistered words is:

[0145]

[0146] Chinese character w iThe contribution of the unregistered word t to the word formation process is:

[0147] R(w i , t) = R left (w i , w k )+R right (w i , w j )

[0148] Normalize the word formation contribution of the left neighbor and the word formation contribution of the right neighbor:

[0149]

[0150]

[0151] Normalized Chinese character w i The calculation method of word formation contribution in the word formation process of unregistered word t is:

[0152]

[0153] In the formula, t[n] represents the nth Chinese character in the unregistered word t, and |t| represents the number of Chinese characters contained in the unregistered word t.

[0154] In the specific implementation, the word vector of the unregistered word is calculated based on the word formation contribution:

[0155]

[0156] In the formula, VT(t) represents the word vector of the unregistered word t, VW(w i ) represents the Chinese character w i The word vector of .

[0157] In addition, combined Figure 1 , Figure 2 , Figure 3 The described method for calculating the word vector of an unregistered word can be implemented by an electronic device. Figure 5 It is a framework diagram of the electronic device of the present invention.

[0158] The electronic device may include a processor 61 and a memory 62 storing computer program instructions.

[0159] Specifically, the processor 61 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0160] Among them, the memory 62 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 62 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 62 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 62 may be inside or outside the data processing device. In a specific embodiment, the memory 62 is a non-volatile memory. In a specific embodiment, the memory 62 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (Programmable Read-Only Memory, PROM for short), an erasable PROM (Erasable ProgrammableRead-Only Memory, EPROM for short), an electrically erasable PROM (Electrically Erasable ProgrammableRead-Only Memory, EEPROM for short), an electrically alterable ROM (Electrically Alterable Read-Only Memory, EAROM for short) or a flash memory (FLASH) or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0161] The memory 62 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 61 .

[0162] The processor 61 implements any unregistered word vector calculation method in the above-mentioned embodiment by reading and executing the computer program instructions stored in the memory 62 .

[0163] In some of the embodiments, the electronic device may further include a communication interface 63 and a bus 60. Figure 5 As shown, the processor 61, the memory 62, and the communication interface 63 are connected via a bus 60 and communicate with each other.

[0164] The communication port 63 can realize data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.

[0165] The bus 60 includes hardware, software or both, and couples the components of the electronic device to each other. The bus 60 includes but is not limited to at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 60 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, bus 60 may include one or more buses. Although embodiments of the present application describe and illustrate a particular bus, the present application contemplates any suitable bus or interconnect.

[0166] The electronic device can execute the method for calculating the word vector of the unregistered word in the embodiment of the present application.

[0167] In addition, in combination with the unregistered word vector calculation method in the above embodiment, the present application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any unregistered word vector calculation method in the above embodiment is implemented.

[0168] The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program codes.

[0169] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0170] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. Calculation method of word vector for unregistered words, It is characterized in that include: A priori knowledge acquisition step, acquiring priori knowledge for pre-training, wherein the priori knowledge includes a dictionary, an unlabeled text corpus, and unregistered words; A corpus preprocessing step, preprocessing the unlabeled text corpus; A character co-occurrence counting step, counting the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus; A character entropy data calculation step, calculating the entropy data of the Chinese characters according to the co-occurrence data; A word vector calculation step, calculating the word formation contribution of the Chinese character to the unregistered word according to the entropy data, and calculating the word vector of the unregistered word according to the word formation contribution; The text co-occurrence statistics steps include: A character occurrence counting step, counting the number of times the Chinese character appears in the unlabeled text corpus; A left and right neighbor character acquisition step, acquiring the left neighbor characters and the right neighbor characters that co-occur on the left and right sides of the Chinese character in the unlabeled text corpus; A co-occurrence counting step, obtaining the co-occurrence times of the Chinese character with the left neighboring character and the right neighboring character; The steps of calculating word entropy data include: an information entropy calculation step, calculating the left information entropy and the right information entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character; a conditional entropy calculation step, calculating the left conditional entropy and the right conditional entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character; The word vector calculation steps include: Calculate the word-forming contribution of the left neighboring character to the left neighboring character of the unregistered word according to the left information entropy and the left conditional entropy, and calculate the word-forming contribution of the right neighboring character to the right neighboring character of the unregistered word according to the right information entropy and the right conditional entropy; Normalizing the word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character, and calculating the word-forming contribution of the Chinese character to the unregistered word according to the normalized word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character; The word vector of the unregistered word is calculated according to the word formation contribution of the Chinese character to the unregistered word and the word vector of the Chinese character.

2. A system for calculating word vectors of unregistered words, It is characterized in that include: A priori knowledge acquisition module, which acquires priori knowledge for pre-training, wherein the priori knowledge includes a dictionary, an unlabeled text corpus, and unregistered words; A corpus preprocessing module, which preprocesses the unlabeled text corpus; A character co-occurrence statistics module, which counts the co-occurrence data of a Chinese character in the dictionary in the unlabeled text corpus; A character entropy data calculation module, which calculates the entropy data of the Chinese characters according to the co-occurrence data; A word vector calculation module, which calculates the word formation contribution of the Chinese character to the unregistered word according to the entropy data, and calculates the word vector of the unregistered word according to the word formation contribution; Among them, the text co-occurrence statistics module includes: A character occurrence counting unit, which counts the number of times the Chinese character appears in the unlabeled text corpus; A left and right neighbor character acquisition unit is used to acquire the left and right neighbor characters of the Chinese character that co-occur on the left and right sides in the unlabeled text corpus; A co-occurrence counting unit is used to obtain the co-occurrence times of the Chinese character with the left neighboring character and the right neighboring character; Among them, the word entropy data calculation module includes: An information entropy calculation unit, which calculates the left information entropy and the right information entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character; A conditional entropy calculation unit, which calculates the left conditional entropy and the right conditional entropy of the Chinese character according to the number of times the Chinese character appears in the unlabeled text corpus and the number of times the Chinese character co-occurs with the left neighboring character and the right neighboring character; Among them, the word vector calculation module includes: Calculate the word-forming contribution of the left neighboring character to the left neighboring character of the unregistered word according to the left information entropy and the left conditional entropy, and calculate the word-forming contribution of the right neighboring character to the right neighboring character of the unregistered word according to the right information entropy and the right conditional entropy; Normalizing the word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character, and calculating the word-forming contribution of the Chinese character to the unregistered word according to the normalized word-forming contribution of the left neighbor character and the word-forming contribution of the right neighbor character; The word vector of the unregistered word is calculated according to the word formation contribution of the Chinese character to the unregistered word and the word vector of the Chinese character.

3. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method for calculating the word vector of an unregistered word as described in claim 1 is implemented.

4. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method for calculating the word vector of an unregistered word as described in claim 1 is implemented.

Citation Information

Patent Citations

  • Chinese unregistered word recognition system and method based on improvement information entropy characteristics

    CN103020022A

  • Text entity recognition method and device, electronic device, and storage medium

    CN109145294A