Method and apparatus for constructing a glossary of archaic words and phrases, electronic device, storage medium
By splitting and encoding radicals of ancient characters, combining the character sequences with the highest frequency, and constructing the vocabulary of ancient characters, the problems of low recognition efficiency and high misrecognition rate of ancient characters are solved, and adaptive construction of the vocabulary of ancient characters is realized.
Patent Information
- Application Number
- CN202110742880.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-06-30
AI Technical Summary
When identifying and constructing ancient font vocabulary, the prior art has problems of high misrecognition rate, high cost and low efficiency, especially because the various forms of ancient font writing and the image recognition effect are affected by paper, resulting in poor recognition effect.
By splitting radicals of ancient characters, the preliminary character sequence is obtained, the frequency of continuous radical pairs is counted, and the sequence with the highest frequency is merged to construct a target vocabulary, and the efficiency of the algorithm is optimized using four-corner numbers and suffix identifiers.
Adaptively constructing ancient vocabulary vocabulary is realized, which improves recognition accuracy and efficiency, reduces calculation complexity, and adapts to the vocabulary requirements of downstream tasks.
Smart Images

Figure CN113435189B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for constructing a vocabulary list of ancient Chinese characters, an electronic device, and a storage medium. Background Art
[0002] Most of the current recognition of ancient characters (such as ancient poems and ancient classics) is based on image recognition, and its recognition model itself is derived from an image segmentation task or an image classification task based on a deep learning algorithm model. However, since ancient poems and characters are handwritten and are affected by calligraphy styles of different periods and eras, there may be many different ways to write the same ancient character, resulting in a higher probability of the same ancient character being misidentified as different characters, so the recognition effect is not good. In addition, the effect of image recognition is affected by the paper quality and scanning clarity of ancient classics, and it is easy to be unrecognizable. Although manual recognition verification can be performed by experts in related industries, the recognition efficiency is extremely low and the cost is extremely high. Therefore, if the current technology is used to construct a vocabulary of ancient characters, the result may not be ideal. Summary of the invention
[0003] The main purpose of the embodiments of the present disclosure is to propose a method and device for constructing a vocabulary of ancient Chinese characters, an electronic device, and a storage medium to construct a desired vocabulary.
[0004] To achieve the above-mentioned purpose, a first aspect of the embodiments of the present disclosure proposes a method for constructing a vocabulary of ancient Chinese characters, comprising:
[0005] Obtain a data set of a preset archaic word list to be constructed;
[0006] Splitting the radicals of each ancient character in the data set to obtain the target radicals of each ancient character;
[0007] Encoding each of the target radicals to obtain a preliminary character sequence of each ancient character;
[0008] Counting the frequencies of consecutive radical pairs according to the preliminary character sequence; wherein the consecutive radical pairs include at least two adjacent target radicals;
[0009] Merge the preliminary character sequences corresponding to the continuous radical pairs with the highest frequency to obtain the target coding sequence number;
[0010] A target vocabulary is constructed according to the target coding sequence number.
[0011] In some embodiments, encoding each of the target radicals to obtain a preliminary character sequence of each archaic character includes:
[0012] Encode each target radical to obtain a coding number of each target radical;
[0013] The coding numbers corresponding to the target radicals of each ancient character are concatenated.
[0014] In some embodiments, encoding each target radical to obtain a coding number of each target radical includes:
[0015] A suffix identifier is added to the end of the concatenated code numbers to merge them to obtain the preliminary character sequence.
[0016] In some embodiments, encoding each target radical to obtain a coding number of each target radical includes:
[0017] Each target radical is encoded using the preset four-corner numbers to obtain the encoding number of each target radical.
[0018] In some embodiments, the step of calculating the frequency of consecutive radical pairs based on the preliminary character sequence includes:
[0019] Summarize all preliminary character sequences to obtain a character sequence set;
[0020] The character sequence set is traversed in reverse from the end to the beginning, and the frequencies of the consecutive radical pairs are counted.
[0021] In some embodiments, the method further comprises:
[0022] sorting the preliminary character sequences corresponding to the consecutive radical pairs according to frequency;
[0023] A preliminary character sequence with the highest frequency is determined based on the ranking of the preliminary character sequences.
[0024] In some embodiments, merging the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain the target coding sequence number includes:
[0025] Character sequence acquisition step: obtaining a preliminary character sequence with the highest current frequency;
[0026] Merging step: merging the preliminary character sequences with the highest current frequency and updating the target encoding sequence number;
[0027] The character sequence acquisition step and the merging step are repeatedly performed until the frequency of all current preliminary character sequences is 1.
[0028] To achieve the above-mentioned purpose, a second aspect of the embodiment of the present disclosure proposes a vocabulary building device for ancient Chinese characters, comprising:
[0029] A data acquisition module for acquiring a dataset of a preset archaic word list to be constructed;
[0030] A splitting module for splitting each archaic character in the dataset into radicals to obtain the target radicals of each archaic character;
[0031] An encoding module for encoding each of the target radicals to obtain a preliminary character sequence for each archaic character;
[0032] A frequency statistics module for statistically calculating the frequency of consecutive radical pairs based on the preliminary character sequence; wherein the consecutive radical pairs include at least two adjacent target radicals;
[0033] A merging module for merging the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target encoding sequence number;
[0034] A construction module for constructing a target vocabulary based on the target encoding sequence number.
[0035] To achieve the above object, a third aspect of the embodiments of the present disclosure provides an electronic device, including:
[0036] At least one memory;
[0037] At least one processor;
[0038] At least one program;
[0039] The program is stored in the memory, and the processor executes the at least one program to implement the method described in the first aspect of the present disclosure as above.
[0040] To achieve the above object, a fourth aspect of the embodiments of the present disclosure provides a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute:
[0041] The method described in the first aspect as above.
[0042] The method and apparatus for constructing a vocabulary of archaic Chinese words, electronic device, and storage medium provided by the embodiments of the present disclosure obtain a data set of a preset archaic word list to be constructed, split each archaic Chinese character in the data set into radicals, obtain the target radicals of each archaic Chinese character, encode each target radical to obtain a preliminary character sequence of each archaic Chinese character, count the frequency of consecutive radical pairs based on the preliminary character sequence, merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding sequence number, and construct a target vocabulary based on the target coding sequence number. Through the technical solution provided by the embodiments of the present disclosure, an adaptive construction of the vocabulary of archaic Chinese words can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flowchart of the method for constructing a vocabulary of archaic Chinese words provided by the embodiments of the present disclosure.
[0044] Figure 2 is Figure 1 a flowchart of step 103 in
[0045] Figure 3 is Figure 1 a flowchart of step 104 in
[0046] Figure 4 is a partial flowchart of the method for constructing a vocabulary of archaic Chinese words provided by another embodiment of the present disclosure.
[0047] Figure 5 is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0051] First, some nouns involved in this application are analyzed:
[0052] Radicals: Radicals are the basic units of Chinese characters composed of strokes; radicals are all radicals; but among radicals, only the radicals that represent meaning are called radicals. Chinese characters can be divided into single-character and compound characters. Single-character, also called monomer character, is a character composed of a radical alone, and is no longer called a radical. Compound characters are characters composed of two or more radicals.
[0053] Word segmentation: Word segmentation is also called word cutting, which means identifying and separating the words in a sentence in some way, so that the text is upgraded from the representation of "character sequence" to "word sequence" representation; word segmentation technology is not only applicable to Chinese, but also to English, Japanese, Korean and other languages. Words are the smallest meaningful language components that can act independently. Generally, word segmentation is the first core technology of natural language processing; in English, each sentence separates words with spaces or punctuation marks, while in Chinese, it is difficult to define the boundaries of words and to separate words. In Chinese, although the smallest unit is a character, the semantic expression of an article is still divided by words. Therefore, when processing Chinese text, word segmentation is required to convert sentences into word representations, which is Chinese word segmentation. However, in Chinese sentences, many words are ambiguous, and there may be multiple word segmentation methods in a sentence. For example: "married / of / monk / unmarried / of", "married / of / and / not yet married / of".
[0054] The main idea of statistical word segmentation is to regard each word as composed of characters. If the number of times connected characters appear in different texts is greater, it proves that this segment of connected characters is likely to be one word.
[0055] Recurrent Neural Network (RNN): RNN refers to a structure that repeats over time. RNN can be seen as a neural network that is transmitted in time, and the depth of RNN is the length of time. The biggest difference between RNN and other networks is that RNN can achieve a certain "memory function", which is the best choice for time series analysis. Just as humans can better understand the world with their past memories, RNN also implements a mechanism similar to the human brain, retaining a certain memory of the processed information, unlike other types of neural networks that cannot retain memory of the processed information.
[0056] Bidirectional Encoder Representations from Transformers (Bert): Bert is a method for pre-training language representations. It is a language representation model that uses the transformer encoder for pre-training. It is a new type of language model that pre-trains deep bidirectional representations by jointly adjusting the bidirectional Transformers in all layers.
[0057] Currently, most of the recognition of ancient characters (such as ancient poems and ancient classics) is based on image recognition. The recognition model itself is derived from the image segmentation task or image classification task based on the deep learning algorithm model. However, since the ancient poems are handwritten and are influenced by the calligraphy styles of different periods and eras, the same ancient character may be written in many different ways, resulting in a high probability that the same ancient character will be misidentified as different characters, so the recognition effect is not good. In addition, the effect of image recognition is affected by the paper quality and scanning clarity of ancient classics, and it is easy to fail to recognize. Although manual recognition verification can be performed by experts in related industries, the recognition efficiency is extremely low and the cost is extremely high.
[0058] In addition, the current machine translation accuracy of ancient poetry and dictionary books is not high. On the one hand, this is because the image recognition algorithm is not effective. On the other hand, ancient characters often have multiple meanings. For example, the "zheng" in traditional Chinese medicine classics means in traditional Chinese medicine diagnosis and syndrome dialectics: a summary of the body's response to symptoms at a certain stage. According to the description of the context, "zheng" can be understood as "zheng element", "zhenghou" or "zheng name". Therefore, the recognition and translation of ancient characters also need to refer to the context to solve the problem of one word having multiple meanings.
[0059] Current natural language processing algorithm models based on deep learning and machine learning cannot handle the recognition and translation of ancient characters well, because the minimum granularity of the vocabulary used for training by the current pre-trained model is at the character level. In addition, due to the poor effect of the image recognition algorithm and the fact that ancient characters often have multiple meanings, if the relevant models are used directly, it can be seen that both the word encoding quality and the results of applying the encoding vector to downstream tasks are not ideal.
[0060] Traditional English word segmentation methods are not suitable for the segmentation of ancient words, and English word segmentation usually splits the spaces in the sentence. It was subsequently optimized to the Sentence Piece algorithm, which can split words in English into sub-words to build a word list, but the algorithm principle of this method has been applied to a large number of Berts based on deep learning.
[0061] Based on this, the embodiments of the present disclosure provide a method and apparatus for constructing a vocabulary of archaic Chinese words, an electronic device, and a storage medium, which can realize the adaptive construction of the vocabulary of archaic Chinese words.
[0062] The embodiments of the present disclosure provide a method and apparatus for constructing a vocabulary of archaic Chinese words, an electronic device, and a storage medium, which will be specifically described through the following embodiments. First, the method for constructing the vocabulary of archaic Chinese words in the embodiments of the present disclosure will be described.
[0063] The method for constructing the vocabulary of archaic Chinese words provided by the embodiments of the present disclosure relates to the technical field of data processing. The method for constructing the vocabulary of archaic Chinese words provided by the embodiments of the present disclosure can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the method for constructing the vocabulary of archaic Chinese words, etc., but is not limited to the above forms.
[0064] Figure 1 is an optional flowchart of the method for constructing the vocabulary of archaic Chinese words provided by the embodiments of the present disclosure. Figure 1 The method in can include but is not limited to steps 101 to 106.
[0065] Step 101: Obtain a data set of a preset archaic Chinese word list to be constructed;
[0066] Step 102: Split each archaic Chinese character in the data set into radicals to obtain the target radicals of each archaic Chinese character;
[0067] Step 103: Encode each target radical to obtain a preliminary character sequence of each archaic Chinese character;
[0068] Step 104: Statistically calculate the frequency of consecutive radical pairs according to the preliminary character sequence; the consecutive radical pairs include at least two adjacent target radicals;
[0069] Step 105: Merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding sequence number;
[0070] Step 106: Construct a target vocabulary according to the target coding sequence number.
[0071] In some embodiments, the preset archaic word list can be set according to actual needs, for example, it can be a preset archaic word list including 600 archaic characters, it can also be a preset archaic word list including 800 archaic characters, or it can be an ancient text (i.e., an article) including 300 characters. The data set can be a preset archaic word list or sentence including multiple (600 or 800) archaic characters, or an ancient text (i.e., an article including single Chinese characters, phrases, and sentences) including 300 characters, or an ancient text (i.e., an article including hundreds / thousands of ancient texts).
[0072] In step 102 of some embodiments, for a Chinese character that is a single-character (also called a single-character), its radical refers to the character itself, for example, the radical of the Chinese character "月" is the character itself "月".
[0073] See also Figure 2 In some embodiments, step 103 may include but is not limited to steps 201 to 203.
[0074] Step 201, encode each target radical to obtain a code number of each target radical;
[0075] Step 202, concatenating the code numbers corresponding to the target radicals of each ancient character;
[0076] Step 203: Add a suffix mark to the end of the concatenated code numbers to merge them and obtain a preliminary character sequence.
[0077] In some embodiments, step 201 includes:
[0078] Each target radical is encoded using the four-corner numbers to obtain the encoding number of each target radical.
[0079] Specifically, please refer to Table 1. In some embodiments, the four-corner number method for looking up characters is as shown in Table 1. The coding number for the stroke "亠" (i.e., an independent dot-horizontal combination, horizontal strokes with hooks are not counted) is 0, the coding numbers for "一" and "乚" (i.e., "horizontal, pick, right hook") are both 1, the coding numbers for "丨", "丿", and "亅" (i.e., "straight, left-falling stroke, left hook") are all 2, the coding number for "丶" (i.e., "dot, right stroke") is 3, and so on. For details, please refer to Table 1.
[0080]
[0081] Table 1
[0082] Please refer to Table 1 again, the order of picking the four corner numbers: for each character, pick the four corner numbers in the following order (1)-(4): (1) upper left corner, (2) upper right corner, (3) lower left corner, (4) lower right corner. For example: for the ancient Chinese character "端", the conventional principle of picking the four corner numbers based on word granularity coding is: if (1) the upper left corner is "亠", the corresponding code number is 0, if (2) the upper right corner corresponds to "丨丿亅", the corresponding code number is 2, if (3) the lower left corner corresponds to "一乚" (i.e. "horizontal, pick, right hook"), the corresponding code number is 1, if (4) the lower right corner is "亅", the corresponding code number is 2, then the ancient Chinese character "端" is encoded to obtain 0212; by analogy, for the ancient Chinese character "明", the conventional principle of picking the four corner numbers based on word granularity coding is 6702. According to the technical solution provided by the embodiment of the present disclosure, the principle of splitting the radicals of "明" based on the present case is to first split it into two radicals: "日" and "月", and then respectively perform four-corner numbering on these two radicals, that is, "日" is encoded based on 6010 and "月" is encoded based on 7772, thereby obtaining the preliminary character sequence 60107772 corresponding to "明"; by analogy, according to the technical solution provided by the embodiment of the present disclosure, "端" is first split into three radicals: "立", "山" and "而", and then these three radicals are respectively performed four-corner numbering, which will not be repeated in the embodiment of the present disclosure.
[0083] Furthermore, the method of selecting the angles of the four-corner numbers is:
[0084] ① A stroke can be numbered by angle. For example, if the left side is a stroke, the upper left corner corresponds to 2, and the lower left corner corresponds to 7;
[0085] ② If the upper and lower parts of a stroke and a separate stroke form two different strokes, the two corners are numbered. For example, on the left side of the word "水", the upper left corner corresponds to code number 1, and the lower left corner corresponds to code number 9.
[0086] ③ If the lower corner of the character is biased towards one corner, the number is assigned according to the actual position, and the missing corner is 0. For example, if the lower right corner of the character "嫉" is missing, the corresponding code number for the lower right corner is 0.
[0087] ④ For the three types of characters whose outer parts are "口、门(门)", the left and right lower corners are changed to the inner strokes. For example, the initial character sequence corresponding to "田" is 6040.
[0088] ⑤ If the front corner of a pen shape has been used, the back corner is 0. For example, if the upper left corner of "王" is a horizontal line, the corresponding code number of the upper left corner is 1, and the upper right corner has been used before, so the corresponding code number of the upper right corner is 0.
[0089] Furthermore, the method of taking the angle of the four-corner number is as follows:
[0090] 1. when the four-corner code word is more, get a stroke shape near the top of the lower right corner and make " additional sign " again, if this stroke shape has been used by the upper right corner, then make 0.Specifically, take " as " as example and illustrate, the corresponding fork (i.e. two " ten " or " 乂 " that intersect) of the upper left corner, then the corresponding code number is 4; The corresponding mouth (i.e. four sides are neat, the pen tip does not overhang) of the upper left corner, then the corresponding code number is 6; The corresponding fork (i.e. two " ten " or " 乂 " that intersect) of the lower left corner, then the corresponding corresponding code number is 4; The lower right corner is a mouth, but has been used by the upper right corner, then the corresponding code is 0, so " as " is 4640 after being encoded.
[0091] ② The characters with the same four corners and "attachment" are arranged in order according to the number of horizontal strokes contained in each character. Specifically, if the characters with the same four corners and attachment use the same code, they are arranged in order according to the number of horizontal strokes.
[0092] In the disclosed embodiment, for example, after the ancient Chinese character "明" is split into radicals, two target radicals "日" and "月" are obtained. The radical "日" is encoded using four-corner numbers, and the resulting preliminary character sequence is 6010. The preliminary character sequence obtained after encoding "月" using four-corner numbers is 7772. Thus, the preliminary character sequence of "明" is 6010 and 7772.
[0093] In the embodiment of the present disclosure, the suffix identifier can be expressed as:.
[0094] In an application scenario, please refer to Table 2. Taking the phrase "明月几时" as an example, the initial character sequence obtained by encoding the single ancient character "月" through the four-corner number is 7772, the initial character sequence obtained by encoding the single ancient character "几" through the four-corner number is 7721 (i.e., based on the character granularity encoding), the encoding number of the single ancient character "明" after encoding is 6702 (i.e., based on the character granularity encoding), and the single ancient character "明" is split into two radicals after the radical is split: "日" and "月", and the four-corner number encoding of "日" is performed. The preliminary character sequence obtained after row encoding is 6010, the preliminary character sequence obtained after encoding the four-corner number of "月" is 7772, and the preliminary character sequence of the ancient character "明" is 60107772; after the radical splitting of the single ancient character "时", two radicals are obtained: "日" and "寸", the preliminary character sequence obtained after encoding the four-corner number of "日" is 6010, and the preliminary character sequence obtained after encoding the four-corner number of "寸" is 4030, and the preliminary character sequence of the ancient character "时" is 6010 4030.
[0095] phrase bright moon what time character-level encoding 6702 7772 7721 6400 preliminary character sequence 6010 7772 7772 7721 6010 4030
[0096] Table 2
[0097] In the embodiment of the present disclosure, when performing the merging in step 203, for Chinese characters that are single characters (also called single characters), a suffix mark is also added to the end of the preliminary character sequence when merging, for example, "月" can be represented as 7772; for Chinese characters that are compound characters, a suffix mark is added to the end of the preliminary character sequence, for example, "明" is represented as 6010 7772<\end> after adding the suffix mark; and "明月" can be represented as 6010 77727772; for the phrase "明月几时", after adding the suffix mark and merging, it can be represented as:
[0098] 6010 7772777277216010 4030
[0099] In the disclosed embodiment, the computational complexity can be reduced by using suffix identifiers. For example, if “明” is split into “日” and “月”, the combination of “月” needs to add a suffix identifier as an identifier. If the suffix identifier is not added, the increase in the possibility of the combination will cause the computational complexity to increase exponentially.
[0100] The technical solution provided by the disclosed embodiment can effectively improve the calculation efficiency of the algorithm by adding suffix identifiers. Because there is natural language knowledge information in archaic characters, the radicals themselves are unique, and the left and right, up and down positions of the radicals are not interchangeable. Therefore, a stop sign is added as a suffix identifier to avoid adding combinations that do not conform to the knowledge of the archaic character language into the vocabulary, thereby reducing the introduction of noise.
[0101] See also Figure 3 In some embodiments, step 104 includes:
[0102] Step 301: Summarize all preliminary character sequences to obtain a character sequence set;
[0103] Step 302: traverse the character sequence set from the end to the beginning in reverse order, and count the frequencies of consecutive radical pairs.
[0104] Specifically, the preliminary character sequences corresponding to all the archaic characters in the data set are merged and summarized to obtain a character sequence set. For the character sequence set, a reverse traversal method is performed from the end of the set to the beginning of the set to count the frequencies of consecutive radical pairs.
[0105] In some embodiments, step 105 includes:
[0106] Character sequence acquisition step: obtaining a preliminary character sequence with the highest current frequency;
[0107] Merging step: merge the preliminary character sequences with the highest current frequency and update the target encoding sequence number;
[0108] Repeat the above character sequence acquisition step and merging step until the frequency of any one of the current preliminary character sequences is 1.
[0109] Specifically, in an application scenario, if in the dataset, the frequency of the preliminary character sequence 1001 1002 1003 is 2 (where 1001 1002 1003 represents the preliminary character sequence obtained by encoding the first archaic character after splitting it into three radicals), the frequency of the preliminary character sequence 1001 1002 1003 1005 is 3 (where 1001 1002 1003 1005 represents the preliminary character sequence obtained by encoding the second archaic character after splitting it into four radicals), and the frequency of the preliminary character sequence 1006 1004 1007 1004 is 4 (where 1006 1004 1007 1004 represents the preliminary character sequence obtained by encoding the third archaic character after splitting it into four radicals), then the initial character sequence set can be recorded as:
[0110] {'1001 1002 1003': 2, '1001 1002 1003 1005': 3, '1006 1004 1007 1004': 4}
[0111] Taking the above initial character sequence set as the first input, the reverse traversal method is as follows: Traverse from the end of the character sequence set towards 1004 and 1007 until reaching the beginning 1001. The most frequently occurring consecutive radical pair 1001 and 1002 appears 2 + 3 = 5 times, so 1001 and 1002 are merged into "(1001 1002)". After the first traversal, the first output is:
[0112] {'(1001 1002)}1003': 2, '(1001 1002)1003 1005': 3, '1006 1004 1007 1004': 4}
[0113] Taking the first output as the second input, and again using the reverse traversal method: Traverse in reverse from the end of the character sequence set towards the beginning. The most frequently occurring consecutive radical pair 1007 and 1004 appears 4 times, so 1007 1004 are merged into "(1007 1004)". After the second traversal, the second output is:
[0114] {'(1001 1002)}1003': 2, '(1001 1002)1003 1005': 3, '1006 1004(1007 1004)': 4}
[0115] The second output is used as the third input, and the reverse traversal method is used again: reverse traversal is performed from the end of the character sequence set to the beginning, and the highest frequency 1004 (1007 1004) appears 4 times, so 1004 (1007 1004) is merged into "(1004 1007 1004)". After the third traversal, the third output is:
[0116] {'(1001 1002)}1003':2,'(1001 1002)1003 1005':3,'1006(10041007 1004)':4}
[0117] And so on, the iteration is continued until the frequency of any current preliminary character sequence is 1, that is, the iteration is stopped; or, it can also be stopped according to actual needs, for example, until the constructed target vocabulary reaches the expected preset ancient word list size.
[0118] See also Figure 4 In some embodiments, the method for constructing a vocabulary of archaic words further includes:
[0119] Step 401, sorting preliminary character sequences corresponding to consecutive radical pairs according to frequency;
[0120] Step 402: Determine the preliminary character sequence with the highest frequency according to the sorting of the preliminary character sequences.
[0121] Since the radical corresponds to a coding number, each ancient character corresponds to a preliminary character sequence. That is, the mapping association between the radical and the preliminary character sequence can be used to map the statistical frequency of the preliminary character sequence with the frequency of the radical to obtain the frequency of the corresponding radical pair.
[0122] In an application scenario, if the frequency of "明" in the data set is 20 ("明" can be split into "日" and "月" based on radicals, the frequency of "日" is 20, and the frequency of "月" is 20), the frequency of "时" in the data set is 12 ("时" can be split into "日" and "寸" based on radicals, the frequency of "日" is 12, and the frequency of "寸" is 12), the frequency of "日" in the data set is 18, and the frequency of "月" in the data set is 16, and assuming that no other characters in the data set have the three radicals "日", "月", and "寸", then it can be known that in the data set, the frequency of "日" is 20+12+18=50, the frequency of "月" is 20+16=36, and the frequency of "寸" is 12. Therefore, among the frequencies of "日", "月", and "寸", the one with the highest frequency deviation is "日", and the one with the lowest frequency deviation is "寸".
[0123] Specifically, in an application scenario, please refer to Table 3. The frequency of the ancient character "明" is 20. After the ancient character "明" is split into radicals, the two target radicals "日" and "月" are obtained. The radical "日" is encoded using the four-corner number, and the resulting preliminary character sequence is 6010. The preliminary character sequence obtained after encoding the four-corner number for "月" is 7722. The preliminary character sequences corresponding to the two target radicals "日" and "月" are concatenated to obtain 6010 and 7722, and then a suffix mark is added after the concatenation, so that the target coding number is [6010 7722]: 20.
[0124] ancient style character frequency encoding of "sun" encoding of "moon" concatenation merged target encoding serial number bright 20 6010 7772 6010、7772 [6010 7772]:20
[0125] Table 3
[0126] The method for constructing a vocabulary of archaic characters provided by the embodiment of the present disclosure obtains a data set of a preset archaic word list to be constructed, and splits each archaic character in the data set into radicals to obtain a target radical of each archaic character, and then encodes each target radical to obtain a preliminary character sequence of each archaic character, and then counts the frequency of continuous radical pairs according to the preliminary character sequence, and then merges the preliminary character sequences corresponding to the continuous radical pairs with the highest frequency to obtain a target coding sequence number, and constructs a target vocabulary according to the target coding sequence number. The technical solution provided by the embodiment of the present disclosure can realize the adaptive construction of the vocabulary of archaic characters, so as to realize the adaptive construction of the vocabulary of archaic characters. The embodiment of the present disclosure utilizes the encoding of radicals to construct archaic characters from a finer granularity, and utilizes the information of the increased granularity of radicals to increase the amount of information of the archaic subcontext.
[0127] The technical solution provided by the disclosed embodiment can effectively improve the calculation efficiency of the algorithm by adding suffix identifiers. Because there is natural language knowledge information in archaic characters, the radicals themselves are unique, and the left and right, up and down positions of the radicals are not interchangeable. Therefore, a stop sign is added as a suffix identifier to avoid adding combinations that do not conform to the knowledge of the archaic character language into the vocabulary, thereby reducing the introduction of noise.
[0128] Embodiments of the present disclosure use a coding method for archaic Chinese character tables based on the division of radicals, which is different from the word segmentation methods of modern Chinese and English, and can effectively reduce the poor parsing and coding effect of archaic Chinese characters by deep models constructed using modern vocabulary tables. The technical solution of the embodiments of the present disclosure can also achieve the expansion of the character table. For example, if the current character table only includes 100 archaic Chinese characters, and the downstream task requires 300 archaic Chinese characters to be completed (i.e., an additional 200 archaic Chinese characters need to be expanded), the technical solution of the embodiments of the present disclosure can achieve the expansion of the character table. Therefore, the algorithm provided by the embodiments of the present disclosure can adapt to the requirements of downstream tasks for the size of the vocabulary table, which is different from the requirements of traditional RNN (Recurrent Neural Network) models and the deep pre-training model Bert for the size of the vocabulary table. Therefore, the adaptive algorithm can generate a vocabulary table of the expected size according to the requirements of downstream tasks.
[0129] Embodiments of the present disclosure can construct archaic Chinese characters at a finer granularity by encoding radicals, and use the increased information at the radical granularity to increase the amount of context information of archaic Chinese characters.
[0130] Embodiments of the present disclosure also provide a device for constructing a vocabulary table of archaic Chinese words, which can implement the method for constructing a vocabulary table of archaic Chinese words. The device includes:
[0131] A data acquisition module, configured to acquire a dataset of a preset archaic Chinese word table to be constructed;
[0132] A splitting module, configured to split each archaic Chinese character in the dataset into radicals to obtain the target radicals of each archaic Chinese character;
[0133] An encoding module, configured to encode each target radical to obtain a preliminary character sequence of each archaic Chinese character;
[0134] A frequency statistics module, configured to count the frequency of each target radical according to the preliminary character sequence;
[0135] A merging module, configured to merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target encoding sequence number;
[0136] A construction module, configured to construct a target vocabulary table according to the target encoding sequence number.
[0137] Embodiments of the present disclosure also provide an electronic device, including:
[0138] At least one memory;
[0139] At least one processor;
[0140] At least one program;
[0141] The program is stored in a memory, and a processor executes the at least one program to implement the method for constructing a glossary of archaic words as described above in the present disclosure. The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc.
[0142] Please refer to Figure 5 , Figure 5 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0143] A processor 501, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;
[0144] A memory 502, which can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), etc. The memory 502 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 502 and are called by the processor 501 to execute the method for constructing a glossary of archaic words in the embodiments of the present disclosure;
[0145] An input / output interface 503, which is used to implement information input and output;
[0146] A communication interface 504, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.); and
[0147] A bus 505, which transmits information between various components of the device (such as the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);
[0148] Among them, the processor 501, the memory 502, the input / output interface 503, and the communication interface 504 are communicatively connected to each other inside the device through the bus 505.
[0149] Embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned method for constructing a vocabulary of archaic Chinese words.
[0150] The method for constructing a vocabulary of archaic Chinese words, the apparatus for constructing a vocabulary of archaic Chinese words, the electronic device, and the storage medium provided by the embodiments of the present disclosure obtain a data set of a preset archaic Chinese word list to be constructed, split each archaic Chinese character in the data set into radicals, obtain the target radicals of each archaic Chinese character, encode each target radical to obtain a preliminary character sequence of each archaic Chinese character, count the frequency of consecutive radical pairs according to the preliminary character sequence, merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding sequence number, and construct a target vocabulary according to the target coding sequence number. Through the technical solution provided by the embodiments of the present disclosure, an adaptive construction of a vocabulary of archaic Chinese words can be realized. By adding a suffix identifier, the calculation efficiency of the algorithm can be effectively improved. Since archaic Chinese characters contain natural language knowledge information, the radicals themselves are unique, and their left-right and up-down positions are not interchangeable. Therefore, a stop symbol is added as a suffix identifier to prevent combinations that do not conform to the language knowledge of archaic Chinese characters from being added to the word list and reduce the introduction of noise.
[0151] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0152] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0153] Those skilled in the art can understand that Figure 1-4 the technical solutions shown do not constitute a limitation on the embodiments of the present disclosure and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0156] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0157] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0158] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0159] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0161] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0162] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the present disclosure shall be within the scope of rights of the present disclosure.
Claims
1. A method for constructing a vocabulary of archaic words and expressions, characterized in that, Including: Obtain a dataset of a preset archaic Chinese character list to be constructed; Perform radical splitting on each archaic Chinese character in the dataset to obtain the target radicals of each archaic Chinese character; Encode each of the target radicals to obtain a preliminary character sequence for each archaic Chinese character; Statistically calculate the frequency of consecutive radical pairs based on the preliminary character sequence; wherein, the consecutive radical pairs include at least two adjacent target radicals; Merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding serial number; Construct a target vocabulary based on the target coding serial number; Wherein, the encoding each of the target radicals to obtain a preliminary character sequence for each archaic Chinese character includes: Encode each target radical to obtain a coding number for each target radical; specifically including: using a preset four-corner code to encode each target radical to obtain a coding number for each target radical; Concatenate the coding numbers corresponding to the target radicals of each archaic Chinese character; Wherein, the merging the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding serial number includes: Character sequence acquisition step: Obtain the preliminary character sequence with the highest current frequency; Merging step: Merge the preliminary character sequence with the highest current frequency to update the target coding serial number; Repeat the character sequence acquisition step and the merging step until the frequency of all current preliminary character sequences is 1.
2. The method according to claim 1, wherein The encoding each of the target radicals to obtain a preliminary character sequence for each archaic Chinese character further includes: Add a suffix identifier at the end of the concatenated coding numbers for merging to obtain the preliminary character sequence.
3. The method according to any one of claims 1 to 2, characterized in that The statistically calculating the frequency of consecutive radical pairs based on the preliminary character sequence includes: Summarize the preliminary character sequences to obtain a character sequence set; Traverse the character sequence set in reverse order from the end to the beginning to statistically calculate the frequency of the consecutive radical pairs.
4. The method according to any one of claims 1 to 2, characterized in that The method further includes: Sort the preliminary character sequences corresponding to the consecutive radical pairs according to the frequency; Determine the preliminary character sequence with the highest frequency according to the sorting of the preliminary character sequences.
5. A vocabulary building device for archaic words and expressions, characterized in that Including: A data acquisition module for obtaining a dataset of a preset archaic Chinese character list to be constructed; A splitting module for performing radical splitting on each archaic Chinese character in the dataset to obtain the target radicals of each archaic Chinese character; An encoding module for encoding each of the target radicals to obtain a preliminary character sequence for each archaic Chinese character; A frequency statistics module for statistically calculating the frequency of consecutive radical pairs based on the preliminary character sequence; wherein the consecutive radical pairs include at least two adjacent target radicals; A merging module for merging the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain a target coding serial number; A construction module for constructing a target vocabulary based on the target coding serial number; Wherein, the encoding module for encoding each of the target radicals to obtain a preliminary character sequence for each archaic Chinese character includes: Encode each target radical to obtain the encoding number of each target radical; specifically, include: using the preset four-corner numbers to encode each target radical to obtain the encoding number of each target radical; Concatenate the encoding numbers corresponding to the target radicals of each archaic character; Among them, the merging module is used to merge the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequency to obtain the target encoding sequence number, including: Character sequence acquisition step: acquire the preliminary character sequence with the highest current frequency; Merging step: merge the preliminary character sequence with the highest current frequency to update the target encoding sequence number; Repeat the character sequence acquisition step and the merging step until the frequencies of all current preliminary character sequences are 1.
6. An electronic device, characterized in that, Include: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to implement: The method according to any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute: The method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Radical input method
CN101872250A