Oracle bone oracle bone character classification method based on context mask prediction
Patent Information
- Application Number
- CN202610841951.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-28
AI Technical Summary
[0007]为解决现有技术中存在仅依据字形外观难以利用甲骨文语句上下文判断异体字标准字归属、人工映射方式难以适应大规模语料中待归类异体字自动处理需求的缺陷,本发明提供的技术方案为:
将甲骨文原始字符序列映射为字形分类编码序列,使模型处理对象从不统一的原始字形文本转化为具有一致编码规则的字位序列。该特征避免了甲骨文中未编码字、未释字、异体字形混杂造成的输入不稳定问题,使同一语料中的标准字、异体字、破损字和不可辨识字能够在统一词表下参与建模,从而为后续上下文预测提供稳定的数据基础。
Smart Images

Figure CN122657918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ancient script digitization technology, and in particular to a method for classifying variant characters in oracle bone script based on context mask prediction. Background Technology
[0002] The classification of variant characters in oracle bone script falls under the fields of ancient script digitization and natural language processing. It is primarily used in the processes of organizing oracle bone script corpora, character encoding, lexical analysis, and textual interpretation to categorize different forms of the same standard character. As an early form of Chinese characters, oracle bone script was not yet fully standardized. The same standard character may exhibit variations in stroke additions or subtractions, component positions, character orientation, component substitutions, and structural differences across different lexical examples, inscription environments, or materials. Therefore, the construction of digitized oracle bone script corpora typically requires organizing the correspondence between different character forms and the standard character.
[0003] Existing methods for classifying variant characters in oracle bone script mainly include manual interpretation and classification, and automatic recognition based on character images. Manual interpretation typically involves researchers combining character structure, lexical examples, structural rules, historical period, and existing character tables, interpretation tables, and variant relationship tables to determine whether different oracle bone characters belong to the same standard character. Automatic recognition based on character images typically uses oracle bone character images as the processing object, employing image feature extraction, character similarity calculation, or visual recognition models to cluster or classify characters with similar appearances to assist in the work of organizing variant characters.
[0004] However, in practical applications, variant characters in oracle bone script do not always exhibit similar appearances. Some characters, while visually similar, do not belong to the same standard character in lexical contexts and actual usage; others, differing significantly in component forms, structural methods, or overall appearance, may constitute variant characters because they record the same word or exist in similar contexts. Therefore, classifying characters solely based on visual similarity easily overlooks the contextual distribution and usage differences of the target character in oracle bone script sentences. This makes it difficult to accurately handle variant characters with significant structural differences but similar contextual usages, and may also lead to the misclassification of characters with similar appearances but different usages into the same category.
[0005] Meanwhile, while manual interpretation and classification methods can comprehensively utilize character form and lexical examples, their classification process relies heavily on researchers' experience and is quite complex. With the continuous expansion of large-scale digitized oracle bone script corpora, existing character form tables, interpretation tables, and variant character relationship tables may contain character form relationships that need to be supplemented, revised, or are not covered. Relying solely on static manual mapping relationships for classification makes it difficult to adapt to the processing needs of new corpora and variant characters requiring interpretation. Furthermore, oracle bone script sentences are typically short and may contain damaged characters, unrecognizable characters, or missing context. How to effectively utilize contextual information in short and incomplete sentences to assist in determining the standard character attribution of variant characters remains a problem that existing automatic classification methods need to solve.
[0006] In summary, existing technologies have shortcomings, such as difficulty in determining the standard character classification of variant characters based solely on their appearance using the context of oracle bone script sentences, and the inability of manual mapping methods to meet the needs of automatic processing of variant characters to be classified in large-scale corpora. Summary of the Invention
[0007] To address the shortcomings of existing technologies, such as the difficulty in determining the standard character classification of variant characters based solely on their appearance using the context of oracle bone script sentences, and the inadequacy of manual mapping methods to meet the needs of automatic processing of variant characters to be classified in large-scale corpora, the technical solution provided by this invention is as follows: A method for classifying variant characters in oracle bone script based on context mask prediction includes: The steps of obtaining oracle bone sentences from oracle bone script corpus and converting the oracle bone sentences into character shape classification coding sequences according to the character position order; The training samples are constructed based on the character shape classification coding sequence, wherein the character position with a determined standard character category is replaced with a mask coding, the replaced character shape classification coding sequence is used as the model input, the original character shape classification coding sequence is used as the model output, and the target variant character position is not used as the mask position. The steps are as follows: train a context mask prediction model using the training samples, and make the context mask prediction model predict the glyph classification code corresponding to the mask position based on the word context before and after the mask position, thus obtaining the trained context mask prediction model. The steps include: obtaining the oracle bone script sentence to be classified and the target variant character positions within it, and converting the oracle bone script sentence to be classified into a character shape classification encoding sequence to be inferred according to the same character shape classification encoding rules; The steps are as follows: replace the character shape classification code corresponding to the target variant character position in the character shape classification coding sequence to be reasoned with the mask code, and input it into the trained context mask prediction model to obtain the candidate character shape classification code corresponding to the target variant character position; The steps are as follows: determining the corresponding standard character category based on the candidate character shape classification code, and outputting the standard character category as the classification result of the target variant character position.
[0008] Furthermore, in a preferred embodiment, the character shape classification coding sequence includes standard character shape classification coding, variant character shape classification coding, and damaged or unrecognizable character coding, and the character shape classification coding corresponds one-to-one with the oracle bone characters in the oracle bone script sentence according to the character position order.
[0009] Furthermore, in a preferred embodiment, the character-level word segmentation is performed on the character classification encoding sequence, so that the input token obtained by word segmentation corresponds one-to-one with the oracle bone characters in the oracle bone sentence, and the word-segmented encoding sequence is used as the basis for processing training samples and character classification encoding sequences to be inferred.
[0010] Furthermore, in a preferred embodiment, the training samples generate mask positions based on the characters with defined standard character categories. The target variant characters, damaged characters, and unrecognizable characters are not used as mask positions and are not used as objects for calculating prediction error when updating model parameters.
[0011] Furthermore, in a preferred embodiment, the context mask prediction model performs whole-sentence context encoding on the mask input sequence and generates the predicted probability of each candidate glyph classification code corresponding to the mask position based on the whole-sentence context encoding result.
[0012] Furthermore, in a preferred embodiment, the candidate character shape classification codes corresponding to the target variant character position are sorted according to the predicted probability, and the candidate character shape classification code with the highest predicted probability is selected as the target candidate character shape classification code.
[0013] A device for classifying variant characters in oracle bone script based on context mask prediction, comprising: A module that acquires oracle bone sentences from oracle bone script corpus and converts the oracle bone sentences into character shape classification encoding sequences according to the character position order; Training samples are constructed based on the character shape classification coding sequence, wherein the character position with a defined standard character category is replaced with a mask coding, the replaced character shape classification coding sequence is used as the model input, the original character shape classification coding sequence is used as the model output, and the target variant character position is not used as the mask position. The training samples are used to train a context mask prediction model, which predicts the glyph classification code corresponding to the mask position based on the word context before and after the mask position, thus obtaining the module of the trained context mask prediction model. A module that acquires the oracle bone script sentence to be classified and the target variant character positions therein, and converts the oracle bone script sentence to be classified into a character shape classification encoding sequence to be deduced according to the same character shape classification encoding rules; The glyph classification code corresponding to the target variant character position in the glyph classification code sequence to be inferred is replaced with a mask code, and input into the trained context mask prediction model to obtain the module of candidate glyph classification code corresponding to the target variant character position; The module determines the corresponding standard character category based on the candidate character shape classification code and outputs the standard character category as the classification result of the target variant character position.
[0014] A computer storage medium for storing a computer program, which, when read by the computer, is executed by the computer using the method described thereon.
[0015] A computer, including a processor and a storage medium, executes the method when the processor reads a computer program stored in the storage medium.
[0016] A computer program product, which, as a computer program, implements the method when the computer program is executed.
[0017] Compared with the prior art, the advantages of the technical solution provided by the present invention are as follows: Mapping the original oracle bone script character sequences to glyph classification encoding sequences transforms the model's processing object from inconsistent original glyph text into a sequence of character positions with consistent encoding rules. This feature avoids the input instability problem caused by the mixture of unencoded characters, unexplained characters, and variant characters in oracle bone script, allowing standard characters, variant characters, damaged characters, and unrecognizable characters in the same corpus to participate in modeling under a unified vocabulary, thus providing a stable data foundation for subsequent context prediction.
[0018] Mapping glyph encoding to the Unicode private region allows oracle bone script glyphs not included in the general character set to still enter the model as trainable characters. This feature reduces the loss of contextual information caused by treating unencoded characters as unknown symbols, enabling the model to retain the distribution characteristics of specific glyphs in oracle bone script examples, thereby improving the problem that traditional text modeling cannot directly handle special glyphs in oracle bone script.
[0019] By employing SentencePiece character-level segmentation and maintaining a one-to-one correspondence between the input token and the original character position, the target variant character positions in short oracle bone script sentences can be accurately located. This feature avoids the problem of target position shifting caused by changes in character position boundaries in conventional segmentation methods, ensuring consistency between the mask position during training, the target variant character position during inference, and the final predicted position, thereby improving the positional reliability of variant character classification tasks in short sentence corpora.
[0020] A glyphic vocabulary consisting of a standard character set, a variant character set, codes for damaged or unrecognizable characters, and special symbols is established to distinguish and process different types of oracle bone script characters. This feature enables the model to recognize normal standard character contexts while preserving the specific codes of target variant characters, and to independently represent damaged and unrecognizable characters, thereby reducing the interference of incomplete corpora on context modeling.
[0021] The merging mapping from specific variant character forms to standard character categories is limited to target selection, post-inference merging, and result statistics, rather than being used directly as a supervision label at the target variant character training location. This feature avoids the model simply memorizing artificial variant character mapping relationships during training, allowing the model to rely more on the context of sentences to learn the distribution of character form usage, thereby improving its ability to assist in the classification of variant characters that are not fully organized or require further interpretation.
[0022] When constructing training samples, mask positions are randomly selected from ordinary non-empty standard character positions, without using target variant characters or damaged characters as random mask positions. This feature allows the model training to focus on the task of restoring the context of standard characters with identifiable meanings, while preserving the true distribution of target variant characters in sentences. This avoids incorrectly including variant characters or damaged characters whose attribution is not yet certain in the supervised loss, thereby improving the reliability of training labels.
[0023] The BART model is trained based on the input sequence after masking corruption, enabling the model to obtain the entire sentence context representation at the encoding end and recover the original glyph encoding sequence at the decoding end. This feature transforms the problem of variant character classification into a context prediction problem, allowing the model to use the context information of the words before and after the target character to determine its possible corresponding standard character category, thereby making up for the insufficiency of relying solely on the similarity of glyph appearance to determine semantic usage relationships.
[0024] The bidirectional encoder in the BART model performs overall contextual encoding on the masked input sequence, allowing the positional information of characters on both sides of the target position to participate in contextual modeling. This feature is suitable for the characteristics of short oracle bone script sentences and insufficient information from one side, and can comprehensively utilize the limited context before and after the target variant character to enhance the model's ability to determine the standard character's attribution under short sentence conditions.
[0025] The autoregressive decoder, combined with a teacher-forcing mechanism, recovers the target sequence, enabling the decoding process to generate a current position prediction result under the common constraints of the given left-hand target token and the encoder context representation. This feature allows the model to not only rely on the masked context of the input but also utilize the positional continuity relationships within the output sequence, thereby improving the coherence of candidate standard character predictions in oracle bone script phrases.
[0026] The decoder reads the encoder output through a cross-attention mechanism, enabling the current position prediction to assign different attention weights to each character position in the input sequence. This feature allows the model to extract information from contextual characters that are more relevant to the classification of the target variant character, rather than mechanically judging based on fixed distance or overall similarity, thereby enhancing its ability to utilize key word examples in the same context.
[0027] By calculating the predicted probability of each candidate token at the target position through the output layer and softmax, the model can form candidate standard character categories ranked by probability. This feature can not only provide a single classification result, but also output Top-K candidate results, providing multiple candidate references for the interpretation of variant characters in oracle bone inscriptions, and adapting to the actual situation of uncertainty and the need for verification in the classification of ancient characters.
[0028] The mask position cross-entropy loss is calculated only at the mask position corresponding to the standard character Label, and not at the positions of target variant characters or damaged characters. This feature avoids noisy supervision caused by forcibly training the model using uncertain character positions, and makes parameter updates mainly come from character position prediction tasks with clear attribution, thereby improving the training stability of the context prediction model.
[0029] The training samples are categorized by difficulty based on the proportion of damaged characters in oracle bone script sentences, allowing samples with low, medium, and high damage levels to participate in training in order of increasing difficulty. This feature enables the model to first learn the character distribution patterns from samples with relatively complete context, and then gradually adapt to incomplete samples with weaker contextual information, thereby improving the problem of unstable predictions caused by the large number of damaged characters in oracle bone fragment corpora.
[0030] By using the model parameters obtained from training at the previous level as the initialization parameters for training at the next level, the model's capabilities can be progressively accumulated with the increasing difficulty of the corpus. This feature prevents the model from struggling to converge when directly facing a large amount of incomplete corpus in the early stages of training. Instead, it first develops basic context recovery capabilities and then expands to more complex incomplete contexts, thereby improving the final model's adaptability to short sentences and incomplete contexts.
[0031] In the inference stage, the position of the target variant character in the sample to be classified is replaced with a mask symbol, and the trained model generates a candidate sequence based on the context of the entire sentence. This feature enables the system to predict the attribution of the target variant character to its standard character based on the contextual environment of the target variant character without directly relying on the shape image of the target variant character and manual mapping labels, thereby being capable of handling variant relationships where the shapes differ greatly but the contextual usage is similar.
[0032] After uniformly merging the tokens output by the model into standard Labels, the final classification result is generated, so that the prediction result of the model can correspond to the standard character categories required for the classification of oracle bone variant characters. This feature converts the glyph encoding result obtained from context prediction into a standard character attribution result that can be used for variant character collation and statistics, avoiding the problem that the model output stays at the encoding level and is difficult to be directly used for ancient character classification tasks.
[0033] Based on the Top-K candidate results, comparison with standard glyph encoding is performed and the hit situation is calculated, enabling the system to verify the prediction ability under different candidate ranges. This feature adapts to the characteristics of multiple candidates and uninterpreted cases in the classification of oracle bone variant characters, so that the classification result is not limited to Top-1 judgment, but can also provide candidate references such as Top-2 and Top-3, thereby improving the usability of the result in manual review and subsequent interpretation. Description of Drawings
[0034] Figure 1 is a statistical chart of citation volume of variant character documents.
[0035] Figure 2 is a schematic diagram of the overall structure of the system.
[0036] Figure 3 is an example diagram of oracle bone character glyph encoding data.
[0037] Figure 4 is a logic diagram of model training.
[0038] Figure 5 is a schematic diagram of the structure of the adopted BART model.
[0039] Figure 6 is a logic diagram of variant character classification testing.
[0040] Figure 7 is a schematic diagram of the prediction result of variant characters of the character 子 (zi).
[0041] Figure 8 is a schematic diagram of the prediction result of variant characters of the character 巳 (si).
[0042] Figure 9 is a schematic diagram of the prediction result of variant characters of the character 羌 (qiang).
[0043] Figure 10 is a schematic diagram of the prediction result of variant characters of the character 巛 (chuan).
[0044] Figure 11 It is a schematic diagram of the prediction result for variant characters of "Wu (五)".
[0045] Figure 12 It is a schematic diagram of the prediction result for variant characters of "Feng (風)".
[0046] Figure 13 It is a schematic diagram of the prediction result for variant characters of "Bi (畀)".
[0047] Figure 14 It is a schematic diagram of the prediction result for variant characters of "Shu (黍)".
[0048] Figure 15 It is a schematic diagram of the prediction result for variant characters of "Huan (衁)".
[0049] Figure 16 It is a schematic diagram of the prediction result for variant characters of "Zhen (朕)". Detailed Description of Embodiments
[0050] In order to more clearly show the advantages and benefits of the technical solution provided by the present invention, the technical solution provided by the present invention will now be described in further detail with reference to the accompanying drawings. Specifically: Embodiment 1. This embodiment provides an oracle bone inscription variant character classification method based on context mask prediction, comprising: A step of obtaining oracle bone inscription sentences from oracle bone inscription corpus and converting said oracle bone inscription sentences into glyph classification coding sequences according to character position order; A step of constructing training samples based on said glyph classification coding sequences, wherein character positions with determined standard character categories are replaced with mask codes, the replaced glyph classification coding sequences are used as model inputs, the glyph classification coding sequences before replacement are used as model outputs, and target variant character positions are not used as mask positions; A step of training a context mask prediction model using said training samples, enabling said context mask prediction model to predict the glyph classification coding corresponding to the mask position according to the character position context before and after the mask position, so as to obtain a trained context mask prediction model; A step of obtaining an oracle bone inscription sentence to be classified and the target variant character position therein, and converting said oracle bone inscription sentence to be classified into a glyph classification coding sequence to be inferred according to the same glyph classification coding rule; A step of replacing the glyph classification coding corresponding to said target variant character position in said glyph classification coding sequence to be inferred with a mask code, and inputting the result into said trained context mask prediction model to obtain candidate glyph classification codings corresponding to said target variant character position; A step of determining a corresponding standard character category according to the candidate glyph classification coding, and outputting said standard character category as the classification result of said target variant character position.
[0051] The character shape classification coding sequence includes standard character shape classification coding, variant character shape classification coding, and damaged or unrecognizable character coding. The character shape classification coding corresponds one-to-one with the oracle bone characters in the oracle bone script sentence according to the character position order.
[0052] The character-level word segmentation is performed on the character shape classification encoding sequence so that the input token obtained by word segmentation corresponds one-to-one with the oracle bone characters in the oracle bone sentence. The segmented encoding sequence is used as the basis for training samples and processing of the character shape classification encoding sequence to be inferred.
[0053] The training samples generate mask positions based on the characters with defined standard character categories. The characters of target variant characters, damaged characters, and unrecognizable characters are not used as mask positions and are not included in the prediction error calculation when updating model parameters.
[0054] The context mask prediction model performs whole-sentence context encoding on the mask input sequence and generates the predicted probability of each candidate glyph classification code corresponding to the mask position based on the whole-sentence context encoding result.
[0055] The candidate character shape classification codes corresponding to the target variant character position are sorted according to the predicted probability, and the candidate character shape classification code with the highest predicted probability is selected as the target candidate character shape classification code.
[0056] A device for classifying variant characters in oracle bone script based on context mask prediction, comprising: A module that acquires oracle bone sentences from oracle bone script corpus and converts the oracle bone sentences into character shape classification encoding sequences according to the character position order; Training samples are constructed based on the character shape classification coding sequence, wherein the character position with a defined standard character category is replaced with a mask coding, the replaced character shape classification coding sequence is used as the model input, the original character shape classification coding sequence is used as the model output, and the target variant character position is not used as the mask position. The training samples are used to train a context mask prediction model, which predicts the glyph classification code corresponding to the mask position based on the word context before and after the mask position, thus obtaining the module of the trained context mask prediction model. A module that acquires the oracle bone script sentence to be classified and the target variant character positions therein, and converts the oracle bone script sentence to be classified into a character shape classification encoding sequence to be deduced according to the same character shape classification encoding rules; The glyph classification code corresponding to the target variant character position in the glyph classification code sequence to be inferred is replaced with a mask code, and input into the trained context mask prediction model to obtain the module of candidate glyph classification code corresponding to the target variant character position; The module determines the corresponding standard character category based on the candidate character shape classification code and outputs the standard character category as the classification result of the target variant character position.
[0057] A computer storage medium for storing a computer program, which, when read by the computer, is executed by the computer using the method described thereon.
[0058] A computer, including a processor and a storage medium, executes the method when the processor reads a computer program stored in the storage medium.
[0059] A computer program product, which, as a computer program, implements the method when the computer program is executed.
[0060] Implementation Method Two: This implementation method is a further detailed description of the technical solution provided in Implementation Method One, specifically: This embodiment provides a method for classifying variant characters in oracle bone script based on context masking prediction, used to determine the standard character category corresponding to a target variant character based on the context of an oracle bone script statement. The method includes the following steps.
[0061] Oracle bone script sentences are read from an oracle bone script corpus. Each sentence comprises multiple oracle bone script characters arranged in positional order. A glyph classification encoding is performed on each character character, converting the original character sequence into a glyph classification encoding sequence. This glyph classification encoding includes at least standard character glyph classification encoding, variant character glyph classification encoding, and encoding for damaged or unrecognizable characters. This yields an encoding sequence consistent with the positional order of the original oracle bone script sentence, which serves as the foundation for subsequent vocabulary construction and model input processing.
[0062] A glyphic vocabulary is constructed based on a standard character set, a variant character set, and special symbols required for model processing. These special symbols include mask symbols, padding symbols, unknown symbols, sentence-initial symbols, and sentence-final symbols. The glyph codes in the vocabulary are mapped to the Unicode private area, enabling glyphs in oracle bone script not included in the general character set to participate in model training and inference using a unified character encoding. The encoded sequence mapped to the Unicode private area then enters character-level word segmentation processing, ensuring a one-to-one correspondence between the segmented input token and the character positions in the original oracle bone script sentence.
[0063] A merging mapping relationship is established between specific variant character forms and standard character categories. This merging mapping relationship is used to determine the target to be classified, merge inference results, and statistically analyze classification results. It is not directly used as a supervision label at the training position of the target variant character. Through this process, the training process retains the specific character form encoding and contextual distribution of the target variant character in the corpus, rather than having the model perform memory-based predictions solely based on a pre-set mapping relationship.
[0064] Training samples are constructed based on the encoded sequences after character-level word segmentation. In the training samples, mask positions are selected from ordinary non-empty standard character positions, and these selected standard character positions are replaced with mask symbols to generate a corrupted input sequence. Target variant characters, broken characters, or unrecognizable characters are not used as random mask positions. The corrupted input sequence serves as the encoder input, and the original encoded sequence serves as the target sequence to be recovered, enabling the model to learn the correspondence between the masked standard character positions from the context of the oracle bone script sentence.
[0065] The training samples are input into the context mask prediction model for training. The model performs context encoding on the corrupted input sequence to obtain an encoded representation containing context information of each input character position. The model then reconstructs the original encoded sequence position by position based on the encoded representation and the right-shifted target sequence, and outputs the prediction probability of each candidate token in the glyph vocabulary at each position to be predicted. During training, the error between the prediction result and the original standard character encoding is calculated only for the standard character mask position, and the model parameters are updated based on this error; target variant character positions, damaged character positions, or unrecognizable character positions are not included in the loss calculation.
[0066] In one specific training method, the difficulty level of training samples is determined based on the proportion of damaged or unrecognizable characters in oracle bone script sentences. The model is first trained using samples with a lower proportion of damaged characters and more complete context to obtain initial model parameters. Then, the model parameters obtained from the previous difficulty level are used as the initial parameters for the next difficulty level. Training samples with even higher proportions of damaged characters are then introduced for further training until the preset difficulty level is achieved, resulting in a completed model for variant character classification.
[0067] When performing reasoning on oracle bone script sentences to be classified, the sentence and the position of the target variant character are read. Following the same character shape classification encoding, Unicode private area mapping, and character-level word segmentation methods as in the training phase, the sentence is converted into a reasoning-based encoding sequence, and the position of the target variant character is replaced with a mask symbol, resulting in the reasoning-based input sequence. After the model is trained using this input sequence, the model outputs the predicted probability of each candidate token corresponding to the target position based on the context of the oracle bone script sentence before and after the target variant character's position.
[0068] Based on the predicted probabilities of candidate tokens at the target location, the candidate tokens are sorted and then uniformly merged into their corresponding standard character categories according to a merge mapping relationship. The standard character category with the highest predicted probability is selected as the classification result for the target variant character; alternatively, several standard character categories with the highest probability ranking are output as candidate classification results. This yields the standard character attribution information corresponding to the target variant character.
[0069] In an optional implementation, the context masking prediction model employs a sequence-to-sequence model with a bidirectional encoder and an autoregressive decoder. The bidirectional encoder reads the masked entire input sequence and forms a context representation containing information about the characters before and after the target position. The autoregressive decoder combines the given target sequence preceding the token and the encoder output to generate the recovered glyph encoding sequence character by character. The decoder reads the encoder output through a cross-attention mechanism, enabling the target position prediction result to be generated according to the importance of different context characters.
[0070] In an optional implementation, after the model outputs the predicted probability of each candidate token at the target location, in addition to outputting the standard character category with the highest probability, it also outputs Top-K candidate standard character categories. Comparing the Top-K candidate standard character categories with known standard character codes can be used to statistically analyze the hit rate under different candidate ranges and provide candidate basis for manual review or subsequent oracle bone inscription interpretation.
[0071] In an optional implementation, the merging mapping relationship includes the correspondence between the standard Label character set and the SubLabel character set. Candidate tokens obtained during the model inference phase can first be mapped to specific glyph codes, and then converted to standard Label categories according to the merging mapping relationship, ensuring that the model output is consistent with the standard character attribution results in the oracle bone script variant character classification task.
[0072] In an optional implementation, damaged or unrecognizable characters are encoded independently to participate in the input sequence representation, but are not used as mask prediction targets during the training phase. For oracle bone script sentences containing damaged or unrecognizable characters, the model can still use the remaining recognizable characters to form a contextual representation and predict the position of the target variant character during the inference phase.
[0073] Implementation Method 3, in conjunction with Appendix Figure 1-16 This embodiment describes the technical solution provided above in further detail through specific examples. Specifically: System Architecture Overview Oracle data processing For each oracle bone script sentence, the system reads its character position sequence. Let a sentence be:
[0074] in, Indicates the first in the sentence The character position, This indicates the word length of the sentence. The system maps the original character sequence to a glyph classification encoding sequence:
[0075] The encoding rules are as follows:
[0076] in Indicates the classification code of variant characters. This represents the standard character glyph classification encoding. Encoding for damaged or unrecognizable characters. The system also establishes... Formal merge mapping is used to merge specific variant glyphs into their corresponding standard character categories:
[0077] in, For the standard Label character set, This represents the set of SubLabel characters used in the modeling. It should be noted that this mapping is only used during target selection, post-inference merging, and result statistics; it is not used directly as a supervision label at the training positions of variant characters, thus avoiding the model simply memorizing manually generated mapping relationships.
[0078] The system constructs a glossary of characters. This includes the standard character set, the variant character set, and the special symbols required for the model:
[0079] in, For mask symbols, , , , These are symbols for filler, unknown, sentence beginning, and sentence end, respectively. The system maps the glyph codes in the above vocabulary to the Unicode private area, allowing the model to use Unicode encoding for training and inference. This encoding method avoids interference from a large number of unencoded and uninterpreted characters in Unicode, making the contextual information of oracle bone script sentences clearer.
[0080] Since oracle bone script sentences are typically short, averaging about 5 to 6 characters, it is necessary to accurately locate the position of the target variant character. Therefore, using SentencePiece character-level segmentation can ensure a one-to-one correspondence between the model's input token and the original character position, avoiding the shift of the target variant character position after segmentation.
[0081] During training sample construction, the system randomly selects a set of mask positions from ordinary non-empty Label positions. :
[0082] in, This represents the set of target variant character positions in the sentence. Target variant character positions and broken character positions are not used as random mask positions. The system generates a corrupted input sequence based on the mask position set:
[0083] This yields the input sequence for the BART model:
[0084] For oracle bone script sentences with a large number of damaged characters, the system further calculates the percentage of damaged characters and uses this percentage to determine the difficulty level of the course.
[0085] in, This is an indicator function. It returns 1 if the condition within the parentheses is true, and 0 otherwise. The system follows... The training samples are graded by difficulty based on intervals, for example, using 10% as the base interval for statistics, and can be combined into several difficulty levels such as low defect, medium defect, and high defect during course learning and training:
[0086] in, Indicates the preceding A dataset of difficulty levels This represents the threshold for the percentage of damaged characters corresponding to this level. During training, data of increasing difficulty are introduced gradually, starting with easier data and progressing to more challenging data.
[0087]
[0088] Training phase The system divides the input training corpus according to difficulty and trains it using the BART model. This model consists of a bidirectional encoder and an autoregressive decoder, capable of contextually encoding the input sequence after it has been masked and recovering the original glyph encoding sequence at the decoding end. Let the parameters of the BART model be... For input The model first obtains the context representation through the encoder. :
[0089] The decoder recovers the target sequence using an autoregressive approach. During training, a teacher-forcing mechanism is employed, using the right-shifted target sequence as the decoder input.
[0090] For the There are several positions to be predicted. The decoder is based on Masked Self-Attention, which means the target token given on the left and the encoder output. Calculate the hidden state at this location:
[0091] in, The decoder is in the first... The hidden representation is generated at each position. Internally, the decoder reads the encoder output through a cross-attention mechanism, allowing the prediction at the current position to utilize the overall context of the masked input sequence. Cross-attention can be represented as:
[0092]
[0093]
[0094] in, Indicates the first A query vector for each decoding position. and These represent the outputs from the encoder. The keys and values obtained from the mapping, Indicates the first The attention weights of each position for each token in the input sequence. This represents the context vector obtained through cross attention.
[0095] After obtaining the current position representation, the model maps it to the vocabulary dimension through the output layer to obtain the [current position]... The logits of each candidate token at each position. For any candidate token In its first The predicted probabilities at each position are calculated using softmax as follows:
[0096] in, and They represent candidate tokens respectively. Weights and biases in the output layer.
[0097] The training phase employs the masked positional cross-entropy loss function. For sentences... Its training loss is:
[0098] The loss only The mask position of the corresponding standard character label is calculated. The positions of the target variant characters and the positions of the damaged characters are characters whose meanings are not yet known to the model during the training phase, and no mask operation is performed, so they are not included in the loss calculation.
[0099] The system performs backpropagation based on the training loss, calculates the gradient of the model parameters, and updates the parameters. The system uses the loss metric on the validation set to determine if the early stopping criterion has been met. If the criterion has not been met, the next epoch begins training; if the criterion has been met, the next difficulty level begins training. Training.
[0100] This allows the model to first be implemented on a low-difficulty dataset. Training was performed to obtain parameters. ; then with Initialize model parameters on more challenging datasets. Continue training and get ;......; then with Initialize model parameters on more challenging datasets. Continue training and get ...and so on, until the optimal model parameters are obtained. :
[0101] In this way, the model first learns from samples with fewer damaged characters and more complete context, and then gradually adapts to samples with more damaged characters and weaker context information, thereby improving the model's stability in short oracle bone sentences and incomplete contexts.
[0102] Testing phase The system reads the location of the target variant character in the sample to be classified and replaces that location with a mask symbol. Then, the model generates candidate sequences based on the context of the entire sentence. The system outputs the token and merges it into a standard label, and then provides the final classification result.
[0103] The system reads the sample to be classified and the location of its target variant characters. Let the sentence to be classified be:
[0104] The target variant word position is The system uses the same encoding, Unicode private region mapping, and SentencePiece character-level segmentation as in the training phase to convert the sentence into an encoded sequence. Then, the target variant character position is replaced with a mask symbol to obtain the inference input:
[0105] The system will Input the trained BART model and at the target location The vocabulary list was obtained. Predicted probabilities of each candidate token:
[0106] The system sorts the candidate standard character categories according to the above probabilities and takes the category with the highest probability as the classification result for the variant character:
[0107] At the same time, the system can output Top-K candidate results:
[0108] The candidate results are compared with the standard glyph encoding to determine whether they are a match, and the Top-k accuracy is calculated to verify the model's predictive ability.
[0109] In terms of technical effectiveness: The context masking prediction method employed in this embodiment can classify oracle bone script variant characters using corpus contextual information without directly supervising the mapping from target variant characters to standard characters. Compared to most computer vision-based methods, this embodiment preserves the contextual distribution of specific character forms; compared to manual variant character classification methods that rely on expert experience, this embodiment can reduce the time and experience costs of manual classification to a certain extent, and provides a new solution for the problem of classifying unknown variant characters.
[0110] Furthermore, this implementation method is not only applicable to the classification of variant characters in oracle bone script, but logically it can also be extended to other ancient script corpora with variant character classification problems, such as bronze inscriptions and bamboo and silk manuscripts.
[0111] Results Display: The following tests were conducted using recently deciphered variant characters and typical variant characters:
[0112] Summary of results:
[0113] The accuracy rates for classifying the 10 variant character examples listed above are as follows:
[0114] The above description of several specific embodiments further details the technical solution provided by the present invention in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above-described specific embodiments are not intended to limit the present invention. Any reasonable modifications and improvements to the present invention, combinations of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for classifying variant characters in oracle bone script based on context mask prediction, characterized in that, include: The steps of obtaining oracle bone sentences from oracle bone script corpus and converting the oracle bone sentences into character shape classification coding sequences according to the character position order; The training samples are constructed based on the character shape classification coding sequence, wherein the character position with a determined standard character category is replaced with a mask coding, the replaced character shape classification coding sequence is used as the model input, the original character shape classification coding sequence is used as the model output, and the target variant character position is not used as the mask position. The steps are as follows: train a context mask prediction model using the training samples, and make the context mask prediction model predict the glyph classification code corresponding to the mask position based on the word context before and after the mask position, thus obtaining the trained context mask prediction model. The steps include: obtaining the oracle bone script sentence to be classified and the target variant character positions within it, and converting the oracle bone script sentence to be classified into a character shape classification encoding sequence to be inferred according to the same character shape classification encoding rules; The steps are as follows: replace the character shape classification code corresponding to the target variant character position in the character shape classification coding sequence to be reasoned with the mask code, and input it into the trained context mask prediction model to obtain the candidate character shape classification code corresponding to the target variant character position; The steps are as follows: determining the corresponding standard character category based on the candidate character shape classification code, and outputting the standard character category as the classification result of the target variant character position.
2. The method for classifying variant characters in oracle bone script based on context mask prediction according to claim 1, characterized in that, The character shape classification coding sequence includes standard character shape classification coding, variant character shape classification coding, and damaged or unrecognizable character coding. The character shape classification coding corresponds one-to-one with the oracle bone characters in the oracle bone script sentence according to the character position order.
3. The method for classifying variant characters in oracle bone script based on context mask prediction according to claim 1, characterized in that, The character-level word segmentation is performed on the character shape classification encoding sequence so that the input token obtained by word segmentation corresponds one-to-one with the oracle bone characters in the oracle bone sentence. The segmented encoding sequence is used as the basis for training samples and processing of the character shape classification encoding sequence to be inferred.
4. The method for classifying variant characters in oracle bone script based on context mask prediction according to claim 1, characterized in that, The training samples generate mask positions based on the characters with defined standard character categories. The characters of target variant characters, damaged characters, and unrecognizable characters are not used as mask positions and are not included in the prediction error calculation when updating model parameters.
5. The method for classifying variant characters in oracle bone script based on context mask prediction according to claim 1, characterized in that, The context mask prediction model performs whole-sentence context encoding on the mask input sequence and generates the predicted probability of each candidate glyph classification code corresponding to the mask position based on the whole-sentence context encoding result.
6. The method for classifying variant characters in oracle bone script based on context mask prediction according to claim 1, characterized in that, The candidate character shape classification codes corresponding to the target variant character position are sorted according to the predicted probability, and the candidate character shape classification code with the highest predicted probability is selected as the target candidate character shape classification code.
7. A device for classifying variant characters in oracle bone script based on context mask prediction, characterized in that, include: A module that acquires oracle bone sentences from oracle bone script corpus and converts the oracle bone sentences into character shape classification encoding sequences according to the character position order; Training samples are constructed based on the character shape classification coding sequence, wherein the character position with a defined standard character category is replaced with a mask coding, the replaced character shape classification coding sequence is used as the model input, the original character shape classification coding sequence is used as the model output, and the target variant character position is not used as the mask position. The training samples are used to train a context mask prediction model, which predicts the glyph classification code corresponding to the mask position based on the word context before and after the mask position, thus obtaining the module of the trained context mask prediction model. A module that acquires the oracle bone script sentence to be classified and the target variant character positions therein, and converts the oracle bone script sentence to be classified into a character shape classification encoding sequence to be deduced according to the same character shape classification encoding rules; The glyph classification code corresponding to the target variant character position in the glyph classification code sequence to be inferred is replaced with a mask code, and input into the trained context mask prediction model to obtain the module of candidate glyph classification code corresponding to the target variant character position; The module determines the corresponding standard character category based on the candidate character shape classification code and outputs the standard character category as the classification result of the target variant character position.
8. A computer storage medium for storing computer programs, characterized in that, When the computer program is read by the computer, the computer executes the method of claim 1.
9. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method of claim 1.
10. A computer program product, as a computer program, is characterized by: When the computer program is executed, it implements the method of claim 1.