A method and system for deciphering oracle bone characters based on contrastive learning and visual ideographic description sequences

Through comparative learning and visual representational description methods, combined with oracle component recognition and feature enhancement, the problem of high cost and insufficient label-side information in oracle character decoding is solved, achieving end-to-end efficient decoding and recognition of unknown categories.

CN119625753BActive Publication Date: 2025-09-02SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411671924.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-09-02
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The existing oracle character deciphering methods have high implementation costs and insufficient information mining on the label side, and it is difficult to effectively use the ideographic description sequence (IDS) of modern Chinese characters to assist end-to-end deciphering.

Method used

Using a method based on contrast learning and visual representation description sequences, the oracle component recognizer is trained to extract features, combined with visual encoder and feature enhancer, feature enhancement is used to use IDS interquery technology, and deciphering results are obtained through contrast learning.

Benefits of technology

It realizes the effective identification of unseen Oracle components in the case of scarce training data, improves the end-to-end capability of oracle characterization and IDS representation ability, and avoids the problem of missing clues in the middle stage of evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625753B_ABST
    Figure CN119625753B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of pattern recognition and artificial intelligence, and discloses a method and system for deciphering oracle bone characters based on contrastive learning and visual ideographic description sequences. The method comprises: using a trained oracle bone character component identifier to extract features of an input oracle bone character image component to be deciphered, so as to obtain visual embedding features of the oracle bone character image; using a visual encoder to extract overall visual features of the input oracle bone character image to be deciphered and a modern Chinese character set, so as to obtain overall visual features of the oracle bone character and overall visual features of the Chinese character; using a feature enhancer based on IDS mutual inquiry to represent and enhance the visual embedding features of the oracle bone character image and the overall visual features of the oracle bone character, the visual embedding features of the Chinese character image and the overall visual features of the Chinese character, so as to obtain enhanced features of the oracle bone character and enhanced features of the Chinese character; using contrastive learning to compare and calculate the enhanced features of the oracle bone character and the enhanced features of the Chinese character, so as to obtain a contrast vector, and obtaining an oracle bone character deciphering result based on the contrast vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pattern recognition and artificial intelligence, and in particular to an oracle bone inscription deciphering method and system based on contrastive learning and visual ideographic description sequence. Background Art

[0002] As one of the oldest pictographic scripts in the world, the recognition, composition and deciphering of oracle bone script characters have received increasing attention from researchers in recent years. These ancient characters are key to exploring ancient China (especially the Shang Dynasty culture, approximately 1600-1046 BC), but most of the discovered categories have not yet been deciphered. Deciphering these characters will not only help their preservation, but also contribute to the inheritance of cultural heritage. Existing research on the interpretation of oracle bone characters mainly focuses on reasoning and reconstruction of evolutionary stages. Previously, researchers used generative models such as CycleGAN to simulate the evolutionary process of oracle bone characters. At each intermediate stage of evolution (such as bronze inscriptions, seal scripts, etc.), they generated reconstructed glyphs as input clues for subsequent stages to advance their simulation process.

[0003] However, this type of method has two major limitations: (1) The method process involves expensive and complex implementation costs. It requires the compilation of ground truth data from all intermediate evolutionary stages from oracle bone inscriptions to modern Chinese characters for supervision. Due to the extensive and complex nature of the evolutionary process, many character categories may have key clues that are blurred or lost in the intermediate stages, so this method often fails in this scenario. (2) The information mining on the label side is not deep enough. In this type of deciphering method, the utilization of the label side (i.e., modern Chinese characters) is only introduced in the form of glyph images in the final evolutionary stage. However, for the complex and challenging oracle bone deciphering task, other important clues of modern Chinese characters on the label side should also be deeply mined to assist in deciphering, such as ideographic description sequences (IDS). IDS includes components and structural symbols, which can provide valuable knowledge and clues for understanding the structural information and composition details of Chinese characters. Therefore, how to overcome the problem of insufficient evolutionary clues that the existing methods are limited to, and how to introduce IDS knowledge to assist model reasoning and complete end-to-end oracle bone deciphering has become an urgent problem that needs to be solved today. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides an oracle bone inscription deciphering method based on contrastive learning and visual ideographic description sequence, the method comprising:

[0005] Step S1: using a trained oracle bone script component identifier to extract features of the input oracle bone script image components to be deciphered, to obtain visual embedding features of the oracle bone script components;

[0006] Step S2: using a visual encoder to extract overall visual features from the input oracle bone inscription image to be deciphered and the modern Chinese character set, to obtain overall visual features of the oracle bone inscription and overall visual features of the Chinese characters;

[0007] Step S3, using a feature enhancer based on IDS mutual query to characterize and enhance the oracle bone character component visual embedding features and the oracle bone character overall visual features, the Chinese character component visual embedding features and the Chinese character overall visual features, to obtain oracle bone character enhanced features and Chinese character enhanced features;

[0008] Step S4: using contrastive learning to perform contrast calculation on the oracle bone script enhanced features and the Chinese character enhanced features to obtain a contrast vector, and obtaining an oracle bone script deciphering result based on the contrast vector.

[0009] Optionally, in step S1, the training process of the oracle bone script component identifier specifically includes:

[0010] A visual encoder is used to extract the overall visual features of the oracle bone script image to be deciphered;

[0011] A visual embedding layer is used to extract visual embedding features from the component set of the oracle bone script image to be deciphered and the visual IDS of the component set;

[0012] Using a contrastive learner based on cosine distance, the similarity between the feature sequence of the overall visual features of the oracle bone script image to be deciphered and the feature sequence of the visual embedding features is calculated, and a contrast vector is obtained based on the similarity;

[0013] A loss value of the comparison vector is calculated using a counting loss function, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

[0014] Optionally, in step S1, the process of extracting features of the input oracle bone script image components to be deciphered using a trained oracle bone script component identifier to obtain visual embedding features of the oracle bone script components specifically includes:

[0015] The visualization embedding layer of the oracle bone script component identifier includes a CNN feature extractor, a one-hot encoder, and a fused convolutional layer;

[0016] Inputting the IDS of the oracle bone character image component to be deciphered into a one-hot encoder to obtain the IDS feature of the oracle bone character component;

[0017] Inputting the IDS of the oracle bone inscription image component to be deciphered into a CNN feature extractor to obtain visual IDS features of the oracle bone inscription component;

[0018] The oracle bone character component IDS features and the oracle bone character component visual IDS features are input into the fusion convolution layer for feature fusion to obtain the oracle bone character component visual embedding features.

[0019] Optionally, in step S3, the process of using a feature enhancer based on IDS mutual query to characterize and enhance the oracle bone inscription image visual embedding features and the oracle bone inscription overall visual features, the Chinese character image visual embedding features and the Chinese character overall visual features to obtain the oracle bone inscription enhanced features and the Chinese character enhanced features specifically includes:

[0020] Fusing the oracle bone inscription image visual embedding features with the overall visual features of the oracle bone inscription to obtain oracle bone inscription fusion features;

[0021] Fusing the Chinese character image visual embedding features with the overall visual features of the Chinese character to obtain a Chinese character fusion feature;

[0022] Use the tag dictionary to perform one-hot encoding on the TopK candidate items of the comparison vector and the IDS of the oracle bone script fusion feature and the IDS of the Chinese character fusion feature to obtain the TopK candidate item embedding features and the label item embedding features;

[0023] Preprocessing the tag item embedding features and the Chinese character visual overall features after splicing to obtain tag item fusion features;

[0024] Preprocessing the embedded features of the TopK candidate items and the overall visual features of the oracle bone inscriptions to obtain the fused features of the TopK candidate items;

[0025] The tag item fusion feature and the TopK candidate item fusion feature are dually input into the IDS mutual query based feature enhancer to output Chinese character enhancement features and oracle bone script enhancement features.

[0026] The present invention also discloses an oracle bone inscription deciphering system based on contrastive learning and visual ideographic description sequence, the system comprising:

[0027] The component feature extraction module is used to extract features of the components of the input oracle bone script image to be deciphered using the trained oracle bone script component identifier to obtain visual embedding features of the oracle bone script components;

[0028] The overall feature extraction module is used to extract overall visual features from the input oracle bone inscription image to be deciphered and the modern Chinese character set using a visual encoder to obtain overall visual features of the oracle bone inscription and the Chinese character set;

[0029] A feature enhancement module is used to enhance the visual embedding features of the oracle bone inscription components and the overall visual features of the oracle bone inscriptions, and the visual embedding features of the Chinese character components and the overall visual features of the Chinese character using a feature enhancer based on IDS mutual query, so as to obtain enhanced features of the oracle bone inscriptions and enhanced features of the Chinese character;

[0030] The oracle bone deciphering module is used to perform comparative calculation on the oracle bone inscription enhanced features and the Chinese character enhanced features using contrastive learning to obtain a contrast vector, and obtain an oracle bone inscription deciphering result based on the contrast vector.

[0031] Optionally, in the component feature extraction module, the training process of the Oracle component identifier specifically includes:

[0032] A visual encoder is used to extract the overall visual features of the oracle bone script image to be deciphered;

[0033] A visual embedding layer is used to extract visual embedding features from the component set of the oracle bone script image to be deciphered and the visual IDS of the component set;

[0034] Using a contrastive learner based on cosine distance, the similarity between the feature sequence of the overall visual features of the oracle bone script image to be deciphered and the feature sequence of the visual embedding features is calculated, and a contrast vector is obtained based on the similarity;

[0035] A loss value of the comparison vector is calculated using a counting loss function, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

[0036] Optionally, the workflow of the component feature extraction module specifically includes:

[0037] The visualization embedding layer of the oracle bone script component identifier includes a CNN feature extractor, a one-hot encoder, and a fused convolutional layer;

[0038] Inputting the IDS of the oracle bone character image component to be deciphered into a one-hot encoder to obtain the IDS feature of the oracle bone character component;

[0039] Inputting the IDS of the oracle bone inscription image component to be deciphered into a CNN feature extractor to obtain visual IDS features of the oracle bone inscription component;

[0040] The oracle bone character component IDS features and the oracle bone character component visual IDS features are input into the fusion convolution layer for feature fusion to obtain the oracle bone character component visual embedding features.

[0041] Optionally, the workflow of the feature enhancement module specifically includes:

[0042] Fusing the oracle bone inscription image visual embedding features with the overall visual features of the oracle bone inscription to obtain oracle bone inscription fusion features;

[0043] Fusing the Chinese character image visual embedding features with the overall visual features of the Chinese character to obtain a Chinese character fusion feature;

[0044] Use the tag dictionary to perform one-hot encoding on the TopK candidate items of the comparison vector and the IDS of the oracle bone script fusion feature and the IDS of the Chinese character fusion feature to obtain the TopK candidate item embedding features and the label item embedding features;

[0045] Preprocessing the tag item embedding features and the Chinese character visual overall features after splicing to obtain tag item fusion features;

[0046] Preprocessing the embedded features of the TopK candidate items and the overall visual features of the oracle bone inscriptions to obtain the fused features of the TopK candidate items;

[0047] The tag item fusion feature and the TopK candidate item fusion feature are dually input into the IDS mutual query based feature enhancer to output Chinese character enhancement features and oracle bone script enhancement features.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] The present invention predicts the deciphering relationship based on the feature comparison algorithm between the oracle bone inscriptions and the modern Chinese characters, which can effectively avoid the problem of missing clues in the intermediate stage of evolution, so that the entire model can be trained end-to-end and output the final deciphering result; the oracle bone inscription component identifier based on visual IDS and contrastive learning can effectively identify components of unseen categories, which has positive significance in the oracle bone inscription scenario with scarce training data; the feature enhancement based on IDS mutual inquiry can fully expand the evaluation range of TopK component candidates, effectively improving the IDS representation capability of oracle bone inscriptions. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 A diagram showing the steps of a method for deciphering oracle bone characters based on contrastive learning and visual ideographic description sequences according to an embodiment of the present invention;

[0052] Figure 2 This is a pre-training flow chart of the oracle bone script component identifier according to an embodiment of the present invention;

[0053] Figure 3 This is a flow chart of a feature enhancer based on IDS mutual query according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.

[0056] First, some technical terms used in this invention are explained:

[0057] IDS (Ideographic Description Sequence) is a method used to describe the structure of Chinese characters, particularly when digitizing, decomposing, reconstructing, or analyzing them. IDS is part of the Unicode encoding standard and is primarily used to represent structural information about complex or non-directly encoded Chinese characters. It uses a set of special symbols and basic components to describe the composition and arrangement of Chinese characters. Key Features: 1. Structured Representation: IDS is component-based and uses descriptors (such as left-right structure, top-bottom structure, etc.) to break down Chinese characters into their basic building blocks, providing a modular, structured representation of Chinese characters.

[0058] Descriptors support a variety of structural relationships:

[0059] IDS uses specific descriptors, such as:

[0060] Indicates left-right structure

[0061] Indicates a top-down structure

[0062] Represents a cross structure

[0063] These descriptors can be flexibly combined to express the diverse structures of Chinese characters. 3. Applicable to non-encoded Chinese characters: IDS is particularly suitable for rare or newly created Chinese characters that are not included in the Unicode encoding. By combining descriptors and components, users can create digital representations of these characters. 4. Support for Chinese character component decomposition and recognition: IDS facilitates component-level decomposition and analysis of Chinese characters, facilitating computer-based tasks such as pattern recognition, Chinese character learning, and glyph retrieval. It also provides technical support for philological research and the compilation of historical documents.

[0064] Contrastive learning is an unsupervised or self-supervised learning method whose primary goal is to extract useful feature representations by learning the similarities and differences between samples. This method has been widely used in fields such as computer vision, natural language processing, and recommender systems. Basic Principles: 1. Positive and Negative Pairs: A positive pair consists of two samples that are similar in some sense, such as from the same category or different perspectives of the same instance. A negative pair consists of two samples that are different in some sense, such as from different categories or different instances. 2. Loss Function: Contrastive learning optimizes model parameters by designing a specific loss function to make the feature representations of the positive pair closer and the feature representations of the negative pair farther apart. Common loss functions include contrastive loss and triplet loss. Common Contrastive Learning Methods: 1. SimCLR: SimCLR is a self-supervised learning method that generates positive pairs by applying different data augmentation techniques to the same image. It then optimizes the model using a contrastive loss function. The specific steps include data augmentation, feature extraction, projection head, and contrastive loss calculation. 2. MoCo (Momentum Contrast): MoCo addresses the limited number of negative samples by maintaining a momentum-updated queue to store them. This momentum update mechanism gradually updates the negative samples in the queue, maintaining both diversity and stability. 3. BYOL (Bootstrap Your Own Latent): BYOL does not use negative sample pairs. Instead, it learns feature representations using two independent encoders: an online encoder and a target encoder. The online encoder updates its parameters by optimizing the predicted output of the target encoder, while the target encoder is gradually updated using a momentum update mechanism. 4. The "HUST-OBS" dataset is used to evaluate and compare the performance of different OCR algorithms, particularly in the field of handwritten Chinese character recognition. It is a commonly used dataset for oracle bone script deciphering tasks. 5. CLIP (Contrastive Language-Image Pretraining) is a multimodal model proposed by OpenAI that jointly trains text and image embeddings through contrastive learning. It leverages large-scale image-text pair data to achieve cross-modal matching and efficiently recognize image content in zero-shot tasks. CLIP performs well in unsupervised learning and multi-task adaptation, and has the ability to recognize unseen categories.

[0065] Application Scenarios: 1. Image Recognition: In computer vision, contrastive learning can help models learn more discriminative feature representations, improving performance in tasks such as image classification and object detection. 2. Natural Language Processing: In natural language processing, contrastive learning can be used for tasks such as word embedding and sentence representation, helping models better understand the semantics of text. 3. Recommender Systems: In recommendation systems, contrastive learning can be used to learn feature representations for users and items, improving the accuracy and personalization of recommendations.

[0066] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0067] Example 1

[0068] A method for deciphering oracle bone characters based on contrastive learning and visual ideographic description sequences, such as Figure 1 As shown, the method includes:

[0069] Step S1: Use the trained oracle bone script component identifier to extract features of the input oracle bone script image components to be deciphered, and obtain visual embedding features of the oracle bone script components.

[0070] like Figure 2 As shown in the figure, the training process of the oracle bone script component identifier specifically includes: using a visual encoder based on DenseNet-44 to extract visual features from the input oracle bone script image, and the feature is recorded as F rad ; Use the visualization embedding layer VE(·) to embed the component set π of the existing oracle bone inscription image and the visual IDS(Q π ) Extract visual embedding features G vt ;

[0071] G vt =VE(π,Q π )=Conv(Enc or (Q π )||Embed(π))

[0072] Among them, Conv(·), Enc or (·), Embed(·), and (·||·) refer to a 1*1 convolutional layer, a ResNet-18 network, a one-hot encoder, and a vector concatenation operation, respectively. The implementation of visual IDS here directly uses the skeleton image corresponding to the component.

[0073] Use the contrastive learner based on cosine distance to calculate the visual feature F of the existing oracle bone inscription image rad And the visual embedding feature G vt The cosine distance of , and the comparison vector is obtained based on the cosine distance;

[0074]

[0075] Where M and T represent the preset component set size and F respectively. rad The length after vectorization.

[0076] A counting loss function is used to calculate a loss value of the comparison vector, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

[0077] The process of using a trained oracle bone script component identifier to extract features of the input oracle bone script image component to be deciphered to obtain the visual embedding features of the oracle bone script component specifically includes: the visual embedding layer of the oracle bone script component identifier includes a CNN feature extractor, a one-hot encoder and a fusion convolutional layer; the IDS of the oracle bone script image component to be deciphered is input into the one-hot encoder to obtain the oracle bone script component IDS features; the IDS of the oracle bone script image component to be deciphered is input into the CNN feature extractor to obtain the oracle bone script component visual IDS features; the oracle bone script component IDS features and the oracle bone script component visual IDS features are input into the fusion convolutional layer for feature fusion to obtain the oracle bone script component visual embedding features.

[0078] Step S2: Using a visual encoder to extract overall visual features from the input oracle bone inscription image and the modern Chinese character set to be deciphered, obtaining overall visual features of the oracle bone inscription and the Chinese character set. The IDs of the modern Chinese character set are input into a one-hot encoder to obtain modern Chinese character IDS features; the IDS of the visual features are input into a CNN feature extractor to obtain visual IDS features; the modern Chinese character IDS features and the visual IDS features are input into the fusion convolutional layer for feature fusion to obtain overall visual features of the oracle bone inscription and the Chinese character set.

[0079] The DenseNet-44 framework is used to construct two parameter-independent visual encoders to extract visual features from the input oracle bone script image and the overall glyph graph of the modern Chinese character set, respectively.

[0080] Step S3: Use the feature enhancer based on IDS mutual query to characterize and enhance the oracle bone character component visual embedding features and the oracle bone character overall visual features, the Chinese character component visual embedding features and the Chinese character overall visual features to obtain oracle bone character enhanced features and Chinese character enhanced features.

[0081] like Figure 3As shown, the visual embedding features of the oracle bone script components and the overall visual features of the oracle bone script are fused to obtain the oracle bone script fusion features; the visual embedding features of the Chinese character components and the overall visual features of the Chinese character are fused to obtain the Chinese character fusion features; the tag dictionary is used to perform one-hot encoding on the TopK candidates of the comparison vector and the IDS of the token features to obtain the TopK candidate embedding features and the label item embedding features; the tag dictionary is used to perform one-hot encoding on the output TopK component candidates (oracle bone script) and IDS label items (Chinese characters), wherein the role of the tag dictionary is to apply the existing deciphering knowledge so that the tags of the determined oracle bone script components and the corresponding modern Chinese character components are consistent, so that the one-hot encodings generated by the two are consistent;

[0082] The tag item embedding features and token features are concatenated and preprocessed to obtain the tag item fusion features. A linear mapping layer is used to map the output TopK component candidate embeddings (oracle bone characters) and tag item embeddings (Chinese characters). The linear mapping layer is composed of a fully connected layer.

[0083] The embedding features of the TopK candidate items are concatenated with the token features and preprocessed to obtain the fusion features of the TopK candidate items; the embedding features and token features of the label items are concatenated (denoted as C), reshaped (denoted as R) and dimensionally expanded (denoted as E) to obtain the fusion features of the label items (Chinese characters), denoted as Where N, L N , K and C represent the number of labels, the preset maximum length of Chinese character IDS, the preset number of candidates and the hidden feature length of Transformer respectively; similarly, the embedding features and token features of TopK candidates are spliced ​​and reshaped to obtain the TopK candidate fusion features (oracle bone characters), which are recorded as Where M and L M They represent the size of the Oracle component set and the preset maximum length of the Oracle IDS respectively; according to experimental results, the setting of K is 3 to achieve better results;

[0084] The dual input of the label item fusion feature and the TopK candidate item fusion feature is input into the feature enhancer based on IDS mutual inquiry, and the oracle bone character enhanced feature and the Chinese character enhanced feature are output. The Chinese character enhanced feature and the TopK candidate item fusion feature (oracle bone character enhanced feature) are used as query (denoted as Q), key (denoted as K) and value (denoted as V) respectively, and input into two Transformer-based mutual inquiry enhancers (with the same internal design, but the parameters are not shared) for mutual inquiry modeling. The two Transformers are each designed with an encoder layer (Encoder Layer) and a decoder layer (DecoderLayer). When the number of layers is set to 2, the experimental effect is better; the label item enhanced feature (Chinese character) obtained after the reshaping operation and self-attention calculation is denoted as The input item enhancement feature (oracle bone characters) obtained after the reshaping operation and the linear pooling layer is denoted as

[0085] Step S4: using contrastive learning to perform contrast calculation on the oracle bone script enhanced features and the Chinese character enhanced features to obtain a contrast vector, and obtaining an oracle bone script deciphering result based on the contrast vector.

[0086] The contrast model in the cosine distance-based contrast learning module of the present invention includes the following steps:

[0087] The cosine distance-based comparison model has two steps. The first step is to splice the label item enhanced features (Chinese characters, i.e., Chinese character enhanced features) and the output features of the visual encoder of the corresponding Chinese character glyph image, and at the same time splice the input item enhanced features (oracle bone characters, i.e., oracle bone character enhanced features) and the output features of the visual encoder of the oracle bone characters; the second step is to use the cosine distance to calculate the similarity between the input features after splicing in the first step and the label end features to obtain the similarity vector; the cross entropy is used to calculate the loss value of the similarity vector of the previous step and its true value, and the model is optimized based on the loss value; in the inference stage, the argmax function is used to calculate the similarity vector, and the calculation result is the final decryption result.

[0088] Example 2

[0089] A system for deciphering oracle bone characters based on contrastive learning and visual ideographic description sequences, comprising:

[0090] The component feature extraction module is used to extract features of the components of the input oracle bone script image to be deciphered using the trained oracle bone script component identifier to obtain visual embedding features of the oracle bone script components;

[0091] The training process of the oracle bone script component recognition includes: using the DenseNet-44-based visual encoder to extract visual features from the input oracle bone script image, and the feature is recorded as Frad ; Use the visualization embedding layer VE(·) to embed the component set π of the existing oracle bone inscription image and the visual IDS(Q π ) Extract visual embedding features G vt ;

[0092] G vt =VE(π,Q π )=Conv(Enc or (Q π )||Embed(π))

[0093] Among them, Conv(·), Enc or (·), Embed(·), and (·||·) refer to a 1*1 convolutional layer, a ResNet-18 network, a one-hot encoder, and a vector concatenation operation, respectively. The implementation of visual IDS here directly uses the skeleton image corresponding to the component.

[0094] Use the contrastive learner based on cosine distance to calculate the visual feature F of the existing oracle bone inscription image rad And the visual embedding feature G vt The cosine distance of , and the comparison vector is obtained based on the cosine distance;

[0095]

[0096] Where M and T represent the preset component set size and F respectively. rad The length after vectorization.

[0097] A counting loss function is used to calculate a loss value of the comparison vector, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

[0098] The overall feature extraction module uses a visual encoder to extract overall visual features from the input oracle bone inscriptions to be deciphered and the modern Chinese character set to obtain overall visual features of the oracle bone inscriptions and the overall visual features of the Chinese characters. Specifically, the module includes: a visual embedding layer including a CNN feature extractor, a one-hot encoder and a fusion convolution layer; the IDS of the modern Chinese character set is input into the one-hot encoder to obtain the IDS features of the modern Chinese characters; the IDS of the visual features is input into the CNN feature extractor to obtain the visual IDS features; the IDS features of the modern Chinese characters and the visual IDS features are input into the fusion convolution layer for feature fusion to obtain the token features.

[0099] The DenseNet-44 framework is used to construct two visual encoders with independent parameters to extract visual features from the input oracle bone script image and the glyph graph of the modern Chinese character set respectively.

[0100] The feature enhancement module is used to use a feature enhancer based on IDS mutual inquiry to characterize and enhance the visual embedding features of the oracle bone character components and the overall visual features of the oracle bone character, the visual embedding features of the Chinese character components and the overall visual features of the Chinese character, so as to obtain enhanced features of the oracle bone character and enhanced features of the Chinese character.

[0101] The oracle bone script image visual embedding feature and the oracle bone script overall visual feature are fused to obtain the oracle bone script fusion feature; the Chinese character image visual embedding feature and the Chinese character overall visual feature are fused to obtain the Chinese character fusion feature; the TopK candidate items of the comparison vector and the IDS of the oracle bone script fusion feature and the IDS of the Chinese character fusion feature are one-hot encoded using a tag dictionary to obtain the TopK candidate item embedding feature and the label item embedding feature; the output TopK component candidate items (oracle bone script) and the IDS label item (Chinese character) are one-hot encoded using a tag dictionary, wherein the role of the tag dictionary is to apply existing deciphering knowledge so that the tags of the determined oracle bone script components and the corresponding modern Chinese character components are consistent, so that the one-hot encodings generated by the two are consistent;

[0102] The tag item embedding features and token features are concatenated and preprocessed to obtain the tag item fusion features. A linear mapping layer is used to map the output TopK component candidate embeddings (oracle bone characters) and tag item embeddings (Chinese characters). The linear mapping layer is composed of a fully connected layer.

[0103] The embedding features of the TopK candidate items are concatenated with the token features and preprocessed to obtain the fusion features of the TopK candidate items; the embedding features and token features of the label items are concatenated (denoted as C), reshaped (denoted as R) and dimensionally expanded (denoted as E) to obtain the fusion features of the label items (Chinese characters), denoted as Where N, L N , K and C represent the number of labels, the preset maximum length of Chinese character IDS, the preset number of candidates and the hidden feature length of Transformer respectively; similarly, the embedding features and token features of TopK candidates are spliced ​​and reshaped to obtain the TopK candidate fusion features (oracle bone characters), which are recorded as Where M and L M They represent the size of the Oracle component set and the preset maximum length of the Oracle IDS respectively; according to experimental results, the setting of K is 3 to achieve better results;

[0104] Based on the tag item fusion feature and the TopK candidate item fusion feature, the enhanced original output feature and the enhanced token feature are obtained. The tag item fusion feature and the TopK candidate item fusion feature are used as Query (denoted as Q), Key (denoted as K) and Value (denoted as V) respectively, and input into two Transformer-based mutual query enhancers (with the same internal design, but the parameters are not shared) for mutual query modeling. The two Transformers each design an encoder layer (Encoder Layer) and a decoder layer (Decoder Layer). When the number of layers is set to 2, the experimental effect is better; the tag item enhanced feature (Chinese character) obtained after the reshaping operation and self-attention calculation is denoted as The input item enhancement feature (oracle bone characters) obtained after the reshaping operation and the linear pooling layer is denoted as

[0105] The oracle bone deciphering module is used to perform comparative calculation on the oracle bone inscription enhanced features and the Chinese character enhanced features using contrastive learning to obtain a contrast vector, and obtain an oracle bone inscription deciphering result based on the contrast vector.

[0106] The contrast model in the cosine distance-based contrast learning module of the present invention includes the following steps:

[0107] The cosine distance-based comparison model has two steps. The first step is to splice the label item enhanced features (Chinese characters) and the output features of the visual encoder of the corresponding Chinese character glyph image, and at the same time splice the input item enhanced features (oracle bone characters) and the output features of the visual encoder of the oracle bone characters; the second step is to use the cosine distance to calculate the similarity between the input features after splicing in the first step and the label end features to obtain the similarity vector; the cross entropy is used to calculate the loss value of the similarity vector of the previous step and its true value, and the model is optimized based on the loss value; in the inference stage, the argmax function is used to calculate the similarity vector, and the calculation result is the final decryption result.

[0108] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for deciphering oracle bone inscriptions based on contrastive learning and visual ideographic description sequences, characterized in that: The method comprises: Step S1: using a trained oracle bone script component identifier to extract features of the input oracle bone script image components to be deciphered, to obtain visual embedding features of the oracle bone script components; Step S2: using a visual encoder to extract overall visual features from the input oracle bone inscription image to be deciphered and the modern Chinese character set, to obtain overall visual features of the oracle bone inscription and overall visual features of the Chinese characters; Step S3: Using a feature enhancer based on IDS mutual query to enhance the visual embedding features of the oracle bone character components and the overall visual features of the oracle bone character, the visual embedding features of the Chinese character components and the overall visual features of the Chinese character, to obtain enhanced features of the oracle bone character and enhanced features of the Chinese character, specifically including: Fusing the oracle bone inscription image visual embedding features with the overall visual features of the oracle bone inscription to obtain oracle bone inscription fusion features; Fusing the Chinese character image visual embedding features with the overall visual features of the Chinese character to obtain a Chinese character fusion feature; Use the tag dictionary to perform one-hot encoding on the TopK candidate items of the comparison vector and the IDS of the oracle bone script fusion feature and the IDS of the Chinese character fusion feature to obtain the TopK candidate item embedding features and the label item embedding features; Preprocessing the tag item embedding features and the Chinese character visual overall features after splicing to obtain tag item fusion features; Preprocessing the embedded features of the TopK candidate items and the overall visual features of the oracle bone inscriptions to obtain the fused features of the TopK candidate items; Based on the dual input of the label item fusion feature and the TopK candidate item fusion feature, the feature enhancer based on IDS mutual query is outputted as Chinese character enhancement feature and Oracle bone character enhancement feature; Step S4: using contrastive learning to perform contrast calculation on the oracle bone script enhanced features and the Chinese character enhanced features to obtain a contrast vector, and obtaining an oracle bone script deciphering result based on the contrast vector.

2. The oracle bone inscription deciphering method based on contrastive learning and visual ideographic description sequence according to claim 1 is characterized in that: In step S1, the training process of the oracle bone script component identifier specifically includes: A visual encoder is used to extract the overall visual features of the oracle bone script image to be deciphered; A visual embedding layer is used to extract visual embedding features from the component set of the oracle bone script image to be deciphered and the visual IDS of the component set; Using a contrastive learner based on cosine distance, the similarity between the feature sequence of the overall visual features of the oracle bone script image to be deciphered and the feature sequence of the visual embedding features is calculated, and a contrast vector is obtained based on the similarity; A counting loss function is used to calculate a loss value of the comparison vector, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

3. The oracle bone inscription deciphering method based on contrastive learning and visual ideographic description sequence according to claim 2, characterized in that: In step S1, the process of extracting features of the input oracle bone script image components to be deciphered using the trained oracle bone script component identifier to obtain visual embedding features of the oracle bone script components specifically includes: The visualization embedding layer of the oracle bone script component identifier includes a CNN feature extractor, a one-hot encoder, and a fused convolutional layer; Inputting the IDS of the oracle bone character image component to be deciphered into a one-hot encoder to obtain the IDS feature of the oracle bone character component; Inputting the IDS of the oracle bone inscription image component to be deciphered into a CNN feature extractor to obtain visual IDS features of the oracle bone inscription component; The oracle bone character component IDS features and the oracle bone character component visual IDS features are input into the fusion convolution layer for feature fusion to obtain the oracle bone character component visual embedding features.

4. A system for deciphering oracle bone characters by contrastive learning and visual ideographic description sequences, the system being used to implement the oracle bone character deciphering method according to any one of claims 1 to 3, characterized in that: The system comprises: The component feature extraction module is used to extract features of the components of the input oracle bone script image to be deciphered using the trained oracle bone script component identifier to obtain visual embedding features of the oracle bone script components; The overall feature extraction module is used to extract overall visual features from the input oracle bone inscription image to be deciphered and the modern Chinese character set using a visual encoder to obtain overall visual features of the oracle bone inscription and the Chinese character set; A feature enhancement module is used to enhance the visual embedding features of the oracle bone inscription components and the overall visual features of the oracle bone inscriptions, and the visual embedding features of the Chinese character components and the overall visual features of the Chinese character using a feature enhancer based on IDS mutual query, so as to obtain enhanced features of the oracle bone inscriptions and enhanced features of the Chinese character; The oracle bone deciphering module is used to perform comparative calculation on the oracle bone inscription enhanced features and the Chinese character enhanced features using contrastive learning to obtain a contrast vector, and obtain an oracle bone inscription deciphering result based on the contrast vector.

5. The oracle bone inscription deciphering system based on contrastive learning and visual ideographic description sequence according to claim 4, characterized in that: In the component feature extraction module, the training process of the Oracle component identifier specifically includes: A visual encoder is used to extract the overall visual features of the oracle bone script image to be deciphered; A visual embedding layer is used to extract visual embedding features from the component set of the oracle bone script image to be deciphered and the visual IDS of the component set; Using a contrastive learner based on cosine distance, the similarity between the feature sequence of the overall visual features of the oracle bone script image to be deciphered and the feature sequence of the visual embedding features is calculated, and a contrast vector is obtained based on the similarity; A loss value of the comparison vector is calculated using a counting loss function, and the oracle bone script component identifier is optimized and trained based on the loss value to obtain a trained oracle bone script component identifier.

6. The oracle bone inscription deciphering system based on contrastive learning and visual ideographic description sequence according to claim 5, characterized in that: The workflow of the component feature extraction module specifically includes: The visualization embedding layer of the oracle bone script component identifier includes a CNN feature extractor, a one-hot encoder, and a fused convolutional layer; Inputting the IDS of the oracle bone character image component to be deciphered into a one-hot encoder to obtain the IDS feature of the oracle bone character component; Inputting the IDS of the oracle bone inscription image component to be deciphered into a CNN feature extractor to obtain visual IDS features of the oracle bone inscription component; The oracle bone character component IDS features and the oracle bone character component visual IDS features are input into the fusion convolution layer for feature fusion to obtain the oracle bone character component visual embedding features.

7. The oracle bone inscription deciphering system based on contrastive learning and visual ideographic description sequence according to claim 6, characterized in that: The workflow of the feature enhancement module specifically includes: Fusing the oracle bone inscription image visual embedding features with the overall visual features of the oracle bone inscription to obtain oracle bone inscription fusion features; Fusing the Chinese character image visual embedding features with the overall visual features of the Chinese character to obtain a Chinese character fusion feature; Use the tag dictionary to perform one-hot encoding on the TopK candidate items of the comparison vector and the IDS of the oracle bone script fusion feature and the IDS of the Chinese character fusion feature to obtain the TopK candidate item embedding features and the label item embedding features; Preprocessing the tag item embedding features and the Chinese character visual overall features after splicing to obtain tag item fusion features; Preprocessing the embedded features of the TopK candidate items and the overall visual features of the oracle bone inscriptions to obtain the fused features of the TopK candidate items; The tag item fusion feature and the TopK candidate item fusion feature are dually input into the IDS mutual query based feature enhancer to output Chinese character enhancement features and oracle bone script enhancement features.

Citation Information

Patent Citations

  • Deep learning-based oracle bone character part detection and recognition method

    CN111539437A

  • Deep learning-based oracle radical splitting and matching method

    CN117333882A