AI-driven ancient coin identification method and system
Through the methods of text feature mapping and model parameter update, the subjectivity and data dependence of traditional ancient currency identification methods are solved, and efficient and accurate ancient currency era identification is achieved.
Patent Information
- Application Number
- CN202510984350.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Traditional ancient coins identification methods rely on manual observation by experts, and have strong subjectivity, low efficiency, and experience inheritance. AI identification methods based on image recognition face the problems of difficulty in obtaining physical image data and high labeling costs.
By obtaining text training examples carrying age marks, the first feature mapper is used to map text features to the image feature alignment domain, combining the ancient coin identification model to be trained, the age is inferred based on text features, and the model parameters are updated through deviations, and finally, the ancient coin identification model that can be used for physical image identification is constructed.
It effectively improves the efficiency and accuracy of ancient coin AI identification, reduces dependence on physical image data, and achieves efficient and accurate ancient coin era identification.
Smart Images

Figure CN120510620A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an AI-driven ancient coin identification method and system. Background Art
[0002] Authentication of ancient coins is a crucial component of archaeological research, cultural relic preservation, and collection. Traditional methods rely on expert observation, such as determining the age of coins based on features like inscriptions, shape, and rust color. These methods are subject to limitations such as high subjectivity, low efficiency, and reliance on experience-based methods. With the development of artificial intelligence (AI) technology, AI authentication methods based on image recognition are gaining popularity, but they face challenges such as the difficulty of acquiring physical image data and the high cost of annotation. Summary of the Invention
[0003] The purpose of the present invention is to provide an AI-driven ancient coin identification method and system.
[0004] In a first aspect, an embodiment of the present invention provides an AI-driven ancient coin identification method, comprising:
[0005] Obtaining a text training instance of a text class, wherein the text training instance carries a year identifier, and the year identifier is used to represent the actual year of the text training instance;
[0006] Performing a feature extraction operation on the text training instance by a first feature mapper to obtain a text feature representation, wherein the first feature mapper is used to match the text-type ancient coin information to an image feature alignment domain, wherein the image feature alignment domain is a domain where the feature representation of the physical image-type ancient coin information resides, and the information content of the text-type ancient coin information is lower than the information content of the physical image-type ancient coin information;
[0007] Based on the text feature representation, the ancient coin identification model to be trained is used for identification to obtain an estimated age;
[0008] According to the deviation between the estimated age and the age mark, the parameter configuration of the ancient coin identification model to be trained is updated to obtain an ancient coin identification model, which is used to identify the age of the ancient coin information of the physical image type based on the ancient coin identification model.
[0009] In one possible implementation, the method further includes:
[0010] Acquire first inferred data of the physical object image;
[0011] Performing feature mapping on the first inferred data by a second feature mapper to obtain an image feature representation, wherein the second feature mapper is used to match the ancient coin information of the physical image class to the image feature alignment domain;
[0012] According to the image feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
[0013] In a possible implementation, the first feature mapper includes a text feature mapper and an image feature mapper, the text feature mapper is used to match the text-type ancient coin information to the image feature alignment domain, and the image feature mapper is used to match the physical image-type ancient coin information to the image feature alignment domain;
[0014] The step of performing a feature extraction operation on the text training instance by a first feature mapper to obtain a text feature representation includes:
[0015] Performing a feature extraction operation on the text training instance by the text feature mapper to obtain a text feature representation;
[0016] The method further comprises:
[0017] Acquire first inferred data of the physical object image;
[0018] performing a feature extraction operation on the first inferred data by the image feature mapper to obtain an image feature representation;
[0019] Based on the image feature representation, the ancient coin identification model is used to perform identification to obtain the age of the first inferred data;
[0020] The method comprises:
[0021] Acquire multiple training instance arrays, wherein the training instance arrays include a first target instance array of the text class and a second target instance array of the physical image class, wherein the first target instance array and the second target instance array of the same training instance array describe the same ancient coin;
[0022] For a target training instance array in the plurality of training instance arrays, performing a feature extraction operation on a first target instance array in the target training instance array by a text feature mapper to be trained to obtain a first undetermined feature representation;
[0023] Performing a feature extraction operation on a second target instance array in the target training instance array by the image feature mapper to be trained to obtain a second undetermined feature representation;
[0024] Using the plurality of training instance arrays as the target training instance arrays, respectively, to obtain a plurality of first undetermined feature representations and a plurality of second undetermined feature representations;
[0025] Based on the optimization goal of improving the feature aggregation of the target undetermined feature representation array and improving the feature discreteness of other undetermined feature representation pairs, the parameter configuration of the text feature mapper to be trained and the parameter configuration of the image feature mapper to be trained are updated to obtain the text feature mapper and the image feature mapper. The first undetermined feature representation and the second undetermined feature representation included in the target undetermined feature representation array are obtained based on the same training instance array, and the first undetermined feature representation and the second undetermined feature representation included in the other undetermined feature representation pairs are not obtained based on the same training instance array.
[0026] In one possible implementation, the method further includes:
[0027] Obtaining at least two age identification items and a first ancient coin text generation paradigm including an age placeholder, wherein the age identification item is used to describe the age of the ancient coin information of the physical image type, and the first ancient coin text generation paradigm is used to drive an NLP model fine-tuned based on the ancient coin field to generate multiple ancient coin text fragments, wherein multiple ancient coin rubbing samples indicated by the multiple ancient coin text fragments constitute an age feature map associated with the age corresponding to the age placeholder;
[0028] Adding the at least two age identification items to the age placeholder area in the first ancient coin description generation paradigm to obtain a first guiding template;
[0029] Generating the text training instance using the NLP model fine-tuned based on the ancient coin field according to the first guidance template, wherein the text training instance is a fragment of ancient coin text related to the at least two chronological identification items;
[0030] The at least two era identification items are determined as era identifications of the text training instance.
[0031] In a possible implementation, the method of performing identification based on the text feature representation using a to-be-trained ancient coin identification model to obtain an estimated age includes:
[0032] According to the text feature representation, a first adapted feature representation is obtained by projecting it through a cross-modal adaptation network, wherein the cross-modal adaptation network is used to project the feature representation from the image feature alignment domain to the model feature alignment domain, wherein the model feature alignment domain is a feature alignment domain that can be identified by the ancient coin identification model;
[0033] According to the first adapted feature representation, the ancient coin identification model to be trained is used for identification to obtain an estimated age;
[0034] The method further comprises:
[0035] Acquire first inferred data of the physical object image;
[0036] performing a feature extraction operation on the first inferred data by the first feature mapper to obtain an image feature representation;
[0037] Projecting the image feature representation through the cross-modal adaptation network to obtain a second adapted feature representation;
[0038] According to the second adaptive feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
[0039] In one possible implementation, the method further includes:
[0040] Obtaining an auxiliary training instance of the text class, wherein the auxiliary training instance carries an auxiliary age annotation;
[0041] According to the auxiliary training instance, performing a feature extraction operation on the auxiliary training instance by the first feature mapper to obtain a cross-modal intermediate feature representation;
[0042] Projecting the cross-modal intermediate feature representation through a cross-modal adaptation network to be trained to obtain an inferred adaptation feature representation;
[0043] According to the inferred adaptation feature representation, identification is performed using the ancient coin identification model to be trained to obtain second inferred data;
[0044] According to the deviation between the second inferred data and the auxiliary age annotation, the parameter configuration of the ancient coin identification model to be trained is fixed, and the parameter configuration of the cross-modal adaptation network to be trained is updated to obtain the cross-modal adaptation network.
[0045] In one possible implementation, the method further includes:
[0046] Obtaining a second ancient coin description generation paradigm and a cross-modal adaptation command, wherein the second ancient coin description generation paradigm includes a feature representation placeholder area and a cross-modal adaptation command placeholder area, and the cross-modal adaptation command is used to drive the ancient coin identification model to be trained to generate the second inference data;
[0047] Adding the inferred adaptation feature representation to the feature representation placeholder in the second ancient coin text generation paradigm, and adding the cross-modal adaptation command to the cross-modal adaptation command placeholder in the second ancient coin text generation paradigm, to obtain a second guidance template;
[0048] The method of obtaining second inferred data by performing identification based on the inferred adaptation feature representation using the ancient coin identification model to be trained includes:
[0049] According to the second guiding template, identification is performed using the ancient coin identification model to be trained to obtain the second inferred data.
[0050] In a possible implementation, updating the parameter configuration of the ancient coin identification model to be trained based on the deviation between the estimated age and the age identifier to obtain the ancient coin identification model includes:
[0051] According to the deviation between the inferred age and the age identifier, the parameter configuration of the ancient coin identification model to be trained is updated to obtain the ancient coin identification model, and according to the deviation between the inferred age and the age identifier, the parameter configuration of the cross-modal adaptation network is updated to obtain the cross-modal adaptation network that has been iteratively tuned.
[0052] In one possible implementation, the method further includes:
[0053] Obtaining the core dating basis of the text training instance;
[0054] Performing a feature extraction operation on the core generation basis to obtain a core feature representation;
[0055] The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes:
[0056] According to the text feature representation and the core feature representation, the ancient coin identification model to be trained is used for identification to obtain the estimated age;
[0057] The method further comprises:
[0058] Obtaining a third ancient coin text generation paradigm and a guiding command, wherein the third ancient coin text generation paradigm includes a feature representation placeholder area and a guiding command placeholder area, and the guiding command is used to drive the ancient coin identification model to be trained to generate the estimated age;
[0059] Adding the text feature representation to the feature representation placeholder area of the third ancient coin text generation paradigm, and adding the guide command to the guide command placeholder area of the third ancient coin text generation paradigm to obtain a third guide template;
[0060] The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes:
[0061] According to the third guiding template, the ancient coin identification model to be trained is used to perform identification to obtain an estimated age;
[0062] The third ancient coin description generation paradigm also includes a core dating basis placeholder area, and the method further includes:
[0063] Obtaining the core dating basis of the text training instance;
[0064] Performing a feature extraction operation on the core generation basis to obtain a core feature representation;
[0065] The step of adding the text feature representation to the feature representation placeholder area in the third ancient coin text generation paradigm, and adding the guide command to the guide command placeholder area in the third ancient coin text generation paradigm to obtain a third guide template includes:
[0066] Add the character feature representation to the third ancient coin character
[0067] The feature representation placeholder area in the generation paradigm is added, the guide command is added to the guide command placeholder area in the third ancient coin text generation paradigm, and the core feature representation is added to the core dating basis placeholder area in the third ancient coin text generation paradigm to obtain a third guide template.
[0068] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is configured to execute the method described in the first aspect.
[0069] Compared with the existing technology, the beneficial effects provided by the present invention include: using the AI-driven ancient coin identification method and system disclosed in the present invention, by obtaining text training examples carrying age identification; mapping the text features to the image feature alignment domain where the physical image features are located through a first feature mapper to obtain a text feature representation; using the ancient coin identification model to be trained to infer the age based on the text feature representation, updating the model parameters through the deviation from the true age, and finally obtaining an ancient coin identification model that can be used for physical image age identification. This method effectively improves the efficiency and accuracy of ancient coin AI identification through cross-modal feature alignment and text data drive, and reduces dependence on physical image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.
[0071] Figure 1 A schematic diagram of the steps of the AI-driven ancient coin identification method provided in an embodiment of the present invention;
[0072] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0074] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0075] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of the AI-driven ancient coin identification method provided in an embodiment of the present disclosure. The AI-driven ancient coin identification method is introduced in detail below.
[0076] Step S201: obtaining a text training instance of a text class, wherein the text training instance carries a year identifier, and the year identifier is used to represent the actual year of the text training instance;
[0077] Step S202: performing a feature extraction operation on the text training instance by a first feature mapper to obtain a text feature representation, wherein the first feature mapper is used to match the text-type ancient coin information to an image feature alignment domain, wherein the image feature alignment domain is the domain where the feature representation of the physical image-type ancient coin information resides, and the information content of the text-type ancient coin information is lower than the information content of the physical image-type ancient coin information;
[0078] Step S203, performing identification based on the text feature representation using the ancient coin identification model to be trained to obtain an estimated age;
[0079] Step S204, based on the deviation between the estimated age and the age identifier, update the parameter configuration of the ancient coin identification model to be trained to obtain an ancient coin identification model, so as to identify the age of the ancient coin information of the physical image type based on the ancient coin identification model.
[0080] In an embodiment of the present invention, the server, as the execution entity, first needs to construct text training instances for model training. Considering the heterogeneity of text data (such as ancient book records, archaeological reports, rubbing descriptions, etc.) and physical image data (such as high-definition photos and scans of ancient coins) in the field of ancient coin identification, and the fact that the amount of information in text-based ancient coin information (such as "Han Wuzhu coin, diameter 2.5 cm, square hole, the coin text 'Wuzhu' has thicker strokes") is usually lower than that of the physical image (including visual details such as texture, rust color, and wear marks), the server needs to obtain text training instances through a combination of active generation and passive collection.
[0081] Specifically, the server is pre-configured with a natural language processing (NLP) model fine-tuned for the ancient coin domain and trained on specialized ancient coin data. The server first obtains at least two date identifiers (e.g., "Western Han Dynasty," "Eastern Han Dynasty," "Tang Dynasty," or "Song Dynasty") and invokes a pre-designed "first ancient coin description generation paradigm," a structured text generation template consisting of a "date placeholder" and a coin description framework (e.g., "The typical currency of the [date placeholder] period is [coin name], characterized by: [coin diameter], [hole size], [coin inscription features], [casting process]").
[0082] The server fills in the "era placeholder area" with era identifiers such as "Western Han Dynasty" and "Tang Dynasty" to generate a "first guidance template" (for example, "The typical currency of the Western Han Dynasty was the Wuzhu coin, which has the following characteristics: a diameter of approximately 2.3-2.5 cm, a square hole with a side length of approximately 0.9 cm, the two characters "Wuzhu" on the coin have relatively thick strokes, and the casting process is mold casting"; "The typical currency of the Tang Dynasty was the Kaiyuan Tongbao coin, which has the following characteristics: a diameter of approximately 2.4-2.5 cm, a square hole with a side length of approximately 0.8 cm, and the coin character "Kaiyuan Tongbao" is in the Ouyang Xun style with dignified strokes"). Subsequently, the server uses a fine-tuned NLP model to generate a large number of ancient coin text fragments based on the guidance template. These fragments are "text training instances", and the era identifier of each fragment directly corresponds to the filled-in era items such as "Western Han Dynasty" and "Tang Dynasty".
[0083] The server also passively collects dated text data from public archaeological reports and museum collection description databases to supplement text training examples. (For example, an archaeological report clearly states, "Three round coins with square holes, inscribed 'half liang', lightly corroded, were unearthed. Carbon-14 dating indicates they are from the late Warring States period, Qin Dynasty.") Ultimately, the server integrates this actively generated and passively collected text data to form a text training example set containing tens of thousands of records, each with a corresponding date identifier (e.g., "late Warring States period," "Western Han Dynasty," etc.).
[0084] To map textual coin information and physical imaged coin information to the same "image feature alignment domain" (i.e., the feature space where physical image features reside), the server needs to train a "first feature mapper." This first feature mapper is designed to consist of a "text feature mapper" and an "image feature mapper." The former is responsible for mapping textual features to the image feature alignment domain, while the latter is responsible for mapping physical image features to the same domain, thus achieving cross-modal feature alignment.
[0085] The server first constructs a "training instance array": for a single ancient coin (e.g., a Western Han Dynasty Wuzhu coin), it collects its corresponding text description (e.g., "Western Han Dynasty Wuzhu, diameter 2.4 cm, square hole, clear inscription 'Wuzhu', wide hole") as the "first target instance array," and collects high-definition images of the coin (including multi-angle photos of the front, back, and side) as the "second target instance array." These two together form a training instance array describing the same ancient coin. The server needs to collect thousands of such training instance arrays, covering ancient coins from different eras and types (e.g., knife coins, cloth coins, and round coins with square holes).
[0086] Next, the server initializes the text feature mapper (e.g., a BERT-based text encoder) and image feature mapper (e.g., a ResNet-based visual encoder) to be trained and optimizes their parameters through comparative learning. The specific process is as follows: For each training instance array (denoted as the target training instance array), the server uses the text feature mapper to extract features from the text description (the first target instance array) to obtain the "first undetermined feature representation" (a vector of 256 dimensions). Simultaneously, the image feature mapper extracts features from the physical image (the second target instance array) to obtain the "second undetermined feature representation" (also a vector of 256 dimensions).
[0087] The server's optimization goal is to ensure that the first and second undetermined feature representations corresponding to the same training instance array (i.e., the target undetermined feature representation array) are as similar as possible in the image feature alignment domain (increased aggregation), while the feature pairs corresponding to different training instance arrays (i.e., other undetermined feature representation pairs) are as dissimilar as possible (increased discreteness). This is achieved through a contrast loss function, such as using the InfoNCE loss, calculating the cosine similarity of the same set of features as the positive sample score, and the cosine similarity of different sets of features as the negative sample score, and optimizing the mapper parameters by maximizing the difference between the positive and negative sample scores. After tens of thousands of iterative training cycles, the parameters of the text feature mapper and the image feature mapper converged, enabling the unified mapping of text and image features to the image feature alignment domain.
[0088] Based on this, when the server needs to process a text training instance, it calls the trained text feature mapper to perform feature extraction on it. For example, for a text training instance such as "Tang Kaiyuan Tongbao, the coin inscription is in Ouyang Xun's style, with strong and straight strokes and a moon pattern on the back," the text feature mapper extracts its key features (such as "coin inscription style" and "back pattern features") and maps them to the image feature alignment domain, resulting in a 256-dimensional "text feature representation" vector. This vector is in the same feature space as the feature vector of the physical image processed by the image feature mapper.
[0089] After obtaining the text feature representation, the server needs to predict its age through the ancient coin identification model to be trained (such as a Transformer-based classification model) and update the model parameters based on the deviation from the actual age identification.
[0090] Taking into account that the features of the image feature alignment domain may differ from the input requirements of the ancient coin identification model (i.e., the "model feature alignment domain"), the server introduces a "cross-modal adaptation network" (for example, a projection network consisting of two fully connected layers) to project the features of the image feature alignment domain into the model feature alignment domain. Specifically, the server first inputs the text feature representation into the cross-modal adaptation network to obtain a "first adapted feature representation" (for example, a vector with a dimension adjusted to 512). This adapted feature is then input into the ancient coin identification model to be trained. The model outputs the probability distribution of each age category (such as "Warring States" 0.1, "Han" 0.7, "Tang" 0.2, etc.), and the category with the highest probability is the "estimated age."
[0091] To optimize the cross-modal adaptation network and the ancient coin identification model, the server employs a multi-stage training strategy: 1. Pre-training the cross-modal adaptation network: The server obtains additional "auxiliary training examples" (e.g., text data not included in the main training, but with auxiliary dating annotations) and extracts their text feature representations (in the image feature alignment domain) through a text feature mapper. This is then fed into the cross-modal adaptation network to be trained to obtain the "inferred adapted feature representation," which is then fed into the ancient coin identification model to be trained to obtain the "second inferred data" (i.e., the inferred dating). The server calculates the deviation (e.g., cross-entropy loss) between the second inferred data and the auxiliary dating annotations, fixes the parameters of the ancient coin identification model, and only updates the parameters of the cross-modal adaptation network to effectively project the features in the image feature alignment domain into a space that the model can process. 2. Jointly training the ancient coin identification model and the cross-modal adaptation network: During the main training phase, the server feeds the text feature representations of the text training examples into the cross-modal adaptation network to obtain the first adapted feature representation, which is then fed into the ancient coin identification model to obtain the inferred dating. The server calculates the deviation between the inferred age and the age of the training text (e.g., cross-entropy loss), and simultaneously backpropagates the parameters of the ancient coin authentication model and the cross-modal adaptation network. For example, if the model infers the Han Dynasty for a training text with the age of the Tang Dynasty, the loss will be large, triggering parameter adjustments. If the inference is correct, the loss will be small, and the parameter adjustments will be reduced.
[0092] The server also considers the guiding role of core dating criteria for ancient coin identification (such as the style of the coin inscription, its form, and the casting process) in guiding the model. For example, for the text training example "Northern Song Chongning Tongbao, inscription in Song Huizong's Slender Gold Script, iron-stroked and silver-hooked, coin diameter 3.5 cm," the core dating criteria are "Slender Gold Script" and "coin diameter." The server extracts features (such as keyword extraction or entity recognition) to obtain "core feature representations" (e.g., feature vectors for "Slender Gold Script" and "3.5 cm") and inputs these into the ancient coin identification model along with the text feature representations. To enhance the model's focus on core criteria, the server uses the "Third Ancient Coin Description Generation Paradigm," a template consisting of a "feature representation placeholder," a "core dating criteria placeholder," and a "guiding command placeholder" (e.g., "Based on feature [feature representation placeholder] and core dating criteria [core dating criteria placeholder], the age of this ancient coin is inferred to be: [guiding command placeholder]"). The server fills the text feature representation, core feature representation and "age inference" guidance command into the template to form a "third guidance template". The ancient coin identification model performs inference based on this template, thereby focusing more on key features.
[0093] Once the ancient coin identification model is trained, the server can be used to identify ancient coins using physical images. The specific process is as follows: 1. The server obtains the physical image to be identified (e.g., a user-uploaded photo of the front of an ancient coin) as the "first inferred data"; 2. The trained image feature mapper is used to extract features from the image, generating an "image feature representation" (in the image feature alignment domain); 3. The image feature representation is projected into the model feature alignment domain using a cross-modal adaptation network, generating a "second adapted feature representation"; 4. The second adapted feature representation is input into the ancient coin identification model, which outputs a probability distribution for each age. The age with the highest probability is the inferred age of the coin.
[0094] For example, a user uploads an image of an ancient coin. The server uses an image feature mapper to extract its features (such as the font shape of the inscription "Daguan Tongbao," the coin border width, and the rust color distribution). After mapping these features to the image feature alignment domain, the cross-modal adaptation network adjusts the dimensionality and feature representation before inputting them into the ancient coin authentication model. Based on the knowledge learned during the training phase, such as "Daguan Tongbao was a currency from the Huizong period of the Northern Song Dynasty and the inscription is in the Slender Gold script," the model outputs "Northern Song Dynasty" as the estimated date, completing the authentication.
[0095] In summary, this method effectively addresses the shortage of physical image data in ancient coin authentication through cross-modal alignment of text and image features, generative data augmentation, and guidance based on core dating criteria. By leveraging this rich text data to train models, the method ultimately enables accurate dating of physical images. The server, acting as the execution agent, builds a complete AI-driven ancient coin authentication system through steps such as data acquisition, feature mapper training, and model optimization. This system provides efficient and accurate technical support for ancient coin research and collection.
[0096] In the embodiments of the present invention, the following implementation modes are also provided.
[0097] Acquire first inferred data of the physical object image;
[0098] Performing feature mapping on the first inferred data by a second feature mapper to obtain an image feature representation, wherein the second feature mapper is used to match the ancient coin information of the physical image class to the image feature alignment domain;
[0099] According to the image feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
[0100] In an embodiment of the present invention, for example, when a user needs to identify the age of an ancient coin, they upload the coin's physical image data through a client (e.g., an ancient coin identification app or a museum digitization system). The server, acting as the data receiver and processor, first obtains these images as "first inferred data." For example, a user takes a photo of the obverse of an ancient coin with a resolution of 4032×3024 pixels, clearly showing the inscription "Chongning Tongbao," the outline of the coin's outer edge, and details such as the rust color distribution. The server receives this image file via an HTTP interface, stores it in a temporary data buffer, and performs preprocessing on it, including cropping (removing irrelevant background to retain only the coin's main body), scaling (adjusting to a uniform 224×224 pixels to accommodate model input requirements), and normalizing (scaling pixel values to the range [0,1]). This ultimately yields standardized physical image data, the "first inferred data." The server then invokes a pre-trained "second feature mapper" (i.e., an image feature mapper) to extract and map features from the first inferred data. The image feature mapper is a visual encoder trained through contrastive learning (for example, based on the ResNet-50 architecture). It has the ability to map the features of physical images to the "image feature alignment domain", which is a unified feature space of text features and image features, ensuring that the features of text descriptions and physical images are comparable in the same dimension. During specific execution, the server inputs the pre-processed ancient coin image into the image feature mapper: first, the image is passed through the convolutional layer of ResNet-50 to extract local features (such as the stroke edges of the coin text and the texture gradient of the rust color), and then the local features are aggregated into a global feature vector (with a dimension of 256) through the global average pooling layer; then, the projection head inside the mapper (consisting of two fully connected layers) performs a nonlinear transformation on the vector, and finally outputs an "image feature representation" with a dimension of 256. This vector is in the image feature alignment domain and is in the same space as the feature vector of the text training instance after being processed by the text feature mapper. For example, for an image of a Chongning Tongbao coin, the image feature mapper extracts key features such as "the coin's inscription is in a thin gold script, with strokes as slender as iron wire," "the coin's diameter is approximately 3.5 cm," and "there's no border on the back." It encodes these features into a 256-dimensional vector. This vector is highly similar to the vector of the text description of a "Northern Song Chongning Tongbao coin" processed by the text feature mapper in the image feature alignment domain. After obtaining the image feature representation, the server uses the trained ancient coin authentication model to infer the coin's age.Given the dimensionality difference between the features in the image feature alignment domain (256 dimensions) and the input requirements of the ancient coin authentication model (e.g., the 512-dimensional model feature alignment domain), the server first invokes a "cross-modal adaptation network" (pre-trained and fine-tuned using auxiliary training examples) to project the image feature representation. The cross-modal adaptation network consists of two fully connected layers. The first layer maps the 256-dimensional vector to 512 dimensions. The second layer enhances the nonlinear representation using the ReLU activation function, ultimately outputting a 512-dimensional "second adapted feature representation" that meets the input requirements of the ancient coin authentication model. The server then inputs the second adapted feature representation into the ancient coin authentication model (a Transformer-based classification model). The encoder layer within the model performs contextual modeling on this feature, capturing correlations between features across dimensions (e.g., the strong correlation between "Slender Gold Coin Inscription" and the period of Emperor Huizong of the Northern Song Dynasty, and the typical size of a Chongning Tongbao coin, "3.5 cm in diameter"). Finally, the classification head (fully connected layer with a softmax activation function) outputs a probability distribution for each era category (e.g., 0.92 for "Northern Song Dynasty," 0.05 for "Southern Song Dynasty," 0.03 for "Tang Dynasty," etc.). The highest probability, "Northern Song Dynasty," is the estimated era for the coin. For example, given a user-uploaded image of a Chongning Tongbao coin, after the server processes it through the aforementioned process, the coin identification model outputs a 92% probability of "Northern Song Dynasty," based on the knowledge learned during training, such as "Slender Gold Coin Inscription is a typical feature of the period of Emperor Huizong of the Northern Song Dynasty" and "Chongning Tongbao coin diameter is typically 3.3-3.7 cm," ultimately determining the coin's era as the Northern Song Dynasty. In summary, the server achieves efficient dating of ancient coin information from physical images by receiving physical images, feature map alignment, cross-modal adaptation, and model inference, providing users with accurate technical support.
[0101] In an embodiment of the present invention, the first feature mapper includes a text feature mapper and an image feature mapper, the text feature mapper is used to match the text-type ancient coin information to the image feature alignment domain, and the image feature mapper is used to match the physical image-type ancient coin information to the image feature alignment domain;
[0102] The feature extraction operation performed on the text training instance by the first feature mapper to obtain text feature representation can be implemented through the following example.
[0103] Performing a feature extraction operation on the text training instance by the text feature mapper to obtain a text feature representation;
[0104] The present invention also provides the following implementation modes.
[0105] Acquire first inferred data of the physical object image;
[0106] performing a feature extraction operation on the first inferred data by the image feature mapper to obtain an image feature representation;
[0107] According to the image feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
[0108] In an embodiment of the present invention, the server first processes a text training instance for the text category. For example, the server retrieves a text training instance from an archaeological database: "Western Han Dynasty Wuzhu coin, diameter approximately 2.4 cm, square hole side length 0.9 cm, the inscription 'Wuzhu' in bold and clear strokes, wide hole." Its date is indicated as "Western Han Dynasty." The server then invokes a trained text feature mapper to extract features from this instance. The text feature mapper is a text encoder based on the BERT architecture, trained through contrastive learning to map text features to an "image feature alignment domain" that aligns with physical image features. The specific processing flow is as follows: the server first performs word segmentation on the text training instance (e.g., "Western Han Dynasty," "Wuzhu coin," "diameter," "2.4 cm," etc.), generating a sequence of word vectors that are input to BERT. BERT's 12-layer Transformer encoder then contextualizes the word vectors, capturing the semantic associations of key information such as "coin inscription features" and "dimensional parameters." Finally, a pooling layer (e.g., CLS vector extraction) is used to generate a 256-dimensional "text feature representation" vector. This vector is not a simple semantic representation of text, but rather an alignment feature in the same space as the physical image features. For example, the cosine similarity between this vector and the feature vector of a physical image of a "Western Han Dynasty Wu Zhu Coin" processed by the image feature mapper is calculated to be above 0.85, indicating a high correlation between the two in the image feature alignment domain. When a user uploads a physical image of an ancient coin (such as a high-definition photo of a "Western Han Dynasty Wu Zhu Coin," including the obverse text, reverse outline, and rust details) as "first guess data," the server invokes the image feature mapper to extract features. The image feature mapper is a visual encoder based on ResNet-50, also trained through contrastive learning, and is able to map image features to the image feature alignment domain. The specific processing flow is as follows: the server first preprocesses the image (cropping to the coin body, scaling to 224×224 pixels, and normalizing pixel values). The preprocessed image is then fed into a ResNet-50 network, where five convolutional blocks extract local features (such as the edges of the inscription, the outline of the square hole, and the rust texture gradient). A global average pooling layer then aggregates these local features into a 256-dimensional global feature vector. Finally, a nonlinear transformation is performed within the mapper's projection head (two fully connected layers) to output a 256-dimensional "image feature representation" vector. For example, for an image of a "Western Han Dynasty Wuzhu coin," this vector encodes key visual features such as "coin diameter 2.4 cm," "coin inscription strokes thick," and "hole width." These features highly overlap with the text feature representations of the corresponding text training examples in the image feature alignment domain. After obtaining the image feature representation, the server then applies it to a trained ancient coin authentication model for dating.Due to the dimensionality difference between the features in the image feature alignment domain (256 dimensions) and the model input requirement (e.g., the 512-dimensional model feature alignment domain), the server first uses a cross-modal adaptation network to project the image feature representation. The cross-modal adaptation network maps the 256-dimensional vector into a 512-dimensional "second-adapted feature representation" and enhances the feature representation through Reinforced Unit (ReLU) activation. This second-adapted feature representation is then fed into the ancient coin authentication model (a Transformer-based classification model). The model's encoder layer contextualizes the features, associating "thick strokes" with typical characteristics of "Western Han Dynasty Wuzhu coins" and "hole width" with the Western Han Dynasty casting process. Finally, the classification head outputs a probability distribution for each era (e.g., 0.91 for "Western Han Dynasty," 0.06 for "Eastern Han Dynasty," and 0.03 for "Warring States Period"). The highest probability, "Western Han Dynasty," is the estimated era for the coin. In summary, the server unifies heterogeneous text and image data into the same feature space through text feature mappers and image feature mappers, and combines cross-modal adaptation networks with ancient coin identification models to achieve full process coverage from text training to physical image identification, providing efficient and accurate technical support for ancient coin age identification.
[0109] In the embodiments of the present invention, the following implementation modes are also provided.
[0110] Acquire multiple training instance arrays, wherein the training instance arrays include a first target instance array of the text class and a second target instance array of the physical image class, wherein the first target instance array and the second target instance array of the same training instance array describe the same ancient coin;
[0111] For a target training instance array in the plurality of training instance arrays, performing a feature extraction operation on a first target instance array in the target training instance array by a text feature mapper to be trained to obtain a first undetermined feature representation;
[0112] Performing a feature extraction operation on a second target instance array in the target training instance array by the image feature mapper to be trained to obtain a second undetermined feature representation;
[0113] Using the plurality of training instance arrays as the target training instance arrays, respectively, to obtain a plurality of first undetermined feature representations and a plurality of second undetermined feature representations;
[0114] Based on the optimization goal of improving the feature aggregation of the target undetermined feature representation array and improving the feature discreteness of other undetermined feature representation pairs, the parameter configuration of the text feature mapper to be trained and the parameter configuration of the image feature mapper to be trained are updated to obtain the text feature mapper and the image feature mapper. The first undetermined feature representation and the second undetermined feature representation included in the target undetermined feature representation array are obtained based on the same training instance array, and the first undetermined feature representation and the second undetermined feature representation included in the other undetermined feature representation pairs are not obtained based on the same training instance array.
[0115] In an embodiment of the present invention, the server first constructs a "training instance array" for training the feature mapper. For example, the server collects text descriptions and physical image data of the same ancient coin from museum digitization systems, archaeological report databases, and user-shared ancient coin datasets (user authorization required). For example, the corresponding "first target instance array" (text category) for the "Northern Song Chongning Tongbao" coin might read: "Chongning Tongbao, inscription in Song Huizong's Slender Gold Script, iron-stroked and silver-hooked, coin diameter approximately 3.5 cm, no inner rim on the back." The "second target instance array" (physical image category) consists of a high-definition photo of the coin's front (including inscription details), a photo of the back (showing the absence of an inner rim), and a side scan (for measuring the coin's diameter). The server must collect thousands of such instance arrays, covering ancient coins of different eras and types (e.g., Warring States knife-shaped coins, Tang Dynasty Kaiyuan Tongbao, Ming Dynasty Hongwu Tongbao, etc.), with each instance array strictly corresponding to heterogeneous data from the same ancient coin. The server initializes the text feature mapper (BERT-based text encoder, with randomly initialized parameters) and image feature mapper (ResNet-50-based visual encoder, with randomly initialized parameters) to be trained. For each training instance array (referred to as the "target training instance array"), the server performs the following operations: Processing text data: The "first target instance array" (e.g., "Chongning Tongbao, the coin inscription is in the thin gold style...") is fed into the text feature mapper to be trained. The mapper first segments the text into words (e.g., "Chongning Tongbao," "Slender Gold Style," "3.5 centimeters"), generating a sequence of word vectors. It then extracts contextual features using BERT's 12-layer Transformer encoder, ultimately outputting a 256-dimensional "first undetermined feature representation" (e.g., vector V1). Processing image data: The "second target instance array" (e.g., a front-facing image of a Chongning Tongbao coin) is fed into the image feature mapper to be trained. The mapper first preprocesses the image (cropping to the coin body and scaling to 224×224 pixels). It then extracts local features (such as the edges of the inscription strokes and the coin diameter outline) using five convolutional blocks of ResNet-50. After global average pooling and projection, it outputs a 256-dimensional "second undetermined feature representation" (e.g., vector V2). The server iterates through all training instance arrays, generating thousands of (V1, V2) pairs of undetermined feature representations, each corresponding to the text and image features of the same coin. The server's optimization goal is to ensure that V1 and V2 (the target undetermined feature representation array) within the same training instance array are as similar as possible in feature space (improving convergence), and that V1 and Vj (j ≠ the current array, other undetermined feature representation pairs) between different training instance arrays are as dissimilar as possible (improving divergence). This is achieved using a contrastive loss function (such as the InfoNCE loss). The server calculates the gradient of the loss with respect to the parameters of the text and image feature mappers through backpropagation and updates the parameters using the Adam optimizer.For example, if the current target array's V1 has low similarity to V2 but high similarity to a negative sample Vj, the loss value is high, triggering parameter adjustments. This causes the text mapper to focus more on key text features such as "Slender Gold" and "3.5 cm," while the image mapper focuses more on visual features such as the coin's font shape and coin diameter. This improves the similarity between V1 and V2 and reduces the similarity with the negative sample. After tens of thousands of training iterations (e.g., 100,000 steps), when the loss value converges to a stable range (e.g., below 0.1), the parameters of the text feature mapper and image feature mapper no longer change significantly. The resulting mapper is now able to map the text and image features of the same ancient coin to similar locations in the image feature alignment domain (e.g., cosine similarity ≥ 0.8), while the features of different ancient coins are spatially dispersed (cosine similarity ≤ 0.4), completing the mapper training. In summary, by constructing heterogeneous data pairs, extracting features, and optimizing for contrastive learning, the server achieves joint training of the text and image feature mappers, laying the foundation for feature alignment in subsequent cross-modal learning of ancient coin authentication models.
[0116] In the embodiments of the present invention, the following implementation modes are also provided.
[0117] Obtaining at least two age identification items and a first ancient coin text generation paradigm including an age placeholder, wherein the age identification item is used to describe the age of the ancient coin information of the physical image type, and the first ancient coin text generation paradigm is used to drive an NLP model fine-tuned based on the ancient coin field to generate multiple ancient coin text fragments, wherein multiple ancient coin rubbing samples indicated by the multiple ancient coin text fragments constitute an age feature map associated with the age corresponding to the age placeholder;
[0118] Adding the at least two age identification items to the age placeholder area in the first ancient coin description generation paradigm to obtain a first guiding template;
[0119] Generating the text training instance using the NLP model fine-tuned based on the ancient coin field according to the first guidance template, wherein the text training instance is a fragment of ancient coin text related to the at least two chronological identifiers;
[0120] The at least two era identification items are determined as era identifications of the text training instance.
[0121] In an embodiment of the present invention, illustratively, the server first needs to determine the "era identification items" used to generate text training instances. These era items need to cover the key periods commonly seen in ancient coin research, such as typical dynasties such as the "Western Han Dynasty", "Tang Dynasty", "Northern Song Dynasty" and "Ming Dynasty" extracted from authoritative historical documents and archaeological databases. Assume that the server currently selects "Western Han Dynasty" and "Tang Dynasty" as at least two era identification items to generate ancient coin text fragments of the corresponding eras. At the same time, the server calls the pre-designed "first ancient coin text generation paradigm", which is a structured text generation template that contains a "era placeholder" and a fixed framework for describing the characteristics of ancient coins. For example, the structure of the generation paradigm is: "The typical currency of the [era placeholder] period is [coin name], and its core features include: [coin text style], [coin diameter range], [casting process], [special mark] (if none, mark 'none')". The design of this paradigm is based on the core dating criteria for ancient coin identification (coin inscriptions, shapes, craftsmanship, etc.), and aims to drive the NLP model to generate text fragments containing key features. These fragments will subsequently constitute a "chronological feature map" associated with the era (for example, "Western Han Dynasty" corresponds to "Wuzhu coins, coarse brush inscriptions, and casting technology", and "Tang Dynasty" corresponds to "Kaiyuan Tongbao, Ouyang Xun's calligraphy, and sand casting").
[0122] The server fills the selected "Western Han Dynasty" and "Tang Dynasty" into the "era placeholder" of the generative model, generating the "first guiding template." For example, the template for the "Western Han Dynasty" reads: "The typical currency of the Western Han Dynasty is [coin name], and its core features include: [coin inscription style], [coin diameter range], [casting process], and [special markings] (if none, mark 'none')"; the template for the "Tang Dynasty" reads: "The typical currency of the Tang Dynasty is [coin name], and its core features include: [coin inscription style], [coin diameter range], [casting process], and [special markings] (if none, mark 'none')." The server invokes an NLP model fine-tuned in the field of ancient coins (for example, a generative model based on RoBERTa trained on professional corpus and familiar with ancient coin terminology such as "Wu Zhu," "Kai Yuan Tong Bao," and "Slender Gold Script") to generate specific ancient coin text snippets based on the first guiding template. Taking the "Western Han Dynasty" template as an example, the model generates content based on slots such as "Coin Name" and "Coin Inscription Style" in the template, combined with the ancient coin knowledge base: "The typical currency of the Western Han Dynasty was the Wuzhu coin. Its core characteristics include: the inscription 'Wuzhu' has thick and deep strokes, the coin diameter ranges from 2.3-2.5 cm, it was cast using a clay mold, and there are no special markings." For the "Tang Dynasty" template, the model generates: "The typical currency of the Tang Dynasty was the Kaiyuan Tongbao coin. Its core characteristics include: the inscription 'Kaiyuan Tongbao' is in Ouyang Xun's style, with dignified and symmetrical strokes, the coin diameter ranges from 2.4-2.5 cm, it was cast using the sand casting method, and special markings often include a moon pattern on the back." The server batch-calls the model to generate thousands of similar text snippets (e.g., "Western Han Wuzhu coin, shallow inscription, 2.4 cm diameter, cast in a mold," "Tang Kaiyuan Tongbao coin, star pattern on the back, clear inscription"), etc. These snippets serve as "text training instances." Each instance focuses on a typical ancient coin from a specific era, covering the core characteristics of that era (inscription, dimensions, craftsmanship, and markings), thus forming a "period feature map." For example, 90% of the text snippets in the "Western Han" map mention "Wuzhu coin," "coarse brush inscription," and "cast in a mold," while 85% of the snippets in the "Tang" map mention "Kaiyuan Tongbao coin," "Ouyang Xun calligraphy," and "sand casting." The server uses the "Western Han" and "Tang" entries when generating the text snippets as the "period identifiers" for the corresponding text training instances. For example, the "Western Han Wuzhu coin..." instance has the period identifier "Western Han," while the "Tang Kaiyuan Tongbao coin..." instance has the period identifier "Tang." To ensure the quality of the generated data, the server also filters out abnormal fragments through manual review or rule-based verification (for example, checking whether the "coin name" matches the era and whether the "casting process" is consistent with the historical period). For example, "Wu Zhu coins from the Tang Dynasty" would be marked as invalid and discarded because Wu Zhu coins were primarily popular in the Han Dynasty and conflict with the Tang Dynasty's typical currency, Kaiyuan Tongbao. Ultimately, the server obtained tens of thousands of accurately labeled and clearly defined text training examples, providing rich text data support for the subsequent training of ancient coin identification models.
[0123] In an embodiment of the present invention, if the first ancient coin text generation paradigm is also used to drive the NLP model fine-tuned based on the ancient coin field to generate auxiliary target values related to the age feature map, then the determination of the at least two age identification items as the age identification of the text training instance can be implemented through the following example.
[0124] The at least two era identification items and the auxiliary target value are determined as the era identification of the text training instance.
[0125] In this embodiment of the present invention, the server, based on the existing "First Ancient Coin Description Generation Paradigm," is further designed to drive the NLP model to generate "auxiliary target values" related to the "chronological feature map." These auxiliary target values are fine-grained supplements to the chronological identifier. For example, if the chronological identifier is "Western Han Dynasty," the auxiliary target values could be "early," "middle," or "late" (subdivided by time) or "official casting" or "private casting" (subdivided by casting nature); if the chronological identifier is "Tang Dynasty," the auxiliary target values could be "early Tang Dynasty," "high Tang Dynasty," or "middle Tang Dynasty" (subdivided by historical period) or "cast in Yangzhou" or "cast in Luoyang" (subdivided by casting location). These auxiliary target values are strongly correlated with the chronological feature map. For example, Wuzhu coins from the "early Western Han Dynasty" typically have shallow inscriptions and crude casting, while Wuzhu coins from the "middle Western Han Dynasty" (the reign of Emperor Wu of Han) have deep inscriptions and regular craftsmanship, forming differentiated feature labels. The structure of the generative paradigm is therefore expanded to: "[era placeholder] - [auxiliary target area] The typical currency of the period is [coin name], and its core features include: [coin inscription style], [casting process], [special mark] (if none, mark 'none')". Among them, the "auxiliary target area" is a placeholder that drives the NLP model to generate auxiliary target values. The server selects at least two era identification items (such as "Western Han Dynasty" and "Tang Dynasty"), and determines the corresponding auxiliary target values based on the ancient coin knowledge base (such as "Western Han Dynasty" corresponds to "early" and "middle", and "Tang Dynasty" corresponds to "early Tang Dynasty" and "prosperous Tang Dynasty"). Subsequently, the era identification items and auxiliary target values are filled into the "era placeholder" and "auxiliary target area" of the generative paradigm to obtain the expanded "first guidance template". For example, the template for the "Western Han Dynasty - Early Period" reads: "The typical currency of the Western Han Dynasty - Early Period is [coin name], and its core features include: [coin inscription style], [casting technique], and [special markings] (marked 'none' if none exist)"; the template for the "Tang Dynasty - High Tang Dynasty" reads: "The typical currency of the Tang Dynasty - High Tang Dynasty" reads: "The typical currency of the Tang Dynasty - High Tang Dynasty" is [coin name], and its core features include: [coin inscription style], [casting technique], and [special markings] (marked 'none' if none exist)." The server invokes an NLP model fine-tuned in the ancient coin field (familiar with specific knowledge such as "the inscription on the Wuzhu coin in the early Western Han Dynasty is shallow" and "the Kaiyuan Tongbao coin in the High Tang Dynasty has multiple moon patterns on the back") to generate specific ancient coin inscription snippets based on these templates. For example, the model generated from the "Western Han Dynasty - Early Period" template: "The typical currency of the Western Han Dynasty - Early Period was the Wuzhu coin, whose core characteristics include: the inscription 'Wuzhu' had shallow strokes, the casting process used crude clay molds, and no special markings." The model also generated from the "Tang Dynasty - High Tang Dynasty" template: "The typical currency of the Tang Dynasty - High Tang Dynasty was the Kaiyuan Tongbao coin, whose core characteristics include: the inscription 'Kaiyuan Tongbao' was a variant of Ouyang Xun's script, with rounded and full strokes, the casting process used a fine sand casting method, and the special marking often included a moon pattern on the back." These textual fragments contain not only chronological identifiers (such as "Western Han Dynasty" and "Tang Dynasty"), but also auxiliary target values (such as "Early Period" and "High Tang Dynasty"), and these auxiliary target values are strictly correlated with the detailed features in the chronological feature map (coin inscription style and craftsmanship differences).After generating a text snippet, the server determines the combination of "era identifier + auxiliary target value" as the "era identifier" for the text training instance. For example, the aforementioned "Western Han Dynasty - Early Wuzhu Coin..." instance has the era identifier "Western Han Dynasty - Early"; the "Tang Dynasty - High Tang Kaiyuan Tongbao Coin..." instance has the era identifier "Tang Dynasty - High Tang". To ensure the accuracy of the auxiliary target value, the server implements a dual verification mechanism: first, rule verification, checking whether the auxiliary target value conforms to historical facts (e.g., "Early Western Han Dynasty" corresponds to the clay mold rough casting process); second, manual review, in which ancient coin experts randomly check the generated snippets to confirm the match between the "auxiliary target value - feature description" (e.g., whether the inscription style of the "High Tang Kaiyuan Tongbao Coin" coin conforms to the evolution of Ouyang Xun's calligraphy). This approach refines the dating of text training instances from a single "dynasty" to a "dynasty + detailed attributes," such as "Western Han Dynasty - Early" and "Tang Dynasty - High Tang Dynasty." This enables subsequent ancient coin identification models to learn more fine-grained feature associations (e.g., "shallow inscription + coarse clay mold - Early Western Han Dynasty" and "round inscription + fine sand casting - High Tang Dynasty"), significantly improving the model's accuracy in dating physical imagery. In summary, by expanding the generation paradigm and driving the NLP model to generate text fragments with auxiliary target values, the server uses the "age identifier + auxiliary target value" as the dating identifier, enriching the annotation dimension of text training instances, providing the model with more accurate supervision signals, and ultimately improving the fine-grained classification capabilities of ancient coin identification.
[0126] In an embodiment of the present invention, the first ancient coin text generation paradigm is also used to drive the NLP model fine-tuned based on the ancient coin field to stop generating text training instances related to the age feature map when there is ambiguity between the at least two age identification items. The embodiment of the present invention also provides the following implementation method.
[0127] The at least two era identification items are obtained by screening from a plurality of era identification items.
[0128] In an embodiment of the present invention, illustratively, before generating a text training instance, the server must first ensure that the selected "era identification item" is unambiguous. For example, the server extracts "Wu Zhu Qian" from a candidate list (such as "Wu Zhu Qian", "Kaiyuan Tong Bao", and "Chongning Tong Bao") as the era identification item to be verified. At this time, the server calls the built-in "ancient coin era association knowledge base" for verification: the knowledge base shows that "Wu Zhu Qian" was first minted by Emperor Wu of the Western Han Dynasty, and continued to be minted in the Eastern Han Dynasty, Shu Han Dynasty, and Sui Dynasty. The characteristics of Wu Zhu Qian in different periods (such as the depth of the coin text and the size of the coin diameter) overlap but there is no absolute distinguishing mark. Therefore, "Wu Zhu Qian" as an era identification item will lead to ambiguity and cannot clearly correspond to a single or strongly associated era. When the server determines through the knowledge base that a certain era identification item is ambiguous, the first ancient coin text generation paradigm will trigger the "stop generation" mechanism. For example, the server attempts to insert "Wu Zhu Qian" into the "era placeholder" of the generation model (e.g., the template reads: "The typical currency of the [era placeholder] period is the Wu Zhu Qian, with characteristics including..."). Before generating the data, the NLP model invokes the ambiguity detection module and discovers that "Wu Zhu Qian" corresponds to multiple eras (Western Han Dynasty, Eastern Han Dynasty, and Sui Dynasty) without sufficient distinguishing features. Therefore, the model refuses to generate the relevant text fragment and returns a message to the server stating, "Age identifier ambiguous, generation aborted." The server must filter unambiguous age identifiers from the original candidate list. The filtering rules are based on the following criteria: 1. Age uniqueness: The name or type of the ancient coin is strongly associated with only one major historical period. For example, although the "Kai Yuan Tong Bao" coin lasted for nearly three hundred years during the Tang Dynasty, it is generally considered by archaeologists to be a typical Tang Dynasty currency (no large-scale minting occurred during other dynasties). Furthermore, the characteristic differences between Kai Yuan Tong Bao coins from the early, mid-Tang, and late Tang dynasties have been refined using auxiliary target values (e.g., "Early Tang - narrow rim" and "Mid-Tang - wide rim"). Therefore, "Kai Yuan Tong Bao" is unambiguous as a age identifier. 2. Feature Distinction: The core features of the coin (inscription, craftsmanship, markings) are significantly different from similar coins of other eras. For example, the "Chongning Tong Bao" coin was minted during the reign of Emperor Huizong of the Northern Song Dynasty. The inscription is in the "Slender Gold" script (a style not found in other dynasties). The coin diameter (approximately 3.5 cm) and the casting process (the iron mother coin is fine) are significantly different from Tongbao coins of other dynasties. Therefore, "Chongning Tong Bao" is unambiguous. 3. Documentary Clarity: Authoritative historical documents or archaeological reports clearly record that the coin is a "typical representative" of a specific period. For example, the "Hongwu Tong Bao" coin was minted during the reign of Emperor Taizu Zhu Yuanzhang of the Ming Dynasty. "Ming History: Food and Goods Records" clearly records it as the standard currency of the early Ming Dynasty. There is no possibility of confusion with Tongbao coins of other dynasties. Therefore, "Hongwu Tong Bao" is unambiguous. The server uses the above rules to filter the candidate list. For example, the original candidate list was "Wuzhuqian", "Kaiyuan Tongbao", "Chongning Tongbao" and "Hongwu Tongbao". After screening: "Wuzhuqian" was eliminated because it spanned multiple dynasties and had overlapping features; "Kaiyuan Tongbao", "Chongning Tongbao" and "Hongwu Tongbao" were retained as "at least two unambiguous date identifiers" because they met the requirements of uniqueness, distinctiveness and documentary clarity.The server will fill the selected unambiguous date identifiers (such as "Kaiyuan Tong Bao" and "Chongning Tong Bao") into the "date placeholder" of the first ancient coin description generation paradigm to generate a guiding template (such as "The typical currency of the Kaiyuan Tong Bao period is Kaiyuan Tong Bao, and its core features include: the coin text is in Ouyang Xun's style, the casting process is the sand casting method, and the back moon pattern is common"; "The typical currency of the Chongning Tong Bao period is Chongning Tong Bao, and its core features include: the coin text is in Slender Gold script, the coin diameter is about 3.5 cm, and the casting process is iron mother coin precision casting"). Subsequently, the NLP model fine-tuned in the ancient coin field is called to generate text fragments based on the template (such as "Kaiyuan Tong Bao, dignified coin text, moon pattern on the back, cast by sand casting method"; "Chongning Tong Bao, clear Slender Gold script coin text, coin diameter is 3.5 cm, iron mother coin casting method"), and use "Kaiyuan Tong Bao" and "Chongning Tong Bao" as the date identifiers of the corresponding instances. In summary, the server ensures that the age identification items of text training instances are unambiguous through ambiguity detection, generation termination and strict screening, avoids model training errors caused by age confusion, and provides a reliable data foundation for the accuracy of subsequent ancient coin identification models.
[0129] In the embodiment of the present invention, the characterization of text features is performed by using an ancient coin identification model to be trained to obtain an estimated age, which can be implemented through the following examples.
[0130] According to the text feature representation, a first adapted feature representation is obtained by projecting it through a cross-modal adaptation network, wherein the cross-modal adaptation network is used to project the feature representation from the image feature alignment domain to the model feature alignment domain, wherein the model feature alignment domain is a feature alignment domain that can be identified by the ancient coin identification model;
[0131] According to the first adapted feature representation, the ancient coin identification model to be trained is used for identification to obtain an estimated age;
[0132] The method further comprises:
[0133] Acquire first inferred data of the physical object image;
[0134] performing a feature extraction operation on the first inferred data by the first feature mapper to obtain an image feature representation;
[0135] Projecting the image feature representation through the cross-modal adaptation network to obtain a second adapted feature representation;
[0136] According to the second adaptive feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
[0137] In an embodiment of the present invention, exemplarily, after the server obtains the "text feature representation" of the text training instance (for example, for the text training instance of "Tang Kaiyuan Tongbao, the coin text is in Ouyang Xun's style, with a moon pattern on the back", a 256-dimensional vector V_text is obtained after processing by the text feature mapper, which is in the image feature alignment domain), it needs to be projected to the "model feature alignment domain" that can be processed by the ancient coin identification model through the "cross-modal adaptation network". The cross-modal adaptation network is a projection network consisting of two fully connected layers: the first layer maps the 256-dimensional V_text to 512 dimensions (e.g., W1×V_text+b1, where W1 is a 256×512 weight matrix and b1 is a bias vector). The second layer enhances the nonlinear representation through the ReLU activation function (e.g., ReLU(W2×(first layer output)+b2), where W2 is a 512×512 weight matrix and b2 is a bias vector). The final output is a 512-dimensional "first adaptation feature representation" V_ada_text, which is in the model feature alignment domain (the input requirement of the ancient coin identification model is 512 dimensions). The server then inputs V_ada_text into the ancient coin identification model to be trained (a Transformer-based classification model). The model's encoder layer performs contextual modeling on V_ada_text, capturing the association between features like "Ouyang Xun's calligraphy" and "back moon pattern" and the Tang Dynasty (for example, the model has learned that "Ouyang Xun's calligraphy" is a typical feature of Tang Dynasty Kaiyuan Tongbao coins, and that "back moon pattern" was common during the heyday of the Tang Dynasty). Finally, the classification head (fully connected layer with softmax activation) outputs a probability distribution for each era (e.g., 0.93 for "Tang," 0.04 for "Sui," and 0.03 for "Five Dynasties"). The highest probability, "Tang," is the "estimated era" for the training text instance. When a user uploads a physical image of an ancient coin (e.g., a high-definition photo of a Tang Kaiyuan Tongbao coin, including the font details of the inscription "Kaiyuan Tongbao" and the features of the back moon pattern) as the "first estimated data," the server first extracts features using the image feature mapper, generating a 256-dimensional "image feature representation" V_image (in the image feature alignment domain). The server uses the same cross-modal adaptation network to project V_image: The first layer maps the 256-dimensional V_image to a 512-dimensional representation (W1 × V_image + b1). The second layer, after ReLU activation, outputs a 512-dimensional "second-adapted feature representation," V_ada_image, which is also within the model's feature alignment domain. V_ada_image is then fed into the trained ancient coin authentication model. The model's encoder layer analyzes features in V_ada_image (e.g., the compatibility of the coin's inscription's font with Ouyang Xun's, or the position of the moon pattern with the typical distribution of Kaiyuan Tongbao coins from the heyday of the Tang Dynasty). The classification head then outputs a probability distribution for each era (e.g., 0.95 for "Tang," 0.02 for "Sui," and 0.03 for "Song"). The highest probability, "Tang," is designated as the "estimated era" of the image.The core function of the cross-modal adaptation network is to solve the spatial mismatch problem between the "image feature alignment domain" and the "model feature alignment domain". For example, the features output by the text feature mapper and the image feature mapper are 256 dimensions (to adapt to the cross-modal alignment requirements), while the input requirement of the ancient coin identification model is 512 dimensions (to adapt to the complex feature modeling requirements of the deep model). Through the projection of the cross-modal adaptation network, the text and image features are uniformly converted into dimensions that the model can process, ensuring that the model can use text and image data for training and reasoning at the same time. In summary, the server projects the text and image features from the alignment domain to a space that the model can process through the cross-modal adaptation network. Combined with the context modeling capabilities of the ancient coin identification model, it realizes the unified age identification of text training instances and physical image data, effectively improving the model's generalization ability for heterogeneous data.
[0138] In the embodiments of the present invention, the following implementation modes are also provided.
[0139] Obtaining an auxiliary training instance of the text class, wherein the auxiliary training instance carries an auxiliary age annotation;
[0140] According to the auxiliary training instance, performing a feature extraction operation on the auxiliary training instance by the first feature mapper to obtain a cross-modal intermediate feature representation;
[0141] Projecting the cross-modal intermediate feature representation through a cross-modal adaptation network to be trained to obtain an inferred adaptation feature representation;
[0142] According to the inferred adaptation feature representation, identification is performed using the ancient coin identification model to be trained to obtain second inferred data;
[0143] According to the deviation between the second inferred data and the auxiliary age annotation, the parameter configuration of the ancient coin identification model to be trained is fixed, and the parameter configuration of the cross-modal adaptation network to be trained is updated to obtain the cross-modal adaptation network.
[0144] In an embodiment of the present invention, illustratively, the server first obtains "auxiliary training instances" from an ancient coin research database (such as an archaeological report text library, a museum collection description system). These instances are text descriptions of ancient coins, and have been annotated by experts with "auxiliary age annotations" (i.e., clear age labels). For example, the server obtains an auxiliary training instance: "Northern Song Chongning Tongbao, the coin text is in thin gold font, iron stroke and silver hook, the coin diameter is about 3.5 cm", and its auxiliary age annotation is "Northern Song Dynasty". The number of such instances is usually thousands, covering different eras (such as "Tang" and "Ming", etc.), which are used for special training of cross-modal adaptation networks. The server calls the trained "first feature mapper" (text feature mapper) to extract features from the auxiliary training instances. The text feature mapper is a BERT-based text encoder that can map text features to the "image feature alignment domain". Taking the example of "Northern Song Chongning Tongbao coin...", the server first segments the text into words (e.g., "Northern Song Dynasty," "Chongning Tongbao," "Slender Gold Script," "3.5 centimeters"), generating a sequence of word vectors. BERT's 12-layer Transformer encoder then contextualizes the word vectors, extracting semantic associations between key features such as "Slender Gold Script Coin Inscription" and "Coin Diameter." Finally, a pooling layer outputs a 256-dimensional "cross-modal intermediate feature representation" vector V_mid (in the image feature alignment domain). The server then inputs V_mid into the "cross-modal adaptation network to be trained" (with initial random parameters). The network consists of two fully connected layers: the first maps the 256-dimensional V_mid to 512 dimensions (e.g., W1×V_mid+b1, where W1 is a 256×512 weight matrix and b1 is a bias vector). The second layer enhances the nonlinear representation using the ReLU activation function (e.g., ReLU(W2×(first layer output)+b2), where W2 is a 512×512 weight matrix and b2 is a bias vector). The final output is a 512-dimensional "inferred adaptive feature representation" V_ada (which is in the model feature alignment domain and meets the input requirements of the ancient coin identification model). The server inputs V_ada into the "ancient coin identification model to be trained" (a Transformer-based classification model). The model's encoder layer performs contextual modeling on V_ada, capturing the association between "Slender Gold" and "Northern Song" and the matching degree between "3.5 cm coin diameter" and "Chongning Tongbao." The classification head (fully connected layer + Softmax) outputs a probability distribution for each era (e.g., 0.6 for "Northern Song," 0.2 for "Southern Song," and 0.2 for "Tang"). The "Northern Song" with the highest probability is the "second guess." The server calculates the deviation between the second guess and the auxiliary era annotation ("Northern Song") using the cross-entropy loss function: ,in, is a one-hot encoding for auxiliary era labels ("Northern Song" corresponds to c=1, others are 0), is the probability distribution of the model output. To ensure that the cross-modal adaptation network learns to correctly project features from the image feature alignment domain into a space the model can process, the server fixes the parameters of the ancient coin authentication model (does not update them), and only calculates the gradient of the loss with respect to the cross-modal adaptation network parameters (W1, b1, W2, and b2) through backpropagation. These parameters are then updated using the Adam optimizer. For example, if the cross-modal adaptation network's projection performance is poor during initial training (for example, V_ada fails to effectively reflect the characteristics of the "Slender Gold" script), resulting in the model outputting a probability of only 0.3 for "Northern Song Dynasty" and a large loss, the server adjusts the weights of W1 and W2 to make V_ada more prominent in expressing key features such as the "Slender Gold" script and the length of 3.5 cm. After multiple iterations (e.g., 1000 steps), when the loss drops below 0.1, the cross-modal adaptation network's parameters converge and are able to accurately project features from the image feature alignment domain into the model feature alignment domain, completing training. In summary, the server pre-trains a cross-modal adaptation network through supervised learning of auxiliary training instances, solving the spatial mismatch problem between the image feature alignment domain and the model feature alignment domain, laying a key foundation for the subsequent joint training of ancient coin identification models and physical image identification.
[0145] In the embodiments of the present invention, the following implementation modes are also provided.
[0146] Obtaining a second ancient coin description generation paradigm and a cross-modal adaptation command, wherein the second ancient coin description generation paradigm includes a feature representation placeholder area and a cross-modal adaptation command placeholder area, and the cross-modal adaptation command is used to drive the ancient coin identification model to be trained to generate the second inference data;
[0147] Adding the inferred adaptation feature representation to the feature representation placeholder in the second ancient coin text generation paradigm, and adding the cross-modal adaptation command to the cross-modal adaptation command placeholder in the second ancient coin text generation paradigm, to obtain a second guidance template;
[0148] The method of obtaining second inferred data by performing identification based on the inferred adaptation feature representation using the ancient coin identification model to be trained includes:
[0149] According to the second guiding template, identification is performed using the ancient coin identification model to be trained to obtain the second inferred data.
[0150] In an embodiment of the present invention, exemplarily, the server first obtains the "second ancient coin text generation paradigm" from a predefined template library. This is a structured text generation template used to guide the ancient coin identification model to perform reasoning based on feature representation. The paradigm contains two key placeholders: the "feature representation placeholder" (used to fill in the feature information after cross-modal adaptation) and the "cross-modal adaptation command placeholder" (used to clarify the reasoning task of the model). For example, the structure of the second ancient coin text generation paradigm is: "[cross-modal adaptation command placeholder]: [feature representation placeholder]". At the same time, the server obtains the "cross-modal adaptation command", which is an instruction that clarifies the model task, such as "infer the age of the ancient coin based on the following features:". The function of this command is to drive the ancient coin identification model to be trained to focus on the analysis of feature representation and generate corresponding age inference results. After obtaining the "inferred adaptation feature representation" (for example, the 512-dimensional vector V_ada obtained after projecting the auxiliary training example "Northern Song Chongning Tongbao, coin inscription is in Slender Gold Script..." through the cross-modal adaptation network), the server needs to combine it with the cross-modal adaptation command to construct the "second guidance template." The specific operation is as follows: the server converts V_ada into a structured description that can be understood by the model (for example, extracting a textual representation of key features: "Coin inscription features: Slender Gold Script; Coin diameter features: 3.5 cm; Craftsmanship features: Iron mother coin, precision casting"). This textual description is then entered into the "feature representation placeholder" of the second ancient coin description generation paradigm. The cross-modal adaptation command "Infer the age of the ancient coin based on the following features:" is then entered into the "cross-modal adaptation command placeholder." The resulting second guidance template is: "Infer the age of the ancient coin based on the following features: Coin inscription features: Slender Gold Script; Coin diameter features: 3.5 cm; Craftsmanship features: Iron mother coin, precision casting." The server then inputs this second guidance template into the ancient coin authentication model to be trained (a Transformer-based generative model with textual reasoning capabilities). The model generates "second inferred data" through the following steps: 1. Input Parsing: The model first parses the cross-modal adaptation command ("infer era") and feature representations ("Slender Gold Script," "3.5 cm," "Iron Mother Coin Precision") in the template, clarifying the task objective of "determining the era of an ancient coin based on given features." 2. Feature Association: The model draws on its internal knowledge base (learned during training, such as "Slender Gold Script Coin Inscriptions are a typical feature of the Northern Song Dynasty Huizong period," "Iron Mother Coin Precision Casting Techniques Commonly Used on Official Northern Song Dynasty Currency," and "The Diameter of Chongning Tongbao Coins is Typically 3.3-3.7 cm") to match input features with era labels. 3. Result Generation: Based on the strength of the correlation between the features and the era, the model outputs the most likely era label (e.g., "Northern Song Dynasty") as the "second inferred data." The server compares the second inferred data ("Northern Song Dynasty") with the auxiliary era label ("Northern Song Dynasty"). If the two agree, the loss is low, and the cross-modal adaptation network parameters are adjusted sparingly. If they disagree (e.g., the model mistakenly identifies "Southern Song Dynasty"), the loss is high, triggering a parameter update.For example, if during initial training the model mistakenly identifies a text as "Southern Song" because V_ada fails to adequately express the characteristics of the "Slender Gold" script, the server calculates the deviation using cross-entropy loss and uses backpropagation to adjust the weight matrix (W1, W2) and bias vector (b1, b2) of the cross-modal adaptation network, allowing V_ada to more prominently express key features like the "Slender Gold" script, thereby improving the model's inference accuracy. In summary, the server uses the second ancient coin description generation paradigm to transform abstract feature representations into textual instructions that the model can understand. Combined with cross-modal adaptation commands, this guides the model to focus on the inference task, effectively improving the training efficiency of the cross-modal adaptation network and the inference accuracy of the ancient coin authentication model.
[0151] In an embodiment of the present invention, updating the parameter configuration of the ancient coin identification model to be trained based on the deviation between the estimated age and the age identifier to obtain the ancient coin identification model includes:
[0152] According to the deviation between the inferred age and the age identifier, the parameter configuration of the ancient coin identification model to be trained is updated to obtain the ancient coin identification model, and according to the deviation between the inferred age and the age identifier, the parameter configuration of the cross-modal adaptation network is updated to obtain the cross-modal adaptation network that has been iteratively tuned.
[0153] In an embodiment of the present invention, for example, the server selects a text training example: "Tang Kaiyuan Tongbao, the coin text is in Ouyang Xun's style, with elegant strokes and a moon pattern on the back." Its era is identified as "Tang." The server first extracts features from this example using a text feature mapper, obtaining a 256-dimensional "text feature representation" vector V_text (in the image feature alignment domain). V_text is then input into a cross-modal adaptation network, which, after projection through two fully connected layers, outputs a 512-dimensional "first adaptation feature representation" V_ada (in the model feature alignment domain). The server inputs V_ada into the trained ancient coin authentication model (a Transformer-based classification model). The model's encoder layer performs contextual modeling on V_ada, capturing the associations between features such as "Ouyang Xun's style" and "moon pattern on the back." The classification head outputs a probability distribution for each era (e.g., in the initial training phase, it might be 0.4 for "Tang," 0.5 for "Sui," and 0.1 for "Five Dynasties"). The highest probability, "Sui," is designated as the "estimated era." The server calculates the deviation between the estimated date ("Sui") and the date marker ("Tang") using a cross-entropy loss function. The server calculates the gradient of the loss with respect to the parameters of the ancient coin authentication model and the cross-modal adaptation network through backpropagation, and simultaneously updates the parameters of both: 1. Parameter Update of the Ancient Coin Authentication Model: The loss generates gradients for the model's Transformer encoder layer weights (such as the Q, K, and V matrices of the attention head) and the classification head weights (W and b in the fully connected layers). For example, if the model misclassifies because it has not fully learned the association between "Ouyang Xun's calligraphy style" and "Tang," the gradient will adjust the encoder layer's attention weights to focus more on the "Ouyang Xun's calligraphy style" feature; it will also adjust the classification head weights to increase the output probability of the "Tang" category. 2. Parameter Update of the Cross-Modal Adaptation Network: The loss is backpropagated through the model to the cross-modal adaptation network, generating gradients for the weights (W1, W2) and biases (b1, b2) of its two fully connected layers. For example, if V_ada fails to effectively represent the feature "Ouyang Xun's calligraphy" (e.g., if the dimension of this feature in V_ada is low), the gradient will adjust W1 and W2 to increase the dimension corresponding to "Ouyang Xun's calligraphy" in V_ada, thereby improving the model's perception of this feature. The server repeats this process (e.g., processing tens of thousands of text training instances), and each iteration simultaneously updates the parameters of both models using the loss. For example, at the 100th iteration, processing another text training instance, "Tang Kaiyuan Tongbao, the coin text is a variant of Ouyang Xun's, with a double moon pattern on the back" (marked with the "Tang" era), the ancient coin authentication model's probability of outputting "Tang" increased to 0.7, and the loss dropped to 0.36. By the 1000th iteration, the model's probability of outputting "Tang" reached 0.95, and the loss dropped to 0.05, indicating that the model has accurately captured the strong correlation between "Ouyang Xun's calligraphy," "moon pattern on the back," and "Tang."Finally, when the loss stabilized below 0.1 and no longer decreased significantly, the server completed iterative tuning: the parameters of the ancient coin identification model (Transformer encoder, classification head) had learned the key feature associations of each era; the parameters of the cross-modal adaptation network (W1, W2, b1, b2) were able to accurately project the features of the image feature alignment domain (256-dimensional vectors of text or images) to the model feature alignment domain (512-dimensional vectors), ensuring that the model can efficiently use these features for reasoning. In summary, by jointly updating the ancient coin identification model and the cross-modal adaptation network, the server solved the problem of collaborative optimization of feature space adaptation and model reasoning capabilities, ultimately resulting in a complete system capable of accurately identifying the age of ancient coins.
[0154] In an embodiment of the present invention, the auxiliary training instance includes multiple ancient coin text fragments, and the auxiliary age annotation is the full ancient coin text of the multi-dimensional ancient coin archive corresponding to the auxiliary training instance or the multiple ancient coin text fragments.
[0155] In an embodiment of the present invention, exemplarily, the server obtains auxiliary training examples from the archaeological database, such as a plurality of ancient coin text fragments including "Tang Kaiyuan Tongbao with moon pattern on the back" and "Tang Kaiyuan Tongbao with narrow edge", and the auxiliary age annotation is not a single label, but the full text of the corresponding multi-dimensional ancient coin archive. For example, the auxiliary age annotation of the "Tang Kaiyuan Tongbao with moon pattern on the back" fragment is the complete document of the ancient coin: "officially cast during the Kaiyuan period of the Tang Dynasty, unearthed from the Hejiacun cellar in Xi'an, and the "New Book of Tang·Food and Goods Records" records that "Kaiyuan Tongbao was cast in the fourth year of Wude, with a diameter of eight fen and a weight of two zhu and four si". The server associates each text fragment with the full archive to ensure that the auxiliary annotation covers multi-dimensional information such as documents of the age and unearthed records, providing comprehensive supervision signals for cross-modal adaptive network training.
[0156] In the embodiments of the present invention, the following implementation modes are also provided.
[0157] Obtaining the core dating basis of the text training instance;
[0158] Performing a feature extraction operation on the core generation basis to obtain a core feature representation;
[0159] The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes:
[0160] According to the text feature representation and the core feature representation, the ancient coin identification model to be trained is used for identification to obtain the estimated age.
[0161] In an embodiment of the present invention, the server first extracts "core dating criteria" from text training examples, i.e., the most critical features that distinguish ancient coins from different eras. For example, the server obtains a text training example: "Northern Song Chongning Tongbao, coin text in Slender Gold script, iron-stroked and silver-hooked, coin diameter approximately 3.5 cm, no inner rim on the back," which is dated "Northern Song." The server then uses a built-in "Ancient Coin Dating Rule Library" (which records the core features of ancient coins from different eras) to identify the core dating criteria: the rule library indicates that "Slender Gold script coin text" is a typical feature of the Northern Song Huizong period (other dynasties do not have this calligraphy style); "coin diameter of 3.5 cm" is the standard size of Chongning Tongbao (significantly different from Tongbao coins of other eras); and "no inner rim on the back" is a common feature of official Northern Song Chongning Tongbao (unlike privately minted or other dynasties). Therefore, the core dating criteria for this example are extracted as: "Slender Gold script coin text," "coin diameter of 3.5 cm," and "no inner rim on the back." The server invokes the "core feature extraction module" (a tool based on NLP-based entity recognition and value extraction) to extract features from the core dating criteria. The following steps are performed: Semantic feature extraction is performed on the "Slender Gold Coin Inscription": Using a pre-trained word embedding model for ancient coins (e.g., Word2Vec), "Slender Gold Coin Inscription" is mapped into a 128-dimensional semantic vector (e.g., V_style); Numerical feature extraction is performed on the "Coin Diameter 3.5 cm": "3.5 cm" is converted to a normalized numeric value (e.g., 3.5) and expanded into a 128-dimensional numeric feature vector (e.g., V_size); Boolean feature extraction is performed on the "No Inner Rib on the Back": "No Inner Rib" is converted to a binary flag ("absent" corresponds to 1, "present" corresponds to 0) and expanded into a 128-dimensional Boolean feature vector (e.g., V_rim). Finally, the server concatenates these three feature vectors into a 384-dimensional "core feature representation," V_core ([V_style; V_size; V_rim]), which focuses on the key dating features of the ancient coin. The server extracts features from the original text training instance using a text feature mapper, generating a 256-dimensional "text feature representation" V_text (in the image feature alignment domain). Subsequently, V_text is projected into a 512-dimensional "first adaptation feature representation" V_ada_text (in the model feature alignment domain) via a cross-modal adaptation network. The server concatenates V_ada_text and V_core into an 896-dimensional fused feature vector V_fused ([V_ada_text;V_core]), which is then fed into the ancient coin authentication model to be trained (a Transformer-based classification model).The model's encoder layer uses a multi-head attention mechanism to focus on the features of "Slender Gold Script," "3.5 cm," and "no inner rim" in V_core (for example, the dimension weighted for "Slender Gold Script" is 0.8). It also incorporates the contextual semantics in V_ada_text (such as the association between "Chongning Tong Bao" and "Northern Song Dynasty"). The classification head outputs a probability distribution for each era (e.g., 0.96 for "Northern Song Dynasty," 0.03 for "Southern Song Dynasty," and 0.01 for "Tang Dynasty"). The "Northern Song Dynasty" with the highest probability is designated as the "estimated era." The server compares the estimated era ("Northern Song Dynasty") with the era identifier ("Northern Song Dynasty"). If they match, the server receives a smaller loss (e.g., a cross-entropy loss of 0.04) and fine-tunes the model parameters. If they don't match (e.g., the model mistakenly classified it as "Southern Song Dynasty" during initial training), the server receives a larger loss (e.g., 0.8), triggering a parameter update. The model's attention weights are adjusted to prioritize core features like "Slender Gold Script." The server also optimizes the projection parameters of the cross-modal adaptation network, making V_ada_text more prominent in the era association of "Chongning Tong Bao." In summary, the server extracts core dating evidence and integrates its feature representations, guiding the ancient coin identification model to focus on key information, significantly improving the accuracy of dating speculation and ensuring that the model can make decisions based on the most discriminative features.
[0162] In the embodiments of the present invention, the following implementation modes are also provided.
[0163] Obtaining a third ancient coin text generation paradigm and a guiding command, wherein the third ancient coin text generation paradigm includes a feature representation placeholder area and a guiding command placeholder area, and the guiding command is used to drive the ancient coin identification model to be trained to generate the estimated age;
[0164] Adding the text feature representation to the feature representation placeholder area of the third ancient coin text generation paradigm, and adding the guide command to the guide command placeholder area of the third ancient coin text generation paradigm to obtain a third guide template;
[0165] The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes:
[0166] According to the third guiding template, the ancient coin identification model to be trained is used for identification to obtain the estimated age.
[0167] In an embodiment of the present invention, exemplarily, the server obtains the "third ancient coin text generation paradigm" from a predefined template library. This is a structured text generation template used to guide the ancient coin identification model to perform age reasoning based on feature representation. The paradigm contains two key placeholders: the "feature representation placeholder" (used to fill in the textual description of the text features) and the "guidance command placeholder" (used to clarify the reasoning task of the model). For example, the structure of the third ancient coin text generation paradigm is: "[guidance command placeholder]: [feature representation placeholder]". At the same time, the server obtains the "guidance command", which is an instruction that clarifies the model task, such as "infer the age of the ancient coin based on the following features:". The function of this command is to drive the ancient coin identification model to be trained to focus on feature analysis and generate corresponding age inference results. After obtaining the text feature representation of a training example (for example, for the example "Tang Kaiyuan Tongbao, inscription in Ouyang Xun's script, with moon pattern on the back"), the server processes the text feature mapper to produce a 256-dimensional vector V_text, whose textual feature description is "Inscription: Ouyang Xun's script; Back pattern: Moon pattern"), it then combines this vector with the guidance command to construct a "third guidance template." The server performs the following operations: The server populates V_text's textual feature description ("Inscription: Ouyang Xun's script; Back pattern: Moon pattern") into the "feature representation placeholder" of the third ancient coin description generation paradigm; and the guidance command ("Based on the following ancient coin features, infer its age:") into the "guidance command placeholder." The resulting third guidance template is: "Based on the following ancient coin features, infer its age: Inscription: Ouyang Xun's script; Back pattern: Moon pattern." The server then inputs this third guidance template into the ancient coin authentication model to be trained (a Transformer-based generative model that already possesses reasoning capabilities in the ancient coin domain). The model generates the "estimated age" through the following steps: 1. Instruction parsing: The model first identifies the guiding command ("infer age") in the template and clarifies the task objective as "determining the age of ancient coins based on given features"; 2. Feature extraction: The model parses the content of the "feature representation placeholder area" ("coin text: Ouyang Xun's script; back pattern: moon pattern") and extracts the key features "Ouyang Xun's script" and "moon pattern"; 3. Knowledge association: The model calls the internal knowledge base (knowledge learned during the training phase, such as "Ouyang Xun's script is a typical feature of Kaiyuan Tongbao coins in the Tang Dynasty" and "back moon pattern is common in Kaiyuan Tongbao coins in the heyday of the Tang Dynasty") and associates the extracted features with the age label; 4. Result generation: Based on the strength of the association between the feature and the age, the model outputs the most likely age label (such as "Tang") as the "estimated age".The server compares the inferred era ("Tang") with the era identifier ("Tang") of the training text. If they match, the loss is small (e.g., a cross-entropy loss of 0.05), and the model parameters are fine-tuned. If they don't match (e.g., during the initial training phase, the model misclassifies it as "Sui"), the loss is large (e.g., 0.8), triggering a parameter update: the model's attention weights are adjusted to prioritize key features like "Ouyang Xun's calligraphy." Simultaneously, the projection parameters of the cross-modal adaptation network are optimized to ensure that the text feature representation more accurately reflects the era association. For example, during the initial training phase, the model might misclassify "Qian Wen: Ouyang Xun's calligraphy; Back Pattern: Moon Pattern" as "Sui" because it hasn't fully learned the association between "Ouyang Xun's calligraphy" and "Tang." In this case, the server calculates the bias using a loss function and backpropagates the model's Transformer encoder weights (e.g., increasing the weight of the attention head corresponding to "Ouyang Xun's calligraphy") and the classification head weights (increasing the output probability of the "Tang" category), ultimately enabling the model to accurately associate features with eras. In summary, the server converts abstract feature representations into text instructions that can be understood by the model through the third ancient coin description generation paradigm, and combines it with guiding commands to clarify reasoning tasks, effectively improving the accuracy of the age estimation of the ancient coin identification model and ensuring that the model can make accurate decisions based on structured feature descriptions.
[0168] In an embodiment of the present invention, the third ancient coin text generation paradigm also includes a core dating basis placeholder area, and the embodiment of the present invention also provides the following implementation method.
[0169] Obtaining the core dating basis of the text training instance;
[0170] Performing a feature extraction operation on the core generation basis to obtain a core feature representation;
[0171] The step of adding the text feature representation to the feature representation placeholder area in the third ancient coin text generation paradigm, and adding the guide command to the guide command placeholder area in the third ancient coin text generation paradigm to obtain a third guide template includes:
[0172] Add the character feature representation to the third ancient coin character
[0173] The feature representation placeholder area in the generation paradigm is added, the guide command is added to the guide command placeholder area in the third ancient coin text generation paradigm, and the core feature representation is added to the core dating basis placeholder area in the third ancient coin text generation paradigm to obtain a third guide template.
[0174] In an embodiment of the present invention, illustratively, the server first obtains the expanded "Third Ancient Coin Description Generation Paradigm" from the template library. The paradigm contains three placeholders: "Feature Representation Placeholder" (used to fill in the textual description of text features), "Core Dating Basis Placeholder" (used to fill in the key features of the core dating basis), and "Guiding Command Placeholder" (used to clarify the model reasoning task). For example, the paradigm structure is: "[Guiding Command Placeholder]: Based on the feature [feature representation placeholder] and the core dating basis [core dating basis placeholder], what is the age of the ancient coin?" At the same time, the server obtains "guiding commands", such as "infer the age of the ancient coin based on the following information", which is used to drive the ancient coin identification model to focus on the reasoning task. Taking the training example text "Tang Kaiyuan Tongbao, inscription in Ouyang Xun's calligraphy, elegant strokes, moon pattern on the back, cast using the sand casting method" (the era is marked with "Tang") as an example, the server extracts core dating criteria from the "Ancient Coin Dating Rule Library": The rule library shows that "Ouyang Xun's calligraphy" is a hallmark of Tang Kaiyuan Tongbao coins (other dynasties do not have this calligraphy style); "moon pattern on the back" is a typical mark of Kaiyuan Tongbao coins from the prosperous Tang Dynasty (rarely seen in the early Tang Dynasty); and "sand casting method" was the mainstream casting method in the Tang Dynasty (after mold casting in the Han Dynasty and mother coin casting in the Song Dynasty). Therefore, the core dating criteria are extracted as: "Inscription: Ouyang Xun's calligraphy; Moon pattern on the back; Sand casting method." The server invokes the "core feature extraction module" to extract features from the core dating criteria: It performs semantic extraction on "Coin Inscription: Ouyang Xun's Calligraphy," generating a 128-dimensional semantic vector (e.g., V_style) using the ancient coin domain word embedding model. It also performs mark extraction on "Back Pattern: Moon Pattern," converting it to a binary flag (1 if "moon pattern" is present, 0 if it isn't) and expanding it to a 128-dimensional Boolean vector (e.g., V_mark). It also extracts the craft type from "Craftsmanship: Sand-casting Method," using a predefined craft coding table ("Sand-casting Method" corresponds to code 001) and expanding it to a 128-dimensional categorical vector (e.g., V_tech). Ultimately, the core dating criteria are converted into a 384-dimensional "core feature representation" V_core ([V_style; V_mark; V_tech]), which is further textualized as "Coin Inscription: Ouyang Xun's Calligraphy; Back Pattern: Moon Pattern; Craftsmanship: Sand-casting Method." The server textualizes the "text feature representation" of the text training instance (for example, after being processed by the text feature mapper, the feature description of the original instance is "Coin inscription: Ouyang Xun's calligraphy; Back pattern: Moon pattern"), combines it with the textual content of the core feature representation ("Coin inscription: Ouyang Xun's calligraphy; Back pattern: Moon pattern; Craftsmanship: Sand casting") and the guiding command ("Infer the age of the ancient coin based on the following information"), and fills it into the corresponding placeholder area of the third ancient coin description generation paradigm to obtain the third guiding template: "Infer the age of the ancient coin based on the following information: Based on the features [Coin inscription: Ouyang Xun's calligraphy; Back pattern: Moon pattern] and the core dating basis [Coin inscription: Ouyang Xun's calligraphy; Back pattern: Moon pattern; Craftsmanship: Sand casting], what is the age of the ancient coin?".The server inputs the third guidance template into the ancient coin authentication model to be trained (a Transformer-based generative model). The model performs inference through the following steps: 1. Instruction parsing: Identifying the task objective of "inferring the age" and clarifying the need to combine "features" with "core dating criteria" for judgment; 2. Feature association: Analyzing the core dating criteria of "Ouyang Xun's calligraphy" (exclusive to the Tang Dynasty), "sand-casting" (a mainstream Tang Dynasty technique), and "moon pattern" (a hallmark of the heyday of the Tang Dynasty), confirming their strong correlation with "Tang"; 3. Result generation: The model outputs "Tang" as the estimated age, consistent with the dating identifier in the text training example. If the model incorrectly identifies "Sui" as the estimated age due to insufficient learning of the "sand-casting"-Tang association during initial training, the server calculates the deviation between the estimated age and the identifier (e.g., a cross-entropy loss of 0.7) and adjusts the model parameters through backpropagation: increasing the attention weight for the "sand-casting" feature and optimizing the projection parameters of the cross-modal adaptation network so that the text feature representation better emphasizes the core dating criteria. After iterative training, the model ultimately accurately combines features with the core dating criteria to output the correct age. In summary, the server expands the third ancient coin description generation paradigm, integrates the core dating basis into the guidance template, guides the model to focus on key features, and significantly improves the accuracy and interpretability of age speculation.
[0175] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned AI-driven ancient coin identification method. Figure 2 As shown, Figure 2 This is a block diagram of the structure of a computer device 100 provided in an embodiment of the present invention. Computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or exchange, memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines.
[0176] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.
Claims
1. An AI-driven ancient coin identification method, characterized in that: include: Obtaining a text training instance of a text class, wherein the text training instance carries a year identifier, and the year identifier is used to represent the actual year of the text training instance; Performing a feature extraction operation on the text training instance by a first feature mapper to obtain a text feature representation, wherein the first feature mapper is used to match the text-type ancient coin information to an image feature alignment domain, wherein the image feature alignment domain is a domain where the feature representation of the physical image-type ancient coin information resides, and the information content of the text-type ancient coin information is lower than the information content of the physical image-type ancient coin information; Based on the text feature representation, the ancient coin identification model to be trained is used for identification to obtain an estimated age; According to the deviation between the estimated age and the age mark, the parameter configuration of the ancient coin identification model to be trained is updated to obtain an ancient coin identification model, which is used to identify the age of the ancient coin information of the physical image type based on the ancient coin identification model.
2. The method according to claim 1, characterized in that The method further comprises: Acquire first inferred data of the physical object image; Performing feature mapping on the first inferred data by a second feature mapper to obtain an image feature representation, wherein the second feature mapper is used to match the ancient coin information of the physical image class to the image feature alignment domain; According to the image feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
3. The method according to claim 1, characterized in that The first feature mapper includes a text feature mapper and an image feature mapper, wherein the text feature mapper is used to match the ancient coin information of the text type to the image feature alignment domain, and the image feature mapper is used to match the ancient coin information of the physical image type to the image feature alignment domain; The step of performing a feature extraction operation on the text training instance by a first feature mapper to obtain a text feature representation includes: Performing a feature extraction operation on the text training instance by the text feature mapper to obtain a text feature representation; The method further comprises: Acquire first inferred data of the physical object image; performing a feature extraction operation on the first inferred data by the image feature mapper to obtain an image feature representation; According to the image feature representation, the age of the first inferred data is obtained by performing identification using the ancient coin identification model; The method comprises: Acquire multiple training instance arrays, wherein the training instance arrays include a first target instance array of the text class and a second target instance array of the physical image class, wherein the first target instance array and the second target instance array of the same training instance array describe the same ancient coin; For a target training instance array in the plurality of training instance arrays, performing a feature extraction operation on a first target instance array in the target training instance array by a text feature mapper to be trained to obtain a first undetermined feature representation; Performing a feature extraction operation on a second target instance array in the target training instance array by the image feature mapper to be trained to obtain a second undetermined feature representation; Using the plurality of training instance arrays as the target training instance arrays, respectively, to obtain a plurality of first undetermined feature representations and a plurality of second undetermined feature representations; Based on the optimization goal of improving the feature aggregation of the target undetermined feature representation array and improving the feature discreteness of other undetermined feature representation pairs, the parameter configuration of the text feature mapper to be trained and the parameter configuration of the image feature mapper to be trained are updated to obtain the text feature mapper and the image feature mapper. The first undetermined feature representation and the second undetermined feature representation included in the target undetermined feature representation array are obtained based on the same training instance array, and the first undetermined feature representation and the second undetermined feature representation included in the other undetermined feature representation pairs are not obtained based on the same training instance array.
4. The method according to claim 1, wherein The method further comprises: Obtaining at least two age identification items and a first ancient coin text generation paradigm including an age placeholder, wherein the age identification item is used to describe the age of the ancient coin information of the physical image type, and the first ancient coin text generation paradigm is used to drive an NLP model fine-tuned based on the ancient coin field to generate multiple ancient coin text fragments, wherein multiple ancient coin rubbing samples indicated by the multiple ancient coin text fragments constitute an age feature map associated with the age corresponding to the age placeholder; Adding the at least two age identification items to the age placeholder area in the first ancient coin description generation paradigm to obtain a first guiding template; Generating the text training instance using the NLP model fine-tuned based on the ancient coin field according to the first guidance template, wherein the text training instance is a fragment of ancient coin text related to the at least two chronological identifiers; The at least two era identification items are determined as era identifications of the text training instance.
5. The method according to claim 1, wherein The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes: According to the text feature representation, a first adapted feature representation is obtained by projecting it through a cross-modal adaptation network, wherein the cross-modal adaptation network is used to project the feature representation from the image feature alignment domain to the model feature alignment domain, wherein the model feature alignment domain is a feature alignment domain that can be identified by the ancient coin identification model; According to the first adapted feature representation, the ancient coin identification model to be trained is used for identification to obtain an estimated age; The method further comprises: Acquire first inferred data of the physical object image; performing a feature extraction operation on the first inferred data by the first feature mapper to obtain an image feature representation; Projecting the image feature representation through the cross-modal adaptation network to obtain a second adapted feature representation; According to the second adaptive feature representation, identification is performed using the ancient coin identification model to obtain the age of the first inferred data.
6. The method according to claim 5, characterized in that The method further comprises: Obtaining an auxiliary training instance of the text class, wherein the auxiliary training instance carries an auxiliary age annotation; According to the auxiliary training instance, performing a feature extraction operation on the auxiliary training instance by the first feature mapper to obtain a cross-modal intermediate feature representation; Projecting the cross-modal intermediate feature representation through a cross-modal adaptation network to be trained to obtain an inferred adaptation feature representation; According to the inferred adaptation feature representation, identification is performed using the ancient coin identification model to be trained to obtain second inferred data; According to the deviation between the second inferred data and the auxiliary age annotation, the parameter configuration of the ancient coin identification model to be trained is fixed, and the parameter configuration of the cross-modal adaptation network to be trained is updated to obtain the cross-modal adaptation network.
7. The method according to claim 6, characterized in that The method further comprises: Obtaining a second ancient coin description generation paradigm and a cross-modal adaptation command, wherein the second ancient coin description generation paradigm includes a feature representation placeholder area and a cross-modal adaptation command placeholder area, and the cross-modal adaptation command is used to drive the ancient coin identification model to be trained to generate the second inference data; Adding the inferred adaptation feature representation to the feature representation placeholder in the second ancient coin text generation paradigm, and adding the cross-modal adaptation command to the cross-modal adaptation command placeholder in the second ancient coin text generation paradigm, to obtain a second guidance template; The method of obtaining second inferred data by performing identification based on the inferred adaptation feature representation using the ancient coin identification model to be trained includes: According to the second guiding template, identification is performed using the ancient coin identification model to be trained to obtain the second inferred data.
8. The method according to claim 6, characterized in that The method of updating the parameter configuration of the ancient coin identification model to be trained based on the deviation between the estimated age and the age identifier to obtain the ancient coin identification model includes: According to the deviation between the inferred age and the age identifier, the parameter configuration of the ancient coin identification model to be trained is updated to obtain the ancient coin identification model, and according to the deviation between the inferred age and the age identifier, the parameter configuration of the cross-modal adaptation network is updated to obtain the cross-modal adaptation network that has been iteratively tuned.
9. The method according to claim 1, characterized in that The method further comprises: Obtaining the core dating basis of the text training instance; Performing a feature extraction operation on the core generation basis to obtain a core feature representation; The method of performing identification based on the text feature representation and using the ancient coin identification model to be trained to obtain an estimated age includes: According to the text feature representation and the core feature representation, the ancient coin identification model to be trained is used for identification to obtain the estimated age.
10. A server system, characterized in that: The method comprises a server, wherein the server is configured to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent evaluation method and system based on ancient coin image
CN111046883A
Model training method and device, open set classification method and device, equipment and medium
CN118196556A
Method for detecting coin attributes based on neural network model
CN118711033A
Cultural relic classification and identification system and method based on image recognition and natural language processing
CN119048854A
Archaeological cultural relic model training method and system
CN119443284A